Multi-modal interactive smart home central control screen system and method based on large model
Through a multi-modal interactive smart home central screen control system based on large models, integrating multiple interaction methods and learning algorithms, the problem of single interaction and intelligence of traditional central screen control systems is solved, and efficient and personalized management and control of smart home devices is realized.
Patent Information
- Application Number
- CN202510436893.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-11
AI Technical Summary
The traditional smart home central control system has a single interaction method, insufficient intelligence, and it is difficult to achieve whole-house intelligence, inconvenient user operation and insufficient equipment linkage.
A multimodal interactive smart home central control system based on large models is adopted, integrating touch, voice, gesture and visual recognition and combining deep learning and machine learning algorithms to generate personalized control strategies and uniformly manage smart home devices.
It realizes a natural and convenient multi-modal interactive experience, improves the intelligence level and user satisfaction of the smart home system, simplifies the operation process, and has good scalability and personalized service capabilities.
Smart Images

Figure CN120295154A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of smart home central control screens, and particularly to a multi-modal interaction smart home central control screen system and method based on a large model. Background Art
[0002] With the continuous development of technology, smart home has become an indispensable part of modern families. However, traditional smart home central control screen systems often have problems such as single interaction methods and insufficient intelligence. Users need to operate through specific remote controls or mobile phone apps, and the interaction experience is not natural and convenient enough. At the same time, there is a lack of effective linkage and coordination between smart home devices, making it difficult to achieve true whole-house intelligence. Considering the above situation, this application proposes a multi-modal interaction smart home central control screen system and method based on a large model. Summary of the Invention
[0003] Based on the technical problems existing in the background art, the present invention proposes a multi-modal interaction smart home central control screen system and method based on a large model.
[0004] The multi-modal interaction smart home central control screen system based on a large model proposed by the present invention includes an instruction receiving module, an instruction processing module, a strategy generation and optimization module, an instruction sending and execution module, a status feedback and monitoring module, and a learning and adaptation module;
[0005] The instruction receiving module is connected to the instruction processing module. The strategy generation and optimization module is connected to the instruction processing module, the instruction sending and execution module, and the learning and adaptation module. The status feedback and monitoring module is connected to the learning and adaptation module. Both the instruction sending and execution module and the status feedback and monitoring module are connected to external smart home devices.
[0006] Preferably, the instruction receiving module is responsible for interacting with users and receiving users' instruction inputs through touch, voice, gesture, and image. For touch, it uses touch screen technology to detect users' touch operations. For voice, it captures sound through a microphone and uses speech recognition technology to convert the sound into text. For gesture and image, it captures images through a camera and uses computer vision technology to recognize gesture and image content.
[0007] Preferably, the instruction processing module is used to process the received instructions, including speech recognition, image recognition, and gesture recognition, and convert them into control instructions that the system can understand and execute;
[0008] Speech recognition uses natural language processing technology to convert speech text into instructions that can be understood by the system. The natural language processing technology adopts advanced NLP model structures such as the sequence-to-sequence (Seq2Seq) model and Transformer, utilizes the deep learning capabilities of large models, combines attention mechanism technology, and performs text conversion, semantic understanding, and intent recognition processing on speech instructions. By training a large amount of language data, the large model can learn the grammar, vocabulary, and semantic rules of the language, so as to accurately understand the user's speech instructions;
[0009] Image recognition uses image recognition algorithms to extract features and classify images, and identify the user's intent. The image recognition algorithms adopt advanced image recognition model structures such as YOLO (You Only Look Once) and ResNet, utilize the deep learning technology of the convolutional neural network (CNN) of large models, combine data augmentation and transfer learning technology, and perform feature extraction, classification, and recognition on the images captured by the camera. By training a large amount of image data, the large model can learn the feature patterns in the images, so as to accurately identify the user's gesture and face image information.
[0010] Preferably, the policy generation and optimization module generates personalized control policies according to the processed instructions and user needs, and optimizes the policies through optimization algorithms. Among them, policy generation uses machine learning algorithms to generate control policies according to the user's historical behavior and preferences. The machine learning algorithms are used to build prediction models. Policy optimization uses optimization algorithms to adjust the parameters of the prediction models to improve the prediction performance and optimize the generated control policies to improve the overall performance of the system and user satisfaction;
[0011] The implementation steps of the machine learning algorithm are as follows:
[0012] (1). Data preparation: Collect and clean the processed instruction and user need data, and divide them into training sets and test sets;
[0013] (2). Feature engineering: Select or extract features that have a significant impact on the model performance;
[0014] (3). Model selection: Select appropriate machine learning algorithms according to the problem type and data characteristics;
[0015] (4). Model training: Use the training data to train the model and adjust the model parameters to minimize the loss function;
[0016] (5). Model evaluation: Use the test data to evaluate the performance of the model and select the best model;
[0017] (6). Model deployment: Deploy the trained model to the actual application for prediction or classification;
[0018] The implementation steps of the optimization algorithm are as follows:
[0019] (1) Initialize parameters: Set the initial values of the model parameters;
[0020] (2) Calculate gradients: Use the training data to calculate the gradients of the objective function with respect to the model parameters;
[0021] (3) Update parameters: Update the model parameters according to the optimization algorithm;
[0022] (4) Check convergence condition: Determine whether the convergence condition "the value of the loss function is less than a certain threshold" is satisfied. If satisfied, stop the iteration; otherwise, return to step (2) to continue the iteration.
[0023] Preferably, the instruction sending and execution module is responsible for sending the optimized control instructions to the smart home device through wireless communication technology, and the wireless communication technology can be one of Wi-Fi, Bluetooth or Zigbee.
[0024] Preferably, after receiving the instruction, the smart home device performs corresponding operations and feeds back the device status to the status feedback and monitoring module through sensors and wireless communication technology. The sensor is built into the smart home device and is used to monitor the status of the smart home device, including temperature, humidity, and brightness. The wireless communication technology is the MQTT protocol, which can send the status information of the smart home device to the central control screen system.
[0025] Preferably, the status feedback and monitoring module is used to receive the device status feedback from the smart home device, display it to the user, and monitor the operation of the central control screen system at the same time. Using data analysis technology, it extracts monitoring information to provide data support for the learning and adaptation module. A user interface is designed on the status feedback and monitoring module to display the device status through a graphical interface, and monitor the operation of the central control screen system through heartbeat detection and logging technology to ensure that the smart home device is online and working properly;
[0026] Among them, the data analysis technology adopts the LSTM (Long Short-Term Memory) and GRU (Gated Recurrent Unit) recurrent neural network model structures. Utilizing the data analysis ability of the large model, combined with time series analysis and clustering analysis technologies, it mines and analyzes the user's historical operation data and environmental data, discovers the user's preferences and behavior patterns. Through the prediction model, the large model can predict the user's future needs and behaviors, thereby generating personalized control strategies.
[0027] Preferably, the learning and adaptation module analyzes and learns user behavior based on feedback information and monitoring information, using machine learning algorithms and data analysis techniques, continuously learns and adapts to the user's usage habits and preferences, and uses the learning results to optimize policy generation, enabling the central control screen system to better meet the user's needs and preferences.
[0028] The present invention also proposes a multi-modal interaction smart home central control screen method based on a large model, including the following steps:
[0029] S1: Instruction reception: Receive touch instructions, voice instructions, gesture instructions, and image data;
[0030] Among them, for touch instruction reception: The user performs click and slide operations through the touch screen of the central control screen to issue touch instructions. The touch screen of the central control screen senses the user's operation, converts the touch instructions into electrical signals, and transmits them to the processing unit of the central control screen;
[0031] For voice instruction reception: The user issues instructions to the central control screen through voice, such as "turn on the lights in the living room". The microphone array of the central control screen receives the user's voice signal and performs preliminary voice signal processing, such as noise reduction and gain, to improve the accuracy of voice recognition;
[0032] For gesture instruction and image data reception: The user issues instructions to the central control screen through gestures or specific actions, such as waving to indicate turning off the device. The camera of the central control screen captures the user's gesture actions or image information and performs preprocessing, such as image enhancement and cropping, to prepare for subsequent image recognition;
[0033] S2: Instruction processing: Perform voice instruction, gesture instruction, and image data processing;
[0034] The specific steps of voice instruction processing are as follows:
[0035] S2011: After the large model module receives the processed voice signal, it uses natural language processing technology to convert the voice instruction into text form;
[0036] S2012: Perform semantic understanding and intention recognition on the text to extract the user's true needs. "Turn on the lights in the living room" is understood as the need to turn on the lighting equipment in the living room;
[0037] S2013: Generate corresponding control instructions or query requests according to the user's needs and context information, and generate a control instruction to turn on the lights in the living room;
[0038] The specific steps of gesture instruction and image data processing are as follows:
[0039] S2021: After receiving the preprocessed image data, the large model module uses image recognition technology to extract features and classify gesture actions or image information;
[0040] S2022: Identify the user's gesture type or specific object in the image, identify the action of the user waving to indicate turning off the device, and identify the user pointing to a certain smart device;
[0041] S2023: Generate corresponding control instructions or feedback information according to the recognition results, and generate control instructions to turn off the specified device;
[0042] S3: Policy Generation and Optimization: The large model module generates personalized control policies based on the user's historical operation data, environmental data (such as indoor temperature, humidity, light intensity, etc.), and current instructions and data, and uses optimization algorithms (such as genetic algorithms, particle swarm optimization, etc.) to optimize the control policies. The control policies include adjusting the brightness and color temperature of smart lights, turning on or off smart air conditioners, and locking or unlocking smart door locks to meet the user's personalized needs; When considering the optimization algorithm, multiple factors such as energy consumption, comfort, and the service life of smart home devices need to be considered to find the optimal combination of control policies;
[0043] S4: Instruction Sending and Execution: The central control screen sends the generated control instructions to the corresponding smart home devices through wireless communication technology. After receiving the control instructions, the smart home devices parse the instructions and execute the corresponding operations. For example, the smart lights adjust the brightness and color temperature according to the instructions, the smart air conditioner turns on or off according to the instructions, and the smart door lock locks or unlocks according to the instructions;
[0044] S5: Status Feedback and Monitoring: After the smart home devices execute the operations, they feedback the execution results or device status to the central control screen. The central control screen displays the feedback information on the touch screen or feedbacks the information to the user through voice synthesis technology, so that the user can understand the current status of the device. At the same time, the central control screen system continuously monitors the status of the smart home devices, such as whether the device is online and whether it is working properly. In case of an abnormality, the central control screen system alarms in time or takes corresponding measures, such as sending a notice to the user and trying to reconnect the device;
[0045] S6: Learning and Adaptation Stage: The large model module continuously learns and adapts to the user's usage habits and preferences. By analyzing the user's historical operation data and feedback information, it optimizes and improves the control policies. The central control screen system can automatically identify new users or changes in user behavior and adjust the control policies accordingly to provide more personalized services. At the same time, the central control screen system continuously collects data and feedback information during operation to optimize the performance and accuracy of the large model. Through continuous learning and self-optimization, the central control screen system can continuously improve the fluency and intelligence level of the interaction experience.
[0046] Compared with the existing technologies, the beneficial effects of the present invention are as follows:
[0047] 1. The central control screen system integrates multiple interaction methods, including touch, voice, gesture, and even visual recognition, forming a comprehensive multi-modal interaction system. Users can choose the most convenient interaction method according to their habits and needs to control smart home devices. The multi-modal interaction design greatly improves the user experience and makes the operation of the smart home system more natural and smooth;
[0048] 2. The central control screen system adopts advanced large model technology, enabling the central control screen system to have a high level of intelligence. The large model can learn the user's usage habits, preferences, and changes in the home environment, and thus automatically adjust the status of smart home devices to provide personalized services;
[0049] 3. As the core control device of the smart home, the central control screen system highly integrates the control functions of various smart home devices. Users can manage all the smart devices at home through the central control screen only, without separately operating the remote control or mobile APP of each device. The highly integrated design simplifies the operation process and improves the management efficiency;
[0050] 4. The central control screen system has good scalability and can easily access new smart home devices. And users can upgrade the central control screen system at any time to add new devices or functions without replacing the entire system. This scalable design enables the central control screen system to adapt to the development trend of future smart homes and meet the changing needs of users;
[0051] The present invention integrates multiple interaction methods such as touch, voice, gesture, and visual recognition to form a comprehensive multi-modal interaction system, which is applied to the smart home central control screen system. By adopting large model technology, it realizes the intelligent service of the smart home central control screen system, and can uniformly manage all the smart devices at home without separately operating the remote control or mobile APP of each device, providing more intelligent and personalized services for the smart home central control screen system, and realizing the efficient, intelligent, and personalized control and management of smart home devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 is a block diagram of the multi-modal interaction smart home central control screen system based on a large model proposed by the present invention;
[0053] Figure 2 is a flowchart of the multi-modal interaction smart home central control screen method based on a large model proposed by the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0054] The present invention will be further explained below with reference to specific embodiments.
[0055] Embodiment
[0056] Reference Figure 1-2 , this embodiment proposes a multi-modal interactive smart home central control screen system based on a large model, including an instruction receiving module, an instruction processing module, a policy generation and optimization module, an instruction sending and execution module, a status feedback and monitoring module, and a learning and adaptation module;
[0057] The instruction receiving module is connected to the instruction processing module. The policy generation and optimization module is connected to the instruction processing module, the instruction sending and execution module, and the learning and adaptation module. The status feedback and monitoring module is connected to the learning and adaptation module. Both the instruction sending and execution module and the status feedback and monitoring module are connected to external smart home devices;
[0058] The instruction receiving module is responsible for interacting with the user and receiving the user's instruction input through touch, voice, gesture, and image. For touch, it uses touch screen technology to detect the user's touch operation. For voice, it captures sound through a microphone and uses speech recognition technology to convert the sound into text. For gesture and image, it captures images through a camera and uses computer vision technology to recognize gesture and image content;
[0059] The instruction processing module is used to process the received instructions, including speech recognition, image recognition, and gesture recognition, and convert them into control instructions that the system can understand and execute;
[0060] Speech recognition uses natural language processing technology to convert speech text into instructions that the system can understand. The natural language processing technology adopts the sequence-to-sequence (Seq2Seq) model and the advanced NLP model structure of Transformer. Utilizing the deep learning ability of the large model and combining attention mechanism technology, it performs text conversion, semantic understanding, and intent recognition processing on speech instructions. By training a large amount of language data, the large model can learn the grammar, vocabulary, and semantic rules of the language, so as to accurately understand the user's speech instructions;
[0061] Image recognition uses image recognition algorithms to extract features and classify images to recognize the user's intent. The image recognition algorithms adopt advanced image recognition model structures such as YOLO (You Only Look Once) and ResNet. Utilizing the deep learning technology of the large model's convolutional neural network (CNN) and combining data augmentation and transfer learning technology, it extracts features, classifies, and recognizes the images captured by the camera. By training a large amount of image data, the large model can learn the feature patterns in the images, so as to accurately recognize the user's gesture and face image information;
[0062] The strategy generation and optimization module generates personalized control strategies based on the processed instructions and user requirements, and optimizes the strategies through optimization algorithms. Among them, machine learning algorithms are used for strategy generation to generate control strategies according to the user's historical behavior and preferences. The machine learning algorithms are used to build prediction models. Optimization algorithms are used for strategy optimization to adjust the parameters of the prediction model to improve the prediction performance and optimize the generated control strategies to improve the overall performance of the system and user satisfaction.
[0063] The implementation steps of the machine learning algorithm are as follows:
[0064] (1). Data preparation: Collect and clean the processed instruction and user requirement data, and divide it into a training set and a test set;
[0065] (2). Feature engineering: Select or extract features that have a significant impact on the model performance;
[0066] (3). Model selection: Select a suitable machine learning algorithm according to the problem type and data characteristics;
[0067] (4). Model training: Use the training data to train the model and adjust the model parameters to minimize the loss function;
[0068] (5). Model evaluation: Use the test data to evaluate the performance of the model and select the best model;
[0069] (6). Model deployment: Deploy the trained model to the actual application for prediction or classification;
[0070] The implementation steps of the optimization algorithm are as follows:
[0071] (1). Initialize parameters: Set the initial values of the model parameters;
[0072] (2). Calculate gradients: Use the training data to calculate the gradients of the objective function with respect to the model parameters;
[0073] (3). Update parameters: Update the model parameters according to the optimization algorithm;
[0074] (4). Check convergence conditions: Judge whether the convergence condition "the value of the loss function is less than a certain threshold" is satisfied. If it is satisfied, stop the iteration; otherwise, return to step (2) to continue the iteration;
[0075] The instruction sending and execution module is responsible for sending the optimized control instructions to the smart home devices through wireless communication technology, and the wireless communication technology can be one of Wi-Fi, Bluetooth or Zigbee;
[0076] The smart home device is used to execute corresponding operations after receiving an instruction, and feedback the device status to the status feedback and monitoring module through sensors and wireless communication technology. The sensors are built into the smart home device and are used to monitor the status of the smart home device, including temperature, humidity, and brightness. Its wireless communication technology is the MQTT protocol, which can send the status information of the smart home device to the central control screen system;
[0077] The status feedback and monitoring module is used to receive the device status feedback from the smart home device, display it to the user, and monitor the operation of the central control screen system at the same time. Using data analysis technology, it extracts monitoring information to provide data support for the learning and adaptation module. A user interface is designed on the status feedback and monitoring module to display the device status through a graphical interface, and monitor the operation of the central control screen system through heartbeat detection and logging technology to ensure that the smart home device is online and working properly;
[0078] Among them, the data analysis technology adopts the LSTM (Long Short-Term Memory) and GRU (Gated Recurrent Unit) recurrent neural network model structures. Utilizing the data analysis ability of the large model, combined with time series analysis and clustering analysis technologies, it mines and analyzes the user's historical operation data and environmental data, discovers the user's preferences and behavior patterns. Through the prediction model, the large model can predict the user's future needs and behaviors, thereby generating personalized control strategies;
[0079] The learning and adaptation module analyzes and learns the user's behavior using machine learning algorithms and data analysis technology based on the feedback information and monitoring information, continuously learns and adapts to the user's usage habits and preferences, and uses the learning results to optimize strategy generation, so that the central control screen system can better meet the user's needs and preferences.
[0080] This embodiment also proposes a multi-modal interaction smart home central control screen method based on a large model, including the following steps:
[0081] S1: Instruction reception: Receive touch instructions, voice instructions, gesture instructions, and image data;
[0082] Among them, for touch instruction reception: The user clicks and swipes on the touch screen of the central control screen to issue a touch instruction. The touch screen of the central control screen senses the user's operation, converts the touch instruction into an electrical signal, and transmits it to the processing unit of the central control screen;
[0083] For voice instruction reception: The user issues an instruction to the central control screen by voice, such as "Turn on the lights in the living room". The microphone array of the central control screen receives the user's voice signal and performs preliminary voice signal processing, such as noise reduction and gain, to improve the accuracy of voice recognition;
[0084] Gesture instruction and image data reception: The user issues instructions to the central control screen through gestures or specific actions. For example, waving the hand indicates turning off the device. The camera on the central control screen captures the user's gesture actions or image information and performs preprocessing, such as image enhancement and cropping, to prepare for subsequent image recognition;
[0085] S2: Instruction processing: Perform voice instruction, gesture instruction, and image data processing;
[0086] The specific steps of voice instruction processing are as follows:
[0087] S2011: After the large model module receives the processed voice signal, it uses natural language processing technology to convert the voice instruction into text form;
[0088] S2012: Perform semantic understanding and intention recognition on the text, and extract the user's real needs. "Turn on the lights in the living room" is understood as the need to turn on the lighting equipment in the living room;
[0089] S2013: According to the user's needs and context information, generate corresponding control instructions or query requests, and generate a control instruction to turn on the lights in the living room;
[0090] The specific steps of gesture instruction and image data processing are as follows:
[0091] S2021: After the large model module receives the preprocessed image data, it uses image recognition technology to extract features and classify gesture actions or image information;
[0092] S2022: Identify the user's gesture type or specific object in the image, identify the action of the user waving the hand to indicate turning off the device, and identify the user pointing to a certain smart device;
[0093] S2023: Generate corresponding control instructions or feedback information according to the recognition result, and generate a control instruction to turn off the specified device;
[0094] S3: Strategy generation and optimization: The large model module generates personalized control strategies based on the user's historical operation data, environmental data (such as indoor temperature, humidity, light intensity, etc.), as well as the current instructions and data, and uses optimization algorithms (such as genetic algorithms, particle swarm optimization, etc.) to optimize the control strategies. Its control strategies include adjusting the brightness and color temperature of smart lights, turning on or off smart air conditioners, locking or unlocking smart door locks to meet the user's personalized needs; When using the optimization algorithm, multiple factors such as energy consumption, comfort, and the service life of smart home devices need to be considered to find the optimal combination of control strategies;
[0095] S4: Instruction Sending and Execution: The central control screen sends the generated control instructions to the corresponding smart home devices through wireless communication technology. After receiving the control instructions, the smart home devices parse the instructions and perform corresponding operations. For example, the smart lights adjust the brightness and color temperature according to the instructions, the smart air conditioner turns on or off according to the instructions, and the smart door lock locks or unlocks according to the instructions;
[0096] S5: Status Feedback and Monitoring: After performing the operations, the smart home devices feedback the execution results or device status to the central control screen. The central control screen displays the feedback information on the touch screen or feeds the information back to the user through voice synthesis technology, enabling the user to understand the current status of the devices. At the same time, the central control screen system continuously monitors the status of the smart home devices, such as whether the devices are online and whether they are working properly. In case of abnormalities, the central control screen system alarms in a timely manner or takes corresponding measures, such as sending notifications to the user and attempting to reconnect to the devices;
[0097] S6: Learning and Adaptation Phase: The large model module continuously learns and adapts to the user's usage habits and preferences. By analyzing the user's historical operation data and feedback information, it optimizes and improves the control strategy. The central control screen system can automatically identify new users or changes in user behavior and accordingly adjust the control strategy to provide more personalized services. At the same time, the central control screen system continuously collects data and feedback information during operation for optimizing the performance and accuracy of the large model. Through continuous learning and self-optimization, the central control screen system can continuously improve the fluency and intelligence level of the interaction experience;
[0098] In this embodiment, by integrating various interaction methods such as touch, voice, gesture, and visual recognition, a comprehensive multi-modal interaction system is formed and applied to the smart home central control screen system. By adopting the large model technology, the intelligent service of the smart home central control screen system is realized, and all smart devices in the home can be uniformly managed without separately operating the remote control or mobile APP of each device, providing more intelligent and personalized services for the smart home central control screen system and realizing the efficient, intelligent, and personalized control and management of smart home devices.
[0099] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and all should be covered by the protection scope of the present invention.
Claims
1. A multi-modal interactive smart home central control screen system based on a large model, characterized in that, It includes an instruction receiving module, an instruction processing module, a policy generation and optimization module, an instruction sending and execution module, a status feedback and monitoring module, and a learning and adaptation module; The instruction receiving module is connected to the instruction processing module. The policy generation and optimization module is connected to the instruction processing module, the instruction sending and execution module, and the learning and adaptation module. The status feedback and monitoring module is connected to the learning and adaptation module. Both the instruction sending and execution module and the status feedback and monitoring module are connected to external smart home devices.
2. The multimodal interaction smart home central control screen system based on a large model according to claim 1, wherein The instruction receiving module is responsible for interacting with the user and receiving the user's instruction input through touch, voice, gesture, and image. For touch, it uses touch screen technology to detect the user's touch operation. For voice, it uses a microphone to capture sound and voice recognition technology to convert the sound into text. For gesture and image, it uses a camera to capture images and computer vision technology to recognize gesture and image content.
3. The multimodal interaction smart home central control screen system based on a large model according to claim 1, characterized in that, The instruction processing module is used to process the received instructions, including voice recognition, image recognition, and gesture recognition, and convert them into control instructions that the system can understand and execute; For voice recognition, it uses natural language processing technology to convert the voice text into instructions that the system can understand. The natural language processing technology adopts advanced NLP model structures such as sequence-to-sequence models and Transformers, utilizes the deep learning ability of large models, combines attention mechanism technology, and performs text conversion, semantic understanding, and intention recognition processing on voice instructions. By training a large amount of language data, the large model can learn the grammar, vocabulary, and semantic rules of the language, so as to accurately understand the user's voice instructions; For image recognition, it uses image recognition algorithms to extract features and classify images, and recognize the user's intention. The image recognition algorithms adopt advanced image recognition model structures such as YOLO and ResNet, utilize the convolutional neural network deep learning technology of large models, combine data augmentation and transfer learning technology, and perform feature extraction, classification, and recognition on the images captured by the camera. By training a large amount of image data, the large model can learn the feature patterns in the images, so as to accurately recognize the user's gesture and face image information.
4. The multi-modal interactive smart home central control screen system based on a large model according to claim 1, wherein The policy generation and optimization module generates personalized control policies according to the processed instructions and user requirements, and optimizes the policies through optimization algorithms. Among them, policy generation uses machine learning algorithms to generate control policies according to the user's historical behavior and preferences. The machine learning algorithms are used to build prediction models. Policy optimization uses optimization algorithms to adjust the parameters of the prediction models to improve the prediction performance and optimize the generated control policies to improve the overall performance of the system and user satisfaction; The implementation steps of the machine learning algorithms are as follows: (1) Data preparation: Collect and clean the processed instruction and user requirement data, and divide them into a training set and a test set; (2) Feature engineering: Select or extract features that have a significant impact on the model performance; (3) Model selection: Select appropriate machine learning algorithms according to the problem type and data characteristics; (4), Model training: Use the training data to train the model and adjust the model parameters to minimize the loss function; (5), Model evaluation: Use the test data to evaluate the performance of the model and select the best model; (6), Model deployment: Deploy the trained model into the actual application for prediction or classification; The implementation steps of the optimization algorithm are as follows: (1), Initialize parameters: Set the initial values of the model parameters; (2), Calculate gradients: Use the training data to calculate the gradients of the objective function with respect to the model parameters; (3), Update parameters: Update the model parameters according to the optimization algorithm; (4), Check the convergence condition: Determine whether the convergence condition "the value of the loss function is less than a certain threshold" is satisfied. If it is satisfied, stop the iteration; otherwise, return to step (2) to continue the iteration.
5. The multi-modal interactive smart home central control screen system based on a large model according to claim 1, characterized in that The instruction sending and execution module is responsible for sending the optimized control instructions to the smart home device through wireless communication technology, and the wireless communication technology can be one of Wi-Fi, Bluetooth, or Zigbee.
6. The multi-modal interactive smart home central control screen system based on a large model according to claim 1, characterized in that, The smart home device is used to execute corresponding operations after receiving the instructions, and feedback the device status to the status feedback and monitoring module through sensors and wireless communication technology. The sensors are built into the smart home device and are used to monitor the status of the smart home device, including temperature, humidity, and brightness. Its wireless communication technology is the MQTT protocol, and it can send the status information of the smart home device to the central control screen system.
7. The multimodal interaction smart home central control screen system based on a large model according to claim 1, wherein The status feedback and monitoring module is used to receive the device status feedback from the smart home device and display it to the user. At the same time, it monitors the operation of the central control screen system, uses data analysis technology to extract monitoring information, and provides data support for the learning and adaptation module. A user interface is designed on the status feedback and monitoring module to display the device status through a graphical interface. Through heartbeat detection and logging technology, it monitors the operation of the central control screen system to ensure that the smart home device is online and working properly; Among them, the data analysis technology adopts the LSTM and GRU recurrent neural network model structures. Utilizing the data analysis ability of the large model, combined with time series analysis and clustering analysis technologies, it mines and analyzes the user's historical operation data and environmental data, discovers the user's preferences and behavior patterns. Through the prediction model, the large model can predict the user's future needs and behaviors, thereby generating personalized control strategies.
8. The multi-modal interaction smart home central control screen system based on a large model according to claim 1, characterized in that, The learning and adaptation module analyzes and learns the user's behavior according to the feedback information and monitoring information, uses machine learning algorithms and data analysis technology, continuously learns and adapts to the user's usage habits and preferences, and uses the learning results for optimizing strategy generation, so that the central control screen system can better meet the user's needs and preferences.
9. A multi-modal interaction smart home central control screen method based on a large model, characterized in that, It includes the following steps: S1: Instruction reception: Receive touch instructions, voice instructions, gesture instructions, and image data; Among them, for touch instruction reception: The user clicks and swipes on the touch screen of the central control screen to send touch instructions. The touch screen of the central control screen senses the user's operation, converts the touch instructions into electrical signals, and transmits them to the processing unit of the central control screen; Voice command reception: The user issues a command to the central control screen via voice. The microphone array of the central control screen receives the user's voice signal and performs preliminary voice signal processing to improve the accuracy of voice recognition; Gesture command and image data reception: The user issues a command to the central control screen via gestures or specific actions. The camera of the central control screen captures the user's gesture actions or image information and performs preprocessing to prepare for subsequent image recognition; S2: Command processing: Perform voice command, gesture command, and image data processing; The specific steps of voice command processing are as follows: S2011: After the large model module receives the processed voice signal, it uses natural language processing technology to convert the voice command into text form; S2012: Perform semantic understanding and intent recognition on the text, extract the user's real needs, and "turn on the lights in the living room" is understood as the need to turn on the lighting equipment in the living room; S2013: According to the user's needs and context information, generate corresponding control commands or query requests, and generate a control command to turn on the lights in the living room; The specific steps of gesture command and image data processing are as follows: S2021: After the large model module receives the preprocessed image data, it uses image recognition technology to extract features and classify gesture actions or image information; S2022: Identify the type of user's gesture or specific object in the image, identify the action of the user waving to indicate turning off the device, and identify the user pointing to a certain smart device; S2023: Generate corresponding control commands or feedback information according to the recognition result, and generate a control command to turn off the specified device; S3: Strategy generation and optimization: The large model module generates a personalized control strategy based on the user's historical operation data, environmental data, and current commands and data, and uses an optimization algorithm to optimize the control strategy. Its control strategy includes adjusting the brightness and color temperature of smart lights, turning on or off smart air conditioners, locking or unlocking smart door locks to meet the user's personalized needs; When using the optimization algorithm, multiple factors such as energy consumption, comfort, and the service life of smart home devices need to be considered to find the optimal combination of control strategies; S4: Command sending and execution: The central control screen sends the generated control command to the corresponding smart home device via wireless communication technology. After receiving the control command, the smart home device parses the command and executes the corresponding operation; S5: Status feedback and monitoring: After the smart home device executes the operation, it feeds back the execution result or device status to the central control screen. The central control screen displays the feedback information on the touch screen or feeds back the information to the user through voice synthesis technology, so that the user can understand the current state of the device. At the same time, the central control screen system continuously monitors the status of the smart home device. In case of an abnormality, the central control screen system alarms in time or takes corresponding measures; S6: Learning and Adaptation Phase: The large model module continuously learns and adapts to the user's usage habits and preferences. By analyzing the user's historical operation data and feedback information, it optimizes and improves the control strategy. The central control screen system can automatically identify new users or changes in user behavior and accordingly adjust the control strategy to provide more personalized services. At the same time, the central control screen system continuously collects data and feedback information during operation for optimizing the performance and accuracy of the large model.
Citation Information
Cited By
Smart home display module driving method based on Internet of Things
CN120802660A
Home intelligent question and answer and design method and device based on large model and medium
CN121051850A
A large model-based home intelligent question answering and design method, device and medium
CN121051850B