Hierarchical feature fusion type device end lightweight recommendation method for smart home

Through hierarchical feature extraction and multimodal feature expression, combined with user portraits and dynamic UI, the problem of insufficient device-side data correlation in the smart home recommendation system is solved, personalized, real-time and lightweight recommendations are achieved, and the accuracy of the recommendation system and user experience are improved.

CN120633844APending Publication Date: 2025-09-12ULTIMATE IOT (HENAN) TECHNOLOGY LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510720209.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing smart home recommendation systems do not fully consider the hierarchical feature associations of multi-source heterogeneous data on the device side, resulting in the loss of semantic association information between multimodal features, static user intent modeling, computing architecture design defects leading to recommendation delays and accuracy loss, and device-side computing resource constraints leading to insufficient recommendation robustness.

Method used

Adopting hierarchical feature extraction and multimodal feature expression methods, the device data is processed through a neural network model, combined with user portraits and dynamic UI, the recommendation strategy is adjusted in real time to generate personalized recommendations.

Benefits of technology

It achieves the accuracy and real-time performance of lightweight recommendations on the device side, improves user experience and system performance, and ensures the personalization and response speed of recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633844A_ABST
    Figure CN120633844A_ABST
Patent Text Reader

Abstract

The invention provides a hierarchical feature fusion type equipment end lightweight recommendation method for smart home. The method is applied to the technical field of data processing, and comprises the steps of collecting basic data of various devices, including device data, device historical operation data and device environment data, and arranging and mapping the collected data into system layer data in an application layer; after hierarchical feature extraction and multi-modal feature expression are carried out on system layer data based on a user scene, context features are generated on an application layer, and the context features are recommended to a specified user through an active intelligent service based on a user portrait and a dynamic UI; and the specified user generates feedback information based on the context features and transmits the feedback information to the system layer. In this way, the recommendation accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to a layered feature fusion device-side lightweight recommendation method for smart homes. Background Art

[0002] With the widespread adoption of smart home devices and the rapid development of IoT technology, users are increasingly demanding personalized scenario recommendation services. Existing smart home recommendation systems typically utilize a centralized recommendation architecture based on cloud computing, collecting diverse information such as device operating data and user operation logs to build a unified recommendation model. However, in practical applications, these technical solutions have been found to have the following key flaws that affect recommendation accuracy: At the data feature processing level, existing systems fail to fully consider the hierarchical feature correlations of heterogeneous multi-source data on the device side. The environmental sensor data, device status time series data, user operation behavior data, etc. generated by smart home devices have significant modal differences and scene dependencies. Traditional feature fusion methods use simple feature splicing or weighted averaging strategies, which leads to the loss of semantic correlation information between multimodal features and cannot accurately characterize the dynamic interaction relationship between user, device and environment, directly reducing the completeness of scene feature expression; in terms of user intent modeling, existing technologies mostly use static portrait update mechanisms, which fail to effectively combine real-time interaction data on the device side for dynamic correction; at the computing architecture level, the design defects of the cloud-edge collaboration mechanism exacerbate recommendation delays and accuracy losses. Existing solutions transmit raw data to the cloud for centralized processing. When encountering network fluctuations or high concurrent requests, the device side cannot obtain the latest recommendation strategy in time, forcing the system to adopt a degraded recommendation mode. In addition, due to the computing resource constraints on the device side, existing lightweight models often oversimplify the feature interaction structure, further weakening the recommendation robustness in complex scenarios and resulting in low recommendation accuracy. Summary of the Invention

[0003] This disclosure provides a lightweight device-side recommendation method based on hierarchical feature fusion for smart homes. The method includes:

[0004] S1, collects basic data of various devices, including device data, device historical operation data, and device environment data, and organizes and maps the collected data into system layer data at the application layer;

[0005] S2, based on user scenarios, generates contextual features at the application layer after hierarchical feature extraction and multimodal feature expression of system-level data. These features are then recommended to designated users through proactive intelligent services based on user profiles and dynamic UIs.

[0006] S3, the designated user generates feedback information based on context features and transmits the feedback information to the system layer.

[0007] Furthermore, S1 includes:

[0008] Users interact with smart homes through voice or APP. The application layer organizes, analyzes, and maps the data of the above interactions, digitizes the interaction data, and maps it into system layer data.

[0009] Based on the above data, the application layer trains the application layer data to obtain models of various devices.

[0010] Furthermore, S2 includes:

[0011] The collected data is organized into system-layer data at the application layer, including: voice data processing, using a neural network model and TTS to process voice data. The TTS processing process is as follows: the neural network model converts voice data into text data. The neural network model uses an ASR model trained based on the Bert model. The processing process is: obtain voice data, convert it into MFCC data, use the ASR model for pre-training, and generate text data from the voice data. When training the ASR model, a large amount of voice data and training vocabulary are required. The generation process of voice data and training vocabulary is as follows: the voice data set is obtained through the Internet platform. The voice data set contains voices from different scenarios and different speakers; the mainstream Internet vocabulary dictionary is selected as the original dictionary, and the professional vocabulary that appears in smart home applications is combined to complete the vocabulary building;

[0012] Equipment historical operation data processing: Equipment historical operation data and equipment environment data are divided into two layers, of which the first layer is the original data of equipment historical operation data and equipment environment data, and the second layer is feature data, that is, data processing is performed on the original data. The data processing flow is as follows: Remove dirty data: For incomplete original data, remove the noise in the data, and fill the missing data with the mean; Characterize the original data, characterize the original data with dirty data removed, and convert the original data with dirty data removed into a multi-dimensional matrix or multi-dimensional vector, where each row of the matrix or vector corresponds to the characteristics of each specific equipment historical operation data sample and equipment environment data sample, including equipment status, time, environmental sensor data and other data; Data screening: For the original sample data set, select different data according to its characteristics;

[0013] Performing feature expression on system layer data, including: S230a text data feature expression, the application layer receives text data, the text data is passed through TinyBERT to obtain corresponding word embedding and position embedding, the word embedding and position embedding are integrated to generate a global context representation, wherein the TinyBERT model is: processing text data to obtain word embedding and position embedding specifically: sentence segmentation, each sequence after sentence segmentation, mapping each word using a word embedding matrix to generate a word embedding vector with a dimension corresponding to the word, using a glove model or a dictionary generated by BERT to perform vocabulary mapping on each word after sentence segmentation to generate a word embedding vector, assigning a value to the position embedding, using the same embedding vector for each position, the first element of each sequence is a special classifier, the last element is a special classifier, and the other elements are ordinary classifiers; training the word segmentation to obtain the corresponding word embedding and position embedding through training; integrating the word embedding and position embedding, training based on the corresponding embedded text data, and obtaining the context representation and global representation through training; using cross entropy to calculate the loss for training;

[0014] Sensor data feature expression, sensor data is data collected by the device in the form of a one-dimensional array; 1DCNN is used to learn the features of the original sensor data, specifically: the input data is a sequence of original sensors, and the sequence of original sensors is expressed as, where is the various types of data collected by device X at time i; the 1DCNN neural network includes 3 convolutional layers, 3 pooling layers, and 2 fully connected layers, and a residual network is added after each convolutional layer to accelerate training; 1DCNN extracts key features from sensor data, obtains raw data, performs feature analysis, converts the raw data into a fixed length, and converts each sensor data into a multidimensional vector of a specific dimension; trains the neural network model, and inputs the converted multidimensional vector of a specific dimension into the neural network model; obtains output data, and the neural network model outputs a multidimensional vector of a specific dimension; the activation function of the 1DCNN uses the Relu function, and the loss function uses the MSE loss function;

[0015] Acquire device features, where device features are multidimensional vectors of a specific device in a fixed scenario; acquire perception features, where the perception features are the abstract states of multiple devices in the environment; acquire text features, where the text features are the user's voice or text interaction information; extract the corresponding feature vectors of each feature through a hierarchical feature extraction algorithm, where the hierarchical feature extraction algorithm is as follows: based on the BERT model, voice text is converted into text features; based on the CNN model, device features are converted into device features; based on the LSTM layer, perception features are converted into multidimensional vectors of a specific dimension; fuse the multidimensional vectors of each feature, generate recommendation data after feature fusion, and make recommendations.

[0016] Furthermore, S3 includes:

[0017] Determine whether the user has received the contextual features recommended by the system. If so, continue; otherwise, end;

[0018] Receive user feedback, pass it to the system layer, and proceed to the next step;

[0019] Integrate feedback information and re-recommend context features.

[0020] Furthermore, it also includes:

[0021] Acquire sensor data, discretize the sensor data, obtain discretized sensor data, normalize the original sensor data so that the data value range of the original sensor data is mapped within a specific interval of [0,1] or [-1,1]; discretize the processed data and divide the data into K dimensional features, where each dimension is data in a specific range.

[0022] Furthermore, it also includes:

[0023] After 1DCNN feature extraction, the output is obtained, wherein the output data is the key features of the sensor data, including trend features and statistical features, wherein the trend features include: trend turning point, trend gradient, and trend cycle; the statistical features are statistical data and probability data.

[0024] Furthermore, the method for processing speech data includes: feature extraction, sampling, value taking and filtering of speech data; speech recognition, converting speech data into text data through a neural network model.

[0025] Furthermore, the mapping process includes: obtaining user semantic expressions and obtaining user interaction language in different environments; organizing the above interaction language to generate data; abstracting the interaction subjects in the above interaction language and mapping them into device layer data; abstracting the interaction objects in the above data and mapping them into system layer data; abstracting the above interaction scenarios and mapping them into system layer data.

[0026] The present disclosure effectively integrates device data, device historical operation data and device environment data through hierarchical feature extraction and multimodal feature expression, which can more comprehensively capture user needs and behaviors, making the recommendation system more accurate; after hierarchical feature extraction, combined with user scenarios and contextual features, it can provide users with personalized recommendations in real time and dynamically, and can flexibly adjust recommendation strategies according to users' real-time behaviors and environmental changes to improve user experience; through user portraits, the system can provide users with more personalized and accurate recommendations based on users' historical behaviors, preferences and other information, and combined with dynamic UI adjustments, it can ensure that the user interface is always adjusted according to users' needs and the current environment.

[0027] It should be understood that the contents described in the Summary of the Invention are not intended to limit the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. The accompanying drawings are provided for a better understanding of the present disclosure and do not constitute a limitation of the present disclosure. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0029] Figure 1 A flowchart of a hierarchical feature fusion device-side lightweight recommendation method for smart homes according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0030] To make the purpose, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of the present disclosure.

[0031] In this document, the term "and / or" simply describes a relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the related objects are in an "or" relationship.

[0032] Figure 1 A flowchart of a hierarchical feature fusion device-side lightweight recommendation method for smart homes according to an embodiment of the present disclosure is shown. The method includes:

[0033] S1, collects basic data of various devices, including device data, device historical operation data, and device environment data, and organizes and maps the collected data into system layer data at the application layer;

[0034] S2, based on user scenarios, generates contextual features at the application layer after hierarchical feature extraction and multimodal feature expression of system-level data. These features are then recommended to designated users through proactive intelligent services based on user profiles and dynamic UIs.

[0035] S3, the designated user generates feedback information based on context features and transmits the feedback information to the system layer.

[0036] According to an embodiment of the present invention, through hierarchical feature extraction and multimodal feature expression, device data, device historical operation data and device environment data are effectively integrated, which can capture user needs and behaviors more comprehensively, making the recommendation system more accurate; after hierarchical feature extraction, combined with user scenarios and contextual features, personalized recommendations can be provided to users in real time and dynamically, and recommendation strategies can be flexibly adjusted according to users' real-time behaviors and environmental changes to improve user experience; through user portraits, the system can provide users with more personalized and accurate recommendations based on users' historical behaviors, preferences and other information, and combined with dynamic UI adjustments, it can ensure that the user interface is always adjusted according to user needs and the current environment.

[0037] In some embodiments, S1 includes: the user interacts with the smart home through voice or APP, the application layer organizes, analyzes, and maps the data of the above interaction, digitizes the interaction data, and maps it into system layer data; based on the above data, the application layer trains the application layer data to obtain models of various types of devices. According to an embodiment of the present invention, training the data by the application layer to obtain models of various types of devices helps the devices respond to user needs more accurately. The establishment of device models can improve the system's control efficiency and accuracy of different devices, making the smart home system more intelligent and better able to adapt to the changing needs of users; by mapping the interaction data into system layer data, the standardization and consistency of the data can be ensured. This structured data helps to optimize the processing efficiency and response speed of the system layer, and improve the performance and stability of the entire smart home system.

[0038] In some embodiments, S2 includes: arranging the collected data into system layer data at the application layer, including: voice data processing, using a neural network model and TTS to process the voice data, the TTS processing process is as follows: the neural network model converts the voice data into text data, the neural network model uses an ASR model trained based on the Bert model, and the processing process is: obtaining voice data, converting it into MFCC data, using the ASR model for pre-training, generating text data through voice data, and when training the ASR model, a large amount of voice data and training vocabulary are required. The generation process of voice data and training vocabulary is as follows: the voice data set is obtained through the Internet platform, and the voice data set contains different scenes and The speech of different speakers; select the Internet mainstream vocabulary dictionary as the original dictionary, and combine it with the professional vocabulary that appears in smart home applications to complete the vocabulary list construction; device history operation data processing: device history operation data and device environment data are divided into two layers, of which the first layer is the original data of device history operation data and device environment data, and the second layer is feature data, that is, data processing is performed on the original data. The data processing flow is as follows: Remove dirty data: For incomplete original data, remove the noise in the data, and fill the missing data with the mean; feature the original data, feature the original data with the dirty data removed, and convert the original data with the dirty data removed into a multi-dimensional matrix or a multi-dimensional vector, where each row of the matrix or vector Corresponding to the characteristics of each specific device historical operation data sample and device environment data sample, including device status, time, environmental sensor data and other data; Data screening: For the original sample data set, different data are selected according to its characteristics; Feature expression of system layer data, including: S230a text data feature expression, the application layer receives text data, and the text data is passed through TinyBERT to obtain the corresponding word embedding and position embedding, and the word embedding and position embedding are integrated to generate a global context representation, where the TinyBERT model is: Processing text data to obtain word embedding and position embedding is specifically: Sentence, each sequence after sentence, each word is mapped using the word embedding matrix to generate a sum The word embedding vector of the corresponding dimension of the word is mapped to the vocabulary of each word after the sentence segmentation using the glove model or the dictionary generated by BERT to generate a word embedding vector and assign a value to the position embedding. Each position uses the same embedding vector. The first and last elements of each sequence are special classifiers, and the other elements are common classifiers. The word segmentation is trained to obtain the corresponding word embedding and position embedding through training. The word embedding and position embedding are fused and trained based on the corresponding embedded text data to obtain contextual representation and global representation through training. The cross entropy loss is used for training. Sensor data feature expression. Sensor data is data collected by the device in the form of a one-dimensional array.1DCNN is used to learn the features of raw sensor data. Specifically, the input data is a sequence of raw sensors, and the sequence of raw sensors is expressed as , where is the various types of data collected by device X at time i; the 1DCNN neural network includes 3 convolutional layers, 3 pooling layers, and 2 fully connected layers. A residual network is added after each convolutional layer to accelerate training; 1DCNN extracts key features from sensor data, obtains raw data, performs feature analysis, converts raw data into a fixed length, and converts each sensor data into a multidimensional vector of a specific dimension; trains the neural network model, and inputs the converted multidimensional vector of a specific dimension into the neural network model; obtains output data, and the neural network model outputs a multidimensional vector of a specific dimension; the 1DCNN The NN uses the Relu function as its activation function and the MSE loss function as its loss function. Device features are obtained, where device features are multidimensional vectors of a specific device in a fixed scenario. Perception features are obtained, where the perception features are the abstract states of multiple devices in an environment. Text features are obtained, where the text features are user voice or text interaction information. The corresponding feature vectors of each feature are extracted using a hierarchical feature extraction algorithm. The hierarchical feature extraction algorithm converts voice text into text features based on the BERT model, converts device features into device features based on the CNN model, and converts perception features into multidimensional vectors of specific dimensions using the LSTM layer. The multidimensional vectors of each feature are fused to generate recommendation data after feature fusion, and recommendations are made.

[0039] In some embodiments, S3 includes: determining whether the user has received the contextual features recommended by the system; if so, continuing; otherwise ending; receiving user feedback, transmitting the feedback information to the system layer, and proceeding to the next step; integrating the feedback information and re-recommending the contextual features. According to an embodiment of the present invention, by determining whether the user has received the contextual features recommended by the system, the accuracy of the recommended content can be ensured; by receiving user feedback and transmitting it to the system, the user feedback can be used as an important basis to enhance the interaction between the system and the user, and improve the user's sense of participation and satisfaction; by adjusting the recommended contextual features in real time, the recommended content can be dynamically adjusted according to the user's feedback, thereby making the recommendation more relevant and personalized, avoiding the blind push of irrelevant information, and improving the user experience.

[0040] In some embodiments, the method further includes: acquiring sensor data, discretizing the sensor data to obtain discretized sensor data, normalizing the raw sensor data so that the data value range of the raw sensor data is mapped within a specific interval of [0, 1] or [-1, 1]; discretizing the processed data, dividing the data into K dimensional features, and each dimension is data in a specific range.

[0041] In some embodiments, it also includes: after 1DCNN feature extraction, an output is obtained, wherein the output data is the key features of the sensor data, including trend features and statistical features, wherein the trend features include: trend turning points, trend gradients, and trend cycles; the statistical features are statistical data and probability data.

[0042] In some embodiments, the method for processing speech data includes: feature extraction, sampling, value taking and filtering of speech data; and speech recognition, converting speech data into text data through a neural network model.

[0043] In some embodiments, the mapping process includes: obtaining user semantic expressions, obtaining user interaction speech in different environments; organizing the above interaction speech to generate data; abstracting the interaction subject in the above interaction speech and mapping it to device layer data; abstracting the interaction object in the above data and mapping it to system layer data; abstracting the above interaction scene and mapping it to system layer data. According to an embodiment of the present invention, by capturing and understanding the user's interaction speech in different environments, it is possible to more accurately reflect user needs and provide users with more personalized and intelligent feedback, reduce misunderstandings and friction in user-device interaction, and thus improve user experience; by abstracting the interaction subject, object, and scene and mapping them to device layer and system layer data, this design enables better collaboration between devices and systems. In multi-device, cross-platform scenarios, this layered abstraction can effectively reduce the coupling between different devices and improve the scalability and maintainability of the system.

[0044] It should be noted that, for simplicity of description, the aforementioned method embodiments are presented as a series of combined actions. However, those skilled in the art should be aware that the present disclosure is not limited by the order of the actions described, as certain steps may be performed in a different order or simultaneously, according to the present disclosure. Furthermore, those skilled in the art should also be aware that the embodiments described in this specification are all optional embodiments, and the actions and modules involved are not necessarily required for the present disclosure. In the technical solutions of the present disclosure, the acquisition, storage, and application of user personal information involved comply with relevant laws and regulations and do not violate public order and good morals. It should be understood that the various forms of the above-mentioned processes can be used, and steps can be reordered, added, or deleted. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions of the present disclosure are achieved. This is not intended to limit the scope of protection of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present disclosure are intended to be included within the scope of protection of the present disclosure.

Claims

1. A hierarchical feature fusion device-side lightweight recommendation method for smart homes, characterized by: include: S1, collects basic data of various devices, including device data, device historical operation data, and device environment data, and organizes and maps the collected data into system layer data at the application layer; S2, based on user scenarios, generates contextual features at the application layer after hierarchical feature extraction and multimodal feature expression of system-level data. These features are then recommended to designated users through proactive intelligent services based on user profiles and dynamic UIs. S3, the designated user generates feedback information based on context features and transmits the feedback information to the system layer.

2. The device-side lightweight recommendation method for hierarchical feature fusion for smart home according to claim 1 is characterized in that: S1 includes: Users interact with smart homes through voice or APP. The application layer organizes, analyzes, and maps the data of the above interactions, digitizes the interaction data, and maps it into system layer data. Based on the above data, the application layer trains the application layer data to obtain models of various devices.

3. The device-side lightweight recommendation method for hierarchical feature fusion for smart home according to claim 2 is characterized in that: S2 includes: The collected data is organized into system-layer data at the application layer, including: voice data processing, using a neural network model and TTS to process voice data. The TTS processing process is as follows: the neural network model converts voice data into text data. The neural network model uses an ASR model trained based on the Bert model. The processing process is: obtain voice data, convert it into MFCC data, use the ASR model for pre-training, and generate text data from the voice data. When training the ASR model, a large amount of voice data and training vocabulary are required. The generation process of voice data and training vocabulary is as follows: the voice data set is obtained through the Internet platform. The voice data set contains voices from different scenarios and different speakers; the mainstream Internet vocabulary dictionary is selected as the original dictionary, and the professional vocabulary that appears in smart home applications is combined to complete the vocabulary building; Equipment historical operation data processing: Equipment historical operation data and equipment environment data are divided into two layers, of which the first layer is the original data of equipment historical operation data and equipment environment data, and the second layer is feature data, that is, data processing is performed on the original data. The data processing flow is as follows: Remove dirty data: For incomplete original data, remove the noise in the data, and fill the missing data with the mean; Characterize the original data, characterize the original data with dirty data removed, and convert the original data with dirty data removed into a multi-dimensional matrix or multi-dimensional vector, where each row of the matrix or vector corresponds to the characteristics of each specific equipment historical operation data sample and equipment environment data sample, including equipment status, time, environmental sensor data and other data; Data screening: For the original sample data set, select different data according to its characteristics; Performing feature expression on system layer data, including: S230a text data feature expression, the application layer receives text data, the text data is passed through TinyBERT to obtain corresponding word embedding and position embedding, the word embedding and position embedding are integrated to generate a global context representation, wherein the TinyBERT model is: processing text data to obtain word embedding and position embedding specifically: sentence segmentation, each sequence after sentence segmentation, mapping each word using a word embedding matrix to generate a word embedding vector with a dimension corresponding to the word, using a glove model or a dictionary generated by BERT to perform vocabulary mapping on each word after sentence segmentation to generate a word embedding vector, assigning a value to the position embedding, using the same embedding vector for each position, the first element of each sequence is a special classifier, the last element is a special classifier, and the other elements are ordinary classifiers; training the word segmentation to obtain the corresponding word embedding and position embedding through training; integrating the word embedding and position embedding, training based on the corresponding embedded text data, and obtaining the context representation and global representation through training; using cross entropy to calculate the loss for training; Sensor data feature expression, sensor data is data collected by the device in the form of a one-dimensional array; 1DCNN is used to learn the features of the original sensor data, specifically: the input data is a sequence of original sensors, and the sequence of original sensors is expressed as, where is the various types of data collected by device X at time i; the 1DCNN neural network includes 3 convolutional layers, 3 pooling layers, and 2 fully connected layers, and a residual network is added after each convolutional layer to accelerate training; 1DCNN extracts key features from sensor data, obtains raw data, performs feature analysis, converts the raw data into a fixed length, and converts each sensor data into a multidimensional vector of a specific dimension; trains the neural network model, and inputs the converted multidimensional vector of a specific dimension into the neural network model; obtains output data, and the neural network model outputs a multidimensional vector of a specific dimension; the activation function of the 1DCNN uses the Relu function, and the loss function uses the MSE loss function; Acquire device features, where device features are multidimensional vectors of a specific device in a fixed scenario; acquire perception features, where the perception features are the abstract states of multiple devices in the environment; acquire text features, where the text features are the user's voice or text interaction information; extract the corresponding feature vectors of each feature through a hierarchical feature extraction algorithm, where the hierarchical feature extraction algorithm is as follows: based on the BERT model, voice text is converted into text features; based on the CNN model, device features are converted into device features; based on the LSTM layer, perception features are converted into multidimensional vectors of a specific dimension; fuse the multidimensional vectors of each feature, generate recommendation data after feature fusion, and make recommendations.

4. The device-side lightweight recommendation method for hierarchical feature fusion for smart home according to claim 3 is characterized in that: S3 includes: Determine whether the user has received the contextual features recommended by the system. If so, continue; otherwise, end; Receive user feedback, pass it to the system layer, and proceed to the next step; Integrate feedback information and re-recommend context features.

5. The device-side lightweight recommendation method for hierarchical feature fusion for smart home according to claim 4 is characterized in that: Also includes: Acquire sensor data, discretize the sensor data, obtain discretized sensor data, normalize the original sensor data so that the data value range of the original sensor data is mapped within a specific interval of [0,1] or [-1,1]; discretize the processed data and divide the data into K dimensional features, where each dimension is data in a specific range.

6. The device-side lightweight recommendation method for hierarchical feature fusion for smart home according to claim 5 is characterized in that: Also includes: After 1DCNN feature extraction, the output is obtained, wherein the output data is the key features of the sensor data, including trend features and statistical features, wherein the trend features include: trend turning point, trend gradient, and trend cycle; the statistical features are statistical data and probability data.

7. The device-side lightweight recommendation method for hierarchical feature fusion for smart home according to claim 6 is characterized in that: The methods for processing speech data include: feature extraction, sampling, value taking and filtering of speech data; speech recognition, converting speech data into text data through a neural network model.

8. The device-side lightweight recommendation method for layered feature fusion for smart home according to claim 7 is characterized in that: The mapping process includes: obtaining user semantic expressions and user interaction scripts in different environments; organizing the above interaction scripts to generate data; abstracting the interaction subjects in the above interaction scripts and mapping them into device-layer data; abstracting the interaction objects in the above data and mapping them into system-layer data; abstracting the above interaction scenarios and mapping them into system-layer data.