A method and system for optimizing intelligent furniture control based on user preferences
By acquiring user monitoring images and analyzing user intentions using predictive analysis models, and outputting furniture control strategies, the problem of inductive intelligent furniture control system is solved, the system's practicality and user experience are improved, and the model is optimized through real-time training.
Patent Information
- Application Number
- CN202411263534.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-10
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-09-10
AI Technical Summary
The existing induction smart furniture control system is too sensitive to the non-purpose movement of users, which can easily lead to misoperation and reduce the practicality of the system and user experience.
By obtaining the user's monitoring images and analyzing the user's intentions using predictive analysis models, the corresponding furniture control strategy is output to avoid misoperation. The predictive analysis model is processed based on the character feature extraction layer, user behavior feature extraction layer, marker feature extraction layer, etc., and outputs furniture control strategies.
It effectively avoids the error control problems caused by users under the induction control of traditional furniture, improves the practicality and user experience of the system, and makes the predictive analysis model more in line with user preferences through real-time training.
Smart Images

Figure CN119202944B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of device control, and in particular to a method and system for optimizing smart furniture control based on user preferences. Background Art
[0002] With the rapid development of smart home technology, smart furniture has gradually become a part of modern home life. Traditional smart furniture mostly uses inductive control systems, such as infrared sensors and motion sensors, to achieve automatic adjustment and control functions. These systems can easily respond to users' physical activities, such as automatically adjusting light brightness, adjusting the height of seats and desks, etc. However, the existing inductive smart furniture control technology also has some shortcomings, especially in the problem of misoperation.
[0003] In actual use, inductive control systems are often overly sensitive to the user's non-purposeful movements. For example, the user may just pass by or do other activities near the furniture unintentionally, but the system will mistakenly interpret these actions as control instructions, thus triggering unnecessary adjustments. This misoperation not only causes trouble to the user, but also reduces the practicality of the system and the user experience. Summary of the invention
[0004] The present invention obtains the user's monitoring image, and analyzes the user's intention through the user's monitoring image and the predictive analysis model, and then outputs the corresponding furniture control strategy. It can control the furniture based on the user's intention, avoiding the problem of miscontrol caused by the user passing by under traditional furniture sensing control; at the same time, the predictive analysis model will also perform real-time training based on the user's operation feedback, so that the predictive analysis model's analysis of the user's intention is more in line with the user's preference.
[0005] A smart furniture control optimization method based on user preference, comprising:
[0006] Acquire a monitoring image of the user at a monitoring time point;
[0007] When the smart furniture is triggered to start, the monitoring images acquired in the previous N times are combined into a time series analysis data set, which is then sent to the prediction analysis model for processing, and the furniture control strategy is output;
[0008] Acquire the user's operation feedback data, and determine the reward value corresponding to the furniture control strategy based on the user's operation feedback data; perform real-time training on the prediction analysis model based on the reward value corresponding to the furniture control strategy;
[0009] The prediction analysis model includes a character feature extraction layer, a user behavior feature extraction layer, a marker feature extraction layer, a fully connected layer, a feature splicing layer and a furniture control strategy output layer. The character feature extraction layer is used to extract character features from the last monitoring image in the time series analysis data set to construct a character feature map; the user behavior feature extraction layer is used to extract user behavior features based on the time series analysis data set to construct a user behavior feature map; the marker feature extraction layer is used to extract marker features from the last monitoring image in the time series analysis data set to construct a marker feature map; the fully connected layer has built-in character feature fully connected units, user behavior The feature fully connected unit and the marker feature fully connected unit, the character feature fully connected unit, the user behavior feature fully connected unit and the marker feature fully connected unit are respectively used to perform fully connected operations on the character feature map, the user behavior feature map and the marker feature map to construct the corresponding character feature fully connected vector, the user behavior feature fully connected vector and the marker feature fully connected vector; the feature concatenation layer is used to concatenate the character feature fully connected vector, the user behavior feature fully connected vector and the marker feature fully connected vector to construct a prediction analysis comprehensive vector; the furniture control strategy output layer is used to process the prediction analysis comprehensive vector to output the furniture control strategy.
[0010] As a preferred aspect of the present invention, the time series analysis data set is sent to the prediction analysis model for processing, and the furniture control strategy is output, which specifically includes the following steps:
[0011] Send the last monitoring image in the time series analysis data set to the character feature extraction layer for processing to construct a character feature map;
[0012] The user behavior feature extraction layer has built-in user feature analysis unit, feature enhancement unit and time series feature extraction unit. The character feature map is sent to the user feature analysis unit for processing to construct the user feature map. In the built-in feature enhancement unit, each monitoring image in the time series analysis data set is traversed. For each selected monitoring image, the following operations are performed: the selected monitoring image is multiplied by the key weight matrix and the value weight matrix respectively to construct the monitoring key feature map K and the monitoring value feature map V, and the user feature map is multiplied by the query weight matrix to construct the monitoring query matrix. The attention weight matrix ATT=softmax(QK T / (d) 0.5 ), and then multiply the attention weight matrix ATT by the monitoring value feature map V to construct a monitoring enhancement image; until all monitoring images in the time series analysis data set are traversed, all monitoring enhancement images are sorted in chronological order to form a time series analysis enhancement data set; the time series analysis enhancement data set is sent to the time series feature extraction unit for processing to construct a user behavior feature map;
[0013] The last monitoring image in the time series analysis data set is sent to the marker feature extraction layer for processing to construct a marker feature map;
[0014] In the fully connected layer, the character feature map, user behavior feature map and marker feature map are fully connected through the character feature fully connected unit, user behavior feature fully connected unit and marker feature fully connected unit respectively, so as to construct the corresponding character feature fully connected vector, user behavior feature fully connected vector and marker feature fully connected vector;
[0015] In the feature concatenation layer, the fully connected vectors of character features, user behavior features, and landmark features are concatenated to construct a comprehensive prediction and analysis vector.
[0016] The prediction analysis comprehensive vector is sent to the furniture control strategy output layer for processing to construct the furniture control strategy.
[0017] As a preferred aspect of the present invention, training the prediction analysis model specifically includes the following steps:
[0018] Pre-train the character feature extraction layer through ImageNet data to initialize the parameters of the character feature extraction layer;
[0019] Acquire a number of user feature training samples, the user feature training samples including monitoring images and their corresponding user feature maps; pre-train the user feature analysis unit through all the user feature training samples to initialize the parameters of the user feature analysis unit;
[0020] Acquire several prediction analysis samples, which include monitoring images, and label the prediction analysis samples through furniture control strategies; combine all labeled prediction analysis samples into a training set, and then send the training set to the prediction analysis model with parameter initialization for training, calculate the loss value, and determine whether the loss value is within a preset range. If the loss value is within the preset range, output the trained prediction analysis model; otherwise, continue to train the prediction analysis model through the training set.
[0021] As a preferred aspect of the present invention, the prediction analysis model is trained in real time based on the reward value corresponding to the furniture control strategy, which specifically includes the following steps:
[0022] At the monitoring time point, the time series analysis data set is sent to the target furniture control strategy output model for processing, the target furniture control strategy is output, and then the furniture control strategy is spliced with the target furniture control strategy to construct reward evaluation data, and the reward evaluation data is sent to the reward evaluation network model for processing. The reward evaluation network model is established based on the BP neural network model, and the reward evaluation value is output. The policy gradient value of the reward evaluation value for the target furniture control strategy is calculated, and the calculated derivative is used as the policy gradient value. Then, the parameters of the furniture control strategy output model are adjusted using the gradient ascent method based on the policy gradient value, so as to realize the real-time training of the furniture control strategy output model; at the same time, it also includes the real-time training of the reward evaluation network model based on the reward value corresponding to the furniture control strategy;
[0023] At the target network update time point, the furniture control strategy output model is directly used as the target furniture control strategy output model to replace the previous target furniture control strategy output model. The two adjacent target network update time points include N monitoring time points, and the time period between the two adjacent target network update time points is recorded as the analysis period.
[0024] As a preferred aspect of the present invention, the real-time training of the reward evaluation network model based on the reward value corresponding to the furniture control strategy specifically includes the following steps:
[0025] The reward evaluation data is labeled with the reward value corresponding to the furniture control strategy, and the labeled reward evaluation data is sent to the reward evaluation network model for processing. The reward evaluation loss value is calculated, and then the reward evaluation data parameters are adjusted through the back propagation algorithm based on the reward evaluation loss value to achieve real-time training of the reward evaluation network model.
[0026] As a preferred aspect of the present invention, the user feature analysis unit is established based on the U-net model.
[0027] As a preferred aspect of the present invention, it also includes regular cleaning of monitoring images.
[0028] The present invention also provides a smart furniture control optimization system based on user preferences, comprising:
[0029] A monitoring image acquisition module is used to acquire the user's monitoring image at the monitoring time point;
[0030] The furniture control strategy output module is used to combine the monitoring images acquired for the previous N times into a time series analysis data set when the smart furniture is triggered to start, send the time series analysis data set to the prediction analysis model for processing, and output the furniture control strategy;
[0031] A real-time training module for the prediction analysis model is used to obtain the user's operation feedback data, determine the reward value corresponding to the furniture control strategy based on the user's operation feedback data, and perform real-time training on the prediction analysis model based on the reward value corresponding to the furniture control strategy;
[0032] The prediction analysis model includes a character feature extraction layer, a user behavior feature extraction layer, a marker feature extraction layer, a fully connected layer, a feature splicing layer and a furniture control strategy output layer. The character feature extraction layer is used to extract character features from the last monitoring image in the time series analysis data set to construct a character feature map; the user behavior feature extraction layer is used to extract user behavior features based on the time series analysis data set to construct a user behavior feature map; the marker feature extraction layer is used to extract marker features from the last monitoring image in the time series analysis data set to construct a marker feature map; the fully connected layer has built-in character feature fully connected units, user behavior The feature fully connected unit and the marker feature fully connected unit, the character feature fully connected unit, the user behavior feature fully connected unit and the marker feature fully connected unit are respectively used to perform fully connected operations on the character feature map, the user behavior feature map and the marker feature map to construct the corresponding character feature fully connected vector, the user behavior feature fully connected vector and the marker feature fully connected vector; the feature concatenation layer is used to concatenate the character feature fully connected vector, the user behavior feature fully connected vector and the marker feature fully connected vector to construct a prediction analysis comprehensive vector; the furniture control strategy output layer is used to process the prediction analysis comprehensive vector to output the furniture control strategy.
[0033] The present invention has the following advantages:
[0034] The present invention obtains the user's monitoring image, and analyzes the user's intention through the user's monitoring image and the predictive analysis model, and then outputs the corresponding furniture control strategy. It can control the furniture based on the user's intention, avoiding the problem of miscontrol caused by the user passing by under traditional furniture sensing control; at the same time, the predictive analysis model will also perform real-time training based on the user's operation feedback, so that the predictive analysis model's analysis of the user's intention is more in line with the user's preference. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 This is a schematic diagram of the structure of the intelligent furniture control optimization system based on user preferences adopted in an embodiment of the present invention.
[0036] Figure 2 A schematic diagram of the structure of the prediction analysis model used in an embodiment of the present invention. DETAILED DESCRIPTION
[0037] In order to enable persons skilled in the art to better understand the technical solutions in the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0038] Embodiment 1, a smart furniture control optimization method based on user preference, comprising:
[0039] The monitoring image of the user is obtained at the monitoring time point. It should be noted that this embodiment is aimed at the smart cabinet in the bedroom. By installing a micro camera on the smart cabinet, the monitoring image of the user can be obtained. By analyzing the monitoring image, the user's intention can be obtained, and then the control of the smart furniture can be realized; and the monitoring time point is set by the user, generally set to 30s; and the obtained monitoring image is grayed;
[0040] When the smart furniture is triggered to start, the monitoring images acquired N times before are combined into a time series analysis data set. It should be noted that there are many ways to be triggered to start. For example, in the infrared sensing method adopted in this embodiment, when the user approaches the smart cabinet, the infrared will sense the user's approach and then start the control program. However, since the user may just walk normally when approaching the smart cabinet, it is necessary to analyze the user's intention; the time series analysis data set is sent to the prediction analysis model for processing, and the furniture control strategy is output. The furniture control strategy can specifically be the corresponding code of any door opened in the smart cabinet, or the corresponding code of no operation; and the smart furniture is controlled based on the furniture control strategy. The specific control method can be to select the corresponding control program based on the furniture control strategy and execute the control program;
[0041] Obtaining the user's operation feedback data, and determining the reward value corresponding to the furniture control strategy based on the user's operation feedback data. It should be added that the user's operation feedback generally refers to the user's operation after the user has controlled the furniture, such as accepting the furniture control strategy and accessing the corresponding items, the reward value can be 1; or rejecting the furniture control strategy and opening other doors in the smart cabinet by itself, the reward value can be 0; based on the reward value corresponding to the furniture control strategy, the predictive analysis model is trained in real time. In the process of real-time training, the predictive analysis model can make the user's intention analysis more in line with the user's preference, thereby realizing the optimization of the control of smart furniture;
[0042] like Figure 2As shown in the figure, the prediction and analysis model includes a character feature extraction layer, a user behavior feature extraction layer, a marker feature extraction layer, a fully connected layer, a feature splicing layer and a furniture control strategy output layer, wherein the character feature extraction layer is established based on a pre-trained convolutional neural network, and is used to extract character features from the last monitoring image in the time series analysis data set to construct a character feature map. By extracting character features from the last monitoring image in the time series analysis data set, the user who triggers the start of the smart furniture can be determined; the user behavior feature extraction layer is used to extract user behavior features based on the time series analysis data set to construct a user behavior feature map; the marker feature extraction layer is also established based on a convolutional neural network, and is used to extract marker features from the last monitoring image in the time series analysis data set to construct a marker feature map. The marker feature map, where the marker can be an item that the user needs to deposit, such as clothing, etc.; the fully connected layer has built-in character feature fully connected units, user behavior feature fully connected units and marker feature fully connected units, which are used to perform full connection operations on the character feature map, user behavior feature map and marker feature map, respectively, to construct corresponding character feature fully connected vectors, user behavior feature fully connected vectors and marker feature fully connected vectors; the feature concatenation layer is used to concatenate the character feature fully connected vectors, user behavior feature fully connected vectors and marker feature fully connected vectors to construct a prediction analysis comprehensive vector; the furniture control strategy output layer is used to process the prediction analysis comprehensive vector to output the furniture control strategy;
[0043] The present application obtains the user's monitoring image, and analyzes the user's intention through the user's monitoring image and the predictive analysis model, and then outputs the corresponding furniture control strategy. It can control the furniture based on the user's intention, avoiding the problem of miscontrol caused by the user passing by under the traditional furniture sensing control; at the same time, the predictive analysis model will also perform real-time training based on the user's operation feedback, so that the predictive analysis model's analysis of the user's intention is more in line with the user's preferences.
[0044] See also Figure 2 , the time series analysis data set is sent to the predictive analysis model for processing, and the furniture control strategy is output, which specifically includes the following steps:
[0045] The last monitoring image in the time series analysis dataset is sent to the character feature extraction layer for processing, during which specific operations include convolution and pooling to construct a character feature map;
[0046] The user behavior feature extraction layer has built-in user feature analysis unit, feature enhancement unit and time series feature extraction unit. The user feature analysis unit is established based on the U-net model. The character feature map is sent to the user feature analysis unit for processing to construct a user feature map. The user feature map only includes the user and background that triggers the start of the smart furniture. In the built-in feature enhancement unit, each monitoring image in the time series analysis data set is traversed. For each selected monitoring image, the following operations are performed: the selected monitoring image is multiplied by the key weight matrix and the value weight matrix respectively to construct the monitoring key feature map K and the monitoring value feature map V, and the user feature map is multiplied by the query weight matrix to construct the monitoring query matrix. The attention weight matrix ATT=softmax(QK T / (d) 0.5 ), and then multiply the attention weight matrix ATT by the monitoring value feature map V to construct a monitoring enhancement image. It should be noted that the query weight matrix, the key weight matrix, and the value weight matrix are all set based on the self-attention mechanism of the Transformer model; until all monitoring images in the time series analysis dataset are traversed, all monitoring enhancement images are sorted in chronological order to form a time series analysis enhancement dataset; the time series feature extraction unit is established based on TCN, and the time series analysis enhancement dataset is sent to the time series feature extraction unit for processing to construct a user behavior feature map; the user feature map can be used to enhance the features corresponding to the user who triggers the start of the smart furniture in each monitoring image in the time series analysis dataset, thereby improving the accuracy of user behavior feature extraction;
[0047] The last monitoring image in the time series analysis data set is sent to the marker feature extraction layer for processing to construct a marker feature map;
[0048] In the fully connected layer, the character feature map, user behavior feature map and marker feature map are fully connected through the character feature fully connected unit, user behavior feature fully connected unit and marker feature fully connected unit respectively, so as to construct the corresponding character feature fully connected vector, user behavior feature fully connected vector and marker feature fully connected vector;
[0049] In the feature concatenation layer, the fully connected vectors of character features, user behavior features, and landmark features are concatenated to construct a comprehensive prediction and analysis vector.
[0050] The prediction analysis comprehensive vector is sent to the furniture control strategy output layer for processing to construct the furniture control strategy. It should be added that the furniture control strategy output layer has a built-in softmax function, which will output the probabilities corresponding to different furniture control strategies and output the furniture control strategy with the highest probability.
[0051] Training the predictive analysis model includes the following steps:
[0052] Pre-train the character feature extraction layer through ImageNet data to initialize the parameters of the character feature extraction layer;
[0053] Obtain several user feature training samples, the user feature training samples include monitoring images and their corresponding user feature maps. It should be noted that the corresponding user feature map refers to a monitoring image with only one user left, and the remaining users can be manually removed, or only one user is photographed during shooting; pre-train the user feature analysis unit through all user feature training samples to initialize the parameters of the user feature analysis unit;
[0054] A number of prediction analysis samples are obtained, including monitoring images, which are obtained by developers according to actual conditions, and the prediction analysis samples are labeled by the furniture control strategy. During the labeling process, for each prediction analysis sample, the prediction analysis sample is labeled by the furniture control strategy actually selected; all labeled prediction analysis samples are combined into a training set, and then the training set is sent to the prediction analysis model with parameter initialization for training, and the loss value is calculated to determine whether the loss value is within a preset range. The preset range is set manually. If the loss value is within the preset range, the trained prediction analysis model is output; otherwise, the prediction analysis model is continuously trained through the training set.
[0055] The prediction analysis model is trained in real time based on the reward value corresponding to the furniture control strategy, which specifically includes the following steps:
[0056] At the monitoring time point, the time series analysis data set is sent to the target furniture control strategy output model for processing, the target furniture control strategy is output, and then the furniture control strategy is spliced with the target furniture control strategy to construct reward evaluation data, and the reward evaluation data is sent to the reward evaluation network model for processing. The reward evaluation network model is established based on the BP neural network model, and the reward evaluation value is output. The policy gradient value of the reward evaluation value for the target furniture control strategy is calculated. The specific calculation method can be to calculate the derivative of the reward evaluation value for the target furniture control strategy, and use the calculated derivative as the policy gradient value, and then use the gradient ascent method based on the policy gradient value to adjust the parameters of the furniture control strategy output model to realize real-time training of the furniture control strategy output model; at the same time, it also includes real-time training of the reward evaluation network model based on the reward value corresponding to the furniture control strategy;
[0057] At the target network update time point, the furniture control strategy output model is directly used as the target furniture control strategy output model to replace the previous target furniture control strategy output model, and the time between two adjacent target network update time points includes N monitoring time points, and the time period between two adjacent target network update time points is recorded as the analysis period, and the value of N is determined by the developer. It should be noted that between two adjacent target network update time points, the furniture control strategy output model will be updated based on the target furniture control strategy output model. Under the correction of the target furniture control strategy output model, it can effectively avoid the furniture control strategy output model from updating too much, which will affect the accuracy of the furniture control strategy output;
[0058] The real-time training of the reward evaluation network model based on the reward value corresponding to the furniture control strategy includes the following steps:
[0059] The reward evaluation data is labeled with the reward value corresponding to the furniture control strategy, and the labeled reward evaluation data is sent to the reward evaluation network model for processing. The reward evaluation loss value is calculated, and then the reward evaluation data parameters are adjusted through the back propagation algorithm based on the reward evaluation loss value to achieve real-time training of the reward evaluation network model.
[0060] Since monitoring images are continuously acquired, in order to avoid insufficient storage space, the monitoring images are also cleaned up regularly. Specifically, once the smart furniture is triggered to start, the monitoring images except for the time series analysis data set are deleted, or when the storage space reaches 80% of the maximum storage space, the N monitoring images at the end of the time sort are retained and the remaining monitoring images are deleted.
[0061] Embodiment 2, a smart furniture control optimization system based on user preferences, such as Figure 1 As shown, including:
[0062] The monitoring image acquisition module is used to acquire the user's monitoring image at the monitoring time point. It should be noted that this embodiment is aimed at the smart cabinet in the bedroom. By installing a micro camera on the smart cabinet, the user's monitoring image can be acquired. By analyzing the monitoring image, the user's intention can be acquired, and then the control of the smart furniture can be realized; and the monitoring time point is set by the user, generally set to 30s; and the acquired monitoring image is grayed;
[0063] The furniture control strategy output module is used to form a time series analysis data set from the monitoring images acquired for the previous N times when the smart furniture is triggered to start. It should be noted that there are many ways to be triggered to start. For example, in the infrared sensing method adopted in this embodiment, when the user approaches the smart cabinet, the infrared will sense the user's approach and then start the control program. However, since the user may just walk normally when approaching the smart cabinet, it is necessary to analyze the user's intention; the time series analysis data set is sent to the prediction analysis model for processing, and the furniture control strategy is output. The furniture control strategy can specifically be the corresponding code for opening any door in the smart cabinet, or the corresponding code for no operation; and the smart furniture is controlled based on the furniture control strategy. The specific control method can be selecting the corresponding control program based on the furniture control strategy and executing the control program;
[0064] The real-time training module of the prediction analysis model is used to obtain the user's operation feedback data and determine the reward value corresponding to the furniture control strategy based on the user's operation feedback data. It should be added that the user's operation feedback generally refers to the user's operation after the user controls the furniture. For example, if the user accepts this furniture control strategy and performs the storage and access of the corresponding items, the reward value can be 1; or if the user rejects this furniture control strategy and opens other doors in the smart cabinet by himself, the reward value can be 0; the prediction analysis model is trained in real time based on the reward value corresponding to the furniture control strategy. In the process of real-time training, the prediction analysis model can make the user's intention analysis more in line with the user's preference, thereby realizing the optimization of the control of smart furniture;
[0065] The prediction analysis model includes a character feature extraction layer, a user behavior feature extraction layer, a marker feature extraction layer, a fully connected layer, a feature splicing layer and a furniture control strategy output layer, wherein the character feature extraction layer is established based on a pre-trained convolutional neural network, and is used to extract character features from the last monitoring image in the time series analysis data set to construct a character feature map. By extracting character features from the last monitoring image in the time series analysis data set, the user who triggers the start of the smart furniture can be determined; the user behavior feature extraction layer is used to extract user behavior features based on the time series analysis data set to construct a user behavior feature map; the marker feature extraction layer is also established based on a convolutional neural network, and is used to extract marker features from the last monitoring image in the time series analysis data set to construct a marker Feature map, where the marker can be the item that the user needs to deposit, such as clothing; the fully connected layer has built-in character feature fully connected units, user behavior feature fully connected units and marker feature fully connected units, which are used to perform full connection operations on the character feature map, user behavior feature map and marker feature map, respectively, to construct corresponding character feature fully connected vectors, user behavior feature fully connected vectors and marker feature fully connected vectors; the feature concatenation layer is used to concatenate the character feature fully connected vectors, user behavior feature fully connected vectors and marker feature fully connected vectors to construct a comprehensive prediction and analysis vector; the furniture control strategy output layer is used to process the comprehensive prediction and analysis vector to output the furniture control strategy.
[0066] It should be understood that those skilled in the art can make improvements or changes based on the above description, and all such improvements and changes should fall within the scope of protection of the appended claims of the present invention. Parts not described in detail in this specification belong to the prior art known to those skilled in the art.
Claims
1. A smart furniture control optimization method based on user preferences, characterized in that: include: Acquire a monitoring image of the user at a monitoring time point; When the smart furniture is triggered to start, the monitoring images acquired in the previous N times are combined into a time series analysis data set, which is then sent to the prediction analysis model for processing, and the furniture control strategy is output; Acquire the user's operation feedback data, and determine the reward value corresponding to the furniture control strategy based on the user's operation feedback data; perform real-time training on the prediction analysis model based on the reward value corresponding to the furniture control strategy; The prediction analysis model includes a character feature extraction layer, a user behavior feature extraction layer, a landmark feature extraction layer, a fully connected layer, a feature concatenation layer, and a furniture control strategy output layer. The character feature extraction layer is used to extract character features from the last monitoring image in the time series analysis data set to construct a character feature graph; the user behavior feature extraction layer is used to extract user behavior features based on the time series analysis data set to construct a user behavior feature graph; The marker feature extraction layer is used to extract the marker features of the last monitoring image in the time series analysis data set to construct a marker feature map; the fully connected layer has built-in character feature fully connected units, user behavior feature fully connected units and marker feature fully connected units, and the character feature fully connected units, user behavior feature fully connected units and marker feature fully connected units are used to perform full connection operations on the character feature map, user behavior feature map and marker feature map respectively to construct corresponding character feature fully connected vectors, user behavior feature fully connected vectors and marker feature fully connected vectors; the feature splicing layer is used to splice the character feature fully connected vectors, user behavior feature fully connected vectors and marker feature fully connected vectors to construct a prediction analysis comprehensive vector; the furniture control strategy output layer is used to process the prediction analysis comprehensive vector to output the furniture control strategy; The time series analysis data set is sent to the predictive analysis model for processing, and the furniture control strategy is output, which specifically includes the following steps: Send the last monitoring image in the time series analysis data set to the character feature extraction layer for processing to construct a character feature map; The user behavior feature extraction layer has built-in user feature analysis units, feature enhancement units, and time series feature extraction units. The character feature map is sent to the user feature analysis unit for processing to construct a user feature map. In the built-in feature enhancement unit, each monitoring image in the time series analysis data set is traversed. For each selected monitoring image, the following operations are performed: the selected monitoring image is multiplied by the key weight matrix and the value weight matrix respectively to construct a monitoring key feature map K and a monitoring value feature map V, and the user feature map is multiplied by the query weight matrix to construct a monitoring query matrix. Calculate the attention weight matrix ATT=softmax(QK T / (d) 0.5 ), and then multiply the attention weight matrix ATT by the monitoring value feature map V to construct a monitoring enhancement image; until all monitoring images in the time series analysis data set are traversed, all monitoring enhancement images are sorted in chronological order to form a time series analysis enhancement data set; the time series analysis enhancement data set is sent to the time series feature extraction unit for processing to construct a user behavior feature map; The last monitoring image in the time series analysis data set is sent to the marker feature extraction layer for processing to construct a marker feature map; In the fully connected layer, the character feature map, user behavior feature map and marker feature map are fully connected through the character feature fully connected unit, user behavior feature fully connected unit and marker feature fully connected unit respectively, so as to construct the corresponding character feature fully connected vector, user behavior feature fully connected vector and marker feature fully connected vector; In the feature concatenation layer, the fully connected vectors of character features, user behavior features, and landmark features are concatenated to construct a comprehensive prediction and analysis vector. The prediction analysis comprehensive vector is sent to the furniture control strategy output layer for processing to construct the furniture control strategy.
2. The method for optimizing smart furniture control based on user preference according to claim 1, characterized in that: Training the predictive analysis model includes the following steps: Pre-train the character feature extraction layer through ImageNet data to initialize the parameters of the character feature extraction layer; Acquire a number of user feature training samples, the user feature training samples including monitoring images and their corresponding user feature maps; pre-train the user feature analysis unit through all the user feature training samples to initialize the parameters of the user feature analysis unit; Acquire several prediction analysis samples, which include monitoring images, and label the prediction analysis samples through furniture control strategies; combine all labeled prediction analysis samples into a training set, and then send the training set to the prediction analysis model with parameter initialization for training, calculate the loss value, and determine whether the loss value is within a preset range. If the loss value is within the preset range, output the trained prediction analysis model; otherwise, continue to train the prediction analysis model through the training set.
3. The method for optimizing smart furniture control based on user preference according to claim 2, characterized in that: The prediction analysis model is trained in real time based on the reward value corresponding to the furniture control strategy, which specifically includes the following steps: At the monitoring time point, the time series analysis data set is sent to the target furniture control strategy output model for processing, the target furniture control strategy is output, and then the furniture control strategy is spliced with the target furniture control strategy to construct reward evaluation data, and the reward evaluation data is sent to the reward evaluation network model for processing. The reward evaluation network model is established based on the BP neural network model, and the reward evaluation value is output. The policy gradient value of the reward evaluation value for the target furniture control strategy is calculated, and the calculated derivative is used as the policy gradient value. Then, the parameters of the furniture control strategy output model are adjusted using the gradient ascent method based on the policy gradient value, so as to realize the real-time training of the furniture control strategy output model; at the same time, it also includes the real-time training of the reward evaluation network model based on the reward value corresponding to the furniture control strategy; At the target network update time point, the furniture control strategy output model is directly used as the target furniture control strategy output model to replace the previous target furniture control strategy output model. The two adjacent target network update time points include N monitoring time points, and the time period between the two adjacent target network update time points is recorded as the analysis period.
4. The method for optimizing smart furniture control based on user preference according to claim 3, characterized in that: The real-time training of the reward evaluation network model based on the reward value corresponding to the furniture control strategy includes the following steps: The reward evaluation data is labeled with the reward value corresponding to the furniture control strategy, and the labeled reward evaluation data is sent to the reward evaluation network model for processing. The reward evaluation loss value is calculated, and then the reward evaluation data parameters are adjusted through the back propagation algorithm based on the reward evaluation loss value to achieve real-time training of the reward evaluation network model.
5. The method for optimizing intelligent furniture control based on user preference according to claim 4, characterized in that: The user feature analysis unit is established based on the U-net model.
6. The method for optimizing intelligent furniture control based on user preference according to claim 5, characterized in that: It also includes regular cleaning of monitoring images.
7. A smart furniture control optimization system based on user preferences, characterized in that: The system applies a smart furniture control optimization method based on user preference as described in any one of claims 1 to 6, including: A monitoring image acquisition module is used to acquire the user's monitoring image at the monitoring time point; The furniture control strategy output module is used to combine the monitoring images acquired for the previous N times into a time series analysis data set when the smart furniture is triggered to start, send the time series analysis data set to the prediction analysis model for processing, and output the furniture control strategy; A real-time training module for the prediction analysis model is used to obtain the user's operation feedback data, determine the reward value corresponding to the furniture control strategy based on the user's operation feedback data, and perform real-time training on the prediction analysis model based on the reward value corresponding to the furniture control strategy; The prediction analysis model includes a character feature extraction layer, a user behavior feature extraction layer, a marker feature extraction layer, a fully connected layer, a feature splicing layer and a furniture control strategy output layer. The character feature extraction layer is used to extract character features from the last monitoring image in the time series analysis data set to construct a character feature map; the user behavior feature extraction layer is used to extract user behavior features based on the time series analysis data set to construct a user behavior feature map; the marker feature extraction layer is used to extract marker features from the last monitoring image in the time series analysis data set to construct a marker feature map; the fully connected layer has built-in character feature fully connected units, user behavior The feature fully connected unit and the marker feature fully connected unit, the character feature fully connected unit, the user behavior feature fully connected unit and the marker feature fully connected unit are respectively used to perform fully connected operations on the character feature map, the user behavior feature map and the marker feature map to construct the corresponding character feature fully connected vector, the user behavior feature fully connected vector and the marker feature fully connected vector; the feature concatenation layer is used to concatenate the character feature fully connected vector, the user behavior feature fully connected vector and the marker feature fully connected vector to construct a prediction analysis comprehensive vector; the furniture control strategy output layer is used to process the prediction analysis comprehensive vector to output the furniture control strategy.
Citation Information
Patent Citations
Air conditioner control method and device based on user behaviors, air conditioner and storage medium
CN116734411A