Smart home control method and system based on multi-modal fusion and deep learning

Through the smart home control method of multimodal fusion and deep learning, infrared and radar sensors are used to detect human body data, and particle filtering and graph neural networks are combined to predict human body trajectories. This solves the problem that existing smart home control methods cannot meet multi-scene and multi-space coverage in complex indoor environments, and realizes precise control of smart appliances and user privacy protection.

CN120630745APending Publication Date: 2025-09-12SICHUAN ZHIYUANJI TECH CO LTD

Patent Information

Application Number
CN202511006949.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing smart home control methods cannot meet the multi-scene and multi-space coverage requirements of complex indoor environments. There are problems such as complex multi-APP management, risk of privacy data abuse, and limited compatibility.

Method used

A smart home control method based on multimodal fusion and deep learning is adopted. Human body data is detected through infrared and radar sensors, and human body trajectory is predicted using particle filtering and graph neural network. Combined with the user's historical appliance usage habit model, intelligent startup and standby operation of appliances are realized.

Benefits of technology

It realizes intelligent control of multiple scenes in complex indoor environments, meets the needs of multi-space coverage, and at the same time ensures user privacy and security, avoids the abuse of privacy data and improves the compatibility of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120630745A_ABST
    Figure CN120630745A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of smart home, in particular to a smart home control method and system based on multi-modal fusion and deep learning. The method comprises the steps of triggering a data acquisition instruction, and performing human body data acquisition; acquiring human body data, and generating a global human body motion track sequence according to the preprocessed human body data; predicting a human body track target point and a space region; acquiring a human body track target point and spatial region data, and predicting a target electric appliance and a use probability thereof; the use probability of the target electric appliances is obtained, the target electric appliances are sorted, and electric appliance starting and standby operation is carried out; the system comprises a data acquisition module, a motion trail sequence generation module, a prediction module, a target electric appliance prediction module and an electric appliance operation module. Based on the human body track generated by sensor data, the corresponding electric appliances are pre-started by predicting the room where the user is about to arrive and the electric appliances possibly used by the user, so that multi-scene and multi-space coverage is realized, and a complex indoor environment can be met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of smart home technology, and in particular to a smart home control method and system based on multimodal fusion and deep learning. Background Art

[0002] Smart homes, a product of the deep integration of IoT technology and home living, are reshaping the modern family living experience through intelligent devices, system integration, and scenario-based applications. Currently, smart home control primarily involves remote control via mobile phone / tablet apps, interactive voice and video surveillance, automated scene linkage via pre-programmed settings, and physical control panels or touchscreen terminals. Among these, remote control via mobile phone / tablet apps presents complex multi-app management and advertising interference; voice commands or gesture and posture commands during interactive voice and video surveillance control may be stored on the device, posing the risk of data misuse or leakage through hacker attacks; automated scene linkage via pre-programmed settings lacks flexibility; and physical control panels or touchscreen terminals have limited compatibility.

[0003] In summary, the existing smart home control methods cannot meet the needs of complex indoor environments; therefore, it is very necessary to propose a smart home control method and system that can achieve multi-scene and multi-space coverage and meet the needs of complex indoor environments. Summary of the Invention

[0004] The purpose of the present invention is to provide a smart home control method and system based on multimodal fusion and deep learning, aiming to solve the technical problem that the smart home control methods in the existing technology cannot meet the needs of complex indoor environments.

[0005] To achieve the above objectives, the present invention adopts a smart home control method based on multimodal fusion and deep learning, which includes the following steps:

[0006] Establish an activation threshold, detect the sensor signal strength, compare the activation threshold with the sensor signal strength, and trigger the data collection instruction to collect human body data;

[0007] Acquire human body data, preprocess the human body data, generate a global human body motion trajectory sequence based on the preprocessed human body data, and output it;

[0008] Obtain the global human motion trajectory sequence and spatial graph construction, and predict the target points and spatial areas of the human trajectory;

[0009] Obtain human trajectory target points and spatial area data, collect historical appliance startup times and usage durations, and predict target appliances and their usage probabilities;

[0010] Obtain the usage probability of the target electrical appliances, sort the target electrical appliances according to the usage probability, and perform electrical appliance startup and standby operations based on the sorting results.

[0011] Among them, in the steps of setting an activation threshold, detecting the sensor signal strength, comparing the activation threshold with the sensor signal strength, and triggering a data collection instruction to collect human body data:

[0012] An infrared activation threshold is set. When the infrared intensity signal data exceeds the activation threshold, it is determined that a human body is detected, and an activation instruction for the infrared sensor and radar sensor is issued, and human body data is collected.

[0013] Among them, in the steps of acquiring human body data, preprocessing the human body data, generating a global human body motion trajectory sequence according to the preprocessed human body data, and outputting the sequence:

[0014] Use filtering tools to preprocess human body data;

[0015] Human body data and space model data are obtained respectively, and particle filtering and trajectory smoothing output are performed for indoor scenes.

[0016] Among them, in the steps of obtaining the global human motion trajectory sequence and constructing the spatial graph, and predicting the human trajectory target point and spatial area:

[0017] Obtain the global human motion trajectory sequence and spatial graph construction, convert between architectural space and graph data, and output the converted data;

[0018] Obtain conversion data and perform feature extraction;

[0019] Predict the target points and spatial areas of human body trajectories, and obtain the predicted data of the target points and spatial areas of human body trajectories.

[0020] Among them, after obtaining the conversion data and performing feature extraction steps:

[0021] Perform feature aggregation operations and output aggregated data.

[0022] Among them, in the steps of obtaining human trajectory target points and spatial area data, collecting historical appliance startup times and usage durations, and predicting target appliances and their usage probabilities:

[0023] Obtain human body trajectory target point and spatial area data;

[0024] Build a model of users' historical appliance usage habits and output target appliances and their usage probabilities.

[0025] Among them, in the step of obtaining the human body trajectory target point and spatial area data:

[0026] At the same time, the distribution map of electrical appliance types in the spatial area, historical appliance startup times and usage duration are collected.

[0027] Among them, in the steps of obtaining the usage probability of the target electrical appliances, sorting the target electrical appliances according to the usage probability, and performing electrical appliance startup and standby operations according to the sorting results:

[0028] Set the startup probability threshold, obtain the usage probability of the target appliance, compare the startup probability threshold and the usage probability, output the target appliances to be sorted, and sort multiple target appliances. The target appliance ranked first in probability enters the startup mode, and the target appliance ranked second in probability enters the standby mode.

[0029] Among them, in the step of comparing the startup probability threshold and the usage probability and outputting the target appliances to be sorted:

[0030] When the usage probability of the target appliance is greater than the startup probability threshold, the target appliance is output.

[0031] The present invention also provides a smart home control system based on multimodal fusion and deep learning, including a data acquisition module, a motion trajectory sequence generation module, a prediction module, a target appliance prediction module, and an appliance operation module; wherein:

[0032] The data acquisition module is used to set an activation threshold, detect the sensor signal strength, compare the activation threshold with the sensor signal strength, and trigger a data acquisition instruction to collect human body data;

[0033] The motion trajectory sequence generation module is used to obtain human body data, pre-process the human body data, generate a global human body motion trajectory sequence based on the pre-processed human body data, and output the sequence;

[0034] The prediction module is used to obtain the global human motion trajectory sequence and spatial graph construction, and predict the human trajectory target point and spatial area;

[0035] The target appliance prediction module is used to obtain human body trajectory target points and spatial area data, collect historical appliance startup times and usage durations, and predict target appliances and their usage probabilities;

[0036] The appliance operation module is used to obtain the usage probability of the target appliance, sort the target appliances according to the usage probability, and perform appliance startup and standby operations according to the sorting result.

[0037] The present invention provides a smart home control method and system based on multimodal fusion and deep learning, which respectively uses the data acquisition module, the motion trajectory sequence generation module, the prediction module, the target appliance prediction module, and the appliance operation module to perform the following steps: setting an activation threshold, detecting the sensor signal strength, comparing the activation threshold with the sensor signal strength, and triggering a data acquisition instruction to collect human body data; acquiring human body data, preprocessing the human body data, generating a global human body motion trajectory sequence based on the preprocessed human body data, and outputting it; acquiring the global human body motion trajectory sequence and constructing a spatial map, predicting the target point and spatial area of ​​the human body trajectory; acquiring the human body trajectory Target point and spatial area data are collected to collect historical appliance startup times and usage durations, and target appliances and their usage probabilities are predicted; the usage probabilities of target appliances are obtained, and target appliances are sorted according to the usage probabilities. Based on the sorting results, the appliances are started and put into standby mode. Based on the human body trajectory generated by sensor data, the corresponding appliances are pre-started by predicting the room the user is about to arrive at and the appliances that may be used, achieving multi-scene and multi-space coverage that can meet the needs of complex indoor environments. At the same time, based on the human body trajectory generated by sensor data, human body data is extracted through sensors. This human body data is basic physical characteristics, and the user's privacy characteristics will not be extracted, ensuring the user's privacy security. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0039] Figure 1 It is a flowchart of the steps of the smart home control method based on multimodal fusion and deep learning of the present invention.

[0040] Figure 2 Schematic diagram of the layout of the sensor of the present invention.

[0041] Figure 3 It is a step flow chart of S200 of the present invention.

[0042] Figure 4 It is a step flow chart of S300 of the present invention.

[0043] Figure 5 It is a step flow chart of S400 of the present invention.

[0044] Figure 6 It is a spatial layout diagram of electrical appliances with attached coordinates according to the present invention.

[0045] Figure 7 It is a schematic diagram of the trajectory prediction process of the present invention.

[0046] Figure 8 It is a schematic diagram of the core prediction architecture for implementing spatiotemporal-semantic-physiological joint modeling of the present invention.

[0047] Figure 9 It is a flow chart of the core prediction architecture for implementing spatiotemporal-semantic-physiological joint modeling of the present invention.

[0048] Figure 10 Schematic diagram of the model architecture for predicting appliances to be started based on deep learning and multimodal fusion layer of the present invention.

[0049] Figure 11 It is a schematic diagram of the dynamic adjustment of the weights α, β, and γ of the present invention based on historical AUC.

[0050] Figure 12 It is a schematic diagram of the process of dynamically adjusting the weights α, β, and γ according to the historical AUC of the present invention.

[0051] Figure 13 This is a structural principle diagram of the smart home control system based on multimodal fusion and deep learning of the present invention.

[0052] Figure 14 It is a structural principle diagram of the electronic device of the present invention.

[0053] 601-data acquisition module, 602-motion trajectory sequence generation module, 603-human motion target point prediction module, 604-target electrical appliance prediction module, 605-electrical appliance operation module. DETAILED DESCRIPTION

[0054] Exemplary embodiments are described in detail herein, with examples illustrated in the accompanying drawings. When the following description refers to the drawings, identical numerals in different drawings represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with this application.

[0055] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. As used in this application and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0056] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0057] See also Figures 1 to 12 The present invention provides a smart home control method based on multimodal fusion and deep learning, comprising the following steps:

[0058] S100: Setting an activation threshold, detecting the sensor signal strength, comparing the activation threshold with the sensor signal strength, and triggering a data collection instruction to collect human body data.

[0059] In this embodiment, an activation threshold is set, the sensor signal strength is detected, the activation threshold is compared with the sensor signal strength, and a data collection instruction is triggered to collect human body data. The specific process is as follows:

[0060] An infrared activation threshold is set. When the infrared intensity signal data exceeds the activation threshold, it is preliminarily judged that a human body is detected, and an activation command for the infrared sensor and radar sensor is issued, and human body data is collected.

[0061] In the above process, infrared sensors and radar sensors are used to collect data. At the same time, in order to reduce power consumption, an infrared activation threshold is set. When the infrared signal intensity detected by the infrared sensor exceeds the activation threshold, the command activation is performed, and the infrared sensor and radar sensor are used to collect human body information; when it is judged again that it is not a human body, the radar sensor enters sleep mode and the infrared sensor continues to collect infrared signals.

[0062] This link adds target recognition and judgment. Family members include adults, babies and pets. The infrared characteristics of babies and pets have similar infrared characteristics. At this time, the stored baby data is integrated to identify babies and pets to avoid misidentification of pet activation system. When different human bodies are close together, the infrared and radar data are fused, and the point cloud obtained by the radar is used to reconstruct the body contour (including but not limited to height and shoulder width). The infrared signal assists in judging the stacking of body shapes, so as to achieve the distinction and judgment of different human bodies.

[0063] S200: Acquire human body data, pre-process the human body data, generate a global human body motion trajectory sequence according to the pre-processed human body data, and output the sequence.

[0064] In this embodiment, human body data is acquired and preprocessed, and a global human body motion trajectory sequence is generated based on the preprocessed human body data and output. The specific process is as follows:

[0065] S201: Preprocessing human body data using filtering tools;

[0066] S202: Acquire human body data and space model data respectively, perform particle filtering and trajectory smoothing output for indoor scenes.

[0067] In the above process, filtering operations are used to preprocess human body data. In indoor environments, human movement is often accompanied by complex turns, non-uniform motion, and multipath interference in radar data. In this embodiment, first, radars use different frequency bands to avoid signal overlap. Second, the radars use a frequency band with low wall penetration to reduce interference between radar signals in different rooms in the radar array. Third, algorithms are used. First, a particle filter is used to effectively handle nonlinear and non-Gaussian problems through multi-hypothesis tracking and probability weighting. Second, it incorporates a building space model to filter out false echoes generated by multiple reflections from walls, ceilings, etc. Fourth, time diversity is used to allocate independent time slices to each radar to transmit signals in turn, or random transmission delays are added to the radars to reduce the probability of signal collisions. Fourth, multi-sensor fusion technology is used to fuse information from different sensors, including but not limited to the infrared signal strength and azimuth obtained by the infrared sensor, and the target range, azimuth, pitch angle, velocity, and acceleration obtained by the radar sensor. Based on this data, a global estimated position of the target is obtained, and a human trajectory sequence is smoothly output and stored.

[0068] In the process of selecting particle filters to effectively handle nonlinear and non-Gaussian problems through multi-hypothesis tracking and probability weighting, the following are the specific implementation steps for particle filtering and trajectory smoothing output for indoor scenes: parameter setting and information input, system initialization of particles, radar spatial position correction, calculation of particle weights, systematic resampling, environmental constraint processing, particle weight update, position estimation, and generation of the final trajectory.

[0069] Parameter setting and information input: including infrared, radar information input, and building space model input.

[0070] System initialization particle: Create an all-zero array of specified shape and data type, convert the radar information into polar coordinates and Cartesian coordinates, add random noise, and initialize the velocity and acceleration.

[0071] The formula for converting polar coordinates to Cartesian coordinates is:

[0072] x=rcosθcosφ+Tx

[0073] y=rsinθsinφ+T y

[0074] z=rsinφ+T z

[0075] In the above, r is the distance (meters); θ is the azimuth (radians); φ is the elevation (radians); T x ,T y ,T z is the position of the radar in the global coordinate system.

[0076] Considering the suddenness of indoor motion, a 9-dimensional state vector with random acceleration is used to simulate the complex motion of the human body. The state vector contains position, velocity, and acceleration. The state formula is as follows:

[0077]

[0078] In the above: x, y, z are positions (meters); v x ,v y ,v z is the speed (m / s); a x ,a y ,a z is the speed (m / s 2 ).

[0079] Set up a state transition model (for indoor motion) and use a random acceleration model to simulate sudden motion (including sudden stops and turns indoors). Adjust the acceleration noise parameters based on actual data, increasing the number of particles during intense motion. If the velocity is close to 0 for multiple consecutive frames, reduce the acceleration noise to avoid excessive noise diffusion. To ensure temporal consistency, adjust the state transition model based on the actual time difference Δt, ensuring that the velocity and acceleration updates correctly correspond to the time interval. At the same time, the output trajectory sequence needs to accurately record the estimated position at each time point to ensure the consistency of the timestamp. The formula is as follows:

[0080]

[0081] In the above: Δ t is the time interval; is the random disturbance of acceleration (simulating indoor emergency stop and turn); the observation noise is Gaussian distribution:

[0082] Radar spatial position correction: Extract special position features (including walls, wall intersections, and corners) from CAD architectural drawings, align the real-time point cloud of the radar to be calibrated with the point cloud of the building model, and use a variant of the ICP algorithm to correct the pose deviation (R, t) by taking into account the normal vector of the target point cloud. This allows the point cloud to coincide with the building model, and outputs the corrected radar global coordinates. The formula is as follows:

[0083]

[0084] In the above: p i is the radar point cloud to be calibrated; q i is the building model point cloud; R is the rotation matrix; t is the translation vector (this vector is a three-dimensional vector); n i is the target point q i Normal vector; T is the rigid body transformation matrix, R is the 3x3 rotation matrix (orthogonal matrix, satisfying R T R=I and det(R)=1); t is a 3x1 translation vector, indicating the offset of the coordinate origin; 0 is a 1x3 zero vector, used to keep the dimensions of the homogeneous coordinates consistent.

[0085] T=p corrected =R·p raw +t, where: T is the output radar global coordinate; P corrected is the global coordinate of the radar after correction, P raw are the initial radar coordinates.

[0086] Calculate particle weights: When calculating weights, the degree of match between each particle and all observations is calculated, and the best match is selected. To mitigate multipath effects, if there is multipath interference in the radar, a multi-hypothesis observation model can be introduced to allow particles to simultaneously match multiple possible reflection paths.

[0087] Environmental constraint processing and particle weight update: This step sets environmental constraints (including using the input building model to set indoor boundary constraints) and removes particles that penetrate walls. If there are too few valid particles, they are dynamically supplemented to maintain particle diversity and ensure that the trajectory conforms to physical space constraints. After removing environmental particles, the weight is updated.

[0088] Systematic Resampling: Systematic resampling to improve diversity.

[0089] Position estimation: weighted average method is used.

[0090] Multi-sensor data fusion: includes coordinate system alignment (unifying the time and position data of multiple sensors into the same coordinate system), data fusion, and output. The fusion formula is as follows:

[0091] X k =ω A ·XA +ω B ·X B +ω N ·X N

[0092] In the above, ω is the weight of the weighted sensor, which is dynamically adjusted according to the sensor confidence; A, B, and N are the sensor numbers.

[0093] Generate the final history and current moment trajectory: This includes extracting trajectory coordinates (t, x, y, z), applying a Savitzky-Golay filter to reduce estimation noise, and generating the trajectory.

[0094] Different family members have different movement characteristics. To improve the accuracy of subsequent predictions, the member features entered into the system are encoded and learned using a convolutional neural network (CNN) to improve the accuracy of member identification. Deep learning is used to extract target features and compare them with features of members already entered into the system. If a target is not identified as a member of the system, it is labeled as a "visitor." Based on the identified user number, the historical trajectories of different users are classified, stored, and numbered. Target recognition and judgment are added again in this step. When different human bodies are close together, the infrared and radar data are fused, and the point cloud obtained by the radar is used to reconstruct the body contour (including but not limited to height and shoulder width). The infrared signal assists in determining the stacking characteristics of the body, enabling the distinction between different human bodies.

[0095] In the human body trajectory sequence storage, first, different users are stored separately; second, user trajectory data is marked as used and unused labels, and trajectory data is filtered. The smart device activation data is imported, and the motion trajectory of the first minute of using the smart device is marked with the "used" label, and the remaining data is marked with the "unused" label; third, the human motion trajectory marked with the "used" label is further labeled according to the time period and the name of the smart device used.

[0096] When family members are making judgments, the infrared and radar signal data of different members are first entered into the system, and corresponding radar point cloud models and infrared feature models are formed according to different members, and the data is stored; at the same time, the infrared and radar signal features of common pets can be preset in advance. When the collected infrared and radar features match the infrared and radar signal features of the preset pets in the database, it can be judged as a non-human body. At this time, the radar sensor enters sleep mode, and the infrared sensor continues to collect infrared signals.

[0097] S300: Obtaining a global human motion trajectory sequence and constructing a spatial graph, and predicting the human trajectory target point and spatial area.

[0098] In this embodiment, the global human motion trajectory sequence and spatial graph are obtained to predict the target point and spatial area of ​​the human trajectory. The specific process is as follows:

[0099] S301: Obtaining a global human motion trajectory sequence and constructing a spatial graph, performing conversion between architectural space and graph data, and outputting the converted data;

[0100] S302: Obtain conversion data and perform feature extraction;

[0101] S303: Perform feature aggregation operation and output aggregated data;

[0102] S304: predicting the human body trajectory target point and spatial area based on the aggregated data, and obtaining the human body trajectory target point and spatial area prediction data.

[0103] In the above process, the trajectory prediction process is as follows Figure 7 As shown in the figure, the motion trajectory sequence and spatial graph are input respectively. After constructing the spatial model data, a room connectivity graph is constructed based on the graph neural network (GNN) (nodes are rooms, edges are transition probabilities). The graph attention network (GAT) is used to capture spatial dependencies, adjacency relationships between rooms, and spatial constraints. The rooms are labeled by function and location, and the labels are one-hot encoded. The spatial graph is constructed. Based on the modeling information of each room, the rooms are labeled by function or location and the labels are one-hot encoded. The features (coordinates, velocity, acceleration) are normalized. The normalization formula is:

[0104]

[0105] Convert the direction to a unit vector (sinθ, cosθ).

[0106] The architectural space is modeled as a graph structure. The architectural space graph is the basis of GAT. Each room is divided into a node (including room name, room function, and room floor). The connection relationship or functional relationship between nodes is defined as an edge, and corresponding node features (including: geometric attributes, functional attributes, topological attributes) and edge features (including: connection type, distance, and frequency of passage) are established.

[0107] GAT design (multi-layer GAT): The first layer learns room-level relationships, the second layer learns floor-level relationships, and the third layer learns functional area-level relationships.

[0108] Input layer: node feature matrix: X∈R N·F , construct a real number matrix of N·F, where: X is a matrix, R is a mathematical symbol representing a real number set, N is the number of nodes, and F is the feature dimension.

[0109] Middle layer: Each layer aggregates node features and neighbor features by weight, and gradually extracts higher-level semantic information.

[0110] Attention coefficient calculation: For each pair of adjacent nodes (i, j), calculate the attention score:

[0111] e ij =LeakyReLU(a T [Wx i ‖Wx j ])

[0112] Among them, W is the weight matrix, a is the attention vector, and ‖ represents splicing.

[0113] Use softmax to normalize the attention weights:

[0114]

[0115] Feature aggregation: The new feature of node i is the weighted sum of its neighbors:

[0116]

[0117] Output layer: The embedding of the final layer is the final refinement of the node features by the network, which can capture the global architectural space semantics. The node embedding h in the final GAT layer i , capturing the node's own attributes and contextual relationships, the formula is as follows:

[0118] The final node embedding of the GAT layer = node attributes + contextual relationships of multi-hop neighbors (dynamically weighted by the attention mechanism).

[0119] Using LSTM calculation, the historical and current motion trajectory time series (including: coordinates [x, y], speed, acceleration, direction, timestamp [t]) are used as input; the room sequence (including: room code, coordinates [x, y], spatial dependencies, adjacency between rooms and spatial constraints) is input into the fully connected layer; and the softmax activation function is used.

[0120] Fully connected layer g(x) formula:

[0121] g(x)=σ·(W x+b )

[0122] Where: g(x) is the fully connected layer, W is the weight vector of the gate, b is the bias term, and σ is the sigmoid function.

[0123] Input gate formula:

[0124] i t =σ·(W i ·[h t-1 ,xt ]+b i )

[0125] Where: W i is the weight matrix of the input gate, b i is the input gate bias term, [h t-1 ,x t ] is to connect two vectors into a longer vector, σ is the sigmoid function

[0126] Forget gate formula:

[0127] f t =σ(W f ·[h t-1 ,x t ]+b f )

[0128] Where: W f is the forget gate weight matrix, b f is the forget gate bias term, [h t-1 ,x t ] is to connect two vectors into a longer vector, and σ is the sigmoid function.

[0129] Compute the cell state that describes the current input Used to describe the current input unit state It is calculated based on the previous output and this input:

[0130] C t =tanh(W c ·[h t-1 ,x t ]+b c )

[0131] Calculate the current cell state C t , the current unit state C t The last cell state C t-1 Multiply the forget gate f element-wise t , and then use the current input unit state C t Element-wise multiplication of the input gate i t , and then add the two products to produce, as shown below:

[0132]

[0133] Jiang LSTM's current memory and long-term memory C t-1 Combined together, a new unit state C is formed tDue to the control of the forget gate, information from a long time ago can be saved, and due to the control of the input gate, irrelevant content can be prevented from entering the memory. The output gate controls the influence of long-term memory on the current output. The formula is:

[0134] O t =σ(W o ·[h t-1 ,x t ]+b o )

[0135] Output gate calculation formula:

[0136] h n =O n tanh(C n )

[0137] Feature fusion: The LSTM output h of the last time step of the trajectory n Matrix multiplication is performed with the GNN embeddings of all rooms to generate a spatiotemporal association matrix. Cross-Attention is used to dynamically fuse timing, connections between appliance types, spatial features, and body features. The prediction layer outputs the target room probability distribution and the coordinates of the next time point. Finally, based on the target room probability distribution map output by the prediction layer, the room with the highest probability and the corresponding coordinates are used as the output result.

[0138] The connection between the types of electrical appliances is to first build a graph structure to represent the relationship between different electrical appliances, use the graph convolutional network (GCN) to extract features, and then fuse them with other features; the graph convolutional network (GCN) is a two-layer GCN structure, the first layer aggregates neighbor features, the second layer high-order aggregation, generates the final embedding, and adds an attention mechanism, and finally weighted aggregation output, the node is the electrical appliance type, the edge relationship is the adjacency matrix (such as lamp TV), and appliances with switches, such as lights and heaters, extract the switches; at this time, incorporate the reasons for the connection between appliance types, for example: after the user turns on the air conditioner in the room, he walks to the humidifier again and turns on the humidifier in the room.

[0139] Actual arrival verification: The feedback data from infrared and radar sensors is used to determine whether the target has arrived at the predicted target room and coordinate points for verification. When the user actually enters the room, the true label "Room-True" is recorded; when the user does not enter the room, the error label "Room-Error" is recorded.

[0140] Corrective learning: Utilize the stored recent prediction error data and regularly fine-tune the model with accumulated sample data.

[0141] S400: Acquire human trajectory target points and spatial area data, collect historical appliance startup times and usage durations, and predict target appliances and their usage probabilities.

[0142] In this embodiment, the target point and spatial area data of the human trajectory are obtained, the historical start-up time and usage time of electrical appliances are collected, and the target electrical appliances and their usage probability are predicted. The specific process is as follows:

[0143] S401: Obtain human trajectory target points and spatial area data, and collect the appliance type distribution map in the spatial area, historical appliance startup time, usage duration, external environment at startup time, whether a human body is present at startup, and the relationship between appliance types;

[0144] S402: Construct a model of the user's historical appliance usage habits, and dynamically adjust and output the target appliance and its usage probability based on the model of the user's historical appliance usage habits.

[0145] In the above process, the type of appliance to be started is predicted based on a model architecture of deep learning and multimodal fusion layer. The predicted room and trajectory target points, as well as the appliance type distribution map near the predicted coordinate points, the time series appliance usage matrix, and body features are used to deeply learn the relationship between appliance types to build a user's historical appliance usage habit model. The convolutional neural network (CNN) is used to output the target appliance usage probability through convolution training and feature fusion, thus realizing spatiotemporal-semantic-physiological joint modeling.

[0146] Implement the core prediction architecture of spatiotemporal-semantic-physiological joint modeling, such as Figure 8 and Figure 9 shown.

[0147] Among them, the spatiotemporal-semantic-physiological joint modeling is realized, such as Figure 10 As shown in the figure, LSTM captures movement patterns, CNN locates spatial hotspots, GCN associates electrical appliances to collaborate, and body features reflect real-time status; Input represents input, Timedata represents time data, Room represents room data, Conv represents convolution layer, Relu represents activation function, MaxPoling represents maximum pooling, concat represents feature connection, Softmax represents activation function, Output represents output, Dense represents weighted summation, Trajectory represents predicted trajectory input, and DrawingSheet represents the spatial layout diagram of electrical appliance types with coordinates. Figure 6 As shown, Appliance represents appliance usage information, and Physical represents basic human characteristics.

[0148] Constructing a user's historical appliance usage habit model:

[0149] Obtain the usage probability of the target appliance, predict the startup probability of the target appliance based on the usage probability, and perform the appliance startup and standby operations based on the probability interval. The specific process is:

[0150] According to different family members, the user's appliance usage habits are modeled separately, and the usage probability of various appliances in different modes is obtained.

[0151] Obtaining the usage probability of the target appliance: The user appliance usage habit model is categorized into the following six core habits, each of which contains quantifiable behavioral pattern characteristics:

[0152] Timing rules and habits:

[0153] Fixed time type: For example, turning on the coffee machine at 7:00 on weekdays (standard deviation < 5 minutes)

[0154] Identification method: Gaussian mixture model (GMM)

[0155] 2. Time interval type: For example, the air conditioner runs continuously from 13:00 to 16:00 (probability > 80%)

[0156] Identification method: time window sliding statistics

[0157] 3. Periodic burst type: For example, multiple devices are used concurrently from 20:00 to 23:00 on weekends

[0158] Identification method: For example, Fourier transform detection of periodic peaks

[0159] Duration of habit

[0160] 1. Short-term high-frequency type

[0161] Microwave oven is used for 3-5 minutes at a time (daily frequency > 4 times)

[0162] Identification method: Weibull distribution fitting duration distribution

[0163] 2. Long-term steady-state type

[0164] Refrigerator operates 24 hours a day (downtime <0.1%)

[0165] Identification method: Survival analysis (Kaplan-Meier curve)

[0166] 3. Intermittent working type

[0167] The air conditioner runs for 45 minutes and then stops for 15 minutes.

[0168] Identification method: Hidden Markov Model (HMM) state transition

[0169] Environmental Response Habits:

[0170]

[0171] For example, for every 1°C increase in temperature, the probability of turning on the air conditioner increases by 35% (95% CI: 32-38%).

[0172] Device collaboration habits

[0173] 1. Causal Chain

[0174] For example, the sequence "Turn on TV → Turn off main light → Turn on ambient light" (support > 75%) is implemented using the Apriori association rule mining algorithm.

[0175] 2. Mutually exclusive mode

[0176] For example: the dishwasher and range hood or integrated stove should not be operated at the same time (conflict probability <5%)

[0177] Implementation algorithm: mutual information entropy detection

[0178] 3. Scene bundling

[0179] For example, "breakfast mode": concurrent startup of coffee machine + toaster + radio. Implementation algorithm: Graph Neural Network (GNN) device relationship modeling.

[0180] Human comfort characteristics and habits:

[0181]

[0182] Abnormal deviation habits

[0183] 1. Timing anomaly

[0184] For example: suddenly starting the washing machine at 3 am

[0185] Detection method: Isolation Forest

[0186] 2. Sequence breakage

[0187] For example, the habit of turning on the lights when you get home, which lasted for 5 days, suddenly disappears.

[0188] Detection method: Sequential pattern matching algorithm

[0189] User preferences and habits:

[0190]

[0191] Calculate the activation strength of a specific habit:

[0192] Get historical pattern data;

[0193] Calculate time matching (Gaussian kernel function);

[0194] Computing environment matching

[0195] Quantifying Habit Strength: Calculating the Habit Stability Index

[0196] habit strength =frequency*(1-entropy)*duration tatio

[0197] Meaning: frequency is frequency; 1-entropy behavior entropy; durationratio is duration ratio

[0198] The modeling is achieved by adopting the relevant technical key points of multi-granularity modeling, habit strength quantification, and personalized update mechanism.

[0199] Multi-granularity modeling: (macro: seasonal habits (such as summer air conditioning mode); micro: minute-level operation sequences (such as microwave button combinations));

[0200] Personalized update mechanisms, such as Figure 11 and Figure 12 shown.

[0201] The target appliance and its usage probability are dynamically adjusted and output based on the user's historical appliance usage habit model:

[0202] 1. Core Prediction Architecture

[0203] 1. Obtain the target room and arrival time predicted in step S300;

[0204] 2. Retrieve the list of electrical appliances in the room;

[0205] 3. Load the user's appliance usage habit model;

[0206] 4. Conduct temporal habit analysis, environmental response pattern analysis, relationship analysis between appliance types, and user personal preference analysis on the model;

[0207] 5. Model probability fusion to predict the probability of electrical appliance startup;

[0208] 6. Set the probability threshold judgment model (when the threshold interval setting is reached, the appliance will be pre-started; when the threshold interval setting is not reached, the appliance will remain in standby or sleep mode);

[0209] 7. Record actual appliance usage;

[0210] 8. Determine whether the prediction is accurate based on actual appliance usage;

[0211] 9. Based on the model’s judgment results, if the prediction is correct, the habit model is reinforced; if the judgment is wrong, the prediction model is revised;

[0212] 10. Update model weights based on the judgment results and model enhancement or correction results;

[0213] 11. Optimize future appliance pre-start predictions based on the updated model weights.

[0214] 2. Implementation steps

[0215] 1. Feature spatiotemporal alignment:

[0216] 1) Time feature coding: such as periodic time coding

[0217] Sinusoidally encode the minutes and hours of each day; set the weekday identifier; and normalize the month.

[0218] 2) Environmental feature fusion: Real-time temperature and humidity / light / family member presence data is matched with historical patterns.

[0219] Retrieving the list of appliances in the room: Inputting the coordinate position drawing of the appliances predicted to arrive in the room, using CNN to extract spatial features, and flattening the feature map for use in the fully connected layer;

[0220] The models are analyzed for time sequence habits: the established user history usage habit models are analyzed for time sequence habits, environmental response mode, device coordination relationship, and user personal preference.

[0221] Probabilistic prediction model setting: This model has a three-layer prediction structure, which is divided into basic layer, core layer and enhancement layer.

[0222] The CNN extracts spatial features of the appliance distribution map near the captured coordinates, the GCN captures associations between appliance types, and the LSTM-Attention constructs trajectory semantic vectors and body features (combined with the calculation results for the target person to determine whether the person is a family member and based on user-defined regional appliance activation restrictions). These features are then combined. The attention mechanism dynamically weights the contributions of different modalities (for example, paying more attention to lights than air conditioning at night). Joint decision-making is learned through a fully connected layer, and the final output generates an activation probability distribution. (Based on the user's historical usage history, if multiple related appliances are used simultaneously, they are combined into a unified ranking output.)

[0223]

[0224] Three-layer prediction fusion formula: P 启动 =α·P 生存 +β·P LSTM +γ·P_ 贝叶斯

[0225] (Weights α, β, γ are dynamically adjusted based on historical AUC)

[0226] The weights α, β, and γ are dynamically adjusted based on the historical AUC. Weights are assigned based on the prediction performance of each sub-model on recent historical data (measured by AUC). The better the performance, the greater the weight. A time decay factor is also introduced to make the model focus more on recent performance. The specific implementation steps are as follows:

[0227] Data slicing: Record the prediction results and actual labels of each sub-model by time window (such as daily);

[0228] Calculate window AUC: For each sub-model, calculate its AUC in the most recent N windows;

[0229] Time decay weighting: assign higher weights to recent windows (exponential decay);

[0230] Normalized weight: Calculate the weight ratio of each model based on the weighted AUC;

[0231] Smooth update: To avoid drastic fluctuations in weights, use weighted moving average.

[0232] Algorithm execution process: collect data daily → record the prediction results of each model → determine whether the update time has been reached. If the update time has been reached, calculate the time-attenuated weight → calculate the historical AUC of each model → calculate the weight ratio → smoothly update the weight → apply the new weight prediction. If the update time has not been reached, continue to use the current weight.

[0233] The specific algorithm formula is as follows:

[0234] Let the AUC of the i-th sub-model in time window t be ai,t, and the current time be T.

[0235] Time decay factor: w t =exp(-λ(Tt)), where λ is the decay rate (e.g., λ = 0.1)

[0236] The weighted AUC of the model in the most recent M windows is:

[0237]

[0238] Weight calculation (unnormalized): using softmax transformation

[0239]

[0240] Where τ is a temperature parameter that controls the concentration of the weight distribution (the smaller τ is, the larger the weight of the best performing model is).

[0241] Final weight smooth update (avoid mutation):

[0242]

[0243] Where: η is the smoothing factor.

[0244] Dynamic compensation adjustment strategy

[0245] Probability calibration mechanism:

[0246] Temperature compensation: When the measured temperature deviates from the historical mean ΔT: P adj =P*(1+K·△T)

[0247] K is the constant between the appliance and temperature, such as: air conditioner: k = 0.03 / ℃, refrigerator: k = -0.01 / ℃

[0248] 2) Holiday correction: If during holidays, different probability adjustment settings are made for different appliances

[0249] P * =holiday factors [device], such as TV:1.5

[0250] 2. Corrective learning: Utilize the stored recent prediction error data and regularly fine-tune the model with accumulated sample data.

[0251] 3. Uncertainty Quantification

[0252] 1) Output probability interval,

[0253] For example: P=0.65±0.08 (95% CI)

[0254] 2) Confidence calculation: confidence = 1-(σ habit / μ habit )*duration factor

[0255] Time decay correction: the predicted probability decays exponentially with the time interval

[0256] P adj (t+Δt)=P original exp (λ△t) , (λ is fitted by historical data)

[0257] Business rule injection, such as:

[0258] The started equipment does not need to be predicted to start again;

[0259] Corrected the air conditioner activation probability in low temperature environments.

[0260] S500: Obtain usage probabilities of target electrical appliances, sort the target electrical appliances according to the usage probabilities, and perform electrical appliance startup and standby operations according to the sorting results.

[0261] In this embodiment, the usage probability of the target electrical appliance is obtained, and a decision and execution are made on the target electrical appliance based on the usage probability and the probability threshold range.

[0262] High probability equipment: Startup mode (such as a water heater heating to a set temperature)

[0263] Medium probability devices: Partial preparation mode (e.g., a coffee machine that fills water but doesn't heat it)

[0264] Low probability but high confidence: Preheat (e.g. preheat oven to 100°C)

[0265] For devices with predicted probabilities in the same range, preparation is carried out in the order of predicted probability (the target appliance ranked first in probability enters the start-up mode, and the target appliance ranked second in probability enters the standby mode).

[0266] Based on the user's historical usage records, if multiple related appliances are used at the same time, they will be combined into a unified ranking output.

[0267] Example 1:

[0268] When the user opens the door and prepares to enter the kitchen carrying vegetables, he / she enters the recognition range of the infrared and radar sensors. Through the infrared threshold setting, the infrared sensor preliminarily determines that the target is a human body, starts to start and integrates the radar sensor to accurately identify the target. Through the infrared and radar data of the mobile phone, the target human body is distinguished as a pet or a human body, and the features of the human body are further judged as an adult or a baby. Finally, the features of the adult target are identified and the relevant human body trajectory is generated.

[0269] Based on historical and current human trajectory models (trajectory turning points, speed, acceleration, direction, and other parameters), combined with building models and time data, the system trains and predicts the user's target location (predicting the target's upcoming arrival point to be the kitchen).

[0270] The predicted user destination is combined with the time the user arrives at the destination (kitchen), the type of kitchen appliances, the relationship between appliance types, historical user usage habits, and user-defined area appliance restrictions to output the predicted appliance usage probability value.

[0271] According to the output target appliance usage probability and probability threshold, the target appliance ranked first in probability enters the start mode, and the target appliance ranked second in probability enters the standby mode. As shown in Table 1, the water heater is started and the direct drinking machine enters the standby mode at the same time:

[0272]

[0273]

[0274] Table 1

[0275] Example 2:

[0276] On a winter night in the north, when a user gets up from bed due to the urgent need to go to the bathroom, infrared and radar recognition are used to collect relevant data and generate a human body trajectory model.

[0277] Based on the historical and current human trajectory models (turning points in the trajectory, speed, acceleration, direction and other parameters), combined with the building model, time data, and infrared data (judging the environment as night), training is performed to predict the user's target location (predicting that the target is about to arrive at the bathroom).

[0278] The system uses target point prediction, combined with the user's arrival time at the target point (bathroom), the type of kitchen appliances, the relationships between appliance types, historical user usage habits, and user-defined area appliance restrictions, to output a predicted probability of appliance usage. When a user enters the bathroom at night in winter, they are highly likely to use the lights and smart toilet, and wash their hands with hot water.

[0279] Based on the output target appliance usage probability and probability threshold, the first-ranked appliance is started and the second-ranked appliance is started and put into standby mode; the bathroom night light is started in advance for auxiliary lighting, the smart toilet seat is preheated in advance, and the household zero-cold water water heater is started and hot water is delivered to the washbasin in advance according to the user's body temperature.

[0280] Ranking Appliance Type Probability TOP1 Night light, smart toilet 0.95 TOP2 water heater 0.85 ··· ··· ···

[0281] Table 2

[0282] Corresponding to the aforementioned embodiment of the smart home control method based on multimodal fusion and deep learning, the present application also provides an embodiment of a smart home control system based on multimodal fusion and deep learning.

[0283] Figure 13 This is a block diagram of a smart home control system based on multimodal fusion and deep learning according to an exemplary embodiment. Figure 13 The system may include: a data acquisition module 601, a motion trajectory sequence generation module 602, a human motion target point prediction module 603, a target electrical appliance prediction module 604, and an electrical appliance operation module 605; wherein:

[0284] The data acquisition module 601 is used to set an activation threshold, detect the sensor signal strength, compare the activation threshold with the sensor signal strength, and trigger a data acquisition instruction to collect human body data;

[0285] The motion trajectory sequence generation module 602 is used to obtain human body data, pre-process the human body data, generate a global human body motion trajectory sequence based on the pre-processed human body data, and output the sequence;

[0286] The human motion target point prediction module 603 is used to obtain the global human motion trajectory sequence and spatial graph construction, and predict the human trajectory target point and spatial area;

[0287] The target appliance prediction module 604 is used to obtain human body trajectory target points and spatial area data, collect historical appliance startup times and usage durations, and predict target appliances and their usage probabilities;

[0288] The appliance operation module 605 is used to obtain the usage probability of the target appliance, sort the target appliances according to the usage probability, and perform appliance startup and standby operations according to the sorting result.

[0289] In this embodiment, the data acquisition module 601 sets an activation threshold, detects the sensor signal strength, compares the activation threshold with the sensor signal strength, and triggers the data acquisition instruction to collect human body data; the motion trajectory sequence generation module 602 obtains human body data, pre-processes the human body data, generates a global human body motion trajectory sequence based on the pre-processed human body data, and outputs it; the human body motion target point prediction module 603 obtains the global human body motion trajectory sequence and spatial graph construction, predicts the human body trajectory target point and spatial area; the target appliance prediction module 604 obtains the human body trajectory target point and spatial area data, collects historical appliance startup and shutdown information, and generates a global human body motion trajectory sequence based on the pre-processed human body data ... The target appliance and its usage probability are predicted based on the time and usage duration; the appliance operation module 605 obtains the usage probability of the target appliance, sorts the target appliances according to the usage probability, and performs appliance startup and standby operations according to the sorting results; based on the human body trajectory generated by the sensor data, the corresponding appliance is pre-started by predicting the room the user is about to arrive at and the appliances that may be used, thereby achieving multi-scene and multi-space coverage and being able to meet the needs of complex indoor environments; at the same time, based on the human body trajectory generated by the sensor data, the human body data is extracted by the sensor, and the human body data is basic body characteristics, and the user's privacy characteristics will not be extracted, thereby ensuring the user's privacy security.

[0290] Regarding the system in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0291] For the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described above is only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this application. A person of ordinary skill in the art can understand and implement it without paying any creative work.

[0292] Accordingly, the present application also provides an electronic device, comprising: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the above-mentioned smart home control method based on multimodal fusion and deep learning. Figure 14 As shown in FIG, a hardware structure diagram of a smart home control system based on multimodal fusion and deep learning provided by an embodiment of the present invention is provided, in which any device with data processing capability is located. Figure 14 In addition to the processor, memory, and network interface shown, any device with data processing capabilities in which the apparatus in the embodiment is located may also include other hardware, generally based on the actual functions of the device with data processing capabilities, which will not be described in detail.

[0293] Accordingly, the present application also provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the smart home control method based on multimodal fusion and deep learning as described above. The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the aforementioned embodiments, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), an SD card, a flash card (Flash Card), etc. equipped on the device. Furthermore, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store data that has been output or is to be output.

[0294] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the contents disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of this application and include common knowledge or customary techniques in the art that are not disclosed in this application.

[0295] It will be understood that the present application is not limited to the exact construction that has been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof.

Claims

1. A smart home control method based on multimodal fusion and deep learning, characterized in that: The steps include: Establish an activation threshold, detect the sensor signal strength, compare the activation threshold with the sensor signal strength, and trigger the data collection instruction to collect human body data; Acquire human body data, preprocess the human body data, generate a global human body motion trajectory sequence based on the preprocessed human body data, and output it; Obtain the global human motion trajectory sequence and spatial graph construction, and predict the target points and spatial areas of the human trajectory; Obtain human trajectory target points and spatial area data, collect historical appliance startup times and usage durations, and predict target appliances and their usage probabilities; Obtain the usage probability of the target electrical appliances, sort the target electrical appliances according to the usage probability, and perform electrical appliance startup and standby operations based on the sorting results.

2. The smart home control method based on multimodal fusion and deep learning according to claim 1, characterized in that: In the steps of setting the activation threshold, detecting the sensor signal strength, comparing the activation threshold with the sensor signal strength, and triggering the data collection instruction to collect human body data: An infrared activation threshold is set. When the infrared intensity signal data exceeds the activation threshold, it is determined that a human body is detected, and an activation instruction for the infrared sensor and radar sensor is issued, and human body data is collected.

3. The smart home control method based on multimodal fusion and deep learning according to claim 1, characterized in that: In the steps of acquiring human body data, preprocessing the human body data, generating a global human body motion trajectory sequence based on the preprocessed human body data, and outputting the sequence: Use filtering tools to preprocess human body data; Human body data and space model data are obtained respectively, and particle filtering and trajectory smoothing output are performed for indoor scenes.

4. The smart home control method based on multimodal fusion and deep learning according to claim 1, characterized in that: In the steps of obtaining the global human motion trajectory sequence and constructing the spatial graph, and predicting the human trajectory target points and spatial regions: Obtain the global human motion trajectory sequence and spatial graph construction, convert between architectural space and graph data, and output the converted data; Obtain conversion data and perform feature extraction; Predict the target points and spatial areas of human body trajectories, and obtain the predicted data of the target points and spatial areas of human body trajectories.

5. The smart home control method based on multimodal fusion and deep learning according to claim 4, characterized in that: After obtaining the transformed data and performing the feature extraction steps: Perform feature aggregation operations and output aggregated data.

6. The smart home control method based on multimodal fusion and deep learning according to claim 1, characterized in that: In the steps of obtaining human trajectory target points and spatial area data, collecting historical appliance startup times and usage durations, and predicting target appliances and their usage probabilities: Obtain human body trajectory target point and spatial area data; Build a model of users' historical appliance usage habits and output target appliances and their usage probabilities.

7. The smart home control method based on multimodal fusion and deep learning according to claim 6, characterized in that: In the steps of obtaining the human body trajectory target point and spatial area data: At the same time, the distribution map of electrical appliance types in the spatial area, historical appliance startup times and usage duration are collected.

8. The smart home control method based on multimodal fusion and deep learning according to claim 1, characterized in that: In the steps of obtaining the usage probability of the target electrical appliances, sorting the target electrical appliances according to the usage probability, and performing electrical appliance startup and standby operations according to the sorting results: Set the startup probability threshold, obtain the usage probability of the target appliance, compare the startup probability threshold and the usage probability, output the target appliances to be sorted, and sort multiple target appliances. The target appliance ranked first in probability enters the startup mode, and the target appliance ranked second in probability enters the standby mode.

9. The smart home control method based on multimodal fusion and deep learning according to claim 8, characterized in that: In the step of comparing the startup probability threshold and the usage probability and outputting the target appliances to be sorted: When the usage probability of the target appliance is greater than the startup probability threshold, the target appliance is output.

10. A smart home control system based on multimodal fusion and deep learning, applied to the smart home control method based on multimodal fusion and deep learning as claimed in claim 1, characterized in that: It includes data acquisition module, motion trajectory sequence generation module, prediction module, target appliance prediction module, and appliance operation module; among which: The data acquisition module is used to set an activation threshold, detect the sensor signal strength, compare the activation threshold with the sensor signal strength, and trigger a data acquisition instruction to collect human body data; The motion trajectory sequence generation module is used to obtain human body data, pre-process the human body data, generate a global human body motion trajectory sequence based on the pre-processed human body data, and output the sequence; The prediction module is used to obtain the global human motion trajectory sequence and spatial graph construction, and predict the human trajectory target point and spatial area; The target appliance prediction module is used to obtain human body trajectory target points and spatial area data, collect historical appliance startup times and usage durations, and predict target appliances and their usage probabilities; The appliance operation module is used to obtain the usage probability of the target appliance, sort the target appliances according to the usage probability, and perform appliance startup and standby operations according to the sorting result.

Citation Information

Patent Citations

  • Method and system for controlling indoor electrical appliance

    CN107065591A

  • Urban scene-oriented pedestrian trajectory prediction method, model and storage medium

    CN115071762A

  • Dynamic graph generative adversarial network track prediction method and system for multi-source heterogeneous data

    CN116523002A

  • Load data enhancement method and device based on power utilization habits of users

    CN117851846A

  • Equipment control method and device, smart home equipment and storage medium

    CN118672149A

Cited By

  • Intelligent temperature control energy-saving system and method based on personnel position detection and prediction

    CN120909374A

  • Hot water kettle control method based on change of Internet of Things and hot water kettle

    CN121300192A