Smart home user behavior prediction method and system

By combining multimodal fusion datasets and indoor user behavior models with 3D indoor environment and biomechanical modeling, the problems of insufficient behavior recognition accuracy and response lag in traditional smart home systems are solved, enabling accurate prediction of user behavior and proactive services.

CN121880848APending Publication Date: 2026-04-17E-SURFING DIGITAL LIFE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-22
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Traditional smart home systems suffer from insufficient accuracy in behavior recognition, significant response lag, and a lack of physical modeling and temporal correlation analysis, resulting in inaccurate prediction of user behavior and an inability to provide proactive services.

Method used

By employing a multimodal fusion dataset and combining it with an indoor user behavior model, and through a 3D indoor environment model and biomechanical modeling, multiple sensor data are acquired in real time. The instantaneous action fit value and behavior fit index of user behavior purpose are calculated, and the user behavior purpose is predicted using a relationship tree, thus achieving accurate prediction.

Benefits of technology

It enables accurate prediction of user behavior, breaks through the traditional passive response mode, provides proactive services, and improves the responsiveness of smart home systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121880848A_ABST
    Figure CN121880848A_ABST
Patent Text Reader

Abstract

The invention relates to a smart home user behavior prediction method and system, and belongs to the technical field of behavior prediction, and the method comprises the steps: obtaining one or more multi-modal fusion data sets in real time; performing behavior action verification of a user behavior purpose on the multi-modal fusion data set based on an indoor user behavior model, and determining an instantaneous action fitting value of the user behavior purpose based on the similarity between the multi-modal fusion data set and the multi-modal fusion data set in the indoor user behavior model; determining a behavior fitting index of a user behavior target according to the plurality of continuously generated instantaneous action fitting values, and determining the user behavior execution target of the user based on the behavior fitting index and a preset index; and determining a possible prediction purpose and a determined behavior prediction purpose of the user based on the number of the execution user behavior purposes and the relationship tree of the behavior purposes, and starting the smart home corresponding to the user behavior of the determined behavior prediction purpose. According to the method, accurate prediction of user behaviors is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of behavior prediction technology, and in particular relates to a method and system for predicting the behavior of smart home users. Background Technology

[0002] With the rapid development of IoT and AI technologies, smart home systems have evolved from single-device control to scenario-based and proactive services. However, traditional smart home systems generally suffer from insufficient accuracy in behavior recognition and significant response delays. Therefore, how to better predict user behavior has become an urgent problem to be solved. Summary of the Invention

[0003] In view of the shortcomings of the existing technology, the purpose of the invention is to provide a method and system for predicting smart home user behavior.

[0004] In a first aspect, the present invention proposes a method for predicting user behavior in smart homes, comprising: S1, acquiring one or more multimodal fusion datasets in real time, each of the multimodal fusion datasets including real-time data from multiple types of sensors at the same time; S2, verifying the user's behavioral purpose using the multimodal fusion datasets based on an indoor user behavior model, and determining an instantaneous action fit value for the user's behavioral purpose based on the similarity between the multimodal fusion datasets and the multimodal fusion datasets in the indoor user behavior model; S3, determining a behavioral fit index for the user's behavioral purpose based on multiple consecutively generated instantaneous action fit values, and determining the user's intended user behavior based on the behavioral fit index and a preset index; S4, determining possible predicted purposes and the user's determined behavioral prediction purpose based on the relationship tree between the number of executed user behavioral purposes and the behavioral purposes, and activating the smart home corresponding to the user behavior with the determined behavioral prediction purpose.

[0005] Furthermore, the user behavior verification based on the multimodal fusion dataset using the indoor user behavior model includes: inputting the multimodal fusion dataset into the indoor user behavior model to obtain the user behavior verification; wherein, the indoor user behavior model is a three-dimensional indoor environment model constructed using CAD tools, and the three-dimensional indoor environment model includes a human body, furniture, and sensors. The construction process of the three-dimensional indoor environment model includes: converting the indoor space into a three-dimensional geometric model; converting the three-dimensional geometric model into a mathematical model with computable physical equations; creating a controllable, biomechanically conforming human body in the mathematical model and defining it as standard behavior; and repeatedly adjusting the parameters of the mathematical model using experimental data to determine whether the virtual space is consistent with the real space.

[0006] Further, based on the similarity between the multimodal fusion dataset and the multimodal fusion dataset in the indoor user behavior model, the instantaneous action fit value for the user behavior purpose is determined, including: converting the multimodal fusion dataset into a first data vector, and converting the multimodal fusion dataset in the indoor user behavior model into a second data vector; based on the similarity formula... Calculate the target similarity between the first data vector and the second data vector, where A represents the first data vector and B represents the second data vector; when there are multiple multimodal fusion datasets and the target similarity is determined to be greater than a preset similarity threshold, determine the average value of the sum of the multiple target similarities and use the average value as the average fitting similarity; multiply the number of action fitting targets by the average fitting similarity as the instantaneous action fitting value for the user behavior purpose; wherein, for each time the target similarity is determined to be greater than the preset similarity threshold, the number of action fitting targets is increased by one, until the target similarity calculations for multiple multimodal fusion datasets are completed, and the number of action fitting targets is obtained.

[0007] Further, determining the user behavior target behavior fit index based on the continuously generated multiple instantaneous action fit values ​​includes: obtaining multiple instantaneous action fit values ​​continuously generated for the user behavior target within a preset time period; sorting all the instantaneous action fit values ​​according to the order of their generation time based on timestamps, and calculating the absolute difference between adjacent instantaneous action fit values ​​after sorting to obtain multiple instantaneous action fit fluctuation values; taking the average of the sum of the multiple instantaneous action fit fluctuation values ​​as the average instantaneous action fit fluctuation value; taking the average of the sum of the multiple instantaneous action fit values ​​as the average instantaneous action fit value; based on Calculate the user behavior target behavior fit index, where DCA represents the user behavior target behavior fit index, ER(s) represents the average instantaneous action fit value, and PE(q) represents the average instantaneous action fit fluctuation value.

[0008] Furthermore, based on the behavior fit index and the preset index, the user's purpose for performing the user behavior is determined, including: determining whether the behavior fit index is greater than the preset index; if the behavior fit index is greater than the preset index, the user's purpose for performing the user behavior is taken as the purpose for performing the user behavior.

[0009] Further, determining possible predicted purposes based on the number of user behavior purposes and the relational tree of the behavior purposes includes: obtaining the relational tree of the behavior purposes, the relational tree including user behavior purposes and the temporal dependencies of the user behavior purposes, wherein a first user behavior purpose in the temporal dependency relationship points to a second user behavior purpose with a directed edge, and the second user behavior purpose is used as a dependent target purpose; determining the dependent target purpose corresponding to the user behavior purpose based on the number of user behavior purposes, and determining the number of dependent target purposes that can be depended on; if it is determined that the number of dependent target purposes is greater than the dependency quantity threshold, the dependent target purpose is used as the possible predicted purpose.

[0010] Further, determining the user's intended behavior includes: obtaining the target number of times the possible intended behavior is marked within a preset time period; and if the target number is greater than a preset number, using the possible intended behavior as the intended behavior.

[0011] A second aspect of the present invention provides a smart home user behavior prediction system, comprising: an acquisition module for real-time acquisition of one or more multimodal fusion datasets, each of the multimodal fusion datasets including real-time data from multiple types of sensors at the same time; a first determination module for verifying the user behavior purpose of the multimodal fusion datasets based on an indoor user behavior model, and determining an instantaneous action fit value for the user behavior purpose based on the similarity between the multimodal fusion datasets and the multimodal fusion datasets in the indoor user behavior model; a second determination module for determining a behavior fit index for the user behavior purpose based on multiple consecutively generated instantaneous action fit values, and determining the user's execution user behavior purpose based on the behavior fit index and a preset index; and a third determination module for determining possible predicted purposes and the user's determined behavior prediction purpose based on the number of executed user behavior purposes and a relationship tree of behavior purposes, and activating the smart home corresponding to the user behavior with the determined behavior prediction purpose.

[0012] A third aspect of the present invention provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the method described in any one aspect of the present invention.

[0013] A fourth aspect of the present invention provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to perform the method described in any one of the first aspects of the present invention.

[0014] The beneficial effects of this invention are as follows:

[0015] The smart home user behavior prediction method and system of this invention acquires one or more multimodal fusion datasets in real time, each of which includes real-time data from multiple types of sensors at the same time. Based on an indoor user behavior model, the system verifies the user's behavioral purpose using the multimodal fusion datasets, and determines the instantaneous action fit value based on the similarity between the multimodal fusion datasets and the multimodal fusion datasets in the indoor user behavior model. It then determines a behavioral fit index based on multiple consecutively generated instantaneous action fit values, and determines the user's intended behavior purpose based on the behavioral fit index and a preset index. Finally, it determines possible predicted purposes and the user's confirmed behavioral prediction purpose based on the relationship tree between the number of executed user behavioral purposes and the behavioral purposes, and activates the smart home corresponding to the user behavior with the confirmed behavioral prediction purpose. This method breaks through the passive response mode of traditional smart homes and achieves accurate prediction of user behavior. Attached Figure Description

[0016] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts. It is obvious that the drawings described below are merely some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings.

[0017] Figure 1 This is a flowchart of a smart home user behavior prediction method according to an embodiment of the present invention;

[0018] Figure 2 This is a flowchart of a smart home user behavior prediction method according to a specific embodiment of the present invention;

[0019] Figure 3 This is a schematic diagram of a smart home user behavior prediction system according to an embodiment of the present invention;

[0020] Figure 4 This is a structural block diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0021] To enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. It should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0022] Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts disclosed in this invention.

[0023] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance. The terms "installed," "connected," and "linked" should be interpreted broadly; for example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0024] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of methods and systems consistent with some aspects of the invention as detailed in the appended claims.

[0025] With the rapid development of IoT and AI technologies, smart home systems have evolved from single-device control to scenario-based and proactive services. However, traditional smart home systems generally suffer from the following technical bottlenecks, limiting the accuracy of user behavior prediction and the proactivity of services:

[0026] 1. Insufficient accuracy in behavior recognition: Single modalities (such as pressure sensors) have difficulty distinguishing between similar actions such as "bending down to pick up a book" and "tying shoelaces", resulting in a high misjudgment rate.

[0027] 2. Significantly delayed response: It relies on users to actively trigger commands (such as "turn on the lights") and cannot anticipate needs in advance.

[0028] 3. Lack of physical modeling: Existing solutions are mostly based on statistical learning and lack modeling of the physical characteristics of sensors (such as the multipath effect of millimeter-wave radar), resulting in a large deviation between simulation data and real-world scenarios.

[0029] 4. Weak temporal correlation analysis: It is unable to capture the temporal dependencies between behaviors such as "eating → washing dishes", making it difficult to achieve automatic device linkage (such as preheating the dishwasher before the end of the meal).

[0030] Although existing technologies attempt to improve predictive capabilities through multimodal data fusion, a complete technological loop of "perception-modeling-reasoning" has not yet been formed, especially in physical field simulation, dynamic threshold adaptation, and time-series dependent modeling, where there are significant shortcomings.

[0031] To this end, the present invention proposes a method, system and related devices for predicting smart home user behavior. Specifically, the following describes the method, system and related devices for predicting smart home user behavior according to embodiments of the present invention with reference to the accompanying drawings.

[0032] Figure 1 This is a flowchart of a smart home user behavior prediction method according to an embodiment of the present invention. It should be noted that the smart home user behavior prediction method of this embodiment can be applied to the smart home user behavior prediction system of this embodiment. This smart home user behavior prediction system can be configured on an electronic device or in a server. This application does not limit the scope of the application.

[0033] like Figure 1 As shown, the smart home user behavior prediction method includes:

[0034] S110 acquires one or more multimodal fusion datasets in real time, each multimodal fusion dataset including real-time data from multiple types of sensors at the same time.

[0035] In embodiments of the present invention, various types of sensors include, but are not limited to, millimeter-wave radar, thermal imaging cameras, pressure sensors, etc. The installation locations and functions of different types of sensors vary.

[0036] For example, millimeter-wave radar, installed on the ceiling indoors, detects reflected signals from users by emitting radio waves and outputs point cloud data (including position, velocity, and angle). In other words, millimeter-wave radar acts like a "spatial perception nerve." It can penetrate non-metallic obstructions, providing precise millimeter-level micro-motion information, something cameras and pressure sensors cannot do. It can solve problems such as "whether a user is still moving behind curtains."

[0037] For example, thermal imaging cameras are installed on indoor walls to capture the temperature of various parts of a user's body. In other words, a thermal imaging camera acts like a "metabolic sensory nerve." It doesn't rely on visible light, distinguishes between living organisms and objects through temperature distribution, and can sense changes in thermal signals from muscle activity. For instance, the pattern of temperature rise in the relevant muscle groups when "lifting dumbbells" is fundamentally different from that when "holding a water cup."

[0038] For example, a pressure sensor installed on the floor indoors is used to detect changes in floor pressure. In other words, the pressure sensor acts like a "ground contact sensory nerve." It provides the most direct and stable contact and center of gravity information to determine postures with strong interaction with the ground, such as sitting, lying down, or falling.

[0039] For example, at the same moment, when radar detects rapid arm movement, thermal imaging shows no significant change in hand temperature, and pressure sensors show that the feet are standing stably, a multimodal fusion dataset can be obtained.

[0040] Compared to traditional single-modal sensor data acquisition, multi-type sensors can collect user behavior-related data from multiple dimensions such as location, temperature, and pressure. The multi-modal fusion dataset generated in real time can more comprehensively and accurately reflect user behavior characteristics, providing high-quality data support for subsequent accurate identification of user behavior. This solves the problem of low recognition accuracy caused by the inability of single-modal data to fully characterize user behavior.

[0041] S120, Based on the indoor user behavior model, the behavior action verification of the user behavior purpose of the multimodal fusion dataset is performed, and the instantaneous action fit value of the user behavior purpose is determined based on the similarity between the multimodal fusion dataset and the multimodal fusion dataset in the indoor user behavior model.

[0042] In an embodiment of the present invention, when a multimodal fusion dataset is obtained, it can be input into an indoor user behavior model to obtain behavioral action verification for the user's behavioral purpose. Once the behavioral action verification is complete, the similarity between the multimodal fusion dataset and the multimodal fusion dataset in the indoor user behavior model is calculated to determine the instantaneous action fit value for the user's behavioral purpose. Specific implementation details can be found in subsequent embodiments.

[0043] S130, determine the user behavior purpose behavior fit index based on multiple continuously generated instantaneous action fit values, and determine the user's execution user behavior purpose based on the behavior fit index and the preset index.

[0044] In embodiments of the present invention, given the instantaneous action fit value for the user's behavioral purpose, a behavioral fit index for the user's behavioral purpose can be determined based on multiple instantaneous action fit values ​​continuously generated within a preset time period. The behavioral fit index is then compared with a preset index, and the user's intended behavioral purpose is determined based on the comparison result. Specific implementation details can be found in subsequent embodiments.

[0045] S140: Based on the number of user behavior objectives and the relationship tree between the behavior objectives, determine the possible predicted objectives and the user's determined behavior prediction objectives, and start the smart home corresponding to the user behavior with the determined behavior prediction objectives.

[0046] In an embodiment of the present invention, when the user's intended action is determined, possible predicted actions can be determined by the number of intended actions performed and a relationship tree between the actions and the predicted actions. Then, based on a comparison of the number of times possible predicted actions are marked with a preset number, the user's intended action is determined, and the smart home corresponding to the user's action with the determined action is activated. Specific implementation details can be found in subsequent embodiments.

[0047] According to an embodiment of the present invention, a smart home user behavior prediction method acquires one or more multimodal fusion datasets in real time, each multimodal fusion dataset including real-time data from multiple types of sensors at the same time. Based on an indoor user behavior model, the method verifies the user's behavioral purpose using the multimodal fusion datasets, and determines the instantaneous action fit value based on the similarity between the multimodal fusion datasets and the multimodal fusion datasets in the indoor user behavior model. Based on multiple consecutively generated instantaneous action fit values, a user behavior purpose fit index is determined, and based on the fit index and a preset index, the user's intended user behavior is determined. Based on the relationship tree between the number of intended user behavior actions and the intended behavior, possible predicted purposes and the user's confirmed predicted behavior purpose are determined, and the smart home corresponding to the user behavior with the confirmed predicted behavior purpose is activated. This method breaks through the passive response mode of traditional smart homes and achieves accurate prediction of user behavior.

[0048] To enable those skilled in the art to more readily understand the present invention, Figure 2 This is a smart home user behavior prediction method according to a specific embodiment of the present invention, such as... Figure 2 As shown, the smart home user behavior prediction method includes:

[0049] S210 acquires one or more multimodal fusion datasets in real time, each multimodal fusion dataset including real-time data from multiple types of sensors at the same time.

[0050] In embodiments of the present invention, step S210 can be implemented with reference to the implementation of step S110 described above. Further details will not be provided here.

[0051] S220 is used to verify user behavior actions based on an indoor user behavior model and a multimodal fusion dataset.

[0052] In an embodiment of the present invention, when a multimodal fusion dataset is obtained, the multimodal fusion dataset is input into an indoor user behavior model. The indoor user behavior model drives a virtual human to perform a complete behavioral action for the user's behavioral purpose, such as the complete behavioral action of opening a refrigerator.

[0053] The types of user behavior purposes include multiple categories, including but not limited to opening the refrigerator, reading, and leaving home.

[0054] In an embodiment of the present invention, the indoor user behavior model is a three-dimensional indoor environment model constructed using CAD tools, wherein the three-dimensional indoor environment model includes human body, furniture, and sensors.

[0055] In an embodiment of the present invention, the construction process of a three-dimensional indoor environment model includes: converting the indoor space into a three-dimensional geometric model; converting the three-dimensional geometric model into a mathematical model with computable physical equations; creating a human body with controllable movement and conforming to biomechanics in the mathematical model and defining it as standard behavior; and repeatedly adjusting the parameters of the mathematical model through experimental data to determine whether the virtual space is consistent with the real space.

[0056] Transforming an interior space into a three-dimensional geometric model can be achieved in the following ways:

[0057] 1. Scene mapping and data acquisition:

[0058] Laser scanning: Using a 3D laser scanner to acquire point cloud data of a room, accurately capturing the spatial position and geometry of walls, doors, windows, and furniture.

[0059] Import architectural drawings: If you have CAD architectural drawings, import them directly as a reference to ensure the accuracy of dimensions (down to the millimeter level).

[0060] Sensor calibration record: Accurately measure and record parameters such as the three-dimensional coordinates, installation orientation (pitch angle, yaw angle), and field of view (FOV) of each real sensor.

[0061] 2. CAD Modeling and Assembly:

[0062] Structural modeling: In CAD, reconstruct walls, floors, and ceilings based on survey data, and assign thickness and material labels (such as "concrete wall" and "wood floor").

[0063] Furniture and appliance inventory: Create or call 3D models of furniture and appliances (sofas, dining tables, refrigerators, etc.) from the standard library and place them according to the actual measured positions.

[0064] Virtual sensor deployment: Place the "virtual sensor" geometry at the corresponding location in the model. Its position and orientation parameters must be strictly consistent with those of the real sensor.

[0065] 3. Model inspection and repair:

[0066] Water tightness check: Ensure the model is a closed space with no broken surfaces or gaps; otherwise, subsequent meshing and fluid / electromagnetic field calculations will fail.

[0067] Interference check: Confirm that there is no impenetrable geometric overlap between furniture, human movement paths and walls.

[0068] Output: A precise 3D geometry assembly file with material labels (e.g., STEP format).

[0069] The transformation of a three-dimensional geometric model into a mathematical model with computable physical equations can be achieved in the following ways:

[0070] 1. Import into the finite element / multiphysics simulation platform:

[0071] Import the CAD model into professional simulation software.

[0072] Clean up geometry: Simplify unnecessary details (such as decorative patterns) but retain key features.

[0073] 2. Define the material property library:

[0074] Assign physical properties to each material (label) in the model:

[0075] Mechanical properties: density, Young's modulus, Poisson's ratio (used to calculate deformation and vibration).

[0076] Electromagnetic properties: dielectric constant, conductivity, and permeability (used for radar wave simulation).

[0077] Thermal properties: thermal conductivity, specific heat capacity, emissivity (used for thermal imaging simulation).

[0078] 3. Create a multiphysics coupling configuration:

[0079] 1) Select the physics interface:

[0080] Solid mechanical field: Simulates the rigid / flexible motion of the human body and furniture.

[0081] Electromagnetic wave frequency domain or time domain field: Simulates the transmission, propagation and reception of millimeter-wave radar.

[0082] Heat transfer field: Simulates thermal radiation, conduction, and convection.

[0083] 2) Set up inter-field coupling:

[0084] Electromagnetic-structural coupling: When the human body moves, its position and posture changes will affect the electromagnetic scattering boundary in real time.

[0085] Thermo-structural coupling: Muscle activity generates heat, and the heat is conducted to the body surface, affecting thermal radiation.

[0086] One-way / two-way coupling selection: determined based on computing resources and accuracy requirements.

[0087] 4. Mesh generation (spatial discretization)

[0088] Strategy: Divide the entire computational domain (air, objects, human body) into tetrahedral or hexahedral meshes.

[0089] Localized mesh refinement: The mesh is refined near sensors, on the human body surface, and in areas expected to have high electromagnetic wave reflectivity to capture details. The mesh is also refined at joints to accurately simulate flexible deformation.

[0090] Mesh quality check: Ensure that the mesh distortion and aspect ratio are within a reasonable range, otherwise the calculation results will be distorted.

[0091] Output: A simulation project file with physical property definitions, field coupling settings, and high-quality mesh generation completed, ready for solution calculation.

[0092] In this process, creating a controllable, biomechanically compliant human body within the mathematical model and defining it as standard behavior can be achieved in the following ways:

[0093] 1. Human multibody dynamics modeling:

[0094] Rigid body segment division: The human body is divided into 15-21 rigid body segments (head, trunk, upper arm, forearm, thigh, lower leg, etc.).

[0095] Joint definition:

[0096] Joint types: ball joints (shoulder, hip), hinges (elbow, knee), plane joints (spine), etc.

[0097] Joint properties: Define the range of motion, stiffness, and damping coefficient of each joint to simulate physiological limitations and the viscoelasticity of muscles.

[0098] Mass attribute assignment: Based on anthropometry data, accurately assign mass, center of mass, and moment of inertia to each segment.

[0099] 2. Standard behavior and action library programming:

[0100] Behavior definition: For each user behavior purpose to be identified (such as "getting up from the sofa"), define a standard sequence of actions.

[0101] Motion trajectory generation:

[0102] Keyframe definition: Inverse kinematics (IK) is used to define the key poses at the start, middle and end of an action.

[0103] Trajectory interpolation: Spline curve interpolation is used between keyframes to generate a smooth joint angle-time function θ(t).

[0104] Driving function: θ(t) is applied as a displacement boundary condition to the corresponding joints of the human body model to drive its movement.

[0105] 3. Embedding of sensor signal generation model

[0106] Virtual sensor configuration: In the simulation software, a mathematical model is configured for each virtual sensor to match the actual hardware.

[0107] Millimeter-wave radar model: Set up the excitation of the transmitting antenna port and the receiving antenna port, and define the signal processing chain (analog mixing, filtering, FFT to generate point cloud).

[0108] Thermal imaging camera model: configured as an infrared thermal radiation receiver, with its spectral response and spatial resolution set.

[0109] Pressure sensor model: set as a surface force / pressure probe.

[0110] Output definition: Specifies the data type and format (such as point cloud, thermal image matrix, pressure value sequence) that each virtual sensor needs to record during the simulation. It must be consistent with the output format of the real sensor.

[0111] Output: A parameterized, dynamic simulation model that accepts behavioral instructions and outputs a multimodal simulation dataset.

[0112] Among these methods, repeatedly adjusting the parameters of the mathematical model using experimental data to ensure consistency between the virtual and real spaces can be achieved through the following:

[0113] 1. Design calibration experiment:

[0114] Controlled experiment: In a real room, test subjects (of different body types) are asked to perform a series of standardized actions (such as walking, sitting, and picking up objects).

[0115] Synchronous data acquisition:

[0116] Use motion capture systems (such as Vicon) to obtain high-precision ground truth values ​​of human joint movements.

[0117] Simultaneously record the raw data from all real sensors.

[0118] 2. Parametric calibration process:

[0119] Simulation reproduction: In the model, the same motion instructions are used to drive the virtual human body to perform the same standardized actions.

[0120] Error calculation and parameter adjustment:

[0121] First-level calibration (kinematics): Compare the simulated human joint trajectory with motion capture data, and adjust the kinematic parameters of the joint (such as the position of the rotation center).

[0122] Second-level calibration (sensor signal): Compare the simulated sensor output with the actual sensor data, and adjust accordingly.

[0123] Electromagnetic parameters of materials (such as the dielectric constant of human tissue).

[0124] Environmental clutter parameters (such as the reflection coefficient of the wall).

[0125] Sensor noise model parameters (with Gaussian noise levels matched to the actual levels).

[0126] Algorithm-assisted optimization: Using Bayesian optimization or genetic algorithms, automatically search for parameter combinations that minimize the difference between simulation data and real data.

[0127] 3. Model Validation and Performance Evaluation:

[0128] Reserve a validation set: Test using a set of new actions and new tester data that were not involved in calibration.

[0129] Quantitative evaluation: Calculate similarity indices between simulation data and real data at the feature level (such as Wasserstein distance of point cloud distribution, SSIM of thermal images).

[0130] Judgment criteria: The model is considered to have passed validation only when all key indicators exceed the preset threshold (e.g., SSIM>0.85).

[0131] 4. Establish model version management:

[0132] Save the final model parameters after calibration as the "baseline model".

[0133] When the environment changes (such as replacing furniture or adding sensors) or systematic errors are detected, the incremental calibration process is initiated to update the model and create a new version.

[0134] In an embodiment of the present invention, during the process of driving a virtual human to perform a complete action for the purpose of user behavior, the indoor user behavior model also generates simulation data, and then uses the simulation data generated at the same time as a multimodal fusion dataset in the indoor user behavior model.

[0135] In the embodiments of the present invention, the indoor user behavior model takes into account the physical characteristics of sensors and the details of human behavior. Through quantitative calculation (similarity, instantaneous action fit value, behavior fit index), it achieves accurate identification of the user's behavioral purpose, avoiding the problem of large deviation between simulation data and real scene caused by the lack of physical modeling in traditional statistical learning. It can effectively distinguish similar behaviors (such as opening a refrigerator and opening a wardrobe) and significantly improve the accuracy of behavior recognition.

[0136] S230, Based on the similarity between the multimodal fusion dataset and the multimodal fusion dataset in the indoor user behavior model, determine the instantaneous action fit value for the user behavior purpose.

[0137] In embodiments of the present invention, upon obtaining the multimodal fusion dataset in the indoor user behavior model, the multimodal fusion dataset is transformed into a first data vector, and the multimodal fusion dataset in the indoor user behavior model is transformed into a second data vector; based on the similarity formula... Calculate the target similarity between the first data vector and the second data vector, where A represents the first data vector and B represents the second data vector; when there are multiple multimodal fusion datasets and the target similarity is determined to be greater than the preset similarity threshold, determine the average value of the sum of multiple target similarities and use the average value as the average fitting similarity; multiply the number of times the action fits the target by the average fitting similarity as the instantaneous action fitting value for the user behavior purpose.

[0138] Specifically, for each target whose similarity is greater than a preset similarity threshold, the number of action fittings is increased by one, until the target similarity of multiple multimodal fusion datasets is calculated, and the number of action fitting targets is obtained.

[0139] The specific implementation method for transforming the multimodal fusion dataset into the first data vector is as follows:

[0140] 1. Regarding millimeter-wave radar data:

[0141] Point cloud preprocessing: Filter static background point clouds and retain dynamic target (human body) point clouds.

[0142] Key Feature Extraction: Centroid Motion Trajectory: Calculate the displacement, velocity, and acceleration (3D vector) of the centroid of the human point cloud between consecutive frames. Micro-motion Features: Extract the frequency and amplitude of periodic signals related to vital signs such as chest rise and fall, and limb tremors. Posture Contour Features: Project the point cloud onto a 2D plane and calculate its bounding box aspect ratio, principal axis direction, and point cloud density distribution histogram. Joint Angle Estimation: Estimate the approximate angles of large joints (such as elbows and knees) through the spatial relationships of multiple point cloud clusters.

[0143] Output: A fixed-length numerical vector consisting of the above statistics (mean, variance, peak value, etc.).

[0144] 2. Regarding thermal imaging camera data:

[0145] Image preprocessing: background subtraction, segmentation of the human body thermal outline region.

[0146] Key Feature Extraction: Regional Temperature Statistics: Calculate the average temperature, maximum temperature, and temperature gradient of the head, torso, and limbs regions. Thermal Motion History Image: Superimpose temperature changes over a period of time to form a "heat flow" pattern, and extract its direction and intensity features. Abnormal Hot Spot Detection: Detect areas of sudden localized temperature increases (such as a hand touching a hot water cup), and record their location and rate of temperature change.

[0147] Output: Arrange the above features into a hot feature vector.

[0148] 3. Regarding pressure sensor data:

[0149] Signal preprocessing: filtering to remove noise and detecting stress events (such as stepping on something or sitting down).

[0150] Key Feature Extraction: Gait Features: If the sensor is laid on the ground, stride length, gait speed, standing / swinging ratio, and pressure center trajectory can be extracted. Posture Recognition Features: Based on the pressure distribution map, identify whether the posture is "standing on two feet," "standing on one foot," "sitting," or "lying down." Impact Force Features: Record the peak value, duration, and rise slope of pressure changes (used to distinguish between "gentle drop" and "severe fall").

[0151] Output: A mechanical eigenvector.

[0152] Given the feature vectors output from millimeter-wave radar data, thermal imaging camera data, and pressure sensor data, the radar feature vector, thermal feature vector, and pressure feature vector are concatenated end-to-end. The concatenated high-dimensional vector is then subjected to PCA to retain the few principal components that best explain the data variance, forming a more compact and decorrelated final vector. A pre-trained neural network model (autoencoder) is used, and the concatenated vector is input into its encoder part, outputting a low-dimensional embedding vector. By performing attention-weighted fusion on the embedding vector, the first data vector is obtained.

[0153] In embodiments of the present invention, the method of converting the multimodal fusion dataset in the indoor user behavior model into a second data vector can refer to the above-described method of converting the multimodal fusion dataset into a first data vector, and the present invention will not repeat it here.

[0154] S240, determine the user behavior target behavior fit index based on multiple instantaneous action fit values ​​generated continuously.

[0155] In an embodiment of the present invention, multiple instantaneous action fit values ​​continuously generated for user behavior within a preset time period are obtained; based on timestamps, all instantaneous action fit values ​​are sorted according to the order of their generation time, and the absolute difference between adjacent instantaneous action fit values ​​after sorting is calculated to obtain multiple instantaneous action fit fluctuation values; the average of the sum of multiple instantaneous action fit fluctuation values ​​is taken as the average instantaneous action fit fluctuation value; the average of the sum of multiple instantaneous action fit values ​​is taken as the average instantaneous action fit value; based on Calculate the user behavior purpose behavior fit index, where DCA represents the user behavior purpose behavior fit index, ER(s) represents the average instantaneous action fit value, and PE(q) represents the average instantaneous action fit fluctuation value.

[0156] S250 determines the user's intended purpose for performing user behaviors based on the behavior fit index and preset index.

[0157] In an embodiment of the present invention, after calculating the user behavior purpose behavior fit index, it is determined whether the behavior fit index is greater than a preset index; if the behavior fit index is greater than the preset index, the user behavior purpose is taken as the user behavior purpose to be executed.

[0158] Among them, there may be one or multiple purposes for executing user behavior. For example, there may be only one purpose for executing user behavior, such as "leaving home", or two purposes for executing user behavior, such as "watching TV" and "eating snacks".

[0159] S260, determine possible predictable objectives based on the number of user behavior objectives and the relationship tree between the behavior objectives.

[0160] In an embodiment of the present invention, a relation tree of behavioral objectives is obtained. The relation tree includes user behavioral objectives and temporal dependencies of user behavioral objectives. In the temporal dependencies, a first user behavioral objective points to a second user behavioral objective with a directed edge, and the second user behavioral objective is used as a dependent target. Based on the number of user behavioral objectives executed, the dependent target corresponding to the executed user behavioral objective is determined, and the number of dependent targets is determined. If the number of dependent targets is greater than the number of dependencies threshold, the dependent target is used as a possible predicted objective.

[0161] The relationship tree is not an automatically generated association graph through data mining, but a semantic network embedded with prior knowledge of human behavior. It explicitly encodes the logic and temporal sequence between actions. For example: cooking → eating, eating → washing dishes, finding keys & putting on shoes → leaving home. Arrows represent strong temporal dependencies and progressive intent relationships.

[0162] In other words, whenever one or more "user behavior purposes" are confirmed, all next-level behaviors (i.e., dependent purposes) directly pointed to by the current behavior are found from the relationship tree. A dependent purpose (such as "washing dishes") may be pointed to by multiple current behaviors (such as "eating" or "clearing the table"). The number of times it is pointed to is its potential dependency count, which is equivalent to an instant vote. The more votes it receives, the more reasonable its occurrence is based on the current context. Only behaviors that receive more votes (i.e., potential dependency counts) than a certain threshold (i.e., dependency count threshold) can become potential predictable purposes. This filters out those accidental or weakly related options.

[0163] In embodiments of the present invention, the relationship tree of behavior purpose clearly displays the temporal dependencies between behaviors. Through statistical analysis and threshold setting, reasoning from the executed behavior to the predicted behavior is realized, solving the problem that traditional solutions cannot capture the temporal correlation of behaviors, making it difficult to predict subsequent behaviors. This enables forward prediction of user behavior, thereby activating corresponding smart home devices, breaking through the traditional passive response mode, providing proactive services, and improving the user experience. For example, it can prepare kitchen equipment in advance for the user's subsequent cooking.

[0164] S270, Determine the purpose of predicting specific user behaviors.

[0165] In an embodiment of the present invention, the number of times a target with a possible predictable purpose is marked is obtained within a preset time period; if the number of target counts is greater than the preset number, the possible predictable purpose is used as the determined behavior prediction purpose.

[0166] In other words, it is determined by counting the number of times each "potential predictive purpose" has been nominated as a candidate within a recent period T. If the number exceeds a preset number, it indicates that this is not a coincidence, but a frequently occurring pattern recently. Therefore, it is ultimately determined as "determining the predictive purpose of the behavior".

[0167] The preset number of times can be adjusted according to different scenarios. For example, kitchen behaviors are highly repetitive, so the preset number of times can be low, while living room behaviors are highly random, so the preset number of times needs to be high.

[0168] In an embodiment of the present invention, if the number of times the target is determined is not greater than the preset number, the "possible prediction purpose" cannot be determined as the user's definite behavior prediction purpose. If all "possible prediction purposes" cannot be determined as the user's definite behavior prediction purpose, then the user does not have a definite behavior prediction purpose.

[0169] S280, initiates the smart home corresponding to the user behavior for which the behavior prediction purpose is determined.

[0170] In embodiments of the present invention, when the user's intended behavior is determined, the smart home system corresponding to that behavior can be activated. For example, if the intended behavior is "washing dishes," the dishwasher can be started 5 minutes in advance to preheat; if the intended behavior is "leaving home," all windows can be checked to ensure they are closed and the security system can be activated.

[0171] The smart home user behavior prediction method according to embodiments of the present invention fundamentally improves the accuracy and semantic understanding of predictions by comparing real-time multimodal fusion data with data generated from simulated "standard behaviors" in an indoor user behavior model to determine whether the current action is physically reasonable and conforms to the expected behavioral purpose. The indoor user behavior model repeatedly adjusts parameters (material properties, noise models, etc.) through experimental data to ensure the physical consistency between the virtual and real environments, making the comparison benchmark more reliable. By calculating the behavior fit index and utilizing relationship trees, the system moves from instantaneous action recognition to continuous intent inference, smoothing out transient anomalies and making predictions based on behavioral logic (e.g., "cooking" is likely followed by "eating"), achieving high robustness and adaptability to complex scenarios. The determination of the behavior prediction purpose is a high-confidence result verified through multiple rounds (instantaneous fit, temporal stability, logical association, historical frequency). This allows the system to activate relevant devices before the user actually performs the final action. For example, the system not only turns on the lights when it detects a user approaching the door, but also predicts their intention to leave home as soon as they begin preparatory actions such as "looking for keys" and "putting on shoes," and proactively activates security systems and adjusts the air conditioning to energy-saving mode. This closed-loop system covers the entire process of user behavior, from data collection to recognition and prediction services. It effectively solves many problems of traditional smart home systems, such as insufficient behavior recognition accuracy, slow response, lack of physical modeling, and weak temporal correlation analysis. It achieves accurate prediction of user behavior and proactive smart home services, comprehensively improving the intelligence level and practicality of smart home systems.

[0172] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.

[0173] According to one aspect of the present invention, a smart home user behavior prediction system is also proposed. Figure 3 This is a schematic diagram of a smart home user behavior prediction system according to an embodiment of the present invention; as shown below. Figure 3 As shown, it includes:

[0174] The acquisition module 310 is used to acquire one or more multimodal fusion datasets in real time, each of the multimodal fusion datasets including real-time data from multiple types of sensors at the same time.

[0175] The first determining module 320 is used to verify the behavioral actions of the user behavior purpose in the multimodal fusion dataset based on the indoor user behavior model, and to determine the instantaneous action fitting value of the user behavior purpose based on the similarity between the multimodal fusion dataset and the multimodal fusion dataset in the indoor user behavior model.

[0176] The second determining module 330 is used to determine the user behavior purpose behavior fitting index based on the multiple instantaneous action fitting values ​​generated continuously, and to determine the user's execution user behavior purpose based on the behavior fitting index and a preset index.

[0177] The third determining module 340 is used to determine possible predicted purposes and the user's determined predicted behavior purpose based on the number of user behavior purposes and the relationship tree of behavior purposes, and to activate the smart home corresponding to the user behavior with the determined predicted behavior purpose.

[0178] According to an embodiment of the present invention, a smart home user behavior prediction system acquires one or more multimodal fusion datasets in real time. Each multimodal fusion dataset includes real-time data from multiple types of sensors at the same time. Based on an indoor user behavior model, the system verifies the user's behavioral purpose using the multimodal fusion datasets. Based on the similarity between the multimodal fusion datasets and the multimodal fusion datasets in the indoor user behavior model, it determines the instantaneous action fit value for the user's behavioral purpose. Based on multiple consecutively generated instantaneous action fit values, it determines the user's behavioral purpose fit index. Based on the fit index and a preset index, it determines the user's intended behavioral action. Based on the relationship tree between the number of executed user behavioral purposes and the behavioral purposes, it determines possible predicted purposes and the user's confirmed behavioral prediction purpose, and then activates the smart home corresponding to the user's behavior with the confirmed behavioral prediction purpose. This breaks through the passive response mode of traditional smart homes and achieves accurate prediction of user behavior.

[0179] Optionally, the first determining module 320 is specifically used to input the multimodal fusion dataset into the indoor user behavior model to obtain the behavioral action verification of the user behavior purpose; wherein, the indoor user behavior model is a three-dimensional indoor environment model constructed using CAD tools, wherein the three-dimensional indoor environment model includes a human body, furniture, and sensors, and the construction process of the three-dimensional indoor environment model includes: converting the indoor space into a three-dimensional geometric model; converting the three-dimensional geometric model into a mathematical model with computable physical equations; creating a human body with controllable movement and conforming to biomechanics in the mathematical model and defining it as standard behavior; and repeatedly adjusting the parameters of the mathematical model through experimental data to determine whether the virtual space is consistent with the real space.

[0180] Optionally, the first determining module 320 is specifically used to convert the multimodal fusion dataset into a first data vector, and to convert the multimodal fusion dataset in the indoor user behavior model into a second data vector; based on the similarity formula Calculate the target similarity between the first data vector and the second data vector, where A represents the first data vector and B represents the second data vector; when there are multiple multimodal fusion datasets and the target similarity is determined to be greater than a preset similarity threshold, determine the average value of the sum of the multiple target similarities and use the average value as the average fitting similarity; multiply the number of action fitting targets by the average fitting similarity as the instantaneous action fitting value for the user behavior purpose; wherein, for each time the target similarity is determined to be greater than the preset similarity threshold, the number of action fitting targets is increased by one, until the target similarity calculations for multiple multimodal fusion datasets are completed, and the number of action fitting targets is obtained.

[0181] Optionally, the second determining module 330 is specifically used to obtain multiple instantaneous action fitting values ​​continuously generated for the user behavior purpose within a preset time period; based on timestamps, sort all the instantaneous action fitting values ​​according to the order of their generation time, and calculate the absolute difference between adjacent instantaneous action fitting values ​​after sorting to obtain multiple instantaneous action fitting fluctuation values; take the average of the sum of the multiple instantaneous action fitting fluctuation values ​​as the average instantaneous action fitting fluctuation value; take the average of the sum of the multiple instantaneous action fitting values ​​as the average instantaneous action fitting value; based on Calculate the user behavior target behavior fit index, where DCA represents the user behavior target behavior fit index, ER(s) represents the average instantaneous action fit value, and PE(q) represents the average instantaneous action fit fluctuation value.

[0182] Optionally, the second determining module 330 is specifically used to determine whether the behavior fit index is greater than the preset index; if the behavior fit index is greater than the preset index, the user behavior purpose is taken as the execution user behavior purpose.

[0183] Optionally, the third determining module 340 is specifically used to obtain the relation tree of the behavioral purpose, the relation tree including the user behavioral purpose and the temporal dependency relationship of the user behavioral purpose, wherein the first user behavioral purpose in the temporal dependency relationship points to the second user behavioral purpose with a directed edge, and the second user behavioral purpose is used as the dependent target purpose; based on the number of executed user behavioral purposes, the dependent target purpose corresponding to the executed user behavioral purpose is determined, and the number of dependent target purposes can be determined; if it is determined that the number of dependent purposes is greater than the dependency number threshold, the dependent target purpose is used as the possible predicted purpose.

[0184] Optionally, the third determining module 340 is specifically used to obtain the number of times the possible prediction purpose is marked within a preset time period; and if it is determined that the number of times the target is greater than the preset number, the possible prediction purpose is used as the determined behavior prediction purpose.

[0185] According to one aspect of the present invention, an electronic device is provided.

[0186] Figure 4 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Figure 4 As shown, an electronic device may include one or more ( Figure 4 Only one is shown in the image. A processor 102 (which may include, but is not limited to, a microprocessor unit (MPU) or a programmable logic device (PLD)) and a memory 104 for storing data are also shown. In one exemplary embodiment, the electronic device may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 4 The structure shown is for illustrative purposes only and does not limit the structure of the terminal device described above. For example, the terminal device may also include components that are more... Figure 4 The more or fewer components shown, or having the same Figure 4 Equivalent functions or ratios shown Figure 4 The functions shown have more different configurations.

[0187] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the smart home user behavior prediction method in this embodiment of the invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thereby implementing the above-described method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to terminal devices via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0188] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the switching device. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module used for wireless communication with the Internet.

[0189] This invention proposes a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute a smart home user behavior prediction method.

[0190] The applicant of this invention has provided a detailed description of the embodiments of the invention in conjunction with the accompanying drawings. However, those skilled in the art should understand that the above embodiments are merely preferred embodiments of the invention. The detailed description is only intended to help readers better understand the spirit of the invention and is not intended to limit the scope of protection of the invention. On the contrary, any improvements or modifications made based on the inventive spirit of the invention should fall within the scope of protection of the invention.

[0191] Furthermore, the functional units in the various embodiments of the present invention can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0192] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

[0193] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the protection scope of the present invention.

Claims

1. A method for predicting user behavior in smart homes, characterized in that, include: S1, acquire one or more multimodal fusion datasets in real time, each of the multimodal fusion datasets including real-time data from multiple types of sensors at the same time; S2, Based on the indoor user behavior model, the user behavior purpose of the multimodal fusion dataset is verified, and based on the similarity between the multimodal fusion dataset and the multimodal fusion dataset in the indoor user behavior model, the instantaneous action fit value of the user behavior purpose is determined; S3, determine the user behavior purpose behavior fit index based on the multiple instantaneous action fit values ​​generated continuously, and determine the user's execution user behavior purpose based on the behavior fit index and the preset index; S4. Based on the number of user behavior objectives and the relationship tree of behavior objectives, determine the possible predicted objectives and the user's determined behavior prediction objectives, and start the smart home corresponding to the user behavior with the determined behavior prediction objectives.

2. The smart home user behavior prediction method according to claim 1, characterized in that, The multimodal fusion dataset is used to verify user behavior actions based on an indoor user behavior model, including: The multimodal fusion dataset is input into the indoor user behavior model to obtain the behavioral action verification of the user's behavioral purpose; The indoor user behavior model is a three-dimensional indoor environment model constructed using CAD tools. This three-dimensional indoor environment model includes human figures, furniture, and sensors. The construction process of the three-dimensional indoor environment model includes: Transform the interior space into a three-dimensional geometric model; The three-dimensional geometric model is transformed into a mathematical model with computable physical equations; In the mathematical model, a biomechanically compliant human body with controlled movement is created and defined as standard behavior; The parameters of the mathematical model were repeatedly adjusted using experimental data to ensure that the virtual space matched the real space.

3. The smart home user behavior prediction method according to claim 1, characterized in that, Based on the similarity between the multimodal fusion dataset and the multimodal fusion dataset in the indoor user behavior model, the instantaneous action fit value for the user behavior objective is determined, including: The multimodal fusion dataset is transformed into a first data vector, and the multimodal fusion dataset in the indoor user behavior model is transformed into a second data vector; Based on similarity formula Calculate the target similarity between the first data vector and the second data vector, where A represents the first data vector and B represents the second data vector; If there are multiple multimodal fusion datasets and the target similarity is determined to be greater than a preset similarity threshold, the average value of the sum of the multiple target similarities is determined and the average value is used as the average fitting similarity. The product of the number of times the action fits the target and the average fitting similarity is used as the instantaneous action fitting value for the user behavior purpose; Specifically, for each target similarity value determined to be greater than a preset similarity threshold, the number of action fitting attempts is increased by one, until the target similarity calculations for multiple multimodal fusion datasets are completed, thus obtaining the number of action fitting targets.

4. The smart home user behavior prediction method according to claim 1, characterized in that, The user behavior target behavior fitting index is determined based on multiple consecutively generated instantaneous action fitting values, including: Obtain multiple instantaneous action fitting values ​​continuously generated for the user's behavioral purpose within a preset time period; Based on the timestamp, all the instantaneous motion fitting values ​​are sorted in chronological order of their generation, and the absolute difference between adjacent instantaneous motion fitting values ​​after sorting is calculated to obtain multiple instantaneous motion fitting sway values. The average value of the sum of multiple instantaneous motion-fitting sway values ​​is taken as the average instantaneous motion-fitting sway value. The average value of the sum of multiple instantaneous motion fit values ​​is taken as the average instantaneous motion fit value; based on Calculate the user behavior target behavior fit index, where DCA represents the user behavior target behavior fit index, ER(s) represents the average instantaneous action fit value, and PE(q) represents the average instantaneous action fit fluctuation value.

5. The smart home user behavior prediction method according to claim 1, characterized in that, Based on the behavior fit index and the preset index, the user's purpose for performing the user behavior is determined, including: Determine whether the behavior fit index is greater than the preset index; If the behavior fit index is greater than the preset index, the user behavior purpose will be taken as the execution user behavior purpose.

6. The smart home user behavior prediction method according to claim 1, characterized in that, Based on the number of user behavior objectives and the relationship tree of the behavior objectives, possible predicted objectives are determined, including: Obtain the relation tree of the behavioral purpose, the relation tree includes the user behavioral purpose and the temporal dependency relationship of the user behavioral purpose, wherein the first user behavioral purpose in the temporal dependency relationship points to the second user behavioral purpose with a directed edge, and the second user behavioral purpose is used as the dependent target purpose; Based on the number of user behavior objectives, determine the dependent target corresponding to the user behavior objectives, and determine the number of dependent targets that can be relied upon. If the number of dependents is greater than the number of dependents threshold, the dependent destination is taken as the possible predicted destination.

7. The smart home user behavior prediction method according to claim 6, characterized in that, Determining the purpose of predicting the user's specific behavior includes: Based on a preset time period, obtain the number of times the target of the possible prediction purpose is marked; If the number of target occurrences is greater than the preset number, the possible prediction purpose is taken as the determined behavior prediction purpose.

8. A smart home user behavior prediction system, characterized in that, include: The acquisition module is used to acquire one or more multimodal fusion datasets in real time, each of which includes real-time data from multiple types of sensors at the same time. The first determining module is used to verify the behavioral actions of the user behavior purpose in the multimodal fusion dataset based on the indoor user behavior model, and to determine the instantaneous action fitting value of the user behavior purpose based on the similarity between the multimodal fusion dataset and the multimodal fusion dataset in the indoor user behavior model. The second determining module is used to determine the user behavior purpose behavior fitting index based on the multiple instantaneous action fitting values ​​generated continuously, and to determine the user's execution user behavior purpose based on the behavior fitting index and a preset index. The third determining module is used to determine possible predicted purposes and the user's determined predicted behavior purpose based on the number of user behavior purposes and the relationship tree of behavior purposes, and to activate the smart home corresponding to the user behavior with the determined predicted behavior purpose.

9. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 7.