Family multi-user real-time event prediction method and device, event prediction system and computer readable storage medium

By generating user trajectory sequences and identifying users using low-resolution millimeter-wave radar and inertial measurement units, and combining this with graph neural networks for event prediction, the problem of privacy leaks and insufficient data support in home scenarios is solved, achieving high-precision, cross-scenario personalized event prediction.

CN121634083APending Publication Date: 2026-03-10QINGDAO HAIER MULTI MEDIA CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-26
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

In home settings, there are risks of privacy leaks and insufficient data support. Existing technologies are unable to achieve high-precision, cross-scenario, and personalized multi-user real-time event prediction while protecting privacy.

Method used

By combining low-resolution millimeter-wave radar and inertial measurement unit to collect user motion data, generate trajectory sequences and identify users, and use graph neural networks for event prediction, including the generation of fine-grained behavioral events and coarse-grained aggregation, and construct heterogeneous graphs for event modeling.

Benefits of technology

It achieves high-precision, cross-scenario, and personalized real-time event prediction for multiple users in the home while protecting user privacy. It improves the ability to abstract events and the quality of semantic representation, and overcomes the problems of weak transferability and poor personalization adaptation of traditional solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121634083A_ABST
    Figure CN121634083A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of smart families, and discloses a family multi-user real-time event prediction method, which comprises the following steps: acquiring motion data of a user based on a low-resolution millimeter wave radar and an inertial measurement unit deployed in a family environment, and generating a trajectory sequence of the user; generating a fine-grained behavior event sequence for describing user behaviors based on the user identity and the track sequence identified by the motion data; aggregating a plurality of fine-grained behavior events meeting a preset aggregation condition into coarse-grained behavior events in the fine-grained behavior event sequence so as to obtain a coarse-grained behavior event sequence arranged according to a time sequence; and based on the coarse-grained behavior event sequence, predicting an event to occur at the next moment of the user through a graph neural network. According to the scheme, high-real-time event prediction considering privacy protection, accurate abstraction and cross-scene generalization is realized. The invention further discloses a family multi-user real-time event prediction device, an event prediction system and a computer readable storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of smart home technology, and for example to a method, apparatus, event prediction system, and computer-readable storage medium for real-time event prediction for multiple users in a home. Background Technology

[0002] Currently, with the development and popularization of smart home technology, accurate perception and event prediction of user behavior in the home environment has become crucial for improving the level of intelligent services. However, the home environment is highly private, which makes traditional behavior recognition technologies that rely on cameras or large-scale collection of sensitive data such as user videos and images face serious privacy risks and application limitations. At the same time, due to cost and deployment convenience considerations, the low-resolution radar sensors widely used in homes have sparse data features and weak semantic information, making it difficult to directly support the abstraction and reasoning of fine-grained behavioral events. This leads to the core challenges of insufficient data support and difficulty in event abstraction for existing prediction systems in home scenarios.

[0003] To address these issues, various alternative solutions have been disclosed in related technologies. For example, to circumvent privacy concerns, some solutions employ behavior recognition based on wearable sensors, inferring activity by analyzing the wearer's movement patterns. On the other hand, to enhance the understanding of sensor data, other solutions introduce complex multimodal fusion models and transfer learning techniques, attempting to integrate environmental sensor data, such as temperature data, sound data, and partial trajectory information, and utilize models pre-trained on public datasets to adapt to new home environments and uncover behavioral patterns.

[0004] In the process of implementing the embodiments of this disclosure, at least the following problems were found in the related art: While related technologies have reduced reliance on visual privacy data to some extent and attempted to improve the generalization ability of models, significant limitations remain. Wearable device-based methods rely on continuous user wear, resulting in a poor user experience and susceptibility to failure due to forgotten devices or depleted battery. Other multimodal fusion and transfer learning solutions, when processing low-resolution radar data from homes, still struggle to effectively extract strong correlations between the data and specific behavioral events, leading to low accuracy in event abstraction. Furthermore, these models remain poorly adaptable to individualized factors such as differences in home layouts and diverse user habits, making it difficult to achieve high-accuracy real-time event prediction across home scenarios while protecting privacy.

[0005] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this application, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0006] To provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. This summary is not intended as a general commentary, nor is it intended to identify key / important components or describe the scope of protection of these embodiments, but rather as a prelude to the detailed description that follows.

[0007] This disclosure provides a method, apparatus, system, and computer-readable storage medium for real-time event prediction for multiple users in a home, enabling high-precision, cross-scenario, and personalized real-time event prediction for multiple users in a home based on low-resolution radar data while protecting user privacy.

[0008] In some embodiments, the method for real-time event prediction for multiple users in a home includes: collecting motion data of users based on low-resolution millimeter-wave radar and inertial measurement units deployed in the home environment, and generating a trajectory sequence of users; generating a fine-grained behavioral event sequence describing user behavior based on the user identity identified from the motion data and the trajectory sequence; aggregating multiple fine-grained behavioral events that meet preset aggregation conditions into coarse-grained behavioral events in the fine-grained behavioral event sequence to obtain a coarse-grained behavioral event sequence arranged in chronological order; and predicting the events that will occur to the user in the next moment based on the coarse-grained behavioral event sequence using a graph neural network.

[0009] In some embodiments, the device for real-time event prediction for multiple users in a home includes a processor and a memory storing program instructions, the processor being configured to execute the aforementioned method for real-time event prediction for multiple users in a home when the program instructions are executed.

[0010] In some embodiments, the event prediction system includes: a system body; and the aforementioned real-time event prediction device for multiple users in a home, installed on the system body.

[0011] In some embodiments, the computer-readable storage medium stores program instructions that, when executed, cause the computer to perform the aforementioned method for real-time event prediction for multi-user homes.

[0012] The method, apparatus, system, and computer-readable storage medium for real-time event prediction for multi-user homes provided in this disclosure can achieve the following technical effects: This solution integrates low-resolution millimeter-wave radar with an inertial measurement unit (IMU) to effectively construct user trajectory sequences and identify users without collecting their visual privacy data, thus solving the privacy leakage problem inherent in traditional camera solutions. Simultaneously, through fine-grained event generation and a coarse-grained aggregation mechanism based on spatiotemporal behavior consistency, it significantly improves the event abstraction capability and semantic representation quality of low-resolution radar data in complex home scenarios. Furthermore, by combining graph neural networks to model the aggregated event sequences, it effectively captures user behavior patterns and scene contextual relationships, overcoming the weaknesses of traditional solutions in cross-home scenarios and poor personalization adaptation. This achieves high real-time event prediction that balances privacy protection, accurate abstraction, and cross-scenario generalization.

[0013] The above general description and the description below are exemplary and illustrative only and are not intended to limit this application. Attached Figure Description

[0014] One or more embodiments are illustrated by way of example with reference to the accompanying drawings. These illustrations and drawings do not constitute a limitation on the embodiments. Elements having the same reference numerals in the drawings are shown as similar elements. The drawings are not to be scaled. And wherein: Figure 1 This is a schematic diagram of a method for real-time event prediction for multiple users in a home, provided in an embodiment of this disclosure; Figure 2 This is a schematic diagram of a method for generating a user's trajectory sequence provided in an embodiment of this disclosure; Figure 3 This is a schematic diagram of a method for generating fine-grained behavioral event sequences provided in an embodiment of this disclosure; Figure 4 This is a schematic diagram of a method for generating coarse-grained behavioral events provided in an embodiment of this disclosure; Figure 5 This is a schematic diagram of another method for real-time event prediction for multi-user homes provided in this embodiment of the disclosure; Figure 6 This is a schematic diagram of a home multi-user real-time event prediction device provided in an embodiment of this disclosure. Detailed Implementation

[0015] To provide a more detailed understanding of the features and technical content of the embodiments of this disclosure, the implementation of the embodiments of this disclosure will be described in detail below with reference to the accompanying drawings. The accompanying drawings are for illustrative purposes only and are not intended to limit the embodiments of this disclosure. In the following technical description, for ease of explanation, several details are used to provide a full understanding of the disclosed embodiments. However, one or more embodiments may still be implemented without these details. In other cases, well-known structures and devices may be simplified in their depiction to simplify the drawings.

[0016] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this disclosure described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion.

[0017] Unless otherwise stated, the term "multiple" means two or more.

[0018] In this embodiment of the disclosure, the character " / " indicates that the objects before and after it are in an "or" relationship. For example, A / B means: A or B.

[0019] The term "and / or" describes an association between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or A and B.

[0020] The term "correspondence" can refer to an association or binding relationship. The correspondence between A and B means that there is an association or binding relationship between A and B.

[0021] Combination Figure 1 As shown, optionally, embodiments of this disclosure provide a method for real-time event prediction for multi-user homes, including: S11, the event prediction system generates user trajectory sequences by collecting user motion data based on low-resolution millimeter-wave radar and inertial measurement units deployed in the home environment.

[0022] S12, the event prediction system generates a fine-grained sequence of behavioral events describing user behavior based on the user identity and trajectory sequence identified from motion data.

[0023] S13, the event prediction system aggregates multiple fine-grained behavioral events that meet preset aggregation conditions into coarse-grained behavioral events in the fine-grained behavioral event sequence, so as to obtain a coarse-grained behavioral event sequence arranged in chronological order.

[0024] S14, the event prediction system predicts the events that will happen to the user in the next moment based on coarse-grained behavioral event sequences and graph neural networks.

[0025] In this scheme, the event prediction system includes the system body and a device installed on the system body for real-time event prediction for multiple users in a home. The system body includes, but is not limited to, a low-resolution millimeter-wave radar and an inertial measurement unit (IMU) that are interconnected. Specifically, the low-resolution millimeter-wave radar collects data on the distance, movement speed, and spatial angle between the user and the radar at a fixed frequency to determine the user's approximate location. Simultaneously, the IMU collects data on the user's acceleration and angular velocity at a slightly higher frequency to aid in gait analysis. The radar data and IMU data are then aligned using timestamps. The IMU's sampling frequency is resampled using linear interpolation to align with the radar's frame rate, ensuring temporal and spatial matching. Further, based on the aligned data, a spatiotemporal matching algorithm is used for multimodal fusion. The delay between the radar radial velocity time series (after low-pass and normalization processing) and the estimated sequence obtained by integrating the velocity from the IMU at the point where the discrete cross-correlation reaches its maximum value is calculated, and the optimal synchronization interval is selected. Finally, a Kalman filter is used to dynamically correct the data within this interval. The filter introduces a bias component, allowing the radar's position and velocity observations to correct these biases during measurement updates. This ensures that errors in the inertial measurement unit (IMU) do not accumulate over time, and that IMU data compensates for temporary radar signal obstruction or frame loss, ultimately forming a unified and continuous sequence of user motion descriptions, i.e., a trajectory sequence. In this way, through the collaborative and complementary fusion of multimodal sensors, high-quality and continuous trajectory input is provided for subsequent event abstraction, effectively overcoming the inherent defects of low-resolution radar data being sparse and susceptible to interference.

[0026] Furthermore, the event prediction system can identify user identities based on motion data. In one example, the event prediction system can perform identity recognition using a lightweight dual-branch temporal coding network. This network extracts short-term temporal motion templates from radar trajectory data and inertial measurement unit (IMU) gait features, respectively. The radar branch uses a three-layer residual temporal convolutional network to process trajectory data, while the IMU branch uses a hybrid structure consisting of a two-layer one-dimensional convolutional neural network and a single-layer gated recurrent unit to process gait data. Subsequently, the two feature streams are adaptively weighted and fused using a cross-modal multi-head self-attention mechanism. The fused feature stream is then concatenated with a long-term vector generated based on the statistical analysis of activity period preferences over the past seven days. Finally, the system outputs the user identity probability distribution and identity embedding vector through a fully connected layer and a normalized exponential function to complete identity recognition. In this way, after the event prediction system identifies the user's identity information, the user's identity information and trajectory sequence are input together into a spatiotemporal constraint model based on a Kolmogorov-Arnold network. This model has a built-in predefined family spatial topology graph, which uses room layout and connectivity to form a spatial constraint matrix. It also applies a temporal continuity check to the trajectory sequence, requiring that the positional changes between adjacent moments meet the theoretical minimum and maximum reasonable movement time constraints. At the same time, a spatial connectivity check is applied to prevent cross-room jumps without a reasonable path. Finally, based on the continuous and reasonable trajectory sequence output after verification by this spatiotemporal constraint model, the piecewise function approximation capability of the Kolmogorov-Arnold network is used to extract the user's dwell time, movement trajectory pattern, and movement speed characteristics in a specific room. These movement characteristics are combined with room function labels and mapped to a predefined fine-grained behavior label library, thereby generating a fine-grained behavior event sequence with the structure of identity information, spatial location, behavior action, and time interval. In this way, through the deep fusion of identity and trajectory and rigorous spatiotemporal logic verification, the original low-resolution sensor data is transformed into semantically clear and logically consistent behavior events, laying a reliable foundation for subsequent event aggregation and accurate prediction.

[0027] Furthermore, the event prediction system can aggregate multiple fine-grained behavioral events that meet preset aggregation conditions into coarse-grained behavioral events within a fine-grained behavioral event sequence, thus obtaining a coarse-grained behavioral event sequence arranged chronologically. Specifically, the event prediction system can first perform preliminary screening based on rules of temporal proximity and spatial similarity, identifying multiple consecutive events in the fine-grained behavioral event sequence that occur within the same household spatial location and with an interval of less than a set threshold of five minutes. Then, it determines whether the behavioral actions of these events are consistent in the secondary behavioral categories in a predefined event library, i.e., checking whether their tertiary behaviors belong to the same parent behavioral category. For multiple fine-grained events that satisfy the above-mentioned temporal and spatial proximity and behavioral category consistency, their respective time intervals are merged to form a continuous time period, and their behavioral actions are uniformly represented as this common secondary behavioral category, thereby aggregating and generating a coarse-grained behavioral event with complete semantics. This aggregation mechanism, by integrating fragmented user behavior sequences into semantically coherent coarse-grained events, significantly reduces the amount of data and computational complexity required by the subsequent recommendation inference module, providing a structured input foundation for efficient real-time event prediction.

[0028] Furthermore, the event prediction system can convert coarse-grained behavioral events into high-dimensional event semantic vectors by embedding a semantic encoder into a large language model. It then constructs a heterogeneous graph containing user nodes, event nodes, and scene location nodes. The edges in the graph represent the execution relationship between users and events, the spatial association between events and locations, and the temporal relationship between events. The event semantic vector is used as the initial feature of the event nodes, while a long-term preference vector generated based on the user's activity distribution over the past 30 days and a short-term state vector generated based on the behavior sequence of the most recent hour are used as feature inputs for the user nodes. A graph convolutional network is used to learn the representation of this heterogeneous graph. The features of neighboring nodes are weighted and aggregated using an adjacency matrix and a degree matrix containing self-loops to update the embedding representation of all nodes. Based on the updated node embeddings, dual-path analysis is performed, and the representation vectors output from the two paths are concatenated. Then, a learnable parameter matrix and a sigmoid activation function are used to calculate the probability of the user executing each candidate event. Finally, the event with the highest probability is determined as the predicted event for the next time step. This process, through the collaborative reasoning of semantic association and behavioral patterns, effectively improves the accuracy of event prediction and the ability to personalize across family scenarios while protecting user privacy.

[0029] The real-time event prediction method for multi-user homes provided in this disclosure integrates low-resolution millimeter-wave radar and inertial measurement units. This effectively constructs user trajectory sequences and identifies users without collecting their visual privacy data, solving the privacy leakage problem inherent in traditional camera solutions. Simultaneously, through fine-grained event generation and a coarse-grained aggregation mechanism based on spatiotemporal behavior consistency, it significantly improves the event abstraction capability and semantic representation quality of low-resolution radar data in complex home scenarios. Furthermore, by combining graph neural networks to model the aggregated event sequences, it effectively captures user behavior patterns and scene contextual relationships, overcoming the weaknesses of traditional solutions in cross-home scenarios, such as weak transferability and poor personalization. This achieves high-real-time event prediction that balances privacy protection, accurate abstraction, and cross-scenario generalization.

[0030] In practical applications, the event prediction system detected the father's movement from the sofa to the kitchen using a low-resolution millimeter-wave radar deployed in the living room. Simultaneously, the system captured his distinctive, gentle gait using an inertial measurement unit in his wearable device. The system then confirmed the user as the father through an identity recognition module and, combined with his continuous trajectory from the living room to the kitchen, generated a fine-grained event: Father - Kitchen - Fruit Retrieval - 19:05:19:08. Next, the system identified that the father washed fruit in the kitchen between 19:09 and 19:12. Since these two events were less than 5 minutes apart, both occurred in the kitchen, and both belonged to the eating behavior category, the system aggregated them into a coarse-grained event: Father - Kitchen - Food Preparation - 19:05:19:12. Finally, based on this event sequence, the system analyzed the father's evening behavioral habits using a graph neural network, successfully predicting that he was most likely to watch television next, thus achieving accurate and proactive perception of user behavior.

[0031] Combination Figure 2 As shown, optionally, in S11, the event prediction system collects the user's motion data based on a low-resolution millimeter-wave radar and inertial measurement unit deployed in the home environment, and generates the user's trajectory sequence, including: S21, the event prediction system acquires user position and velocity data collected by low-resolution millimeter-wave radar, as well as user acceleration and angular velocity data collected by inertial measurement unit.

[0032] S22, the event prediction system timestamps the position and velocity data, as well as the acceleration and angular velocity data.

[0033] S23, the event prediction system generates the user's trajectory sequence by performing data fusion through a Kalman filter based on the aligned data.

[0034] In this scheme, the user's motion data includes user position and velocity data acquired by a low-resolution millimeter-wave radar, and user acceleration and angular velocity data acquired by an inertial measurement unit (IMU). Thus, the event prediction system can timestamp-align the position and velocity data and the acceleration and angular velocity data after acquiring these data. In one example, the event prediction system achieves timestamp alignment of multi-source motion data by linearly interpolating and resampling the IMU's sampling frequency to match the frame rate of the low-resolution millimeter-wave radar. By unifying the data sampling times of the radar and the IMU, the timing deviation caused by differences in hardware acquisition frequencies is fundamentally eliminated, laying a foundation for accurate fusion of multimodal data and high-quality trajectory generation based on timing consistency.

[0035] Furthermore, the event prediction system uses a Kalman filter to fuse the aligned data to generate the user's trajectory sequence. Specifically, a spatiotemporal matching algorithm is employed to calculate the delay between the radar radial velocity time series (after low-pass and normalization processing) and the estimated sequence obtained from the velocity integration of the inertial measurement unit (IMU) at the point where the discrete cross-correlation reaches its maximum. The system then selects the optimal synchronization interval and applies a Kalman filter for dynamic correction within this interval. This correction process introduces specific bias components into the filter's state vector, enabling the radar to reverse-correct these biases based on the actual position and velocity observations during the measurement update phase. This effectively suppresses error drift caused by IMU integration accumulation. Simultaneously, this fusion mechanism ensures seamless compensation and trajectory repair when radar signals are temporarily lost due to obstacles or hardware limitations. The final output is a spatiotemporally consistent and continuously smooth user motion description sequence, i.e., the trajectory sequence. Thus, through the complementary measurements and collaborative error correction between radar and IMU, the reliability and continuity of trajectory generation based on low-precision sensing devices in the complex home environment are fundamentally improved.

[0036] Optionally, the event prediction system can also perform multimodal trajectory fusion using a particle filter based on the aligned data. Specifically, a large number of particles are constructed to represent the user's possible state space distribution, and the weights of each particle are calculated using position and velocity data observed by radar and motion states inferred by the inertial measurement unit. Through a continuous resampling process, the system approximates the user's optimal trajectory. This approach, through probabilistic approximation, exhibits better adaptability to nonlinear motion and complex noise environments, providing another feasible technical path for the system's implementation in specific scenarios.

[0037] Combination Figure 3As shown, optionally, in S12, the event prediction system generates a fine-grained sequence of behavioral events describing user behavior based on the user identity and trajectory sequence identified from the motion data, including: S31, the event prediction system uses motion data and a time-series coding network to identify the identity of the corresponding user.

[0038] S32, the event prediction system inputs the user's identity information and trajectory sequence into the spatiotemporal constraint model; the spatiotemporal constraint model has a built-in predefined family spatial topology relationship, which is used to verify the temporal continuity and spatial connectivity of the trajectory sequence.

[0039] S33, the event prediction system maps the verified trajectory sequence output by the spatiotemporal constraint model to generate a fine-grained behavioral event sequence.

[0040] The fine-grained behavioral event sequence includes one or more fine-grained behavioral events, each of which is defined by the user's identity information, spatial location, behavioral action, and time interval.

[0041] In this scheme, identity recognition based on motion data using a temporal coding network to obtain the corresponding user's identity information can be achieved through a lightweight dual-branch temporal coding network. Specifically, this network extracts short-term temporal motion templates from millimeter-wave radar trajectory data and inertial measurement unit (IMU) gait features, respectively. The millimeter-wave radar trajectory data consists of position and velocity data, while the IMU gait features consist of acceleration and angular velocity data. The millimeter-wave radar branch uses a three-layer residual temporal convolutional network to process the trajectory data. The convolutional kernel size is set to 3×3, with dilation rates of 1, 2, and 4, and the number of intermediate feature channels is 64, 64, and 128, respectively. Each layer uses the GELU activation function and batch normalization. Simultaneously, the IMU branch uses a hybrid structure consisting of two layers of one-dimensional convolutional neural networks and one layer of gated recurrent units to process the gait data. In this design, both convolutional layers have a kernel size of 3, with 32 and 64 channels respectively. The hidden feature dimension of the GRU layer is set to 64. Then, an adaptive weighted fusion of the two features is performed using a cross-modal multi-head self-attention mechanism. First, the features of the millimeter-wave radar and inertial measurement unit are linearly projected into a 256-dimensional vector, and then split into three groups of vectors: Q, K, and V. Each group is divided into four parallel subspaces according to the rule of 64 dimensions per head. For each subspace, the temporal correlation weights between the gait features of the inertial measurement unit and the trajectory features of the millimeter-wave radar are calculated. A scaled dot product attention mechanism is used to concatenate the attention output features of the four subspaces along the feature dimensions to obtain a preliminary fused feature matrix. Then, a linear integration layer is used for feature optimization, finally outputting the cross-modal fused feature matrix. This fused feature is further concatenated with a long-term statistical vector. Here, the long-term vector is constructed by selecting the 10 most frequent time periods from the individual's activity frequency over the past 7 days to form a time period sequence, which is then converted into a 144-dimensional long-term statistical vector through one-hot encoding. Finally, global average pooling is performed on the cross-modal fused feature matrix to obtain a static fused vector. The static vector is then concatenated with the long-term statistical vector to obtain the final fused feature vector. This fused feature vector is then output through a fully connected layer and a normalized exponential function to complete the user identity probability distribution and identity embedding vector, thus achieving identity recognition. This approach, by fusing dual-modal temporal features of millimeter-wave radar trajectory and inertial measurement unit gait, and combining them with long-term user behavior patterns, achieves high-precision, cross-scenario family member identity recognition while protecting privacy, providing a reliable identity information foundation for subsequent personalized event prediction.

[0042] Furthermore, the user's identity information and trajectory sequence can be input into the spatiotemporal constraint model for verification. Specifically, during verification, this spatiotemporal constraint model incorporates a predefined family spatial topology graph, using room layout and connectivity to construct a spatial constraint matrix. The spatial connectivity verification requires that an element S(i,j)=1 in the spatial constraint matrix indicates that room i and room j are directly connected. The spatial connectivity constraint is: if the user's current room is i and the next room is j, S(i,j)=1 must be satisfied, or there must exist an intermediate room k such that S(i,k)=1 and S(k,j)=1. Jumping between rooms with S(i,j)=0 and no intermediate room is prohibited. The temporal continuity verification requires sorting and deduplicating the input trajectory timestamps to construct a continuous time axis. The temporal continuity constraint is: if the user is in room A at time t1 and in room B at time t2, the theoretical shortest movement time from A to B must be less than or equal to t2-t1, and t2-t1 must be less than or equal to the maximum reasonable movement time, avoiding temporal fragmentation. Furthermore, after obtaining the verified trajectory sequence output by the spatiotemporal constraint model, a fine-grained behavioral event sequence can be generated based on the verified trajectory sequence. In one example, the piecewise function approximation capability of the Kolmogorov-Arnold network can be used to extract the user's dwell time, movement trajectory pattern, and movement speed features in a specific room. These movement features, along with room function labels, are then mapped to a predefined fine-grained behavioral label library, thereby generating a fine-grained behavioral event sequence structured as identity information, spatial location, behavioral action, and time interval. This approach, through multimodal feature fusion and strict spatiotemporal logical constraints, transforms low-resolution sensor data into semantically clear and logically consistent behavioral events, providing a reliable foundation for subsequent accurate prediction. In practical applications, the event prediction system, through the above process, can accurately identify the behavior of an elderly person slowly moving from the bedroom to the bathroom, and generate a fine-grained event sequence of elderly person – bathroom – nighttime toilet use – 02:30~02:32, providing effective data support for health monitoring.

[0043] Combination Figure 4 As shown, optionally, in S13, the event prediction system aggregates multiple fine-grained behavioral events that meet preset aggregation conditions into coarse-grained behavioral events in a fine-grained behavioral event sequence, including: S41, the event prediction system identifies multiple fine-grained behavioral events from the fine-grained behavioral event sequence that have a time interval of less than a preset threshold and occur in the same spatial location.

[0044] S42, when multiple fine-grained behavioral events belong to the same behavioral category, the event prediction system merges the time intervals of multiple fine-grained behavioral events and uniformly represents the behavioral actions as the behavioral category, so as to aggregate and generate coarse-grained behavioral events.

[0045] In this scheme, the event prediction system can aggregate multiple fine-grained behavioral events that meet preset aggregation conditions into coarse-grained behavioral events from a fine-grained behavioral event sequence. Specifically, the event prediction system filters events based on the rules of temporal proximity and spatial similarity, identifying multiple consecutive events in the fine-grained behavioral event sequence that occur within the same household spatial location and whose time interval is less than a set threshold. As an example, the set threshold can be five minutes. Furthermore, the event prediction system can determine whether the behavioral actions of these events are consistent in the secondary behavioral categories in a predefined event library, that is, check whether the parent behavioral category to which their tertiary behaviors belong is the same. Thus, for multiple fine-grained events that simultaneously satisfy spatiotemporal proximity and behavioral category consistency, the event prediction system merges their respective time intervals to form a continuous time period, and uniformly represents their behavioral actions as this common secondary behavioral category, thereby aggregating and generating a coarse-grained behavioral event with complete semantics. This aggregation mechanism, by integrating fragmented user behavior sequences into semantically coherent higher-level events, significantly reduces the amount of data and computational complexity required by the subsequent recommendation inference module, providing a structured input foundation for efficient real-time event prediction.

[0046] In practical applications, the event prediction system identifies a user's mother performing a fine-grained behavior event of preparing breakfast in the kitchen from 07:05 to 07:08, followed by a fine-grained behavior event of heating milk from 07:10 to 07:12. The system first determines that the time interval between these two events is less than the preset 5-minute threshold and that they both occur in the same kitchen space. Then, it confirms through the event database that preparing breakfast and heating milk belong to the same secondary behavior category of food in the behavior action classification. Therefore, the time interval of the two events is merged into 07:05 to 07:12, and the behavior action is uniformly represented as food preparation. Finally, a coarse-grained behavior event is generated, consisting of mother-kitchen-food preparation-07:05~07:12, realizing multi-level event abstraction based on spatiotemporal proximity and behavioral consistency.

[0047] Optionally, the event prediction system determines whether multiple fine-grained behavioral events belong to the same behavioral category by: The event prediction system determines the parent behavior category corresponding to the behavior action of each fine-grained behavior event based on pre-stored event classification rules.

[0048] If each fine-grained behavioral event corresponds to the same parent behavioral category, the event prediction system determines that each fine-grained behavioral event belongs to the same behavioral category.

[0049] In this solution, the event prediction system can determine the parent behavior category corresponding to each fine-grained behavioral event based on pre-stored event classification rules. These rules are built upon a pre-constructed structured event database, which categorizes behavioral actions into multiple levels. For example, the primary indicators include six categories: rest, tidying, diet, entertainment, housework, and care. Each primary category is further subdivided into several secondary behavioral categories, and each secondary category is further defined with specific tertiary behavioral tags. The event prediction system queries this database, matching the specific behavioral actions recorded in each fine-grained behavioral event with the tertiary behavioral tags in the database, and then identifies the corresponding secondary behavioral category as the parent behavior category for that event.

[0050] In this way, after identifying the parent behavior category of each event, the event prediction system compares whether the secondary behavior categories corresponding to the selected multiple fine-grained behavior events are completely identical. If these events have the same secondary behavior category identifier in the event database, the event prediction system determines that they belong to the same behavior category; if the secondary behavior category of any event is different from that of other events, it is determined that they do not belong to the same behavior category. In this way, a structured behavior classification system achieves a standardized understanding of user behavior semantics, ensuring the accuracy and reliability of behavior consistency judgment in the event aggregation process, and providing key support for generating coarse-grained events with complete semantics.

[0051] In practical applications, when the event prediction system detects that a user performs two actions in succession, namely folding clothes and tidying up the desktop, it queries the event database to confirm that both actions belong to the second-level behavior category of tidying up, thus determining them to be the same behavior category and completing event aggregation.

[0052] Combination Figure 5 As shown in S14, the event prediction system, based on a coarse-grained sequence of behavioral events, uses a graph neural network to predict the events that will occur to the user in the next moment, including: S51, the event prediction system performs semantic encoding on events in the coarse-grained behavioral event sequence to obtain the corresponding event semantic vector.

[0053] S52, the event prediction system uses users and events as nodes, and the execution relationship between users and events and the temporal relationship between events as edges to construct a heterogeneous graph.

[0054] S53, the event prediction system uses the event semantic vector as the initial feature of the event node in the graph, and uses the graph neural network to perform representation learning on the heterogeneous graph in order to update the embedding representation of user node and event node.

[0055] S54, the event prediction system predicts the events that will happen to the user in the next moment based on the updated embedded representations of user nodes and event nodes.

[0056] In this scheme, the event prediction system can semantically encode events in a coarse-grained sequence of behavioral events to obtain corresponding event semantic vectors. In one example, this can be achieved by embedding a semantic encoder into a large language model, converting the structured description of events into high-dimensional vector representations. These vectors not only preserve the surface semantics of events but also implicitly capture the contextual relationships and semantic similarities between events.

[0057] Furthermore, the event prediction system constructs a heterogeneous graph by using users and events as nodes and the execution relationship between users and events, as well as the temporal relationship between events, as edges. Here, user nodes represent individual members of a family, event nodes represent coarse-grained behavioral events that have occurred, scene location nodes represent the physical space of the family, and edge relationships include the execution relationship between users and events, the relationship of events occurring in specific locations, and the temporal association between events based on their order of occurrence.

[0058] Furthermore, the event prediction system uses the event semantic vector as the initial feature of the event node in the graph, and uses the long-term preference vector generated based on the user's historical behavior statistics and the short-term state vector generated based on recent activities as the feature input of the user node. It uses a graph convolutional network to learn the representation of the heterogeneous graph, and uses the adjacency matrix and degree matrix containing self-loops to perform weighted aggregation of the features of neighboring nodes. Through multi-layer graph convolutional operations, the embedding representations of user nodes and event nodes are gradually updated, so that the embedding of each node can integrate the correlation information in its local graph structure.

[0059] Furthermore, the event prediction system predicts the events that will occur to the user in the next moment based on the updated node embedding representation, thus achieving accurate and context-aware predictions of the user's future activities. This approach significantly improves the accuracy and personalization of event recommendations in home scenarios.

[0060] In practical applications, the event prediction system first converts the mother's dinner preparation event into a high-dimensional event semantic vector by embedding a semantic encoder using a large language model, while using the mother's long-term preference vector and short-term state vector as feature inputs. Then, it constructs a heterogeneous graph containing the mother's user node, the dinner preparation event node, and the kitchen location node, where edge relationships include the execution relationship between the mother and the dinner preparation event, and the spatial association between the dinner preparation event and the kitchen location. Next, it uses a graph convolutional network to learn the representation of this heterogeneous graph, weighting and aggregating the features of neighboring nodes using adjacency and degree matrices containing self-loops, and updating the user node and event node embeddings that fuse graph structure information. Finally, based on the updated node embeddings, it analyzes the data by calculating semantic similarity and behavioral regularity to predict that the mother will perform the dishwashing event next, thus fully realizing the entire process from semantic encoding, heterogeneous graph construction, graph network representation to event prediction.

[0061] Optionally, the event prediction system predicts the events that will occur to the user in the next moment based on the updated embedded representations of user nodes and event nodes, including: The event prediction system performs dual-path analysis based on the updated embedded representations of user nodes and event nodes.

[0062] The event prediction system calculates the probability of a user executing each candidate event based on the results of dual-path analysis.

[0063] The event prediction system identifies the event with the highest probability as the event that will occur to the user in the next moment.

[0064] The dual-path approach includes a semantic path and a behavioral path. The semantic path is used to capture the semantic similarity between events by calculating the similarity between the embedded representations of event nodes, while the behavioral path is used to model the behavioral patterns of users by analyzing the embedded representations of user nodes and their historical event sequences.

[0065] In this scheme, the event prediction system performs dual-path analysis based on the updated embedding representations of user nodes and event nodes. The semantic path inputs the event semantic vector into a graph convolutional network to capture the semantic relationships between events by calculating the similarity between the event node embedding representations. The behavioral path inputs the user's long-term preference vector and short-term state vector into the graph convolutional network to model user behavior patterns by analyzing the association patterns between the user node embedding representations and their historical event sequences. Further, the event prediction system calculates the probability of the user executing each candidate event based on the dual-path analysis results. Specifically, it concatenates the semantic representation vector output from the semantic path with the behavioral representation vector output from the behavioral path, and then calculates the execution probability distribution for each candidate event using a learnable parameter matrix and a sigmoid growth function. In this way, the event with the highest probability is identified as the event that will occur in the user's next moment. By traversing the execution probabilities of all candidate events and selecting the event with the highest probability value, the final prediction result is output.

[0066] This approach, by simultaneously considering the rationality of event semantics and the regularity of user behavior, achieves personalized event prediction across scenarios while protecting user privacy, significantly improving the accuracy and adaptability of behavior prediction in the home environment.

[0067] In practical applications, the event prediction system analyzes the father's historical behavior sequence in the living room, identifies the semantic relationship between watching TV and reading through semantic path recognition, and combines the behavior path with his long-term evening entertainment preferences and recent activity status to successfully predict that the father will continue to watch TV in the next moment.

[0068] Combination Figure 6As shown, this disclosure provides a home multi-user real-time event prediction device 300, including a processor 301 and a memory 302. Optionally, the device 300 may further include a communication interface 303 and a bus 304. The processor 301, communication interface 303, and memory 302 can communicate with each other via the bus 304. The communication interface 303 can be used for information transmission. The processor 301 can call logical instructions in the memory 302 to execute the home multi-user real-time event prediction method described in the above embodiment.

[0069] Furthermore, the logic instructions in the aforementioned memory 302 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium.

[0070] The memory 302, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods in the embodiments of this disclosure. The processor 301 executes functional applications and data processing by running the program instructions / modules stored in the memory 302, thereby implementing the real-time event prediction method for multi-user homes described in the above embodiments.

[0071] The memory 302 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the terminal device. Furthermore, the memory 302 may include high-speed random access memory and may also include non-volatile memory.

[0072] This disclosure provides an event prediction system, including a system body and the aforementioned real-time event prediction device 300 for multi-user homes. The real-time event prediction device 300 for multi-user homes is installed in the system body. The installation relationship described herein is not limited to placement within the system body, but also includes installation connections with other components of the event prediction system, including but not limited to physical connections, electrical connections, or signal transmission connections. Those skilled in the art will understand that the real-time event prediction device 300 for multi-user homes can be adapted to feasible system bodies to achieve other feasible embodiments.

[0073] This disclosure provides a computer-readable storage medium storing computer-executable instructions configured to perform the above-described method for real-time event prediction for multi-user homes.

[0074] The technical solutions of this disclosure can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes one or more instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in this disclosure. The aforementioned storage medium can be a non-transitory storage medium, such as a USB flash drive, external hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, etc., and other media capable of storing program code.

[0075] The foregoing description and accompanying drawings fully illustrate embodiments of this disclosure to enable those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, procedural, and other changes. The embodiments represent only possible variations. Individual components and functions are optional unless explicitly required, and the order of operation may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the terminology used in this application is for describing embodiments only and is not intended to limit the claims. As used in the description of embodiments and claims, the singular forms “a,” “an,” and “the” are intended to equally include the plural forms unless the context clearly indicates otherwise. Similarly, the term “and / or” as used in this application means including one or more of the associated listed items and all possible combinations thereof. Additionally, when used in this application, the term "comprise" and its variations "comprises" and / or "comprising" refer to the presence of stated features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof. Without further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the process, method, or apparatus that includes said element. In this document, each embodiment may focus on the differences from other embodiments, and similar or identical parts between embodiments can be referred to mutually. For methods, products, etc., disclosed in the embodiments, if they correspond to the method section disclosed in the embodiments, the relevant parts can be referred to the description of the method section.

[0076] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this disclosure. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0077] The methods and products disclosed in the embodiments herein (including but not limited to devices and equipment) can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of units may be merely a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, and the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to implement this embodiment according to actual needs. In addition, the functional units in the embodiments of this disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0078] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than that shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. Each block in a block diagram and / or flowchart, and combinations of blocks in a block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

Claims

1. A method for household multi-user real-time event prediction, characterized in that, The method comprises the following steps: Collecting motion data of a user based on a low-resolution millimeter wave radar and an inertial measurement unit deployed in a home environment, and generating a trajectory sequence of the user; Based on the user identity identified from the motion data and the trajectory sequence, a fine-grained behavior event sequence describing the user's behavior is generated; In the fine-grained behavior event sequence, multiple fine-grained behavior events that meet the preset aggregation conditions are aggregated into coarse-grained behavior events, so as to obtain a coarse-grained behavior event sequence arranged in time sequence; Based on the coarse-grained behavior event sequence, a graph neural network is used to predict the event to be occurred by the user at the next moment.

2. The method of claim 1, wherein, Collecting motion data of a user based on a low-resolution millimeter wave radar and an inertial measurement unit deployed in a home environment, and generating a trajectory sequence of the user, comprising: Obtaining user position and velocity data collected by the low-resolution millimeter wave radar, and user acceleration and angular velocity data collected by the inertial measurement unit; Timestamp alignment is performed on the position and velocity data and the acceleration and angular velocity data; Based on the aligned data, data fusion is performed by a Kalman filter to generate a trajectory sequence of the user.

3. The method of claim 1, wherein, Based on the user identity identified from the motion data and the trajectory sequence, a fine-grained behavior event sequence describing the user's behavior is generated, comprising: Based on the motion data, identity recognition is performed by a time sequence encoding network to obtain identity information of the user; The identity information of the user and the trajectory sequence are input into a space-time constraint model; wherein the space-time constraint model has a pre-defined home space topology relationship built in, which is used to check the time continuity and spatial connectivity of the trajectory sequence; Based on the trajectory sequence output by the space-time constraint model after the check, a fine-grained behavior event sequence is mapped and generated; Wherein, the fine-grained behavior event sequence includes one or more fine-grained behavior events, and each fine-grained behavior event is defined by the identity information, spatial position, behavior action and time interval of the user.

4. The method of claim 1, wherein, In the fine-grained behavior event sequence, multiple fine-grained behavior events that meet the preset aggregation conditions are aggregated into coarse-grained behavior events, comprising: From the fine-grained behavior event sequence, multiple fine-grained behavior events occurring at the same spatial position and having a time interval less than a preset threshold are identified; In the case that the multiple fine-grained behavior events belong to the same behavior category, the time intervals of the multiple fine-grained behavior events are combined, and the behavior actions are uniformly represented as the behavior category, to aggregate and generate a coarse-grained behavior event.

5. The method of claim 4, wherein, Whether the multiple fine-grained behavior events belong to the same behavior category is determined in the following way: According to the pre-stored event classification rules, the upper behavior category corresponding to the behavior action of each fine-grained behavior event is determined; If the upper behavior categories corresponding to each fine-grained behavior event are the same, it is determined that each fine-grained behavior event belongs to the same behavior category.

6. The method of claim 1, wherein, Based on the coarse-grained behavior event sequence, a graph neural network is used to predict the event to be occurred by the user at the next moment, comprising: Semantic encoding is performed on the events in the coarse-grained behavior event sequence to obtain corresponding event semantic vectors; constructing a heterogeneous graph by taking users and events as nodes, and taking execution relationships between users and events and time sequence relationships between events as edges; taking the event semantic vectors as initial features of event nodes in the graph, and performing representation learning on the heterogeneous graph by using a graph neural network to update embedding representations of the user nodes and the event nodes; predicting an event to be occurred by the user at the next time based on the updated embedding representations of the user nodes and the event nodes.

7. The method of claim 6, wherein, The method for predicting a real-time event of a multi-user family, comprises: performing double-path analysis based on the updated embedding representations of the user nodes and the event nodes; calculating probabilities of the user performing each candidate event based on the double-path analysis result; determining an event with the highest probability as the event to be occurred by the user at the next time; wherein the double path includes a semantic path and a behavior path, the semantic path is used to capture semantic similarity between events by calculating similarity between embedding representations of event nodes, and the behavior path is used to model behavior rules of the user by analyzing embedding representations of user nodes and historical event sequences of the user nodes.

8. A device for household multi-user real-time event prediction comprising a processor and a memory having stored program instructions, characterized in that, The processor is configured to execute the method for predicting a real-time event of a multi-user family when running the program instructions.

9. An event prediction system characterized by, comprises: a system body; The device for predicting a real-time event of a multi-user family according to claim 8 is installed in the system body.

10. A computer-readable storage medium storing program instructions, characterized in that, The program instructions are used to make the computer execute the method for predicting a real-time event of a multi-user family when running.