Image recognition-based abnormality recognition method and system for home intelligent agent
By constructing a three-dimensional semantic map and a multi-dimensional behavioral model, and combining depth vision and semantic segmentation technologies, the problem of distinguishing between normal and abnormal behaviors of the elderly in existing technologies has been solved, enabling the home-based intelligent agent to accurately identify and promptly warn of the elderly's behavior.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 陈赐锦
- Filing Date
- 2026-03-06
- Publication Date
- 2026-06-09
AI Technical Summary
Existing smart home care systems for the elderly struggle to distinguish between normal and abnormal behaviors, resulting in low accuracy in early warning systems and an inability to effectively link indoor environment models with the elderly’s personalized home information.
By collecting indoor models and home information, and combining depth vision sensors and semantic segmentation networks to construct a 3D semantic map, we can dynamically track the elderly and build a multi-dimensional behavior model. We can identify abnormal behavior nodes by combining facial expressions and body morphology, and determine the warning body areas and emergency events through multi-level iterative interaction.
It improves the accuracy of abnormal behavior nodes and the accuracy of warning body areas, enabling precise identification and timely response to abnormal behaviors in the elderly.
Smart Images

Figure CN122176380A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image recognition technology, and in particular to an anomaly recognition method and system for home intelligent agents based on image recognition. Background Technology
[0002] Existing smart home care systems for the elderly typically use cameras deployed indoors to collect environmental images, monitor the elderly’s daily activities through image recognition technology, and issue alarms when abnormal situations (such as falls) are detected.
[0003] Existing visual detection methods mostly rely on image analysis from a single perspective or simple skeletal key point extraction, failing to effectively link indoor environment models with the elderly's personalized home information. This makes it difficult for home intelligence agents to distinguish between the elderly's normal behaviors (such as bending over to pick up objects) and abnormal behaviors (such as falling), affecting the accuracy of multiple abnormal behavior nodes and resulting in low accuracy of the elderly's warning body areas. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of the prior art. This invention provides an anomaly identification method and system for home intelligent agents based on image recognition.
[0005] This invention provides an anomaly detection method for home intelligent agents based on image recognition, including: Collect an indoor model of the target location, match the corresponding home intelligent agent based on the indoor model and the corresponding home information, and trigger the home intelligent agent to dynamically track the elderly in the room in order to determine multiple behavioral images of the elderly. The dynamic spatiotemporal events of the elderly are determined by image recognition based on multiple behavioral images, and a multidimensional behavioral model of the elderly is constructed by combining the elderly's facial expressions and body postures. Multiple abnormal behavioral nodes are determined by dynamic monitoring of this multidimensional behavioral model. Based on the tracing of each abnormal behavior node, the corresponding abnormal behavior information is determined. Based on the abnormal behavior information, the elderly person's previous warning information and the corresponding behavior image, a warning review mechanism is determined. The elderly person is then interacted with in different dimensions along the warning review mechanism to determine the combination of the elderly person's response information. The home smart agent determines multiple interaction events of the elderly based on the combination of response information, iterates through multiple levels to trigger each interaction event, marks the warning body area of the elderly, and marks the real-time image corresponding to the warning body area. Based on the image recognition of the real-time image, the corresponding emergency event is determined.
[0006] This invention provides an anomaly detection system for home intelligent agents based on image recognition. This system is applied to the aforementioned anomaly detection method for home intelligent agents based on image recognition. The anomaly detection system for home intelligent agents based on image recognition includes: The behavior image module is used to collect indoor models of the target location, match the corresponding home intelligent agent based on the indoor model and the corresponding home information, and trigger the home intelligent agent to dynamically track the elderly in the room in order to determine multiple behavior images of the elderly. The image recognition module is used to determine the dynamic spatiotemporal events of the elderly based on image recognition of multiple behavioral images, and to construct a multi-dimensional behavioral model of the elderly by combining the elderly's facial expressions and body postures. Based on the dynamic monitoring of this multi-dimensional behavioral model, multiple abnormal behavioral nodes are identified. The early warning review module is used to determine the corresponding abnormal behavior information based on the tracing of each abnormal behavior node. Based on the abnormal behavior information, the elderly's previous early warning information and the corresponding behavior image, the early warning review mechanism is determined, and the elderly are interacted with in different dimensions along the early warning review mechanism to determine the combination of the elderly's response information. The emergency event module is used by the home smart agent to determine multiple interaction events of the elderly based on the combination of response information, and to perform multi-level iterations to trigger each interaction event in order to mark the warning body area of the elderly and mark the real-time image corresponding to the warning body area. The corresponding emergency event is determined based on the image recognition of the real-time image.
[0007] Compared with the prior art, the beneficial effects of the present invention are: (1) Collect an indoor model of the target location, match the corresponding home intelligent agent according to the indoor model and the corresponding home information, and trigger the home intelligent agent to dynamically track the elderly in the room to determine multiple behavioral images of the elderly; determine the dynamic spatiotemporal events of the elderly based on the image recognition of multiple behavioral images, and construct a multidimensional behavioral model of the elderly by combining the elderly's facial expressions and body shapes, determine multiple abnormal behavioral nodes based on the dynamic monitoring of the multidimensional behavioral model, introduce multiple behavioral images, further trigger the image recognition of multiple behavioral images to present the dynamic spatiotemporal events of the elderly, and improve the accuracy of multiple abnormal behavioral nodes.
[0008] (2) Based on the tracing of each abnormal behavior node, the corresponding abnormal behavior information is determined. Based on the abnormal behavior information, the elderly’s previous warning information and the corresponding behavior image, the warning review mechanism is determined. The elderly are interacted with in different dimensions along the warning review mechanism to determine the elderly’s response information combination. The home intelligent agent determines multiple interaction events of the elderly based on the response information combination. The multi-level iteration of triggering each interaction event is performed to mark the elderly’s warning body area and mark the real-time image corresponding to the warning body area. Based on the image recognition of the real-time image, the corresponding emergency event is determined. The warning review mechanism is further controlled. The multi-level iteration of multiple interaction events is performed to improve the accuracy of the elderly’s warning body area and trigger the corresponding emergency event around the real-time image corresponding to the warning body area. Attached Figure Description
[0009] Figure 1 This is a flowchart illustrating the anomaly detection method for a home intelligent agent based on image recognition in an embodiment of the present invention. Figure 2 This is a flowchart illustrating step S11 of the anomaly identification method for home intelligent agents based on image recognition in an embodiment of the present invention. Figure 3 This is a flowchart illustrating step S12 in the anomaly identification method for home intelligent agents based on image recognition in an embodiment of the present invention. Figure 4 This is a flowchart illustrating step S13 of the anomaly identification method for home intelligent agents based on image recognition in an embodiment of the present invention. Figure 5 This is a flowchart illustrating step S14 of the anomaly identification method for home intelligent agents based on image recognition in an embodiment of the present invention. Figure 6 This is a schematic diagram of the structural composition of an anomaly recognition system for home intelligent agents based on image recognition in an embodiment of the present invention. Detailed Implementation
[0010] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0011] Please see Figures 1 to 6 An anomaly detection method for home intelligent agents based on image recognition is applied to image recognition scenarios. The anomaly detection method for home intelligent agents based on image recognition includes: Step S11: Collect an indoor model of the target location, match the corresponding home smart agent based on the indoor model and the corresponding home information, and trigger the home smart agent to dynamically track the elderly person in the room in order to determine multiple behavioral images of the elderly person. Step S12: Based on image recognition of multiple behavioral images, determine the dynamic spatiotemporal events of the elderly, and construct a multidimensional behavioral model of the elderly by combining the elderly's facial expressions and body postures. Based on the dynamic monitoring of this multidimensional behavioral model, determine multiple abnormal behavioral nodes. Step S13: Based on the tracing of each abnormal behavior node, determine the corresponding abnormal behavior information. Based on the abnormal behavior information, the elderly person's previous warning information and the corresponding behavior image, determine the warning review mechanism, and interact with the elderly person in different dimensions along the warning review mechanism to determine the elderly person's response information combination. Step S14: The home smart agent determines multiple interaction events of the elderly based on the combination of response information, performs multi-level iterations to trigger each interaction event, marks the warning body area of the elderly, and marks the real-time image corresponding to the warning body area. Based on the image recognition of the real-time image, the corresponding emergency event is determined.
[0012] refer to Figure 2 In step S11, the specific steps are as follows: S111: Mark the target location and perform real-time 3D semantic reconstruction of the target location based on a depth vision sensor. Combine with a semantic segmentation network to generate an indoor model of the target location. Present the corresponding semantic map in the indoor model. Align the semantic map with the home information of the indoor model to determine the perception strategy of the indoor model and match the corresponding home intelligent agent. S112: Collect multiple target information, determine the elderly person indoors based on the tracking of multiple target information, mark the elderly person with the corresponding dynamic tag, and dynamically track the elderly person indoors, and take pictures of the elderly person from multiple angles, and automatically extract multiple behavioral images containing key action frames according to time slices.
[0013] In the embodiments of this application, the target location is marked, and the target location is reconstructed in real time using a depth vision sensor. An indoor model of the target location is generated by combining a semantic segmentation network. The corresponding semantic map is presented in the indoor model. The semantic map is aligned with the home information of the indoor model to determine the perception strategy of the indoor model and match the corresponding home intelligent agent.
[0014] At this point, the system marks the coordinates of key spatial nodes within the indoor monitoring area (such as the center of the living room, the entrance to the bathroom, and in front of the medicine cabinet) to establish the regions of interest for perception. Utilizing an RGB-D depth vision sensor, it continuously collects spatial point cloud data and texture information of the environment. By applying SLAM (Simultaneous Localization and Mapping), the system calculates the pose changes of the sensors in real time, registers and fuses the discrete point cloud data, and constructs a dense three-dimensional geometric model of the indoor environment. Based on this, a semantic segmentation network is introduced to perform pixel-level classification of two-dimensional image frames, and the semantic labels are back-projected onto three-dimensional spatial voxels, thereby realizing the transformation from a "geometric model" to a "semantic model," enabling the model to understand spatial attributes.
[0015] The system maps the reconstructed 3D semantic model into a topological semantic map, clearly marking the category labels (such as "sofa" and "dining table"), spatial locations, and reachable paths of various objects on the map. It then performs semantic alignment operations at the feature level: the system reads a pre-set home information database, extracts the elderly's health feature vectors (such as medical history and mobility level) and lifestyle habit labels, and binds this unstructured home information to physical nodes in the semantic map through spatial feature fusion. For example, the "medicine cabinet" node is associated with the "medication adherence" feature, and the "wet area of the bathroom" node is associated with the "fall risk level" feature. This alignment mechanism endows the static physical environment with dynamic risk perception attributes.
[0016] Based on the semantic map after feature alignment, the system dynamically generates differentiated perception strategies through a risk weight calculation model. For high-risk associated areas, a high-frame-rate, high-precision dense sampling strategy is configured. For regular activity areas, a low-power sparse monitoring strategy is adopted. According to the computational load requirements, response latency requirements, and specific health profiles of the elderly, the system schedules and matches the optimal logical instance from the service pool, loads the corresponding home-based intelligent agent model (such as a gait anomaly detection model for users with mobility impairments), and completes the virtual-real mapping binding between the intelligent agent and the physical environment.
[0017] Specifically, Mr. Zhang, who lives in an age-friendly apartment in the community, suffers from hypertension and has mild leg problems. The home intelligent system marks the "bathroom door" and "living room medicine cabinet" in Mr. Zhang's apartment as key target locations. The depth vision sensor scans the interior of the apartment in real time, and the system constructs a high-precision three-dimensional geometric grid that includes the sofa, TV cabinet, medicine cabinet and bathroom door frame by processing the depth point cloud data.
[0018] The semantic segmentation network identifies various objects in the image and constructs a semantic map in the generated indoor model. The system then performs key feature alignment operations: when the "medication cabinet" node is identified in the model, the system automatically anchors it to the "hypertension history" feature in Mr. Zhang's file and sets the node as a "key area for medication monitoring"; when the "toilet" node is identified, the system combines Mr. Zhang's "slight leg weakness" movement feature and marks the area as a "high-risk area for falls" and assigns it a high risk weight parameter.
[0019] Based on the alignment results, the system automatically generates a customized perception strategy: for the "medication cabinet" area, the intelligent agent activates a timed inspection mode to monitor whether medication has been missed; for the "bathroom" area, it activates a 24 / 7 high-sensitivity posture tracking mode. The system matches and loads a home intelligent agent instance specifically optimized for "slow movement and balance disorders". This intelligent agent is pre-loaded with abnormal behavior judgment logic for elderly people with hypertension, thus completing the precise adaptation from the physical environment to the digital intelligent agent.
[0020] Furthermore, multiple target information is collected, and the elderly person is identified indoors based on the tracking of multiple target information. The home smart agent marks the elderly person with a corresponding dynamic tag and dynamically tracks the elderly person indoors. It also takes pictures of the elderly person from multiple angles and automatically extracts multiple behavioral images containing key action frames according to time slices. This overall consideration of tracking multiple target information ensures the accuracy of identifying the elderly person indoors.
[0021] At this point, the system continuously collects video stream data using visual sensors, extracts feature vectors of all dynamic targets within the field of view through target detection, and generates a list of multiple target information. In complex indoor environments, there may be multiple sources of interference, such as pet movement, changes in light and shadow, or visitors. The system uses deep learning-based pedestrian re-identification technology to perform spatiotemporal tracing of multiple target information. By comparing pre-stored elderly biometric features (such as facial feature point distribution, body contour features, and gait cycle features), it calculates the similarity score between each target and the preset profile. When the score exceeds the confidence threshold, the system accurately removes interference items from the multiple targets, uniquely locking onto the elderly subject indoors, and establishing the core object for subsequent monitoring.
[0022] After locking onto the target, the home smart agent assigns a unique dynamic identity tag to the elderly person and establishes an independent tracking thread in the background data stream. To achieve comprehensive monitoring, the system adopts an adaptive gimbal control and multi-camera collaborative strategy. Based on the elderly person's real-time location coordinates in the 3D semantic map, the smart agent dynamically adjusts the camera's attitude angle and focal length, driving the sensor to follow the elderly person's movement trajectory for smooth shooting. At the same time, the system constructs a multi-view perception matrix, collecting image data from multiple angles such as front, side, and top views along the elderly person's activity path, effectively solving the problems of limb occlusion and blind spots under a single viewpoint, ensuring the integrity of behavioral data and spatial three-dimensionality.
[0023] To reduce data redundancy and improve real-time processing efficiency, the system abandons the full-video analysis mode and instead adopts a keyframe extraction mechanism based on time slices. The system calculates the changes in optical flow field or the displacement velocity of key points of posture between video frames in real time and sets dynamic threshold trigger logic. When the rate of change of the elderly's movement amplitude (such as sudden changes in joint velocity or body rotation) exceeds the preset threshold, the system automatically marks the time node and performs slice extraction. Within this slice window, the system uses image sharpness evaluation functions (such as the Laplacian gradient operator) to select the frames with the least occlusion and the highest sharpness as key action frames, forming a behavioral image sequence that includes the state before, during and after the action, providing high-quality data input for the subsequent construction of a multi-dimensional behavioral model.
[0024] Specifically, at 2:05 PM, in Mr. Zhang's living room, not only was Mr. Zhang walking around, but his pet dog was also running. The visual sensor captured a complex scene including the person, the pet, and changes in light and shadow outside the window. The home AI system initiated multi-target tracking. By analyzing the gait characteristics and outlines of each target, it found that target A exhibited the characteristics of "slight limping and slow gait," which highly matched the description of "slight leg inconvenience" in Mr. Zhang's file. Target B (the pet dog) was determined to be a non-monitoring target. Based on this, the system eliminated the interference information from the pet dog and accurately locked Mr. Zhang as the only tracking subject.
[0025] The system binds a dynamic tag "Target_Zhang_001" to Mr. Zhang; when Mr. Zhang gets up to go to the bathroom, the home smart device immediately controls the camera pan-tilt to start the follow mode; when Mr. Zhang walks into a narrow area of the corridor, the system automatically adjusts the wide-angle view to cover his whole body; when he stops at the bathroom door, the system quickly adjusts the focus and switches to a side angle to capture his body posture from multiple angles. This multi-angle tracking ensures that even if Mr. Zhang's hands cover his torso, the side view can still completely record his center of gravity shift.
[0026] Mr. Zhang suddenly swayed at the bathroom door and then fell. The system monitored in real time that the vertical displacement velocity of his skeletal key points increased sharply in a very short time, exceeding the threshold for normal walking or sitting. The agent immediately triggered the time slicing mechanism, capturing the video stream from 2 seconds before the fall to 3 seconds after the fall. Within this slice, the system automatically filtered out blurry frames and accurately extracted key frame images containing continuous actions such as "Mr. Zhang holding onto the wall", "knees bending uncontrollably", and "body hitting the ground". These high-value behavioral images were then encapsulated and transmitted to subsequent modules to build a multi-dimensional behavioral model and determine the fall anomaly.
[0027] refer to Figure 3 In step S12, the specific steps are as follows: S121: Map multiple behavioral images to a high-dimensional spatiotemporal coordinate system, analyze the elderly person's position change trajectory, velocity vector and its interaction logic with environmental objects, determine the spatiotemporal framework based on the multi-factor synthesis of the elderly person's position change trajectory and velocity vector, and load the corresponding interaction logic to generate dynamic spatiotemporal events with semantic depth. S122: In the elderly's limb morphology flow, the key points of the elderly's skeleton are extracted using a graph convolutional network to construct a limb motion topology map. Abnormal limb features are determined based on the recognition of this limb motion topology map. In the elderly's facial expression morphology flow, the facial micro-features of the elderly are captured. Abnormal limb features and facial micro-features are spatiotemporally aligned and fused to construct a multidimensional behavioral model of the elderly. S123: Dynamically monitor the multidimensional behavior model and output the behavior deviation degree in real time during the monitoring process. Based on the dynamic response of the behavior deviation degree, mark the corresponding abnormal behavior node. The abnormal behavior node presents abnormal facial expression content and abnormal body content.
[0028] In the embodiments of this application, multiple behavioral images are mapped to a high-dimensional spatiotemporal coordinate system, the elderly person's position change trajectory, velocity vector and their interaction logic with environmental objects are analyzed, the spatiotemporal framework is determined based on the multi-factor synthesis of the elderly person's position change trajectory and velocity vector, and the corresponding interaction logic is loaded to generate dynamic spatiotemporal events with semantic depth. This approach is compatible with the overall consideration of the multi-factor synthesis of the elderly person's position change trajectory and velocity vector, ensuring the accuracy of the spatiotemporal framework.
[0029] At this point, the system constructs a four-dimensional spatiotemporal coordinate system that integrates spatial three-dimensional coordinates and time dimension; the discrete behavior image sequence extracted in step S112 is accurately mapped into this coordinate system based on its collected timestamp and depth information to realize the vectorized representation of behavior data; on this basis, the system uses motion estimation to continuously interpolate and smooth the key skeletal points of the elderly in the sequence to analyze the trajectory of the elderly's position change in the indoor space; at the same time, by combining the displacement difference and time difference between adjacent frames, the system calculates the elderly's motion velocity vector, including physical parameters such as horizontal movement velocity, vertical settling velocity and acceleration, thereby reproducing the complete motion process of the elderly in digital space.
[0030] The system projects the parsed position trajectory and velocity vector into the semantic map generated in step S111, calculates the overlap between the elderly person's spatial coordinates and static objects (such as sofas, walls, and medicine cabinets) and dynamic areas in the environment, and analyzes the interaction logic between the elderly person and environmental objects. Furthermore, the system adopts a multi-factor synthesis method to fuse kinematic parameters (velocity vectors) and topological relationships (position trajectory) to construct a dynamic spatiotemporal framework. This framework not only defines "where" and "how" the elderly person moves, but also defines the physical boundaries of the current action and the expected behavior pattern. For example, it determines whether the current action is in a "passage state" or an "interaction state", thereby defining the spatiotemporal boundaries of the current event.
[0031] Based on the established spatiotemporal framework, the system loads the corresponding interaction logic rules from the knowledge base to provide a deep semantic interpretation of the elderly's behavior. The system combines the original physical motion data with specific scene semantics to generate dynamic spatiotemporal events with clear intentions and attributes. This process realizes the leap from "data flow" to "event flow". The system no longer only sees the movement of coordinate points, but understands socially significant behavioral events such as "taking medicine", "going to the toilet", and "falling", and provides semantic support for subsequent anomaly judgment.
[0032] Specifically, the system maps the continuous behavioral images of Mr. Zhang walking from the corridor to the bathroom onto a four-dimensional spatiotemporal coordinate system. It analyzes that Mr. Zhang's movement trajectory is a smooth curve pointing from the center of the living room to the bathroom door, but the trajectory undergoes an unexpected sharp angle near the door. At the same time, velocity vector analysis shows that Mr. Zhang's horizontal movement speed drops sharply from 0.5 m / s to 0 m / s in a short period of time, and the vertical velocity vector suddenly shows a large negative value (downward sinking), and the acceleration data shows an abnormal abrupt peak.
[0033] The system compared the trajectory and velocity vector with the semantic map. Analysis revealed that Mr. Zhang's abnormal position trajectory highly overlapped with the elevation difference area marked as "toilet threshold" in the semantic map, and the distance between his body's center of gravity and the static object "wall" rapidly decreased to within the collision threshold. Based on the synthesis of multiple factors, the system constructed the current spatiotemporal framework: Mr. Zhang is in a complex physical state of "crossing the threshold - losing balance - leaning against the wall". This framework clearly defines the spatial boundary of the current event as the high-risk area at the toilet entrance.
[0034] Based on the spatiotemporal framework, the system loaded the logical rules of "elderly movement-obstacle interaction". Combined with the physiological characteristics of Mr. Zhang's "leg and foot difficulties", the system determined that the current physical movement was not a normal "passage" behavior, but a "tripping-support" interaction process. The system generated a dynamic spatiotemporal event with semantic depth - "Mr. Zhang tripped at the bathroom door and tried to maintain his balance by holding onto the wall". This event not only recorded the action itself, but also contained semantic information of potential fall risk, which directly triggered the abnormal behavior node judgment in the subsequent step S12.
[0035] Furthermore, in the elderly's limb morphology flow, graph convolutional networks are used to extract the skeletal key points of the elderly to construct a limb motion topology map. Abnormal limb features are determined based on the recognition of this limb motion topology map. In the elderly's facial expression morphology flow, facial micro-features of the elderly are captured. Abnormal limb features and facial micro-features are spatiotemporally aligned and fused to construct a multidimensional behavioral model of the elderly. This approach incorporates the overall considerations of limb motion topology map recognition and ensures the accuracy of abnormal limb features.
[0036] At this point, the system treats the sequence of behavioral images obtained in step S112 as a continuous limb morphology stream and inputs it into a graph convolutional network (GCN). The GCN uses convolution operations in the non-Euclidean domain to accurately locate and extract key skeletal points of various parts of the human body (such as head, neck, shoulder, elbow, wrist, hip, knee, and ankle). Based on this, the system constructs a limb motion topology map according to the human anatomical structure. The nodes in the map represent the coordinates of key points, and the edges represent the skeletal connections. By analyzing the evolution of this topology map over time, the system can quantify the changes in joint angles, the extension range of the limbs, and the relative motion speed. The system compares the extracted real-time topology features with the standard human kinematics model and calculates the deviation of each node. When the motion trajectory of a certain joint violates physiological constraints (such as the knee bending backward) or the motion parameters exceed the normal threshold (such as instantaneous instability of the center of gravity), it is determined and marked as an abnormal limb feature.
[0037] The system initiates the Facial Behavior Coding System (FACS) analysis module within the facial expression flow; it uses face detection to locate facial regions, and then captures the displacement of subtle feature points such as the corners of the eyes, mouth, and brow through a key point localization network (such as FAN); the system focuses on the activation state of facial action units, and tracks the micro-tremors and texture changes of facial muscles through optical flow, thereby identifying facial micro-features with specific physiological indicators such as pain, fear, and confusion, eliminating interference caused by daily expression changes, and extracting high-dimensional emotional feature vectors that can characterize the current physiological and psychological state of the elderly.
[0038] Because limb movements and facial expressions differ in sampling frequency and spatial scale, the system performs spatiotemporal alignment. Through timestamp synchronization and interpolation, it maps low-frequency motion data from the limb morphology stream and high-frequency micro-feature data from the facial expression morphology stream onto a unified time axis. A multimodal feature fusion strategy is adopted to concatenate or weightedly fuse abnormal limb feature vectors with facial micro-feature vectors, constructing a multidimensional behavioral model that includes physiological state, movement patterns, and psychological emotions. This model is no longer a single-dimensional physical observation, but a three-dimensional digital twin model that can comprehensively reflect the elderly's "loss of body control accompanied by painful expressions," providing a holographic perspective for subsequent anomaly detection.
[0039] Specifically, in the limb morphology flow analysis stage, the home-based intelligent agent inputs the sequence of behavioral images recording Mr. Zhang's fall into the graph convolutional network. The system successfully extracted 25 key skeletal points from Mr. Zhang's entire body and constructed a real-time limb motion topology map. The analysis showed that Mr. Zhang's left hip and left knee joints underwent abnormal angle reversal in a very short time, and the Z-axis coordinate (height) of his body's center of gravity showed a sudden drop in a free fall. In addition, the skeletal topology chain of the right upper limb showed a "rapidly extended and attempting to find a support point" shape. The system marked these features that violated the normal walking mechanics logic as abnormal limb features of "center of gravity instability" and "failure of balance compensation".
[0040] During the facial expression flow analysis phase, the system intensively tracked Mr. Zhang's facial area; it captured that the feature points in Mr. Zhang's brow area converged sharply towards the center, the opening and closing of his eyelids increased instantaneously, and the muscles at the corners of his mouth showed tense contraction. These facial micro-feature combinations conformed to the "fright" and "pain" action unit patterns in FACS coding; the system determined that Mr. Zhang's expression at that time was not calm, but accompanied by strong physiological pain and fright signals.
[0041] The system spatiotemporally aligned the two sets of data mentioned above, confirming that the time point of the limb "instability and fall" highly overlapped with the time point of the face "fear and pain". Through feature fusion, a multidimensional behavioral model of Mr. Zhang was constructed: the model not only included the physical fact of "body falling", but also integrated the physiological feedback of "facial pain". Based on this, the system confirmed that Mr. Zhang did not lie down to rest voluntarily, but had a fall event accompanied by physiological pain, thus providing a conclusive multimodal evidence chain for determining abnormal behavior nodes in subsequent steps S123.
[0042] Therefore, the multidimensional behavior model is dynamically monitored, and the behavior deviation is output in real time during the monitoring process. Based on the dynamic response of the behavior deviation, the corresponding abnormal behavior nodes are marked. The abnormal behavior nodes present abnormal facial expressions and abnormal body movements. The introduction of abnormal behavior nodes presents abnormal facial expressions and abnormal body movements. At the same time, multiple behavior images are introduced to further trigger image recognition of multiple behavior images to present the dynamic spatiotemporal events of the elderly, thereby improving the accuracy of multiple abnormal behavior nodes.
[0043] At this point, the system integrates the constructed multidimensional behavior model into the real-time monitoring pipeline, using a sliding time window mechanism to track the elderly person's continuous behavioral state throughout the process. During monitoring, the system inputs the real-time collected behavioral feature vectors into a pre-set normal behavior benchmark model for comparison. This benchmark model is a probability distribution space constructed based on the elderly person's historical health data and standard human kinematic parameters. The system quantifies the behavioral deviation by calculating the Euclidean distance or Mahalanobis distance between the real-time feature vector and the benchmark vector. This deviation is a dynamic scalar that comprehensively reflects the degree of abnormality in limb movement trajectory (such as gait deviation) and the level of abnormality in facial emotional state (such as the pain index). The higher the value, the less the current behavior conforms to the expected pattern.
[0044] The system performs real-time filtering and trend analysis on behavioral deviations to eliminate false alarms caused by environmental noise or slight movement fluctuations. A multi-level response mechanism is established: when the deviation is in a low range, the system maintains routine monitoring; when the deviation curve rises sharply and exceeds the preset warning threshold, the anomaly locking logic is triggered. This dynamic response mechanism ensures that the system can not only capture sudden and severe anomalies (such as falls) but also perceive the gradual accumulation of anomalies (such as the early signs of a sudden stroke), achieving a logical leap from "status monitoring" to "risk warning".
[0045] Once the deviation exceeds the threshold, the system immediately marks the abnormal behavior node on both the timeline and semantic map. This node is not only a timestamp but also a data packet encapsulating complete evidence of the abnormality. The system automatically extracts key data at the moment the warning is triggered and instantiates the abnormal behavior node into structured information containing specific content: on the one hand, it analyzes the abnormal content in the limb dimension (such as the specific joint misalignment angle and the direction of the body falling), and on the other hand, it analyzes the abnormal content in the facial expression dimension (such as the specific facial motion unit activation state and emotion category), thereby forming a visual snapshot of the abnormality and providing accurate contextual basis for subsequent interaction confirmation.
[0046] Specifically, the home AI system tracks Mr. Zhang's multi-dimensional behavioral model in real time. Before the fall, although Mr. Zhang's gait was slow, his behavioral deviation remained in the low range of 5%-10% (which is within the normal fluctuation range for elderly people with mobility issues). However, when Mr. Zhang suddenly lost his balance while crossing the bathroom threshold, the system detected that his center of gravity height dropped by 80 centimeters within 0.5 seconds, and his knee flexion angle exceeded the physiological limit. The system calculated the distance between the current behavioral feature vector and the baseline model in real time, and the behavioral deviation instantly soared to 85%.
[0047] The system sets the fall risk warning threshold to 60%. When the behavior deviation curve is detected to break through the 60% threshold with a near-vertical slope, the dynamic response mechanism is immediately triggered. The system determines that the sudden change in value is not an occasional body sway, but an instability event with clear physical damage. It then freezes the data stream of the current time window and starts the anomaly marking program.
[0048] The system accurately marked the "Node_Fall_1430" anomalous behavior node in the spatiotemporal coordinate system. This node presented detailed anomalous content in two dimensions: in terms of limb anomalous content, it recorded "abnormal twisting of the hip joint angle", "body in a side-lying curled-up posture", and "right upper limb exhibiting a protective support action"; in terms of facial expression anomalous content, it recorded "high contraction of the glabella muscles (frowning)", "eyelid enlargement", and "downturned corners of the mouth" as painful facial features. The successful marking of this node marks the system's completion of a key leap from data perception to anomaly recognition, providing solid judgment basis for the agent to initiate interactive inquiries in the subsequent S13 step.
[0049] refer to Figure 4 In step S13, the specific steps are as follows: S131: The home intelligent agent performs time-series backtracking on abnormal behavior nodes and deeply analyzes multiple behavior images within the previous T seconds. Based on the analysis of multiple behavior images within the previous T seconds, it determines the corresponding abnormal behavior factors and generates abnormal behavior information containing a complete timeline, action chain, and abnormal triggers based on the multi-factor fusion of multiple abnormal behavior factors. S132: Collect the elderly person's previous warning information, match the elderly person's previous warning information with the abnormal behavior information, determine multiple information matching combinations, and construct a warning review mechanism for the elderly person based on the key warning content, corresponding priority and corresponding behavior image of the multiple information matching combinations. S133: In this early warning review mechanism, a humanized interactive verification process is initiated, and multi-dimensional early warning interactions such as facial expression guidance, voice inquiry, and behavior review are triggered along the verification sequence. Corresponding early warning interaction information is output, and the various early warning interaction information is deeply integrated to output the elderly's response information combination and determine whether the elderly have the ability to save themselves or whether it is a misoperation.
[0050] In the embodiments of this application, the home intelligent agent performs time-series backtracking on abnormal behavior nodes and deeply analyzes multiple behavior images within the previous T seconds. Based on the analysis of multiple behavior images within the previous T seconds, the corresponding abnormal behavior factors are determined. Based on the multi-factor fusion of multiple abnormal behavior factors, abnormal behavior information containing a complete timeline, action chain, and abnormal triggers is generated. This approach is compatible with the overall consideration of analyzing multiple behavior images within the previous T seconds, ensuring the accuracy of the corresponding abnormal behavior factors.
[0051] At this point, after marking the abnormal behavior node, the system immediately initiates a time-series backtracking mechanism; a dynamic time window T (usually 5-10 seconds) is set, and the system retrieves and extracts the continuous behavior image sequence within T seconds before the occurrence of the abnormal node from the cache database; using image parsing, each frame of the image within the time window is analyzed pixel by pixel, focusing on changes in the environmental background before the anomaly occurs, details of the elderly person's posture transition, and potential physical interference sources; by comparing the differences between adjacent frames, the system captures those weak precursor signals that may be ignored in real-time monitoring, and reconstructs the entire physical process of the anomaly.
[0052] Based on deep analysis of preceding images, the system uses a causal reasoning engine to extract specific abnormal behavioral factors from the underlying data. These factors are categorized into environmental triggers (such as slippery ground or obstructions), physiological triggers (such as gait disturbances or sudden fainting), and behavioral triggers (such as operational errors or distraction). By calculating the confidence level of the association between each factor and the abnormal outcome, the system eliminates occasional noise and accurately identifies the core causes of the abnormality. For example, it can identify abnormal ground reflection to infer the risk of slipperiness, or identify sudden changes in the elderly person's gait frequency to infer sudden physical discomfort.
[0053] The system performs deep feature-level fusion of multiple extracted abnormal behavior factors; based on the timeline logic, it connects the precursory actions, environmental state changes, and the final abnormal node within the first T seconds to construct a complete chain of evidence; the system generates structured abnormal behavior information, which not only includes a complete record in the time dimension, but also the action chain in the logical dimension and the abnormal cause in the causal dimension. This comprehensive information provides a decision-making basis for the agent to formulate accurate interaction strategies in subsequent steps, solving the core problem of "what caused the abnormality".
[0054] Specifically, after the home AI detected that Mr. Zhang had fallen, it immediately locked the "Node_Fall_1430" node and backtracked the behavioral image cache within T=8 seconds. The system re-analyzed the keyframes within these 8 seconds and found that 6 seconds before the fall, Mr. Zhang's walking gait was still stable, but his gaze was not on the ground; 3 seconds before the fall, Mr. Zhang's right toe hesitated slightly and did not lift enough when stepping into the bathroom threshold; 1 second before the fall, his right toe touched the edge of the threshold and his body center of gravity began to lean forward. These details are fleeting in the real-time stream, but they were accurately captured in the backtracking analysis.
[0055] The system performed an attribution analysis on the above analysis results; the system identified a "threshold height difference" at the bathroom entrance, and the ground had reflective spots due to recent showering (judged as an environmental factor); secondly, combined with Mr. Zhang's "history of hypertension" in his file, the system noticed that he had a slight dragging gait before falling (judged as a physiological factor, possibly due to lower limb weakness caused by blood pressure fluctuations), and determined that his "not looking at the ground" was a behavioral factor; the system's comprehensive calculation concluded that tripping over the threshold was the direct cause, while insufficient lower limb muscle strength was a potential risk factor.
[0056] The system integrates these scattered factors to generate a complete abnormal behavior information report; the report presents a clear logical chain: the timeline shows the entire process from 14:05:02 (approaching the threshold) to 14:05:10 (falling); the action chain is described as "normal walking → toes hitting the threshold → loss of balance → falling"; the abnormal cause is marked as "environmental obstacle (threshold) coupled with lower limb muscle weakness", which clearly indicates that Mr. Zhang did not faint due to a sudden heart attack, but was tripped by the threshold due to leg inconvenience, thus providing key decision support for the agent to decide "ask if there is a fracture" rather than "call for cardiac emergency" in the subsequent step S13.
[0057] Furthermore, the system collects the elderly person's past warning information, matches this information with the abnormal behavior information, and identifies multiple matching combinations. Based on the key warning content, corresponding priority, and corresponding behavioral images of these multiple matching combinations, a warning review mechanism is constructed for the elderly person. This mechanism takes into account the key warning content, corresponding priority, and corresponding behavioral images of multiple matching combinations, ensuring the accuracy of the warning review mechanism for the elderly person.
[0058] At this point, the system retrieves and collects the elderly person's historical early warning records from the distributed database, extracting past anomaly types, frequencies, handling results, and false alarm records. Using semantic similarity calculation and feature vector matching, the system performs a deep comparison between the currently generated abnormal behavior information and historical early warning information. The system searches for the correlation between the two in terms of spatiotemporal patterns, action characteristics, and causal logic, and pairs the current anomaly with similar historical cases to determine multiple information matching combinations. This process aims to use historical experience to help determine the nature of the current anomaly and distinguish between occasional accidents and recurring risks.
[0059] For each identified combination of information, the system extracts key warning information, covering specific anomaly categories (such as falls, palpitations, and stagnation), risk level labels, and associated context. The system introduces a dynamic priority assessment model, combining the elderly person's real-time health status with historical treatment feedback to assign response weights to each matching combination. If the current anomaly is highly similar to a past high-risk event (such as a fall requiring hospitalization), it is assigned extremely high priority. If it matches the characteristics of a past false alarm event (such as bending down to pick up an object and being mistakenly identified as a fall), it is assigned review priority. The priority setting directly determines the urgency and strategy depth of subsequent interactions.
[0060] Based on the key warning content and their priority ranking, the system integrates corresponding behavioral image evidence to dynamically construct a warning review mechanism for the elderly person. This mechanism is a logically rigorous interactive judgment process: for high-priority risks, the mechanism is set to a "rapid confirmation-instant response" mode, with a concise and direct interactive process; for matching combinations with ambiguity or a history of false alarms, the mechanism is set to a "multi-dimensional verification-interference elimination" mode, which confirms the authenticity through multi-angle inquiries and image verification. The final warning review mechanism is essentially an "optimal decision tree" formulated by the intelligent agent based on historical experience when facing uncertainty.
[0061] Specifically, the home smart device quickly retrieved Mr. Zhang's warning records from the past six months. The system found that Mr. Zhang had a record of "slipping in the bathroom" two months ago, and two records of "bending over to look for dropped items" last month. Through information matching, the current abnormal information of "tripping and falling" is highly similar to the "slipping in the bathroom" two months ago in terms of posture characteristics (both involve the body leaning to the side and the limbs touching the ground), forming a high-risk matching combination. At the same time, the "curling up" movement in the current information and the "bending over to look for dropped items" have some overlap in skeletal topology, forming an ambiguous matching combination.
[0062] The system analyzes the above combinations; the key warning content of the high-risk matching combination is "fall resulting in soft tissue contusion", and the historical record shows that family intervention was required at the time, so the system determines its priority as P0 level (highest level); the key warning content of the ambiguous matching combination is "daily bending over activities", and the historical record shows that they are all false alarms, so the priority is determined as P2 level (review level); based on Mr. Zhang's physiological characteristic of "inconvenient legs", the system confirms that the weight of the high-risk combination is much higher than that of the ambiguous combination, but also retains the review process as a precaution.
[0063] Based on the above analysis, the system constructed a hierarchical early warning and review mechanism. For P0-level risks, the mechanism set an instant interaction strategy: locking onto the facial micro-features and supporting movements of Mr. Zhang in pain in real-time images, and generating review logic to "ask if he is injured and whether he can stand." For P2-level ambiguity, the mechanism preset an exclusion logic: if Mr. Zhang responds clearly and his posture returns to normal in the subsequent interaction, the level is downgraded, and the agent generates a specific review execution plan: displaying the skeletal posture diagram of Mr. Zhang at the moment of the fall to confirm that it was not a picking-up action; then initiating voice interaction, focusing on asking "Mr. Zhang, we detected that you fell. Do you have any fractures or severe pain?", thus achieving a precise mapping from historical experience to current decision-making.
[0064] Therefore, in this early warning and review mechanism, a humanized interactive verification process is initiated, and multi-dimensional early warning interactions such as facial expression guidance, voice inquiry, and behavior review are triggered along the verification sequence. Corresponding early warning interaction information is output, and the various early warning interaction information is deeply integrated to output the elderly's response information combination and determine whether the elderly have the ability to save themselves or whether it is a misoperation. This introduces the ability to save themselves or whether it is a misoperation.
[0065] At this point, based on the early warning and review mechanism built by S132, the system officially initiates the humanized interactive verification process. This process is not a mechanical one-way inquiry, but rather a 3D virtual avatar driven by the UE5 engine, which intervenes in a friendly and non-intrusive manner. The system strictly follows the preset verification sequence, triggering multimodal interaction nodes in turn: facial expression guidance, with the virtual avatar attracting the elderly person's attention and calming their emotions through gentle facial animation; then, voice inquiry, initiating targeted inquiries on the core suspicious points of abnormal behavior; and finally, behavioral review, requiring the elderly person to perform specific instructions to verify bodily functions. The system records the input and feedback of each interaction node in real time and outputs early warning interaction information containing the elderly person's real-time status data.
[0066] The system inputs the collected multi-dimensional early warning interaction information into a multimodal fusion network; using the attention mechanism, it performs feature-level weighted fusion of facial expression feedback (such as pain, calmness), voice response (such as clear, vague, silent), and behavioral review results (such as action completion and limb coordination); the system constructs a comprehensive response vector to eliminate the ambiguity that may exist in a single modality (such as an elderly person verbally replying "I'm fine" but with a facial expression of pain), thereby generating a combination of response information that can truly reflect the elderly person's current physiological and psychological state. This combination not only includes explicit answer content but also implicit health indicators.
[0067] Based on the deeply fused response information, the system calls the decision classification model for final judgment. By comparing the elderly person's actual response status with the preset health / abnormal template, the system executes a two-way judgment logic: on the one hand, if the elderly person responds to the instructions quickly, performs complete limb movements, and shows no signs of pain, the system marks it as "misoperation" (such as bending over to pick up an object being mistakenly judged as a fall); on the other hand, if the elderly person shows delayed response, speech interruption, limb movement impairment, or persistent pain expression, the system calculates their autonomous action index to determine whether they have the ability to save themselves. This judgment result will directly determine whether to initiate the emergency contact procedure.
[0068] Specifically, after the home smart system detected that Mr. Zhang had fallen, it immediately initiated a follow-up mechanism. The virtual avatar on the living room screen quickly switched to "concerned" mode, guiding Mr. Zhang's expression with gentle eyes to alleviate his fear after the fall. The system then triggered a voice inquiry: "Mr. Zhang, we detected that you have fallen. Are you in any severe pain? Can you stand up by yourself?" The system performed a behavioral review and issued the instruction: "Please try raising your right hand to indicate this." During this process, the system recorded Mr. Zhang's eye focus, voice response, and limb tremors when he tried to raise his hand, generating the original warning interaction information.
[0069] The system performs in-depth analysis of the collected information. Although Mr. Zhang verbally replied, "It's nothing, I just bumped into something," the speech recognition module detected obvious tremors in his voiceprint. At the same time, facial expression analysis showed a persistent "painful micro-expression" between his eyebrows. More importantly, behavioral re-examination data showed that when Mr. Zhang tried to raise his hand, his right shoulder joint movement was restricted, and he could not lift his center of gravity on his own. The system integrated these contradictory pieces of information, eliminated the masking elements in his verbal reply, and extracted the core feature vector of "voice tremor + facial pain + restricted limb movement," generating an accurate combination of response information.
[0070] Based on the fused response information, the system determined that although Mr. Zhang was conscious, he was unable to stand independently due to limb pain and leg weakness, and his autonomous action index was below the safety threshold. The system ruled out the possibility of "misoperation" and marked his status as "conscious but unable to save himself". This determination directly triggered the emergency event marking in the subsequent step S14. The system prepared to send a distress signal containing the tag "suspected fracture from fall" to the emergency contact.
[0071] refer to Figure 5 In step S14, the specific steps are as follows: S141: The home intelligent agent obtains the response information combination, performs context parsing on the response information combination to output multiple sub-interaction information, determines the elderly's interaction events based on the combination of multiple sub-interaction information and corresponding abnormal behavior information, marks multiple interaction events, and triggers the corresponding multi-level iterative reasoning mechanism along multiple interaction events. S142: In the multi-level iterative reasoning mechanism, each interactive event undergoes multi-level iteration and multi-level iteration along different event dimensions. The corresponding abnormal state is marked in different event dimensions. The warning body area of the elderly is determined based on the state content of each abnormal state, the corresponding iteration level, and background interference. The warning body area is dynamically tracked to mark the real-time image corresponding to the warning body area, thereby realizing the dynamic capture of injury features. The corresponding emergency event is determined based on the mapping relationship between multiple injury features and emergency events.
[0072] In the embodiments of this application, the home intelligent agent obtains the response information combination, performs context parsing on the response information combination to output multiple sub-interaction information, determines the elderly's interaction events based on the combination of multiple sub-interaction information and corresponding abnormal behavior information, marks multiple interaction events, and triggers the corresponding multi-level iterative reasoning mechanism along multiple interaction events. This approach takes into account the overall consideration of the combination of multiple sub-interaction information and corresponding abnormal behavior information, ensuring the accuracy of the elderly's interaction events.
[0073] At this point, the system obtains the response information combination output in step S133. This combination typically contains unstructured multimodal data (such as speech streams, facial image sequences, and pose coordinate sequences). Using natural language understanding (NLU) and behavioral semantic analysis, the system performs fine-grained context parsing on this combination. The parsing process aims to decompose the mixed information into independent semantic units, extracting key entity words (such as "leg pain" and "dizziness"), action descriptions (such as "failed to stand up"), and emotion labels (such as "pain"). Based on this, the system outputs multiple structured sub-interaction information, each sub-information corresponding to a specific aspect of the elderly person's feedback during the interaction, forming a discretized description of the elderly person's state.
[0074] The system performs semantic mapping and correlation analysis on the multiple sub-interaction information obtained from parsing and the abnormal behavior information (including triggers and action chains) generated in step S131. By combining "historical triggers" with "current feedback", the system can define and label specific interaction events. Each interaction event represents a clear judgment proposition. For example, "limb restriction feedback" and "fall action" are combined and labeled as "sports injury event", and "vague language feedback" and "history of hypertension" are combined and labeled as "sudden pathological event". The system classifies and labels the interaction events to establish logical anchors for subsequent reasoning.
[0075] For multiple marked interactive events, the system constructs and triggers a multi-level iterative reasoning mechanism. This mechanism simulates the logical chain of medical diagnosis and adopts a progressive analysis strategy: the initial iteration verifies the authenticity of the event and eliminates environmental noise or misoperation; the intermediate iteration assesses the severity and scope of the event; the advanced iteration predicts the evolution trend and potential risks of the event; the system activates the corresponding reasoning logic in sequence according to the priority of the events, forming a dynamic analysis loop, which provides logical support for the subsequent accurate location of the body's early warning area.
[0076] Specifically, the home AI agent acquired a combination of Mr. Zhang's response information. The system used a context parsing engine to break down the mixed data stream into three key sub-interaction information: first, voice sub-information, extracting the keywords "leg pain" and "unable to move," representing the main complaint of pain; second, posture sub-information, parsing the movement features of "effective support of the right upper limb and no displacement of the lower limb," representing lower limb motor dysfunction; and third, facial expression sub-information, extracting "painful facial micro-features," representing the physiological stress state.
[0077] The system logically combines the above sub-interaction information with the "threshold tripping" abnormal behavior information determined in S131; the system spatially matches the "posture sub-information (no displacement of lower limbs)" with the "tripping action chain (right foot lands first)" and marks it as "right lower limb movement restriction interaction event"; it combines the "voice sub-information (leg pain)" with the "fall impact force data" and marks it as "pain feedback confirmation event"; at the same time, combined with his history of hypertension, it combines the "facial expression sub-information (pain)" with the fall stress response and marks it as "blood pressure fluctuation risk monitoring event". These three events together constitute a three-dimensional description of Mr. Zhang's current state.
[0078] The system triggers multi-level iterative reasoning along these three interaction events. For the "pain feedback confirmation event," the system initiates a primary iteration, comparing Mr. Zhang's current pain expression with his historical pain tolerance threshold to confirm that the pain level is high. For the "right lower limb movement restriction interaction event," the system performs an intermediate iteration, combining the skeletal model to analyze the mechanical reasons for his inability to stand up, inferring that there may be a fracture or severe soft tissue contusion. For the "blood pressure fluctuation risk monitoring event," the system performs a high-level iteration in parallel, predicting the risk of a sharp increase in blood pressure that may be induced by the fall pain. Through this series of iterative reasoning, the system eliminates the simple "tripping and resting" scenario and locks in the emergency situation of "traumatic injury accompanied by potential pathological risks."
[0079] Furthermore, in the multi-level iterative reasoning mechanism, each interactive event undergoes multi-level iteration along different event dimensions, marking corresponding abnormal states in different event dimensions. Based on the state content of each abnormal state, the corresponding iteration level, and background interference, the warning body area of the elderly is determined. The warning body area is dynamically tracked to mark the real-time image corresponding to the warning body area, realizing the dynamic capture of injury features. Based on the mapping relationship between multiple injury features and emergency events, the corresponding emergency event is determined, which takes into account the overall consideration of the mapping relationship between multiple injury features and emergency events, ensuring the accuracy of the corresponding emergency event. At the same time, the warning review mechanism is further controlled, and multi-level iteration is performed for multiple interactive events to improve the accuracy of the warning body area of the elderly, and the corresponding emergency event is triggered around the real-time image corresponding to the warning body area.
[0080] At this point, within the multi-level iterative reasoning mechanism, the system conducts deep iterations along different event dimensions for each interactive event (such as movement restriction events and pain feedback events) marked by S141. During each iteration, the system utilizes a spatiotemporal consistency verification method to eliminate background interference (such as visual noise caused by changes in ambient lighting and clothing wrinkles), accurately extracting effective features related to the event. Based on this, the system marks the corresponding abnormal states in different dimensions: in the physiological dimension, it marks surface features such as "redness and swelling" and "bruising"; in the movement dimension, it marks functional features such as "stiffness" and "restriction"; and in the emotional dimension, it marks psychological features such as "pain" and "anxiety." These abnormal states together constitute the basic data layer for locating the injury.
[0081] The system integrates abnormal state content from various dimensions, combines iterative levels (prioritizing high-level fatal or disabling risks) and background interference analysis, and uses spatial positioning to calculate the geometric center of abnormal features, thereby accurately determining the elderly person's warning body area (such as the right calf or left hip). Once the area is locked, the system immediately starts a dynamic tracking mode, using target tracking methods (such as KCF or correlation filtering) to continuously lock the warning area in the video stream. Even if the elderly person moves due to pain, the system can still adaptively adjust the tracking box to ensure that the warning area is always within the core field of vision of visual perception, providing a stable image source for subsequent analysis.
[0082] Based on dynamic tracking, the system performs high-precision feature engineering on real-time images to extract injury features such as color gradient (identifying redness and swelling), texture changes (identifying abrasions), and contour anomalies (identifying deformations), enabling dynamic capture of the injury's development. The system constructs a mapping model of "injury features - emergency events" based on a pre-set medical emergency knowledge graph. Multiple extracted injury feature vectors are input into the model for matching, and the confidence level of the event is calculated to determine the specific type of emergency event (such as "suspected fracture" or "open wound"), thus completing the logical closed loop from phenomenon observation to event characterization.
[0083] Specifically, the home smart agent initiated a deep iteration for the "right lower limb movement restriction interaction event"; in the movement dimension iteration, the system captured that Mr. Zhang's right leg was in an unnatural stiff state and did not have any voluntary retraction action after contact with the ground, which was marked as an abnormal state of "loss of motor function"; in the body surface dimension iteration, the system filtered out the background interference caused by the reflection of the ground tiles and extracted the abnormal color gradient change of the skin on the front of the right tibia, which was marked as an abnormal state of "suspected skin lesion redness and swelling".
[0084] The system integrates the abnormal states of "loss of function" and "skin lesions and swelling" with the iteration level judgment, and accurately locks the warning body area to the "right tibia region"; the system controls the camera pan-tilt to focus on this area and starts dynamic tracking mode; when Mr. Zhang tries to adjust his sitting posture due to pain and his body twists slightly, the intelligent body vision system corrects the tracking coordinates in real time to ensure that the right tibia is always in the central ROI (region of interest) of the image, and continuously outputs a stable real-time image stream.
[0085] The system performs detailed analysis of real-time images and dynamically captures two key injury features: first, the skin on the anterior side of the tibia shows linear abrasions with slight bleeding (characteristic of open injury); second, the right ankle joint shows an abnormal eversion angle (characteristic of skeletal deformity). The system inputs these two features into a mapping model and compares them with the emergency event database. The feature vector highly matches "lower limb fracture and soft tissue injury caused by fall", with a confidence level of 92%. Based on this, the system officially determines the current emergency event as "suspected fracture of the right lower leg with soft tissue injury" and immediately triggers the subsequent emergency contact procedure S14.
[0086] Please see Figure 6 , Figure 6 This is a schematic diagram of the structural composition of an anomaly detection system for a home intelligent agent based on image recognition, as described in this embodiment of the invention. The anomaly detection system for the home intelligent agent based on image recognition is applied to the aforementioned anomaly detection method for a home intelligent agent based on image recognition. The anomaly detection system for the home intelligent agent based on image recognition includes: The behavior image module 21 is used to collect an indoor model of the target location, match the corresponding home intelligent agent based on the indoor model and the corresponding home information, and trigger the home intelligent agent to dynamically track the elderly in the room in order to determine multiple behavior images of the elderly. The image recognition module 22 is used to determine the dynamic spatiotemporal events of the elderly based on image recognition of multiple behavioral images, and to construct a multi-dimensional behavioral model of the elderly by combining the elderly's facial expressions and body postures, and to determine multiple abnormal behavioral nodes based on the dynamic monitoring of the multi-dimensional behavioral model. The early warning review module 23 is used to determine the corresponding abnormal behavior information based on the tracing of each abnormal behavior node, determine the early warning review mechanism based on the abnormal behavior information, the elderly’s previous early warning information and the corresponding behavior image, and interact with the elderly in different dimensions along the early warning review mechanism to determine the elderly’s response information combination. The emergency event module 24 is used by the home smart agent to determine multiple interaction events of the elderly based on the combination of response information, to perform multi-level iterations to trigger each interaction event, to mark the warning body area of the elderly, and to mark the real-time image corresponding to the warning body area, and to determine the corresponding emergency event based on the image recognition of the real-time image.
[0087] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
Claims
1. An anomaly detection method for home intelligent agents based on image recognition, characterized in that, include: Collect an indoor model of the target location, match the corresponding home intelligent agent based on the indoor model and the corresponding home information, and trigger the home intelligent agent to dynamically track the elderly in the room in order to determine multiple behavioral images of the elderly. The dynamic spatiotemporal events of the elderly are determined by image recognition based on multiple behavioral images, and a multidimensional behavioral model of the elderly is constructed by combining the elderly's facial expressions and body postures. Multiple abnormal behavioral nodes are determined by dynamic monitoring of this multidimensional behavioral model. Based on the tracing of each abnormal behavior node, the corresponding abnormal behavior information is determined. Based on the abnormal behavior information, the elderly person's previous warning information and the corresponding behavior image, a warning review mechanism is determined. The elderly person is then interacted with in different dimensions along the warning review mechanism to determine the combination of the elderly person's response information. The home smart agent determines multiple interaction events of the elderly based on the combination of response information, iterates through multiple levels to trigger each interaction event, marks the warning body area of the elderly, and marks the real-time image corresponding to the warning body area. Based on the image recognition of the real-time image, the corresponding emergency event is determined.
2. The anomaly identification method for home intelligent agents based on image recognition according to claim 1, characterized in that, The indoor model of the target location is collected, and a corresponding home-based intelligent agent is matched based on the indoor model and the corresponding home information. The home-based intelligent agent is then triggered to dynamically track the elderly person indoors to determine multiple behavioral images of the elderly person, including: The target location is marked, and a real-time 3D semantic reconstruction of the target location is performed based on a depth vision sensor. An indoor model of the target location is generated by combining a semantic segmentation network. The corresponding semantic map is presented in the indoor model. The semantic map is aligned with the home information of the indoor model to determine the perception strategy of the indoor model and match the corresponding home intelligent agent.
3. The anomaly identification method for home intelligent agents based on image recognition according to claim 2, characterized in that, The method of collecting an indoor model of the target location, matching the corresponding home intelligent agent based on the indoor model and corresponding home information, and triggering the home intelligent agent to dynamically track the elderly person indoors to determine multiple behavioral images of the elderly person, also includes: Collect information from multiple targets, identify elderly people indoors based on the tracking of multiple target information, mark the elderly person with a corresponding dynamic tag, dynamically track the elderly person indoors, and take pictures of the elderly person from multiple angles. Automatically extract multiple behavioral images containing key action frames according to time slices.
4. The anomaly identification method for home intelligent agents based on image recognition according to claim 1, characterized in that, The method involves determining the elderly person's dynamic spatiotemporal events based on image recognition of multiple behavioral images, and constructing a multidimensional behavioral model of the elderly person by combining their facial expressions and body postures. Based on the dynamic monitoring of this multidimensional behavioral model, multiple abnormal behavioral nodes are identified, including: Multiple behavioral images are mapped to a high-dimensional spatiotemporal coordinate system. The trajectory of the elderly person's position change, velocity vector and interaction logic with environmental objects are analyzed. The spatiotemporal framework is determined based on the multi-factor synthesis of the elderly person's position change trajectory and velocity vector, and the corresponding interaction logic is loaded to generate dynamic spatiotemporal events with semantic depth.
5. The anomaly identification method for home intelligent agents based on image recognition according to claim 4, characterized in that, The method of determining the elderly person's dynamic spatiotemporal events based on image recognition of multiple behavioral images, constructing a multidimensional behavioral model of the elderly person by combining the elderly person's facial expressions and body postures, and determining multiple abnormal behavioral nodes based on dynamic monitoring of this multidimensional behavioral model, also includes: In the elderly's limb morphology flow, graph convolutional networks are used to extract the skeletal key points of the elderly to construct a limb motion topology map. Abnormal limb features are identified based on the recognition of this limb motion topology map. In the elderly's facial expression morphology flow, facial micro-features of the elderly are captured. Abnormal limb features and facial micro-features are spatiotemporally aligned and fused to construct a multidimensional behavioral model of the elderly. The multidimensional behavior model is dynamically monitored, and the behavior deviation is output in real time during the monitoring process. Based on the dynamic response of the behavior deviation, the corresponding abnormal behavior nodes are marked. The abnormal behavior nodes present abnormal facial expressions and abnormal body language.
6. The anomaly identification method for home intelligent agents based on image recognition according to claim 1, characterized in that, The process involves tracing each abnormal behavior node to determine the corresponding abnormal behavior information, establishing a warning review mechanism based on this abnormal behavior information, the elderly person's previous warning information, and the corresponding behavior image, and then interacting with the elderly person in different dimensions along this warning review mechanism to determine the elderly person's response information combination, including: The home-based intelligent agent performs time-series backtracking on abnormal behavior nodes and deeply analyzes multiple behavior images within the previous T seconds. Based on the analysis of multiple behavior images within the previous T seconds, it determines the corresponding abnormal behavior factors. Based on the multi-factor fusion of multiple abnormal behavior factors, it generates abnormal behavior information containing a complete timeline, action chain, and abnormal triggers.
7. The anomaly identification method for home intelligent agents based on image recognition according to claim 6, characterized in that, The process of determining corresponding abnormal behavior information based on the tracing of each abnormal behavior node, establishing an early warning review mechanism based on this abnormal behavior information, the elderly person's previous warning information, and corresponding behavior images, and then interacting with the elderly person in different dimensions along this early warning review mechanism to determine the elderly person's response information combination, also includes: Collect the elderly person's past warning information, match the elderly person's past warning information with the abnormal behavior information, and determine multiple information matching combinations. Based on the key warning content, corresponding priority and corresponding behavior image of multiple information matching combinations, construct a warning review mechanism for the elderly person. In this early warning and review mechanism, a humanized interactive verification process is initiated, and multi-dimensional early warning interactions, including facial expression guidance, voice inquiry, and behavior review, are triggered along the verification sequence. Corresponding early warning interaction information is output, and the various early warning interaction information is deeply integrated to output the elderly person's response information combination and determine whether the elderly person has the ability to save themselves or whether it is a misoperation.
8. The anomaly identification method for home intelligent agents based on image recognition according to claim 1, characterized in that, The home-based intelligent agent determines multiple interaction events of the elderly based on the combination of response information, iterates through multiple levels to trigger each interaction event, marks the elderly's warning body areas, and marks the real-time images corresponding to the warning body areas. Based on image recognition of these real-time images, it determines the corresponding emergency events, including: The home-based intelligent agent acquires the response information combination, performs context parsing on the response information combination to output multiple sub-interaction information, determines the elderly's interaction events based on the combination of multiple sub-interaction information and corresponding abnormal behavior information, marks multiple interaction events, and triggers the corresponding multi-level iterative reasoning mechanism along multiple interaction events.
9. The anomaly identification method for home intelligent agents based on image recognition according to claim 8, characterized in that, The home-based intelligent agent determines multiple interaction events of the elderly based on the combination of response information, iterates through multiple levels to trigger each interaction event, marks the elderly's warning body areas, and marks the real-time images corresponding to the warning body areas. Based on image recognition of the real-time images, it determines the corresponding emergency events, and also includes: In the multi-level iterative reasoning mechanism, each interactive event undergoes multi-level iteration along different event dimensions. The corresponding abnormal states are marked in different event dimensions. Based on the state content of each abnormal state, the corresponding iteration level, and background interference, the warning body area of the elderly is determined. The warning body area is dynamically tracked to mark the real-time image corresponding to the warning body area, thereby realizing the dynamic capture of injury features. The corresponding emergency event is determined based on the mapping relationship between multiple injury features and emergency events.
10. An anomaly detection system for a home intelligent agent based on image recognition, characterized in that, The image recognition-based home intelligent agent anomaly detection system is applied to the image recognition-based home intelligent agent anomaly detection method as described in any one of claims 1-9; The anomaly detection system for the image recognition-based home smart agent includes: The behavior image module is used to collect indoor models of the target location, match the corresponding home intelligent agent based on the indoor model and the corresponding home information, and trigger the home intelligent agent to dynamically track the elderly in the room in order to determine multiple behavior images of the elderly. The image recognition module is used to determine the dynamic spatiotemporal events of the elderly based on image recognition of multiple behavioral images, and to construct a multi-dimensional behavioral model of the elderly by combining the elderly's facial expressions and body postures. Based on the dynamic monitoring of this multi-dimensional behavioral model, multiple abnormal behavioral nodes are identified. The early warning review module is used to determine the corresponding abnormal behavior information based on the tracing of each abnormal behavior node. Based on the abnormal behavior information, the elderly's previous early warning information and the corresponding behavior image, the early warning review mechanism is determined, and the elderly are interacted with in different dimensions along the early warning review mechanism to determine the combination of the elderly's response information. The emergency event module is used by the home smart agent to determine multiple interaction events of the elderly based on the combination of response information, and to perform multi-level iterations to trigger each interaction event in order to mark the warning body area of the elderly and mark the real-time image corresponding to the warning body area. The corresponding emergency event is determined based on the image recognition of the real-time image.