Smart home intrusion accurate identification and early warning method based on multi-sensor fusion

Through multi-sensor fusion and timing coding analysis, combined with visual features, the confidence level of smart home intrusion is evaluated, and accurate identification and hierarchical response is achieved, which solves the problem of high false alarm rates in complex scenarios of existing systems, and improves the accuracy and user experience of intrusion recognition.

CN120199005APending Publication Date: 2025-06-24JIANGXI YANGNING TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510463670.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing smart home intrusion detection system has a high false alarm rate in complex scenarios, making it difficult to accurately identify intrusion behavior, especially under low illumination, occlusion or environmental interference.

Method used

The multi-sensor fusion method is adopted to collect data through gate magnetic sensors, PIR infrared sensors and cameras, perform timing encoding and interactive analysis, and obtain switch-mobile state timing interaction characteristics and full-time domain monitoring image characteristics. Combined with these characteristics, the mobile behavior analysis is carried out to evaluate the confidence level, and realize the three-level response mechanism of high confidence intrusion immediate alarm, medium confidence manual review, and low confidence log recording.

Benefits of technology

It significantly improves the accuracy and environmental adaptability of intrusion recognition, reduces the misjudgment rate, optimizes the user experience, and takes into account the real-time security and operational friendliness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199005A_ABST
    Figure CN120199005A_ABST
Patent Text Reader

Abstract

The invention provides a smart home intrusion accurate identification and early warning method based on multi-sensor fusion, and relates to the field of intrusion early warning, and the method comprises the steps: firstly obtaining the time sequence distribution of a home switch state and the time sequence distribution of an infrared radiation value, and carrying out the time sequence coding and interaction analysis, so as to obtain a switch-moving state time sequence interaction feature; and then a time sequence of a monitoring image frame is acquired to carry out movement behavior global analysis, and movement behavior analysis is carried out in combination with switch-movement state time sequence interaction characteristics, so that the confidence level is evaluated, and a three-level response mechanism of high-confidence intrusion immediate alarm, medium-confidence artificial review and low-confidence log recording is realized. Thus, the misjudgment problem of a traditional method is solved, the accuracy and environmental adaptability of intrusion recognition are remarkably improved, meanwhile, the user experience is optimized through hierarchical response, and security real-time performance and operation friendliness are both considered.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intrusion warning, and more particularly, in the embodiments of this application, it relates to a method for accurate identification and warning of smart home intrusion based on multi-sensor fusion. Background Art

[0002] With the rapid development of Internet of Things technology, smart home has become an important part of modern home security systems. In the scenario of home intrusion detection, how to accurately identify abnormal intrusion behaviors and trigger the warning mechanism in a timely manner is directly related to the safety of users' lives and property.

[0003] Current mainstream methods mostly use single sensing modalities such as door magnetic sensors, infrared sensors or cameras for monitoring, but there are significant technical bottlenecks in practical applications: single-modal sensors are limited by their inherent physical characteristics. Door magnetic sensors can only detect the mechanical displacement of doors and windows and cannot sense the movement trajectory of the human body. Infrared sensors are easily interfered by environmental temperature changes and the activities of small pets. Cameras have monitoring blind spots in low-light, occlusion or intruder disguise scenarios. Although traditional multi-sensor fusion schemes attempt to combine time-series signals and visual information, they still use shallow feature splicing or simple decision fusion at the cross-modal feature interaction level, and fail to effectively capture the spatio-temporal correlation between mechanical displacement, thermal radiation characteristics and visual dynamics in intrusion behaviors, resulting in a high false alarm rate in complex scenarios, frequent problems such as false triggers at night and misidentification of pets.

[0004] Therefore, there is a need for a scheme for accurate identification and warning of smart home intrusion based on multi-sensor fusion. Summary of the Invention

[0005] To solve the above technical problems, this application is proposed. The embodiments of this application provide a method for accurate identification and warning of smart home intrusion based on multi-sensor fusion. First, it obtains the temporal distribution of home switch states and the temporal distribution of infrared radiation values, and performs temporal encoding and interaction analysis to obtain switch-movement state temporal interaction features. Then, it obtains the time series of monitoring image frames for global analysis of movement behaviors, and combines the switch-movement state temporal interaction features for movement behavior analysis to evaluate the confidence level, realizing a three-level response mechanism of immediate alarm for high-confidence intrusion, manual review for medium-confidence, and log recording for low-confidence. In this way, the misjudgment problem of traditional methods is solved, the accuracy and environmental adaptability of intrusion identification are significantly improved, and at the same time, the user experience is optimized through hierarchical response, taking into account both the real-time nature of security and the user-friendliness of operation.

[0006] According to one aspect of this application, there is provided a method for accurate identification and warning of smart home intrusion based on multi-sensor fusion, which includes:

[0007] Collect the timing distribution of the switch state, the timing distribution of the infrared radiation value, and the time series of the monitoring image frames within a predetermined time period through a door magnetic sensor, a PIR infrared sensor, and a camera respectively;

[0008] Perform timing encoding and interaction analysis on the timing distribution of the switch state and the timing distribution of the infrared radiation value to obtain the switch-movement state timing interaction feature;

[0009] Perform global analysis of the movement behavior on the time series of the monitoring image frames to obtain the full-time domain monitoring image feature;

[0010] Perform movement behavior analysis on the switch-movement state timing interaction feature and the full-time domain monitoring image feature to obtain the confidence intrusion level, including: performing fine-grained cross-modal movement behavior spatio-temporal interaction characterization on the switch-movement state timing interaction feature and the full-time domain monitoring image feature to obtain the movement behavior spatio-temporal fine-grained interaction coding feature; determining the confidence intrusion level based on the movement behavior spatio-temporal fine-grained interaction coding feature.

[0011] Compared with the prior art, a method for accurate identification and early warning of smart home intrusion based on multi-sensor fusion provided by the present application first obtains the timing distribution of the home switch state and the timing distribution of the infrared radiation value and performs timing encoding and interaction analysis to obtain the switch-movement state timing interaction feature, then obtains the time series of the monitoring image frames for global analysis of the movement behavior, and combines the switch-movement state timing interaction feature for movement behavior analysis to evaluate the confidence level, realizing a three-level response mechanism of immediate alarm for high-confidence intrusion, manual review for medium-confidence, and log recording for low-confidence. In this way, the misjudgment problem of the traditional method is solved, the accuracy and environmental adaptability of intrusion recognition are significantly improved, and at the same time, the user experience is optimized through hierarchical response, taking into account the real-time security and operation friendliness. Brief Description of the Drawings

[0012] By describing the embodiments of the present application in more detail in conjunction with the drawings, the above and other objects, features, and advantages of the present application will become more obvious. The drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification, and are used to explain the present application together with the embodiments of the present application, and do not constitute a limitation to the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0013] Figure 1 It is a flowchart of a method for accurate identification and early warning of smart home intrusion based on multi-sensor fusion according to an embodiment of the present application.

[0014] Figure 2 It is a schematic diagram of data flow of a method for accurate identification and early warning of smart home intrusion based on multi-sensor fusion according to an embodiment of the present application.

[0015] Figure 3 It is a flowchart for performing mobile behavior analysis on the switch-movement state temporal interaction feature and the full-time domain monitoring image feature to obtain a confidence intrusion level in the intelligent home intrusion precise recognition and warning method based on multi-sensor fusion according to an embodiment of the present application.

[0016] Figure 4 It is a flowchart for performing fine-grained cross-modal mobile behavior spatio-temporal interaction characterization on the switch-movement state temporal interaction feature and the full-time domain monitoring image feature to obtain a mobile behavior spatio-temporal fine-grained interaction coding feature in the intelligent home intrusion precise recognition and warning method based on multi-sensor fusion according to an embodiment of the present application.

[0017] Figure 5 It is a flowchart for calculating fine-grained saliency descriptors of each monitoring image local region feature coding vector in the set of monitoring image local region feature coding vectors to obtain a set of mobile behavior local region spatio-temporal feature fine-grained saliency descriptors in the intelligent home intrusion precise recognition and warning method based on multi-sensor fusion according to an embodiment of the present application. Detailed implementation manners

[0018] Various exemplary embodiments, features and aspects of the present application will be described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.

[0019] The specifically used word "exemplary" herein means "serving as an example, embodiment or illustration". Any embodiment described as "exemplary" herein need not be construed as superior to or better than other embodiments.

[0020] In addition, for better explaining the present application, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present application can also be implemented without some specific details. In some instances, methods, means, elements and circuits well-known to those skilled in the art are not described in detail so as to highlight the gist of the present application.

[0021] Furthermore, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, "a plurality" means two or more unless otherwise specifically defined.

[0022] With the development of Internet of Things technology, smart home has become an important part of modern home security. In particular, it is crucial to accurately identify abnormal behaviors and give early warnings in home intrusion detection. However, currently, mainstream single-sensing modalities such as door magnetic sensors, infrared sensors, or cameras face challenges in practical applications due to problems such as only being able to detect door and window displacements, being vulnerable to environmental interference, and having monitoring blind spots. Traditional multi-sensor fusion solutions attempt to combine time-series signals and visual information. However, due to limitations in cross-modal feature interaction, they fail to effectively capture the spatio-temporal correlations in intrusion behaviors, resulting in a still relatively high false alarm rate in complex scenarios, such as frequent false triggers at night and misidentifications of pets.

[0023] Based on this, the technical concept of this application is to adopt an artificial intelligence-based data analysis and encoding method. First, synchronously collect the time-series signals of the switch states of door magnetic sensors, the time-series data of the thermal radiation values of infrared sensors, and the video streams of cameras. Respectively perform time-series embedding encoding on the door magnetic switch states to capture the time-correlation patterns of mechanical displacements, perform time-series correlation encoding on the infrared radiation values to extract the changing laws of the thermal characteristics of human movements, and then perform interactive analysis on the two to establish the spatio-temporal correlation relationship between door and window opening / closing and human movement. At the same time, extract the moving behavior characteristics of the video frame sequence and generate a continuous spatial trajectory representation through full-time-domain aggregation. Subsequently, perform fine-grained cross-modal interaction between the sensor time-series interaction vectors and the visual spatio-temporal features, and use saliency modeling to align the correlation nodes of mechanical displacements, thermal radiation, and visual dynamics to generate fused multi-dimensional spatio-temporal features. Finally, evaluate the confidence level and implement a three-level response mechanism of immediate alarm for high-confidence intrusions, manual review for medium-confidence, and log recording for low-confidence. This solution effectively solves the misjudgment problems of traditional methods in scenarios such as complex lighting and pet interference, significantly improves the accuracy and environmental adaptability of intrusion recognition, and at the same time optimizes the user experience through a hierarchical response mechanism, taking into account both the real-time nature of security and the user-friendliness of operation.

[0024] This application proposes a method for accurate identification and early warning of smart home intrusion based on multi-sensor fusion. Figure 1 The flowchart of the method for accurate identification and early warning of smart home intrusion based on multi-sensor fusion according to an embodiment of this application. Figure 2 The schematic diagram of data flow of the method for accurate identification and early warning of smart home intrusion based on multi-sensor fusion according to an embodiment of this application. As Figure 1 and Figure 2As shown, the intelligent home intrusion precise recognition and warning method based on multi-sensor fusion according to an embodiment of the present application includes: S110, respectively collecting the time series distribution of the switch state, the time series distribution of the infrared radiation value, and the time series of the monitoring image frames in a predetermined time period through a door magnetic sensor, a PIR infrared sensor, and a camera; S120, performing time series encoding and interaction analysis on the time series distribution of the switch state and the time series distribution of the infrared radiation value to obtain the switch-movement state time series interaction feature; S130, performing a global analysis of the movement behavior on the time series of the monitoring image frames to obtain the full-time domain monitoring image feature; S140, performing a movement behavior analysis on the switch-movement state time series interaction feature and the full-time domain monitoring image feature to obtain the confidence intrusion level.

[0025] In the above-mentioned intelligent home intrusion precise recognition and early warning method based on multi-sensor fusion, in step S110, the time series distribution of the switch state, the time series distribution of the infrared radiation value, and the time series of the monitoring image frames in a predetermined time period are respectively collected through a door magnetic sensor, a PIR infrared sensor, and a camera. It should be understood that a door magnetic sensor is a device used to detect the switch state of objects such as doors and windows. It determines whether the doors and windows are opened or closed by sensing the relative position change between two components, thus providing key information about the security status of the entrances and exits. Considering that any unauthorized operation of doors and windows often indicates the existence of a security threat, the door magnetic sensor is used to monitor the change of the switch state of entrances and exits such as doors and windows. It can provide accurate time series information, reflecting whether the doors and windows are opened or closed and their dynamic change process. At the same time, a passive infrared (PIR) sensor is a device that can detect the infrared radiation emitted by the human body or other warm-blooded animals. It uses the sensitivity of pyroelectric materials to temperature changes to identify moving targets in the monitoring area, and is particularly suitable for capturing signs of intrusion behavior based on the difference in human body heat. The PIR infrared sensor detects the infrared rays radiated by the human body to sense whether there are moving human targets in the space. The time series distribution of the infrared radiation value output by this sensor not only reflects the change of the heat source, but also indirectly indicates the position and path of personnel activities. By recording the change of the infrared radiation value over a period of time, a detailed map of the internal temperature fluctuation of the environment can be constructed, which is of great significance for distinguishing real intruders from false alarms caused by environmental factors. In addition, as the main tool for visual monitoring, the camera is responsible for continuously recording the video stream of the monitoring area and extracting the time series of the monitoring image frames from it. These image frames contain rich spatial information, such as human outlines, action postures, and scene details, providing an intuitive basis for subsequent behavior analysis. Integrating the data of these three different types of sensors, namely the change history of the door and window switch state, the fluctuation record of the infrared radiation intensity, and the visual content in the video picture, can comprehensively cover information in multiple dimensions from physical displacement to thermal energy signals to visual features. This comprehensive data collection method can not only enhance the ability to identify complex intrusion behaviors, but also effectively reduce the misjudgment problem caused by the limitations of a single sensor. For example, the door magnetic can only detect mechanical displacement and cannot identify the specific actor, the PIR is vulnerable to non-target heat sources, and the camera may fail under low light conditions. In this way, potential security threats can be captured in the first time and a solid foundation can be laid for further in-depth analysis.

[0026] In the embodiments of the present application, step S120 of performing temporal encoding and interaction analysis on the temporal distribution of the switch state and the temporal distribution of the infrared radiation value to obtain the switch-movement state temporal interaction feature includes: S121, performing temporal embedding encoding on the temporal distribution of the switch state to obtain a switch state temporal embedding encoding feature vector; S122, performing temporal correlation encoding on the temporal distribution of the infrared radiation value to obtain an infrared radiation value temporal correlation encoding feature vector; S123, performing interaction analysis on the switch state temporal embedding encoding feature vector and the infrared radiation value temporal correlation encoding feature vector to obtain a switch-movement state temporal interaction vector as the switch-movement state temporal interaction feature.

[0027] Specifically, in step S121, temporal embedding encoding is performed on the temporal distribution of the switch state to obtain a switch state temporal embedding encoding feature vector. It should be understood that considering that the door magnetic sensor is a basic component of physical security, the time series of its switch state often contains key clues of intrusion behavior. Traditional solutions only regard the door magnetic signal as a discrete switch event (such as the binary state of "open" or "closed"), ignoring the temporal correlation characteristics between sensor actions. For example, when an intruder tentatively pries the doors and windows, it may trigger multiple short-term slightly open states, or irregular opening behaviors during abnormal periods (such as early morning). The mechanical displacement characteristics implicit in these temporal patterns are essentially different from the stable temporal sequence of normal family members opening and closing doors and windows. Therefore, in order to better capture the progressive abnormal state during the process of the doors and windows being damaged (such as the non-continuous switch jitter caused by continuous pressure), the present application performs temporal embedding encoding on the temporal distribution of the switch state to obtain a switch state temporal embedding encoding feature vector. Particularly, in a specific example of the present application, a temporal neural network (such as Transformer or LSTM) can be introduced to perform continuous representation learning on the original switch signal, mapping the discrete switch event sequence into a dense vector in a high-dimensional feature space. Specifically, the encoding process not only retains the jump time points of the switch state, but also captures the dependence relationship of state changes within different time windows (such as the difference between short-term high-frequency triggers and long-term continuous openings) through the attention mechanism or gated recurrent unit, thereby converting the physical signal of mechanical displacement into an embedding vector containing temporal dynamic semantics. For example, multiple abnormal triggers of the door magnetic sensor within a short period of time may be encoded as a "tentative intrusion" feature pattern, while the switch behaviors of normal family members show stable periodicity or occasional single triggers.

[0028] Specifically, in step S122, temporal correlation encoding is performed on the temporal distribution of the infrared radiation values to obtain an infrared radiation value temporal correlation encoding feature vector. It should be understood that considering that the PIR infrared sensor realizes behavior perception by capturing the thermal radiation changes caused by human movement, there are significant environmental sensitivity and semantic ambiguity in the temporal distribution of the originally output infrared radiation values. Since the working principle of the infrared sensor depends on the dynamic difference between the environmental temperature and the target heat source, the temporal data collected by it not only contains real human movement characteristics, but may also be mixed with interference signals caused by the start and stop of air conditioners, sunlight irradiation, or pet activities. Traditional methods usually adopt shallow analysis means such as threshold triggering or sliding window statistics, and can only extract local features such as short-term radiation intensity mutations, but cannot analyze the dynamic association patterns of radiation values in the time dimension. For example, real intruders often present a continuous moving thermal radiation fluctuation curve (such as the heat source gradually spreading from the outside to the inside after the door or window is opened), while pet interference shows short-term high-frequency local radiation spikes. However, due to the lack of long-range modeling ability for temporal context, traditional methods are difficult to distinguish such essential differences. Based on this, the present application performs temporal correlation encoding on the temporal distribution of the infrared radiation values to obtain an infrared radiation value temporal correlation encoding feature vector. In particular, in a specific example of the present application, by designing a neural network structure with time-dependent modeling ability (such as causal convolution or temporal attention mechanism), multi-level feature extraction is performed on the continuously collected infrared radiation value sequence. The encoding process not only focuses on the radiation intensity at a single time point, but also emphasizes exploring the causal relationship and dynamic evolution law of the radiation value changes within different time windows. For example, for the continuous movement behavior of an intruder after entering through a door or window, the encoder can capture the progressive enhancement trend of the thermal radiation waveform and its spatial propagation characteristics; for the instantaneous radiation peak caused by a pet's short-term scurrying, its isolated and non-persistent characteristics are identified through temporal correlation analysis.

[0029] Specifically, in step S123, an interaction analysis is performed on the switch state time-series embedded encoding feature vector and the infrared radiation value time-series correlation encoding feature vector to obtain a switch-movement state time-series interaction vector as the switch-movement state time-series interaction feature, including: performing an interaction analysis based on a cross-attention mechanism on the switch state time-series embedded encoding feature vector and the infrared radiation value time-series correlation encoding feature vector to obtain the switch-movement state time-series interaction vector. It should be understood that although the independent sensing information of the door magnetic sensor and the infrared sensor can reflect the door and window state and thermal radiation fluctuations, the spatio-temporal correlation between the two in intrusion behavior has not been effectively explored. Traditional fusion methods use simple feature splicing or weighted summation, and cannot distinguish the essential differences between "normal opening of doors and windows accompanied by family member movement" and "abnormal opening accompanied by intruder movement". For example, when a pet triggers the infrared sensor while the door magnetic is in the closed state, traditional methods are prone to misjudgment due to the lack of cross-modal correlation analysis; conversely, when an intruder uses a tool to slowly pry open the doors and windows (the door magnetic generates intermittent slightly open signals), it is difficult to distinguish whether it is wind interference or a real intrusion without synchronous infrared radiation feature evidence. This lack of cross-modal time-series correlation directly leads to insufficient discrimination ability for complex intrusion patterns. Therefore, in order to capture the dynamic coupling relationship between the two types of sensor data in the time domain, in the technical solution of this application, an interaction analysis based on a cross-attention mechanism is performed on the switch state time-series embedded encoding feature vector and the infrared radiation value time-series correlation encoding feature vector to obtain the switch-movement state time-series interaction vector. Specifically, the switch state time-series embedded encoding feature (representing the mechanical displacement pattern) and the infrared radiation value time-series correlation encoding feature (representing the thermal dynamic law) are input into a bidirectional attention network, and through an adaptive weight allocation mechanism, the spatio-temporal key nodes of the two types of signals are aligned in the feature dimension. For example, when the door magnetic detects a long-term opening signal during an abnormal period, the attention mechanism will focus on the continuously increasing area of the infrared radiation value during this period (corresponding to the heat diffusion process of the intruder moving from the door and window into the room), while suppressing the interference of the slow global radiation change caused by the heating. During this process, the model dynamically constructs a correlation map of "abnormal opening of doors and windows - spatial gradient change of thermal radiation", and accurately captures the co-evolution law of mechanical displacement and human thermal characteristics in intrusion behavior.

[0030] In an embodiment of this application, in step S130, a full-domain analysis of the movement behavior is performed on the time series of the monitoring image frames to obtain the full-time-domain monitoring image features, including: S131, performing extraction of movement behavior image features based on MobileNet on the time series of the monitoring image frames to obtain a time series of monitoring image feature maps; S132, performing full-time-domain aggregation on the time series of the monitoring image feature maps to obtain a full-time-domain monitoring image feature map as the full-time-domain monitoring image features.

[0031] Specifically, in step S131, perform MobileNet-based mobile behavior image feature extraction on the time series of the monitored image frames to obtain a time series of monitored image feature maps. It should be understood that although the monitored video stream collected by the camera can provide intuitive visual information, traditional image processing methods (such as fixed-region dynamic detection or background subtraction method) face multiple challenges in complex home environments. Invasive behaviors often have unstructured characteristics. For example, an intruder may use means such as slow movement, partial occlusion (such as covering with fabric), or sneaking in a low-illumination environment to avoid detection. Static image features extracted by ordinary convolutional neural networks (CNNs) are difficult to capture the tiny changes in moving trajectories between consecutive frames. In addition, there are a large number of dynamic interference sources in the home environment, such as curtain fluttering, light flickering, or pet running. Traditional methods are prone to misjudging them as intrusion targets. More critically, smart home devices are usually limited by computing resources. If a complex video analysis model is directly adopted, it will lead to response delays and cannot meet the real-time security requirements. Therefore, in the technical solution of this application, perform MobileNet-based mobile behavior image feature extraction on the time series of the monitored image frames to obtain a time series of monitored image feature maps. That is to say, the depthwise separable convolution structure of MobileNet significantly reduces the number of parameters, enabling it to process high-resolution video streams in real time on embedded devices. During the feature extraction process, the network not only focuses on the target position and contour in a single-frame image but also captures the continuous pattern of human movement (such as the moving direction and speed change of an intruder after entering through a door or window) through temporal convolution or cross-frame feature correlation mechanisms. For example, in a low-illumination scene, the network enhances the gradient feature extraction ability in the shadow area to identify the weak optical flow changes related to human movement; in the case of occlusion, it uses temporal context to infer the movement trajectory of the occluded target.

[0032] Specifically, in step S132, the time series of the monitored image feature maps is aggregated over the entire time domain to obtain the full-time-domain monitored image feature map as the full-time-domain monitored image feature. It should be understood that although the time series of the monitored image feature maps contains rich spatio-temporal behavior information, traditional video analysis methods usually adopt frame-by-frame detection or short-time window sliding processing, resulting in the continuous features of intrusion behaviors being fragmented into discrete representations. For example, after an intruder breaks in through a door or window, their behavior pattern may show intermittent visibility (such as being blocked by furniture) or cross-region movement (such as moving from the living room to the bedroom) in the monitored video. The image features within a single frame or a local time window are difficult to completely depict such long-term dynamic trajectories. In addition, common dynamic interferences in the home environment (such as curtain swings and light and shadow changes) can generate local responses similar to real intrusions in a single-frame feature map, and isolated features without temporal coherence verification are prone to false alarms. Based on this, in this application, the time series of the monitored image feature maps is aggregated over the entire time domain to obtain the full-time-domain monitored image feature map. Specifically, the time series of the monitored image feature maps is cascaded here to obtain the full-time-domain monitored image feature map. In this way, not only the spatial details of a single-frame image (such as human body contours and moving directions) are retained, but also the position migration law and behavior continuity of cross-frame targets are captured through long-term dependence modeling.

[0033] Figure 3 FIG. is a flowchart for performing mobile behavior analysis on the switch-movement state time-series interaction feature and the full-time-domain monitored image feature to obtain a confidence intrusion level in the intelligent home intrusion precise recognition and warning method based on multi-sensor fusion according to an embodiment of the present application. As Figure 3 shown, in step S140, performing mobile behavior analysis on the switch-movement state time-series interaction feature and the full-time-domain monitored image feature to obtain a confidence intrusion level includes: S141, performing fine-grained cross-modal mobile behavior spatio-temporal interaction characterization on the switch-movement state time-series interaction feature and the full-time-domain monitored image feature to obtain a mobile behavior spatio-temporal fine-grained interaction coding feature; S142, determining the confidence intrusion level based on the mobile behavior spatio-temporal fine-grained interaction coding feature.

[0034] Figure 4 FIG. is a flowchart for performing fine-grained cross-modal mobile behavior spatio-temporal interaction characterization on the switch-movement state time-series interaction feature and the full-time-domain monitored image feature to obtain a mobile behavior spatio-temporal fine-grained interaction coding feature in the intelligent home intrusion precise recognition and warning method based on multi-sensor fusion according to an embodiment of the present application. As Figure 4As shown, in the embodiment of the present application, in step S141, performing fine-grained cross-modal spatio-temporal interaction characterization on the switch-movement state temporal interaction feature and the full-time domain monitoring image feature to obtain a spatio-temporal fine-grained interaction coding feature of the movement behavior, including: S1411, performing feature fine-grained deconstruction on the full-time domain monitoring image feature map to obtain a set of monitoring image local area feature coding vectors; S1412, calculating fine-grained saliency descriptors of each monitoring image local area feature coding vector in the set of monitoring image local area feature coding vectors to obtain a set of spatio-temporal fine-grained saliency descriptors of the movement behavior local area; S1413, based on the set of spatio-temporal fine-grained saliency descriptors of the movement behavior local area, performing cross-domain significant interaction aggregation on the switch-movement state temporal interaction vector and the set of monitoring image local area feature coding vectors to obtain a spatio-temporal fine-grained interaction coding feature vector as the spatio-temporal fine-grained interaction coding feature of the movement behavior. It should be understood that considering that the switch-movement state temporal interaction feature mainly fuses the data information of the door magnetic sensor and the PIR infrared sensor, reflects the time series relationship between the opening and closing of doors and windows and human movement, and focuses on capturing information from the sensor signal level. The full-time domain monitoring image feature is a comprehensive feature representation of the monitoring video image in the time dimension and contains rich visual information. These two data modalities are highly complementary. If the sensor temporal feature and the visual global feature are directly concatenated or weighted and fused, it is easy to ignore the local correlation characteristics of mechanical displacement, thermal radiation dynamics, and visual trajectory in the spatio-temporal dimension. For example, when the door magnetic sensor detects an abnormal opening of the door or window and the infrared shows continuous movement, the camera may only capture local limb movements (such as a hand reaching into the door crack) due to the perspective limitation. At this time, it is difficult to establish a fine-grained correlation of "mechanical trigger-thermal radiation enhancement-local visual anomaly" through global feature fusion. At the same time, the intruder may create a spatio-temporal discontinuous behavior pattern by moving across regions (such as quickly leaving the monitoring field of view after invading through the door or window), resulting in an asymmetric distribution of the sensor temporal signal and the visual trajectory on the time axis, and shallow fusion is prone to feature confusion. Based on this, the present application performs fine-grained cross-modal spatio-temporal interaction characterization on the switch-movement state temporal interaction feature and the full-time domain monitoring image feature to obtain a spatio-temporal fine-grained interaction coding feature of the movement behavior.

[0035] That is, first, the spatio-temporal structure of the full-time domain monitoring image features is deconstructed, and it is decomposed into a sequence of local region features with semantic continuity (such as the dynamic of door and window edges, indoor path movement segments). Then, the interaction matrix between the switching-movement state temporal interaction features and the visual local features is calculated through a fine-grained explicit modeling unit, and the correlation strength and semantic consistency of the two at the spatio-temporal nodes are explicitly encoded. For example, when the door magnetic code shows "abnormal continuous opening" during a certain period and the infrared code represents "progressive heat radiation diffusion", the feature weight of the corresponding visual local region (such as around the doors and windows) during this period can be automatically enhanced, and the interference response of irrelevant regions (such as stationary indoor furniture) can be suppressed. Further, through the collaborative analysis of the semantic association significance and spatial isolation dual descriptors, the key anchors in the cross-modal interaction are dynamically identified (such as the spatio-temporal coupling of the moment when the mechanical displacement of the door and window is triggered and the appearance of the human contour in the visual image), and the sensor-visual short-term pseudo-association caused by the scurrying of pets is suppressed, so as to understand the intrusion behavior more comprehensively.

[0036] Specifically, in step S1411, the feature map of the full-time domain monitoring image is deconstructed with fine-grained features to obtain a set of local region feature encoding vectors of the monitoring image, which is represented by the fine-grained feature deconstruction formula of the monitoring image:

[0037]

[0038] Among them, F2 represents the feature map of the full-time domain monitoring image, Partition(·) represents the fine-grained feature deconstruction operation, and are the first, second, i-th, and n-th local region feature encoding vectors of the monitoring image in the set of local region feature encoding vectors of the monitoring image respectively, represents the operation on Linear transformation. It should be understood that since the full-time domain monitoring image features present continuous dynamic features in the time dimension, the spatial structure information it contains (such as the thermal distribution in the door and window area, the local morphology of the human movement trajectory) and the temporal correlation features (such as the progressive changes in the actions of intruders) are often concentrated in specific regional segments. If global feature interaction is directly adopted, the model will cause key local features to be submerged in the overall background noise due to the averaging perspective. For example, the local infrared fluctuations caused by the short-term movement of pets may be misjudged as the whole-body thermal radiation features of intruders, while the key mechanical displacements and visual dynamics of actual intrusion behaviors are often significantly correlated only in specific spatio-temporal regions. Feature fine-grained deconstruction transforms the originally chaotic global feature stream into resolvable local semantic units through spatio-temporal decoupling and reconstruction. Specifically, a multi-scale spatio-temporal segmentation strategy is adopted to divide the continuous video features into local regional segments with spatio-temporal continuity (such as the dynamic area at the edge of the door and window, the moving segment on the indoor path), and independent local region feature encoding vectors of the monitoring images are extracted for each segment. This deconstruction is not simply a spatial dimensionality reduction, but dynamically identifies key regions through a spatio-temporal attention mechanism. For example, when an intruder climbs a window, the algorithm will automatically focus on the deformed area of the glass frame and the asymptotic coverage process of the human contour, while suppressing the interference features of the vibration of the outdoor air conditioner in the background. In this way, the ability of the model to capture small but key abnormal behaviors is significantly improved through local feature decoupling.

[0039] Figure 5 It is a flowchart for calculating the fine-grained saliency descriptors of each local region feature encoding vector in the set of local region feature encoding vectors of the monitoring images in the intelligent home intrusion precise recognition and warning method based on multi-sensor fusion according to the embodiments of the present application to obtain a set of spatio-temporal features of local regions of moving behaviors. As Figure 5As shown, in the embodiment of the present application, in step S1412, calculating the fine-grained saliency descriptors of each monitoring image local region feature coding vector in the set of monitoring image local region feature coding vectors to obtain a set of fine-grained spatio-temporal feature saliency descriptors for the moving behavior local region includes: S1412-1, inputting each monitoring image local region feature coding vector in the set of monitoring image local region feature coding vectors and the switch-movement state time series interaction vector into a fine-grained saliency modeling unit to obtain a set of fine-grained saliency modeling coding matrices for the moving behavior local region spatio-temporal features; S1412-2, based on the set of fine-grained saliency modeling coding matrices for the moving behavior local region spatio-temporal features, determining a set of spatio-temporal semantic association saliency descriptors for the moving behavior local region and a set of spatio-temporal isolation descriptors for the moving behavior local region; S1412-3, based on the set of spatio-temporal semantic association saliency descriptors for the moving behavior local region and the set of spatio-temporal isolation descriptors for the moving behavior local region, determining the set of fine-grained spatio-temporal feature saliency descriptors for the moving behavior local region.

[0040] Specifically, in step S1412-1, inputting each monitoring image local region feature coding vector in the set of monitoring image local region feature coding vectors and the switch-movement state time series interaction vector into a fine-grained saliency modeling unit to obtain a set of fine-grained saliency modeling coding matrices for the moving behavior local region spatio-temporal features, which is represented by the fine-grained saliency modeling coding formula for the moving behavior local region spatio-temporal features as:

[0041]

[0042] Where, T represents the transpose operation, and φ(·) is a feature mapping function, such as a linear mapping or a non-linear kernel function, v1 represents the switch-movement state time series interaction vector, L represents the length of, M i represents The fine-grained explicit modeling encoding matrix for the spatio-temporal features of the local area of the movement behavior between and v1, that is, the i-th fine-grained explicit modeling encoding matrix for the spatio-temporal features of the local area of the movement behavior in the set of fine-grained explicit modeling encoding matrices for the spatio-temporal features of the local area of the movement behavior. It should be understood that the fine-grained explicit modeling unit realizes the structured representation of the inter-modal causal relationship through explicit interaction modeling. Specifically, the fine-grained explicit modeling unit adopts a parameterized interaction function (such as an attention mechanism or a graph neural network), and calculates the coupling strength and pattern between each local area feature encoding vector of the monitored image (such as the movement trajectory encoding of the door and window area) and the global switch-movement state time-series interaction vector (such as the door lock state switching sequence) in the spatio-temporal dimension. For example, when an intruder tries to pry open the door lock, the door magnetic sensor will generate a specific time-series pulse signal (such as three consecutive short-time triggers), and the visual system detects severe deformation features at the local pixel level around the door frame. The explicit modeling unit can quantify the co-occurrence relationship between these two types of features through matrix operations, rather than simply adding or splicing the two types of feature vectors. Those of ordinary skill in the art should be aware that traditional methods for processing cross-modal fusion, such as direct splicing or weighted summation, are usually regarded as implicit fusion methods. Under this method, it is difficult for the model to effectively learn and display the interaction patterns between different modalities. In contrast, the fine-grained explicit modeling unit aims to break through the limitations of implicit fusion and explicitly represent the interaction relationship between modalities in matrix form. Such an explicit expression not only enhances the interpretability of the model, but also concretizes the local information in the local area feature encoding vector of the monitored image and the global information in the switch-movement state time-series interaction vector, reflects their correlation in the feature space, and provides a structured information basis for further saliency analysis.

[0043] In an embodiment of the present application, step S1412-2, based on the set of fine-grained explicit modeling encoding matrices of the spatio-temporal characteristics of the local area of the movement behavior, determining the set of spatio-temporal semantic association significant descriptors of the local area of the movement behavior and the set of spatio-temporal isolation descriptors of the local area of the movement behavior, includes: S1412-21, performing matrix optimization based on the entanglement structure on each fine-grained explicit modeling encoding matrix of the spatio-temporal characteristics of the local area of the movement behavior in the set of fine-grained explicit modeling encoding matrices of the spatio-temporal characteristics of the local area of the movement behavior to obtain a set of optimized fine-grained explicit modeling encoding matrices of the spatio-temporal characteristics of the local area of the movement behavior; S1412-22, performing activation processing based on the trace metric on each optimized fine-grained explicit modeling encoding matrix of the spatio-temporal characteristics of the local area of the movement behavior in the set of optimized fine-grained explicit modeling encoding matrices of the spatio-temporal characteristics of the local area of the movement behavior to obtain the set of spatio-temporal semantic association significant descriptors of the local area of the movement behavior; S1412-23, performing constraint processing based on the F-norm on each optimized fine-grained explicit modeling encoding matrix of the spatio-temporal characteristics of the local area of the movement behavior in the set of optimized fine-grained explicit modeling encoding matrices of the spatio-temporal characteristics of the local area of the movement behavior to obtain the set of spatio-temporal isolation descriptors of the local area of the movement behavior.

[0044] Among them, step S1412-2, based on the set of fine-grained explicit modeling encoding matrices of the spatio-temporal characteristics of the local area of the movement behavior, determining the set of spatio-temporal semantic association significant descriptors of the local area of the movement behavior and the set of spatio-temporal isolation descriptors of the local area of the movement behavior, is expressed by the movement behavior optimization formula as:

[0045]

[0046] Semantic j =σ(Tr(M′ i ))

[0047]

[0048] Among them, log2 represents the logarithmic function value with base 2, represents matrix multiplication, represents pointwise addition by position, M I is the unit matrix with all eigenvalues being one, M′ i represents M i The optimized fine-grained explicit modeling encoding matrix of the spatio-temporal characteristics of the local area of the optimized movement behavior, Tr(·) is the trace metric of the matrix, σ is the sigmoid activation function, Semantic i represents the spatio-temporal semantic association significant descriptor corresponding to M′ i represents calculating the square of the F-norm, Spatiali Denote M' i The corresponding spatio-temporal isolation descriptor for the local region of the movement behavior. It should be understood that in this step, the spatio-temporal semantic association significant descriptor for the local region of the movement behavior quantifies the association strength at the semantic level between modalities. By calculating the matching degree of the two modal features in the local interaction region in terms of semantic content, it reflects their similarity at the high-level concept. The spatio-temporal isolation descriptor for the local region of the movement behavior focuses on the spatial heterogeneity of the modal interaction. On the premise that the feature map of the full-time domain monitoring image contains spatial coordinate information, it focuses on the discreteness of the spatial distribution of the local interaction region or the difference in the interaction patterns of different regions. It should be noted that in this application, the tensor network structure can be characterized by the entanglement entropy, and the cross-modal interaction significance is decomposed into two orthogonal dimensions: the semantic association strength and the spatial heterogeneity. Among them, the semantic association strength depicts the similarity at the content level, and the spatial heterogeneity supplements the difference in the spatial structure. The two cooperate to construct a fine-grained multi-dimensional analysis framework for the cross-modal interaction features.

[0049] Here, for each spatio-temporal feature fine-grained explicit modeling encoding matrix M in the set of spatio-temporal feature fine-grained explicit modeling encoding matrices for the local region of the movement behavior i , under the assumption of the spatio-temporal feature fine-grained explicit modeling encoding matrix , where is the feature space to which each spatio-temporal feature fine-grained explicit modeling encoding matrix M i belongs, then the network entanglement structure of the tensor formed by the set of spatio-temporal feature fine-grained explicit modeling encoding matrices M i can be expressed as Thus, the trace metric Tr(M i ) of the spatio-temporal feature fine-grained explicit modeling encoding matrix M i can be regarded as the entanglement entropy representation of the tensor network.

[0050] By optimizing the entanglement entropy saturation of the tensor network through entropy saturation support, perform eigenvalue activation processing on each spatio-temporal feature fine-grained explicit modeling encoding matrix M i , that is, apply the sigmoid function to each element of the spatio-temporal feature fine-grained explicit modeling encoding matrix M i to map it to the interval [0,1], and then perform element-wise multiplication with the all-1 matrix. In this way, optimize the alignment of the low-rank structure and the high-order interaction features to ensure that when calculating the trace metric (i.e., the subsystem entropy) Tr(M i) When this occurs, the entire tensor network reaches the entanglement entropy saturation state. Through this optimization, the measurement accuracy of the semantic correlation strength between modalities in the local interaction region is significantly enhanced, providing a more discriminative basis for significant weight allocation for subsequent fused features.

[0051] Specifically, in step S1412-3, based on the set of spatio-temporal semantic correlation significant descriptors of the local region of the movement behavior and the set of spatio-temporal isolation descriptors of the local region of the movement behavior, a set of fine-grained significant descriptors of the spatio-temporal characteristics of the local region of the movement behavior is determined, which is represented by the formula for determining the fine-grained significant descriptors of the spatio-temporal characteristics of the local region of the movement behavior:

[0052]

[0053] where α, β, and γ are weight factors for adjusting the Semantic i contribution, and δ, ∈, and μ are weight factors for adjusting the Spatial i contribution, and o i represents the i-th fine-grained significant descriptor of the spatio-temporal characteristics of the local region of the movement behavior in the set of fine-grained significant descriptors of the spatio-temporal characteristics of the local region of the movement behavior. It should be understood that considering that single-modal or single-dimensional features cannot comprehensively characterize the multiple attributes of complex intrusion behaviors. In practical applications, intruders may evade detection by disguising their movement trajectories (such as slowly pushing a door to reduce the frequency of door magnetic triggers) or using environmental interference (such as heating heat sources to confuse infrared sensing). Relying solely on semantic correlation is likely to overlook the spatial abnormal distribution characteristics of abnormal behaviors, while simply considering spatial isolation may misjudge the normal activities of family members as intrusions. By fusing the spatio-temporal semantic correlation significant descriptor of the local region of the movement behavior (characterizing the temporal coupling strength between the mechanical displacement of doors and windows and the human body's thermal radiation) and the spatio-temporal isolation descriptor of the local region of the movement behavior (reflecting the spatial abnormality between the dynamics of the local region of the surveillance video and the overall environment), the limitations of traditional single-dimensional judgment can be broken through. By constructing a fine-grained discriminant basis with dual spatio-temporal sensitivity, it is possible to dynamically balance the temporal coherence and spatial abnormality of behaviors, especially for complex scenarios such as short-term infrared triggers caused by pets scurrying (high spatial isolation but low semantic correlation) or doors and windows shaking due to wind (high semantic correlation but no spatial abnormality). False positives are suppressed through two-dimensional fusion. The resulting set of fine-grained significant descriptors of the spatio-temporal characteristics of the local region of the movement behavior not only retains the microscopic correlation nodes between the sensor temporal interaction features and visual dynamics but also strengthens the deviation characteristics of abnormal behaviors in spatial distribution. Specifically, the fusion is not a simple linear combination but rather coordinates the contribution weights of the two types of descriptors, namely spatio-temporal semantic correlation degree and spatial isolation, to form a synergistic enhancement effect by exploring their inherent complementary relationship.

[0054] Specifically, in step S1413, based on the set of spatio-temporal feature fine-grained saliency descriptors of the local area of the movement behavior, cross-domain significant interaction aggregation is performed on the switch-movement state temporal interaction vector and the set of local area feature encoding vectors of the monitoring image to obtain a spatio-temporal fine-grained interaction encoding feature vector of the movement behavior as the spatio-temporal fine-grained interaction encoding feature of the movement behavior, which is expressed by the spatio-temporal cross-domain significant interaction aggregation formula of the movement behavior as:

[0055]

[0056] where n represents the number of vectors in the set of local area feature encoding vectors of the monitoring image, and v f represents the spatio-temporal fine-grained interaction encoding feature vector. It should be understood that considering the spatio-temporal misalignment and information redundancy in traditional uniform fusion, through the dynamic weight allocation mechanism of the spatio-temporal feature fine-grained saliency descriptor of the local area of the movement behavior, the spatio-temporal alignment relationship between the physical state of the sensor and the visual dynamics can be established. For example, the spatio-temporal coupling intensity (high-saliency area) of the displacement of the door and window hinge and the light and shadow of the curtain swing can be identified, the fusion weight of the deformation hot area of the door hinge and the corresponding pixel block feature vector can be automatically increased, and at the same time, the interference of low-correlation image features generated by the natural fluttering of the curtain can be suppressed. This can strengthen the cross-modal correlation, reduce the participation of irrelevant image features, make the feature vectors in the high-saliency area be assigned a larger weight, and reduce the fusion coefficient in the low-saliency area. This adaptive mechanism realizes the screening and strengthening of cross-modal key information, and at the same time automatically filters out weakly correlated or interfering feature data, thereby improving the effectiveness of multi-source information fusion.

[0057] In an embodiment of the present application, step S142, determining the confidence intrusion level based on the spatio-temporal fine-grained interaction coding features of the movement behavior, includes: inputting the spatio-temporal fine-grained interaction coding feature vector of the movement behavior into a warning confidence evaluator based on a classifier to obtain an evaluation result, and the evaluation result is used to represent the confidence intrusion level. That is, the spatio-temporal fine-grained interaction coding feature vector obtained after a series of previous feature extraction, coding, and cross-modal interaction analysis contains rich information about the movement behavior and possible intrusion situations. However, these feature vectors themselves cannot be directly converted into clear intrusion judgment results. A classifier is an effective tool that can process and analyze these feature vectors, and according to the patterns and rules learned through pre-training, map them to different intrusion levels, thus realizing the conversion from features to decisions. In particular, in the scenario of smart home intrusion detection, a quantitative representation is required for whether an intrusion occurs and the severity of the intrusion. Just judging the intrusion result as "yes" or "no" is too simple to meet the needs of practical applications. Through the warning confidence evaluator based on a classifier, the possibility of intrusion can be quantified into different confidence intrusion levels, such as high, medium, low, etc., so that the risk level of intrusion can be more accurately conveyed, providing a more targeted basis for subsequent response measures.

[0058] In an embodiment of the present application, in step S142, the spatio-temporal fine-grained interaction coding feature vector of the movement behavior is input into a classifier-based early warning confidence evaluator to obtain an evaluation result, and the evaluation result is used to represent the confidence intrusion level, including: in response to the confidence intrusion level being a high-confidence intrusion, triggering an alarm and sending an emergency notification to the user; in response to the confidence intrusion level being a medium-confidence intrusion, the user is required to make a manual confirmation to determine whether it is a false alarm; in response to the confidence intrusion level being a low-confidence intrusion, recording an abnormal event log. It should be understood that when the spatio-temporal fine-grained interaction coding feature vector of the movement behavior is fed into the classifier-based early warning confidence evaluator, the classifier-based early warning confidence evaluator comprehensively analyzes each spatio-temporal fine-grained interaction coding feature vector of the movement behavior according to the pre-trained model parameters and algorithm logic, and maps it into different intrusion level intervals. If the early warning confidence evaluator determines that the current situation belongs to a high-confidence intrusion, this means that it is highly certain that an unauthorized intrusion behavior has been detected. At this time, the alarm program will be immediately triggered, the alarm device will be activated to emit a sound and light warning, and at the same time, an emergency notification will be quickly sent to the user's mobile device through the smart home network to ensure that the user can be aware of the abnormal situation at home in the first time and take corresponding measures. If the evaluation result is a medium-confidence intrusion, it indicates that although there are signs of suspicious activities, it is still not sufficient to fully confirm whether an intrusion event has actually occurred. In this case, the alarm operation will not be automatically executed, but the user will be prompted to perform a manual review. The user can view the real-time monitoring video or other relevant data through the smart home platform to further determine whether action needs to be taken. This method helps to avoid unnecessary panic or resource waste caused by false alarms. For those situations considered to be low-confidence intrusions, that is, some minor abnormalities are identified but are not sufficient to pose a major threat, such as sensor fluctuations caused by the normal activities of pets at home, only the corresponding abnormal event log will be recorded in the background for future reference or data analysis. This hierarchical response strategy not only improves the accuracy and practicality, but also can flexibly adjust the response measures according to the actual threat level, thereby effectively improving the efficiency of the overall security and the user experience. Through such a design, smart home security can more intelligently and efficiently protect the lives and property of family members.

[0059] In summary, the intelligent home intrusion precise recognition and early warning method based on multi-sensor fusion according to the embodiments of the present application is elucidated. Firstly, the temporal distribution of home switch states and the temporal distribution of infrared radiation values are obtained and subjected to temporal encoding and interaction analysis to obtain the switch-movement state temporal interaction features. Then, the time series of monitoring image frames are obtained for global analysis of movement behaviors, and combined with the switch-movement state temporal interaction features for movement behavior analysis to evaluate the confidence level, realizing a three-level response mechanism of immediate alarm for high-confidence intrusion, manual review for medium-confidence, and log recording for low-confidence. In this way, the misjudgment problem of traditional methods is solved, the accuracy and environmental adaptability of intrusion recognition are significantly improved, and at the same time, the user experience is optimized through hierarchical response, taking into account the real-time security and operation friendliness.

Claims

1. A smart home intrusion accurate identification and early warning method based on multi-sensor fusion, characterized in that: include: The time series distribution of the switch state, the time series distribution of the infrared radiation value and the time series of the monitoring image frame in a predetermined time period are collected through the door magnetic sensor, the PIR infrared sensor and the camera respectively; Performing time series coding and interaction analysis on the time series distribution of the switch state and the time series distribution of the infrared radiation value to obtain a switch-movement state time series interaction feature; Performing a full-domain analysis of movement behavior on the time series of the monitoring image frames to obtain full-time-domain monitoring image features; Performing mobile behavior analysis on the switch-mobile state temporal interaction feature and the full-time domain monitoring image feature to obtain a confidence intrusion level, including: performing fine-grained cross-modal mobile behavior spatiotemporal interaction characterization on the switch-mobile state temporal interaction feature and the full-time domain monitoring image feature to obtain a mobile behavior spatiotemporal fine-grained interaction coding feature; The confidence intrusion level is determined based on the spatiotemporal fine-grained interactive coding features of the movement behavior.

2. The smart home intrusion accurate identification and early warning method based on multi-sensor fusion according to claim 1 is characterized in that: Performing time series coding and interactive analysis on the time series distribution of the switch state and the time series distribution of the infrared radiation value to obtain switch-movement state time series interactive features, including: Performing time-series embedding coding on the time-series distribution of the switch state to obtain a switch state time-series embedding coding feature vector; Performing time series correlation coding on the time series distribution of the infrared radiation value to obtain a time series correlation coding feature vector of the infrared radiation value; The switch state time sequence embedded coding feature vector and the infrared radiation value time sequence associated coding feature vector are interactively analyzed to obtain a switch-movement state time sequence interactive vector as the switch-movement state time sequence interactive feature.

3. The smart home intrusion accurate identification and early warning method based on multi-sensor fusion according to claim 2 is characterized in that: The switch state timing embedded coding feature vector and the infrared radiation value timing associated coding feature vector are interactively analyzed to obtain the switch-movement state timing interaction vector, including: the switch state timing embedded coding feature vector and the infrared radiation value timing associated coding feature vector are interactively analyzed based on a cross-attention mechanism to obtain the switch-movement state timing interaction vector.

4. The smart home intrusion accurate identification and early warning method based on multi-sensor fusion according to claim 3 is characterized in that: Performing a full-domain analysis of the movement behavior of the time series of the monitoring image frames to obtain full-domain monitoring image features includes: Performing MobileNet-based mobile behavior image feature extraction on the time series of the monitoring image frames to obtain a time series of monitoring image feature graphs; The time series of the monitoring image feature graph is aggregated in the full time domain to obtain a full time domain monitoring image feature graph as the full time domain monitoring image feature.

5. The smart home intrusion accurate identification and early warning method based on multi-sensor fusion according to claim 1 is characterized in that: The switch-mobility state temporal interaction feature and the full-time domain monitoring image feature are subjected to fine-grained cross-modal mobile behavior spatiotemporal interaction characterization to obtain mobile behavior spatiotemporal fine-grained interaction coding features, including: Performing feature fine-grained deconstruction on the full-time-domain surveillance image feature map to obtain a set of feature coding vectors of local regions of the surveillance image; Calculating the fine-grained saliency descriptor of each local region feature coding vector of the monitoring image in the set of local region feature coding vectors of the monitoring image to obtain a set of fine-grained saliency descriptors of the spatiotemporal features of the local region of the mobile behavior; Based on the set of fine-grained significant descriptors of the local area spatiotemporal features of the mobile behavior, the set of the switch-mobile state temporal interaction vector and the local area feature coding vector of the monitoring image are cross-domain significant interaction aggregation to obtain the mobile behavior spatiotemporal fine-grained interaction coding feature vector as the mobile behavior spatiotemporal fine-grained interaction coding feature.

6. The smart home intrusion accurate identification and early warning method based on multi-sensor fusion according to claim 5 is characterized in that: Calculating the fine-grained saliency descriptor of each local region feature coding vector of the monitoring image in the set of local region feature coding vectors of the monitoring image to obtain a set of fine-grained saliency descriptors of the spatiotemporal features of the local region of the mobile behavior, including: Inputting each monitoring image local area feature coding vector in the set of monitoring image local area feature coding vectors and the switch-mobility state temporal interaction vector into a fine-grained explicit modeling unit to obtain a set of mobile behavior local area spatiotemporal feature fine-grained explicit modeling coding matrices; Based on the set of fine-grained explicit modeling encoding matrices of the local area spatiotemporal features of the mobility behavior, determine a set of significant descriptors of spatiotemporal semantic association of the local area of ​​mobility behavior and a set of spatiotemporal isolation descriptors of the local area of ​​mobility behavior; Based on the set of spatiotemporal semantic association salient descriptors of the local area of ​​movement behavior and the set of spatiotemporal isolation descriptors of the local area of ​​movement behavior, a set of fine-grained spatiotemporal feature salient descriptors of the local area of ​​movement behavior is determined.

7. The method for accurate identification and early warning of smart home intrusion based on multi-sensor fusion according to claim 6 is characterized in that: Based on the set of fine-grained explicit modeling encoding matrices of the local area spatiotemporal features of the mobility behavior, a set of significant descriptors of spatiotemporal semantic association of the local area of ​​the mobility behavior and a set of spatiotemporal isolation descriptors of the local area of ​​the mobility behavior are determined, including: Performing matrix optimization based on the entanglement structure on each of the coding matrices for fine-grained explicit modeling of local spatiotemporal features of mobile behaviors in the set of coding matrices for fine-grained explicit modeling of local spatiotemporal features of mobile behaviors to obtain a set of optimized coding matrices for fine-grained explicit modeling of local spatiotemporal features of mobile behaviors; Performing trace-based activation processing on each optimized local area spatiotemporal feature fine-grained explicit modeling encoding matrix of the mobile behavior in the set of the optimized local area spatiotemporal feature fine-grained explicit modeling encoding matrices to obtain a set of spatiotemporal semantic association salient descriptors of the mobile behavior local area; Each optimized mobile behavior local area spatiotemporal feature fine-grained explicit modeling coding matrix in the set of optimized mobile behavior local area spatiotemporal feature fine-grained explicit modeling coding matrices is subjected to F-norm constraint processing to obtain a set of mobile behavior local area spatiotemporal isolation descriptors.

8. The smart home intrusion accurate identification and early warning method based on multi-sensor fusion according to claim 7 is characterized in that: Based on the spatiotemporal fine-grained interactive coding features of the mobile behavior, the confidence intrusion level is determined, including: inputting the spatiotemporal fine-grained interactive coding feature vector of the mobile behavior into a classifier-based warning confidence evaluator to obtain an evaluation result, wherein the evaluation result is used to represent the confidence intrusion level.

9. The smart home intrusion accurate identification and early warning method based on multi-sensor fusion according to claim 8 is characterized in that: The spatiotemporal fine-grained interactive encoding feature vector of the mobile behavior is input into a warning confidence evaluator based on a classifier to obtain an evaluation result, and the evaluation result is used to represent the confidence intrusion level, including: In response to the confidence intrusion level being a high confidence intrusion, triggering an alarm and sending an emergency notification to a user; In response to the confidence intrusion level being a medium confidence intrusion, a user performs manual confirmation to determine whether it is a false alarm; In response to the confidence intrusion level being a low confidence intrusion, an abnormal event log is recorded.

Citation Information

Patent Citations

  • Intelligent door and window anti-invasion system

    CN103198595A

  • Safety monitoring system based on big data

    CN115499627A

  • Security alarm information data interaction system and method

    CN118486152A

  • Intelligent robot with scene perception capability and method

    CN119458410A

  • Video image-based security area intrusion early warning method and system

    CN119541118A

Cited By

  • Emergency state alarm system

    CN120656287A

  • Smart home environment alarm method and device, computer equipment and storage medium

    CN121418254A

  • Video behavior recognition method and device, equipment and storage medium

    CN121583007A

  • Smart home equipment response method and system based on user positioning and voiceprint recognition

    CN122204576A