A multi-dimensional linkage control method of security alarm and video monitoring and lighting system
By extracting time and frequency domain features and performing spatiotemporal correlation processing on vibration signals, infrared trigger data, and video monitoring data, a multi-source event fusion report is generated, solving the problem of independent data processing in existing security systems and realizing the refined integration and linkage control of multi-dimensional information.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI REGLORY TECH CO LTD
- Filing Date
- 2026-04-14
- Publication Date
- 2026-08-04
AI Technical Summary
In existing security alarm, video surveillance and lighting systems, vibration signals, infrared trigger data and video monitoring data are processed independently, making it impossible to achieve time synchronization and spatial registration. This results in fragmented multi-dimensional intrusion monitoring information that cannot fully reflect the actual state of an intrusion event.
By acquiring multi-dimensional synchronous monitoring data streams, extracting time-domain and frequency-domain features, generating vibration event descriptors and infrared event records, and using a pre-trained moving target detection model to generate video moving target trajectory records, a unified spatiotemporal correlation framework is established for time synchronization and spatial registration, and finally a multi-source event fusion report is generated.
It enables refined feature extraction and information integration of vibration signals, infrared trigger data, and video monitoring data, forming a complete multi-source event fusion report, improving the consistency and accuracy of security monitoring, and supporting coordinated control decisions.
Smart Images

Figure CN122511005A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of security intelligent monitoring and linkage technology, specifically a multi-dimensional linkage control method for security alarm, video surveillance and lighting systems. Background Technology
[0002] Existing security alarm, video surveillance, and lighting monitoring systems generally adopt a distributed working mode of vibration fiber optic sensors, infrared intrusion detectors, and panoramic cameras. Vibration fiber optic sensors only make basic judgments on whether the collected vibration signals have been triggered, infrared intrusion detectors only output simple trigger signals and rough position information, and panoramic cameras independently complete the detection and trajectory recording of moving targets within the frame. The monitoring data of various front-end sensing devices are transmitted and processed independently, without any collaborative correlation mechanism between them, and the multi-source sensing data is in a decentralized operating state.
[0003] This type of decentralized security monitoring solution has obvious technical drawbacks. Vibration signals are not subjected to joint feature mining in the time and frequency domains, and can only achieve basic vibration trigger identification. It cannot form refined vibration event information that includes energy distribution, dominant frequency, and timing. Infrared trigger data and video monitoring data are only stored separately. Monitoring data from different sensing devices have problems such as time synchronization and spatial coordinate mismatch. Multi-dimensional intrusion monitoring information cannot be integrated, and the event information formed is fragmented and cannot fully reflect the actual state of the intrusion event.
[0004] It is necessary to extract time-domain and frequency-domain features from vibration signals and generate corresponding vibration event descriptors. At the same time, a unified spatiotemporal correlation carrier should be built to synchronize the time and register the three types of monitoring data (vibration, infrared, and video) in space, thereby generating event reports that integrate multi-source monitoring features and making up for the technical shortcomings of existing security monitoring systems. Summary of the Invention
[0005] This invention aims to solve at least one of the technical problems existing in the prior art; Therefore, this invention proposes a multi-dimensional linkage control method for security alarm, video surveillance, and lighting systems, comprising: Acquire a multidimensional synchronous monitoring data stream from the front-end sensing network. The multidimensional synchronous monitoring data stream includes vibration signal waveforms generated by a vibration fiber optic sensor, infrared trigger signals generated by an infrared intrusion detector, and panoramic video streams generated by a panoramic camera. The vibration signal waveform is subjected to time-domain and frequency-domain feature extraction to generate a vibration event descriptor that includes vibration energy distribution, vibration dominant frequency and vibration timing features; The infrared trigger signal is analyzed for spatial location and trigger timestamp to generate an infrared event record containing the initial coordinates of the intrusion target and the trigger time. The panoramic video stream is input into a pre-trained moving target detection model to identify and track moving targets in the video footage, generating a video moving target trajectory record containing a sequence of moving target trajectory coordinates. A unified spatiotemporal correlation framework is established to synchronize the vibration event descriptor, the infrared event record, and the video moving target trajectory record in time and spatially register them. Under the unified spatiotemporal correlation framework, a multi-source event fusion report is generated that integrates vibration features, infrared position, and video trajectory.
[0006] Furthermore, time-domain and frequency-domain features are extracted from the vibration signal waveform to generate a vibration event descriptor containing vibration energy distribution, vibration dominant frequency, and vibration timing features, including: The vibration signal waveform is segmented and windowed, and the root mean square value of the vibration signal within each time window is calculated. The sequence of the root mean square values on the time axis is used as the vibration energy distribution. Perform a fast Fourier transform on the vibration signal within each time window to obtain the spectral distribution of the vibration signal, identify the frequency component with the strongest energy from the spectral distribution, and record it as the vibration dominant frequency; The zero-crossing density and envelope shape of the vibration signal waveform in the time dimension are analyzed, and the temporal pattern features characterizing the vibration occurrence, duration and decay process are extracted as the vibration temporal features. The vibration energy distribution, the dominant vibration frequency, and the vibration timing characteristics are integrated, and the vibration fiber optic sensor number and its geographical location are added to report the vibration signal waveform, and encoded as the vibration event descriptor.
[0007] Further, the panoramic video stream is input into a pre-trained moving target detection model to identify and track moving targets in the video frame, generating a video moving target trajectory record containing a sequence of moving target trajectory coordinates, including: Video frames are extracted from the panoramic video stream at a fixed frame rate, and each extracted video frame is input into the pre-trained moving target detection model. The pre-trained moving target detection model is used to perform pixel-level analysis on each video frame, segment out the moving foreground regions that have effective differences from the background, and calculate the bounding box coordinates of each moving foreground region in the video frame coordinate system. Based on the similarity of the bounding box coordinates of the moving foreground regions between adjacent video frames, cross-frame association matching is performed on different moving foreground regions to form multiple independent moving target tracking chains; For each moving target tracking chain, the coordinates of its bounding box center point in each video frame are recorded, and the coordinates of the bounding box center points of each moving target tracking chain are arranged in chronological order to form the moving target trajectory coordinate sequence of the moving target; The sequence of coordinates of the moving targets' trajectories is collected, along with the panoramic camera number and geographical location, to generate the video moving target trajectory record.
[0008] Furthermore, the establishment of a unified spatiotemporal correlation framework, which synchronizes and spatially registers the vibration event descriptor, the infrared event record, and the video moving target trajectory record, includes: Establish a timeline based on the system's absolute time, and uniformly convert the time information in the vibration event descriptor, the trigger time point in the infrared event record, and the video frame timestamp in the video moving target trajectory record to the system's absolute time reference. Obtain and load the pre-stored spatial coordinate mapping table of the front-end sensing network devices. The spatial coordinate mapping table of the front-end sensing network devices records the geographical coordinate range of the monitoring segment of each vibration fiber optic sensor, the effective detection area coordinates of each infrared intrusion detector, and the geographical coordinate range of the monitoring field of view of each panoramic camera. Using the spatial coordinate mapping table of the front-end sensing network device, the sensor number in the vibration event descriptor is mapped to a geographic coordinate range, the preliminary coordinates in the infrared event record are mapped to geographic coordinates, and the coordinates of the center point of the bounding box in the video moving target trajectory record are mapped to a real-world geographic coordinate sequence. Under the absolute time reference and unified geographic coordinate system of the system, the converted vibration event descriptor, the infrared event record and the video moving target trajectory record are aligned to determine the temporal and spatial overlap relationship of events from different sources.
[0009] Furthermore, within the unified spatiotemporal correlation framework, a multi-source event fusion report is generated, integrating vibration characteristics, infrared location, and video trajectory, including: The vibration event descriptors, infrared event records, and video moving target trajectory records that overlap in both time window and geographic space are identified as related event clusters. The data from each source within the associated event cluster are fused using a confidence-weighted method. The confidence of the vibration event descriptor is determined based on the stability of its vibration energy distribution and dominant vibration frequency. The confidence of the infrared event record is determined based on its historical detection accuracy. The confidence of the video moving target trajectory record is determined based on the continuity and clarity of its target tracking. Based on the confidence weighting result, a fusion event is generated, which includes the fused target geographic coordinates, target movement speed, target movement direction, and target category probability distribution. Each fusion event is assigned a globally unique event identifier, and the types and numbers of all front-end sensing devices that trigger it are recorded; All generated fusion events and their detailed information are summarized to form the multi-source event fusion report.
[0010] Furthermore, the pre-trained moving target detection model is constructed in the following manner: We acquire a massive amount of labeled video data containing various scenes, lighting conditions, and moving targets to form the original training dataset; Each video frame in the original training dataset is preprocessed, including image size normalization, illumination normalization, and data augmentation, to obtain an augmented training dataset. A feature extraction backbone network based on a deep convolutional neural network is constructed, and a feature pyramid network for multi-scale feature fusion is constructed. The feature extraction backbone network and the feature pyramid network are connected to form the basic network architecture of the moving target detection model. On the basic network architecture, a pixel-level semantic segmentation head for foreground-background segmentation and a regression prediction head for generating moving target bounding boxes are added to form an initial untrained moving target detection model. The initial untrained moving target detection model is iteratively trained using the enhanced training dataset. The cross-entropy loss of foreground-background segmentation and the smooth L1 loss of bounding box regression are used as the joint loss function. The model parameters are optimized through the backpropagation algorithm until the model converges, thus obtaining the pre-trained moving target detection model.
[0011] Furthermore, the method also includes the step of generating linkage control instructions based on the multi-source event fusion report: The multi-source event fusion report is analyzed to extract the target geographic coordinates, target movement direction, and target category probability distribution for each fusion event; Based on the target geographic coordinates, the linkage device spatial configuration database is queried to determine all linkageable video surveillance devices and linkageable lighting devices within the preset influence range of the target geographic coordinates. The linkage device spatial configuration database records the pan-tilt control parameters of all cameras and the dimming and control interface parameters of all lighting fixtures. Based on the target's movement direction, select video surveillance devices from the linked video surveillance devices that cover the area in front of the target's movement direction and use them as preset camera positions; Based on the probability distribution of the target category, the corresponding response rules are matched from the preset linkage strategy rule base. The linkage strategy rule base defines the preset position number of the video surveillance equipment, the video recording resolution and frame rate, and the brightness level and on / off mode of the lighting equipment under different target categories and different confidence levels.
[0012] Further, based on the probability distribution of the target category, the corresponding response rule is matched from a preset linkage strategy rule base, including: From the target category probability distribution, select the target category with the highest probability value as the main discrimination category; Using the main discrimination category and the confidence level of the fusion event as the joint query key, a search is performed in the linkage strategy rule base; The linkage strategy rule base is for different combinations of the main discrimination category and confidence level, and a detailed set of device action parameters is preset. The set of device action parameters includes, but is not limited to: specified preset position number, video recording parameter combination, and lighting brightness curve; If a completely matching entry is found, the set of device action parameters defined by that entry is directly used; If no completely matching entry is found, a fuzzy matching algorithm based on category similarity and confidence proximity is used to select the most similar entry from the linkage strategy rule base, and the device action parameter set is used as the basis for adaptive adjustment to generate the final matching response rule.
[0013] Furthermore, the method also includes the steps of executing linkage control commands and verifying the linkage effect: Based on the matched response rules, a specific sequence of device control instructions is generated. The sequence of device control instructions includes: sending a pan-tilt rotation instruction containing a specified preset position number, a video parameter setting instruction, and a recording start instruction to the preset position camera; and sending a lighting control instruction containing a specified brightness level and on / off mode to the linked lighting device. The sequence of device control commands is sent to the underlying control interfaces of the corresponding video surveillance and lighting equipment. After sending the device control command sequence, an effect verification timer is started. After the effect verification timer reaches a preset delay time, the camera is called from the controlled preset position to obtain the latest real-time video stream. The latest real-time video stream is analyzed to detect whether there is a valid moving target in the video frame within a preset target area; if so, it is determined whether the position of the valid moving target matches the target geographic coordinates of the fused event. The analysis and judgment results are recorded as feedback on the linkage effect and stored in association with the original fusion event and the issued sequence of device control commands.
[0014] Furthermore, the latest real-time video stream is analyzed to detect whether there are any valid moving targets within a preset target area in the video frame, including: Extract one or more frames of images from the latest real-time video stream; Based on the target geographic coordinates of the fusion event and the spatial mapping relationship between the preset location camera, the preset target area corresponding to the video frame is determined; Foreground segmentation processing is performed on the image within the preset target area to calculate pixel change features and identify potential moving pixel blocks; Morphological filtering and connected component analysis are performed on the identified potential moving pixel blocks to eliminate minor interference areas caused by changes in lighting or noise. Features are extracted from the remaining connected components and compared with typical feature templates of the corresponding target category in the linkage strategy rule base. If the feature similarity exceeds the threshold, it is determined that there is a valid moving target.
[0015] Compared with the prior art, the beneficial effects of the present invention are: By performing time-domain and frequency-domain feature extraction on the vibration signal waveform generated by the vibration fiber optic sensor, a vibration event descriptor containing vibration energy distribution, vibration dominant frequency, and vibration timing characteristics can be directly generated. This process fully extracts the feature information of the vibration signal in different dimensions, refines the representation of vibration events, and transforms the vibration signal from a single trigger state into a quantitative description of multiple feature parameters. It fully preserves the detailed information of the vibration signal in waveform changes, energy distribution, frequency composition, and timing changes, making the representation of vibration events more comprehensive and specific, improving the refinement of vibration event information, and enabling vibration monitoring data to reflect the true characteristics of the signal itself, avoiding the loss of event information caused by relying solely on simple trigger judgments.
[0016] By establishing a unified spatiotemporal correlation framework and performing time synchronization and spatial registration processing on vibration event descriptors, infrared event records, and video moving target trajectory records, monitoring data from three different sources can be matched and integrated based on the same spatiotemporal reference. This eliminates spatiotemporal deviations caused by different front-end sensing devices during data acquisition and transmission, and organically integrates vibration characteristic information, infrared intrusion location and time information, and video moving target trajectory coordinate information to form a multi-source event fusion report that integrates multi-source monitoring data. This opens up information interaction channels between data from different sensing devices, changes the state of independent and unrelated monitoring data, and enables multi-dimensional security monitoring information to form a unified whole, fully presenting the correlation relationships of various monitoring data corresponding to intrusion events, and giving the event information obtained from security monitoring a coherent and unified characteristic. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating the steps of a multi-dimensional linkage control method for a security alarm, video surveillance, and lighting system according to the present invention. Figure 2 A flowchart for generating vibration event descriptors; Figure 3 A flowchart for establishing a unified spatiotemporal correlation framework. Detailed Implementation
[0018] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] See Figure 1This invention provides a multi-dimensional linkage control method for security alarm, video surveillance, and lighting systems. The method includes: acquiring a multi-dimensional synchronous monitoring data stream from a front-end sensing network. This data stream synchronously includes vibration signal waveforms generated by a vibration fiber optic sensor, infrared trigger signals generated by an infrared intrusion detector, and a panoramic video stream generated by a panoramic camera. Then, time-domain and frequency-domain features are extracted from the vibration signal waveforms to generate a vibration event descriptor containing vibration energy distribution, vibration dominant frequency, and vibration timing characteristics. The spatial location and trigger timestamp of the infrared trigger signal are analyzed to generate an infrared event record containing the initial coordinates of the intrusion target and the trigger time point. The panoramic video stream is input into a pre-trained moving target detection model to identify and track moving targets in the video frame, generating a video moving target trajectory record containing the trajectory coordinate sequence of the moving target. A unified spatiotemporal correlation framework is established to synchronize and spatially register the aforementioned vibration event descriptor, infrared event record, and video moving target trajectory record. Under this framework, a multi-source event fusion report integrating vibration features, infrared location, and video trajectory is generated, providing core basis for subsequent linkage control decisions.
[0020] In one embodiment of the present invention, time-domain and frequency-domain features are extracted from the vibration signal waveform to generate a vibration event descriptor containing vibration energy distribution, vibration dominant frequency, and vibration timing features. The specific process includes: (See [reference]). Figure 2 The vibration signal waveform is segmented and windowed. The root mean square (RMS) value of the vibration signal within each time window is calculated, and the sequence of these RMS values on the time axis is used as the vibration energy distribution. A Fast Fourier Transform (FFT) is performed on the vibration signal within each time window to obtain its spectral distribution. The frequency component with the strongest energy is identified from this spectral distribution and recorded as the dominant vibration frequency. The zero-crossing density and envelope shape of the vibration signal waveform in the time dimension are analyzed to extract temporal pattern features characterizing the vibration occurrence, duration, and decay process, which are then used as vibration temporal features. The vibration energy distribution, dominant vibration frequency, and vibration temporal features are integrated, along with the vibration fiber optic sensor number and its geographical location, and encoded as a vibration event descriptor.
[0021] The process involves inputting a panoramic video stream into a pre-trained moving target detection model to identify and track moving targets in the video frame, generating a video moving target trajectory record containing a sequence of moving target trajectory coordinates. Specifically, this process includes: extracting video frames from the panoramic video stream at a fixed frame rate; inputting each extracted video frame into the pre-trained moving target detection model; performing pixel-level analysis on each video frame using the pre-trained model to segment moving foreground regions that have effective differences from the background, and calculating the bounding box coordinates of each moving foreground region in the video frame coordinate system; performing cross-frame association matching on different moving foreground regions based on the similarity of their bounding box coordinates between adjacent video frames to form multiple independent moving target tracking chains; recording the bounding box center point coordinates of each moving target tracking chain in each video frame, and arranging these coordinates chronologically to form a moving target trajectory coordinate sequence; and finally, compiling all the moving target trajectory coordinate sequences, attaching the panoramic camera number and geographical location, to generate a video moving target trajectory record. The vibration signal waveform is generated by a vibration fiber optic sensor deployed in the protected area. In practice, the system receives the raw voltage signal sequence from the vibration fiber optic sensor, which constitutes the vibration signal waveform. The first step in extracting time and frequency domain features from the vibration signal waveform is segmented windowing. In practice, the system uses a Hanning window to overlap and segment the continuous time signal, with each time window having a fixed length of 100 milliseconds and an overlap ratio of 50%. Next, the root mean square (RMS) value of the vibration signal within each time window is calculated. This calculation involves squaring, averaging, and then taking the square root of the voltage values at the sampling points within the time window. Arranging the RMS values calculated for all time windows in their corresponding time order forms the sequence of vibration energy distribution.
[0022] After obtaining the vibration energy distribution, a Fast Fourier Transform (FFT) needs to be performed on the vibration signal within each time window. In some embodiments, the system calls the FFT library to execute the FFT algorithm to obtain the spectral distribution of the vibration signal within that time window. The process of identifying the strongest frequency component from the spectral distribution involves the system scanning the amplitudes of all frequency components in the spectrum and recording the frequency corresponding to the point of maximum amplitude. This frequency is recorded as the dominant vibration frequency for that time window. The zero-crossing density and envelope shape of the vibration signal waveform in the time dimension are analyzed to extract vibration timing features. In specific implementations, the zero-crossing density is calculated by counting the number of times the signal crosses the zero level per unit time, and the signal envelope is obtained through a Hilbert transform, which describes the shape characteristics of the envelope. The vibration energy distribution sequence obtained in the aforementioned steps, the dominant vibration frequency of each time window, and the vibration timing features represented by the zero-crossing density and envelope shape are packaged together with the unique number of the vibration fiber optic sensor that reported the vibration signal waveform and its geographical coordinates, and encoded into a structured vibration event descriptor.
[0023] The process of generating video motion target trajectory records from the panoramic video stream is performed in parallel with vibration signal processing. In specific implementations, the panoramic camera outputs a panoramic video stream at a rate of 25 frames per second. The system first extracts video frames from the panoramic video stream at a fixed frame rate, typically consistent with the camera's original frame rate of 25 frames per second. Each extracted video frame is immediately input into a pre-trained motion target detection model. The pre-trained motion target detection model performs pixel-level analysis on each input video frame. Optionally, this model is based on an encoder-decoder architecture and can segment moving foreground regions that have effective differences from the background model. The model output includes the bounding box coordinates of each moving foreground region in the video frame coordinate system. The bounding box coordinates are in pixels and are typically represented by the coordinates of the top-left corner, as well as the width and height of the bounding box.
[0024] Based on the similarity of the bounding box coordinates of moving foreground regions between adjacent video frames, cross-frame association matching is performed on different moving foreground regions to form multiple independent moving target tracking chains. In specific implementation, the similarity is calculated using a weighted sum of the intersection-union ratio (IU) algorithm and the distance to appearance features. (IU) The calculation formula is: in: and These represent the bounding box areas of two moving foreground regions in adjacent frames. It is the area of intersection. It is the area of the union. The coordinates of the bounding box center point of each moving target tracking chain in each video frame can be recorded. The bounding box center point coordinates are obtained by adding half the width and height to the coordinates of the top-left corner of the bounding box. Arranging the bounding box center point coordinates of the same moving target tracking chain in different video frames according to the timestamp sequence constitutes the moving target trajectory coordinate sequence for that moving target.
[0025] In one embodiment of the present invention, a unified spatiotemporal correlation framework is established to synchronize and spatially register vibration event descriptors, infrared event records, and video moving target trajectory records. The specific process includes: (See [reference]). Figure 3 A timeline based on the system's absolute time is established, uniformly converting the time information in vibration event descriptors, trigger timestamps in infrared event records, and video frame timestamps in video moving target trajectory records to the system's absolute time reference. A pre-stored spatial coordinate mapping table for front-end sensing network devices is acquired and loaded. This table records the geographic coordinate range of the monitoring segment for each vibration fiber optic sensor, the effective detection area coordinates for each infrared intrusion detector, and the geographic coordinate range of the monitoring field of view for each panoramic camera. Using this mapping table, sensor numbers in vibration event descriptors are mapped to geographic coordinate ranges, preliminary coordinates in infrared event records are mapped to geographic coordinates, and the center point coordinates of the bounding box in video moving target trajectory records are mapped to a sequence of real-world geographic coordinates. Under the system's absolute time reference and the unified geographic coordinate system, the converted vibration event descriptors, infrared event records, and video moving target trajectory records are aligned to determine the temporal and spatial overlap of events from different sources.
[0026] Within a unified spatiotemporal correlation framework, a multi-source event fusion report is generated, integrating vibration features, infrared location data, and video trajectories. The specific process includes: identifying clusters of vibration event descriptors, infrared event records, and video moving target trajectory records that overlap within both the time window and geographic space. Confidence-weighted fusion is then performed on the data from each source within each cluster. The confidence level of the vibration event descriptor is determined based on its vibration energy distribution and the stability of its dominant vibration frequency; the confidence level of the infrared event record is determined based on its historical detection accuracy; and the confidence level of the video moving target trajectory record is determined based on the continuity and clarity of its target tracking. Based on the confidence-weighted results, a fused event is generated, containing the fused target geographic coordinates, target movement speed, target movement direction, and target category probability distribution. Each fused event is assigned a globally unique event identifier, and the types and numbers of all front-end sensing devices that triggered it are recorded. Finally, all generated fused events and their detailed information are summarized to form a multi-source event fusion report. Establishing a unified spatiotemporal correlation framework first requires establishing a timeline based on the system's absolute time. In practice, the system uses the time provided by a network time protocol server as the system's absolute time reference. The time information in vibration event descriptors, the trigger timestamps in infrared event records, and the video frame timestamps in video moving target trajectory records are all uniformly converted to the system's absolute time reference. In practice, the conversion process involves time zone compensation and clock deviation correction for timestamps from different devices. Ultimately, all time information is represented in Coordinated Universal Time (UTC) format.
[0027] The process of acquiring and loading the pre-stored spatial coordinate mapping table of front-end sensing network devices is as follows: In specific implementations, the spatial coordinate mapping table of front-end sensing network devices is stored in the system configuration management module as a database table. The spatial coordinate mapping table records the geographical coordinate range of the monitoring segment for each vibration fiber optic sensor. In specific implementations, the geographical coordinate range of the monitoring segment for the vibration fiber optic sensor is described by a polyline range formed by connecting multiple latitude and longitude coordinate points. The spatial coordinate mapping table also records the effective detection area coordinates for each infrared intrusion detector. The spatial coordinate mapping table is used to map information from different data sources to a unified geographical coordinate system. In specific implementations, mapping the sensor number in the vibration event descriptor to the geographical coordinate range is achieved by querying the spatial coordinate mapping table to find the geographical coordinate range of the monitoring segment corresponding to the sensor number. Mapping the preliminary coordinates in the infrared event record to geographical coordinates is achieved by querying the spatial coordinate mapping table to find the device number that triggered the infrared intrusion detector, and combining its effective detection area coordinates with the preliminary coordinates to finally resolve a latitude and longitude point coordinate. Map the coordinates of the center point of the bounding box in the video moving target trajectory record to a sequence of real-world geographic coordinates.
[0028] Under the system's absolute time reference and a unified geographic coordinate system, the converted vibration event descriptors, infrared event records, and video moving target trajectory records are aligned to determine the temporal and spatial overlap of events from different sources. In practice, temporal overlap is determined by setting a time synchronization window threshold ΔT. If the absolute value of the difference between the timestamps of events recorded from different data sources is less than ΔT, they are considered to overlap temporally. Spatial overlap involves geometric calculations, such as determining whether the geographic coordinates of the infrared event fall within the geographic coordinate range of the vibration event, or whether the geographic coordinate sequence of the video trajectory intersects with the coordinate range of the vibration event. If both temporal and spatial overlap conditions are met, these events are considered to have a spatiotemporal correlation.
[0029] The first step in generating a multi-source event fusion report within a unified spatiotemporal correlation framework is to identify clusters of related events—including vibration event descriptors, infrared event records, and video moving target trajectory records—that overlap within both the time window and geographic space. Optionally, the identification of related event clusters can be based on spatiotemporal proximity. The calculation formula is: in: and These are the timestamps of the two events. It is the spatial distance between the two. and These are the preset maximum allowable time difference and spatial distance. and These are time weights and spatial weights, and The data from each source within the associated event cluster are fused using a confidence-weighted method. The confidence score of a vibration event descriptor is determined based on its vibration energy distribution and the stability of its dominant vibration frequency. In some embodiments, the variance of the vibration energy distribution within a time window and the drift of the dominant vibration frequency are used to calculate the confidence score of the vibration event. The confidence score of infrared event records is determined based on their historical detection accuracy. In specific implementations, the system records the number of false alarms and missed alarms for each infrared intrusion detector and calculates a long-term statistical confidence value based on this. The confidence score of video moving target trajectory records is determined based on the continuity and sharpness of target tracking.
[0030] A fused event is generated based on the confidence-weighted results. This fused event includes the fused target geographic coordinates, target movement speed, target movement direction, and target category probability distribution. In practice, the target geographic coordinates are estimated using a weighted least squares method, with the confidence levels of each source data as weights, by fusing the center points of their respective coordinates or coordinate ranges. The target movement speed and direction are primarily calculated based on the continuous coordinate sequences recorded in the video tracking of the moving target. The target category probability distribution is obtained by weighted fusion of the classification probabilities given by the vibration event, infrared event, and video analysis model. Each fused event is assigned a globally unique event identifier, which can be a combination of timestamp, geographic location hash, and random number. The types and numbers of all front-end sensing devices triggered by the fused event are also recorded. These front-end sensing devices include vibration fiber optic sensors, infrared intrusion detectors, and panoramic cameras. All generated fused events and their detailed information, including but not limited to event identifiers, a list of triggering devices, and fused attributes, are summarized to form a final multi-source event fusion report.
[0031] In one embodiment of the present invention, the pre-trained moving target detection model is constructed as follows: A massive amount of labeled video data containing various scenes, lighting conditions, and moving targets is acquired to form the original training dataset. Each video frame in the original training dataset is preprocessed, including image size normalization, illumination normalization, and data augmentation, to obtain an augmented training dataset. A feature extraction backbone network based on a deep convolutional neural network is constructed, and a feature pyramid network for multi-scale feature fusion is also constructed. The feature extraction backbone network and the feature pyramid network are connected to form the basic network architecture of the moving target detection model. On the basic network architecture, a pixel-level semantic segmentation head for foreground / background segmentation and a regression prediction head for generating moving target bounding boxes are added to form an initial untrained moving target detection model. The initial untrained moving target detection model is iteratively trained using the augmented training dataset. The cross-entropy loss for foreground / background segmentation and the smooth L1 loss for bounding box regression are used as the joint loss function. The model parameters are optimized through backpropagation until the model converges, resulting in the pre-trained moving target detection model.
[0032] The original training dataset is constructed from a massive amount of annotated video data encompassing various scenes, lighting conditions, and moving targets. In practice, this dataset is derived from both publicly available datasets and self-collected datasets. The publicly available datasets include annotated videos of city streets, park perimeters, and indoor corridors, while the self-collected datasets supplement this with video clips taken at dusk, at night, and in rainy / foggy weather. The annotation information for each video frame in the original training dataset includes pixel-level foreground masks and bounding box coordinates for moving targets. In practice, the pixel-level foreground masks are stored as binary images, and the bounding box coordinates are stored as coordinate pairs of the top-left and bottom-right corners. Each video frame in the original training dataset undergoes preprocessing, including image size normalization, illumination normalization, and data augmentation. Image size normalization scales or crops input video frames of different resolutions to a uniform size; in practice, this is 640 pixels wide and 480 pixels high. Illumination normalization reduces the impact of varying lighting conditions by applying global histogram equalization or local contrast enhancement to the pixel values of the video frames. Data augmentation expands the training samples by applying random transformations to video frames. Optional data augmentation operations include random horizontal flipping, random brightness and contrast fine-tuning, and random addition of salt-and-pepper noise. After the above preprocessing steps, an augmented training dataset is obtained, which is organized as a mapping list of image file paths and corresponding annotation files.
[0033] A feature extraction backbone network based on a deep convolutional neural network (CNN) is constructed. In specific implementations, the CNN can be either a residual network or a visual transformer architecture. A feature pyramid network for multi-scale feature fusion is constructed. In specific implementations, the feature pyramid network receives feature maps from different depths of the feature extraction backbone network and fuses them through top-down and lateral connection paths to generate a set of feature pyramid levels with rich semantic information and decreasing resolution. The feature extraction backbone network and the feature pyramid network are connected to form the basic network architecture of the moving target detection model. The connection method is to use the outputs of multiple intermediate layers of the feature extraction backbone network as inputs to the feature pyramid network. A pixel-level semantic segmentation head for foreground and background segmentation is added to the basic network architecture. The pixel-level semantic segmentation head is usually composed of one or more convolutional layers and upsampling layers, and the final output is a probability map of each pixel belonging to the foreground or background with the same spatial resolution as the input image. A regression prediction head for generating moving target bounding boxes is added to the basic network architecture. The regression prediction head predicts the coordinate offset and size scaling factor of a bounding box for each position of the feature pyramid network output. The pixel-level semantic segmentation head and regression prediction head share features with the basic network architecture to form an initial untrained moving target detection model. The structural parameters of the initial untrained moving target detection model are shown in Table 1.
[0034] Table 1: Structural parameters of the initial untrained moving target detection model The initial untrained moving object detection model was iteratively trained using an augmented training dataset, with the cross-entropy loss for foreground / background segmentation and the smoothed L1 loss for bounding box regression used as the joint loss function. (Cross-entropy loss function for foreground / background segmentation) The cross-entropy is defined as the cross-entropy between the class probability distribution of each pixel predicted by the model and the pixel-level foreground mask of the ground truth annotation. The smoothed L1 loss function for bounding box regression. The joint loss function is defined as the smoothed L1 norm distance between the bounding box coordinates predicted by the model and the coordinates of the ground truth bounding box. It is a weighted sum of the two, and its calculation formula is: in: Indicates the partition loss. Indicates regression loss, This is a weight coefficient between 0 and 1, used to balance the contributions of the two losses to the total loss. The model parameters are optimized using the backpropagation algorithm. In practice, the backpropagation algorithm calculates the gradient based on the joint loss function and updates all weights and bias parameters in the model using stochastic gradient descent or the Adam optimizer. Iterative training continues, with each iteration using a batch of augmented training data; the batch size can be set to 32. During iterative training, the model's performance on independent validation sets is monitored until the model converges. The convergence criterion is typically that the joint loss value on the validation set no longer significantly decreases over multiple consecutive iterations. Finally, the trained model parameters are saved, resulting in a pre-trained moving object detection model. This pre-trained model can be understood as having the ability to segment moving foregrounds and locate their bounding boxes from a new panoramic video stream.
[0035] In one embodiment of the present invention, a multi-source event fusion report is parsed to extract the target geographic coordinates, target movement direction, and target category probability distribution for each fusion event. Based on the target geographic coordinates, a linkage device spatial configuration database is queried to determine all linkageable video surveillance devices and linkageable lighting devices within the preset influence range of the target geographic coordinates. The linkage device spatial configuration database records the pan-tilt control parameters of all cameras and the dimming and control interface parameters of all lighting fixtures. Based on the target movement direction, video surveillance devices whose monitoring field of view covers the area in front of the target movement direction are selected from the linkageable video surveillance devices and used as preset camera positions. Based on the target category probability distribution, corresponding response rules are matched from a preset linkage strategy rule base. The linkage strategy rule base defines the preset position numbers of video surveillance devices, video recording resolution and frame rate, and brightness levels and activation modes of lighting devices for different target categories and confidence levels.
[0036] From the target category probability distribution, the target category with the highest probability value is selected as the primary discriminant category. The primary discriminant category and the confidence level of the fused event are used as the joint query key to search the linkage strategy rule base. The linkage strategy rule base predefines detailed sets of device action parameters for different combinations of primary discriminant categories and confidence levels. These device action parameter sets include, but are not limited to, specified preset position numbers, video recording parameter combinations, and lighting brightness curves. If a completely matching entry is found, the device action parameter set defined for that entry is directly adopted. If no completely matching entry is found, a fuzzy matching algorithm based on category similarity and confidence proximity is used to select the closest entry from the linkage strategy rule base, and its device action parameter set is used as the basis for adaptive adjustments to generate the final matching response rule.
[0037] The multi-source event fusion report is parsed to extract the target geographic coordinates, target movement direction, and target category probability distribution for each fused event. In practice, the multi-source event fusion report is stored in JSON format, and the parsing operation is completed by reading the "target geographic coordinates," "target movement direction," and "target category probability distribution" fields from the JSON structure. Based on the target geographic coordinates, the spatial configuration database of the linked devices is queried to determine all linked video surveillance devices and linked lighting devices within the preset influence range of the target geographic coordinates. The spatial configuration database of the linked devices records the pan / tilt control parameters of all cameras and the dimming and control interface parameters of all lighting fixtures. In practice, the spatial configuration database of the linked devices uses a spatial index structure, such as an R-tree, to accelerate the retrieval of devices within the preset influence range of the target geographic coordinates. The preset influence range can be a circular area with a radius of 50 meters centered on the target geographic coordinates.
[0038] Based on the target's movement direction, video surveillance devices whose monitoring field of view covers the area in front of the target's movement direction are selected from the linked video surveillance devices and used as preset cameras. The selection logic calculates the angle between the target's movement direction vector and the camera's optical axis direction vector on the horizontal plane, and simultaneously determines whether the target's geographical coordinates are within the camera's monitoring field of view. In some embodiments, if the angle is less than a preset threshold and the target's geographical coordinates are within the monitoring range, the camera is determined to be a preset camera. Based on the target category probability distribution, corresponding response rules are matched from a preset linkage strategy rule base. The linkage strategy rule base defines the preset camera numbers, video recording resolution and frame rate, and brightness levels and activation modes of lighting devices for different target categories and confidence levels.
[0039] The target category with the highest probability value is selected as the primary discriminant category from the target category probability distribution. The target category probability distribution is a vector, where each element represents the probability of the corresponding target category. The primary discriminant category and the confidence level of the fused event are used as the joint query key to retrieve data from the linkage strategy rule base. In practice, the confidence level of the fused event is a comprehensive score calculated when generating the multi-source event fusion report. The linkage strategy rule base predefines detailed sets of device action parameters for different combinations of primary discriminant categories and confidence levels. These parameters include, but are not limited to, specified preset position numbers, video recording parameter combinations, and lighting brightness curves. The lighting brightness curve defines the brightness change pattern of the lighting equipment over a period of time after startup. If a completely matching entry is found, the device action parameter set defined in the linkage strategy rule base entry is directly adopted. A complete match means that there exists a record in the linkage strategy rule base whose "target category" field is completely consistent with the primary discriminant category, and whose "confidence level" field defines a confidence range that includes the confidence value of the fused event. If no perfectly matching entry is found, a fuzzy matching algorithm based on category similarity and confidence proximity is used to select the closest entry from the linkage strategy rule base. The algorithm then adaptively adjusts the entry based on its device action parameter set to generate the final matching response rule. Category similarity can be calculated using distance in the category semantic hierarchy or cosine similarity of category feature vectors. Confidence proximity is defined as the reciprocal of the absolute difference between the fused event confidence and the midpoint of the confidence interval corresponding to the entry in the linkage strategy rule base. The fuzzy matching algorithm searches for the entry with the highest comprehensive score. In some embodiments, the comprehensive score... The calculation formula is: in: It is the primary discriminant category. It is a category in the linkage strategy rule base. It is a function for calculating category similarity. It is the confidence level of the fused event. It is the reference confidence level of the entries in the linkage strategy rule base. It is the confidence proximity calculation function. and It is a weighting coefficient, and The fuzzy matching process, as shown in Table 2, illustrates the matching process in the linkage strategy rule base for a fusion event (primary discrimination category is "vehicle", confidence level is 0.85).
[0040] Table 2: Fuzzy Matching Algorithm Selection Table Based on the examples in Table 2, item A has a high similarity to the target category "vehicle" and a confidence interval that includes 0.85, indicating high proximity, thus it has the highest overall score. Item C, although a perfect category match, has a confidence interval of 0.85 that does not fall within its interval [0.9, 1.0], indicating low proximity. Item B has both low category similarity and low confidence proximity. The fuzzy matching algorithm will select item A as the closest item. It can be understood that after selecting the set of device action parameters for item A, its parameters need to be adaptively adjusted. For example, if the lighting brightness curve of item A is the maximum brightness set for "motor vehicles," while the target category is "vehicles," it may be directly adopted or slightly adjusted to ultimately generate a response rule that matches the current fusion event.
[0041] In one embodiment of the present invention, a specific device control command sequence is generated according to the matched response rules. This sequence includes: sending a pan-tilt-zoom command containing a specified preset position number, a video parameter setting command, and a recording start command to a preset position camera; and sending a lighting control command containing a specified brightness level and activation mode to a linked lighting device. The device control command sequence is sent to the underlying control interfaces of the corresponding video surveillance device and lighting device. After sending the device control command sequence, an effect verification timer is started. After the effect verification timer reaches a preset delay, the latest real-time video stream is obtained from the controlled preset position camera. The obtained latest real-time video stream is analyzed to detect whether there is a valid moving target within a preset target area in the video frame. If so, it is determined whether the position of the valid moving target matches the target geographic coordinates of the fusion event. The analysis and judgment results are recorded as a linkage effect feedback record and stored in association with the original fusion event and the issued device control command sequence.
[0042] The latest real-time video stream is analyzed to detect whether there are valid moving targets within a preset target area. The specific process includes: extracting one or more frames from the latest real-time video stream; determining the corresponding preset target area in the video frame based on the target's geographic coordinates and the spatial mapping relationship between preset location cameras based on the fused event; performing foreground segmentation on the image within the preset target area, calculating pixel change features, and identifying potential moving pixel blocks; performing morphological filtering and connected component analysis on the identified potential moving pixel blocks to eliminate minor interference areas caused by lighting changes or noise; extracting features from the remaining connected components and comparing them with typical feature templates of the corresponding target category in the linkage strategy rule base; if the feature similarity exceeds a threshold, it is determined that a valid moving target exists.
[0043] Based on the matched response rules, a specific sequence of device control commands is generated. In practice, this sequence includes: sending a pan-tilt-zoom command containing a specified preset position number, a video parameter setting command, and a recording start command to a preset position calling camera; and sending a lighting control command containing a specified brightness level and activation mode to a linked lighting device. The pan-tilt-zoom command follows the Pelco-D protocol, specifying the target camera's address, the preset position calling command, and the specific preset position number, such as "FF01002F000030". The video parameter setting command is sent via the media service interface in the ONVIF protocol, setting the video encoding format to H.264, the resolution to 1920x1080, and the frame rate to 25 frames per second. The lighting control command is sent via the DALI or DMX512 protocol, containing the target lighting device's address, brightness level value, and activation mode, which can be instantaneous or gradual illumination.
[0044] The system sends a sequence of device control commands to the underlying control interfaces of the corresponding video surveillance and lighting equipment. In practice, the system uses Ethernet or RS-485 bus to send pan-tilt-zoom (PTZ) commands and video parameter setting commands to the control module built into the preset camera, and sends lighting control commands to the driver controller of the lighting equipment. After sending the sequence of device control commands, an effect verification timer is started. This timer is a software timer, and in practice, the preset delay time is set to 2 seconds to allow time for the pan-tilt-zoom, lens focusing, and lighting equipment stabilization to complete. After the effect verification timer reaches the preset delay time, the system retrieves the latest real-time video stream from the controlled preset camera. The real-time video stream is pulled from the camera via RTSP and includes a short segment of video data after the preset delay time, such as the subsequent 5 seconds of video.
[0045] The latest real-time video stream is analyzed to detect whether a valid moving target exists within a preset target area. In practice, if a valid moving target exists, it is determined whether the position of the valid moving target matches the target geographic coordinates of the fused event. The preset target area is determined based on the spatial mapping relationship between the target geographic coordinates of the fused event and the preset position calling camera. In practice, using the calibration parameters and installation posture of the preset position calling camera, the target geographic coordinates (longitude and latitude) of the fused event are back-projected onto the camera's image coordinate system, thereby defining a square area in the video frame centered at the projection point, with a side length one-tenth of the image width. This square area is the preset target area. It can be understood that determining whether the position matches involves checking whether the center coordinates of the bounding box of a valid moving target in the current video frame fall within this preset target area.
[0046] One or more frames are extracted from the latest real-time video stream for analysis. In some embodiments, the first frame of the real-time video stream after a preset delay is extracted as the analysis object. Foreground segmentation is performed on the image within a preset target area, pixel change features are calculated, and potential moving pixel blocks are identified. In specific implementations, the foreground segmentation process uses a Gaussian mixture model background modeling method, comparing the current frame with the background model. Pixels whose pixel values change beyond a set threshold are marked as foreground pixels, and the connected regions of the foreground pixels are the potential moving pixel blocks. Morphological filtering and connected component analysis are performed on the identified potential moving pixel blocks to eliminate minor interference regions caused by lighting changes or noise. Morphological filtering includes an opening operation followed by erosion and dilation to smooth the boundaries of the foreground region and remove small isolated noise points. Features are extracted from the remaining connected components, including the area, aspect ratio, and contour complexity of the connected components. These features are compared with typical feature templates of the corresponding target category in the linkage strategy rule base. The typical feature templates store feature intervals such as area range and aspect ratio range for different target categories. Feature similarity is then calculated. The distance is calculated by comparing the extracted feature vector with the template feature vector. The formula is as follows: in: It is a feature vector extracted from the connected component. These are typical feature vectors of the target category. Let L2 norm represent the vector. If feature similarity... If the threshold is exceeded, it is determined that there is a valid moving target.
[0047] The analysis and judgment results are recorded as feedback on the linkage effect, and stored in association with the original fusion event and the issued device control command sequence. In specific implementation, the linkage effect feedback record includes a boolean field "Whether the linkage is effective" and a text field "Analysis and judgment description". The original fusion event identifier, the detailed content of the issued device control command sequence, and the linkage effect feedback record are stored together in the same transaction record in the log database. This association storage facilitates subsequent auditing and optimization of the linkage strategy's effectiveness.
[0048] The above embodiments are only used to illustrate the technical methods of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of the present invention without departing from the spirit and scope of the technical methods of the present invention.
Claims
1. A multi-dimensional linkage control method of a security alarm and video monitoring and lighting system, characterized in that, The method includes: Acquire a multidimensional synchronous monitoring data stream from the front-end sensing network. The multidimensional synchronous monitoring data stream includes vibration signal waveforms generated by a vibration fiber optic sensor, infrared trigger signals generated by an infrared intrusion detector, and panoramic video streams generated by a panoramic camera. The vibration signal waveform is subjected to time-domain and frequency-domain feature extraction to generate a vibration event descriptor that includes vibration energy distribution, vibration dominant frequency and vibration timing features; The infrared trigger signal is analyzed for spatial location and trigger timestamp to generate an infrared event record containing the initial coordinates of the intrusion target and the trigger time. The panoramic video stream is input into a pre-trained moving target detection model to identify and track moving targets in the video footage, generating a video moving target trajectory record containing a sequence of moving target trajectory coordinates. A unified spatiotemporal correlation framework is established to synchronize the vibration event descriptor, the infrared event record, and the video moving target trajectory record in time and spatially register them. Under the unified spatiotemporal correlation framework, a multi-source event fusion report is generated that integrates vibration features, infrared position, and video trajectory.
2. The multi-dimensional linkage control method of the security alarm and video monitoring and lighting system according to claim 1, characterized in that, The vibration signal waveform is subjected to time-domain and frequency-domain feature extraction to generate a vibration event descriptor containing vibration energy distribution, vibration dominant frequency, and vibration timing features, including: The vibration signal waveform is segmented and windowed, and the root mean square value of the vibration signal within each time window is calculated. The sequence of the root mean square values on the time axis is used as the vibration energy distribution. Perform a fast Fourier transform on the vibration signal within each time window to obtain the spectral distribution of the vibration signal, identify the frequency component with the strongest energy from the spectral distribution, and record it as the vibration dominant frequency; The zero-crossing density and envelope shape of the vibration signal waveform in the time dimension are analyzed, and the temporal pattern features characterizing the vibration occurrence, duration and decay process are extracted as the vibration temporal features. The vibration energy distribution, the dominant vibration frequency, and the vibration timing characteristics are integrated, and the vibration fiber optic sensor number and its geographical location are added to report the vibration signal waveform, and encoded as the vibration event descriptor.
3. The multi-dimensional linkage control method of the security alarm and video monitoring and lighting system according to claim 1, characterized in that, The panoramic video stream is input into a pre-trained moving target detection model to identify and track moving targets in the video frame, generating a video moving target trajectory record containing a sequence of moving target trajectory coordinates, including: Video frames are extracted from the panoramic video stream at a fixed frame rate, and each extracted video frame is input into the pre-trained moving target detection model. The pre-trained moving target detection model is used to perform pixel-level analysis on each video frame, segment out the moving foreground regions that have effective differences from the background, and calculate the bounding box coordinates of each moving foreground region in the video frame coordinate system. Based on the similarity of the bounding box coordinates of the moving foreground regions between adjacent video frames, cross-frame association matching is performed on different moving foreground regions to form multiple independent moving target tracking chains; For each moving target tracking chain, the coordinates of its bounding box center point in each video frame are recorded, and the coordinates of the bounding box center points of each moving target tracking chain are arranged in chronological order to form the moving target trajectory coordinate sequence of the moving target; The sequence of coordinates of the moving targets' trajectories is collected, along with the panoramic camera number and geographical location, to generate the video moving target trajectory record.
4. The multi-dimensional linkage control method of the security alarm and video monitoring and lighting system according to claim 1, characterized in that, The establishment of a unified spatiotemporal correlation framework, which synchronizes and spatially registers the vibration event descriptors, the infrared event records, and the video moving target trajectory records, includes: Establish a timeline based on the system's absolute time, and uniformly convert the time information in the vibration event descriptor, the trigger time point in the infrared event record, and the video frame timestamp in the video moving target trajectory record to the system's absolute time reference. Obtain and load the pre-stored spatial coordinate mapping table of the front-end sensing network devices. The spatial coordinate mapping table of the front-end sensing network devices records the geographical coordinate range of the monitoring segment of each vibration fiber optic sensor, the effective detection area coordinates of each infrared intrusion detector, and the geographical coordinate range of the monitoring field of view of each panoramic camera. Using the spatial coordinate mapping table of the front-end sensing network device, the sensor number in the vibration event descriptor is mapped to a geographic coordinate range, the preliminary coordinates in the infrared event record are mapped to geographic coordinates, and the coordinates of the center point of the bounding box in the video moving target trajectory record are mapped to a real-world geographic coordinate sequence. Under the absolute time reference and unified geographic coordinate system of the system, the converted vibration event descriptor, the infrared event record and the video moving target trajectory record are aligned to determine the temporal and spatial overlap relationship of events from different sources.
5. The multi-dimensional linkage control method of the security alarm and video monitoring and lighting system according to claim 4, characterized in that, Within the unified spatiotemporal correlation framework, a multi-source event fusion report is generated, integrating vibration characteristics, infrared location, and video trajectory, including: The vibration event descriptors, infrared event records, and video moving target trajectory records that overlap in both time window and geographic space are identified as related event clusters. The data from each source within the associated event cluster are fused using a confidence-weighted method. The confidence of the vibration event descriptor is determined based on the stability of its vibration energy distribution and dominant vibration frequency. The confidence of the infrared event record is determined based on its historical detection accuracy. The confidence of the video moving target trajectory record is determined based on the continuity and clarity of its target tracking. Based on the confidence weighting result, a fusion event is generated, which includes the fused target geographic coordinates, target movement speed, target movement direction, and target category probability distribution. Each fusion event is assigned a globally unique event identifier, and the types and numbers of all front-end sensing devices that trigger it are recorded; All generated fusion events and their detailed information are summarized to form the multi-source event fusion report.
6. The multi-dimensional linkage control method for a security alarm, video surveillance, and lighting system according to claim 1, characterized in that, The pre-trained moving target detection model is constructed in the following way: We acquire a massive amount of labeled video data containing various scenes, lighting conditions, and moving targets to form the original training dataset; Each video frame in the original training dataset is preprocessed, including image size normalization, illumination normalization, and data augmentation, to obtain an augmented training dataset. A feature extraction backbone network based on a deep convolutional neural network is constructed, and a feature pyramid network for multi-scale feature fusion is constructed. The feature extraction backbone network and the feature pyramid network are connected to form the basic network architecture of the moving target detection model. On the basic network architecture, a pixel-level semantic segmentation head for foreground-background segmentation and a regression prediction head for generating moving target bounding boxes are added to form an initial untrained moving target detection model. The initial untrained moving target detection model is iteratively trained using the enhanced training dataset. The cross-entropy loss of foreground-background segmentation and the smooth L1 loss of bounding box regression are used as the joint loss function. The model parameters are optimized through the backpropagation algorithm until the model converges, thus obtaining the pre-trained moving target detection model.
7. The multi-dimensional linkage control method of the security alarm and video monitoring and lighting system according to claim 5, characterized in that, The method further includes the step of generating linkage control instructions based on the multi-source event fusion report: The multi-source event fusion report is analyzed to extract the target geographic coordinates, target movement direction, and target category probability distribution for each fusion event; Based on the target geographic coordinates, the linkage device spatial configuration database is queried to determine all linkageable video surveillance devices and linkageable lighting devices within the preset influence range of the target geographic coordinates. The linkage device spatial configuration database records the pan-tilt control parameters of all cameras and the dimming and control interface parameters of all lighting fixtures. Based on the target's movement direction, select video surveillance devices from the linked video surveillance devices that cover the area in front of the target's movement direction and use them as preset camera positions; Based on the probability distribution of the target category, the corresponding response rules are matched from the preset linkage strategy rule base. The linkage strategy rule base defines the preset position number of the video surveillance equipment, the video recording resolution and frame rate, and the brightness level and on / off mode of the lighting equipment under different target categories and different confidence levels.
8. The multi-dimensional linkage control method of a security alarm and video monitoring and lighting system according to claim 7, characterized in that, Based on the probability distribution of the target category, the corresponding response rule is matched from a preset linkage strategy rule base, including: From the target category probability distribution, select the target category with the highest probability value as the main discrimination category; Using the main discrimination category and the confidence level of the fusion event as the joint query key, a search is performed in the linkage strategy rule base; The linkage strategy rule base is for different combinations of the main discrimination category and confidence level, and a detailed set of device action parameters is preset. The set of device action parameters includes, but is not limited to: specified preset position number, video recording parameter combination, and lighting brightness curve; If a completely matching entry is found, the set of device action parameters defined by that entry is directly used; If no completely matching entry is found, a fuzzy matching algorithm based on category similarity and confidence proximity is used to select the most similar entry from the linkage strategy rule base, and the device action parameter set is used as the basis for adaptive adjustment to generate the final matching response rule.
9. The multi-dimensional linkage control method of a security alarm and video monitoring and lighting system according to claim 7, wherein, The method also includes the steps of executing linkage control commands and verifying the linkage effect: Based on the matched response rules, a specific sequence of device control instructions is generated. The sequence of device control instructions includes: sending a pan-tilt rotation instruction containing a specified preset position number, a video parameter setting instruction, and a recording start instruction to the preset position camera; and sending a lighting control instruction containing a specified brightness level and on / off mode to the linked lighting device. The sequence of device control commands is sent to the underlying control interfaces of the corresponding video surveillance and lighting equipment. After sending the device control command sequence, an effect verification timer is started. After the effect verification timer reaches a preset delay time, the camera is called from the controlled preset position to obtain the latest real-time video stream. The latest real-time video stream is analyzed to detect whether there is a valid moving target in the video frame within a preset target area; if so, it is determined whether the position of the valid moving target matches the target geographic coordinates of the fused event. The analysis and judgment results are recorded as feedback on the linkage effect and stored in association with the original fusion event and the issued sequence of device control commands.
10. The multi-dimensional linkage control method of a security alarm and video monitoring and lighting system according to claim 9, wherein, The latest real-time video stream is analyzed to detect whether there are any valid moving targets within a preset target area in the video frame, including: Extract one or more frames of images from the latest real-time video stream; Based on the target geographic coordinates of the fusion event and the spatial mapping relationship between the preset location camera, the preset target area corresponding to the video frame is determined; Foreground segmentation processing is performed on the image within the preset target area to calculate pixel change features and identify potential moving pixel blocks; Morphological filtering and connected component analysis are performed on the identified potential moving pixel blocks to eliminate minor interference areas caused by changes in lighting or noise. Features are extracted from the remaining connected components and compared with typical feature templates of the corresponding target category in the linkage strategy rule base. If the feature similarity exceeds the threshold, it is determined that there is a valid moving target.