Camera linkage alarm method and system for intelligent environment monitoring

By integrating the deep fusion analysis of multi-camera video stream data and environmental sensor data in the environmental monitoring system, the problems of false alarms and missed reports in traditional systems in complex environments are solved, accurate detection and multi-dimensional description of abnormal events are achieved, and the system's adaptability and response efficiency are improved.

CN120088957APending Publication Date: 2025-06-03SHENZHEN NEW SAIBO TECHNOLOGY CO LTD
View PDF 0 Cites 40 Cited by

Patent Information

Application Number
CN202510548979.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Traditional environmental monitoring and alarm systems are prone to false alarms and missed alarms in complex environments, and lack the ability to deeply integrate and analyze camera visual data and multi-source environmental sensor information, so they cannot effectively distinguish between normal environmental changes and real abnormal events.

Method used

By synchronously collecting video stream data of multiple cameras and multi-parameter monitoring data of environmental sensors, performing multiple feature analysis and cross-modal feature fusion, obtaining environmental-visual joint features, and performing adaptive classification of environmental scenes based on this feature, selecting an appropriate anomaly determination model for analysis, and finally generating multi-dimensional anomaly event description information through multi-camera collaborative verification and visual-environment data correlation analysis.

Benefits of technology

It realizes accurate detection and identification of abnormal events in complex environments, reduces false alarm rates and missed alarm rates, and improves the system's scenario adaptability and response efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088957A_ABST
    Figure CN120088957A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of camera linkage alarm, and discloses a camera linkage alarm method and system for intelligent environment monitoring, and the method comprises the steps: synchronously collecting the video stream data of a plurality of cameras and the multi-parameter monitoring data of an environment sensor; executing multi-path feature analysis and cross-modal feature fusion to obtain environment-vision joint features; environment scene self-adaptive classification is executed, and environment-vision joint features are analyzed through an anomaly judgment model to obtain an abnormal event judgment result; executing multi-camera collaborative verification and vision-environment data association analysis to obtain multi-dimensional abnormal event description information; and selecting a multi-camera linkage response strategy and coordinating the plurality of cameras to carry out three-dimensional monitoring on the abnormal area to obtain a linkage response execution result, so that a false alarm mode can be identified and an abnormal judgment rule can be dynamically adjusted, continuous optimization of the performance of the plurality of cameras is realized, and the false alarm rate and the missing report rate of the plurality of cameras in long-term operation are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of camera linkage alarm, and particularly to a method and system for camera linkage alarm for intelligent environmental monitoring. Background Art

[0002] With the wide application of intelligent security systems, the problems of false alarms and missed alarms of traditional fixed-threshold alarm systems in complex environments have become increasingly prominent. Most of the existing technologies rely on independent video analysis or simple sensor threshold judgments, lacking the ability to deeply fuse and analyze camera visual data and multi-source environmental sensor information, resulting in the system's inability to comprehensively perceive abnormal situations in complex environments. Especially in complex scenarios such as changes in lighting conditions and the influence of weather factors, the detection of a single camera or a single sensor often shows great limitations and cannot effectively distinguish normal environmental changes from real abnormal events.

[0003] Traditional environmental monitoring and alarm systems often ignore the important influence of environmental context on abnormal determination, making the alarm determination lack scene adaptability. In practical applications, there are significant differences in the "normal" and "abnormal" standards in different environmental scenarios. For example, the same human activities should trigger different levels of alarm responses in a daytime office environment and a nighttime quiet environment. In addition, there is a common problem of time asynchronization of multi-source heterogeneous data in existing systems, resulting in the inability to establish an accurate temporal correlation between visual events and environmental parameter changes, seriously affecting the judgment accuracy of abnormal events. Summary of the Invention

[0004] The present invention provides a method and system for camera linkage alarm for intelligent environmental monitoring. The present invention can identify false alarm patterns and dynamically adjust abnormal determination rules, realize continuous optimization of the performance of multiple cameras, and reduce the false alarm rate and missed alarm rate during the long-term operation of multiple cameras.

[0005] In a first aspect, the present invention provides a method for camera linkage alarm for intelligent environmental monitoring. The method for camera linkage alarm for intelligent environmental monitoring includes:

[0006] Synchronously collect video stream data of multiple cameras and multi-parameter monitoring data of environmental sensors;

[0007] Perform multi-channel feature analysis and cross-modal feature fusion on the video stream data and the multi-parameter monitoring data to obtain environment-vision joint features;

[0008] Perform environment scene adaptive classification according to the environment-vision joint features, select a corresponding abnormal determination model for the current monitoring area, and analyze the environment-vision joint features through the abnormal determination model to obtain an abnormal event determination result;

[0009] Execute multi-camera collaborative verification and vision-environment data correlation analysis based on the abnormal event determination result to obtain multi-dimensional abnormal event description information;

[0010] Select a corresponding multi-camera linkage response strategy according to the multi-dimensional abnormal event description information and coordinate multiple cameras to perform three-dimensional monitoring of the abnormal area to obtain a linkage response execution result.

[0011] In a second aspect, the present invention provides a camera linkage alarm system for intelligent environmental monitoring. The camera linkage alarm system for intelligent environmental monitoring includes:

[0012] A synchronous acquisition module for synchronously acquiring video stream data of multiple cameras and multi-parameter monitoring data of environmental sensors;

[0013] A feature fusion module for performing multi-channel feature analysis and cross-modal feature fusion on the video stream data and the multi-parameter monitoring data to obtain an environment-vision joint feature;

[0014] An abnormal determination module for performing environment scene adaptive classification according to the environment-vision joint feature, selecting a corresponding abnormal determination model for the current monitoring area, and analyzing the environment-vision joint feature through the abnormal determination model to obtain an abnormal event determination result;

[0015] A correlation analysis module for performing multi-camera collaborative verification and vision-environment data correlation analysis based on the abnormal event determination result to obtain multi-dimensional abnormal event description information;

[0016] A three-dimensional monitoring module for selecting a corresponding multi-camera linkage response strategy according to the multi-dimensional abnormal event description information and coordinating multiple cameras to perform three-dimensional monitoring of the abnormal area to obtain a linkage response execution result.

[0017] In the technical solution provided by the present invention, through the synchronous acquisition and cross-modal feature fusion of video stream data from multiple cameras and multi-parameter monitoring data of environmental sensors, the in-depth correlation analysis between visual information and environmental parameters is realized, overcoming the limitations of single data source analysis, and significantly improving the detection accuracy of the system for abnormal events in complex environments. Based on the scene adaptive classification mechanism of environment-vision joint features, the system can automatically identify the scene type of the current monitoring environment and select an abnormal determination model optimized for a specific scene, solving the problem that traditional fixed threshold systems are difficult to adapt to environmental changes. Through multi-camera collaborative verification and visual-environment data correlation analysis, the system can generate an abnormal event description containing multi-dimensional information such as abnormal type, location, intensity, duration, and credibility, providing a comprehensive basis for subsequent alarm response decisions. By adopting the role assignment mechanism of the core monitoring camera, auxiliary monitoring camera, and environmental perception camera, combined with the optimal viewing angle calculation and complementary angle configuration, all-round three-dimensional monitoring of the abnormal area is realized, eliminating the visual blind area in traditional single-camera monitoring. The system can analyze the potential development trajectory of abnormal events based on their spatio-temporal characteristics and environmental factors, adjust the monitoring parameters of inactive cameras in advance, form a pre-defense monitoring network, and realize the dynamic tracking and predictive coverage of abnormal events. According to the characteristics of abnormal events, a multi-level linkage response is implemented, and different camera parameter adjustments, video acquisition strategies, and alarm notification methods are adopted for different levels of abnormalities, improving the accuracy and efficiency of the system response. Through continuous analysis of the execution results and processing feedback of the linkage response, the system can identify false alarm patterns and dynamically adjust the abnormal determination rules, realizing the continuous optimization of system performance and reducing the false alarm rate and missed alarm rate during long-term operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0019] Figure 1 It is a schematic diagram of the steps of the camera linkage alarm method for intelligent environmental monitoring in the embodiments of the present invention;

[0020] Figure 2 It is a schematic diagram of the structure of the camera linkage alarm system for intelligent environmental monitoring in the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] An embodiment of the present invention provides a camera linkage alarm method and system for intelligent environmental monitoring. The terms "first", "second", "third", "fourth", etc. (if any) in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments described here can be implemented in an order different from that shown or described here. In addition, the term "comprising" or "having" and any variation thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0022] For ease of understanding, the specific process of the embodiment of the present invention will be described below. Please refer to Figure 1 , an embodiment of the camera linkage alarm method for intelligent environmental monitoring in the embodiment of the present invention includes:

[0023] Step S1, synchronously collect video stream data of multiple cameras and multi-parameter monitoring data of environmental sensors;

[0024] It can be understood that the execution subject of the present invention can be a camera linkage alarm system for intelligent environmental monitoring, or a terminal or a server. Specifically, it is not limited here. The embodiment of the present invention takes the server as the execution subject as an example for illustration.

[0025] Specifically, the key video stream acquisition parameters of multiple cameras are configured, including resolution, frame rate, encoding format and key viewing angle coverage. At the same time, according to the specific environment and risk characteristics of the monitoring area, a distributed environmental sensor network is deployed, and the optimal sampling frequency is set for multiple types of sensors such as temperature, humidity, air pressure, light intensity, smoke concentration, harmful gas concentration and motion state, thereby obtaining a multi-source environmental parameter data acquisition strategy. According to the video stream acquisition parameters, the original video is collected from multiple cameras, and each frame of the image is processed by histogram equalization to enhance the contrast and detail expression of the image. Then, the Gaussian filter algorithm is used to remove the high-frequency noise in the image, improve the visual clarity of the video stream and the accuracy of subsequent feature analysis, and then the distortion caused by the camera viewing angle is corrected by perspective transformation, so that the obtained video stream data has stable and unified visual geometric properties. At the same time, for the acquisition of multi-source environmental parameter data, according to the multi-parameter data acquisition strategy set in the system design phase, the distributed environmental sensor network is used to monitor multiple physical quantities in the environment such as temperature, humidity, air pressure, light intensity, smoke concentration, harmful gas concentration and motion state in real time to obtain multi-parameter data streams. In order to eliminate data anomalies and missing data caused by sensor errors, external interference and other factors in the actual environment, sliding window median filtering is performed on the original multi-parameter data to effectively remove mutation values ​​and extreme outliers. Then, the data of various physical quantities are uniformly mapped to the same scale interval through Z-score standardization processing, which is convenient for subsequent cross-modal data analysis and modeling. Finally, the linear interpolation algorithm is used to reasonably supplement the missing points in the acquisition process to obtain multi-parameter monitoring data.

[0026] Step S2, performing multi-channel feature analysis and cross-modal feature fusion on the video stream data and the multi-parameter monitoring data to obtain an environment-vision joint feature;

[0027] Specifically, the video stream data is input into a three-dimensional convolutional neural network, and the spatial and temporal dimensions of the video sequence are simultaneously modeled through multi-layer three-dimensional convolutional operations. A spatio-temporal attention calculation module is introduced into the network structure. The spatial attention mechanism is used to highlight the key regions in the video frame, and the temporal attention mechanism is used to enhance the dynamic expressiveness of the key frames, so as to obtain a video spatio-temporal feature vector that can reflect the dynamic changes and abnormal clues in the video content. At the same time, for the multi-parameter monitoring data collected by the environmental sensor network, it is input into a bidirectional long short-term memory network with sequence modeling ability. The complex change trends of various environmental parameters at historical and current moments are captured through forward and backward information flow transmission. On this basis, an inter-sensor attention mechanism is introduced to dynamically assign weights to different types of environmental data, realizing the adaptive adjustment of the sensitivity to sensor value anomalies or environmental changes, and thus extracting the environmental parameter time series feature vector representing the evolution characteristics of the current environmental state. Through methods such as fully connected layers, the video spatio-temporal feature vector and the environmental parameter time series feature vector are dimensionally mapped to obtain a set of feature vectors of a unified scale. Based on the feature vectors of the unified scale, by constructing a bidirectional cross-attention calculation module, the correlation matrix between the video features and the environmental features is calculated, realizing the complementary enhancement of information, and highlighting the key features that are simultaneously affected by environmental changes and video dynamics, to obtain the inter-modal association enhanced features. The inter-modal association enhanced features are input into the non-linear feature transformation and complementary filling module, and the feature representation ability is enriched through means such as multi-layer non-linear activation, feature reconstruction, and residual enhancement to obtain a fused feature vector. According to the information gain criterion, feature selection is performed on the high-dimensional fused feature vector, effectively screening out the feature subset that makes the most contribution to anomaly discrimination and environmental understanding, and forming an environment-vision joint feature vector.

[0028] Step S3: Perform environment scene adaptive classification according to the environment-vision joint feature, select the corresponding anomaly determination model for the current monitoring area, and analyze the environment-vision joint feature through the anomaly determination model to obtain the anomaly event determination result;

[0029] Specifically, the environment-vision joint features are input into a densely connected network. Through multi-layer feature cascading and high-dimensional non-linear mapping inside the network, various environmental and behavioral manifestations in the current monitoring area are comprehensively analyzed to identify typical environmental scene types, including normal operation, nighttime silence, bad weather, crowded people, etc. After completing the scene type discrimination, according to the current environmental scene label, the most suitable anomaly determination model for the current situation is intelligently selected from a pre-trained model library. The model library covers a variety of anomaly detection models specifically designed and independently trained for different environmental scenes. Among them, the Gaussian mixture model is used to measure the probability distribution of the environment-vision joint features in the normal mode, and the number of its components and the covariance matrix structure will be flexibly adjusted according to different scene complexities; while the one-dimensional convolutional autoencoder adaptively reconstructs the normal temporal structure of video content or environmental parameters, and the number of network layers and convolutional kernel parameters are different under different scenes to adapt to the diversity and variability of feature distributions in various scenes. The environment-vision joint features at the current moment are input into the anomaly determination model. The logarithmic likelihood probability of this feature under the normal distribution is calculated through the Gaussian mixture model. If the probability is low, it indicates that the possibility of this feature being abnormal is relatively high. At the same time, under the one-dimensional convolutional autoencoder path, the reconstruction error between the feature input and the reconstruction output is calculated to measure the deviation degree between the current observation and the "normal mode", and the above probability result and the reconstruction error are jointly integrated into the target anomaly score, so as to realize the fine-grained quantification and dynamic perception of the abnormal state. To prevent false alarms or missed alarms caused by rapid scene changes, the system continuously runs a change point detection algorithm to monitor the temporal change trend of the multi-modal feature stream in real time. When a significant conversion of the environmental scene is detected, the decision threshold of the anomaly score is temporarily adjusted according to the characteristics of the new scene, and combined with the current anomaly score result and the context information of the environmental change, a more refined reclassification of the type of abnormal event is made, and finally the abnormal event determination result is obtained.

[0030] In this embodiment, according to the current environmental scenario type, weight coefficients with scene adaptability are set for the environmental channel and the visual channel respectively. Through the dynamic weight allocation mechanism, the most representative data sources are strengthened in different types of scenarios to ensure that the feature analysis results best fit the actual environment. For example, in scenarios dominated by environmental changes such as meteorological mutations and chemical leaks, the first weight coefficient of the environmental channel is increased, and the environmental weight allocation is performed on the environmental-visual joint features using this weight to obtain the first environmental-visual feature vector that highlights the sensitivity to environmental changes. In scenarios dominated by video dynamics such as abnormal activities at night or personnel gatherings, the second weight coefficient of the visual channel is correspondingly increased, and the second environmental-visual feature vector that emphasizes video feature expression is obtained through visual weight allocation. The first environmental-visual feature vector highlighting the environmental weight is input into the Gaussian mixture model path in the anomaly determination model. Using its multi-modal distribution modeling ability, the probability density of the feature under each component is comprehensively calculated, and by performing a logarithmic transformation on the obtained multi-dimensional probability density values, the numerical scale is compressed and the abnormal points are highlighted to obtain the abnormal probability score of the joint feature, reflecting the statistical deviation degree between the current observed value and the "normal distribution". At the same time, the second environmental-visual feature vector highlighting the visual weight is input into the one-dimensional convolutional autoencoder path of the anomaly determination model. The feature sequence is compressed and mapped to a high-dimensional space through the encoder, and then the decoder reconstructs a feature result as close as possible to the input. By comparing the difference between the reconstruction result and the original input feature, the deviation degree between the current state and the historical normal mode is quantified to obtain the reconstruction error of the joint feature. A non-linear weighted combination is performed on the joint feature abnormal probability score and the joint feature reconstruction error. During the weight allocation process, the scene type and the historical performance of the model are considered, and an adaptive adjustment coefficient is introduced. Through methods such as activation function transformation and dynamic coefficient optimization, the target abnormal score is finally obtained.

[0031] Step S4: Based on the abnormal event determination result, perform multi-camera collaborative verification and visual-environment data correlation analysis to obtain multi-dimensional abnormal event description information;

[0032] Specifically, the event information preliminarily determined as abnormal is synchronously transmitted to each relevant camera unit within the system. Through spatial layout and field of view configuration, spatial positioning analysis of the same abnormal target is carried out using the overlapping area between adjacent cameras. For the projection of the pixel coordinates of the abnormal target under different camera perspectives, through the triangulation algorithm and spatial coordinate mapping technology, the precise three-dimensional coordinate position information of the abnormal event in the monitoring scene is calculated. According to this three-dimensional target position information, all cameras covering the target area are coordinated to jointly perform real-time tracking of the abnormal target. By integrating the video stream data of multiple cameras and combining algorithms such as object detection, region matching, and temporal correlation, the consistency of the dynamic trajectory of the target in each camera and the coherence of the spatial behavior are continuously verified to ensure that the determination result of the abnormal target is not caused by a false alarm of a single camera, but a real abnormal phenomenon under multiple perspectives. On this basis, in-depth correlation calculation is carried out on the spatial behavior data obtained from cross-camera tracking and the abnormal pattern of environmental parameters extracted from the abnormal event determination model. A variety of algorithms such as dynamic time warping, correlation coefficient analysis, and mutual information measure are used to evaluate the synchronization and coupling strength between the visual target behavior and the change of environmental parameters, and cross-modal abnormal association features that can truly reflect the essential characteristics of multi-modal abnormal events are refined. Based on the cross-modal abnormal association features, multi-dimensional abnormal event description information including abnormal type identifiers, abnormal spatial position coordinates, abnormal intensity values, abnormal duration statistics, and abnormal credibility scores is constructed.

[0033] In this embodiment, according to the target three-dimensional coordinate position information, the multi-modal background of the entire monitoring area is perceived and regionally segmented. In this process, the geometric distribution of visual images is considered, and the changes of environmental factors such as temperature, humidity, light, and gas concentration sensed by environmental sensors are fused. Through feature clustering, spatial analysis, and regional adaptive algorithms, the monitoring scene is divided into multiple sub-regions that are highly independent in terms of environmental characteristics, and environmental characteristic weights are dynamically assigned to each environmental characteristic independent region to obtain a regional division map reflecting environmental sensitivity. Based on the environmental sensitive region division map and the target three-dimensional coordinate position information, a multi-dimensional state transition model is constructed to incorporate the movement trajectory of the target in different regions and the environmental evolution process into a unified modeling framework. Deep spatio-temporal features are extracted from the target regions corresponding to the target three-dimensional coordinates in the video stream data collected by all relevant cameras. This feature extraction relies on traditional spatio-temporal convolutional neural networks to extract the appearance, movement trajectory, and behavioral dynamic features of the target, and at the same time introduces a fusion channel for environmental parameters to obtain a set of target characterization feature sets that can reflect environmental sensitive characteristics. The environment-enhanced target characterization feature set and the multi-dimensional state transition model are input into a non-linear extended Kalman filter. Using the adaptive estimation ability of the Kalman filter for multi-sensor data and complex spatio-temporal dynamics, parallel prediction and state correction of the multi-path trajectory of the target are performed, and the prediction results of the multi-path trajectory of the target in different environmental regions and under different cameras are output. According to the multi-path trajectory prediction, a target transfer probability matrix across cameras is constructed to quantify the spatio-temporal probability relationship of the target migrating from the field of view of one camera to that of another camera. Combining the physical layout and topological relationship of the cameras, the Hungarian matching algorithm is used to perform optimal allocation and identity correspondence for all possible target identity migration schemes to obtain a multi-level, globally optimal camera target identity matching scheme. For the multi-level target identity matching scheme, multiple information such as time synchronization, spatial trajectory, and behavioral continuity is combined to perform spatio-temporal consistency constraints. Abnormal situations such as short-term interference and occlusion misjudgment are excluded through cross-validation, and the verification of environmental parameter changes is introduced to determine whether there is a consistent pattern response between the abnormal trajectory of the target under different cameras and environmental events. When the visual, environmental, and spatio-temporal features all match highly, the final cross-camera abnormal target consistency verification result is output.

[0034] In this embodiment, according to the cross-camera abnormal target consistency verification result, a spatio-temporal trajectory map of the abnormal target is constructed to record the dynamic transfer process, movement trajectory, residence duration, and spatial path of the abnormal target between different camera fields of view. At the same time, according to the abnormal pattern of environmental parameters extracted from the abnormal event determination result, a corresponding spatio-temporal distribution map of environmental anomalies is constructed. By analyzing the distribution changes of multi-parameters such as temperature, humidity, gas concentration, and light in the spatial and temporal dimensions, the occurrence area, evolution trend, and diffusion range of environmental anomalies are characterized. The spatio-temporal trajectory map of the abnormal target and the spatio-temporal distribution map of environmental anomalies are deeply fused to generate a multi-modal temporal correlation map, which reflects the time and space synchronization characteristics between different modalities in terms of structure and reveals the potential coupling and synergy relationships between abnormal events in the two modalities of vision and environment. By implementing structured feature extraction methods on this multi-modal temporal correlation map, such as node clustering, connected component analysis, path weight induction, etc., a high-dimensional and discriminative abnormal event correlation feature vector is obtained. The abnormal event correlation feature vector is input into the modal correlation metric model, and using statistical information theory and machine learning methods, the mutual information and conditional probability between visual anomalies and environmental anomalies are calculated to obtain a cross-modal abnormal synergy index that quantifies the degree of abnormal synergy between the two modalities. Based on the cross-modal abnormal synergy index, the temporal influence relationship and diffusion law of abnormal events between different modalities are analyzed. Through techniques such as time series modeling and influence path tracking, it is identified how abnormal events are transmitted from environmental parameter changes to target movement anomalies, or vice versa, from visual anomalies triggering environmental responses, to realize the quantification and modeling of abnormal propagation characteristic information. Causal relationship analysis is performed on the abnormal propagation characteristic information. Through tools such as causal inference algorithms and Granger causality tests, the causal effect direction and logical relationship of abnormal events between the two modalities of vision and environment are systematically determined, clarifying whether the environment changes first to cause target anomalies, or target anomalies induce environmental fluctuations, so as to obtain the causal relationship of abnormal events. Combining the causal relationship of abnormal events and the cross-modal abnormal synergy index, a cross-modal abnormal correlation feature is constructed.

[0035] Step S5: Select a corresponding multi-camera linkage response strategy according to the multi-dimensional abnormal event description information and coordinate multiple cameras to perform three-dimensional monitoring of the abnormal area to obtain the linkage response execution result.

[0036] Specifically, camera role assignment is performed according to the anomaly type, anomaly credibility, and anomaly spatial position coordinates in the multi-dimensional anomaly event description information. Among them, the anomaly type determines the response level and linkage priority, the anomaly credibility reflects the confidence level of event determination, and the spatial position coordinates provide the basis for perspective coverage calculation. Cameras located in the central area of the anomaly event and with good perspective conditions are assigned as "core monitoring cameras", and these devices undertake the main image acquisition and key feature analysis tasks; while devices located in sub-optimal positions but with supplementary perspectives are designated as "auxiliary monitoring cameras", and their tasks are to supplement blind area information and provide horizontal verification; the sensing unit specifically configured as an "environmental perception camera" is used to collect changes in environmental parameters around the anomaly area to enrich the scene background information. The optimal perspective calculation is performed on the core monitoring cameras. Based on the target three-dimensional position, the current camera installation posture, zoom ability, and field of view angle parameters, the best shooting direction, focal length adjustment, and attitude rotation amount are deduced using a geometric model to generate an accurate positioning instruction containing three-dimensional positioning vectors and instruction parameters. To ensure the comprehensive coverage of the anomaly area and multi-angle information acquisition, according to the monitoring perspectives of the core cameras and the spatial distribution characteristics of the anomaly area, the most complementary monitoring angles are calculated for the auxiliary cameras to ensure the optimization of field of view coverage and the minimization of information overlap among multiple cameras, forming a multi-camera linkage response strategy. Based on the multi-camera linkage response strategy, according to the role positioning of the cameras and the computing capabilities of the edge devices, the image analysis tasks are hierarchically and differentially allocated. Among them, the core cameras undertake key analysis tasks such as deep feature extraction and behavior recognition, the auxiliary cameras are responsible for low-latency tasks such as target tracking and boundary judgment, and the environmental perception cameras focus on the structured processing of environmental data, thereby generating a multi-level visual analysis instruction system. Based on this instruction system, the system schedules the image data of the core and auxiliary cameras to perform cross-perspective geometric reconstruction tasks. Through processing means such as multi-view stereo matching, depth map fusion, and three-dimensional model fitting, a high-precision stereo vision model is constructed. This model restores the three-dimensional spatial morphology and target position of the anomaly area, and also integrates the environmental context information provided by the environmental perception cameras, such as light, smoke concentration, temperature and humidity changes, etc., thereby enhancing the expression ability of the model in complex dynamic environments and forming a stereo monitoring view of the anomaly area. Based on the stereo monitoring view, hierarchical control of the camera roles is implemented, including the core cameras automatically adjusting the zoom, exposure, and focus strategies, the auxiliary cameras performing automatic supplementary positioning, angle synchronization, and collaborative acquisition according to the changes in the main perspective, and the environmental perception cameras adjusting the sampling frequency and switching the perception mode according to the environmental anomaly intensity, thereby forming a dynamic collaborative control mechanism with task orientation as the core within the system and outputting the execution results of the linkage response.

[0037] In the embodiments of the present invention, through the synchronous acquisition and cross-modal feature fusion of video stream data from multiple cameras and multi-parameter monitoring data of environmental sensors, the in-depth correlation analysis of visual information and environmental parameters is realized, overcoming the limitations of single data source analysis, and significantly improving the detection accuracy of the system for abnormal events in complex environments. Based on the scene adaptive classification mechanism of environmental-visual joint features, the system can automatically identify the scene type of the current monitored environment and select an abnormal determination model optimized for a specific scene, solving the problem that traditional fixed threshold systems are difficult to adapt to environmental changes. Through multi-camera collaborative verification and visual-environment data correlation analysis, the system can generate an abnormal event description containing multi-dimensional information such as abnormal type, location, intensity, duration, and credibility, providing a comprehensive basis for subsequent alarm response decisions. By adopting the role assignment mechanism of the core monitoring camera, auxiliary monitoring camera, and environmental perception camera, combined with the optimal viewing angle calculation and complementary angle configuration, the all-round three-dimensional monitoring of the abnormal area is realized, eliminating the visual blind area in traditional single-camera monitoring. The system can analyze the potential development trajectory of an abnormal event based on its spatio-temporal characteristics and environmental factors, adjust the monitoring parameters of inactive cameras in advance, form a pre-defense monitoring network, and realize the dynamic tracking and predictive coverage of abnormal events. Implement multi-level linkage response according to the characteristics of abnormal events, adopt different camera parameter adjustments, video acquisition strategies, and alarm notification methods for different levels of abnormalities, improving the accuracy and efficiency of the system response. Through the continuous analysis of the execution results and processing feedback of the linkage response, the system can identify false alarm patterns and dynamically adjust the abnormal determination rules, realizing the continuous optimization of the system performance and reducing the false alarm rate and missed alarm rate during long-term operation.

[0038] In a specific embodiment, the process of executing step S1 may specifically include the following steps:

[0039] Configure the video stream acquisition parameters of multiple cameras, and at the same time set the sampling frequencies of temperature, humidity, air pressure, light intensity, smoke concentration, harmful gas concentration, and motion state sensors in the distributed environmental sensor network to obtain a multi-source environmental parameter data acquisition strategy;

[0040] Collect the original video from multiple cameras according to the video stream acquisition parameters, and perform histogram equalization processing, Gaussian filtering processing, and perspective transformation correction on the original video to obtain video stream data;

[0041] Collect the original multi-parameter data according to the multi-source environmental parameter data acquisition strategy and the distributed environmental sensor network, and perform moving window median filtering, Z-score normalization processing, and linear interpolation supplementation on the original multi-parameter data to obtain multi-parameter monitoring data.

[0042] Specifically, according to the actual requirements of the monitoring scenario, the video stream acquisition parameters of multiple cameras and the working parameters of the distributed environmental sensor network are configured. In the configuration process, based on the area, spatial structure, target monitoring accuracy, lighting conditions, environmental complexity, and expected event types of the monitoring area, parameters such as the resolution, frame rate, exposure time, field of view angle, compression format, and network bandwidth allocation of each camera are set individually. For example, in key areas, a high-definition video resolution of 1920×1080 pixels and a collection frame rate of 25 frames per second are adopted, while in conventional areas, the parameters are reduced to save resources and bandwidth. For various sensing units in the distributed environmental sensor network, for physical parameters such as temperature, humidity, air pressure, light intensity, smoke concentration, harmful gas concentration, and motion state, different sampling frequencies are set according to their respective change rates, the suddenness of physical phenomena, the stability of the environmental background, and the periodic requirements of data fusion modeling. For example, safety risk parameters that change rapidly such as gas leakage and smoke are set to a high-frequency sampling of 1Hz or higher, while slow variables such as temperature and humidity are reduced to a low-frequency collection of 0.1Hz or 0.2Hz. Through this step, a comprehensive strategy for multi-source environmental parameter data acquisition is formed. In the data acquisition stage, according to the set video stream acquisition parameters, the original video stream data is synchronously and real-time collected from each monitoring point through the distributed network of cameras. Histogram equalization processing is performed on each frame of the image to enhance the details in the dark areas and suppress the strong light areas, achieving the smoothing of the overall brightness distribution. For the high-frequency noise and minute particle disturbances that appear in the video, the Gaussian filtering algorithm is used to denoise the video image. By setting the filter kernel size and variance parameters, while ensuring the clarity and structural integrity of the image, the random noise is effectively smoothed, the influence of interference signals is reduced, and the accuracy of feature extraction in the subsequent analysis link is improved. Perspective transformation correction is performed on each frame of the image. By calibrating the internal and external parameters of the camera and using the perspective matrix mapping, the distorted image captured by the camera is restored to a frontal view in the standard space, eliminating the influence of geometric distortion on target positioning and spatial modeling, and ensuring that the video stream data has a unified geometric reference. After the above image enhancement and correction processing, the original video data is transformed into high-quality, spatio-temporally consistent standardized video stream data. At the same time, according to the multi-source environmental parameter data acquisition strategy, the original multi-parameter data stream is synchronously collected from the distributed environmental sensor network, including multiple environmental indicators such as temperature and humidity, air pressure, light, smoke, harmful gas, and motion state. Sliding window median filtering processing is performed on all the collected multi-parameter data streams. According to a certain window length, local statistics are performed on the time series data, automatically removing short-term mutation values and extreme outliers, effectively weakening the outlier noise, and improving the stability of the environmental parameter curve.The Z-score standardization processing algorithm is adopted to convert all multi-parameter data into a dimensionless standard distribution with a mean of zero and a variance of one, improving the fusion modeling ability between different data sources and preventing problems such as feature weight bias and poor model convergence caused by scale imbalance. At the same time, the linear interpolation algorithm is used to continuously supplement the missing points in the data sequence to ensure the integrity of various environmental parameter data without breakpoints in time series, and multi-parameter monitoring data is obtained.

[0043] In a specific embodiment, the process of executing step S2 may specifically include the following steps:

[0044] Input the video stream data into a three-dimensional convolutional neural network for three-dimensional convolutional processing and spatio-temporal attention calculation to obtain a video spatio-temporal feature vector;

[0045] Input the multi-parameter monitoring data into a bidirectional long short-term memory network, and introduce an inter-sensor attention mechanism to assign dynamic weights to different sensor data to obtain an environmental parameter time series feature vector;

[0046] Perform dimensionality mapping on the video spatio-temporal feature vector and the environmental parameter time series feature vector to obtain a unified scale feature vector;

[0047] Perform bidirectional cross-attention calculation on the unified scale feature vector to obtain an inter-modal correlation enhanced feature, and perform non-linear feature transformation and complementary filling on the inter-modal correlation enhanced feature to obtain a fusion feature vector;

[0048] Perform feature selection on the fusion feature vector according to the information gain criterion to obtain an environment-vision joint feature.

[0049] Specifically, the video stream data is input into the three-dimensional convolutional neural network architecture, and the video sequence is modeled in both spatial and temporal dimensions. Through the stacking of multiple layers of 3D convolutional kernels and the adjustment of step size, the deep-level features of object motion, behavior changes and spatial structure between video frames are extracted. At the same time, in order to improve the perception of dynamic spatiotemporal focus, the spatiotemporal attention module is embedded in the high-order feature layer of the three-dimensional convolutional neural network structure. On the one hand, this module automatically identifies the key areas in the video screen through the spatial attention mechanism, focusing the network's attention on abnormal targets, motion hotspots and other areas. On the other hand, it allocates weights in the sequence dimension through the temporal attention mechanism, highlighting the key frames with drastic changes in behavior, important moments such as the occurrence points of emergencies, and forming a video spatiotemporal feature vector that can efficiently express the dynamic information of the video and the law of abnormal occurrence. For multi-source environmental monitoring data, the change law and correlation of various sensor parameters in the time series are analyzed. The synchronously collected multi-parameter data is input into the bidirectional long short-term memory network, and the data is globally modeled using the forward and reverse time series information, respectively, to effectively capture complex time dynamics such as sudden changes in the environment, parameter trend reversal, and periodic fluctuations. In this network structure, in order to prevent a certain type of sensor data from affecting the overall feature expression due to decreased sensitivity or increased noise in special scenarios, an inter-sensor attention mechanism is introduced. By dynamically allocating the weights of each sensor channel, the physical quantity that best represents the abnormal mode in the current scenario is adaptively strengthened. For example, the weights of smoke and harmful gas sensors are greatly increased in scenarios such as fire, while the long-term fluctuation trend of temperature, humidity and air pressure is strengthened in the silent environment at night. Through this step, the time series feature vector of environmental parameters is obtained. Through methods such as fully connected networks, normalization layers or convolutional mapping, the video spatiotemporal feature vector and the time series feature vector of environmental parameters are dimensionally mapped and uniformly converted into a set of scaled feature vectors of equal dimensions, eliminating the heterogeneity of the original feature dimensions, structure and distribution between modalities. After completing the dimensional alignment, a bidirectional cross-attention mechanism is introduced in the feature fusion stage. For the unified scale features of the two major modalities of video and environment, the coupling relationship between video information and environmental information is dynamically mined by constructing a modal correlation matrix and allocating attention weights, realizing information complementarity and enhanced discrimination between modalities, and obtaining inter-modal correlation enhancement features. Perform nonlinear feature transformation and complementary filling operations on the inter-modal correlation enhancement features, including multi-layer activation functions, residual connections, convolution transformations or adaptive completion mechanisms, and explore the potential structural relationships in the high-order feature space through deep nonlinear reconstruction. Effectively fill and correct feature holes caused by short-term loss of a single modality, attenuation of abnormal signals, etc., to obtain a more discriminative and noise-resistant fusion feature vector. According to the information gain criterion, perform a feature selection algorithm on the high-dimensional fusion feature vector to automatically screen out the feature subset with the greatest contribution to abnormal recognition and the least redundancy, and obtain the environment-vision joint feature.

[0050] In a specific embodiment, the process of executing step S3 may specifically include the following steps:

[0051] Input the environment-vision joint feature into a densely connected network for scene type analysis to obtain the environmental scene type of the current monitoring area;

[0052] Select an anomaly determination model for the current environmental scene from the pre-trained model library according to the environmental scene type. The anomaly determination model includes a Gaussian mixture model and a one-dimensional convolutional autoencoder; the pre-trained model library contains anomaly determination models pre-trained for various environmental scene types, and the number of components, covariance matrix structure of the Gaussian mixture model corresponding to different environmental scene types, and the number of network layers and convolutional kernel parameters of the one-dimensional convolutional autoencoder are all different;

[0053] Input the environment-vision joint feature into the anomaly determination model, calculate the log-likelihood probability of the environment-vision joint feature under the Gaussian mixture model and the reconstruction error of the one-dimensional convolutional autoencoder, and obtain the target anomaly score;

[0054] Execute a change point detection algorithm to monitor the environmental scene conversion state, and perform temporary threshold adjustment and anomaly type classification on the target anomaly score when a scene conversion is detected to obtain an anomaly event determination result.

[0055] Specifically, the environment-vision joint feature vector is input into a densely connected network. The densely connected network enhances the richness of feature expression and the multi-level perception ability with its fully connected structure of inter-layer features, and can effectively integrate multi-dimensional heterogeneous information such as space, time, and environment. Through the stacking and fusion of multi-layer features inside the network, the densely connected network can deeply mine the comprehensive environmental features of the monitoring area, so as to perform adaptive analysis and classification on the current scene type. The classification results include typical environmental scene labels such as "normal operation", "night silence", "extreme climate", "environmental mutation", "crowded with people", etc., and are extended to more complex and detailed scene classifications based on historical big data and environmental knowledge bases. According to the discriminated environmental scene labels, an anomaly determination model that best matches the current scene is dynamically selected from the pre-trained model library. This model library serves as the underlying support for the system's intelligent decision-making, integrating a variety of Gaussian mixture models and one-dimensional convolutional autoencoders customized for typical environmental scenes. Each group of models is independently optimized according to the scene characteristics. Different Gaussian mixture models vary in core parameters such as the number of components and the covariance matrix structure, and can accurately fit the normal probability distribution of environment-vision features in the corresponding scene; the one-dimensional convolutional autoencoders in different scenes make adaptive adjustments in terms of network depth, number of convolutional kernels, receptive field, residual structure, etc., so as to ensure the optimal anomaly reconstruction and expression ability in their respective sub-scenes. In the anomaly determination link, the environment-vision joint feature is input into the currently selected anomaly determination model. Through the Gaussian mixture model path, the probability density of the feature vector under the normal distribution is estimated, and then the logarithmic transformation is used to obtain the log-likelihood probability of each observation sample. The lower this probability value means that the current sample deviates more from the normal mode and the higher the possibility of being abnormal. At the same time, through the one-dimensional convolutional autoencoder branch, the joint feature is deeply encoded and decoded, and the reconstruction result of the feature is output. The reconstruction error between the original feature and the reconstructed feature is calculated as the anomaly measure. The larger this error indicates the greater the deviation between the current observation and the known normal behavior. The two combined can achieve multi-dimensional and dual-path anomaly quantification. The probability score and the reconstruction error are fused through non-linear weighting, threshold constraint and other methods to form the target anomaly score. The change point detection algorithm is continuously executed to dynamically monitor indicators such as the distribution, mean, and variance of the multi-modal joint feature on the time axis, and real-time detection of the turning point of the environmental scene is carried out through methods such as sliding window and Bayesian change point detection. When it is detected that the scene has changed, the optimal anomaly determination model is re-selected for the new scene, and at the same time, the threshold parameters of the existing target anomaly score are temporarily adjusted, for example, the alarm threshold is appropriately increased to prevent false alarms during the scene switching process, or the anomaly type classification mechanism is optimized for a short time to adapt to the feature distribution and risk sensitivity in the new scene. According to the adjusted target anomaly score and the new scene classification strategy, a more targeted anomaly event determination result is output.

[0056] In a specific embodiment, the process of inputting the environment-vision joint feature into the anomaly determination model, calculating the log-likelihood probability of the environment-vision joint feature under the Gaussian mixture model and the reconstruction error of the one-dimensional convolutional autoencoder, and obtaining the target anomaly score may specifically include the following steps:

[0057] Set the first weight coefficient of the environment channel according to the environment scene type, and perform environment weight allocation on the environment-vision joint feature according to the first weight coefficient to obtain the first environment-vision feature vector;

[0058] Set the second weight coefficient of the vision channel according to the environment scene type, and perform vision weight allocation on the environment-vision joint feature according to the second weight coefficient to obtain the second environment-vision feature vector;

[0059] Input the first environment-vision feature vector into the Gaussian mixture model of the anomaly determination model, calculate the multi-dimensional probability density value and perform logarithmic transformation to obtain the joint feature anomaly probability score;

[0060] Input the second environment-vision feature vector into the one-dimensional convolutional autoencoder of the anomaly determination model at the same time, and generate the reconstructed feature through the encoding-decoding process to obtain the feature reconstruction result;

[0061] Calculate the joint feature reconstruction error between the feature reconstruction result and the second environment-vision feature vector, and perform non-linear weighted combination on the joint feature anomaly probability score and the joint feature reconstruction error to obtain the target anomaly score.

[0062] Specifically, after the system determines the environmental scenario type of the current monitoring area, it dynamically sets the first weight coefficient of the environmental channel according to the knowledge base or empirical rules. The importance of environmental parameters varies significantly in different scenarios. For example, in scenarios such as extreme climate, gas leakage, and fire, environmental data contributes far more to anomaly detection than video data. Therefore, the weight of the environmental channel should be significantly increased. In scenarios mainly focused on behavior and dynamics, such as night monitoring and crowded areas, the weight of the environmental channel is relatively reduced. The setting of the weight coefficient is dynamically adjusted through expert experience, scenario modeling, or the use of adaptive weight optimization algorithms (such as Bayesian optimization based on historical anomaly recognition accuracy feedback and self-attention mechanism weight learning). Based on the current first weight coefficient, weighted summation or weighted concatenation processing is performed on various environmental feature components in the environmental-visual joint feature vector to strengthen the feature dimensions that are most sensitive and discriminative to environmental changes, forming the first environmental-visual feature vector with scene adaptive capabilities. Similarly, the second weight coefficient of the visual channel is set synchronously according to the environmental scenario type. In scenarios dominated by visual features, such as human behavior recognition, scene intrusion, and item loss, the weight of the visual channel is significantly increased. In situations with limited visual changes, such as long-term static environments and natural scene monitoring, the visual weight is moderately reduced. According to the current second weight coefficient, the visual components in the joint feature are weighted to highlight the feature information in the video data that reflects dynamic changes, behavior patterns, or abnormal targets to the greatest extent, obtaining the second environmental-visual feature vector. The first environmental-visual feature vector is input into the Gaussian mixture model path of the anomaly determination model. As an important tool for probability statistical modeling, the Gaussian mixture model accurately fits the statistical distribution of normal environmental-visual features through multi-component Gaussian distributions, thereby establishing a probability density reference for the "normal mode" in the feature space. In actual calculations, the probability density values of the first environmental-visual feature vector under each component of the Gaussian mixture model are calculated, and the results are logarithmically transformed to amplify the small differences in the high-dimensional probability space into anomaly signals that are more easily distinguishable at the numerical level. The lower the log-likelihood probability score, the greater the deviation of the current feature from the normal mode and the higher the likelihood of an anomaly. At the same time, the second environmental-visual feature vector is input into the one-dimensional convolutional autoencoder of the anomaly determination model. This autoencoder performs multi-layer one-dimensional convolution, dimensionality reduction, and feature reconstruction on the input feature sequence, and restores the high-dimensional features to the original structure through the decoder. The encoding-decoding process is essentially fitting the "normal behavior space" of the input features. If the current observation highly matches the structural pattern of historical normal samples, the reconstruction error is extremely small. If an anomaly occurs, the decoding and restoration ability significantly decreases, resulting in a substantial increase in the reconstruction error. The system directly calculates various reconstruction error metrics such as the Euclidean distance, mean squared error, or cosine similarity between the original second environmental-visual feature vector and the reconstructed features of the autoencoder. This joint feature reconstruction error reflects the deviation intensity between the observed data and the historical normal data pattern.Complementary to probability and statistics modeling, the one-dimensional convolutional autoencoder branch can more sensitively capture structural, sequential, and high-order abnormal signals and is suitable for processing complex time-series and dynamic feature anomalies. A non-linear weighted combination is performed on the joint feature anomaly probability score and the joint feature reconstruction error. The two criteria are normalized, and then the discriminant information of the two is synthesized into a unified target anomaly score through methods such as softmax normalization. The higher the value of this anomaly score, the greater the likelihood that the observed data is abnormal; conversely, it indicates a stronger consistency with normal samples.

[0063] In a specific embodiment, the process of executing step S4 may specifically include the following steps:

[0064] Based on the abnormal event determination result, perform multi-camera spatial positioning analysis, use the overlapping fields of view of adjacent cameras for triangulation and spatial coordinate mapping to obtain the target three-dimensional coordinate position information of the abnormal event;

[0065] According to the target three-dimensional coordinate position information, perform target area collaborative tracking on the video stream data collected by multiple cameras to obtain the cross-camera abnormal target consistency verification result;

[0066] Perform a correlation calculation on the cross-camera abnormal target consistency verification result and the environmental parameter abnormal pattern in the abnormal event determination result to obtain cross-modal abnormal correlation features;

[0067] Based on the cross-modal abnormal correlation features, construct multi-dimensional abnormal event description information including abnormal type identifiers, abnormal spatial position coordinates, abnormal intensity values, abnormal duration statistics, and abnormal credibility scores.

[0068] Specifically, when a suspected abnormal event occurs in a certain monitoring area, according to the multi-camera layout, the calibration results of the internal and external camera parameters, and the scene geometry prior, multiple cameras in the central area of the event and its surrounding area are locked, and a camera view group with overlapping fields of view and complementary observation angles is preferentially selected. At this time, the system performs synchronous time alignment and spatial geometric registration on the video stream data collected by these cameras. Through manual calibration, automatic feature point extraction, and camera projection model establishment, the target pixel coordinates in different views are projected back to a unified three-dimensional world coordinate system. Based on the pixel positions of the target in the images of different cameras and the camera parameters (such as focal length, principal point coordinates, distortion coefficients, and spatial poses, etc.), the three-dimensional coordinates of the target in the actual space are deduced by using the triangulation algorithm. This algorithm comprehensively utilizes the projection information of the same target observed by at least two cameras at different positions on their respective imaging planes, and completes the high-precision estimation of the target's spatial position by calculating the intersection point of the lines of sight or the point of minimum distance. Based on the redundant observations of multiple cameras, the accuracy and robustness of the three-dimensional positioning are improved through weighted least squares method and multi-view consistency optimization, and the three-dimensional coordinate position information of the abnormal event target is output. According to the three-dimensional spatial coordinate information of the target, all cameras covering this area are jointly scheduled and regionally divided, and for the video stream collected by each camera, the region of interest corresponding to the three-dimensional projection of the abnormal target is automatically extracted and focused on. During this process, using spatio-temporal consistency constraints, the dynamic behavior trajectories of the target in the fields of view of each camera are tracked in real time through motion detection, target segmentation, and tracking algorithms (such as multi-object Kalman filtering, SORT / DeepSORT, multi-camera spatio-temporal joint tracking network, etc.). On this basis, by comprehensively considering the tracking results feedback by each camera, through multi-dimensional criteria such as spatial position consistency, appearance feature similarity, behavior pattern matching, and environmental background continuity, the identity consistency and spatio-temporal coherence of the abnormal target in the multi-camera system are verified, and false judgments and tracking interruptions caused by factors such as perspective occlusion, occluder interference, and light changes are excluded. Through multi-view spatial collaborative tracking, rich information such as the global motion path, staying area, and behavior trajectory of the target in the monitoring scene is obtained. The correlation between the cross-camera abnormal target consistency verification result and the environmental parameter abnormal pattern in the abnormal event determination result is calculated. On the one hand, the system constructs visual behavior feature vectors such as the motion trajectory map and activity heat distribution of the target in time and space, and on the other hand, it performs temporal modeling on environmental parameters (such as temperature, humidity, smoke concentration, gas indicators, etc.) to form the environmental abnormal change pattern corresponding to the time period and spatial region. Through correlation analysis algorithms (such as Pearson correlation coefficient, mutual information analysis, dynamic time warping DTW, causal relationship modeling, etc.), the temporal synchronization, spatial coupling degree, and causal influence relationship between the abnormal target behavior and the environmental parameter abnormality are quantitatively characterized.For example, when analyzing whether there are environmental anomalies such as sudden changes in smoke concentration and temperature in the corresponding area when an abnormal target at a certain three-dimensional space position performs a specific behavior, so as to distinguish between "independent behavior anomalies" and "complex anomalies highly coupled with environmental changes". After completing the above cross-modal anomaly correlation modeling, construct multi-dimensional anomaly event description information based on the correlation features. Assign anomaly type identifiers according to the initial determination of the anomaly event and the model output results, such as classifying into categories such as "environmental anomaly", "behavior anomaly", "complex anomaly" or "uncertain anomaly"; then directly extract the three-dimensional space coordinates obtained from the aforementioned triangulation and spatial positioning analysis as the anomaly spatial position coordinate attributes, reflecting the actual occurrence point or activity area of the event in the monitoring scenario; in addition, the system comprehensively quantifies the target abnormal behavior and environmental anomaly parameters, constructs an anomaly intensity value based on model scores, behavior deviation amplitudes, environmental parameter fluctuation amplitudes, etc., and calibrates the severity level of the anomaly phenomenon; for the statistics of the anomaly duration, the system continuously tracks the evolution interval of the target and environmental anomaly states, and accumulatively calculates the total duration from the first determination of the anomaly to the elimination or return to normal of the event to obtain the anomaly duration distribution. Combining comprehensive information such as cross-camera consistency verification results, correlation analysis scores, and model adaptive confidence, assign a credibility score to the anomaly event, which reflects the overall certainty of the system about the true anomaly attribute of the event. Organize all multi-dimensional descriptions into structured data to form an anomaly event description report covering anomaly types, three-dimensional coordinates, intensity values, durations, and credibility scores.

[0069] In a specific embodiment, the process of performing collaborative tracking of the target area on the video stream data collected by multiple cameras according to the target three-dimensional coordinate position information to obtain the cross-camera abnormal target consistency verification result may specifically include the following steps:

[0070] Perform multi-modal background perception area segmentation on the target three-dimensional coordinate position information to obtain multiple environmentally characteristic independent areas, and assign environmental characteristic weights to each environmentally characteristic independent area to obtain an environmentally sensitive area division map;

[0071] Construct a multi-dimensional state transition model based on the environmentally sensitive area division map and the target three-dimensional coordinate position information, and perform deep spatio-temporal feature extraction on the target area in the video stream data collected by multiple cameras to obtain an environmentally enhanced target representation feature set;

[0072] Input the environmentally enhanced target representation feature set and the multi-dimensional state transition model into a non-linear extended Kalman filter to perform multi-trajectory prediction to obtain multi-path trajectory prediction results;

[0073] Generate a cross-camera target migration probability matrix based on the multi-path trajectory prediction results, and execute the Hungarian matching algorithm in combination with the camera topological relationship to obtain a multi-level camera target identity matching scheme;

[0074] Perform spatio-temporal consistency constraints and environmental parameter verification on the multi-level camera target identity matching scheme to obtain the cross-camera abnormal target consistency verification result.

[0075] Specifically, a three-dimensional model of the entire monitoring scenario is built, mapping the three-dimensional coordinates of the target onto the geographical space and environmental topology of the actual scenario. Based on different physical characteristics, functional partitions, and environmental monitoring data within the scenario, the monitoring area is divided into multiple small areas with independent environmental characteristics. The area segmentation can be carried out according to the spatial distribution. For example, a factory area can be divided into a production area, a storage area, a living area, and a dangerous goods storage area. By integrating the results of multi-modal perception, including the spatial variation characteristics of environmental parameters such as temperature, humidity, light, gas concentration, noise intensity, and population density, and using spatial clustering algorithms, region segmentation networks, or multi-modal information entropy distribution models, environmental characteristic areas with significant independence both physically and environmentally are generated. For each independent area, the system dynamically assigns environmental characteristic weights according to its historical risks, sensor monitoring results, and actual business requirements. These weights reflect the importance of the area for anomaly detection and system decision-making. Combining the spatial boundaries, environmental characteristic weights, and three-dimensional coordinate distributions of each area, an environmental sensitive area division map is formed. Based on the environmental sensitive area division map and the target three-dimensional coordinate position information, a multi-dimensional state transition model is constructed. Taking the movement path of the target in the three-dimensional space as the main line, the dynamic state of the target (such as speed, acceleration, direction, behavior pattern, etc.) and the environmental state (such as local temperature, gas concentration, light intensity change, etc.) are jointly embedded in the model. Through Markov processes, Bayesian networks, or deep spatio-temporal graph networks, the state transition probability, behavior switching pattern, and environmental response mechanism of the target in different environmental areas are globally modeled to achieve multi-dimensional prediction of the target's cross-regional and cross-environmental property movement behaviors. On the basis of multi-dimensional state modeling, deep spatio-temporal feature extraction is performed on the ROI areas corresponding to the target three-dimensional coordinate projections in the video streams collected by multiple cameras. The specific process uses 3D convolutional neural networks, spatio-temporal transformation networks, or multi-modal fusion networks to uniformly extract information such as the appearance features, movement trajectories, behavior postures, and environmental responses of the target at different time periods and different camera perspectives into a high-dimensional environment-enhanced target characterization feature set. These feature sets reflect the movement patterns and appearance characteristics of the target itself and integrate the dynamic change information of the local environment, realizing the combination of vision and environmental data. The environment-enhanced target characterization feature set and the aforementioned multi-dimensional state transition model are input into a non-linear extended Kalman filter. As a classic dynamic system state estimation algorithm, the extended Kalman filter can integrate the target historical observations, environmental state, and sensor noise to recursively estimate the temporal movement trajectory of the target and model the uncertainty, and synchronously output multiple predicted trajectory branches in multi-target and multi-path scenarios. Through multi-path trajectory prediction, the system can probabilistically deduce the possible movement states and spatial distributions of the target at future moments and quantify the divergence and potential risks of the target behavior under various environmental states.After obtaining the multi-path trajectory prediction results, based on the trajectory information and combined with the spatial topological relationship of the camera layout, calculate the cross-camera target transfer probability matrix, which quantifies the spatial, temporal, and feature continuity of the target transferring from one camera field of view to another. Using the Hungarian matching algorithm, perform a globally optimal assignment for all possible target identity transfer schemes, assign cost weights to each group of identity matches between cameras according to the spatial overlap degree, appearance feature similarity, behavioral dynamic consistency, and environmental response conformity of the multi-path trajectories, and finally achieve a globally optimal and multi-level camera target identity matching scheme through the Hungarian algorithm. Apply spatio-temporal consistency constraints and environmental parameter verification to the above multi-level identity matching scheme. Use various means such as the spatio-temporal behavior continuity detection algorithm of the target, the appearance and action timing alignment mechanism, and the environmental parameter distribution consistency test to cross-verify the target trajectories under each camera, ensuring that the identity of the same physical target is not confused between cameras, the trajectories do not jump, and the parameter changes and abnormal states in the environmentally sensitive area are closely related. Only when multiple consistency constraints and environmental parameter verification are passed, the system regards the target as a cross-camera consistent abnormal target.

[0076] In a specific embodiment, the process of performing correlation calculation on the environmental parameter abnormal patterns in the cross-camera abnormal target consistency verification results and the abnormal event determination results to obtain cross-modal abnormal correlation features may specifically include the following steps:

[0077] Construct an abnormal target spatio-temporal trajectory graph according to the cross-camera abnormal target consistency verification results, and construct an environmental abnormal spatio-temporal distribution graph according to the environmental parameter abnormal pattern in the abnormal event determination results;

[0078] Fuse the abnormal target spatio-temporal trajectory graph and the environmental abnormal spatio-temporal distribution graph to obtain a multi-modal temporal correlation graph, and perform structured feature extraction on the multi-modal temporal correlation graph to obtain an abnormal event correlation feature vector;

[0079] Input the abnormal event correlation feature vector into the modal correlation metric model, calculate the mutual information and conditional probability between visual abnormality and environmental abnormality, and obtain a cross-modal abnormal cooperation index;

[0080] Analyze the temporal influence relationship and diffusion law of abnormal events between different modalities based on the cross-modal abnormal cooperation index to obtain abnormal propagation characteristic information;

[0081] Perform causal relationship analysis on the abnormal propagation characteristic information to obtain the causal relationship of abnormal events, and construct cross-modal abnormal correlation features according to the causal relationship of abnormal events and the cross-modal abnormal cooperation index.

[0082] Specifically, based on the results of multi-camera collaborative tracking and consistency verification, a spatio-temporal trajectory map of abnormal targets is constructed. This trajectory map uses the spatial coordinates, activity states, behavior labels, etc. of abnormal targets at different times in three-dimensional space as nodes, and the movement paths, behavior switches, and state transitions of the targets between consecutive times as edges. Through the graph structure, the dynamic evolution process of abnormal targets is completely mapped onto the time axis and the spatial coordinate system. The spatio-temporal trajectory map records the starting point of the target's movement, the areas passed through, the stopping points, and the points where abnormalities occur, and reflects the changes in the behavior patterns, sudden speed changes, and spatial diffusion trends of abnormal targets in different monitoring areas. At the same time, based on the abnormal patterns of environmental parameters in the abnormal event determination results, an environmental abnormal spatio-temporal distribution map is constructed. The environmental abnormal spatio-temporal distribution map uses time slices, spatial grid points, or regions as units, and projects the abnormal detection results of multiple parameters such as temperature, humidity, air pressure, light, smoke concentration, and harmful gases onto a unified environmental grid according to the occurrence time and spatial coordinates. Each grid unit includes the type, intensity, and duration of parameter abnormalities at this spatio-temporal point, and combines spatial topology and historical evolution information to form the spatial spread path, influence range, and high-risk intervals of environmental abnormalities, thereby revealing the spatio-temporal distribution law and dynamic evolution trajectory of environmental abnormal events. Perform multi-modal temporal correlation fusion on the abnormal target spatio-temporal trajectory map and the environmental abnormal spatio-temporal distribution map to generate a unified multi-modal temporal correlation map. Specific fusion strategies include joint node mapping (such as the merger of target behavior and environmental abnormalities at the same spatio-temporal point), edge weight coupling (such as the joint diffusion probability of abnormal events between adjacent spatio-temporal nodes), and modal attribute complementarity (such as the joint feature supplementation of target movement attributes and environmental change attributes), etc. Through the comprehensive modeling of the multi-modal attributes of spatio-temporal nodes and associated edges, the dynamic collaborative description between target behavior and environmental changes is realized. Perform structured feature extraction on the multi-modal temporal correlation map, and use graph neural networks, graph embedding algorithms, or manually designed graph feature engineering to extract high-dimensional features such as spatial aggregation, path connectivity, cross-modal synchronization, and propagation rate of abnormal events from the graph structure, forming an abnormal event correlation feature vector for downstream discrimination and analysis. Input the abnormal event correlation feature vector into the modal correlation metric model to calculate statistical indicators such as the mutual information content and conditional probability between visual abnormalities and environmental abnormalities. The mutual information content measures the coupling strength between two modal abnormalities, and the conditional probability reflects the probability level of the occurrence of another modal abnormality when one modal abnormality occurs. These cross-modal abnormal collaboration indicators can reflect the interaction and linkage mechanism of abnormal events in multi-source perception data. For example, if the occurrence of abnormal target movement is highly synchronized with the sudden change in gas concentration, the mutual information content between the two will be extremely high, indicating that the abnormal event has a strong composite characteristic. Analyze the temporal influence relationship and diffusion law of abnormal events between different modalities based on the cross-modal abnormal collaboration indicators.By constructing methods such as the time series transition probability matrix, propagation time delay analysis, and propagation path tracking, reveal how visual anomalies are transmitted over time to the changes in environmental parameters, or conversely, how environmental anomalies induce behavioral anomalies, and identify the temporal diffusion characteristics of abnormal events, including the initial point, outbreak moment, diffusion speed, and extinction trend, etc. On this basis, perform causal relationship analysis on the information of abnormal propagation characteristics, and with the help of algorithms such as Granger causality test, determine the causal sequence, action direction, and influence mechanism between environmental anomalies and target anomalies at the time series level and event logic level. For example, identify that gas leakage (environmental anomaly) occurs prior to the regional personnel gathering (behavioral anomaly), or find that the intrusion of suspicious personnel (behavioral anomaly) induces a sharp fluctuation in environmental parameters. Based on the causal relationship of abnormal events and cross-modal anomaly collaboration indicators, comprehensively construct cross-modal anomaly association features.

[0083] In a specific embodiment, the process of executing step S5 may specifically include the following steps:

[0084] Perform camera role allocation according to the anomaly type, anomaly credibility, and anomaly spatial position coordinates in the multi-dimensional abnormal event description information, and divide the cameras in the monitoring network into core monitoring cameras, auxiliary monitoring cameras, and environmental perception cameras;

[0085] Perform optimal viewing angle calculation on the core monitoring cameras to obtain precise positioning instructions for the core monitoring cameras;

[0086] According to the monitoring viewing angle of the core monitoring cameras and the spatial distribution characteristics of the abnormal areas, calculate complementary monitoring angles for the auxiliary monitoring cameras to obtain a multi-camera linkage response strategy;

[0087] Based on the multi-camera linkage response strategy, differentially allocate the image analysis tasks according to the camera roles and computing capabilities to obtain multi-level visual analysis instructions;

[0088] Based on the multi-level visual analysis instructions, perform multi-view geometric reconstruction of the images of the core monitoring cameras and the auxiliary monitoring cameras, construct a stereoscopic vision model, and fuse the environmental context information of the environmental perception cameras into the stereoscopic vision model to obtain a stereoscopic monitoring view of the abnormal area;

[0089] Execute hierarchical control of the camera roles according to the stereoscopic monitoring view of the abnormal area to obtain the execution result of the linkage response.

[0090] Specifically, based on the multi-dimensional abnormal event description information, role allocation is performed for all cameras in the network. Key indicators such as abnormal type, abnormal credibility, and abnormal spatial position coordinates are comprehensively analyzed. Among them, the abnormal type (such as environmental abnormality, behavior abnormality, compound abnormality, etc.) directly affects the subsequent monitoring response level and focus of attention. The abnormal credibility determines the priority of linkage scheduling, while the spatial position coordinates provide spatial constraints for the selection of area cameras and perspective optimization. Through the scheduling algorithm, cameras within the coverage range of the abnormal spatial coordinates and with excellent observation conditions are preferentially allocated as core monitoring cameras, and are assigned high-priority tasks such as in-depth observation of the main abnormal targets or event areas, fine-grained feature extraction, target tracking, and intelligent recognition; for devices covering the edge of the abnormal area or providing complementary views for the core cameras, they are allocated as auxiliary monitoring cameras, mainly responsible for auxiliary tasks such as perspective supplementation, occlusion elimination, multi-source collaborative verification, and area linkage; while the sensing cameras located outside the event and dedicated to collecting environmental parameters (such as temperature and humidity, gas, light, etc.) are designated as environmental perception cameras, providing environmental context and physical background data when abnormal events occur. Perform the optimal perspective calculation for the core monitoring cameras. This process is based on the three-dimensional spatial coordinates of the target, the spatial distribution of the abnormal area, the physical installation position and attitude parameters of the current camera. Through three-dimensional spatial geometric modeling and reverse optimization, comprehensively considering the field of view angle, zoom ability, pitch / rotation range, imaging resolution, etc., and using optimization algorithms (such as minimum occlusion optimization based on gradient descent, maximum information gain optimization, minimum target distortion optimization, etc.) to deduce the precise perspective and positioning parameters that the camera needs to adjust, including specific instructions such as pitch angle, yaw angle, focal length, aperture, focus distance, etc., and automatically generate control instructions executable by the device end to achieve precise scheduling of the core camera. At the same time, analyze the matching relationship between the core camera and the spatial distribution of the abnormal area, and calculate the optimal deployment plan for the auxiliary monitoring cameras for the monitoring blind spots and weak edge coverage areas in the perspective of the core camera. For each auxiliary camera, based on the spatial overlap relationship and complementary angle between the current field of view and the core camera, use the area division algorithm and the minimum field of view overlap strategy to deduce the best monitoring compensation angle, form a spatial cooperation grid among multiple cameras, and form a multi-camera linkage response strategy including core and auxiliary devices. Based on the above strategy, combined with resource constraints such as the actual computing power of the camera, edge computing power distribution, and communication bandwidth, perform multi-level and differentiated allocation of image analysis tasks.The core monitoring camera mainly undertakes tasks such as in-depth visual analysis and intelligent recognition, including object detection, abnormal behavior analysis, video object segmentation, and temporal feature modeling. The auxiliary camera undertakes low-latency and lightweight tasks such as target re-identification, trajectory tracking, auxiliary feature matching, and multi-source verification. The environmental perception camera focuses on the collection and preprocessing of environmental data, environmental change analysis, and feature fusion of parameter events. Based on this, the system generates hierarchical visual analysis instructions to achieve the optimal configuration of various tasks in terms of space, computing power, bandwidth, and real-time performance. The core and auxiliary cameras synchronously collect and upload their respective high-quality video data according to the visual analysis instructions. The system uses multi-view geometric reconstruction algorithms (such as multi-view stereo reconstruction MVS, spatial point cloud fusion, depth map superposition, etc.) at the central processing end or distributed edge nodes to perform three-dimensional fusion of multi-source video data, comprehensively extract the depth information, spatial structure, motion trajectory, and behavior changes of the target area from different perspectives, and then construct a three-dimensional visual model of the abnormal area. At the same time, during the generation of the three-dimensional visual model, physical environment parameters (such as environmental temperature, gas concentration, light distribution, etc.) collected by the environmental perception camera are introduced. Through the multi-modal feature fusion mechanism, the environmental context information is superimposed into the three-dimensional spatial structure and target behavior model to obtain a three-dimensional monitoring view of the abnormal area that reflects the deep integration of spatial dynamics, environmental evolution, and target behavior. Based on the three-dimensional monitoring view of the abnormal area, hierarchical control of the camera roles is implemented. The core camera realizes automatic zooming, dynamic tracking, and real-time angle adjustment according to the target behavior changes and model feedback. The auxiliary camera automatically switches the focus and tracking object according to the main perspective compensation and target migration trend. The environmental perception camera dynamically adjusts the sampling frequency, perception mode, or warning threshold according to the abnormal intensity. Inside the system, through the real-time data bus and linkage protocol, data sharing, task collaboration, and priority switching among various cameras are coordinated, and finally an intelligent response closed-loop driven by abnormal events, based on the three-dimensional space, and centered on multi-modal fusion is formed, and an efficient and stable linkage response execution result is output.

[0091] The above describes the camera linkage alarm method for intelligent environmental monitoring in the embodiments of the present invention. Next, the camera linkage alarm system for intelligent environmental monitoring in the embodiments of the present invention will be described. Please refer to Figure 2 , an embodiment of the camera linkage alarm system for intelligent environmental monitoring in the embodiments of the present invention includes:

[0092] A synchronous acquisition module for synchronously acquiring video stream data of multiple cameras and multi-parameter monitoring data of environmental sensors;

[0093] A feature fusion module for performing multi-channel feature analysis and cross-modal feature fusion on the video stream data and multi-parameter monitoring data to obtain environment-vision joint features;

[0094] Anomaly determination module, which is used to perform environment scene adaptive classification according to the environment-vision joint features, select the corresponding anomaly determination model for the current monitoring area, and analyze the environment-vision joint features through the anomaly determination model to obtain the anomaly event determination result;

[0095] Association analysis module, which is used to perform multi-camera collaborative verification and vision-environment data association analysis based on the anomaly event determination result to obtain multi-dimensional anomaly event description information;

[0096] Stereo monitoring module, which is used to select the corresponding multi-camera linkage response strategy according to the multi-dimensional anomaly event description information and coordinate multiple cameras to perform stereo monitoring of the anomaly area to obtain the linkage response execution result.

[0097] Through the collaborative cooperation of the above-mentioned various components, through the synchronous acquisition and cross-modal feature fusion of the video stream data of multiple cameras and the multi-parameter monitoring data of environmental sensors, the in-depth association analysis of visual information and environmental parameters is realized, overcoming the limitations of single data source analysis, and significantly improving the detection accuracy of the system for anomaly events in complex environments. The scene adaptive classification mechanism based on the environment-vision joint features enables the system to automatically identify the scene type of the current monitoring environment and select an anomaly determination model optimized for a specific scene, solving the problem that traditional fixed-threshold systems are difficult to adapt to environmental changes. Through multi-camera collaborative verification and vision-environment data association analysis, the system can generate anomaly event descriptions containing multi-dimensional information such as anomaly type, location, intensity, duration, and credibility, providing a comprehensive basis for subsequent alarm response decisions. By adopting the role assignment mechanism of the core monitoring camera, auxiliary monitoring camera, and environmental perception camera, combined with the optimal viewing angle calculation and complementary angle configuration, the all-round stereo monitoring of the anomaly area is realized, eliminating the visual blind area in traditional single-camera monitoring. The system can analyze the potential development trajectory of the anomaly event based on its spatio-temporal characteristics and environmental factors, adjust the monitoring parameters of inactive cameras in advance, form a pre-defensive monitoring network, and realize the dynamic tracking and predictive coverage of the anomaly event. Implement multi-level linkage response according to the characteristics of the anomaly event, adopt different camera parameter adjustments, video acquisition strategies, and alarm notification methods for anomalies at different levels, improving the accuracy and efficiency of the system response. Through the continuous analysis of the linkage response execution result and the processing feedback, the system can identify false alarm patterns and dynamically adjust the anomaly determination rules, realizing the continuous optimization of the system performance and reducing the false alarm rate and missed alarm rate during long-term operation.

[0098] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system, system, and unit can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0099] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0100] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A camera linkage alarm method for intelligent environment monitoring, characterized in that: include: Synchronously collect video stream data from multiple cameras and multi-parameter monitoring data from environmental sensors; Performing multi-channel feature analysis and cross-modal feature fusion on the video stream data and the multi-parameter monitoring data to obtain an environment-vision joint feature; Performing adaptive classification of environmental scenes according to the environmental-visual joint features, selecting a corresponding abnormality determination model for the current monitoring area, and analyzing the environmental-visual joint features through the abnormality determination model to obtain an abnormal event determination result; Based on the abnormal event determination result, multi-camera collaborative verification and visual-environment data association analysis are performed to obtain multi-dimensional abnormal event description information; According to the multi-dimensional abnormal event description information, a corresponding multi-camera linkage response strategy is selected and multiple cameras are coordinated to perform stereoscopic monitoring of the abnormal area to obtain a linkage response execution result.

2. The camera linkage alarm method for intelligent environment monitoring according to claim 1 is characterized in that: The synchronous collection of video stream data from multiple cameras and multi-parameter monitoring data from environmental sensors includes: Configure the video stream acquisition parameters of multiple cameras, and set the sampling frequency of temperature, humidity, air pressure, light intensity, smoke concentration, harmful gas concentration, and motion state sensors in the distributed environmental sensor network to obtain a multi-source environmental parameter data acquisition strategy; Collecting original videos from the multiple cameras according to the video stream acquisition parameters, and performing histogram equalization processing, Gaussian filtering processing and perspective transformation correction on the original videos to obtain video stream data; The original multi-parameter data is collected according to the multi-source environmental parameter data collection strategy and the distributed environmental sensor network, and the original multi-parameter data is subjected to sliding window median filtering, Z-score normalization processing and linear interpolation supplementation to obtain multi-parameter monitoring data.

3. The camera linkage alarm method for intelligent environment monitoring according to claim 1 is characterized in that: The performing multi-path feature analysis and cross-modal feature fusion on the video stream data and the multi-parameter monitoring data to obtain an environment-vision joint feature includes: Inputting the video stream data into a three-dimensional convolutional neural network for three-dimensional convolution processing and spatiotemporal attention calculation to obtain a video spatiotemporal feature vector; The multi-parameter monitoring data is input into a bidirectional long short-term memory network, and an inter-sensor attention mechanism is introduced to assign dynamic weights to different sensor data to obtain a time series feature vector of the environmental parameters; Performing dimension mapping on the video spatiotemporal feature vector and the environmental parameter temporal feature vector to obtain a unified scale feature vector; Performing a bidirectional cross attention calculation on the unified-scale feature vector to obtain an inter-modal correlation enhancement feature, and performing a nonlinear feature transformation and complementary filling on the inter-modal correlation enhancement feature to obtain a fused feature vector; Feature selection is performed on the fused feature vector according to an information gain criterion to obtain an environment-vision joint feature.

4. The camera linkage alarm method for intelligent environment monitoring according to claim 1 is characterized in that: The performing of adaptive classification of environmental scenes according to the environmental-visual joint features, selecting a corresponding abnormality determination model for the current monitoring area, and analyzing the environmental-visual joint features by the abnormality determination model to obtain an abnormal event determination result includes: Inputting the environment-vision joint feature into a densely connected network to perform scene type analysis to obtain the environment scene type of the current monitoring area; Selecting an abnormality determination model for the current environmental scene from a pre-trained model library according to the environmental scene type, the abnormality determination model includes a Gaussian mixture model and a one-dimensional convolutional autoencoder; the pre-trained model library contains abnormality determination models pre-trained for various environmental scene types, and the number of components, covariance matrix structure of the Gaussian mixture model corresponding to different environmental scene types, and the number of network layers and convolution kernel parameters of the one-dimensional convolutional autoencoder are all different; Inputting the environment-visual joint feature into the abnormality determination model, calculating the log-likelihood probability of the environment-visual joint feature under the Gaussian mixture model and the reconstruction error of the one-dimensional convolutional autoencoder, and obtaining a target abnormality score; A change point detection algorithm is executed to monitor the state of environmental scene conversion, and when a scene conversion is detected, a temporary threshold adjustment and anomaly type classification are performed on the target anomaly score to obtain an abnormal event determination result.

5. The camera linkage alarm method for intelligent environment monitoring according to claim 4 is characterized in that: The step of inputting the environment-visual joint feature into the abnormality determination model, calculating the log-likelihood probability of the environment-visual joint feature under the Gaussian mixture model and the reconstruction error of the one-dimensional convolutional autoencoder, and obtaining a target abnormality score includes: Setting a first weight coefficient of an environment channel according to the environment scene type, and performing environment weight allocation on the environment-vision joint feature according to the first weight coefficient to obtain a first environment-vision feature vector; Setting a second weight coefficient of the visual channel according to the environment scene type, and performing visual weight allocation on the environment-vision joint feature according to the second weight coefficient to obtain a second environment-vision feature vector; Inputting the first environment-visual feature vector into the Gaussian mixture model of the abnormality determination model, calculating the multidimensional probability density value and performing a logarithmic transformation to obtain a joint feature abnormality probability score; Inputting the second environment-visual feature vector into the one-dimensional convolutional autoencoder of the abnormality determination model at the same time, generating reconstruction features through an encoding-decoding process, and obtaining a feature reconstruction result; A joint feature reconstruction error between the feature reconstruction result and the second environment-visual feature vector is calculated, and a nonlinear weighted combination is performed on the joint feature anomaly probability score and the joint feature reconstruction error to obtain a target anomaly score.

6. The camera linkage alarm method for intelligent environment monitoring according to claim 1 is characterized in that: The performing of multi-camera collaborative verification and visual-environment data association analysis based on the abnormal event determination result to obtain multi-dimensional abnormal event description information includes: Based on the abnormal event determination result, multi-camera spatial positioning analysis is performed, and overlapping fields of view of adjacent cameras are used for triangulation and spatial coordinate mapping to obtain target three-dimensional coordinate position information of the abnormal event; Performing target area collaborative tracking on video stream data collected by multiple cameras according to the target three-dimensional coordinate position information to obtain a consistency verification result of abnormal targets across cameras; Performing correlation calculation on the abnormal target consistency verification result across cameras and the abnormal mode of environmental parameters in the abnormal event determination result to obtain a cross-modal abnormal correlation feature; Based on the cross-modal anomaly association features, multi-dimensional abnormal event description information including anomaly type identifier, abnormal spatial position coordinates, abnormal intensity value, abnormal duration statistics and abnormal credibility score is constructed.

7. The camera linkage alarm method for intelligent environment monitoring according to claim 6 is characterized in that: The performing target area collaborative tracking on the video stream data collected by multiple cameras according to the target three-dimensional coordinate position information to obtain the consistency verification result of the abnormal target across cameras includes: Performing multimodal background perception area segmentation on the target three-dimensional coordinate position information to obtain a plurality of independent environmental characteristic areas, and assigning an environmental characteristic weight to each independent environmental characteristic area to obtain an environmental sensitive area division map; Constructing a multidimensional state transition model according to the environmental sensitive area division map and the target three-dimensional coordinate position information, and performing deep spatiotemporal feature extraction on the target area in the video stream data collected by the multiple cameras to obtain an environmentally enhanced target representation feature set; Inputting the target characterization feature set of the environment enhancement and the multi-dimensional state transition model into a nonlinear extended Kalman filter to perform multi-trajectory prediction to obtain a multi-path trajectory prediction result; Generate a cross-camera target migration probability matrix according to the multi-path trajectory prediction results, and execute the Hungarian matching algorithm in combination with the camera topological relationship to obtain a multi-level camera target identity matching solution; The multi-level camera target identity matching scheme is subjected to spatiotemporal consistency constraints and environmental parameter verification to obtain a cross-camera abnormal target consistency verification result.

8. The camera linkage alarm method for intelligent environment monitoring according to claim 6 is characterized in that: The performing correlation calculation on the abnormal target consistency verification result across cameras and the abnormal mode of the environmental parameters in the abnormal event determination result to obtain the cross-modal abnormal correlation feature includes: Constructing a spatiotemporal trajectory diagram of an abnormal target according to the abnormal target consistency verification result across cameras, and constructing a spatiotemporal distribution diagram of an environmental anomaly according to an abnormal pattern of environmental parameters in the abnormal event determination result; The abnormal target spatiotemporal trajectory map and the environmental anomaly spatiotemporal distribution map are fused to obtain a multimodal time series association map, and structured feature extraction is performed on the multimodal time series association map to obtain an abnormal event association feature vector; Inputting the abnormal event association feature vector into the modal association measurement model, calculating the mutual information and conditional probability between the visual anomaly and the environmental anomaly, and obtaining the cross-modal anomaly coordination index; Based on the cross-modal anomaly coordination index, the temporal influence relationship and diffusion law of abnormal events between different modes are analyzed to obtain abnormal propagation characteristic information; A causal relationship analysis is performed on the abnormal propagation characteristic information to obtain the causal relationship of the abnormal event, and a cross-modal abnormality association feature is constructed according to the causal relationship of the abnormal event and the cross-modal abnormality collaboration indicator.

9. The camera linkage alarm method for intelligent environment monitoring according to claim 1, characterized in that: The step of selecting a corresponding multi-camera linkage response strategy according to the multi-dimensional abnormal event description information and coordinating multiple cameras to perform stereoscopic monitoring of the abnormal area to obtain a linkage response execution result includes: Perform camera role allocation according to the abnormal type, abnormal credibility and abnormal spatial position coordinates in the multi-dimensional abnormal event description information, and divide the cameras in the monitoring network into core monitoring cameras, auxiliary monitoring cameras and environmental perception cameras; Calculating the optimal viewing angle for the core surveillance camera to obtain precise positioning instructions for the core surveillance camera; According to the monitoring angle of the core monitoring camera and the spatial distribution characteristics of the abnormal area, a complementary monitoring angle is calculated for the auxiliary monitoring camera to obtain a multi-camera linkage response strategy; Based on the multi-camera linkage response strategy, the image analysis tasks are differentially allocated according to the camera roles and computing capabilities to obtain multi-level visual analysis instructions; Based on the multi-level visual analysis instruction, perform multi-view geometric reconstruction of images on the core monitoring camera and the auxiliary monitoring camera, build a stereo vision model, and fuse the environmental context information of the environmental perception camera into the stereo vision model to obtain a stereo monitoring view of the abnormal area; Camera role hierarchical control is performed according to the abnormal area stereoscopic monitoring view to obtain a linkage response execution result.

10. A camera linkage alarm system for intelligent environmental monitoring, characterized in that: The camera linkage alarm method for intelligent environment monitoring according to any one of claims 1 to 9, wherein the camera linkage alarm system for intelligent environment monitoring comprises: Synchronous acquisition module, used to synchronously acquire video stream data from multiple cameras and multi-parameter monitoring data from environmental sensors; A feature fusion module, used to perform multi-channel feature analysis and cross-modal feature fusion on the video stream data and the multi-parameter monitoring data to obtain an environment-vision joint feature; An abnormality determination module, used to perform adaptive classification of environmental scenes according to the environmental-visual joint features, select a corresponding abnormality determination model for the current monitoring area, and analyze the environmental-visual joint features through the abnormality determination model to obtain an abnormal event determination result; A correlation analysis module, used to perform multi-camera collaborative verification and visual-environment data correlation analysis based on the abnormal event determination result to obtain multi-dimensional abnormal event description information; The stereo monitoring module is used to select the corresponding multi-camera linkage response strategy according to the multi-dimensional abnormal event description information and coordinate multiple cameras to perform stereo monitoring of the abnormal area to obtain the linkage response execution result.

Citation Information

Cited By

  • Real-time quality detection method and system based on 5G edge calculation

    CN120355707A

  • Video monitoring system, video monitoring method and spherical camera

    CN120358330A

  • Monitoring and early warning analysis method based on artificial intelligence and server

    CN120408383A

  • Multi-defense-area intelligent linkage alarm method based on AIoT gateway and related equipment

    CN120431699A

  • Monitoring image quality diagnosis method and system based on weather perception

    CN120451690A