Federal vision Transform-based video anomaly detection method, device and equipment and medium
By using a video anomaly detection method based on federated vision Transformer, edge nodes are used to process multi-source video data and perform anomaly detection. This solves the problems of weak privacy protection and insufficient cross-scene generalization ability, and improves the accuracy and real-time performance of video anomaly detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-25
- Publication Date
- 2026-04-03
AI Technical Summary
Existing video anomaly detection technologies suffer from weak privacy protection and insufficient cross-scenario generalization capabilities, making it difficult to meet the dual requirements of data security and real-time performance in the energy, power, and industrial sectors.
A video anomaly detection method based on federated visual Transformer is adopted. Multi-source video data is obtained through edge nodes for preprocessing, block processing and encoding. Anomaly detection is performed using a multi-scale visual Transformer model, and anomaly evaluation and alarm operations are performed based on the anomaly response map.
It improves the accuracy, real-time performance, and stability of video anomaly detection, solves the problems of weak privacy protection and insufficient cross-scenario generalization ability, and achieves high-precision detection and low-latency response in multi-source video surveillance scenarios.
Smart Images

Figure CN121789132A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent video detection technology, and in particular to a video anomaly detection method, apparatus, device, and medium based on Federated Vision Transformer. Background Technology
[0002] With the rapid development of artificial intelligence and video surveillance technology, deep learning-based visual anomaly detection has been widely used in fields such as power, manufacturing, transportation and public safety.
[0003] In recent years, deep learning-based anomaly detection methods have gradually become a core technology of intelligent monitoring systems. In particular, models such as Convolutional Neural Networks (CNNs) and Transformers have significantly improved performance in image and video understanding, enabling monitoring systems to achieve automatic identification, event detection, and risk warning at the visual level. However, most existing intelligent video analysis systems adopt a centralized training architecture, requiring the aggregation of multi-source video data to a central server for unified modeling. This approach not only brings enormous communication and storage pressure but also poses serious privacy and security risks, making it difficult to meet the dual requirements of data security and real-time performance in the energy, power, and industrial sectors. On the other hand, current visual anomaly detection models still face significant limitations in generalization ability across different scenarios. Due to substantial differences in lighting, camera angles, equipment types, and background environments among different monitoring sites, models trained centrally often perform well on the source datasets but their accuracy drops significantly in new scenarios.
[0004] Therefore, existing technologies cannot simultaneously meet the comprehensive needs of privacy protection, cross-scenario generalization, and real-time adaptation. In particular, in multi-source video surveillance scenarios, how to achieve high-precision detection and low-latency response of the model remains a key technical challenge that the industry urgently needs to solve. Summary of the Invention
[0005] This invention provides a video anomaly detection method, apparatus, device, and medium based on federated vision Transformer, which solves the problems of weak privacy protection and insufficient cross-scene generalization ability in existing video anomaly detection technologies, and improves the accuracy, real-time performance, and stability of video anomaly detection.
[0006] According to one aspect of the present invention, a video anomaly detection method based on federated vision Transformer is provided, comprising:
[0007] The target video data is obtained by acquiring multi-source video data collected by multiple monitoring devices through edge nodes and preprocessing the multi-source video data.
[0008] The target video data is segmented and encoded to obtain a target video frame sequence;
[0009] The target video frame sequence is input into the target anomaly detection model, and the corresponding anomaly response map is output; wherein, the target anomaly detection model is trained based on a federated learning mechanism and an initial anomaly detection model; the target anomaly detection model is a multi-scale visual Transformer model;
[0010] An anomaly assessment is performed based on the anomaly response graph to obtain an anomaly assessment result. The corresponding alarm level is determined based on the anomaly assessment result, and an alarm operation is performed based on the alarm level.
[0011] According to another aspect of the present invention, a video anomaly detection device based on a federated vision Transformer is provided, comprising:
[0012] The data acquisition module is used to acquire multi-source video data collected by multiple monitoring devices through edge nodes, and to preprocess the multi-source video data to obtain target video data.
[0013] The data processing module is used to perform block processing and encoding processing on the target video data to obtain a target video frame sequence;
[0014] An anomaly detection module is used to input the target video frame sequence into a target anomaly detection model and output the corresponding anomaly response map; wherein, the target anomaly detection model is trained based on a federated learning mechanism and an initial anomaly detection model; the target anomaly detection model is a multi-scale visual Transformer model;
[0015] The anomaly assessment module is used to perform anomaly assessment based on the anomaly response graph to obtain anomaly assessment results, determine the corresponding alarm level based on the anomaly assessment results, and perform alarm operations based on the alarm level.
[0016] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0017] At least one processor; and
[0018] A memory communicatively connected to the at least one processor; wherein,
[0019] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the video anomaly detection method based on Federated Vision Transformer as described in any embodiment of the present invention.
[0020] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the video anomaly detection method based on federated vision Transformer as described in any embodiment of the present invention.
[0021] The technical solution of this invention acquires multi-source video data collected by multiple monitoring devices through edge nodes, preprocesses the multi-source video data to obtain target video data, performs block processing and encoding on the target video data to obtain a target video frame sequence, inputs the target video frame sequence into a target anomaly detection model to output a corresponding anomaly response map, performs anomaly evaluation based on the anomaly response map to obtain anomaly evaluation results, determines the corresponding alarm level based on the anomaly evaluation results, and performs alarm operations based on the alarm level. This technical solution solves the problems of weak privacy protection and insufficient cross-scene generalization ability in existing video anomaly detection technologies, and improves the accuracy, real-time performance, and stability of video anomaly detection.
[0022] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a flowchart of a video anomaly detection method based on Federated Vision Transformer according to Embodiment 1 of the present invention;
[0025] Figure 2 This is a flowchart of a video anomaly detection method based on Federated Vision Transformer according to Embodiment 2 of the present invention;
[0026] Figure 3 This is a schematic diagram of a video anomaly detection device based on Federated Vision Transformer according to Embodiment 3 of the present invention;
[0027] Figure 4 This is a schematic diagram of the structure of an electronic device provided according to Embodiment 4 of the present invention. Detailed Implementation
[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0029] It should be noted that the terms "first," "second," "initial," and "target," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0030] Example 1
[0031] Figure 1 This is a flowchart of a video anomaly detection method based on Federated Vision Transformer according to Embodiment 1 of the present invention. This embodiment is applicable to detecting abnormal behavior in complex monitoring environments. The method can be executed by a video anomaly detection device based on Federated Vision Transformer, which can be implemented in hardware and / or software and can be configured in an electronic device with data processing capabilities. Figure 1 As shown, the method includes:
[0032] S110. Obtain multi-source video data collected by multiple monitoring devices through edge nodes, and preprocess the multi-source video data to obtain target video data.
[0033] This embodiment can be executed by a video anomaly detection system based on a federated vision Transformer. This system can include an edge side and a federated side. Through federated collaborative decision-making and model evolution methods at both the edge and federated sides, it achieves anomaly identification processing of multi-source video data from multiple monitoring stations or camera terminals. This addresses problems in existing video surveillance systems across various scenarios, such as insufficient privacy protection, weak model generalization, inconsistent features across devices, and poor real-time edge performance. The method consists of multiple edge nodes and a federated coordination server. Each edge node performs feature extraction and model training locally, only uploading model parameter updates to the server. A globally optimized model is formed through a federated aggregation mechanism and periodically updated, thereby achieving a distributed, adaptive, and scalable anomaly detection system.
[0034] In this embodiment, an edge node can refer to a terminal or edge computing device deployed near the monitoring equipment, possessing local data processing and model inference capabilities. Each edge node in this embodiment can correspond to one or more monitoring devices. A monitoring device can refer to a device deployed at the monitoring site for acquiring video streams from the target area. Multi-source video data can be multiple video streams from multiple cameras within the same area corresponding to the edge node. Preprocessing can refer to operations such as illumination and color normalization, noise suppression, geometric image stabilization correction, geometric offset modeling, and ROI region cropping. Target video data can be the individual video frames obtained after preprocessing the multi-source video data. In this embodiment, the monitoring equipment specifically includes, but is not limited to, high-definition network cameras, infrared thermal imaging cameras, and explosion-proof cameras commonly used in industrial scenarios. It can also select appropriate equipment types according to the monitoring scenario requirements (e.g., power distribution rooms and charging stations) to meet imaging requirements in different environments. In this embodiment, the monitoring equipment can be directly deployed in the target monitoring area, such as electrical cabinet doors, charging gun ports, and power distribution panels, for real-time acquisition of on-site video data. It is understandable that multiple monitoring devices in the same monitoring scenario in this embodiment are usually under the jurisdiction of the same edge node, and the multiple video streams they collect together constitute multi-source video data.
[0035] In this embodiment, multiple video streams from multiple cameras within the same area corresponding to the edge node can be obtained through the edge node. Preprocessing operations such as illumination and color normalization, noise suppression, geometric image stabilization correction, geometric offset modeling, and ROI region cropping are performed on the multi-source video data to obtain the preprocessed video frame data.
[0036] Specifically, in this embodiment, based on the video stream sampling of traditional monitoring equipment, a Multi-Source Scene Stability Index (MSSI) is introduced, which can be determined by illumination entropy. Motion intensity entropy Image sharpness It consists of multiple dimensions of indicators, used to measure the dynamism and usability of the current monitoring screen, specifically:
[0037] ;
[0038] In this embodiment, the system can collect frame sequences at a unified clock on the edge node side and perform adaptive frame rate adjustment based on the MSSI value. Specifically, when high dynamic changes occur in the monitored area, such as vehicles, electric arcs, or abnormal intrusions by personnel, the sampling frequency is increased to capture key anomalies; while when the scene is stable (e.g., low motion, low texture changes), the sampling frequency is reduced to decrease unnecessary data volume and improve energy efficiency. Among these, inter-frame motion energy... The definition is as follows:
[0039] ;
[0040] in, and Representing the t-th frame and the t-th frame respectively The original input image of the frame, where H and W are the image height and width, respectively. This represents the L1 norm. The system sets a threshold. ,when At that time, shorten the sampling interval This mechanism increases the frame rate and conversely extends the sampling interval, enabling high-frequency acquisition in dynamic scenes and energy-efficient sampling in static scenes. This ensures the monitoring system has agile response capabilities in highly variable scenarios, while reducing energy consumption and storage pressure during stable periods.
[0041] Furthermore, the system introduces a time synchronization correction module among multiple camera nodes, using edge NTP synchronization signals to align frame timestamps, making data between different nodes comparable and providing temporal consistency for subsequent federated feature alignment.
[0042] In this embodiment, to address the data distribution changes caused by multi-source monitoring equipment such as energy sites, charging stations, and power distribution rooms under cross-time periods and cross-view conditions, a scene stability-driven preprocessing mechanism for federated collaborative modeling is proposed. This mechanism not only completes video acquisition and traditional image processing, but more importantly, it generates quality indicator signals for subsequent federated adaptive aggregation (ACA) and cross-node feature alignment (CFA), thus forming the input basis for end-to-end collaboration in federated learning.
[0043] In this embodiment, optionally, preprocessing of multi-source video data to obtain target video data includes: normalizing, denoising, and stabilizing the multi-source video data to obtain first image data; wherein the first image is geometrically consistent image data aligned by a homography matrix; generating a target region mask using a segmentation algorithm, and performing region blurring processing on the first image data based on the target region mask to obtain target video data.
[0044] Normalization processing can refer to illumination and color normalization. Denoising processing refers to the process of removing noise from video data. Image stabilization processing refers to geometric image stabilization correction processing of video data, used to eliminate cross-frame geometric shifts caused by camera shake and installation angle deviations. The first image data can be geometrically consistent image data obtained after normalization, denoising, and image stabilization processing. The homography matrix can be a 3×3 homogeneous matrix representing the projection transformation relationship between two images. In this embodiment, the cross-frame geometric mapping matrix can be obtained based on the SIFT / ORB feature matching homography matrix estimation algorithm. The segmentation algorithm can be any algorithm for region segmentation. In this embodiment, the segmentation algorithm can be a weakly supervised segmentation algorithm. The target region mask can refer to a specific region mask or a region of interest mask. In this embodiment, the target region can refer to key areas such as electrical cabinet doors, charging gun ports, and power distribution panels, and specific regions can be set according to actual needs. In this embodiment, performing region blurring processing on the first image data based on the target region mask can be achieved by using the target region mask to perform privacy blurring processing on irrelevant regions in the first image data.
[0045] In this embodiment, since the imaging characteristics and illumination variations of different monitoring devices are significant, multi-stage illumination normalization can be performed locally. Therefore, the specific method for performing illumination and color normalization processing on multi-source video data in this embodiment can be to first use the Gray-World white balance algorithm to obtain the pixel value matrix of the input image in the R / G / B channels. Value after balancing color temperature Specifically:
[0046] ;
[0047] in, , For the first The ratio of channel mean to global grayscale mean is used to balance color temperature; This indicates the original input image in color channels. The pixel matrix on; Channel image after white balance.
[0048] Then gamma correction can be performed, specifically as follows:
[0049] ;
[0050] in, It is an adaptive coefficient (usually taken as 0.9–1.1), which can be dynamically adjusted through luminance entropy to balance exposure; It involves recombining the correction results of each channel to form a three-channel image with complete white balance processing; This is the color normalization result after gamma correction. Subsequently, the CLAHE algorithm is used for local contrast enhancement, synchronously correcting low-light and overexposed areas in strong light at night. In this embodiment, by normalizing multi-source video data, the images output by different monitoring devices can be made more uniform in brightness distribution and color space, eliminating domain shifts caused by lighting differences and providing lighting robustness for multi-scene anomaly detection.
[0051] In this embodiment, since the video data stream in the monitoring equipment is often affected by compression artifacts, noise interference, and slight jitter, a combined denoising and image stabilization mechanism can be used to improve inter-frame geometric stability. The degradation model is defined as follows:
[0052] ;
[0053] in, For an ideal, clear image, For point-diffusion nuclei, It is Gaussian noise. This is the degraded observation image. In this embodiment, Wiener filtering can be used to recover high-frequency information, and bilateral filtering can be combined to suppress texture noise.
[0054] In this embodiment, a homography matrix estimation algorithm based on SIFT / ORB feature matching can be used to obtain the cross-frame geometric mapping matrix. It is used to describe the perspective transformation relationship between adjacent image frames. To eliminate camera angle deviation, in this embodiment, a stable image matrix can be further constructed based on the homography matrix, specifically:
[0055] ;
[0056] in, , These are the original and aligned pixel coordinates, respectively. The homography matrix is obtained through SIFT matching estimation. In this embodiment, after noise suppression and geometric image stabilization correction are performed, the homography matrix is obtained. Aligned geometrically consistent image, denoted as That is, the first image data is obtained.
[0057] In this embodiment, after completing cross-frame and cross-viewpoint geometric image stabilization correction, a cross-camera geometric offset modeling mechanism can be introduced at this stage to further characterize the structural differences of different monitoring nodes in terms of camera installation position, viewing angle, and imaging geometry. Furthermore, in this embodiment, the system is based on the estimated homography matrix. By combining the reprojection error of key points with the distribution of geometric transformation parameters, the degree of offset between each monitoring node and the global reference geometry is quantitatively evaluated, and the nodes are defined. The geometric offset index is:
[0058] ;
[0059] in, The geometric difference metric function can be calculated based on homography matrix parameter distance, reprojection residual, or keypoint alignment error. It is a global geometric reference structure formed by statistics of multiple nodes. The smaller the value, the higher the consistency between the node and the global spatial structure. In this embodiment, the geometric offset index does not contain any original image content or reversible visual features; it only reflects the degree of offset between the camera viewpoint and the spatial structure level, serving as an important input signal for geometric consistency constraints in the subsequent federated adaptive aggregation and cross-node feature alignment processes.
[0060] Specifically, in this embodiment, a weakly supervised segmentation algorithm is used to generate region of interest masks. Focusing on key areas such as electrical cabinet doors, charging gun ports, and power distribution panels, while applying privacy blurring to irrelevant areas, specifically:
[0061] ;
[0062] in, For Gaussian blur or mosaic function, Represents pixel-by-pixel product. The image after image stabilization and geometric alignment. This represents a semantic interest region mask. In this embodiment, this processing operation allows the model to focus its learning on regions strongly correlated with anomalies, while simultaneously achieving de-identification at the source, ensuring privacy compliance during data upload and federated training.
[0063] S120. The target video data is segmented and encoded to obtain the target video frame sequence.
[0064] Block processing can be the operation of dividing image data into non-overlapping blocks of fixed size. Encoding processing can be the operation of encoding positional information and geometric consistency description vectors. The target video frame sequence can refer to the final video frame sequence obtained after block processing and encoding of each video frame data.
[0065] In this embodiment, the image can be divided into fixed-size non-overlapping blocks (P×P), the input dimension is standardized, and then position information and geometric consistency description vector encoding operations are performed to obtain the target video frame block sequence for model input.
[0066] S130. Input the target video frame sequence into the target anomaly detection model and output the corresponding anomaly response map.
[0067] The target anomaly detection model is trained based on a federated learning mechanism and an initial anomaly detection model. The target anomaly detection model can be a pre-trained model used to identify abnormal behaviors in a video frame sequence. In this embodiment, the target anomaly detection model can be deployed on edge nodes. The initial anomaly detection model can be an untrained initial model. In this embodiment, the target anomaly detection model can be obtained by federated collaborative training and optimization of the initial anomaly detection model by edge nodes and a federated coordinator based on a federated learning mechanism. The anomaly response map can be a local anomaly map obtained by the target anomaly detection model performing anomaly detection on the target video frame sequence. In this embodiment, there can be one or more anomaly response maps.
[0068] In this embodiment, the target video frame sequence can be input into the target anomaly detection model deployed at each edge node for anomaly detection, thereby outputting one or more local anomaly response maps corresponding to the target video frame sequence. Specifically, in this embodiment, the target anomaly detection model can be a multi-scale visual Transformer model. In this embodiment, the target video frame sequence is input into the multi-scale visual Transformer model. The model learns the correlation between local and global features through a self-attention mechanism, and then introduces a hierarchical downsampling structure to capture fine-grained anomaly features at the lower level and aggregate macroscopic scene information at the higher level. The output feature tensor is input into the embedded anomaly prediction head to generate the corresponding local anomaly response map.
[0069] S140. Based on the abnormal response diagram, perform anomaly assessment to obtain the anomaly assessment result, determine the corresponding alarm level according to the anomaly assessment result, and perform alarm operation according to the alarm level.
[0070] Here, anomaly assessment refers to the processing operation of structured evaluation and stability determination of the anomaly response map output by the model. Anomaly detection result can be an anomaly score signal obtained from the structured evaluation and stability determination of the anomaly response map. Alarm level can refer to the alarm level determined by the anomaly determination result and the corresponding alarm mechanism. In this embodiment, the alarm level can be divided into three levels according to the severity of the anomaly, specifically including level one alarm, level two alarm, and level three alarm.
[0071] In this embodiment, the abnormal response graphs generated by each node are subjected to temporal smoothing and spatial consistency constraints. Then, a final alarm threshold is determined by combining empirical thresholds and adaptive learning thresholds. Dynamic judgment is performed based on the abnormality scoring results. If the abnormality scoring results and the final alarm threshold meet the corresponding judgment conditions, the corresponding alarm level is determined, thereby automatically triggering an early warning and outputting a visual alarm signal, thus performing the corresponding alarm operation. It is understood that in this embodiment, different alarm levels may correspond to different alarm operations, which can be set according to actual needs.
[0072] The technical solution of this invention acquires multi-source video data collected by multiple monitoring devices through edge nodes, preprocesses the multi-source video data to obtain target video data, performs block processing and encoding on the target video data to obtain a target video frame sequence, inputs the target video frame sequence into a target anomaly detection model to output a corresponding anomaly response map, performs anomaly evaluation based on the anomaly response map to obtain anomaly evaluation results, determines the corresponding alarm level based on the anomaly evaluation results, and performs alarm operations according to the alarm level. This technical solution solves the problems of weak privacy protection and insufficient cross-scene generalization ability in existing video anomaly detection technologies, and improves the accuracy, real-time performance, and stability of video anomaly detection.
[0073] Example 2
[0074] Figure 2 This is a flowchart of a video anomaly detection method based on Federated Vision Transformer according to Embodiment 2 of the present invention. This embodiment is an optimization based on the above embodiment. Specifically, the optimization involves: performing block processing and encoding processing on the target video data to obtain a target video frame sequence, including: dividing the target video data according to a set size and performing encoding operations to obtain an initial video frame sequence; determining a geometric description vector; and fusing the geometric description vector and the initial video frame sequence to obtain the target video frame sequence. Figure 2 As shown, the method includes:
[0075] S210. Obtain multi-source video data collected by multiple monitoring devices through edge nodes, and preprocess the multi-source video data to obtain target video data.
[0076] S220. Divide and encode the target video data according to the set size to obtain the initial video frame sequence.
[0077] The set size can be a pre-defined size. For example, the set size can be 16×16 or 32×32, etc. The set size in this embodiment can be set according to actual needs. The encoding operation can be to transform the divided image patches into feature vectors with spatial location information. In this embodiment, the encoding operation can include linear projection encoding and position encoding. The linear projection operation can be to perform a linear transformation on each image patch of the set size through a fully connected layer, mapping the high-dimensional pixel vector to a low-dimensional feature vector of fixed dimension, completing the dimensionality reduction and mapping from pixel space to feature space. The position encoding operation can be to embed a position encoding vector into each divided patch to represent the spatial location information of each patch in the original image. The initial video frame sequence can be a structured feature vector sequence obtained by dividing each video frame data according to the set size and performing linear projection encoding and position encoding operations.
[0078] In this embodiment, after preprocessing the multi-source video data to obtain the target video data, the video frames in the target video data can be divided into non-overlapping segments. Patch and flatten into vectors Then, through linear projection and positional coding, an input embedding sequence is formed, which yields the initial video frame sequence, specifically:
[0079] ;
[0080] in, This represents the input sequence of the Transformer model after patch embedding. For learnable embedding matrices, This represents the patch flattening vector. Location encoding is used to preserve spatial location information. In this embodiment, this setting not only ensures compatibility with the model input interface of the multi-scale Transformer architecture, but also, through preprocessing, image stabilization, and normalization, enables patch embedding to maintain structural comparability across different devices, which helps subsequent hierarchical attention mechanisms accurately identify abnormal patterns.
[0081] Furthermore, to improve the adaptability of subsequent federated aggregation (ACA) in this embodiment, a privacy-free quality metric, including the average brightness, is generated at the edge. Motion intensity entropy (Calculated from optical flow histogram entropy) and sharpness index (Obtained based on Laplace energy), where, Indicates the first The frames, after illumination and color normalization, are used for brightness statistics and quality assessment. These metrics participate in the aggregated weighting as data quality meta-features:
[0082] ;
[0083] in, This represents a multidimensional weighting function that outputs a quality signal. This is for subsequent use by the adaptive aggregation module.
[0084] S230. Determine the geometric description vector, and fuse the geometric description vector with the initial video frame sequence to obtain the target video frame sequence.
[0085] The geometric description vector can be a structured vector representing the geometric features of the image, camera viewpoint relationships, and spatial consistency. In this embodiment, the geometric description vector can be obtained by encoding the homography matrix. This embodiment can utilize the Geometry-Constrained Attention (GCA) mechanism to encode the homography matrix. The encoding is used as geometric reference weights, enabling the model to automatically strengthen stable regions across viewpoints and weaken structural deformations caused by differences in shooting angle and equipment installation position during attention weighting. The fusion process can be either stitching or weighted fusion.
[0086] In this embodiment, after patch division and embedding of a set size, a basic patch embedding sequence with a uniform format can be obtained. This refers to the initial video frame sequence, which has already undergone uniformization in terms of spatial structure and color distribution. Then, an initial representation for multi-scale feature modeling can be constructed based on the basic embeddings, serving as the model input for subsequent hierarchical Transformer structures. In this embodiment, to ensure that the patch features have cross-camera spatial correspondence before entering multi-scale attention modeling, a geometric consistency enhancement mechanism is introduced, which uses the homography matrix estimated in the preprocessing operation... The geometric description vectors are converted into patch-level geometric consistency description vectors and then injected into the initial video frame sequence through splicing or weighted fusion to obtain the target video frame sequence. This ensures that each patch not only contains local visual information in the initial feature space, but also explicitly encodes its spatial stability and correspondence under different camera views.
[0087] S240. Input the target video frame sequence into the target anomaly detection model and output the corresponding anomaly response map.
[0088] The target anomaly detection model is a multi-scale visual Transformer model trained based on a federated learning mechanism and an initial anomaly detection model. In this embodiment, the target anomaly detection model adopts a hierarchical pyramid architecture, with each layer containing a multi-head attention module and a downsampling inference structure, allowing the representation to be compressed layer by layer in the spatial dimension and gradually enriched in the semantic dimension. Furthermore, this embodiment introduces cross-camera geometric consistency constraints on top of the original multi-scale feature extraction framework, enabling the Transformer model to refer to the obtained homography matrix when performing local-global attention. It automatically suppresses feature shifts caused by differences in viewpoint, thereby achieving cross-device uniformity of feature space.
[0089] The core of the multi-scale visual Transformer model in this embodiment lies in its hierarchical self-attention modeling mechanism. This mechanism uses a multi-head self-attention structure to characterize the global dependencies between patch features, enabling the model to adaptively focus on key regions and suppress redundant background information within a unified feature space. To further adapt to the spatial structure shift problem caused by viewpoint differences in multi-camera monitoring scenarios, this embodiment introduces geometric consistency constraints during the self-attention weight calculation process, constructing a Geometry-Constrained Attention (GCA) mechanism to explicitly utilize cross-viewpoint geometric correspondences to guide feature modeling. Specifically, based on the estimated homography matrix... The geometric consistency constraint term between patches is calculated and introduced as a bias into the calculation process of self-attention weights. Its attention form can be represented as follows:
[0090] ;
[0091] in, , , These represent the query, key, and value matrices, respectively. is the scale factor for the key vector. The geometric consistency constraint term, calculated from the homography matrix, is used to quantify the spatial correspondence stability of different patches across camera views. When patches maintain high geometric consistency across different cameras, this constraint term enhances the attentional association between corresponding features; when patches experience significant spatial shifts due to viewpoint differences, the constraint term automatically reduces their attention weight, thereby avoiding erroneous feature aggregation.
[0092] In this embodiment, through the aforementioned geometric consistency attention mechanism, the model can actively perceive and adapt to differences in cross-node geometric structures when performing spatial attention modeling and feature aggregation, enabling the learned feature representations to have stronger stability and alignment in multi-camera scenarios.
[0093] After completing the self-attention modeling at each layer, the model further performs scale compression on the feature maps through pyramid-style downsampling operations (including stride convolution or window pooling), generating multi-level feature representations with progressively decreasing resolution but enhanced semantics. Lower-level features focus on characterizing local fine-grained anomaly patterns, while higher-level features focus on overall structure and scene-level anomalies, thus achieving unified multi-scale modeling of complex monitoring anomalies.
[0094] Furthermore, to fully integrate the feature information extracted by the multi-scale Transformer at different levels, this embodiment introduces a geometric consistency-aware adaptive aggregation mechanism during cross-layer feature fusion, enabling the fusion result to simultaneously ensure semantic integrity and cross-camera spatial consistency. Specifically, after scale alignment of the output features of each layer, the model achieves unified integration of cross-layer features through a weighted approach, and its fusion form is expressed as follows:
[0095] ;
[0096] in, Indicates the first Feature maps output by the Transformer layer This is an upsampling operator used to map features from different levels to a uniform spatial resolution; For the first The fusion weights corresponding to the layer features are used to adjust the contribution of features at each scale to the final fusion result. The fusion weights in this embodiment... It is not set independently, but is adjusted by cross-camera geometric consistency information. Specifically, this weight can be based on the homography matrix. The geometric stability features described, and the geometric consistency measure formed during the attention modeling stage. Adaptive adjustments are made. When a feature at a certain level exhibits high spatial consistency under different camera views, its corresponding weight is enhanced; conversely, when the feature is significantly affected by changes in viewpoint and exhibits significant geometric shift, its fusion weight is reduced accordingly.
[0097] In this embodiment, by introducing a geometric consistency-aware cross-layer aggregation mechanism, the model can actively suppress the interference of geometrically unstable scales during the multi-scale feature integration stage, and prioritize the retention of feature representations with greater spatial alignability in multi-camera scenes. This allows the fused features to maintain multi-scale semantic integrity while significantly enhancing spatial consistency across devices and viewpoints.
[0098] In this embodiment, after completing cross-layer feature fusion based on geometric consistency awareness, the model obtains a fused feature representation F. Based on this, anomaly feature saliency and density-aware detection mechanisms are introduced to map the high-dimensional fused features into spatially interpretable anomaly response representations, achieving hierarchical anomaly detection from pixel-level perception to region-level judgment. Specifically, the fused feature F is first input to a lightweight anomaly detection head, through... Convolution linearly compresses channel-dimensional features and combines them with a non-linear activation function to generate pixel-level anomaly response maps, which can be expressed as follows:
[0099] ;
[0100] in, This represents a pixel-wise channel mapping operator used to integrate multi-channel feature information; The Sigmoid activation function normalizes the output to the [0,1] interval; This is an anomaly significance response map, and its numerical value reflects the relative confidence level of the anomaly occurring at the corresponding spatial location.
[0101] This embodiment introduces density-aware modeling into the anomaly response map. By aggregating and analyzing anomaly responses within a spatial neighborhood, the spatial distribution density of the anomaly responses is characterized. When anomaly responses exhibit a continuous, high-density distribution within a local area, the system determines that the area has a higher anomaly confidence level; conversely, scattered, isolated high-response points are suppressed, thereby effectively reducing the risk of false detections caused by noise or local disturbances. Through the aforementioned anomaly saliency and density-aware detection mechanism, this embodiment enables the model to achieve structured modeling of anomaly regions while maintaining pixel-level positioning accuracy. This elevates the detection results from single-point responses to regional and scene-level early warning signals, enhancing the stability and interpretability of anomaly detection results.
[0102] In this embodiment, optionally, the training steps of the target anomaly detection model include: obtaining a training sample set and an initial anomaly detection model for each edge node; wherein the training sample set is obtained by screening through a scene stability index; training the initial anomaly detection model based on each training sample contained in the training sample set to obtain the model parameters of each edge node in the current training round; wherein the model parameters include model performance indicators and stability indicators; constructing node state vectors corresponding to each edge node based on the model parameters through a federated coordinator, and determining the set of nodes with full participation, the set of nodes with non-full participation, and the set of observation nodes based on the node state vectors; updating the parameters based on the model parameters of the set of nodes with full participation and setting update operators to obtain global model parameters; and distributing the global model parameters to each edge node through the federated coordinator, so that each edge node continues to execute the training operation of the initial anomaly detection model in the next training round based on the global model parameters, until the target anomaly detection model is obtained.
[0103] The training sample set can be composed of video frame sequences obtained by processing historical video data collected by each edge node. In this embodiment, the training sample set of each edge node can be constructed by organizing local data with scene stability awareness and weakly supervised samples. Specifically, in this embodiment, on the edge node side, the local initial anomaly detection model is trained using video frame sequences collected by the node itself. To improve training efficiency under limited annotation conditions, the system organizes the local data in a structured manner, dividing the samples into two categories: manually annotated samples and weakly supervised samples. Manually annotated samples are determined by maintenance personnel or historical rules, and contain clear anomaly location or region annotations; weakly supervised samples originate from model prediction results or implicit anomaly prompts in device operation logs. To avoid noise labels in weakly supervised samples interfering with model training, this embodiment introduces a sample screening mechanism based on the Scene Stability Index (MSSI) to constrain the reliability of weakly supervised samples. Specifically, the system only selects samples with scene stability higher than a set threshold within the current time window. Weakly supervised samples are only allowed to be included in the local training set when certain conditions are met. When drastic changes occur in the monitoring scene (such as large-scale movement, sudden changes in lighting, or viewpoint disturbances) causing a decrease in MSSI, the corresponding weakly supervised samples will be downweighted or temporarily removed to avoid training bias caused by unreliable pseudo-labels. In this embodiment, an adaptive training mechanism driven by the Scene Stability Index (MSSI) is introduced during the local optimization process, enabling the training intensity, sample usage strategy, and update reliability to be adjusted according to the dynamic changes in the monitoring environment.
[0104] Furthermore, in the weakly supervised sample generation process, this embodiment optionally introduces a teacher-student dual-model consistency mechanism as an auxiliary discrimination method. The teacher model is smoothly updated using historical model parameters through an exponential moving average, providing a relatively stable prediction reference; the student model makes real-time predictions based on the current frame sequence. Only when the teacher and student outputs are consistent in spatial location and anomalous response intensity, and the corresponding sample satisfies scene stability constraints, is the sample considered a valid weakly supervised sample for training. By introducing a scene stability-aware data organization strategy, this step transforms the use of weakly supervised samples from "passive introduction" to "conditional triggering," effectively avoiding the problem of false label diffusion caused by scene disturbances in complex monitoring environments, and providing a more reliable and controllable sample foundation for subsequent local model training.
[0105] Among these, model performance metrics can be indicators used to reflect model performance. Stability metrics can be indicators used to characterize the stability of model updates.
[0106] In this embodiment, multi-scale visual Transformer models can be independently trained and adaptively updated locally at each edge node, enabling the model to continuously adapt to the scene characteristics and operating status of the monitoring site. Each edge node corresponds to one or more monitoring points. Internally, the node takes locally acquired video frame sequences as input, that is, each training sample in the training sample set is input into the initial anomaly detection model for training. The input to the initial anomaly detection model undergoes feature extraction and anomaly saliency processes, outputting the corresponding initial anomaly response map, obtaining the model parameters of each edge node in the current training round, and generating model parameter update quantities without uploading the original data. After each round of local training, the node not only generates model parameter update quantities, but also simultaneously calculates quality and reliability indicators reflecting model performance and update stability, used to characterize the effectiveness of the node's update results in the current scene. These indicators are passed to the federated end as additional attributes of parameter updates.
[0107] In this embodiment, during local model training, anomalous samples are typically significantly fewer than normal samples, and anomalous regions exhibit a highly unbalanced spatial distribution. If standard cross-entropy loss is directly applied for optimization, the model may tend to favor learning the background or normal patterns while ignoring key anomalous features. Therefore, this embodiment constructs a stability-modulated reweighted loss function during the local training phase. This function balances class balance and region focus while introducing adaptive adjustment of training intensity based on scene stability. Specifically, let the local training sample set be... The sample label is ,in Indicates an abnormal sample. The model represents the samples If the probability of predicting an anomaly is given, the basic loss function is defined as a weighted binary cross-entropy form:
[0108] ;
[0109] Wherein, the sample weight function Defined as:
[0110] ;
[0111] in, Indicates the category to which the sample belongs. The frequency of occurrence in the current node is used to alleviate the class imbalance problem; It is a numerically stable term; This is the indicator of the region of interest (ROI) corresponding to the sample, and its value is 1 when the sample is located in an anomaly candidate region; This represents a measure of sample difficulty, used to characterize prediction error or gradient magnitude. This is the weighting adjustment coefficient. (Function) This is a modulation factor based on the scene stability index, used to control the effectiveness of the reweighting mechanism under different scene states.
[0112] In this embodiment, when the monitored scene is in a stable state and the MSSI is high... Taking a larger value allows the weighting strategies such as class frequency, ROI focus, and hard example enhancement to take full effect, thereby strengthening the model's learning of key anomaly regions; when the scene fluctuates drastically and the MSSI is low, Automatic shrinkage brings the loss function closer to its smooth base form, preventing noisy samples or unreliable pseudo-labels from being overemphasized during training. By introducing a reweighted loss design with scene stability modulation, local model training is transformed from static weight optimization into an adaptive process dynamically coupled with the monitoring environment. This effectively improves the model's training stability and generalization ability in complex, non-stationary scenarios while ensuring class balance and anomaly sensitivity.
[0113] In this embodiment, the local training sample set defined above is used. Sample Labels and prediction probability Building upon this foundation, to further enhance the model's ability to identify small, ambiguous, and boundary anomalies, a hard example focusing mechanism with stability constraints is introduced. A temporal consistency regularization term is constructed based on video sequence characteristics to dynamically regulate the training process. Firstly, regarding hard example focusing, the system employs a focus loss form based on prediction confidence to suppress the gradient contribution of easily classified samples, thereby highlighting the influence of hard-classified samples in parameter updates. The category-related prediction confidence is defined as follows: when... hour, ,otherwise The difficult example focusing on the loss item is represented as:
[0114] ;
[0115] in, This is a focus adjustment parameter used to control the magnification level of difficult examples. This is the sample weight function, used to maintain the consistency of the loss design across different training stages.
[0116] Secondly, in order to avoid drastic prediction fluctuations in the model due to short-term noise, instantaneous occlusion or scene disturbances in the video stream, a temporal consistency regularization constraint based on adjacent time frames is introduced during the training process to limit the variation of abnormal response results in consecutive frames, so as to enhance the consistency and reliability of the detection results in the time dimension.
[0117] In this embodiment, the Scene Stability Index (MSSI) is used as a modulation factor to adaptively control the intensity of hard example focusing and temporal consistency constraints. When the monitored scene changes drastically and the MSSI is low, the system strengthens the temporal consistency constraints and relatively suppresses the intensity of hard example focusing to avoid excessive amplification of noisy samples or unstable pseudo-labels during training. When the scene is stable and the MSSI is high, the temporal constraints are automatically relaxed, and the hard example focusing mechanism takes full effect, enabling the model to converge faster and accurately characterize the distribution of abnormal features. By introducing hard example focusing with stability constraints and temporal consistency regularization, the local model training process can dynamically switch between "stability priority" and "efficiency priority" under different monitoring scene conditions, improving the sensitivity of anomaly detection while effectively suppressing prediction jitter and false alarms.
[0118] This embodiment integrates the aforementioned multiple loss terms and introduces a stability and quality-aware mechanism to construct a stable and controllable local optimization and update process. Specifically, the comprehensive optimization objective function during the local training phase is defined as:
[0119] ;
[0120] in, For stability modulation, reweighted cross-entropy loss, To focus on the loss items in the above difficult examples, For weakly supervised consistency and time-series regularization, Represents the response diagram to anomalies The total variation regularization is used to constrain the spatial smoothness of outlier region boundaries. For parameter regularization terms, These are the weighting coefficients for each loss term.
[0121] During parameter optimization, the local node uses the AdamW optimization algorithm to iteratively update the objective function, and combines this with a gradient pruning strategy to suppress abnormal gradient fluctuations, ensuring stable convergence under conditions of limited edge computing resources. It is important to emphasize that the amount of model updates generated by the node in this embodiment... It is not a "naked update" obtained directly from gradient accumulation, but a structured update result that integrates training gradients, scene stability, and data quality information.
[0122] Specifically, after completing one local training round, the system modulates the parameter update magnitude and effectiveness based on the current scene stability index (MSSI) and the data quality signal extracted in the preceding steps, thereby obtaining the final local update amount. When the monitoring scene is stable and the training sample quality is high, the parameter updates are fully preserved; when the scene fluctuates greatly or the sample reliability decreases, the update amplitude is automatically reduced to avoid unstable gradients causing excessive perturbation to the model.
[0123] In this embodiment, after completing the stability-aware optimization update of the local model, to characterize the effectiveness and reliability of the model update results of each edge node, the quality of the local training results is further evaluated, and an update stability index for federated collaboration is calculated. The generated evaluation results are not only used for internal node monitoring, but also serve as an additional attribute for parameter updates, providing the federated end with weighted decision-making in the subsequent adaptive aggregation process.
[0124] Specifically, after each round of local training, the node evaluates the performance of the current model on its local validation set and calculates a comprehensive quality index. , used to reflect the The local discriminative ability of a node model in anomaly detection tasks. This metric can be expressed as a weighted combination of various evaluation metrics:
[0125] ;
[0126] in, , and Let represent the area under the curve, mean precision, and F1 score of the model on the local validation set at each node, respectively. , which is a weighting coefficient used to balance the relative importance of different evaluation indicators in the comprehensive quality assessment.
[0127] This embodiment also introduces an update stability metric to characterize the reliability of the local parameter update process. Specifically, the node calculates the update stability by statistically analyzing the variance of the model parameter update amount within a continuous training window. Its formal definition is:
[0128] ;
[0129] in, This represents the parameter update increment of a node in adjacent local training rounds. For windowed variance calculation operators, This is a numerical stability term. This metric measures the stability of model updates over time, preventing highly volatile and low-reliability updates from having an excessive impact on federated aggregation.
[0130] In this embodiment, the aforementioned local model quality assessment and update stability calculation mechanism generates a structured update description for each edge node, which includes information on model update amount, performance quality, and stability. This enables the federation end to distinguish between "high-quality stable updates" and "low-reliability perturbation updates" during the aggregation process, thereby achieving more robust and adaptive cross-node collaborative optimization.
[0131] In this embodiment, to submit model update results to the federated end while ensuring data privacy and security, the local parameter update amount is encapsulated with privacy protection. This embodiment introduces a scene stability-aware modulation strategy during the update encapsulation stage, enabling the privacy protection strength to be dynamically adjusted according to the monitoring environment and the reliability of the model update.
[0132] Specifically, before uploading the local model update, the node first updates the local parameters. Normalization and pruning are performed to limit the impact of a single update on the global model. The privacy-preserving update form is expressed as follows:
[0133] ;
[0134] in, Indicated by threshold The parameter update vector is pruned to control the upper bound of the update norm; The term is a zero-mean Gaussian noise term. It is the identity matrix. For the first The noise intensity parameters corresponding to each node.
[0135] To adapt to changes in computing power, network, and scene conditions faced by edge nodes under different operating states, this embodiment introduces an adaptive training scheduling mechanism during local training to dynamically control the number of training rounds, update frequency, and optimization intensity. Unlike scheduling methods based solely on resource load, this embodiment uses the Scene Stability Index (MSSI) and local update stability index as core decision-making criteria: when the monitored scene is stable and model updates are reliable, nodes are allowed to execute relatively complete training rounds to fully learn the current scene features; when scene fluctuations are significant or update stability decreases, the system automatically reduces the training intensity, suppressing the impact of unreliable training on the model by reducing the number of training rounds or shrinking the update amplitude.
[0136] The federated coordinator can be considered the central hub for managing edge nodes. It does not store any raw video data or intermediate feature data from edge nodes; instead, it receives and processes lightweight information such as model parameters and status metrics uploaded by edge nodes and issues global model update commands. The node state vector is a structured feature vector constructed by the federated coordinator for each edge node, representing its collaborative value and reliability in the current training epoch. The set of fully participating nodes refers to edge nodes with high scene stability, reliable model updates, and strong compatibility with the global model structure. The set of non-fully participating nodes refers to edge nodes whose state performance falls between that of fully participating and observational nodes; it can also be called the set of restricted participating nodes. The set of observational nodes refers to edge nodes with poor scene state, unreliable updates, or extremely low compatibility with the global model.
[0137] In this embodiment, the federation coordinator, without sharing the original monitoring data, uniformly coordinates and globally models the local model updates generated by multiple edge nodes to construct an anomaly detection model with cross-scenario adaptability. This embodiment introduces a collaborative decision-making mechanism based on node scenario status and update reliability at the federation end, upgrading the federation process from a single parameter aggregation operation to an active scheduling of node participation methods and knowledge absorption levels. By dynamically evaluating and selectively absorbing the collaborative value of different nodes in the current scenario, the global model is continuously updated in a controlled evolutionary manner, effectively addressing the scenario heterogeneity and non-independent identical distribution problems commonly found in multi-source monitoring environments.
[0138] In this embodiment, before the federated collaboration begins, the federated end performs a unified modeling of the operational status of each participating edge node to characterize its scenario conditions, model update reliability, and structural consistency level in the current training round. Specifically, for the first... For each edge node, a node state vector can be constructed, which can be:
[0139] ;
[0140] in, This represents the scene stability and input quality signal formed by the node during the data acquisition and preprocessing stage, and is used to reflect the learnability of the current monitoring environment of the node; This represents the parameter update stability index calculated by the node during local model training and update, used to measure the reliability of this round of updates; It represents a geometric or distributional consistency measure between the features of the node model and the global model structure space, and is used to characterize the compatibility between node knowledge and the global model.
[0141] Unlike existing federated learning methods that directly aggregate based on parameters or gradients, this embodiment first maps multi-source heterogeneous node information to a unified state space at the federation end. The collaborative value of each node in the current round is then structurally described using node state vectors. The node state modeling process does not involve the transmission of any raw monitoring data or intermediate features; it only transmits high-level state variables. This ensures data privacy while providing a scene-aware input foundation for subsequent federated collaborative decision-making and model evolution.
[0142] In this embodiment, the state vectors of each edge node are obtained. Subsequently, the federated end executes a federated collaborative decision-making strategy based on the node state awareness results, dynamically determining the participation mode of each node in this round of collaborative training. In this embodiment, nodes are no longer treated as homogeneous individuals, but rather their collaborative roles are differentiated and scheduled based on their scene stability, update reliability, and structural consistency performance in the current round. Specifically, the federated end dynamically divides the set of participating nodes based on the comprehensive performance of each component in the node state vector, forming a set of all participating nodes. Limited Participation Node Set and the set of observation nodes In this system, nodes with high scene stability, stable updates, and good compatibility with the global model structure are prioritized for inclusion in the full set of participating nodes, serving as the primary source of knowledge for the global model evolution. Nodes experiencing drastic scene changes or unstable updates have their participation intensity actively reduced, or participate only as observation nodes in the current round of the federated process. This embodiment introduces a node-state-based collaborative decision-making and participation role allocation mechanism, enabling the federated system to adaptively adjust its collaborative structure in each training round. This allows the global model update to absorb more effective knowledge from reliable nodes while suppressing interference from unstable nodes on model evolution. This achieves more robust and controllable federated collaborative training in a multi-source, heterogeneous, and non-independently distributed monitoring environment.
[0143] The global model parameters can be all learnable parameters of the corresponding target anomaly detection model. The set update operator can be a pre-defined reliable update absorption operator, used to constrain consistency and compatibility before absorbing node updates. In this embodiment, the global model parameters can be obtained by updating the model parameters of the entire set of participating nodes and the set update operator.
[0144] In this embodiment, after completing the collaborative decision-making and participation role allocation based on node status, the federation end only performs model absorption operations on nodes identified as trusted contribution sources. Let the set of nodes selected as full participants in the current round be . , No. The local model update uploaded by each node is represented as: The global model parameters are represented as In this embodiment, the global model update process is modeled as a controlled model evolution process, and its update form can be represented as follows:
[0145] ;
[0146] Among them, symbols This represents the controlled evolution operation of the global model. This represents a trusted update absorber operator, used to constrain consistency and compatibility of the absorber node before it is updated.
[0147] Specifically, the trusted update absorption operator During model updates, the consistency between node update directions and global model evolution directions, as well as the compatibility between node model features and the global structure space, are comprehensively considered. Only update components that meet the credibility conditions are absorbed. For update components with significant deviations or unstable characteristics, their absorption intensity is actively weakened or delayed to avoid adverse effects of local anomalies or noisy updates on the global model.
[0148] In this embodiment, the update of the global model is no longer equivalent to the simple accumulation of node update amounts, but evolves into an evolutionary process with the gradual injection of credible knowledge as the core. This enables the federated model to maintain stable convergence and continuous optimization in multi-source heterogeneous and non-independent and identically distributed monitoring scenarios, thereby significantly improving the robustness and generalization ability in cross-node anomaly detection tasks.
[0149] Furthermore, in this embodiment, before the node update enters the federated collaboration process, the system performs privacy-preserving encapsulation on the model updates uploaded by each edge node to prevent the inference of local monitoring data or scene information through model parameters. In this embodiment, the local model update generated by the i-th node in the current round is set as follows: Its corresponding node state vector is This embodiment combines privacy constraints with node state awareness results, enabling the strength of privacy protection to be dynamically adjusted according to the node's operating state. Specifically, the system adaptively controls the perturbation intensity of the model update based on the scene stability and update reliability reflected in the node state vector, forming a state-aware update encapsulation result. It is represented as:
[0150] ;
[0151] in, This represents a privacy constraint and encapsulation operator, used to prune and perturb node updates while meeting privacy protection requirements. This is particularly relevant when nodes are in a stable, reliable, and structurally consistent state. Apply low-intensity perturbations to updates to preserve effective knowledge to the maximum extent; when node states are unstable or consistency is poor, increase the perturbations or limit the range of information that can be absorbed in updates to reduce potential privacy risks and their interference with the evolution of the global model.
[0152] In this embodiment, the privacy protection mechanism is embedded in the node state awareness and federated collaborative decision-making process, realizing the coordinated unity of privacy constraints and model evolution control. This makes privacy protection no longer an independent security add-on module, but an important constraint for the federated system to schedule and absorb node update credibility, thereby improving the stability and effectiveness of global model updates while ensuring data security.
[0153] In this embodiment, after completing the trusted update absorption and global model evolution, the updated global model parameters are transmitted through the federation coordinator. The updated global model is distributed to each edge node as initialization parameters for the next round of local model training and inference. Each node continues its local anomaly detection and training process based on the updated global model, and regenerates its node state vector in subsequent rounds. This allows the node's operating status to be continuously linked with the model's evolution results. By iterating this operation repeatedly, a target anomaly detection model that meets the requirements is obtained.
[0154] In this embodiment, the federated end comprehensively evaluates the performance trend of the global model across multiple nodes and the distribution of node states. Based on the evaluation results, it dynamically adjusts the federated collaborative decision-making strategy, including key parameters such as node participation role division rules, trusted update absorption strength, and privacy constraint encapsulation strategy. By explicitly introducing global model feedback into the collaborative decision-making process, federated training no longer relies on static strategies but can continuously and adaptively adjust with long-term changes in monitoring scenarios and node states. Furthermore, while ensuring data privacy and system stability, the global model achieves continuous evolution and robust optimization in multi-source heterogeneous monitoring environments, significantly improving the adaptability and long-term generalization performance of the federated anomaly detection model in complex real-world scenarios.
[0155] Furthermore, this embodiment addresses the feature drift problem in multi-camera monitoring systems caused by differences in viewing angles, changes in device installation locations, and inconsistent scene structures. It proposes a contrastive feature alignment mechanism that introduces cross-camera geometric consistency constraints to unify the feature representation space of different edge nodes without sharing the original monitoring data. This embodiment explicitly introduces geometric offset information obtained from the aforementioned geometric modeling and consistency analysis during the feature alignment process. This allows the contrastive learning process to be jointly constrained by semantic and geometric consistency, thereby effectively suppressing feature drift in complex multi-view, multi-device environments and significantly improving the cross-scene generalization ability and stability of the anomaly detection model.
[0156] Specifically, in this embodiment, the high-dimensional feature representation extracted by the aforementioned Multi-Scale Visual Transformer (MSVT) module in each edge node is denoted as follows: ,in This represents the video frame or its corresponding patch sequence input acquired by the node. To construct a unified representation space suitable for cross-node contrastive learning, this step maps features to a low-dimensional contrastive embedding space and performs scale normalization.
[0157] Specifically, a lightweight nonlinear projection function is introduced. , original features Map to contrast space and perform Normalization yields the feature embedding representation on a unit sphere:
[0158] ;
[0159] in, This represents the embedding vector used for subsequent similarity calculation and comparison optimization. It typically consists of a multilayer perceptron and a normalization layer, used to compress feature dimensions and unify the numerical range of outputs from different nodes.
[0160] This embodiment introduces a geometry-aware constraint during the embedding generation stage: the projection function. The parameter updates and embedding distribution are modulated by the geometric consistency metric (GD) calculated in the preceding steps. This ensures that features with relatively stable geometric structures and clear cross-camera spatial correspondences maintain a tighter distribution in the embedding space, while reducing the alignment strength of features with significant geometric shifts. Through this design, the embedding space explicitly encodes cross-camera geometric consistency while maintaining semantic discriminativity, providing a geometrically aware foundation for subsequent comparison sample construction and loss weighting. This is achieved after completing multi-node feature embedding and obtaining normalized embedding vectors. Subsequently, positive and negative sample pairs are constructed for contrastive learning. In this embodiment, cross-camera geometric consistency information is explicitly introduced during the sample relationship modeling process, so that the contrast constraints simultaneously reflect semantic and spatial structural relationships.
[0161] Specifically, two samples from different nodes or different time frames can be embedded as... and Its corresponding semantic tag is The degree of geometric offset across cameras is denoted as .in, The estimation, based on homography or geometric consistency as described in the preceding steps, is used to characterize the spatial correspondence stability of samples under different camera perspectives. When samples satisfy semantic consistency (i.e....),... And the degree of geometric offset is lower than the set threshold. When a pair of samples is geometrically consistent, it is identified as a positive sample pair, used to enhance the ability to aggregate the same semantic pattern under cross-node conditions. This type of positive sample pair emphasizes semantic consistency under the premise of stable spatial structure, which helps to improve the reliability of feature alignment. For sample pairs with semantic inconsistency or significant geometric offset, they are constructed as negative sample pairs. Furthermore, this embodiment performs geometric consistency-aware differentiation on negative samples: when samples are semantically similar but have a large geometric offset, their negative sample weights are appropriately increased to enhance the model's ability to identify "pseudo-similarity caused by perspective changes" and avoid the model mislearning geometric differences as semantic differences.
[0162] In this embodiment, after constructing positive and negative samples for geometric consistency perception, cross-camera geometric offset information is further introduced into the contrast optimization objective to enhance the model's discrimination ability under cross-view conditions. This makes the contrast learning process no longer rely solely on feature similarity itself, but is simultaneously constrained by semantic relations and geometric consistency.
[0163] Specifically, using samples Let be the anchor point, and let its corresponding positive sample embedding be denoted as . Negative sample embedding is represented as The contrast loss function after introducing geometric offset modulation is defined as:
[0164] ;
[0165] in, A similarity function (such as cosine similarity) between embedded vectors. This is a temperature parameter used to adjust the smoothness of the similarity distribution. (Weight term) The geometric modulation weights for negative samples are determined by the degree of geometric offset between sample pairs. The decision is used to characterize the stability of spatial correspondences under different camera perspectives.
[0166] In this embodiment, when the geometric offset between the negative sample and the anchor sample is large, the corresponding By appropriately amplifying these geometrically inconsistent but visually similar challenging negative samples during optimization, the model focuses more on them during the optimization process. Conversely, for sample pairs with relatively stable geometric structures, the contrastive penalty is reduced to avoid mislearning geometric differences as semantic differences. Through this geometric offset modulation mechanism, the contrastive loss function guides the model to form a feature distribution in the embedding space that simultaneously satisfies semantic discriminability and cross-camera geometric consistency.
[0167] In this embodiment, after completing the comparative optimization of geometric offset modulation, the federated end further constructs cross-node shared feature prototypes for different semantic categories based on the low-dimensional feature statistics uploaded by each edge node. This prototype is used to characterize stable semantic center representations in a multi-camera environment. This process does not rely on original features or sample data, but is completed solely based on the embedding statistical information uploaded by the nodes, thereby meeting the privacy constraints under federated learning.
[0168] Specifically, for the first Class semantic target, let the feature embedding statistics of this class from the i-th node be . The corresponding geometric consistency metric is The federal side is building a global semantic prototype. At this time, a geometric consistency constraint is introduced to perform weighted aggregation of the contributions of different nodes:
[0169] ;
[0170] in, The weighting coefficient is related to geometric consistency. Its value is adaptively adjusted according to the consistency between node features and the global geometry, which is used to enhance the influence of geometrically stable nodes in prototype updates. Node features with higher geometric consistency account for a larger proportion in prototype construction, while node features with significant viewpoint shifts or unstable structures are weakened.
[0171] In this embodiment, the contrastive feature alignment mechanism serves as an auxiliary optimization mechanism during the training phase. It works in conjunction with the local anomaly detection loss and the federated collaborative update process to shape a unified feature representation space with cross-camera geometric consistency during the model learning phase. This mechanism does not directly participate in anomaly detection output but indirectly improves the model's generalization ability and stability in multi-node environments through structured constraints on the feature embedding distribution. During model training, the contrastive loss function, incorporating the aforementioned geometric consistency constraints, is used. The spatial structure of the multi-node feature embedding sequence z is optimized to create an alignable embedding distribution for the same semantic target under different camera views. This process, along with the local supervision loss, works on the model parameter update but does not change the model structure during the inference phase.
[0172] In this embodiment, the complex cross-node feature alignment process is restricted to the training phase, which ensures the efficiency of the federated anomaly detection system during the deployment phase and significantly enhances the robustness and transferability of the model in multi-camera environments.
[0173] S250. Based on the abnormal response diagram, perform anomaly assessment to obtain the anomaly assessment result, determine the corresponding alarm level according to the anomaly assessment result, and perform alarm operation according to the alarm level.
[0174] In this embodiment, optionally, an anomaly assessment result is obtained by performing anomaly assessment based on the anomaly response map, including: performing anomaly assessment based on the anomaly response map to obtain pixel-level anomaly scores, and aggregating the pixel-level anomaly scores according to a set region to obtain a region-level anomaly score; determining an anomaly assessment score based on the pixel-level anomaly score and the region-level anomaly score, and using the anomaly assessment score as the anomaly assessment result.
[0175] The pixel-level anomaly score can be determined based on the confidence value of each pixel in the anomaly response image. The defined region can be a predefined key monitoring area (i.e., region of interest, ROI). For example, the defined region could be a location requiring focused monitoring, such as an electrical cabinet door, charging port, or distribution panel, typically marked with its coordinates in the image using masking techniques. The region-level anomaly score can be the anomaly score obtained by aggregating all pixel-level anomaly scores within each defined region. The anomaly evaluation score can be the final score determined by fusing the pixel-level and region-level anomaly scores.
[0176] In this embodiment, the abnormal response map can be structurally evaluated and its stability determined, and the abnormal response map output by the front-end model can be transformed into an abnormal scoring signal that is quantifiable, comparable, and can trigger system decisions.
[0177] In this embodiment, the pixel-level anomaly probability map output by the multi-scale visual Transformer (MSVT) model at time t is... A multi-scale anomaly intensity characterization is constructed to depict the spatial distribution characteristics of anomalies. This represents the pixel location in the image. Pixel-level anomaly scores can be defined as follows:
[0178] ;
[0179] In this embodiment, to further describe the degree of anomaly aggregation in spatial structure, the system considers connected regions or regions of interest (ROIs). The pixel-level anomaly scores are aggregated by region to obtain a region-level anomaly intensity index, i.e., the region-level anomaly score can be:
[0180] ;
[0181] in, Indicates the region The number of pixels. Pixel-level anomaly scores characterize local, fine-grained anomaly responses, while region-level anomaly scores reflect the overall trend of anomalies in spatial structure.
[0182] In this embodiment, the anomaly evaluation score is determined based on the pixel-level anomaly score and the region-level anomaly score. This can be achieved by merging the weighted average of the pixel-level anomaly score and the region-level anomaly score, or by first screening out anomaly pixel clusters whose pixel-level scores exceed a preset threshold, calculating their overlap with the corresponding region of interest, and then weighting and correcting the region-level score based on the overlap to finally obtain the anomaly evaluation score, which is then used as the anomaly evaluation result.
[0183] In this embodiment, by setting up the joint modeling of pixel-level and region-level anomaly scores and determining the final anomaly evaluation score, a multi-scale anomaly representation is formed within a single frame. This not only retains the sensitivity to small local anomalies but also enhances the robust characterization of overall structural anomalies.
[0184] In this embodiment, optionally, determining the corresponding alarm level based on the anomaly assessment result includes: determining the corresponding anomaly judgment threshold and anomaly indicator data based on the anomaly assessment result; determining whether the anomaly judgment threshold and anomaly indicator data meet the alarm triggering conditions; and, if the anomaly judgment threshold and anomaly indicator data meet the alarm triggering conditions, classifying the corresponding alarm level based on the correspondence between the anomaly judgment threshold and the anomaly indicator data.
[0185] The anomaly detection threshold is a learned threshold obtained by statistically analyzing the temporal characteristics of anomaly scores, combined with a predefined threshold to arrive at the final anomaly detection threshold. Anomaly indicator data can refer to predefined indicator data. In this embodiment, the anomaly indicator data can include frame-level anomaly indicator data and region-level anomaly indicator data.
[0186] This embodiment is based on anomaly assessment scores. Define frame-level and region-level anomaly intensity indices as follows:
[0187] ;
[0188] ;
[0189] in, Indicates pixel position, Represents a predefined region of interest (ROI). For a set of regions.
[0190] In this embodiment, the alarm triggering condition can be a pre-defined alarm judgment rule. In this embodiment, during the abnormal triggering phase, a judgment rule with hysteresis characteristics is used to suppress frequent alarms caused by short-term fluctuations. The triggering logic of its alarm triggering condition can be:
[0191] ;
[0192] ;
[0193] ;
[0194] in, This is the minimum number of frames required to ensure that the anomaly has temporal continuity. This is the hysteresis bandwidth, used to prevent repeated switching caused by abnormal scores oscillating around the threshold.
[0195] In this embodiment, the system determines whether the anomaly detection threshold and the abnormal indicator data meet the preset alarm triggering conditions. If the anomaly detection threshold and the abnormal indicator data meet the preset alarm triggering conditions, the corresponding alarm level can be divided based on the correspondence between the anomaly detection threshold and the abnormal indicator data. If the anomaly detection threshold and the abnormal indicator data do not meet the preset alarm triggering conditions, no alarm operation is required.
[0196] In this embodiment, through the collaborative design of multi-scale voting and hysteresis judgment, pixel-level, region-level and frame-level anomaly information are integrated to verify the consistency of abnormal behavior from different spatial scales, so as to avoid unstable alarms caused by misjudgment at a single scale. It can effectively distinguish between real continuous anomalies and short-term noise disturbances, and avoid frequent system actions triggered by local anomalies or short-term drift.
[0197] In this embodiment, when the anomaly detection threshold and anomaly indicator data meet the alarm triggering conditions, the corresponding alarm level is determined based on the correspondence between the anomaly detection threshold and the anomaly indicator data. Specifically, this can be based on the regional anomaly score. With anomaly detection threshold Based on the relative relationship, the severity of anomalies is divided into three alarm levels to differentiate response strategies for different risk levels. These are Level 1, Level 2, and Level 3 alarms, as follows:
[0198] Level I (Mild Alarm): ;
[0199] Level II (Moderate Alert): ;
[0200] Level III (Severe Alert): ;
[0201] in, These are the regional anomaly scores after multi-scale aggregation and time-series stabilization. This is the threshold for anomaly detection. In this embodiment, a Level 1 alarm can indicate the initial appearance of an anomaly or a localized disturbance, triggering enhanced monitoring. In this embodiment, a Level 2 alarm can indicate that the anomaly has a certain degree of persistence and spatial consistency, requiring close monitoring. In this embodiment, a Level 2 alarm can indicate that the anomaly is significant and stable, triggering advanced response or linkage mechanisms.
[0202] Furthermore, in this embodiment, at the system interface level, alarm results can be displayed in real time as anomaly heatmaps, ROI boundary annotations, and timeline overlay curves, while simultaneously recording metadata such as anomaly duration, threshold evolution process, and alarm level. This visualization not only serves as an immediate alarm and manual confirmation but also as a historical record of system operation status, providing a reliable basis for subsequent anomaly tracing analysis, model fine-tuning triggering, or model rollback decisions.
[0203] In this embodiment, by setting up such an alarm, spatial aggregation and temporal stabilization processing can be introduced during the abnormal score generation stage, which can effectively suppress the risk of false alarms caused by local noise or instantaneous fluctuations, ensure the continuity and controllability of alarm triggering, and perform different alarm processing according to different alarm levels, thereby improving the reliability of alarms.
[0204] In this embodiment, optionally, determining the corresponding anomaly judgment threshold based on the anomaly evaluation result includes: performing temporal smoothing and spatial consistency processing on the anomaly evaluation result to obtain the processed anomaly evaluation result; performing data statistics based on the processed anomaly evaluation result to determine the initial learning threshold, and introducing a correction term to compensate the initial learning threshold to obtain the target learning threshold; and determining the anomaly judgment threshold according to the target learning threshold and the set threshold.
[0205] Temporal smoothing can be a temporal filtering operation performed on the anomaly evaluation results of consecutive video frames. Spatial uniformity can be a spatial constraint operation performed on the anomaly evaluation results. The processed anomaly evaluation results can refer to the evaluation results obtained after temporal smoothing and spatial uniformity processing.
[0206] In this embodiment, considering the potential for abnormal score fluctuations in the surveillance video sequence caused by lighting flicker, background disturbances, or momentary occlusion, the system introduces a time-weighted sliding smoothing mechanism at the abnormal scoring level to perform short-term stabilization processing on the abnormal assessment results. In this embodiment, the abnormal assessment score at the current moment can be set as... Its smoothing result is defined as:
[0207] ;
[0208] in, This represents the outlier score after time smoothing. This is a smoothing coefficient used to adjust the relative weight of the current observation and historical states in the anomaly scoring. In this embodiment, it can be... Considered a short-term stability modulation parameter, its value can be adaptively adjusted according to scene complexity, motion intensity, or acquisition frame rate. When the image is detected to be in a relatively stable state, such as low motion or slow structural changes, the system increases... Strengthen time consistency constraints to suppress the interference of random noise on abnormal scores; when the scene exhibits rapid changes, reduce... This improves the system's response speed to real anomalies.
[0209] To avoid false anomaly alarms caused by local noise, high-response isolated points, or imaging interference, this embodiment introduces spatial consistency constraints at the anomaly heatmap level to stabilize the spatial structure of anomaly scores. This embodiment utilizes a Total Variation (TV) regularization mechanism to smooth the local structure of the anomaly heatmap, suppressing random noise while preserving the spatial boundary features of the anomaly region. Specifically, the anomaly heatmap corresponding to the time-smoothed anomaly evaluation result is defined as... The system obtains the spatially consistent anomaly response graph by solving the following optimization problem. :
[0210] ;
[0211] in, Represents the optimization variable. This is the spatial smoothing intensity coefficient, used to balance the fidelity of anomalous responses with spatial continuity; Represents the spatial gradient operator, whose The norm is used to encourage piecewise smooth structures. This optimization process can effectively eliminate spatially isolated high-response noise points while maintaining the contour integrity and connectivity of anomalous regions.
[0212] The initial learning threshold can refer to a learning threshold. In this embodiment, statistical modeling can be performed on the processed anomaly evaluation results to evaluate the initial learning threshold constructed from the mean and fluctuation range. The correction term can refer to an uncertainty correction term, which can be a pre-set correction term. The target learning threshold can be the learning threshold obtained after compensating the initial learning threshold with the correction term. The set threshold can be a pre-set threshold or a threshold set based on long-term operating experience. The anomaly judgment threshold can be the final threshold for anomaly alarm judgment.
[0213] In this embodiment, the statistical distribution of anomaly scores varies significantly under different camera angles, lighting conditions, and operating environments. To avoid the failure of fixed thresholds in complex monitoring scenarios, this embodiment proposes an adaptive learning threshold modeling mechanism based on anomaly distribution statistics. By continuously modeling the temporal statistical characteristics of anomaly scores, dynamic updates to the anomaly judgment criteria and environmental adaptation are achieved.
[0214] Specifically, in length of Within the sliding time window, outlier scores after temporal smoothing and spatial consistency processing are... Statistical modeling is performed to dynamically estimate its mean and fluctuation range. A learned threshold is constructed based on the mean-variance model, thus obtaining the initial learned threshold:
[0215] ;
[0216] in, and Representing windows respectively Mean and standard deviation of scores for internal abnormalities This is the sensitivity adjustment coefficient, used to control the sensitivity of anomaly detection to score fluctuations.
[0217] To further enhance robustness to extreme anomalies, this embodiment also introduces a threshold estimation method based on quantile statistics:
[0218] ;
[0219] in, The quantile threshold function is typically taken as... These thresholds are used to characterize the tail features of abnormal distributions. The two types of thresholds mentioned above can be combined or switched depending on scene stability and operational requirements.
[0220] Furthermore, to characterize the impact of model prediction uncertainty on anomaly detection, an uncertainty correction term is introduced to compensate for the learning threshold, thus obtaining the target learning threshold:
[0221] ;
[0222] in, This indicates the variance or confidence uncertainty of the model prediction at the current moment. This is the risk compensation coefficient. When the model output is unstable or the confidence level is low, the system automatically increases the anomaly detection threshold to reduce the risk of false alarms.
[0223] It should be noted that the above statistics are not only used to calculate the real-time anomaly detection threshold, but also to continuously monitor the evolution of anomaly distribution. When the sliding window... , When the quantile statistics show a continuous shift or a significant increase in fluctuation, the system identifies it as a potential abnormal distribution drift signal and records it as an important decision-making basis for triggering subsequent model fine-tuning, collaborative updates, or rollback mechanisms.
[0224] In this embodiment, in actual engineering deployment, different monitoring nodes typically already have manually set thresholds based on long-term operational experience. To achieve a balance between stability and adaptability, this embodiment can use a dynamic fusion of empirical and learned thresholds to construct anomaly detection thresholds with safety anchor characteristics.
[0225] Specifically, in this embodiment, the adaptive target learning threshold obtained based on anomaly distribution statistics can be used. Compared with human experience threshold We perform weighted fusion to obtain the final anomaly detection threshold:
[0226] ;
[0227] in, The fusion weights are dynamically adjusted to regulate the system's dependence on the learning threshold and the empirical threshold. In this embodiment, the fusion weights... It adaptively adjusts based on the current video quality and scene status, and its calculation form is as follows:
[0228] ;
[0229] in, This is the Sigmoid function, used to restrict the weights to the (0,1) interval; This indicates the intensity of anomalies within the current frame or time window. This indicates the level of image jitter or motion complexity. Indicates brightness or image quality indicators. , , For the corresponding normalization constant, This is the adjustment coefficient.
[0230] In this embodiment, through the aforementioned fusion mechanism, the system can automatically reduce [the quality of the monitored images] when the quality is low, the environmental fluctuations are large, or the abnormal distribution is not yet stable. Enhance the understanding of empirical thresholds This avoids over-adaptation due to short-term statistical biases; and under conditions of stable image quality and reliable anomaly distribution, the system gradually improves... This fully leverages the learning threshold's ability to characterize complex anomaly patterns.
[0231] Furthermore, this embodiment also constructs an edge self-healing update mechanism with drift perception, event triggering, and safe rollback capabilities to ensure the stability and reliability of the anomaly detection model in complex and dynamic scenarios. In this embodiment, the online update process at the edge is elevated to a controlled system-level evolution process. Limited and traceable online fine-tuning operations are only triggered when anomaly distribution or continuous changes in feature structure are detected. The cumulative spread of erroneous updates is avoided through model version evaluation and automatic rollback mechanisms.
[0232] In this embodiment, after the model is deployed to edge nodes and enters a long-term online operation phase, the system continuously monitors the model's operating status and input data distribution to identify potential scene changes and model mismatch risks. This embodiment no longer uses abnormal statistical results merely as alarm criteria, but upgrades them to a prerequisite criterion for model reliability assessment and update decisions, used to determine whether there are distribution drift signals requiring intervention.
[0233] Specifically, the system comprehensively utilizes multidimensional statistical information to construct drift sensing signals, including: the mean of the anomaly score distribution within the time window. With variance The system analyzes the continuous shifts in distribution, structural changes in quantile threshold trajectories, and the degree of shift of the feature centers of abnormal regions relative to historical stable states. Simultaneously, the system incorporates update stability metrics (such as parameter update variance or gradient fluctuation amplitude) output from the local training phase to measure the model's convergence reliability under current data conditions. When these statistics exhibit a consistent shift or amplified fluctuation trend over multiple consecutive time windows, the system identifies them as potential distribution drift signals, indicating a decline in model reliability.
[0234] In this embodiment, to avoid model instability or the spread of accumulated errors due to continuous online learning during long-term operation of edge nodes, an event-triggered online fine-tuning and controlled update mechanism is adopted. Only when a clear risk of distribution drift is detected is a limited-range model adaptive repair process initiated. This mechanism uses identified drift signals as trigger conditions. When abnormal score distributions, feature statistics, or update stability indices continuously deviate from historical steady states within a continuous time window, and the drift intensity exceeds a preset threshold, the system determines that the current model can no longer adequately adapt to the new scenario, thus triggering online fine-tuning. Online fine-tuning employs a controlled strategy of small steps, few iterations, and local parameter updates to avoid drastic disturbances to the overall model structure. Specifically, edge nodes adjust model parameters based on recently cached high-confidence samples. Performing a finite-step gradient update can be represented as follows:
[0235] ;
[0236] in, The online learning rate is significantly lower than that used in the offline training phase; This represents the local detection loss based on the current task; This is a knowledge distillation constraint term used to maintain consistency between the model output and historical stable versions; For importance weights estimated based on the Fisher information matrix, This represents historically stable parameters, used to prevent catastrophic forgetting during online fine-tuning.
[0237] In this embodiment, to further suppress the impact of short-term noise on parameter updates, the system introduces an exponential moving average (EMA) mechanism to smooth the model parameters after each fine-tuning:
[0238] ;
[0239] in, The smoothing coefficient is used. Through the above design, the model update process is transformed from "continuous self-learning" to controlled intervention aimed at drift repair, ensuring that online updates are only used to correct distribution shifts rather than relearning task semantics, thereby achieving rapid adaptation to new scenarios while maintaining model stability.
[0240] In this embodiment, to ensure the security and controllability of the edge model during long-term online operation and adaptive updates, a model version evaluation and safety rollback mechanism is introduced on the edge node side. This mechanism is used to continuously evaluate the new model version generated by online fine-tuning, and automatically perform model rollback or version freeze operations when performance degradation or increased uncertainty risk is detected, thereby preventing erroneous updates from having a cumulative impact on actual monitoring tasks.
[0241] Specifically, after each event-triggered online fine-tuning is completed, the system will adjust the parameters of the current candidate model. Compared to historical stable versions A comparative evaluation was conducted. The evaluation metrics included not only detection performance indicators based on recent samples (such as anomaly score distribution stability and regional anomaly consistency), but also a comprehensive consideration of changes in model prediction uncertainty and update stability. Uncertainty was assessed through prediction entropy. The change in variance of outlier fractions can be characterized, and the stability statistic defined above is used to update the stability.
[0242] In this embodiment, when any of the following conditions are met, the system determines that the current update is risky. The specific conditions are: (1) The detection performance of the new model in the verification window is significantly lower than that of the reference version; (2) The prediction uncertainty continues to increase, indicating that the model is not confident in the current scene; (3) The parameter update amplitude or abnormal distribution fluctuation exceeds the safety threshold, and there is a risk of overfitting or drift amplification.
[0243] In the above situation, the system will automatically trigger a safety rollback strategy, specifically by stopping the use of... Perform online inference and restore the use of the reference model. or its exponential moving average version As the current inference model, this rollback process does not rely on human intervention and does not affect the system's continuous operation capability, thus ensuring that edge nodes maintain a reliable and stable operating state while allowing exploratory adaptive updates.
[0244] This embodiment significantly improves the security and engineering reliability of model updates in online learning scenarios by introducing model version evaluation and security rollback mechanisms, providing key guarantees for long-term deployment in complex dynamic monitoring environments.
[0245] To ensure the stability and reliability of the edge anomaly detection model under long-term, unattended operation, this embodiment introduces a unified health monitoring and self-healing control mechanism at the edge. This mechanism continuously evaluates the model's operating status, inference quality, and system load, triggering a layered self-recovery strategy accordingly. Specifically, the system constructs a comprehensive health index. Used to characterize edge nodes at time 10 ... The overall operating status can be:
[0246] ;
[0247] in, This represents an indicator related to inference latency and real-time performance. This indicates the level of energy consumption and resource usage. Indicators representing model stability and drift (e.g., fluctuations in outlier distribution, feature center shift, or changes in update variance); These are weighting coefficients used to balance the impact of different dimensions on the health assessment. All the above indicators are derived from the operational monitoring and online update feedback in the preceding steps and do not involve raw data or privacy information.
[0248] In this embodiment, when the overall health level When the model's performance falls below a preset threshold, a tiered self-healing response strategy is not adopted instead of a single processing method. For mild anomalies, rapid recovery is prioritized by adjusting quantization intensity, inference resolution, or smoothing parameters. For moderate anomalies, event-driven online fine-tuning or model parameter recalibration is triggered. For continuous deterioration or severe anomalies, the system automatically reverts to a historically stable model version, and if necessary, reports to the cloud for remote diagnosis and model redeployment. Through this mechanism, the system can proactively repair and isolate risks related to model performance degradation without interrupting real-time monitoring tasks. This embodiment constructs a system-level self-healing strategy for edge scenarios, enabling the model to maintain availability and reliability in complex and dynamic environments, thereby supporting the long-term stable operation of the aforementioned federated collaborative learning and anomaly detection system.
[0249] The technical solution of this invention acquires multi-source video data collected by multiple monitoring devices through edge nodes, and preprocesses the multi-source video data to obtain target video data. The target video data is then divided and encoded according to a set size to obtain an initial video frame sequence. A geometric description vector is determined, and the geometric description vector and the initial video frame sequence are fused to obtain the target video frame sequence. The target video frame sequence is input into a target anomaly detection model, which outputs a corresponding anomaly response map. Anomaly evaluation is performed based on the anomaly response map to obtain anomaly evaluation results. The corresponding alarm level is determined based on the anomaly evaluation results, and alarm operations are performed according to the alarm level. This technical solution solves the problems of weak privacy protection and insufficient cross-scene generalization ability in existing video anomaly detection technologies, improving the accuracy, real-time performance, and stability of video anomaly detection.
[0250] Example 3
[0251] Figure 3 This is a schematic diagram of a video anomaly detection device based on a federated vision Transformer according to Embodiment 3 of the present invention. Figure 3 As shown, the device includes:
[0252] The data acquisition module 310 is used to acquire multi-source video data collected by multiple monitoring devices through edge nodes, and to preprocess the multi-source video data to obtain target video data.
[0253] The data processing module 320 is used to perform block processing and encoding processing on the target video data to obtain the target video frame sequence;
[0254] The anomaly detection module 330 is used to input the target video frame sequence into the target anomaly detection model and output the corresponding anomaly response map; wherein, the target anomaly detection model is trained based on the federated learning mechanism and the initial anomaly detection model;
[0255] The anomaly assessment module 340 is used to perform anomaly assessment based on the anomaly response graph to obtain anomaly assessment results, determine the corresponding alarm level based on the anomaly assessment results, and perform alarm operations according to the alarm level.
[0256] Optional, the data processing module 320 is specifically used for:
[0257] The target video data is divided and encoded according to a set size to obtain an initial video frame sequence;
[0258] Determine the geometric description vector, and fuse the geometric description vector with the initial video frame sequence to obtain the target video frame sequence.
[0259] Optionally, the device may also include the following modules for training a target anomaly detection model:
[0260] The sample acquisition module is used to acquire the training sample set and the initial anomaly detection model for each edge node; the training sample set is obtained by filtering through the scene stability index.
[0261] The model parameter acquisition module is used to train the initial anomaly detection model based on each training sample contained in the training sample set, and obtain the model parameters of each edge node in the current training round; among which, the model parameters include model performance indicators and stability indicators.
[0262] The state vector construction module is used to construct the node state vector corresponding to each edge node based on the model parameters through the federated coordinator, and determine the set of fully participating nodes, the set of non-fully participating nodes, and the set of observation nodes based on the node state vectors.
[0263] The parameter update module is used to update the global model parameters based on the model parameters of the full set of participating nodes and the set update operators;
[0264] The parameter distribution module is used to distribute global model parameters to each edge node through the federated coordinator, so that each edge node can continue to perform the training operation of the initial anomaly detection model in the next training round based on the global model parameters, until the target anomaly detection model is obtained.
[0265] Optional, the anomaly assessment module 340 is specifically used for:
[0266] Anomaly assessment is performed based on the anomaly response map to obtain pixel-level anomaly scores, and these pixel-level anomaly scores are aggregated according to a set region to obtain region-level anomaly scores.
[0267] An anomaly assessment score is determined based on pixel-level and region-level anomaly scores, and this score is used as the anomaly assessment result.
[0268] Optional, the anomaly assessment module 340 includes:
[0269] The data determination unit is used to determine the corresponding anomaly judgment threshold and anomaly indicator data based on the anomaly assessment results.
[0270] The judgment unit is used to determine whether the anomaly judgment threshold and anomaly indicator data meet the alarm triggering conditions.
[0271] The alarm level classification unit is used to classify the corresponding alarm level based on the correspondence between the anomaly judgment threshold and the anomaly indicator data when the anomaly judgment threshold and the anomaly indicator data meet the alarm triggering conditions.
[0272] Optional, data determination unit, specifically used for:
[0273] The anomaly assessment results are subjected to temporal smoothing and spatial consistency processing to obtain the processed anomaly assessment results.
[0274] The initial learning threshold is determined by statistical analysis of the processed anomaly evaluation results, and a correction term is introduced to compensate for the initial learning threshold to obtain the target learning threshold.
[0275] The anomaly detection threshold is determined based on the target learning threshold and the set threshold.
[0276] Optionally, the data acquisition module 310 is specifically used for:
[0277] The multi-source video data is normalized, denoised, and stabilized to obtain the first image data; wherein the first image is geometrically consistent image data aligned with the homography matrix.
[0278] A target region mask is generated using a segmentation algorithm. Based on the target region mask, the first image data is subjected to region blurring to obtain the target video data.
[0279] The video anomaly detection device based on federated vision Transformer provided in this embodiment of the invention can execute the video anomaly detection method based on federated vision Transformer provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method execution.
[0280] Example 4
[0281] Figure 4 This is a schematic diagram of an electronic device according to Embodiment 4 of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0282] like Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0283] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0284] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as a video anomaly detection method based on a federated vision Transformer.
[0285] In some embodiments, the video anomaly detection method based on Federated Vision Transformer can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the video anomaly detection method based on Federated Vision Transformer described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the video anomaly detection method based on Federated Vision Transformer by any other suitable means (e.g., by means of firmware).
[0286] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0287] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0288] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0289] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0290] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0291] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0292] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0293] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A video anomaly detection method based on Federated Vision Transformer, characterized in that, include: The target video data is obtained by acquiring multi-source video data collected by multiple monitoring devices through edge nodes and preprocessing the multi-source video data. The target video data is segmented and encoded to obtain a target video frame sequence; The target video frame sequence is input into the target anomaly detection model, and the corresponding anomaly response map is output; wherein, the target anomaly detection model is trained based on a federated learning mechanism and an initial anomaly detection model; the target anomaly detection model is a multi-scale visual Transformer model; An anomaly assessment is performed based on the anomaly response graph to obtain an anomaly assessment result. The corresponding alarm level is determined based on the anomaly assessment result, and an alarm operation is performed based on the alarm level.
2. The method according to claim 1, characterized in that, The target video data is divided into blocks and encoded to obtain a target video frame sequence, including: The target video data is divided and encoded according to a set size to obtain an initial video frame sequence; A geometric description vector is determined, and the geometric description vector and the initial video frame sequence are fused to obtain the target video frame sequence.
3. The method according to claim 1, characterized in that, The training steps of the target anomaly detection model include: Obtain the training sample set for each edge node and the initial anomaly detection model; wherein, the training sample set is obtained by filtering through the scene stability index; The initial anomaly detection model is trained based on each training sample contained in the training sample set to obtain the model parameters of each edge node in the current training round; wherein, the model parameters include model performance indicators and stability indicators. The federated coordinator constructs a node state vector for each edge node based on the model parameters, and determines the set of nodes with full participation, the set of nodes with non-full participation, and the set of observation nodes based on the node state vectors. The global model parameters are obtained by updating the parameters based on the model parameters of the set of all participating nodes and the set of update operators; The global model parameters are distributed to each edge node through the federated coordinator, so that each edge node continues to perform the training operation of the initial anomaly detection model in the next training round based on the global model parameters, until the target anomaly detection model is obtained.
4. The method according to claim 1, characterized in that, Anomaly assessment results are obtained by performing anomaly assessment based on the aforementioned anomaly response graph, including: Anomaly assessment is performed based on the anomaly response map to obtain pixel-level anomaly scores, and the pixel-level anomaly scores are aggregated according to a set region to obtain region-level anomaly scores. An anomaly assessment score is determined based on the pixel-level anomaly score and the region-level anomaly score, and the anomaly assessment score is used as the anomaly assessment result.
5. The method according to claim 1, characterized in that, The corresponding alarm level is determined based on the anomaly assessment results, including: Based on the anomaly assessment results, the corresponding anomaly judgment threshold and anomaly indicator data are determined. Determine whether the anomaly determination threshold and the anomaly indicator data meet the alarm triggering conditions; If the anomaly determination threshold and the anomaly indicator data meet the alarm triggering conditions, the corresponding alarm level is determined based on the correspondence between the anomaly determination threshold and the anomaly indicator data.
6. The method according to claim 5, characterized in that, Based on the anomaly assessment results, the corresponding anomaly determination threshold is determined, including: The anomaly assessment results are subjected to temporal smoothing and spatial consistency processing to obtain the processed anomaly assessment results; Based on the processed anomaly evaluation results, data statistics are performed to determine the initial learning threshold, and a correction term is introduced to compensate for the initial learning threshold to obtain the target learning threshold. The anomaly detection threshold is determined based on the target learning threshold and the set threshold.
7. The method according to claim 1, characterized in that, The multi-source video data is preprocessed to obtain target video data, including: The multi-source video data is normalized, denoised, and stabilized to obtain first image data; wherein, the first image is geometrically consistent image data aligned by a homography matrix. A target region mask is generated using a segmentation algorithm, and the first image data is then subjected to region blurring based on the target region mask to obtain the target video data.
8. A video anomaly detection device based on Federated Vision Transformer, characterized in that, include: The data acquisition module is used to acquire multi-source video data collected by multiple monitoring devices through edge nodes, and to preprocess the multi-source video data to obtain target video data. The data processing module is used to perform block processing and encoding processing on the target video data to obtain a target video frame sequence; An anomaly detection module is used to input the target video frame sequence into a target anomaly detection model and output the corresponding anomaly response map; wherein, the target anomaly detection model is trained based on a federated learning mechanism and an initial anomaly detection model; the target anomaly detection model is a multi-scale visual Transformer model; The anomaly assessment module performs anomaly assessment based on the anomaly response graph to obtain anomaly assessment results, determines the corresponding alarm level based on the anomaly assessment results, and performs alarm operations based on the alarm level.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the video anomaly detection method based on Federated Vision Transformer as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the video anomaly detection method based on any one of claims 1-7 using a federated vision Transformer.
Citation Information
Cited By
A police video image intelligent analysis system based on artificial intelligence
CN122223631A