Scene anomaly detection method, device, storage medium and computer equipment
Through the combination of multimodal data fusion and heterogeneous discriminator groups, the performance instability of single-modal detection method in complex environments is solved, and scene anomaly detection with higher accuracy and interpretability is achieved.
Patent Information
- Application Number
- CN202510663218.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-05-22
AI Technical Summary
The existing single-modal anomaly detection methods have unstable performance in complex environments, making it difficult to cope with ambient light changes, occlusion interference and complex backgrounds, and lack interpretability, resulting in unsatisfactory detection results.
Multimodal data fusion technology is used, combined with infrared and visible images, scene anomaly detection is performed through the methods of timing alignment, spatial registration, convolutional network feature extraction, feature fusion and heterogeneous discriminator group.
It improves the accuracy and interpretability of scene abnormality detection, reduces the false detection rate and missed detection rate, and enhances the transparency and credibility of the detection results.
Smart Images

Figure CN120182901B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of scene detection technology, and in particular to a scene anomaly detection method, device, storage medium and computer equipment. Background Art
[0002] In application scenarios such as intelligent monitoring, industrial inspection and autonomous driving, anomaly detection technology faces challenges brought by complex environments. Traditional single-modal detection methods usually rely on a single sensor input, such as visible light images or infrared images, but these methods are difficult to cope with factors such as ambient lighting changes, occlusion interference and complex backgrounds, resulting in unstable performance. The limitations of single-modality are particularly obvious in low light or bad weather conditions.
[0003] With the development of multi-sensor technology, the combination of multimodal data such as infrared and visible light provides complementary information for scene understanding. Visible light images provide rich texture and color features, but their performance degrades in low light; while infrared images can reflect the thermal radiation characteristics of objects, which are advantageous in abnormal temperature conditions, but lack detailed information. Therefore, how to fuse these heterogeneous modal data and explore synergistic effects has become the key to improving the robustness of anomaly detection.
[0004] Existing anomaly detection methods usually adopt simple feature splicing or decision-level fusion strategies, which are difficult to fully capture the deep correlation between modalities, resulting in unsatisfactory fusion effects; at the same time, the lack of interpretable analysis of abnormal areas restricts its value and usability in practical applications. Although the rapid development of deep learning technology has provided new ideas for multimodal anomaly detection, current solutions still face multiple problems. On the one hand, information redundancy or semantic conflicts are prone to occur during cross-modal feature fusion, affecting the discriminative ability of fused features; on the other hand, most detection methods in related technologies lack interpretability, affecting users' trust in the detection results. Summary of the invention
[0005] The embodiments of the present disclosure at least provide a scene anomaly detection method, apparatus, storage medium and computer equipment, which can effectively improve the accuracy and interpretability of scene anomaly detection by introducing a combination of multimodal data fusion, a preset three-branch generator architecture and a heterogeneous discriminator group.
[0006] The present disclosure provides a method for detecting anomalies in a scene, including:
[0007] Acquire an infrared image and a visible light image corresponding to the scene to be detected, and perform temporal alignment processing and spatial alignment processing on the infrared image and the visible light image to obtain infrared modal data and visible light modal data respectively;
[0008] Extract features from the infrared modality data using a convolutional network to obtain infrared features corresponding to the infrared modality data; and, extract features from the visible light modality data using a convolutional network to obtain visible light features corresponding to the visible light modality data;
[0009] Perform feature fusion processing on the infrared features and the visible light features to obtain multi-modal fusion features;
[0010] Determine a set of generation results based on a preset three-branch generator architecture and the multi-modal fusion features; and perform anomaly detection on the set of generation results through a heterogeneous discriminator group to obtain a scene anomaly detection result for the scene to be detected.
[0011] In some possible embodiments, the performing temporal alignment processing and spatial registration processing on the infrared image and the visible light image includes:
[0012] Taking the acquisition time information of any one of the infrared image and the visible light image as a reference, perform timestamp matching on the other image to complete the temporal alignment processing of the infrared image and the visible light image;
[0013] Extract the image feature point information of the infrared image and the visible light image respectively, and determine the geometric transformation relationship between the two images based on the extraction results and a matching algorithm to determine a transformation matrix;
[0014] Perform geometric transformation on any one of the infrared image and the visible light image according to the transformation matrix to complete the spatial registration processing of the infrared image and the visible light image.
[0015] In some possible embodiments, the performing feature fusion processing on the infrared features and the visible light features includes:
[0016] Calculate the feature cross-correlation between the infrared features and the visible light features to construct a cross-modal attention weight matrix;
[0017] Determine an initial multi-modal fusion feature based on the cross-modal attention weight matrix, the infrared features and the visible light features;
[0018] Perform recalibration processing on the initial multi-modal fusion feature, and perform feature enhancement processing on the recalibration processing result based on a multi-scale spatial enhancement strategy to obtain the multi-modal fusion feature.
[0019] In some possible embodiments, the preset three-branch generator architecture includes a local feature generator, a global feature generator, and a spatio-temporal feature generator, and the set of generation results includes local feature generation results, global feature generation results, and spatio-temporal feature generation results; determining the set of generation results based on the preset three-branch generator architecture and the multi-modal fusion features includes:
[0020] Using the local feature generator to perform local feature extraction processing on the multi-modal fusion features to obtain the local feature generation results; wherein, the local feature generator includes an encoding-decoding framework and a skip connection structure;
[0021] Using the global feature generator to perform global feature extraction processing on the multi-modal fusion features to obtain the global feature generation results; wherein, the global feature generator includes a self-attention mechanism and a neural network;
[0022] Using the spatio-temporal feature generator to perform spatio-temporal dimensional feature extraction processing on the multi-modal fusion features to obtain the spatio-temporal feature generation results; wherein, the spatio-temporal feature generator includes a three-dimensional convolutional neural network and a gated recurrent unit.
[0023] In some possible embodiments, the heterogeneous discriminator group includes a local anomaly detection discriminator, a global semantic discriminator, and a temporal consistency discriminator; performing anomaly detection on the set of generation results through the heterogeneous discriminator group includes:
[0024] Based on the local anomaly detection discriminator, evaluating the anomaly situation of the texture and structure of the local region in the local feature generation results to obtain a local anomaly probability;
[0025] Based on the global semantic discriminator, evaluating the semantic consistency between the global feature generation results and the multi-modal fusion features to obtain a global anomaly probability;
[0026] Based on the temporal consistency discriminator, evaluating the physical rationality of the motion trajectory in the spatio-temporal feature generation results to obtain a temporal anomaly probability.
[0027] In some possible embodiments, after performing anomaly detection on the set of generation results through the heterogeneous discriminator group, it further includes:
[0028] In the case where at least two of the local anomaly probability, the global anomaly probability, and the temporal anomaly probability are greater than the first warning threshold, it is determined as a primary anomaly;
[0029] In the case of being determined as a primary anomaly, the local anomaly probability, the global anomaly probability, and the temporal anomaly probability are weighted and calculated according to preset probability distribution weights to obtain a comprehensive anomaly probability; in the case where the value of the comprehensive anomaly probability is greater than a second warning threshold, it is determined as an ultimate anomaly;
[0030] Based on the anomaly determination situation, determine the scene anomaly detection result regarding the to-be-detected scene.
[0031] In some possible embodiments, after determining the scene anomaly detection result regarding the to-be-detected scene based on the anomaly determination situation, it includes:
[0032] Construct a fully connected graph with three discriminators as nodes and the anomaly detection results of the three discriminators as edge weights;
[0033] Calculate the spectral entropy value of the fully connected graph;
[0034] In the case where the spectral entropy value is lower than a preset entropy threshold, directly output the scene anomaly detection result;
[0035] In the case where the spectral entropy value is not lower than the preset entropy threshold, mark the scene anomaly detection result as a complex anomaly, and based on the anomaly monitoring results of each discriminator and the scene anomaly detection result, generate a complex anomaly result report and send the complex anomaly result report to the manual analysis end.
[0036] An embodiment of the present disclosure provides a scene anomaly detection device, including:
[0037] An image processing module, configured to obtain an infrared image and a visible light image corresponding to a to-be-detected scene, and perform temporal alignment processing and spatial registration processing on the infrared image and the visible light image to obtain infrared modality data and visible light modality data respectively;
[0038] A feature extraction module, configured to use a convolutional network to extract features from the infrared modality data to obtain infrared features corresponding to the infrared modality data; and use a convolutional network to extract features from the visible light modality data to obtain visible light features corresponding to the visible light modality data;
[0039] A feature fusion module, configured to perform feature fusion processing on the infrared features and the visible light features to obtain multi-modal fusion features;
[0040] An anomaly detection module, configured to determine a set of generation results based on a preset three-branch generator architecture and the multi-modal fusion features; and perform anomaly detection on the set of generation results through a heterogeneous discriminator group to obtain a scene anomaly detection result regarding the to-be-detected scene.
[0041] In some possible embodiments, the image processing module is specifically configured to:
[0042] Based on the acquisition time information of any one of the infrared image and the visible light image, perform timestamp matching on the other image to complete the timing alignment process of the infrared image and the visible light image;
[0043] Extract the image feature point information of the infrared image and the visible light image respectively, and determine the geometric transformation relationship between the two images based on the extraction results and the matching algorithm, and determine the transformation matrix;
[0044] Perform geometric transformation on any one of the infrared image and the visible light image according to the transformation matrix to complete the spatial registration process of the infrared image and the visible light image.
[0045] In some possible embodiments, the feature fusion module is specifically configured to:
[0046] Calculate the feature cross-correlation between the infrared feature and the visible light feature, and construct a cross-modal attention weight matrix;
[0047] Determine the initial multi-modal fusion feature based on the cross-modal attention weight matrix, the infrared feature and the visible light feature;
[0048] Perform recalibration processing on the initial multi-modal fusion feature, and perform feature enhancement processing on the recalibration processing result based on the multi-scale spatial enhancement strategy to obtain the multi-modal fusion feature.
[0049] In some possible embodiments, the preset three-branch generator architecture includes a local feature generator, a global feature generator, and a spatio-temporal feature generator, and the generation result set includes local feature generation results, global feature generation results, and spatio-temporal feature generation results; the anomaly detection module is specifically configured to:
[0050] Use the local feature generator to perform local feature extraction processing on the multi-modal fusion feature to obtain the local feature generation result; wherein, the local feature generator includes an encoding-decoding framework and a skip connection structure;
[0051] Use the global feature generator to perform global feature extraction processing on the multi-modal fusion feature to obtain the global feature generation result; wherein, the global feature generator includes a self-attention mechanism and a neural network;
[0052] Use the spatio-temporal feature generator to perform spatio-temporal dimensional feature extraction processing on the multi-modal fusion feature to obtain the spatio-temporal feature generation result; wherein, the spatio-temporal feature generator includes a three-dimensional convolutional neural network and a gated recurrent unit.
[0053] In some possible embodiments, the heterogeneous discriminator group includes a local anomaly detection discriminator, a global semantic discriminator, and a temporal consistency discriminator; specifically, the anomaly detection module is configured to:
[0054] Evaluate the anomalies of the texture and structure of the local regions in the local feature generation result based on the local anomaly detection discriminator to obtain a local anomaly probability;
[0055] Evaluate the semantic consistency between the global feature generation result and the multi-modal fusion feature based on the global semantic discriminator to obtain a global anomaly probability;
[0056] Evaluate the physical rationality of the motion trajectory in the spatio-temporal feature generation result based on the temporal consistency discriminator to obtain a temporal anomaly probability.
[0057] In some possible embodiments, the anomaly detection module is further configured to:
[0058] When at least two of the local anomaly probability, the global anomaly probability, and the temporal anomaly probability are greater than a first warning threshold, it is determined as a primary anomaly;
[0059] When it is determined as a primary anomaly, perform a weighted calculation on the local anomaly probability, the global anomaly probability, and the temporal anomaly probability according to a preset probability distribution weight to obtain a comprehensive anomaly probability; when the value of the comprehensive anomaly probability is greater than a second warning threshold, it is determined as an ultimate anomaly;
[0060] Determine the scene anomaly detection result regarding the scene to be detected based on the anomaly determination situation.
[0061] In some possible embodiments, the anomaly detection module is further configured to:
[0062] Construct a fully connected graph with the three discriminators as nodes and the anomaly detection results of the three discriminators as edge weights;
[0063] Calculate the graph entropy value of the fully connected graph;
[0064] When the graph entropy value is lower than a preset entropy value threshold, directly output the scene anomaly detection result;
[0065] When the graph entropy value is not lower than the preset entropy value threshold, mark the scene anomaly detection result as a complex anomaly, and generate a complex anomaly result report based on the anomaly monitoring results of each discriminator and the scene anomaly detection result, and send the complex anomaly result report to the manual analysis end.
[0066] An embodiment of the present disclosure provides a computer device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the computer device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the scenario anomaly detection method described in any of the above possible embodiments is executed.
[0067] An embodiment of the present disclosure provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, the scenario anomaly detection method described in any of the above possible embodiments is implemented.
[0068] In the scenario anomaly detection method, device, storage medium, and computer device provided by the embodiments of the present disclosure, by introducing the combination of multi-modal data fusion, a preset three-branch generator architecture, and a heterogeneous discriminator group, the accuracy and interpretability of scenario anomaly detection can be effectively improved. Specifically, first, through temporal alignment and spatial registration processing, the infrared image and the visible light image can be compared and analyzed in a unified coordinate framework, thereby eliminating the possible temporal and spatial differences between different modal data and ensuring the precise alignment of the data. Secondly, a convolutional neural network is used to extract features from the infrared and visible light modal data respectively, which can capture the fine-grained information in the two modal data. At the same time, through feature fusion processing, the advantages of the two modalities are combined, so that the final multi-modal fusion features are more comprehensive, which can effectively improve the sensitivity of detection and the reliability of recognition. Then, based on the preset three-branch generator architecture, the expression ability of the model is further enhanced, making the generated result set richer. Finally, through the heterogeneous discriminator group to perform anomaly detection on the generated results, the determination criteria from multiple angles can be comprehensively considered, further reducing the false detection rate and the missed detection rate.
[0069] In addition, the preset three-branch generator architecture and the heterogeneous discriminator group can achieve a more transparent feature extraction and discrimination process, effectively solving the problem of the interpretability of the detection results, and making the anomaly detection results easier to understand and trust.
[0070] To make the above objects, features, and advantages of the present disclosure more obvious and understandable, the following specific preferred embodiments are given in conjunction with the accompanying drawings and are described in detail as follows. Description of the Drawings
[0071] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings required to be cited in the embodiments will be briefly introduced below. The accompanying drawings here are incorporated into the specification and constitute a part of this specification. These accompanying drawings show the embodiments that conform to the present disclosure and are used together with the specification to illustrate the technical solutions of the present disclosure. It should be understood that the following accompanying drawings only show some embodiments of the present disclosure and should not be regarded as a limitation of the scope. For those of ordinary skill in the art, other related accompanying drawings can be obtained based on these accompanying drawings without creative efforts.
[0072] Figure 1 It shows a flowchart of a scenario anomaly detection method provided by an embodiment of the present disclosure;
[0073] Figure 2 It shows a schematic structural diagram of a multi-layer convolutional network provided by an embodiment of the present disclosure;
[0074] Figure 3 It shows a flowchart of a feature fusion processing method provided by an embodiment of the present disclosure;
[0075] Figure 4 It shows a flowchart of an anomaly detection method provided by an embodiment of the present disclosure;
[0076] Figure 5 It shows a flowchart of a heterogeneous graph spectrum analysis method provided by an embodiment of the present disclosure;
[0077] Figure 6 It shows a schematic structural diagram of a scenario anomaly detection device provided by an embodiment of the present disclosure;
[0078] Figure 7 It shows a schematic structural diagram of a computer device provided by an embodiment of the present disclosure. Detailed implementation manners
[0079] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only some of the embodiments of the present disclosure, rather than all the embodiments. Usually, the components of the embodiments of the present disclosure described and shown in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure provided in the accompanying drawings is not intended to limit the scope of the present disclosure to be protected, but only represents the selected embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present disclosure.
[0080] It should be noted that like reference numerals and letters refer to like items in the following figures, and thus, once an item is defined in one figure, it need not be further defined and explained in subsequent figures.
[0081] As used herein, the term "and / or" merely describes an association relationship, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, or B exists alone. In addition, the term "at least one" as used herein means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C may represent including any one or more elements selected from the set consisting of A, B, and C.
[0082] For ease of understanding of this embodiment, the execution subject of the scenario anomaly detection method provided by the embodiments of the present disclosure will be introduced in detail first. The execution subject of the scenario anomaly detection method provided by the embodiments of the present disclosure is a computer device. The computer device may be a terminal device or a server. Among them, the terminal device may also be a mobile device, a user terminal, a terminal, a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The server may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud storage, big data, and artificial intelligence platforms. Optionally, this method may also be applied to an implementation environment composed of a computer device and a server.
[0083] The following will describe in detail the scenario anomaly detection method provided by the embodiments of the present application with reference to the accompanying drawings. Refer to Figure 1 As shown, it is a flowchart of a scenario anomaly detection method provided by the embodiments of the present disclosure. The method includes the following S101 to S104:
[0084] S101, obtain an infrared image and a visible light image corresponding to the to-be-detected scenario, and perform temporal alignment processing and spatial registration processing on the infrared image and the visible light image respectively to obtain infrared modality data and visible light modality data.
[0085] It can be understood that the infrared image is formed based on the thermal radiation of an object, can capture the temperature distribution information on the surface of the object, is not restricted by the lighting conditions, and can clearly image even in a dark environment. It is suitable for detecting abnormal situations with obvious temperature differences, such as detecting equipment overheating faults, personnel activities at night, etc. The visible light image presents the reflection characteristics of an object in the visible light band, has rich colors and clear details, and can intuitively display the appearance characteristics of the scenario, such as the shape, color, texture, etc. of the object, which helps to identify some scenarios with abnormal appearances, such as building damage, illegal placement of objects, etc.
[0086] However, due to the possible differences in the shooting times of the infrared camera and the visible light camera, as well as their differences in spatial position and shooting angle, the directly obtained images may not be aligned in time and space. Therefore, it is necessary to perform temporal alignment processing and spatial registration processing on the obtained infrared images and visible light images.
[0087] On the one hand, the temporal alignment processing mainly ensures that the infrared image and the visible light image are collected at the same time point or at a similar time point to eliminate the image offset problem caused by the shooting time difference. For example, when monitoring a traffic intersection, if the time interval between the collection of the infrared image and the visible light image is too long, situations such as vehicle position movement, pedestrians entering or leaving the scene may occur, thus affecting the accurate judgment of the scene state. During the temporal alignment processing of the infrared image and the visible light image, the collection time information of either the infrared image or the visible light image can be used as a reference to match the time stamps of the other image, and a set of images that are closest in time can be selected for analysis. Here, by calculating the difference in the collection times of the two images, interpolation methods (such as linear interpolation, spline interpolation) can be used to adjust one of the images in time, so that the infrared image and the visible light image are matched on the time axis.
[0088] On the other hand, the spatial registration processing is to align the infrared image and the visible light image in the spatial coordinate system so that they can accurately correspond to the same physical space position. This usually involves geometric transformations of the image, such as translation, rotation, scaling, etc.
[0089] In some possible embodiments, the image feature point information of the infrared image and the visible light image can be extracted separately first, and then through the feature point matching algorithm, the same feature points (such as the corners of buildings, the intersections of roads, etc.) in the infrared image and the visible light image can be found. Then, according to the corresponding relationship of these feature points, the transformation matrix can be calculated, and one of the images can be geometrically transformed (such as homography matrix, affine transformation or perspective transformation, etc.) to align it with the other image in space. Among them, feature extraction methods such as SIFT, SURF or ORB can be selected to capture the significant features in the image. Feature point matching algorithms such as brute-force matching or fast nearest neighbor search matching can be selected to find the corresponding points between the two images.
[0090] Here, in order to improve the matching accuracy, the Random Sample Consensus (RANSAC) algorithm can also be used to eliminate the incorrect matching points to ensure that the obtained geometric transformation matrix is more accurate.
[0091] In this way, the dual-modal data (i.e., infrared modal data and visible light modal data) after time series alignment and spatial registration can be made consistent in time and space, thereby improving the data quality.
[0092] In some other embodiments, after obtaining the infrared image and the visible light image, modal-opposite normalization can also be performed on them respectively. Specifically:
[0093] For the visible light image, channel-level Z-Score normalization is adopted, and the formula is:
[0094] ;
[0095] where , respectively represent the mean and standard deviation of the visible light image;
[0096] For the infrared image, it can be linearly mapped according to the sensor range and noise can be removed (such as non-uniformity correction).
[0097] S102, using a convolutional network to extract features from the infrared modal data to obtain infrared features corresponding to the infrared modal data; and, using a convolutional network to extract features from the visible light modal data to obtain visible light features corresponding to the visible light modal data.
[0098] It can be understood that after obtaining the infrared modal data and the visible light modal data, a convolutional network can be used to extract features from them respectively. Among them, the convolutional network is a deep learning model that can automatically learn features in images through structures such as convolutional layers and pooling layers.
[0099] Specifically, for the infrared modal data, the convolutional network can capture the temperature distribution pattern and thermal features therein. For example, when detecting industrial equipment, the normal operating state and abnormal states (such as overheating, local damage, etc.) of the equipment will show different temperature distributions in the infrared image. The convolutional network can learn a large number of infrared images of normal and abnormal equipment and extract features that can distinguish these two states, such as temperature gradient changes, hot spot area distributions, etc. These features can be called infrared features, which reflect the essential attributes of objects in the infrared band and are helpful for detecting temperature-related abnormalities.
[0100] Similarly, for visible light modality data, a convolutional network can extract the appearance features of an object. For example, when monitoring a warehouse scene, there are obvious appearance differences between the normal placement of goods and abnormal accumulation of goods (such as goods tipping over, chaotic placement, etc.) in visible light images. The convolutional network can learn these appearance features, such as the outline of the object, color distribution, texture details, etc., to obtain visible light features corresponding to the visible light modality data, so as to intuitively reflect the appearance state of the scene.
[0101] Exemplarily, in the present disclosure, a multi-layer convolutional network is adopted to implement the feature extraction tasks for infrared modality data and visible light modality data. Here, the infrared features include low-order infrared feature outputs and high-order infrared feature outputs. Similarly, the visible light features include low-order visible light feature outputs and high-order visible light feature outputs. Referring to Figure 2 as shown, the multi-layer convolutional network in the present disclosure includes a shallow feature extraction module and a deep feature extraction module. Specifically, the shallow feature extraction module is used to capture low-order features of both visible light and infrared modality data. This module is constructed based on the Conv-BN-ReLU basic unit, avoiding the use of pooling operations and using strided convolution for downsampling to maximize the retention of spatial detail information.
[0102] Among them, for infrared modality data, the initial convolutional layer is used to capture the characteristics of the thermal radiation distribution, highlighting the feature expression formed by temperature differences; at the same time, for visible light modality data, rich texture features, including local visual patterns such as edges and corners, can be extracted through 3×3 small convolutional kernels. The feature extraction processes of the two modalities maintain the same network architecture, and finally feature tensors with consistent dimensions are output:
[0103] The low-order infrared feature output of the infrared modality data can be expressed as:
[0104] ;
[0105] where represents the low-order infrared feature output of the infrared modality data; R represents the set of real numbers; represents that after downsampling by strided convolution, the height H and width W of the feature map are 1 / 2 of the original input image respectively; 64 represents the number of channels of the feature map.
[0106] The low-order visible light feature output of the visible light modality data can be expressed as:
[0107] ;
[0108] where represents the low-order visible light feature output of the visible light modality data; the height H and width W of the feature map are 1 / 2 of the original input image respectively; 64 represents the number of channels of the feature map.
[0109] Specifically, in the deep feature extraction stage, the multi-layer convolutional network uses a deep feature extraction module to mine high-order semantic features from infrared modal data and visible light modal data. This module introduces dilated convolution on the basis of standard convolution to expand the receptive field, enabling it to capture context information in a larger range. At the same time, a self-attention mechanism is embedded to dynamically adjust the feature channel weights and enhance the expression ability of key features.
[0110] Exemplarily, for infrared data, after being processed by the deep feature extraction module, abnormal sensitive features can be highlighted, especially having a significant response to mutation regions in the temperature distribution (such as device hot spots, living targets, etc.). For visible light modal data, the deep feature extraction module extracts rich semantic context information through multi-level non-linear transformation, including object categories (such as vehicles, pedestrians) and scene semantic labels (such as roads, buildings). These features have stronger discriminability and can support higher-level visual understanding tasks.
[0111] Among them, the same downsampling strategy is adopted in the feature extraction process of both modalities, and the final output is a 512-dimensional feature tensor with a spatial resolution of 1 / 8 of the input (H / 8×W / 8). That is, the high-order infrared feature output of the infrared modal data can be expressed as:
[0112] ;
[0113] Among them, represents the high-order infrared feature output of the infrared modal data.
[0114] The high-order visible light feature output of the visible light modal data can be expressed as:
[0115] ;
[0116] Among them, represents the high-order visible light feature output of the visible light modal data.
[0117] In this disclosure, a multi-layer convolutional network is introduced to implement the feature extraction task for infrared modal data and visible light modal data. The shallow feature extraction module ensures the spatial alignment of features of each modality and lays a compatibility foundation for subsequent feature fusion; the deep feature extraction module ensures the semantic richness of deep features, and also provides convenience for subsequent multi-modal fusion through a unified feature dimension. At the same time, the introduction of dilated convolution and the attention mechanism further improves the model's ability to model global context and key regions, enhancing its robustness and accuracy in complex scenarios.
[0118] S103, perform feature fusion processing on the infrared features and the visible light features to obtain multi-modal fusion features.
[0119] It is understandable that after obtaining the infrared features and visible light features, in order to make full use of the advantages of the two types of modal data, they can be subjected to feature fusion processing to obtain multi-modal fusion features. Among them, the purpose of feature fusion is to integrate the feature information of different modalities to obtain a more comprehensive and accurate scene representation.
[0120] Exemplarily, common feature fusion methods include early fusion, mid-term fusion, and late fusion, etc. Early fusion is to splice the infrared image and the visible light image in channels before feature extraction, and then input them into the same convolutional network for feature extraction. Mid-term fusion is to fuse the infrared features and visible light features in the middle layer of the network during the feature extraction process. For example, after a certain convolutional layer of the convolutional network, the features of the two modalities are spliced or added, and then subsequent convolutional operations are performed. This method can better retain the information of different modal features and enable the network to learn the association between them during the feature extraction process. Late fusion is to input the infrared features and visible light features into different classifiers respectively after feature extraction is completed, and then fuse the output results of the classifiers.
[0121] In some possible embodiments, in order to fully exploit the complementary information between the infrared image and the visible light image, two different types of modal data, and achieve a more accurate scene feature representation, the present disclosure proposes a feature fusion processing method, as shown in Figure 3 it may include the following S301 to S303:
[0122] S301, calculate the feature cross-correlation between the infrared features and the visible light features, and construct a cross-modal attention weight matrix.
[0123] Here, a cross-modal attention weight matrix can be constructed at each level l of the feature fusion network ( l represents the level number in the network; represents the cross-modal attention weight matrix of the l th layer; C represents the number of feature channels); to quantify the association degree between the infrared features and the visible light features, and highlight the association strength between different modal channels (i.e., the cross-correlation between the visible light features and the infrared features). Each element of this cross-modal attention weight matrix reflects the association strength between different modal channels.
[0124] S302, determine the initial multi-modal fusion features based on the cross-modal attention weight matrix, the infrared features, and the visible light features.
[0125] Specifically, after obtaining the cross-modal attention weight matrix, the infrared features and visible light features can be adaptively weighted and combined using the cross-modal attention weight matrix, and the fusion process is achieved through matrix multiplication:
[0126] ;
[0127] Among them, can be expressed as the initial multi-modal fusion feature of the l th layer; its dimension is the same as that of and ; and are the weight matrices for visible light features and infrared features respectively, and can be obtained by appropriately splitting or transforming the cross-modal attention weight matrix . and are the visible light features and infrared features of the l th layer respectively.
[0128] In this way, through matrix multiplication, the weight matrix can be multiplied by the corresponding feature matrix to achieve weighting of different modal features. The weighted features are then added together, enabling the modal features to be adaptively weighted and combined according to their importance.
[0129] S303, perform recalibration processing on the initial multi-modal fusion feature, and perform feature enhancement processing on the recalibration result based on a multi-scale spatial enhancement strategy to obtain the multi-modal fusion feature.
[0130] Here, in order to further optimize the initial multi-modal fusion feature and solve problems such as feature redundancy and imbalance of feature information at different scales, the present disclosure proposes to perform recalibration processing on the initial multi-modal fusion feature and perform feature enhancement processing based on a multi-scale spatial enhancement strategy.
[0131] Specifically, in order to highlight important channels and suppress redundant channels in the initial multi-modal fusion feature, the fused initial multi-modal fusion feature can be first subjected to recalibration processing. Channel-level statistical features are obtained through global average pooling, where the feature values at all spatial positions on each channel are averaged, and then the channel description vector is input into two fully connected layers for non-linear transformation. Channel attention vectors are generated by the two fully connected layers, which can recalibrate the feature channels. In this way, after channel attention recalibration, the importance of different channels in the fusion feature is readjusted, important channel features are enhanced, and redundant channel features are suppressed, thereby improving the quality and discriminability of the features.
[0132] Furthermore, in order to capture feature information at different spatial scales and enhance the perception ability of the fused features for targets of different sizes, dilated convolutions with different dilation rates (2 / 4 / 6) can be used to process the feature maps in parallel. Dilated convolution expands the receptive field by inserting holes in the convolution kernel, enabling it to obtain more extensive context information without increasing the number of parameters and computational complexity.
[0133] Specifically, the recalibrated fused feature matrix can be input into three dilated convolution branches with different dilation rates respectively. Each branch uses a 3×3 dilated convolution kernel with dilation rates of 2, 4, and 6 respectively. After the dilated convolution operation, three spatial feature maps of different scales are obtained, which focus on different-sized targets and scene regions respectively. Then, these three spatial feature maps of different scales are concatenated, and finally, a 1×1 convolutional layer is used to reduce the dimension of the concatenated feature map, fusing multi-scale context information while maintaining the spatial resolution (H×W) of the feature map. The finally output multi-modal fused features : Here, the multi-modal fused features contain both the detailed texture and semantic information of the visible light features and integrate the temperature distribution features of the infrared features, providing a unified multi-modal representation for the following steps.
[0134] It can be understood that the above fusion process is repeatedly executed at multiple levels of the feature fusion network, forming a progressive feature fusion architecture from local to global. At the low level, feature fusion mainly focuses on local texture and edge information, and associates and fuses the local features in the infrared and visible light images through a cross-modal attention mechanism. At the high level, feature fusion considers more semantic information and global structure, captures different-sized targets and scene regions through a multi-scale spatial enhancement strategy, and forms a more discriminative global feature representation.
[0135] Through feature fusion processing, the obtained multi-modal fused features integrate the temperature information of the infrared modality and the appearance information of the visible light modality, and can more comprehensively describe the state of the scene to be detected. For example, in a factory production scene, the multi-modal fused features can simultaneously reflect the temperature abnormality of the equipment (through infrared features) and the appearance damage of the equipment (through visible light features), thus improving the detection ability of scene abnormalities.
[0136] S104. Determine a set of generation results based on a preset three-branch generator architecture and the multi-modal fused features; and perform anomaly detection on the set of generation results through a heterogeneous discriminator group to obtain a scene anomaly detection result for the scene to be detected.
[0137] Specifically, based on a preset three-branch generator architecture and multimodal fusion features, a set of generation results can be determined. The three-branch generator architecture is a special structure of a Generative Adversarial Network (GAN), which usually includes three different generator branches, and each branch is responsible for generating different types or different scales of features or images. Here, the preset three-branch generator architecture may include a local feature generator, a global feature generator, and a spatio-temporal feature generator.
[0138] Exemplarily, when determining the set of generation results based on the preset three-branch generator architecture and multimodal fusion features, the following (a) to (c) may be included:
[0139] (a) Using the local feature generator to perform local feature extraction processing on the multimodal fusion features to obtain the local feature generation result; wherein, the local feature generator includes an encoder-decoder framework and a skip connection structure;
[0140] (b) Using the global feature generator to perform global feature extraction processing on the multimodal fusion features to obtain the global feature generation result; wherein, the global feature generator includes a self-attention mechanism and a neural network;
[0141] (c) Using the spatio-temporal feature generator to perform spatio-temporal dimension feature extraction processing on the multimodal fusion features to obtain the spatio-temporal feature generation result; wherein, the spatio-temporal feature generator includes a three-dimensional convolutional neural network and a gated recurrent unit.
[0142] Exemplarily, the local feature generator can adopt a U-Net structure. By cooperating the encoder-decoder framework with skip connections to retain spatial high-frequency details, it focuses on high-resolution image reconstruction tasks. The output local feature generation result maintains the same spatial size as the input multimodal fusion features, and focuses on restoring local detail features such as edges and textures to obtain the local feature generation result. The global feature generator can be designed based on the Transformer architecture. Using the multi-head self-attention mechanism to model long-range dependencies, and processing global context information through layer normalization and feed-forward networks, the output global feature generation result is a pixel-level semantic segmentation mask, realizing the semantic consistency expression of the scene. The spatio-temporal feature generator can adopt a structure combining a three-dimensional convolutional neural network and a Gated Recurrent Unit (GRU). The 3D convolutional kernel extracts features in the spatio-temporal dimension, and the GRU module models the inter-frame temporal relationship. The output spatio-temporal feature generation result is a temporal optical flow prediction map, which is used to capture the motion patterns in dynamic scenes. The three generators use the multimodal fusion features as a shared input and generate output results with different attributes through parallel forward propagation.
[0143] In some other embodiments, the local feature generator can also adopt other different encoding-decoding frameworks and skip connection structures; the global feature generator can be based on other self-attention mechanisms and neural networks; similarly, the spatio-temporal feature generator can combine different 3D convolutional neural networks and recurrent unit structures to achieve the extraction of multi-modal features. The above description is only an example and does not limit the implementation details.
[0144] Exemplarily, in the inference stage, the three-branch generator architecture can support three working modes: any generator can be individually invoked to obtain a specific output, or multi-task joint inference can be achieved through a feature sharing mechanism, where the shallow features of U-Net can interact with the attention map of Transformer, and the spatio-temporal features extracted by 3D convolution can assist in processing dynamic objects in the segmentation task. By explicitly decoupling the three major feature dimensions of spatial details, semantic information, and temporal dynamics, this architecture realizes feature-level collaboration while maintaining the independence of each task, providing a flexible and efficient solution for multi-modal visual understanding.
[0145] In some possible embodiments, a multi-task loss function can be adopted for joint optimization during the training process of the three generators: applying pixel-level L1 loss and perceptual loss to the local feature generator to ensure the reconstruction quality; using cross-entropy loss and Dice loss to optimize the segmentation accuracy for the global feature generator; using the endpoint error loss specific to optical flow estimation for the spatio-temporal feature generator.
[0146] Furthermore, to promote feature decoupling, the present disclosure also proposes to introduce modality-specific normalization layers in the design of each generator, separating the feature representations required for different generation tasks at the feature level, and at the same time enabling each generator to focus on its own task domain through an adversarial training strategy.
[0147] It can be understood that after the generation result set (i.e., the local feature generation result, the global feature generation result, and the spatio-temporal feature generation result) is obtained, an anomaly detection can be performed on the generation result set through a heterogeneous discriminator group. Here, the heterogeneous discriminator group is composed of multiple discriminators of different types, each discriminator having a different structure and function, and being able to evaluate the generation results from different perspectives. Specifically, the heterogeneous discriminator group can include a local anomaly detection discriminator, a global semantic discriminator, and a temporal consistency discriminator, which perform anomaly detection on the results generated by the above-mentioned generators respectively.
[0148] Among them, during the anomaly detection process, the heterogeneous discriminator group will score each generation result in the generation result set. If a certain generation result is judged as abnormal by the discriminator, it can be considered that there is an abnormal situation in the scene to be detected. Finally, according to the overall detection results of the heterogeneous discriminator group, the scene anomaly detection result of the scene to be detected can be obtained.
[0149] Specifically, referring to Figure 4 As shown, when performing anomaly detection on the generated result set through the heterogeneous discriminator group, the following S401 - S403 may be included:
[0150] S401, evaluate the anomaly of the texture and structure of the local area in the local feature generation result based on the local anomaly detection discriminator to obtain the local anomaly probability.
[0151] Specifically, after obtaining the local feature generation result generated by the above - mentioned local feature generator, the local anomaly detection discriminator can be used to evaluate the anomaly of the texture and structure of the local area therein to obtain the local anomaly probability. Here, the local anomaly detection discriminator can adopt the PatchGAN (Markov discriminator) architecture. By dividing the input local feature generation result into local blocks of 70×70 pixels through a convolutional network, the anomaly probability of texture details and structural rationality (i.e., the local anomaly probability) is evaluated block by block. This can effectively capture the subtle anomalies in the local area of the image, such as structural distortion or texture distortion. At the same time, the Markov property of the PatchGAN architecture makes the discriminator pay more attention to local information during evaluation, avoiding the interference of global information on local anomaly judgment and improving the sensitivity and detection accuracy of local anomalies.
[0152] S402, evaluate the semantic consistency between the global feature generation result and the multi - modal fusion feature based on the global semantic discriminator to obtain the global anomaly probability.
[0153] Specifically, after obtaining the global feature generation result generated by the above - mentioned global feature generator, the global semantic discriminator can be used to evaluate the semantic consistency between the global feature generation result and the multi - modal fusion feature to obtain the global anomaly probability. Here, the global semantic discriminator can be constructed based on ResNet - 18. Through multi - level feature extraction and global pooling operations, the semantic consistency between the global feature generation result and the multi - modal fusion feature is evaluated from the perspective of the overall scene, and the global anomaly probability is output. This discriminator is particularly good at identifying anomalies at the macroscopic level such as incorrect object categories and unreasonable scene layouts. Among them, the multi - modal fusion feature usually integrates information from different modalities (such as images, texts, sensor data, etc.), and these information describe the features and attributes of the target object from different angles.
[0154] Here, the global semantic discriminator determines whether the semantic representations of elements such as roads, vehicles, and pedestrians in the generated global feature generation results are consistent with the semantic information contained in the multi-modal fusion features. If the road direction in the generated result does not match the map information, or the position and type of the vehicle do not match, it indicates poor semantic consistency. The global semantic discriminator will give a relatively high global anomaly probability, indicating that there is an anomaly in the generated result at the global semantic level.
[0155] S403, based on the temporal consistency discriminator, evaluate the physical rationality of the motion trajectories in the spatio-temporal feature generation results to obtain the temporal anomaly probability.
[0156] Specifically, for the spatio-temporal feature generation results, the main task of the temporal consistency discriminator is to evaluate the physical rationality of the motion trajectories therein, so as to obtain the temporal anomaly probability. In many tasks involving dynamic scenarios, such as video generation, motion capture and reconstruction, etc., the spatio-temporal feature generation results contain the change information of objects in time and space, that is, motion trajectories. Here, the temporal consistency discriminator can adopt a 3D convolutional network structure, extract features in the spatio-temporal dimension through 3D convolutional kernels, pay attention to the position and pose of objects in each frame of the image, combine with the LSTM module to model the time dependence, analyze and evaluate whether the changes in the position and pose of objects in the video sequence conform to physical laws in the time series, and then output the temporal anomaly probability to detect abnormal behaviors that violate the natural motion laws.
[0157] For example, in the task of generating sports event videos, the motion trajectories of athletes should follow the principles of human kinematics. If there are jump-like position changes in the running actions of athletes in the generated results that do not conform to human motion laws between consecutive frames, or the motion speed suddenly changes abnormally, accelerating or decelerating beyond the normal human motion speed range, the temporal consistency discriminator can detect these abnormal situations by analyzing the continuity and physical rationality of the motion trajectories in the time dimension and calculate the corresponding temporal anomaly probability. In addition, the temporal consistency discriminator can also combine physical models, such as Newton's laws of motion, etc., to conduct a more accurate physical rationality evaluation of the motion trajectories, further improving the detection ability for temporal anomalies.
[0158] The heterogeneous discriminator group proposed in the embodiments of the present disclosure can comprehensively and meticulously detect anomalies in the generated result set from three different dimensions: local texture structure, global semantic consistency, and temporal physical rationality.
[0159] In some possible embodiments, after obtaining the anomaly probabilities generated by each discriminator, in the anomaly score calculation stage, the output results of the three discriminators can be comprehensively evaluated, including: first, performing Sigmoid normalization on the anomaly probability values of each discriminator to convert the discrimination results into anomaly probability values in the range of [0, 1]; then calculating the comprehensive anomaly score through weighted summation, where the weight coefficients can be dynamically adjusted according to the importance of each modality.
[0160] In some possible embodiments, to improve the robustness of the evaluation, an anomaly score calibration mechanism is also designed in the present disclosure: constructing a reference distribution using the discriminator outputs of historical normal samples, and measuring the degree of deviation of the current sample from the normal distribution by calculating the Mahalanobis distance; at the same time, introducing an uncertainty estimation module to obtain the score variance through multiple inferences of Monte Carlo Dropout, and specially marking the evaluation results with low confidence. This multi-discriminator system can not only give a comprehensive anomaly score, but also locate the anomaly type through the special outputs of each discriminator. For example, the high-response region of the local discriminator indicates the texture anomaly position, and the peak frame of the temporal discriminator marks the moment of motion anomaly, providing multi-dimensional decision-making basis for subsequent anomaly analysis and processing.
[0161] Exemplarily, in order to more accurately, comprehensively and hierarchically evaluate the anomaly degree of the generated result set in the scene to be detected, so as to timely discover potential problems and take targeted measures, after obtaining the local anomaly probability, global anomaly probability and temporal anomaly probability by performing anomaly detection on the generated result set through a heterogeneous discriminator group, the following (1) to (3) may further be included:
[0162] (1) When at least two of the local anomaly probability, the global anomaly probability and the temporal anomaly probability are greater than the first warning threshold, it is determined as a primary anomaly;
[0163] (2) When it is determined as a primary anomaly, perform weighted calculation on the local anomaly probability, the global anomaly probability and the temporal anomaly probability according to the preset probability allocation weights to obtain the comprehensive anomaly probability; when the value of the comprehensive anomaly probability is greater than the second warning threshold, it is determined as an ultimate anomaly;
[0164] (3) Determine the scene anomaly detection result regarding the scene to be detected based on the anomaly determination situation.
[0165] It can be understood that in actual application scenarios, the abnormal probability in a single dimension may be interfered by various factors, and there is a certain risk of misjudgment. For example, in video generation tasks, the generation results of local features may cause a temporary increase in the local abnormal probability due to local noise or light changes, but this does not necessarily mean that there are serious problems in the entire generation result. Therefore, by setting joint judgment conditions for multiple abnormal probabilities, that is, when at least two of the local abnormal probability, global abnormal probability, and temporal abnormal probability are greater than the first warning threshold, it is determined as a primary anomaly, which can effectively improve the accuracy and reliability of anomaly detection.
[0166] Among them, the setting of the first warning threshold can comprehensively consider the characteristics of the specific task, data distribution, and actual application requirements. For medical image generation tasks with extremely high precision requirements, the first warning threshold can be set relatively low to ensure that any possible abnormal situations can be captured in a timely manner; while for some monitoring video analysis tasks with high real-time requirements and allowing a certain degree of tolerance, the first warning threshold can be appropriately increased to avoid frequent warnings due to minor abnormal fluctuations. When the condition that at least two abnormal probabilities are greater than the first warning threshold is met, it is determined as a primary anomaly.
[0167] Furthermore, the primary anomaly determination is only a preliminary hint that there may be anomalies in the set of generation results. To more accurately evaluate the severity of the anomalies, the abnormal information in the three dimensions of local, global, and temporal can be comprehensively considered. The importance of abnormal probabilities in different dimensions may vary in reflecting the quality of the generation results. Therefore, it is necessary to preset probability distribution weights to perform weighted calculations on the local abnormal probability, global abnormal probability, and temporal abnormal probability to obtain a comprehensive abnormal probability.
[0168] In the embodiments of the present disclosure, the preset probability distribution weights are dynamically allocated based on the performance of each discriminator in the receiver operating characteristic curve on the validation set. In some other embodiments, they can also be set in advance based on expert experience, data analysis, or experiments, which are not specifically limited herein.
[0169] It can be understood that the second warning threshold is an indicator for determining an ultimate anomaly, and it is determined as an ultimate anomaly only when the abnormal situation reaches a certain level. When the value of the comprehensive abnormal probability is greater than the second warning threshold, it is determined as an ultimate anomaly, indicating that there are relatively serious abnormal problems in the set of generation results, and corresponding measures need to be taken immediately, such as manual intervention.
[0170] Specifically, based on the determination results of the above-mentioned primary anomaly and ultimate anomaly, the scene anomaly detection result for the scene to be detected can be determined. If it is not determined as a primary anomaly, it indicates that the generated result set does not show obvious anomalies in the three dimensions of local, global, and temporal, and the scene anomaly detection result is normal. If it is determined as a primary anomaly but not as an ultimate anomaly, it indicates that there is a certain degree of anomaly risk in the scene to be detected, but it has not reached a serious level. At this time, a warning message can be generated to prompt relevant personnel to conduct further manual inspections or analyses on the scene to be detected. At the same time, relevant anomaly information can also be recorded for subsequent optimization and improvement of the generation model. If it is determined as an ultimate anomaly, it means that there are serious anomaly problems in the scene to be detected, and the scene anomaly detection result is abnormal. At this time, emergency measures should be taken immediately, such as sending a detailed anomaly report to relevant personnel, including information such as anomaly type, anomaly degree, and possible impact range, for timely troubleshooting and repair.
[0171] It can be understood that to further improve the decision-making quality, after obtaining the scene anomaly detection result, this disclosure also introduces a heterogeneous graph spectrum analysis method. Referring to Figure 5 as shown, it can include the following steps S501 to S504:
[0172] S501, construct a fully connected graph with three discriminators as nodes and the anomaly detection results of the three discriminators as edge weights.
[0173] It can be understood that when constructing the fully connected graph, these three discriminators are abstracted as nodes in the graph. Each node represents an independent anomaly detection perspective or analysis dimension, and the anomaly detection results of the three discriminators serve as the weights of the edges connecting these nodes. Specifically, the weight of the edge can be determined according to the value, confidence level, or other quantization indicators of the anomaly detection result. For example, if a discriminator detects an anomaly with a high confidence level, the weight of the edge connected to it is relatively large; conversely, if the detection result is close to normal, the weight is relatively small. In this way, the fully connected graph associates the detection results of different discriminators with an intuitive graphical structure, clearly showing the mutual influence and potential correlation between the detection results of each dimension.
[0174] S502, calculate the graph spectrum entropy value of the fully connected graph.
[0175] Here, the graph spectrum entropy value is an important indicator to measure the information complexity and uncertainty in the graph. It can reflect the uniformity of the edge weight distribution in the fully connected graph and the complexity of the relationship between nodes. In the anomaly detection scenario, the level of the graph spectrum entropy value is closely related to the complexity of the scene anomaly.
[0176] Specifically, the process of calculating the graph entropy value involves the statistics and analysis of the edge weight distribution in the fully connected graph. First, the weight situation of the edges connecting each node to other nodes is statistically analyzed to obtain the frequency distribution of the weights; then, according to the calculation formula of information entropy, the graph entropy value is calculated in combination with the weight frequency distribution. The basic idea of information entropy is that when the weight distribution is more uniform, the uncertainty of information is greater, and the graph entropy value is higher; on the contrary, if the weight distribution is concentrated on a few values, it indicates that there are obvious rules or dominant relationships in the graph, and the graph entropy value is lower. By calculating the graph entropy value, the complexity of the current anomaly detection result can be quantitatively evaluated.
[0177] S503. In the case where the graph entropy value is lower than the preset entropy value threshold, directly output the scene anomaly detection result.
[0178] It can be understood that when the calculated graph entropy value is lower than the preset entropy value threshold, it means that the distribution of edge weights in the fully connected graph is relatively concentrated, and there is a high similarity or consistency between the anomaly detection results of different discriminators. This indicates that the current anomaly situation is relatively clear, the judgment directions of each discriminator for the anomaly are basically the same, and there are no obvious contradictions or complex correlations. In this way, the scene anomaly detection result can be directly output, and the existence of the anomaly can be determined without further in-depth analysis.
[0179] S504. In the case where the graph entropy value is not lower than the preset entropy value threshold, mark the scene anomaly detection result as a complex anomaly, and based on the anomaly monitoring results of each discriminator and the scene anomaly detection result, generate a complex anomaly result report and send the complex anomaly result report to the manual analysis terminal.
[0180] Specifically, if the graph entropy value is not lower than the preset entropy value threshold, it means that the distribution of edge weights in the fully connected graph is relatively dispersed, and there are large differences or contradictions between the anomaly detection results of different discriminators. This implies that the current anomaly situation is relatively complex, possibly involving the interweaving of multiple dimensions of anomaly factors, or there are some potential problems that are difficult to be automatically identified by the existing discriminators.
[0181] At this time, the scene anomaly detection result can be marked as a complex anomaly. At the same time, based on the anomaly monitoring results of each discriminator and the scene anomaly detection result, generate a complex anomaly result report, which should detail information such as the detection results of each discriminator, the key indicators during the detection process, and the difference analysis between different results. For example, the report can list the specific anomaly types detected by each discriminator, the regions or time points where the anomalies occur, the quantitative indicators of the anomaly degree, etc., and make a preliminary speculation on the reasons for the inconsistent results of each discriminator.
[0182] Furthermore, send the complex anomaly result report to the manual analysis terminal to further analyze the deep - seated reasons for the anomaly by leveraging the experience and knowledge of relevant technical personnel. Professionals at the manual analysis terminal can conduct in - depth analysis and diagnosis of complex anomalies based on the detailed information in the report, combined with their own domain knowledge and experience. They can find the root cause of the anomaly and propose targeted solutions through different analysis methods such as viewing the original data and adjusting detection parameters.
[0183] In some possible embodiments, when the scene anomaly detection result shows a complex anomaly, the current detection state can also be frozen, all intermediate features can be saved, then the differences in the attention areas of each discriminator can be highlighted in the result visualization window, and finally a review report containing a probability distribution graph and a confidence explanation is generated and submitted for manual analysis.
[0184] The scene anomaly detection method, device, storage medium, and computer equipment provided in the embodiments of the present disclosure can effectively improve the accuracy and interpretability of scene anomaly detection by introducing the combination of multi - modal data fusion, a preset three - branch generator architecture, and a heterogeneous discriminator group. Specifically, first, through temporal alignment and spatial registration processing, the infrared image and the visible - light image can be compared and analyzed in a unified coordinate framework, thereby eliminating the possible temporal and spatial differences between different - modal data and ensuring the precise alignment of the data. Secondly, a convolutional neural network is used to extract features from the infrared and visible - light modal data respectively, which can capture the fine - grained information in both modal data. At the same time, through feature fusion processing, combining the advantages of the two modalities can make the final multi - modal fusion features more comprehensive, effectively improving the detection sensitivity and recognition reliability. Then, based on the preset three - branch generator architecture, the expressive ability of the model is further enhanced, making the generated result set richer. Finally, through the heterogeneous discriminator group to perform anomaly detection on the generated results, the judgment criteria from multiple angles can be comprehensively considered, further reducing the false - detection rate and the missed - detection rate.
[0185] In addition, the preset three - branch generator architecture and the heterogeneous discriminator group can achieve a more transparent feature extraction and discrimination process, effectively solving the problem of the interpretability of detection results, and making the anomaly detection results easier to understand and trust.
[0186] Those skilled in the art can understand that in the above - mentioned method of the specific implementation manner, the writing order of each step does not mean a strict execution order that constitutes any limitation to the implementation process, and the specific execution order of each step should be determined by its function and possible internal logic.
[0187] Based on the same inventive concept, an embodiment of the present disclosure further provides a scene anomaly detection device corresponding to the scene anomaly detection method. Since the principle of problem-solving of the device in the embodiment of the present disclosure is similar to that of the above-mentioned scene anomaly detection method in the embodiment of the present disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0188] Referring to Figure 6 As shown, it is a schematic diagram of a scene anomaly detection device 600 provided by an embodiment of the present disclosure. The device includes:
[0189] An image processing module 601, configured to obtain an infrared image and a visible light image corresponding to a scene to be detected, and perform temporal alignment processing and spatial registration processing on the infrared image and the visible light image respectively to obtain infrared modal data and visible light modal data;
[0190] A feature extraction module 602, configured to extract features from the infrared modal data by using a convolutional network to obtain infrared features corresponding to the infrared modal data; and extract features from the visible light modal data by using a convolutional network to obtain visible light features corresponding to the visible light modal data;
[0191] A feature fusion module 603, configured to perform feature fusion processing on the infrared features and the visible light features to obtain multi-modal fusion features;
[0192] An anomaly detection module 604, configured to determine a set of generation results based on a preset three-branch generator architecture and the multi-modal fusion features; and perform anomaly detection on the set of generation results through a heterogeneous discriminator group to obtain a scene anomaly detection result of the scene to be detected.
[0193] In some possible embodiments, the image processing module 601 is specifically configured to:
[0194] Taking the acquisition time information of any one of the infrared image and the visible light image as a reference, perform timestamp matching on the other image to complete the temporal alignment processing of the infrared image and the visible light image;
[0195] Extract the image feature point information of the infrared image and the visible light image respectively, and determine the geometric transformation relationship between the two images based on the extraction results and the matching algorithm to determine the transformation matrix;
[0196] Perform geometric transformation on any one of the infrared image and the visible light image according to the transformation matrix to complete the spatial registration processing of the infrared image and the visible light image.
[0197] In some possible embodiments, the feature fusion module 603 is specifically configured to:
[0198] Calculate the feature cross-correlation between the infrared feature and the visible light feature, and construct a cross-modal attention weight matrix;
[0199] Determine an initial multi-modal fusion feature based on the cross-modal attention weight matrix, the infrared feature, and the visible light feature;
[0200] Perform recalibration processing on the initial multi-modal fusion feature, and perform feature enhancement processing on the result of the recalibration processing based on a multi-scale spatial enhancement strategy to obtain the multi-modal fusion feature.
[0201] In some possible embodiments, the preset three-branch generator architecture includes a local feature generator, a global feature generator, and a spatio-temporal feature generator, and the set of generation results includes a local feature generation result, a global feature generation result, and a spatio-temporal feature generation result; the anomaly detection module 604 is specifically configured to:
[0202] Use the local feature generator to perform local feature extraction processing on the multi-modal fusion feature to obtain the local feature generation result; wherein, the local feature generator includes an encoding-decoding framework and a skip connection structure;
[0203] Use the global feature generator to perform global feature extraction processing on the multi-modal fusion feature to obtain the global feature generation result; wherein, the global feature generator includes a self-attention mechanism and a neural network;
[0204] Use the spatio-temporal feature generator to perform spatio-temporal dimension feature extraction processing on the multi-modal fusion feature to obtain the spatio-temporal feature generation result; wherein, the spatio-temporal feature generator includes a three-dimensional convolutional neural network and a gated recurrent unit.
[0205] In some possible embodiments, the heterogeneous discriminator group includes a local anomaly detection discriminator, a global semantic discriminator, and a temporal consistency discriminator; the anomaly detection module 604 is specifically configured to:
[0206] Evaluate the anomaly situation of the texture and structure of the local area in the local feature generation result based on the local anomaly detection discriminator to obtain a local anomaly probability;
[0207] Evaluate the semantic consistency between the global feature generation result and the multi-modal fusion feature based on the global semantic discriminator to obtain a global anomaly probability;
[0208] Evaluate the physical rationality of the motion trajectory in the spatio-temporal feature generation result based on the temporal consistency discriminator to obtain a temporal anomaly probability.
[0209] In some possible embodiments, the anomaly detection module 604 is further configured to:
[0210] When at least two of the local anomaly probability, the global anomaly probability, and the temporal anomaly probability have values greater than the first warning threshold, it is determined as a primary anomaly;
[0211] When it is determined as a primary anomaly, the local anomaly probability, the global anomaly probability, and the temporal anomaly probability are weighted and calculated according to a preset probability distribution weight to obtain a comprehensive anomaly probability; when the value of the comprehensive anomaly probability is greater than the second warning threshold, it is determined as an ultimate anomaly;
[0212] Based on the anomaly determination result, determine the scene anomaly detection result for the to-be-detected scene.
[0213] In some possible embodiments, the anomaly detection module 604 is further configured to:
[0214] Construct a fully connected graph with three discriminators as nodes and the anomaly detection results of the three discriminators as edge weights;
[0215] Calculate the spectral entropy value of the fully connected graph;
[0216] When the spectral entropy value is lower than a preset entropy value threshold, directly output the scene anomaly detection result;
[0217] When the spectral entropy value is not lower than the preset entropy value threshold, mark the scene anomaly detection result as a complex anomaly, and generate a complex anomaly result report based on the anomaly monitoring results of each discriminator and the scene anomaly detection result, and send the complex anomaly result report to the manual analysis end.
[0218] Based on the same inventive concept, an embodiment of the present disclosure also provides a computer device. Referring to Figure 7 As shown, it is a schematic structural diagram of a computer device 700 provided by an embodiment of the present disclosure, including a processor 701, a memory 702, and a bus 703. Among them, the memory 702 is used to store execution instructions, including an internal memory 7021 and an external memory 7022; here, the internal memory 7021 is also called the main memory, which is used to temporarily store the operation data in the processor 701 and the data exchanged with the external memory 7022 such as a hard disk, and the processor 701 exchanges data with the external memory 7022 through the internal memory 7021.
[0219] In the embodiments of the present application, the memory 702 is specifically configured to store the application program code for executing the solution of the present application, and the processor 701 is used to control the execution. That is, when the computer device 700 runs, the processor 701 communicates with the memory 702 through the bus 703, so that the processor 701 executes the application program code stored in the memory 702, and then executes the method described in any of the foregoing embodiments.
[0220] Among them, the memory 702 may be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc.
[0221] The processor 701 may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0222] It can be understood that the structure schematically shown in the embodiments of the present application does not constitute a specific limitation on the computer device 700. In other embodiments of the present application, the computer device 700 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0223] An embodiment of the present disclosure also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the scenario anomaly detection method described in the foregoing method embodiment. Among them, the storage medium may be a volatile or non-volatile computer-readable storage medium.
[0224] An embodiment of the present disclosure also provides a computer program product, which carries program code. The instructions included in the program code can be used to execute the steps of the scenario anomaly detection method described in the foregoing method embodiment. For details, reference can be made to the foregoing method embodiment and will not be elaborated here.
[0225] Among them, the above computer program product can be specifically implemented in a manner of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.
[0226] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here. In several embodiments provided by the present disclosure, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some communication interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.
[0227] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0228] In addition, in each embodiment of the present disclosure, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0229] When the above-mentioned functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present disclosure. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0230] Finally, it should be noted that the above-mentioned embodiments are only specific implementation manners of the present disclosure, used to illustrate the technical solutions of the present disclosure, rather than limiting them. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present disclosure can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. A method for detecting scene anomalies, characterized in that, Including: Obtain an infrared image and a visible light image corresponding to a scene to be detected, and perform temporal alignment processing and spatial registration processing on the infrared image and the visible light image to respectively obtain infrared modality data and visible light modality data; Use a convolutional network to extract features from the infrared modality data to obtain infrared features corresponding to the infrared modality data; and use a convolutional network to extract features from the visible light modality data to obtain visible light features corresponding to the visible light modality data; Perform feature fusion processing on the infrared features and the visible light features to obtain multi-modal fusion features; Determine a set of generation results based on a preset three-branch generator architecture and the multi-modal fusion features; wherein, the preset three-branch generator architecture includes a local feature generator, a global feature generator, and a spatio-temporal feature generator, and the set of generation results includes a local feature generation result, a global feature generation result, and a spatio-temporal feature generation result; Perform anomaly detection on the set of generation results through a heterogeneous discriminator group to obtain a scene anomaly detection result for the scene to be detected; wherein, the heterogeneous discriminator group includes a local anomaly detection discriminator, a global semantic discriminator, and a temporal consistency discriminator; The determining the set of generation results based on a preset three-branch generator architecture and the multi-modal fusion features includes: Use the local feature generator to perform local feature extraction processing on the multi-modal fusion features to obtain the local feature generation result; wherein, the local feature generator includes an encoder-decoder framework and a skip connection structure; Use the global feature generator to perform global feature extraction processing on the multi-modal fusion features to obtain the global feature generation result; wherein, the global feature generator includes a self-attention mechanism and a neural network; Use the spatio-temporal feature generator to perform spatio-temporal dimension feature extraction processing on the multi-modal fusion features to obtain the spatio-temporal feature generation result; wherein, the spatio-temporal feature generator includes a three-dimensional convolutional neural network and a gated recurrent unit; The performing anomaly detection on the set of generation results through a heterogeneous discriminator group includes: Evaluate the anomaly situation of the texture and structure of the local area in the local feature generation result based on the local anomaly detection discriminator to obtain a local anomaly probability; Evaluate the semantic consistency between the global feature generation result and the multi-modal fusion features based on the global semantic discriminator to obtain a global anomaly probability; Evaluate the physical rationality of the motion trajectory in the spatio-temporal feature generation result based on the temporal consistency discriminator to obtain a temporal anomaly probability.
2. The method according to claim 1, wherein The performing temporal alignment processing and spatial registration processing on the infrared image and the visible light image includes: Based on the acquisition time information of any one of the infrared image and the visible light image, perform timestamp matching on the other image to complete the temporal alignment processing of the infrared image and the visible light image; Respectively extract the image feature point information of the infrared image and the visible light image, and determine the geometric transformation relationship between the two images based on the extraction results and a matching algorithm to determine a transformation matrix; Geometrically transform any one of the infrared image and the visible light image according to the transformation matrix to complete the spatial registration process of the infrared image and the visible light image.
3. The method according to claim 1, wherein The feature fusion process for the infrared feature and the visible light feature includes: Calculating the feature cross-correlation between the infrared feature and the visible light feature to construct a cross-modal attention weight matrix; Determining an initial multi-modal fusion feature based on the cross-modal attention weight matrix, the infrared feature, and the visible light feature; Performing recalibration processing on the initial multi-modal fusion feature and performing feature enhancement processing on the recalibration processing result based on a multi-scale spatial enhancement strategy to obtain the multi-modal fusion feature.
4. The method according to claim 1, characterized in that After performing anomaly detection on the generated result set through the heterogeneous discriminator group, it further includes: Determining a primary anomaly when at least two of the local anomaly probability, the global anomaly probability, and the temporal anomaly probability are greater than the first warning threshold; When a primary anomaly is determined, calculating a weighted sum of the local anomaly probability, the global anomaly probability, and the temporal anomaly probability according to a preset probability assignment weight to obtain a comprehensive anomaly probability; determining an ultimate anomaly when the value of the comprehensive anomaly probability is greater than the second warning threshold; Determining the scene anomaly detection result for the scene to be detected based on the anomaly determination situation.
5. The method according to claim 4, wherein After determining the scene anomaly detection result for the scene to be detected based on the anomaly determination situation, it includes: Constructing a fully connected graph with three discriminators as nodes and the anomaly detection results of the three discriminators as edge weights; Calculating the spectral entropy value of the fully connected graph; Directly outputting the scene anomaly detection result when the spectral entropy value is lower than the preset entropy threshold; When the spectral entropy value is not lower than the preset entropy threshold, marking the scene anomaly detection result as a complex anomaly, and generating a complex anomaly result report based on the anomaly monitoring results of each discriminator and the scene anomaly detection result, and sending the complex anomaly result report to the manual analysis end.
6. A scene anomaly detection device, characterized in that, It includes: An image processing module for obtaining an infrared image and a visible light image corresponding to the scene to be detected, and performing temporal alignment processing and spatial registration processing on the infrared image and the visible light image to obtain infrared modal data and visible light modal data respectively; A feature extraction module for using a convolutional network to extract features from the infrared modal data to obtain an infrared feature corresponding to the infrared modal data; and using a convolutional network to extract features from the visible light modal data to obtain a visible light feature corresponding to the visible light modal data; A feature fusion module for performing feature fusion processing on the infrared feature and the visible light feature to obtain a multi-modal fusion feature; A result generation module, configured to determine a set of generation results based on a preset three-branch generator architecture and the multi-modal fusion features; wherein, the preset three-branch generator architecture includes a local feature generator, a global feature generator, and a spatio-temporal feature generator, and the set of generation results includes a local feature generation result, a global feature generation result, and a spatio-temporal feature generation result; An anomaly detection module, configured to perform anomaly detection on the set of generation results through a heterogeneous discriminator group to obtain a scene anomaly detection result regarding the scene to be detected; wherein, the heterogeneous discriminator group includes a local anomaly detection discriminator, a global semantic discriminator, and a temporal consistency discriminator; The result generation module is specifically configured to: Use the local feature generator to perform local feature extraction processing on the multi-modal fusion features to obtain the local feature generation result; wherein, the local feature generator includes an encoding-decoding framework and a skip connection structure; Use the global feature generator to perform global feature extraction processing on the multi-modal fusion features to obtain the global feature generation result; wherein, the global feature generator includes a self-attention mechanism and a neural network; Use the spatio-temporal feature generator to perform spatio-temporal dimension feature extraction processing on the multi-modal fusion features to obtain the spatio-temporal feature generation result; wherein, the spatio-temporal feature generator includes a three-dimensional convolutional neural network and a gated recurrent unit; The anomaly detection module is specifically configured to: Evaluate the anomaly situation of the texture and structure of the local area in the local feature generation result based on the local anomaly detection discriminator to obtain a local anomaly probability; Evaluate the semantic consistency between the global feature generation result and the multi-modal fusion features based on the global semantic discriminator to obtain a global anomaly probability; Evaluate the physical rationality of the motion trajectory in the spatio-temporal feature generation result based on the temporal consistency discriminator to obtain a temporal anomaly probability.
7. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
8. A computer device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Infrared and visible light image fusion method based on multi-scale generative adversarial network
CN111145131A
Character action recognition analysis method and system based on infrared laser and deep learning
CN118747911A