Scene anomaly detection method and device, storage medium and computer equipment

Through the combination of multimodal data fusion and the preset three-branch generator architecture and heterogeneous discriminator groups, the problem of unstable scene anomaly detection performance in complex environments is solved, and higher detection accuracy and interpretability are achieved.

CN120182901AActive Publication Date: 2025-06-20XIAN ORDNANCE IND TECH IND DEV CO LTD

Patent Information

Application Number
CN202510663218.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-06-20
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

When performing scene abnormality detection in complex environments, it is difficult to deal with factors such as ambient light changes, occlusion interference and complex backgrounds, resulting in unstable performance, especially in low light or inclement weather conditions.

Method used

By introducing multimodal data fusion, a combination of preset three-branch generator architecture and heterogeneous discriminator groups, infrared images and visible light images are acquired for timing and spatial registration, features are extracted using convolutional networks, and feature fusion and abnormal detection are performed.

Benefits of technology

It effectively improves the accuracy and interpretability of scene abnormality detection, enhances the sensitivity of detection and the reliability of identification, reduces the false detection rate and missed detection rate, and improves the interpretability of detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182901A_ABST
    Figure CN120182901A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of scene detection, and provides a scene anomaly detection method and device, a storage medium and computer equipment, and the method comprises the steps: obtaining an infrared image and a visible light image corresponding to a to-be-detected scene, carrying out the time sequence alignment processing and space registration processing of the two images, and obtaining infrared modal data and visible light modal data; performing feature extraction on the infrared modal data and the visible light modal data by using a convolutional network to obtain infrared features and visible light features; carrying out feature fusion processing on the two features to obtain a multi-modal fusion feature; determining a generation result set based on a preset three-branch generator architecture and the multi-modal fusion features; and performing anomaly detection on the generated result set through the heterogeneous discriminator group to obtain a scene anomaly detection result. According to the embodiment of the invention, by introducing the combination of multi-modal data fusion, the preset three-branch generator architecture and the heterogeneous discriminator group, the accuracy and interpretability of scene anomaly detection can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of scene detection technology, and in particular to a scene anomaly detection method, device, storage medium and computer equipment. Background Art

[0002] In application scenarios such as intelligent monitoring, industrial inspection and autonomous driving, anomaly detection technology faces challenges brought by complex environments. Traditional single-modal detection methods usually rely on a single sensor input, such as visible light images or infrared images, but these methods are difficult to cope with factors such as ambient lighting changes, occlusion interference and complex backgrounds, resulting in unstable performance. The limitations of single-modality are particularly obvious in low light or bad weather conditions.

[0003] With the development of multi-sensor technology, the combination of multimodal data such as infrared and visible light provides complementary information for scene understanding. Visible light images provide rich texture and color features, but their performance degrades in low light; while infrared images can reflect the thermal radiation characteristics of objects, which are advantageous in abnormal temperature conditions, but lack detailed information. Therefore, how to fuse these heterogeneous modal data and explore synergistic effects has become the key to improving the robustness of anomaly detection.

[0004] Existing anomaly detection methods usually adopt simple feature splicing or decision-level fusion strategies, which are difficult to fully capture the deep correlation between modalities, resulting in unsatisfactory fusion effects; at the same time, the lack of interpretable analysis of abnormal areas restricts its value and usability in practical applications. Although the rapid development of deep learning technology has provided new ideas for multimodal anomaly detection, current solutions still face multiple problems. On the one hand, information redundancy or semantic conflicts are prone to occur during cross-modal feature fusion, affecting the discriminative ability of fused features; on the other hand, most detection methods in related technologies lack interpretability, affecting users' trust in the detection results. Summary of the invention

[0005] The embodiments of the present disclosure at least provide a scene anomaly detection method, apparatus, storage medium and computer equipment, which can effectively improve the accuracy and interpretability of scene anomaly detection by introducing a combination of multimodal data fusion, a preset three-branch generator architecture and a heterogeneous discriminator group.

[0006] The present disclosure provides a method for detecting anomalies in a scene, including: Acquire an infrared image and a visible light image corresponding to the scene to be detected, and perform temporal alignment processing and spatial alignment processing on the infrared image and the visible light image to obtain infrared modal data and visible light modal data respectively; Feature extraction is performed on the infrared modality data using a convolutional network to obtain infrared features corresponding to the infrared modality data; and, feature extraction is performed on the visible light modality data using a convolutional network to obtain visible light features corresponding to the visible light modality data; Feature fusion processing is performed on the infrared features and the visible light features to obtain multimodal fusion features; A set of generation results is determined based on a preset three-branch generator architecture and the multimodal fusion features; and an anomaly detection is performed on the set of generation results through a heterogeneous discriminator group to obtain a scene anomaly detection result for the scene to be detected.

[0007] In some possible embodiments, the performing temporal alignment processing and spatial registration processing on the infrared image and the visible light image includes: Taking the acquisition time information of any one of the infrared image and the visible light image as a reference, performing timestamp matching on the other image to complete the temporal alignment processing of the infrared image and the visible light image; Respectively extracting the image feature point information of the infrared image and the visible light image, and determining the geometric transformation relationship between the two images based on the extraction results and a matching algorithm to determine a transformation matrix; Performing geometric transformation on any one of the infrared image and the visible light image according to the transformation matrix to complete the spatial registration processing of the infrared image and the visible light image.

[0008] In some possible embodiments, the performing feature fusion processing on the infrared features and the visible light features includes: Calculating the feature cross-correlation between the infrared features and the visible light features to construct a cross-modal attention weight matrix; Determining an initial multimodal fusion feature based on the cross-modal attention weight matrix, the infrared features, and the visible light features; Performing recalibration processing on the initial multimodal fusion feature, and performing feature enhancement processing on the recalibration processing result based on a multi-scale spatial enhancement strategy to obtain the multimodal fusion feature.

[0009] In some possible embodiments, the preset three-branch generator architecture includes a local feature generator, a global feature generator, and a spatio-temporal feature generator, and the set of generation results includes a local feature generation result, a global feature generation result, and a spatio-temporal feature generation result; the determining the set of generation results based on the preset three-branch generator architecture and the multimodal fusion features includes: The local feature generator is used to perform local feature extraction processing on the multi-modal fusion feature to obtain the local feature generation result; wherein, the local feature generator includes an encoding-decoding framework and a skip connection structure; The global feature generator is used to perform global feature extraction processing on the multi-modal fusion feature to obtain the global feature generation result; wherein, the global feature generator includes a self-attention mechanism and a neural network; The spatio-temporal feature generator is used to perform spatio-temporal dimension feature extraction processing on the multi-modal fusion feature to obtain the spatio-temporal feature generation result; wherein, the spatio-temporal feature generator includes a three-dimensional convolutional neural network and a gated recurrent unit.

[0010] In some possible embodiments, the heterogeneous discriminator group includes a local anomaly detection discriminator, a global semantic discriminator, and a temporal consistency discriminator; the performing anomaly detection on the generation result set by the heterogeneous discriminator group includes: Based on the local anomaly detection discriminator, evaluating the anomaly situation of the texture and structure of the local area in the local feature generation result to obtain a local anomaly probability; Based on the global semantic discriminator, evaluating the semantic consistency between the global feature generation result and the multi-modal fusion feature to obtain a global anomaly probability; Based on the temporal consistency discriminator, evaluating the physical rationality of the motion trajectory in the spatio-temporal feature generation result to obtain a temporal anomaly probability.

[0011] In some possible embodiments, after performing anomaly detection on the generation result set by the heterogeneous discriminator group, it further includes: In the case where at least two of the local anomaly probability, the global anomaly probability, and the temporal anomaly probability are greater than a first warning threshold, it is determined as a primary anomaly; In the case of determining a primary anomaly, perform weighted calculation on the local anomaly probability, the global anomaly probability, and the temporal anomaly probability according to a preset probability distribution weight to obtain a comprehensive anomaly probability; in the case where the value of the comprehensive anomaly probability is greater than a second warning threshold, it is determined as an ultimate anomaly; Determine the scene anomaly detection result regarding the to-be-detected scene based on the anomaly determination situation.

[0012] In some possible embodiments, after determining the scene anomaly detection result regarding the to-be-detected scene based on the anomaly determination situation, it includes: Construct a fully connected graph with three discriminators as nodes and the anomaly detection results of the three discriminators as edge weights; Calculate the spectral entropy value of the fully connected graph; When the entropy value of the spectrum is lower than the preset entropy value threshold, directly output the scene anomaly detection result; When the entropy value of the spectrum is not lower than the preset entropy value threshold, mark the scene anomaly detection result as a complex anomaly, and generate a complex anomaly result report based on the anomaly monitoring results of each discriminator and the scene anomaly detection result, and send the complex anomaly result report to the manual analysis terminal.

[0013] An embodiment of the present disclosure provides a scene anomaly detection device, including: An image processing module, configured to obtain an infrared image and a visible light image corresponding to a scene to be detected, and perform time series alignment processing and spatial registration processing on the infrared image and the visible light image to obtain infrared modality data and visible light modality data respectively; A feature extraction module, configured to use a convolutional network to extract features from the infrared modality data to obtain infrared features corresponding to the infrared modality data; and use a convolutional network to extract features from the visible light modality data to obtain visible light features corresponding to the visible light modality data; A feature fusion module, configured to perform feature fusion processing on the infrared features and the visible light features to obtain multi-modal fusion features; An anomaly detection module, configured to determine a set of generation results based on a preset three-branch generator architecture and the multi-modal fusion features; and perform anomaly detection on the set of generation results through a heterogeneous discriminator group to obtain a scene anomaly detection result regarding the scene to be detected.

[0014] In some possible embodiments, the image processing module is specifically configured to: Taking the acquisition time information of any one of the infrared image and the visible light image as a reference, perform timestamp matching on the other image to complete the time series alignment processing of the infrared image and the visible light image; Extract the image feature point information of the infrared image and the visible light image respectively, and determine the geometric transformation relationship between the two images based on the extraction result and the matching algorithm, and determine the transformation matrix; Perform geometric transformation on any one of the infrared image and the visible light image according to the transformation matrix to complete the spatial registration processing of the infrared image and the visible light image.

[0015] In some possible embodiments, the feature fusion module is specifically configured to: Calculate the feature cross-correlation between the infrared features and the visible light features, and construct a cross-modal attention weight matrix; Determine the initial multi-modal fusion features based on the cross-modal attention weight matrix, the infrared features and the visible light features; Perform recalibration processing on the initial multimodal fusion features, and perform feature enhancement processing on the results of the recalibration processing based on a multi-scale spatial enhancement strategy to obtain the multimodal fusion features.

[0016] In some possible embodiments, the preset three-branch generator architecture includes a local feature generator, a global feature generator, and a spatio-temporal feature generator, and the set of generation results includes local feature generation results, global feature generation results, and spatio-temporal feature generation results; the anomaly detection module is specifically configured to: Use the local feature generator to perform local feature extraction processing on the multimodal fusion features to obtain the local feature generation results; wherein, the local feature generator includes an encoding-decoding framework and a skip connection structure; Use the global feature generator to perform global feature extraction processing on the multimodal fusion features to obtain the global feature generation results; wherein, the global feature generator includes a self-attention mechanism and a neural network; Use the spatio-temporal feature generator to perform spatio-temporal dimension feature extraction processing on the multimodal fusion features to obtain the spatio-temporal feature generation results; wherein, the spatio-temporal feature generator includes a three-dimensional convolutional neural network and a gated recurrent unit.

[0017] In some possible embodiments, the heterogeneous discriminator group includes a local anomaly detection discriminator, a global semantic discriminator, and a temporal consistency discriminator; the anomaly detection module is specifically configured to: Evaluate the anomaly situation of the texture and structure of the local region in the local feature generation results based on the local anomaly detection discriminator to obtain a local anomaly probability; Evaluate the semantic consistency between the global feature generation results and the multimodal fusion features based on the global semantic discriminator to obtain a global anomaly probability; Evaluate the physical rationality of the motion trajectory in the spatio-temporal feature generation results based on the temporal consistency discriminator to obtain a temporal anomaly probability.

[0018] In some possible embodiments, the anomaly detection module is further configured to: If at least two of the local anomaly probability, the global anomaly probability, and the temporal anomaly probability are greater than the first warning threshold, it is determined as a primary anomaly; In the case of determining a primary anomaly, perform weighted calculation on the local anomaly probability, the global anomaly probability, and the temporal anomaly probability according to a preset probability assignment weight to obtain a comprehensive anomaly probability; if the value of the comprehensive anomaly probability is greater than the second warning threshold, it is determined as an ultimate anomaly; Determine the scene anomaly detection result for the to-be-detected scene based on the anomaly determination situation.

[0019] In some possible embodiments, the anomaly detection module is further configured to: Construct a fully connected graph with three discriminators as nodes and the anomaly detection results of the three discriminators as edge weights; Calculate the graph spectrum entropy value of the fully connected graph; When the graph spectrum entropy value is lower than a preset entropy value threshold, directly output the scene anomaly detection result; When the graph spectrum entropy value is not lower than the preset entropy value threshold, mark the scene anomaly detection result as a complex anomaly, and generate a complex anomaly result report based on the anomaly monitoring results of each discriminator and the scene anomaly detection result, and send the complex anomaly result report to the manual analysis end.

[0020] An embodiment of the present disclosure provides a computer device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the computer device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the scene anomaly detection method described in any of the above possible implementation manners is executed.

[0021] An embodiment of the present disclosure provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, the scene anomaly detection method described in any of the above possible implementation manners is implemented.

[0022] In the scene anomaly detection method, device, storage medium, and computer device provided in the embodiments of the present disclosure, by introducing the combination of multi-modal data fusion, a preset three-branch generator architecture, and a heterogeneous discriminator group, the accuracy and interpretability of scene anomaly detection can be effectively improved. Specifically, first, through temporal alignment and spatial registration processing, the infrared image and the visible light image can be compared and analyzed in a unified coordinate framework, thereby eliminating possible time and space differences between different modal data and ensuring the precise alignment of the data. Second, a convolutional neural network is used to extract features from the infrared and visible light modal data respectively, which can capture the fine-grained information in the two modal data. At the same time, through feature fusion processing, the advantages of the two modalities are combined, so that the final multi-modal fusion features are more comprehensive, which can effectively improve the detection sensitivity and recognition reliability. Then, based on the preset three-branch generator architecture, the expression ability of the model is further enhanced, making the generated result set more abundant. Finally, the heterogeneous discriminator group is used to perform anomaly detection on the generated results, and the determination criteria from multiple angles can be comprehensively considered, further reducing the false detection rate and the missed detection rate.

[0023] In addition, the preset three-branch generator architecture and the heterogeneous discriminator group can achieve a more transparent feature extraction and discrimination process, effectively solving the problem of the interpretability of detection results and making the anomaly detection results easier to understand and trust.

[0024] To make the above objects, features, and advantages of the present disclosure more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, gives a detailed description as follows. Description of the Drawings

[0025] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required to be cited in the embodiments. The accompanying drawings are incorporated into the specification and constitute a part of this specification. These drawings show embodiments that conform to the present disclosure and, together with the specification, are used to illustrate the technical solutions of the present disclosure. It should be understood that the following drawings only show some embodiments of the present disclosure and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0026] Figure 1 Shows a flowchart of a method for scene anomaly detection provided by an embodiment of the present disclosure; Figure 2 Shows a schematic structural diagram of a multi-layer convolutional network provided by an embodiment of the present disclosure; Figure 3 Shows a flowchart of a feature fusion processing method provided by an embodiment of the present disclosure; Figure 4 Shows a flowchart of an anomaly detection method provided by an embodiment of the present disclosure; Figure 5 Shows a flowchart of a heterogeneous graph spectrum analysis method provided by an embodiment of the present disclosure; Figure 6 Shows a schematic structural diagram of a scene anomaly detection device provided by an embodiment of the present disclosure; Figure 7 Shows a schematic structural diagram of a computer device provided by an embodiment of the present disclosure. Detailed Embodiments

[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are only a part rather than all of the embodiments of the present disclosure. Components of the embodiments of the present disclosure described and illustrated in the accompanying drawings here generally can be arranged and designed in various different configurations. Therefore, the detailed description of the embodiments of the present disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed present disclosure, but merely represents selected embodiments of the present disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of the present disclosure without creative efforts fall within the scope of protection of the present disclosure.

[0028] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0029] The term "and / or" in this document merely describes an association relationship and indicates that three relationships may exist. For example, A and / or B may represent three cases: A exists alone, both A and B exist simultaneously, and B exists alone. In addition, the term "at least one" in this document means any one of multiple items or any combination of at least two of multiple items. For example, including at least one of A, B, and C may represent selecting any one or more elements from the set composed of A, B, and C.

[0030] To facilitate the understanding of this embodiment, the execution subject of the scenario anomaly detection method provided in the embodiments of the present disclosure will be introduced in detail first. The execution subject of the scenario anomaly detection method provided in the embodiments of the present disclosure is a computer device. This computer device may be a terminal device or a server. Among them, the terminal device may also be a mobile device, a user terminal, a terminal, a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The server may be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud storage, big data, and artificial intelligence platforms. Optionally, this method may also be applied to an implementation environment composed of a computer device and a server.

[0031] The following will describe in detail the scenario anomaly detection method provided in the embodiments of the present application with reference to the accompanying drawings. Refer to Figure 1 As shown, it is a flowchart of a scenario anomaly detection method provided in an embodiment of the present disclosure. This method includes the following S101 to S104: S101. Obtain the infrared image and visible light image corresponding to the scene to be detected, and perform temporal alignment processing and spatial registration processing on the infrared image and the visible light image respectively to obtain infrared modality data and visible light modality data.

[0032] It can be understood that the infrared image is formed based on the thermal radiation of an object, can capture the temperature distribution information on the surface of the object, is not restricted by the lighting conditions, and can clearly image even in a dark environment. It is suitable for detecting abnormal situations with obvious temperature differences, such as detecting overheating faults of equipment, personnel activities at night, etc. The visible light image presents the reflection characteristics of an object in the visible light band, has rich colors and clear details, and can intuitively display the appearance characteristics of the scene, such as the shape, color, texture, etc. of the object, which helps to identify some scenes with abnormal appearances, such as damage to buildings, illegal placement of objects, etc.

[0033] However, due to the possible differences in the shooting times of the infrared camera and the visible light camera, as well as their differences in spatial position and shooting angle, the directly obtained images may not be aligned in time and space. Therefore, it is necessary to perform temporal alignment processing and spatial registration processing on the obtained infrared image and visible light image.

[0034] On the one hand, the temporal alignment processing mainly ensures that the infrared image and the visible light image are collected at the same time point or at a close time point to eliminate the image offset problem caused by the shooting time difference. For example, when monitoring a traffic intersection, if the time interval between the collection of the infrared image and the visible light image is too long, situations such as vehicle position movement, pedestrians entering or leaving the scene may occur, thus affecting the accurate judgment of the scene state. During the temporal alignment processing of the infrared image and the visible light image, the collection time information of any one of the infrared image and the visible light image can be used as a reference to match the time stamps of the other image, and a set of images that are closest in time can be selected for analysis. Here, by calculating the difference in the collection times of the two images, an interpolation method (such as linear interpolation, spline interpolation) can be used to adjust one of the images in time, so that the infrared image and the visible light image are matched on the time axis.

[0035] On the other hand, the spatial registration processing is to align the infrared image and the visible light image in the spatial coordinate system so that they can accurately correspond to the same physical space position. This usually involves geometric transformations of the image, such as translation, rotation, scaling, etc.

[0036] In some possible embodiments, the image feature point information of the infrared image and the visible light image can be extracted separately first, and then through the feature point matching algorithm, the same feature points in the infrared image and the visible light image (such as the corner points of buildings, the intersection points of roads, etc.) can be found, and then the transformation matrix can be calculated according to the corresponding relationship of these feature points, and one of the images can be geometrically transformed (such as homography matrix, affine transformation or perspective transformation, etc.) to align it with the other image in space. Among them, the feature extraction method can choose methods such as SIFT, SURF or ORB to capture the significant features in the image. The feature point matching algorithm can choose matching algorithms such as brute-force matching or fast nearest neighbor search matching to find the corresponding points between the two images.

[0037] Here, in order to improve the matching accuracy, the random sample consensus algorithm (RANSAC) can also be used to eliminate the wrong matching points to ensure that the obtained geometric transformation matrix is more accurate.

[0038] In this way, the dual-modal data (i.e., infrared modal data and visible light modal data) after temporal alignment and spatial registration can be consistent in time and space, thereby improving the data quality.

[0039] In some other embodiments, after obtaining the infrared image and the visible light image, modal contrast normalization can also be performed on them respectively. Specifically: For the visible light image, channel-level Z-Score normalization is adopted, and the formula is: ; Among them, , respectively represent the mean and standard deviation of the visible light image; For the infrared image, it can be linearly mapped according to the sensor range and the noise (such as non-uniformity correction) can be eliminated.

[0040] S102, using a convolutional network to extract features from the infrared modal data to obtain infrared features corresponding to the infrared modal data; and using a convolutional network to extract features from the visible light modal data to obtain visible light features corresponding to the visible light modal data.

[0041] It can be understood that after obtaining the infrared modal data and the visible light modal data, a convolutional network can be used to extract features from them respectively. Among them, the convolutional network is a deep learning model, and it can automatically learn the features in the image through structures such as convolutional layers and pooling layers.

[0042] Specifically, for infrared modality data, the convolutional network can capture the temperature distribution patterns and thermal features therein. For example, when detecting industrial equipment, the normal operating state and abnormal states (such as overheating, local damage, etc.) of the equipment will present different temperature distributions on the infrared image. The convolutional network can learn a large number of infrared images of normal and abnormal equipment and extract features that can distinguish these two states, such as temperature gradient changes, hot spot area distributions, etc. These features can be called infrared features, which reflect the essential attributes of objects in the infrared band and help detect temperature-related abnormalities.

[0043] Similarly, for visible light modality data, the convolutional network can extract the appearance features of objects. For example, when monitoring a warehouse scene, there are obvious appearance differences between normal goods placement and abnormal goods accumulation (such as goods tipping over, chaotic placement, etc.) on the visible light image. The convolutional network can learn these appearance features, such as the outline of objects, color distribution, texture details, etc., to obtain visible light features corresponding to the visible light modality data to intuitively reflect the appearance state of the scene.

[0044] Exemplarily, in the present disclosure, a multi-layer convolutional network is adopted to implement the feature extraction tasks for infrared modality data and visible light modality data. Here, the infrared features include low-order infrared feature outputs and high-order infrared feature outputs. Similarly, the visible light features include low-order visible light feature outputs and high-order visible light feature outputs. Refer to Figure 2 As shown, the multi-layer convolutional network in the present disclosure includes a shallow feature extraction module and a deep feature extraction module. Specifically, the shallow feature extraction module is used to capture low-order features of both visible light and infrared modality data. This module is constructed based on the Conv-BN-ReLU basic unit, and stride convolution is used for downsampling instead of pooling operations to maximize the retention of spatial detail information.

[0045] Among them, for infrared modality data, the initial convolutional layer is used to capture the characteristics of the thermal radiation distribution and highlight the feature expressions formed by temperature differences. At the same time, for visible light modality data, rich texture features, including local visual patterns such as edges and corners, can be extracted through 3×3 small convolutional kernels. The feature extraction processes of the two modalities maintain the same network architecture, and finally feature tensors with consistent dimensions are output: The low-order infrared feature output of the infrared modality data can be expressed as: ; Among them, represents the low-order infrared feature output of the infrared modality data; R represents the set of real numbers; represents that after downsampling by stride convolution, the height H and width W of the feature map are 1 / 2 of the original input image respectively; 64 represents the number of channels of the feature map.

[0046] The low - order visible - light feature output of visible - light modality data can be expressed as: ; Wherein, represents the low - order visible - light feature output of visible - light modality data; the height H and width W of the feature map are respectively 1 / 2 of the original input image; 64 represents the number of channels of the feature map.

[0047] Specifically, in the deep - feature extraction stage, the multi - layer convolutional network uses a deep - feature extraction module to mine high - order semantic features from infrared - modality data and visible - light modality data. This module introduces dilated convolution on the basis of standard convolution to expand the receptive field, enabling it to capture a larger range of context information. At the same time, it embeds a self - attention mechanism to dynamically adjust the feature - channel weights and enhance the expression ability of key features.

[0048] Exemplarily, for infrared data, after being processed by the deep - feature extraction module, it can highlight abnormally sensitive features, especially having a significant response to mutation regions in the temperature distribution (such as equipment hot spots, living targets, etc.). For visible - light modality data, the deep - feature extraction module extracts rich semantic context information through multi - level non - linear transformation, including object categories (such as vehicles, pedestrians) and scene semantic labels (such as roads, buildings). These features have stronger discriminability and can support higher - level visual understanding tasks.

[0049] Wherein, the feature - extraction processes of both modalities adopt the same down - sampling strategy, and the final output is a 512 - dimensional feature tensor with a spatial resolution of 1 / 8 (H / 8×W / 8) of the input. That is, the high - order infrared feature output of infrared - modality data can be expressed as: ; Wherein, represents the high - order infrared feature output of infrared - modality data.

[0050] The high - order visible - light feature output of visible - light modality data can be expressed as: ; Wherein, represents the high - order visible - light feature output of visible - light modality data.

[0051] In the present disclosure, a multi-layer convolutional network is introduced to implement the feature extraction task for infrared modality data and visible light modality data. The shallow feature extraction module ensures the spatial alignment of features in each modality and lays a compatibility foundation for subsequent feature fusion. The deep feature extraction module ensures the semantic richness of deep features. Also, the unified feature dimension facilitates subsequent multi-modal fusion. Meanwhile, the introduction of dilated convolution and attention mechanism further enhances the model's ability to model global context and key regions, strengthening its robustness and accuracy in complex scenarios.

[0052] S103. Perform feature fusion processing on the infrared feature and the visible light feature to obtain a multi-modal fusion feature.

[0053] It can be understood that after obtaining the infrared feature and the visible light feature, in order to make full use of the advantages of the two modality data, feature fusion processing can be performed on them to obtain a multi-modal fusion feature. Among them, the purpose of feature fusion is to integrate the feature information of different modalities to obtain a more comprehensive and accurate scene representation.

[0054] Exemplarily, common feature fusion methods include early fusion, mid-term fusion, and late fusion, etc. Early fusion is to splice the infrared image and the visible light image in channels before feature extraction, and then input them into the same convolutional network for feature extraction. Mid-term fusion is to fuse the infrared feature and the visible light feature in the middle layer of the network during the feature extraction process. For example, after a certain convolutional layer of the convolutional network, the features of the two modalities are spliced or added, and then subsequent convolutional operations are performed. This method can better retain the information of different modality features and enable the network to learn the association between them during the feature extraction process. Late fusion is to input the infrared feature and the visible light feature into different classifiers respectively after feature extraction is completed, and then fuse the output results of the classifiers.

[0055] In some possible embodiments, in order to fully exploit the complementary information between the two different modality data of the infrared image and the visible light image and achieve a more accurate scene feature representation, the present disclosure proposes a feature fusion processing method. Referring to Figure 3 shown, it may include the following S301~S303: S301. Calculate the feature cross-correlation between the infrared feature and the visible light feature, and construct a cross-modal attention weight matrix.

[0056] Here, a cross-modal attention weight matrix can be constructed at each level l of the feature fusion network ( l denotes the level number in the network; denotes thel The cross-modal attention weight matrix of the layer; C represents the number of feature channels); to quantify the correlation degree between the quantized infrared features and visible light features, highlighting the correlation intensity between different modal channels (i.e., the cross-correlation between visible light features and infrared features). Each element of this cross-modal attention weight matrix reflects the correlation intensity between different modal channels.

[0057] S302. Determine the initial multi-modal fusion features based on the cross-modal attention weight matrix, the infrared features, and the visible light features.

[0058] Specifically, after obtaining the cross-modal attention weight matrix, the cross-modal attention weight matrix can be used to perform an adaptive weighted combination of the infrared features and the visible light features, and its fusion process is achieved through matrix multiplication: ; Where, can be expressed as the initial multi-modal fusion feature of the l th layer; its dimension is the same as that of , ; , are the weight matrices for the visible light features and the infrared features respectively, and can be obtained by appropriately splitting or transforming the cross-modal attention weight matrix . , are the visible light features and the infrared features of the l th layer respectively.

[0059] In this way, through matrix multiplication, the weight matrix can be multiplied by the corresponding feature matrix to achieve the weighting of different modal features. The weighted features are then added together, enabling the modal features to be adaptively weighted and combined according to their importance.

[0060] S303. Perform a recalibration process on the initial multi-modal fusion features, and perform a feature enhancement process on the result of the recalibration process based on a multi-scale spatial enhancement strategy to obtain the multi-modal fusion features.

[0061] Here, in order to further optimize the initial multi-modal fusion features and solve problems such as feature redundancy and imbalance of feature information at different scales, the present disclosure proposes to perform a recalibration process on the initial multi-modal fusion features and perform a feature enhancement process based on a multi-scale spatial enhancement strategy.

[0062] Specifically, in order to highlight the important channels and suppress redundant channels in the initial multi-modal fusion features, the fused initial multi-modal fusion features can be first subjected to recalibration processing. Channel-level statistical features are obtained through global average pooling, where the feature values at all spatial positions on each channel are averaged. Then, the channel description vectors are input into two fully connected layers for non-linear transformation. After passing through the two fully connected layers, channel attention vectors are generated, which can recalibrate the feature channels. In this way, after channel attention recalibration, the importance of different channels in the fusion features is readjusted, the important channel features are enhanced, and the redundant channel features are suppressed, thereby improving the quality and discriminative power of the features.

[0063] Furthermore, in order to capture feature information at different spatial scales and enhance the perception ability of the fusion features for targets of different sizes, dilated convolutions with different dilation rates (2 / 4 / 6) can be used to process the feature maps in parallel. Dilated convolution expands the receptive field by inserting holes in the convolution kernel, and can obtain more extensive context information without increasing the number of parameters and computational complexity.

[0064] Specifically, the recalibrated fusion feature matrix can be input into three dilated convolution branches with different dilation rates respectively. Each branch uses a 3×3 dilated convolution kernel with dilation rates of 2, 4, and 6 respectively. After the dilated convolution operation, three spatial feature maps of different scales are obtained, which respectively focus on targets and scene regions of different sizes. Then, these three spatial feature maps of different scales are concatenated, and finally, a 1×1 convolution layer is used to reduce the dimension of the concatenated feature map, while maintaining the spatial resolution (H×W) of the feature map and fusing multi-scale context information. The finally output multi-modal fusion features : . Here, the multi-modal fusion features contain both the detailed texture and semantic information of the visible light features, and integrate the temperature distribution features of the infrared features, providing a unified multi-modal representation for the following steps.

[0065] It can be understood that the above fusion process is repeatedly executed at multiple levels of the feature fusion network, forming a progressive feature fusion architecture from local to global. At the low level, feature fusion mainly focuses on local texture and edge information, and associates and fuses the local features in the infrared and visible light images through a cross-modal attention mechanism. At the high level, feature fusion considers more semantic information and global structure, and captures targets and scene regions of different sizes through a multi-scale spatial enhancement strategy, forming a more discriminative global feature representation.

[0066] Through feature fusion processing, the obtained multi-modal fusion features integrate the temperature information of the infrared modality and the appearance information of the visible light modality, and can more comprehensively describe the state of the scene to be detected. For example, in a factory production scenario, the multi-modal fusion features can simultaneously reflect the temperature anomaly of the equipment (through infrared features) and the appearance damage of the equipment (through visible light features), thereby improving the detection ability of scene anomalies.

[0067] S104. Based on the preset three-branch generator architecture and the multi-modal fusion features, determine a set of generation results; and perform anomaly detection on the set of generation results through a heterogeneous discriminator group to obtain a scene anomaly detection result regarding the scene to be detected.

[0068] Specifically, based on the preset three-branch generator architecture and the multi-modal fusion features, a set of generation results can be determined. The three-branch generator architecture is a special structure of a Generative Adversarial Network (GAN), which usually includes three different generator branches, and each branch is responsible for generating different types or different scales of features or images. Here, the preset three-branch generator architecture can include a local feature generator, a global feature generator, and a spatio-temporal feature generator.

[0069] Exemplarily, when determining the set of generation results based on the preset three-branch generator architecture and the multi-modal fusion features, the following (a) - (c) may be included: (a) Use the local feature generator to perform local feature extraction processing on the multi-modal fusion features to obtain the local feature generation result; wherein, the local feature generator includes an encoding-decoding framework and a skip connection structure; (b) Use the global feature generator to perform global feature extraction processing on the multi-modal fusion features to obtain the global feature generation result; wherein, the global feature generator includes a self-attention mechanism and a neural network; (c) Use the spatio-temporal feature generator to perform spatio-temporal dimension feature extraction processing on the multi-modal fusion features to obtain the spatio-temporal feature generation result; wherein, the spatio-temporal feature generator includes a three-dimensional convolutional neural network and a gated recurrent unit.

[0070] Exemplarily, the local feature generator can adopt a U-Net structure, retain spatial high-frequency details through an encoder-decoder framework with skip connections, focus on high-resolution image reconstruction tasks, and the output local feature generation result maintains the same spatial size as the input multi-modal fusion feature, emphasizing the restoration of local detail features such as edges and textures to obtain the local feature generation result. The global feature generator can be designed based on the Transformer architecture, utilize the multi-head self-attention mechanism to model long-range dependencies, process global context information through layer normalization and feed-forward networks, and the output global feature generation result is a pixel-level semantic segmentation mask to achieve semantic consistency expression of the scene. The spatio-temporal feature generator can adopt a structure combining a 3D convolutional neural network and a gated recurrent unit (GRU). The 3D convolutional kernel extracts features in the spatio-temporal dimension, and the GRU module models the inter-frame temporal relationship. The output spatio-temporal feature generation result is a temporal optical flow prediction map for capturing motion patterns in dynamic scenes. The three generators use the multi-modal fusion feature as a shared input and generate output results with different attributes through parallel forward propagation.

[0071] In some other embodiments, the local feature generator can also adopt other different encoder-decoder frameworks and skip connection structures; the global feature generator can be based on other self-attention mechanisms and neural networks; similarly, the spatio-temporal feature generator can combine different 3D convolutional neural networks and recurrent unit structures to achieve the extraction of multi-modal features. The above description is only an example and does not limit the implementation details.

[0072] Exemplarily, in the inference stage, the three-branch generator architecture can support three working modes: any generator can be called separately to obtain a specific output, or multi-task joint inference can be achieved through a feature sharing mechanism, where the shallow features of the U-Net can interact with the attention map of the Transformer, and the spatio-temporal features extracted by the 3D convolution can assist the segmentation task in processing dynamic objects. By explicitly decoupling the three feature dimensions of spatial details, semantic information, and temporal dynamics, this architecture achieves feature-level collaboration while maintaining the independence of each task, providing a flexible and efficient solution for multi-modal visual understanding.

[0073] In some possible embodiments, a multi-task loss function can be used for joint optimization during the training of the three generators: apply pixel-level L1 loss and perceptual loss to the local feature generator to ensure reconstruction quality; use cross-entropy loss and Dice loss to optimize the segmentation accuracy of the global feature generator; use the endpoint error loss unique to optical flow estimation for the spatio-temporal feature generator.

[0074] Furthermore, to promote feature decoupling, the present disclosure also proposes to introduce modality-specific normalization layers in the design of each generator, separate the feature representations required for different generation tasks at the feature level, and at the same time enable each generator to focus on its own task domain through an adversarial training strategy.

[0075] It can be understood that after the generation result set (i.e., the local feature generation result, the global feature generation result, and the spatio-temporal feature generation result) is generated, the heterogeneous discriminator group can perform anomaly detection on the generation result set. Here, the heterogeneous discriminator group consists of multiple discriminators of different types, each discriminator having a different structure and function, and being able to evaluate the generation results from different perspectives. Specifically, the heterogeneous discriminator group can include a local anomaly detection discriminator, a global semantic discriminator, and a temporal consistency discriminator, which perform anomaly detection on the results generated by the above-mentioned generators respectively.

[0076] Among them, during the anomaly detection process, the heterogeneous discriminator group will score each generation result in the generation result set. If a certain generation result is judged as abnormal by the discriminator, then it can be considered that there is an abnormal situation in the scene to be detected. Finally, according to the overall detection results of the heterogeneous discriminator group, the scene anomaly detection result of the scene to be detected can be obtained.

[0077] Specifically, referring to Figure 4 As shown, when performing anomaly detection on the generation result set through the heterogeneous discriminator group, the following S401~S403 can be included: S401, evaluate the anomaly situation of the texture and structure of the local area in the local feature generation result based on the local anomaly detection discriminator, and obtain the local anomaly probability.

[0078] Specifically, after obtaining the local feature generation result generated by the above local feature generator, the local anomaly detection discriminator can be used to evaluate the anomaly situation of the texture and structure of the local area therein to obtain the local anomaly probability. Here, the local anomaly detection discriminator can adopt the PatchGAN (Markov discriminator) architecture. The input local feature generation result is divided into local blocks of 70×70 pixels through a convolutional network, and the anomaly probability of the texture details and structural rationality (i.e., the local anomaly probability) is evaluated block by block. This can effectively capture the subtle anomalies in the local area of the image, such as structural distortion or texture distortion. At the same time, the Markov property of the PatchGAN architecture enables the discriminator to pay more attention to local information during evaluation, avoiding the interference of global information on local anomaly judgment, and improving the sensitivity and detection accuracy of local anomalies.

[0079] S402, evaluate the semantic consistency between the global feature generation result and the multi-modal fusion feature based on the global semantic discriminator, and obtain the global anomaly probability.

[0080] Specifically, after obtaining the global feature generation result generated by the above global feature generator, the global semantic discriminator can be used to evaluate the semantic consistency between the global feature generation result and the multi-modal fusion feature to obtain the global anomaly probability. Here, the global semantic discriminator can be constructed based on ResNet-18. Through multi-level feature extraction and global pooling operations, it evaluates the semantic consistency between the global feature generation result and the multi-modal fusion feature from the perspective of the overall scene, and outputs the global anomaly probability. This discriminator is particularly good at identifying anomalies at the macroscopic level such as incorrect object categories and unreasonable scene layouts. Among them, the multi-modal fusion feature usually integrates information from different modalities (such as images, texts, sensor data, etc.), and these information describe the features and attributes of the target object from different perspectives.

[0081] Here, the global semantic discriminator will judge whether the semantic representations of elements such as roads, vehicles, and pedestrians in the generated global feature generation result are consistent with the semantic information contained in the multi-modal fusion feature. If the direction of the road in the generation result does not match the map information, or the position and type of the vehicle do not match, it indicates poor semantic consistency. The global semantic discriminator will give a relatively high global anomaly probability, indicating that there is an anomaly in the global semantic level of this generation result.

[0082] S403, based on the temporal consistency discriminator, evaluate the physical reasonableness of the motion trajectory in the spatio-temporal feature generation result to obtain the temporal anomaly probability.

[0083] Specifically, for the spatio-temporal feature generation result, the main task of the temporal consistency discriminator is to evaluate the physical reasonableness of the motion trajectory therein, so as to obtain the temporal anomaly probability. In many tasks involving dynamic scenes, such as video generation, motion capture and reconstruction, etc., the spatio-temporal feature generation result contains the change information of the object in time and space, that is, the motion trajectory. Here, the temporal consistency discriminator can adopt a 3D convolutional network structure. Through 3D convolutional kernels, it extracts features in the spatio-temporal dimension, pays attention to the position and pose of the object in each frame of the image, combines with the LSTM module to model the time dependence, analyzes and evaluates whether the changes in the position and pose of the object in the video sequence conform to the physical laws in the time series, and then outputs the temporal anomaly probability to detect abnormal behaviors that violate the natural motion laws.

[0084] For example, in the task of generating sports event videos, the movement trajectories of athletes should follow the principles of human kinematics. If there are jump-like position changes in the running actions of athletes in the generated results that do not conform to the laws of human movement between consecutive frames, or if the movement speed suddenly increases or decreases abnormally beyond the normal speed range of human movement, the temporal consistency discriminator can detect these abnormal situations by analyzing the continuity and physical rationality of the movement trajectories in the time dimension and calculate the corresponding temporal anomaly probability. In addition, the temporal consistency discriminator can also combine physical models, such as Newton's laws of motion, etc., to conduct a more accurate physical rationality assessment of the movement trajectories, further improving the detection ability of temporal anomalies.

[0085] The heterogeneous discriminator group proposed in the embodiments of the present disclosure can comprehensively and meticulously detect anomalies in the generated result set from three different dimensions: local texture structure, global semantic consistency, and temporal physical rationality.

[0086] In some possible embodiments, after obtaining the anomaly probabilities generated by each discriminator, in the anomaly score calculation stage, a comprehensive evaluation can be performed on the output results of the three discriminators, including: first, performing Sigmoid normalization processing on the anomaly probability values of each discriminator to convert the discrimination results into anomaly probability values in the range of [0, 1]; subsequently, calculating the comprehensive anomaly score through weighted summation, where the weight coefficients can be dynamically adjusted according to the importance of each modality.

[0087] In some possible embodiments, to improve the robustness of the evaluation, an anomaly score calibration mechanism is also designed in the present disclosure: constructing a reference distribution using the discriminator outputs of historical normal samples, and measuring the degree to which the current sample deviates from the normal distribution by calculating the Mahalanobis distance; at the same time, introducing an uncertainty estimation module to obtain the score variance through multiple inferences of Monte Carlo Dropout and specially marking the evaluation results with low confidence. This multi-discriminator system can not only give a comprehensive anomaly score but also locate the anomaly types through the special outputs of each discriminator. For example, the high-response area of the local discriminator indicates the texture anomaly position, and the peak frame of the temporal discriminator marks the moment of movement anomaly, providing multi-dimensional decision-making bases for subsequent anomaly analysis and processing.

[0088] Exemplarily, in order to more accurately, comprehensively, and hierarchically evaluate the anomaly degree of the generated result set in the scene to be detected, so as to timely discover potential problems and take targeted measures, after detecting anomalies in the generated result set through the heterogeneous discriminator group to obtain the local anomaly probability, global anomaly probability, and temporal anomaly probability, the following (1) to (3) may further be included: (1) When at least two of the local anomaly probability, the global anomaly probability, and the temporal anomaly probability are greater than the first warning threshold, it is determined as a primary anomaly; (2) In the case of determining a primary anomaly, the local anomaly probability, the global anomaly probability, and the temporal anomaly probability are weighted and calculated according to preset probability distribution weights to obtain a comprehensive anomaly probability; in the case where the value of the comprehensive anomaly probability is greater than a second warning threshold, it is determined as an ultimate anomaly; (3) Based on the anomaly determination result, determine the scene anomaly detection result for the scene to be detected.

[0089] It can be understood that in an actual application scenario, the anomaly probability in a single dimension may be interfered by various factors, and there is a certain risk of misjudgment. For example, in a video generation task, the local feature generation result may cause a temporary increase in the local anomaly probability due to local noise or light changes, but this does not necessarily mean that there are serious problems in the entire generation result. Therefore, by setting joint judgment conditions for multiple anomaly probabilities, that is, when the values of at least two of the local anomaly probability, the global anomaly probability, and the temporal anomaly probability are greater than a first warning threshold, it is determined as a primary anomaly, which can effectively improve the accuracy and reliability of anomaly detection.

[0090] Among them, the setting of the first warning threshold can comprehensively consider the characteristics of the specific task, data distribution, and actual application requirements. For a medical image generation task with extremely high precision requirements, the first warning threshold can be set relatively low to ensure that any possible anomaly can be captured in time; while for some monitoring video analysis tasks with high real-time requirements and allowing a certain degree of tolerance, the first warning threshold can be appropriately increased to avoid frequent warnings due to minor anomaly fluctuations. When the condition that at least two anomaly probabilities are greater than the first warning threshold is met, it is determined as a primary anomaly.

[0091] Furthermore, the primary anomaly determination is only a preliminary hint that there may be an anomaly in the set of generation results. In order to more accurately evaluate the severity of the anomaly, the anomaly information in the three dimensions of local, global, and temporal can be comprehensively considered. The importance of the anomaly probabilities in different dimensions may be different in reflecting the quality of the generation result. Therefore, it is necessary to preset probability distribution weights to perform weighted calculation on the local anomaly probability, the global anomaly probability, and the temporal anomaly probability to obtain a comprehensive anomaly probability.

[0092] In the embodiments of the present disclosure, the preset probability distribution weights are dynamically allocated based on the performance of each discriminator in the receiver operating characteristic curve on the validation set. In some other embodiments, they can also be set in advance based on expert experience, data analysis, or experiments, which are not specifically limited herein.

[0093] It is understandable that the second warning threshold is an indicator for determining the ultimate anomaly, and it is only determined as the ultimate anomaly when the abnormal situation reaches a certain level. When the value of the comprehensive anomaly probability is greater than the second warning threshold, it will be determined as the ultimate anomaly, which indicates that there are relatively serious abnormal problems in the generated result set, and corresponding measures need to be taken immediately for processing, such as manual intervention, etc.

[0094] Specifically, according to the determination results of the above-mentioned primary anomaly and ultimate anomaly, the scene anomaly detection result for the scene to be detected can be determined. If it is not determined as a primary anomaly, it means that the generated result set does not show obvious abnormal situations in the three dimensions of local, global, and temporal, and the scene anomaly detection result is normal. If it is determined as a primary anomaly but not as an ultimate anomaly, it indicates that there is a certain degree of abnormal risk in the scene to be detected, but it has not reached a serious level. At this time, a warning message can be generated to prompt relevant personnel to conduct further manual inspections or analyses on the scene to be detected, and relevant abnormal information can also be recorded for subsequent optimization and improvement of the generation model. If it is determined as an ultimate anomaly, it means that there are serious abnormal problems in the scene to be detected, and the scene anomaly detection result is abnormal. At this time, emergency measures should be taken immediately, such as sending a detailed abnormal report to relevant personnel, including information such as the type of anomaly, the degree of anomaly, and the possible scope of influence, for timely troubleshooting and repair.

[0095] It is understandable that, to further improve the decision-making quality, after obtaining the scene anomaly detection result, the present disclosure also introduces a heterogeneous graph spectral analysis method, as shown in Figure 5 shown, which may include the following steps S501~S504: S501, construct a fully connected graph with three discriminators as nodes and the anomaly detection results of the three discriminators as edge weights.

[0096] It is understandable that when constructing the fully connected graph, these three discriminators are abstracted as nodes in the graph, each node represents an independent anomaly detection perspective or analysis dimension, and the anomaly detection results of the three discriminators serve as the weights of the edges connecting these nodes. Specifically, the weight of the edge can be determined according to the value, confidence level or other quantization indicators of the anomaly detection result. For example, if a discriminator detects an anomaly with a high confidence level, the weight of the edge connected to it will be large; conversely, if the detection result is close to normal, the weight will be small. In this way, the fully connected graph associates the detection results of different discriminators with an intuitive graphical structure, clearly showing the mutual influence and potential correlation between the detection results of each dimension.

[0097] S502, calculate the spectral entropy value of the fully connected graph.

[0098] Here, the graph entropy value is an important indicator for measuring the complexity and uncertainty of information in a graph. It can reflect the uniformity of the edge weight distribution in a fully connected graph and the complexity of the relationships between nodes. In the context of anomaly detection, the level of the graph entropy value is closely related to the complexity of the scene anomaly.

[0099] Specifically, the process of calculating the graph entropy value involves the statistics and analysis of the edge weight distribution in a fully connected graph. First, the weight situation of the edges connecting each node to other nodes is statistically analyzed to obtain the frequency distribution of the weights; then, according to the calculation formula of information entropy, the graph entropy value is calculated in combination with the weight frequency distribution. The basic idea of information entropy is that when the weight distribution is more uniform, the uncertainty of information is greater, and the graph entropy value is higher; conversely, if the weight distribution is concentrated on a few values, it indicates that there are obvious rules or dominant relationships in the graph, and the graph entropy value is lower. By calculating the graph entropy value, the complexity of the current anomaly detection result can be quantitatively evaluated.

[0100] S503, in the case where the graph entropy value is lower than the preset entropy value threshold, directly output the scene anomaly detection result.

[0101] It can be understood that when the calculated graph entropy value is lower than the preset entropy value threshold, it means that the distribution of edge weights in the fully connected graph is relatively concentrated, and there is a high similarity or consistency between the anomaly detection results of different discriminators. This indicates that the current anomaly situation is relatively clear, and the directions of anomaly judgment by each discriminator are basically the same, without obvious contradictions or complex correlations. In this way, the scene anomaly detection result can be directly output, and the existence of the anomaly can be determined without further in-depth analysis.

[0102] S504, in the case where the graph entropy value is not lower than the preset entropy value threshold, mark the scene anomaly detection result as a complex anomaly, and based on the anomaly monitoring results of each discriminator and the scene anomaly detection result, generate a complex anomaly result report, and send the complex anomaly result report to the manual analysis terminal.

[0103] Specifically, if the graph entropy value is not lower than the preset entropy value threshold, it means that the distribution of edge weights in the fully connected graph is relatively dispersed, and there are significant differences or contradictions between the anomaly detection results of different discriminators. This implies that the current anomaly situation is relatively complex, possibly involving the interweaving of multiple dimensions of anomaly factors, or there are some potential problems that are difficult to automatically identify by existing discriminators.

[0104] At this time, the scene anomaly detection result can be marked as a complex anomaly. Meanwhile, a complex anomaly result report is generated based on the anomaly monitoring results of each discriminator and the scene anomaly detection result. This report should detail information such as the detection results of each discriminator, key metrics during the detection process, and the differential analysis between different results. For example, the report can list the specific anomaly types detected by each discriminator, the regions or time points where the anomalies occur, the quantitative metrics of the anomaly degree, etc., and make a preliminary speculation on the reasons for the inconsistent results of each discriminator.

[0105] Furthermore, the complex anomaly result report is sent to the manual analysis terminal to further analyze the deep - seated reasons for the anomalies using the experience and knowledge of relevant technical personnel. Professionals at the manual analysis terminal can conduct in - depth analysis and diagnosis of the complex anomalies based on the detailed information in the report, combined with their own domain knowledge and experience. They can find the root cause of the anomalies and propose targeted solutions through different analysis methods such as viewing the original data and adjusting detection parameters.

[0106] In some possible embodiments, when the scene anomaly detection result shows a complex anomaly, the current detection state can also be frozen, all intermediate features are saved, then the differences in the attention regions of each discriminator are highlighted in the result visualization window, and finally a review report containing a probability distribution graph and a confidence explanation is generated and submitted for manual analysis.

[0107] In the scene anomaly detection method, device, storage medium, and computer device provided in the embodiments of the present disclosure, by introducing the combination of multi - modal data fusion, a preset three - branch generator architecture, and a heterogeneous discriminator group, the accuracy and interpretability of scene anomaly detection can be effectively improved. Specifically, first, through temporal and spatial alignment processing, the infrared image and the visible - light image can be compared and analyzed in a unified coordinate framework, thereby eliminating the possible temporal and spatial differences between different - modal data and ensuring the precise alignment of the data. Second, a convolutional neural network is used to extract features from the infrared and visible - light modal data respectively, which can capture the fine - grained information in the two - modal data. At the same time, through feature fusion processing, combining the advantages of the two modalities can make the final multi - modal fusion features more comprehensive, effectively improving the detection sensitivity and recognition reliability. Then, based on the preset three - branch generator architecture, the expressive ability of the model is further enhanced, making the generated result set richer. Finally, through the heterogeneous discriminator group for anomaly detection of the generated results, the judgment criteria from multiple perspectives can be comprehensively considered, further reducing the false - detection rate and the missed - detection rate.

[0108] In addition, the preset three - branch generator architecture and the heterogeneous discriminator group can achieve a more transparent feature extraction and discrimination process, effectively solving the problem of the interpretability of detection results and making the anomaly detection results easier to understand and trust.

[0109] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order and does not constitute any limitation to the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.

[0110] Based on the same inventive concept, the embodiments of the present disclosure also provide a scene anomaly detection device corresponding to the scene anomaly detection method. Since the principle of solving problems by the device in the embodiments of the present disclosure is similar to the above scene anomaly detection method of the embodiments of the present disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0111] Referring to Figure 6 As shown, it is a schematic diagram of a scene anomaly detection device 600 provided by an embodiment of the present disclosure. The device includes: An image processing module 601, configured to obtain an infrared image and a visible light image corresponding to a scene to be detected, and perform temporal alignment processing and spatial registration processing on the infrared image and the visible light image respectively to obtain infrared modality data and visible light modality data; A feature extraction module 602, configured to use a convolutional network to extract features from the infrared modality data to obtain infrared features corresponding to the infrared modality data; and use a convolutional network to extract features from the visible light modality data to obtain visible light features corresponding to the visible light modality data; A feature fusion module 603, configured to perform feature fusion processing on the infrared features and the visible light features to obtain multi-modal fusion features; An anomaly detection module 604, configured to determine a set of generation results based on a preset three-branch generator architecture and the multi-modal fusion features; and perform anomaly detection on the set of generation results through a heterogeneous discriminator group to obtain a scene anomaly detection result of the scene to be detected.

[0112] In some possible embodiments, the image processing module 601 is specifically configured to: Based on the acquisition time information of any one of the infrared image and the visible light image, perform timestamp matching on the other image to complete the temporal alignment processing of the infrared image and the visible light image; Extract the image feature point information of the infrared image and the visible light image respectively, and determine the geometric transformation relationship between the two images based on the extraction result and the matching algorithm to determine the transformation matrix; Perform geometric transformation on any one of the infrared image and the visible light image according to the transformation matrix to complete the spatial registration processing of the infrared image and the visible light image.

[0113] In some possible embodiments, the feature fusion module 603 is specifically configured to: Calculate the feature cross-correlation between the infrared feature and the visible light feature, and construct a cross-modal attention weight matrix; Determine an initial multi-modal fusion feature based on the cross-modal attention weight matrix, the infrared feature, and the visible light feature; Perform recalibration processing on the initial multi-modal fusion feature, and perform feature enhancement processing on the result of the recalibration processing based on a multi-scale spatial enhancement strategy to obtain the multi-modal fusion feature.

[0114] In some possible embodiments, the preset three-branch generator architecture includes a local feature generator, a global feature generator, and a spatio-temporal feature generator, and the set of generation results includes local feature generation results, global feature generation results, and spatio-temporal feature generation results; the anomaly detection module 604 is specifically configured to: Use the local feature generator to perform local feature extraction processing on the multi-modal fusion feature to obtain the local feature generation result; wherein, the local feature generator includes an encoding-decoding framework and a skip connection structure; Use the global feature generator to perform global feature extraction processing on the multi-modal fusion feature to obtain the global feature generation result; wherein, the global feature generator includes a self-attention mechanism and a neural network; Use the spatio-temporal feature generator to perform spatio-temporal dimension feature extraction processing on the multi-modal fusion feature to obtain the spatio-temporal feature generation result; wherein, the spatio-temporal feature generator includes a three-dimensional convolutional neural network and a gated recurrent unit.

[0115] In some possible embodiments, the heterogeneous discriminator group includes a local anomaly detection discriminator, a global semantic discriminator, and a temporal consistency discriminator; the anomaly detection module 604 is specifically configured to: Evaluate the anomaly situation of the texture and structure of the local region in the local feature generation result based on the local anomaly detection discriminator to obtain a local anomaly probability; Evaluate the semantic consistency between the global feature generation result and the multi-modal fusion feature based on the global semantic discriminator to obtain a global anomaly probability; Evaluate the physical rationality of the motion trajectory in the spatio-temporal feature generation result based on the temporal consistency discriminator to obtain a temporal anomaly probability.

[0116] In some possible embodiments, the anomaly detection module 604 is further configured to: In the case where at least two of the local anomaly probability, the global anomaly probability, and the temporal anomaly probability have values greater than the first warning threshold, it is determined as a primary anomaly; In the case of being determined as a primary anomaly, the local anomaly probability, the global anomaly probability, and the temporal anomaly probability are weighted and calculated according to a preset probability distribution weight to obtain a comprehensive anomaly probability; in the case where the value of the comprehensive anomaly probability is greater than the second warning threshold, it is determined as an ultimate anomaly; Based on the anomaly determination result, determine the scene anomaly detection result for the to-be-detected scene.

[0117] In some possible embodiments, the anomaly detection module 604 is further configured to: Construct a fully connected graph with three discriminators as nodes and the anomaly detection results of the three discriminators as edge weights; Calculate the spectral entropy value of the fully connected graph; In the case where the spectral entropy value is lower than a preset entropy threshold, directly output the scene anomaly detection result; In the case where the spectral entropy value is not lower than the preset entropy threshold, mark the scene anomaly detection result as a complex anomaly, and generate a complex anomaly result report based on the anomaly monitoring results of each discriminator and the scene anomaly detection result, and send the complex anomaly result report to the manual analysis terminal.

[0118] Based on the same inventive concept, an embodiment of the present disclosure further provides a computer device. Refer to Figure 7 As shown, it is a schematic structural diagram of a computer device 700 provided by an embodiment of the present disclosure, including a processor 701, a memory 702, and a bus 703. Among them, the memory 702 is used to store execution instructions, including an internal memory 7021 and an external memory 7022; here, the internal memory 7021 is also called the main memory, which is used to temporarily store the operation data in the processor 701 and the data exchanged with the external memory 7022 such as a hard disk, and the processor 701 exchanges data with the external memory 7022 through the internal memory 7021.

[0119] In an embodiment of the present application, the memory 702 is specifically used to store the application program code for implementing the solution of the present application, and is controlled by the processor 701 to execute. That is, when the computer device 700 runs, the processor 701 communicates with the memory 702 through the bus 703, so that the processor 701 executes the application program code stored in the memory 702, and further executes the method described in any of the foregoing embodiments.

[0120] Among them, the memory 702 can be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc.

[0121] The processor 701 may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0122] It can be understood that the structure schematically shown in the embodiments of this application does not constitute a specific limitation on the computer device 700. In other embodiments of this application, the computer device 700 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure can be implemented in hardware, software, or a combination of software and hardware.

[0123] The embodiments of the present disclosure also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the scenario anomaly detection method described in the above method embodiments. Among them, the storage medium may be a volatile or non-volatile computer-readable storage medium.

[0124] Embodiments of the present disclosure also provide a computer program product. The computer program product carries program codes, and the instructions included in the program codes can be used to execute the steps of the scenario anomaly detection method described in the above method embodiments. For details, reference can be made to the above method embodiments and will not be elaborated herein.

[0125] Among them, the above computer program product can be specifically implemented in the form of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.

[0126] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein. In several embodiments provided by the present disclosure, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. Also, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some communication interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0127] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0128] In addition, in each embodiment of the present disclosure, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0129] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present disclosure. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.

[0130] Finally, it should be noted that the above-mentioned embodiments are only specific implementation manners of the present disclosure, used to illustrate the technical solutions of the present disclosure, rather than limiting them. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present disclosure can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A method for detecting scene anomalies, characterized in that, Including: Obtain the infrared image and visible light image corresponding to the scene to be detected, and perform temporal alignment processing and spatial registration processing on the infrared image and the visible light image to obtain infrared modality data and visible light modality data respectively; Use a convolutional network to extract features from the infrared modality data to obtain infrared features corresponding to the infrared modality data; and use a convolutional network to extract features from the visible light modality data to obtain visible light features corresponding to the visible light modality data; Perform feature fusion processing on the infrared features and the visible light features to obtain multi-modal fusion features; Determine a set of generation results based on a preset three-branch generator architecture and the multi-modal fusion features; and perform anomaly detection on the set of generation results through a heterogeneous discriminator group to obtain a scene anomaly detection result for the scene to be detected.

2. The method according to claim 1, characterized in that, The performing temporal alignment processing and spatial registration processing on the infrared image and the visible light image includes: Taking the acquisition time information of any one of the infrared image and the visible light image as a reference, perform timestamp matching on the other image to complete the temporal alignment processing of the infrared image and the visible light image; Extract the image feature point information of the infrared image and the visible light image respectively, and determine the geometric transformation relationship between the two images based on the extraction results and the matching algorithm, and determine the transformation matrix; Perform geometric transformation on any one of the infrared image and the visible light image according to the transformation matrix to complete the spatial registration processing of the infrared image and the visible light image.

3. The method according to claim 1, characterized in that, The performing feature fusion processing on the infrared features and the visible light features includes: Calculate the feature cross-correlation between the infrared features and the visible light features, and construct a cross-modal attention weight matrix; Determine an initial multi-modal fusion feature based on the cross-modal attention weight matrix, the infrared features and the visible light features; Perform recalibration processing on the initial multi-modal fusion feature, and perform feature enhancement processing on the recalibration processing result based on a multi-scale spatial enhancement strategy to obtain the multi-modal fusion feature.

4. The method according to claim 3, characterized in that, The preset three-branch generator architecture includes a local feature generator, a global feature generator and a spatio-temporal feature generator, and the set of generation results includes a local feature generation result, a global feature generation result and a spatio-temporal feature generation result; The determining a set of generation results based on a preset three-branch generator architecture and the multi-modal fusion features includes: Use the local feature generator to perform local feature extraction processing on the multi-modal fusion features to obtain the local feature generation result; wherein, the local feature generator includes an encoding-decoding framework and a skip connection structure; Use the global feature generator to perform global feature extraction processing on the multi-modal fusion features to obtain the global feature generation result; wherein, the global feature generator includes a self-attention mechanism and a neural network; The spatio-temporal feature generator is used to perform spatio-temporal dimensional feature extraction processing on the multi-modal fusion feature to obtain the spatio-temporal feature generation result; wherein, the spatio-temporal feature generator includes a three-dimensional convolutional neural network and a gated recurrent unit.

5. The method according to claim 4, characterized in that, The heterogeneous discriminator group includes a local anomaly detection discriminator, a global semantic discriminator, and a temporal consistency discriminator; the performing anomaly detection on the generated result set by the heterogeneous discriminator group includes: Evaluating the anomaly situation of the texture and structure of the local area in the local feature generation result based on the local anomaly detection discriminator to obtain a local anomaly probability; Evaluating the semantic consistency between the global feature generation result and the multi-modal fusion feature based on the global semantic discriminator to obtain a global anomaly probability; Evaluating the physical rationality of the motion trajectory in the spatio-temporal feature generation result based on the temporal consistency discriminator to obtain a temporal anomaly probability.

6. The method according to claim 5, characterized in that, After the performing anomaly detection on the generated result set by the heterogeneous discriminator group, it further includes: In the case that at least two of the local anomaly probability, the global anomaly probability, and the temporal anomaly probability are greater than the first warning threshold, it is determined as a primary anomaly; In the case of being determined as a primary anomaly, the local anomaly probability, the global anomaly probability, and the temporal anomaly probability are weighted and calculated according to a preset probability assignment weight to obtain a comprehensive anomaly probability; in the case that the value of the comprehensive anomaly probability is greater than the second warning threshold, it is determined as an ultimate anomaly; Determining the scene anomaly detection result regarding the to-be-detected scene based on the anomaly determination situation.

7. The method according to claim 6, characterized in that, After the determining the scene anomaly detection result regarding the to-be-detected scene based on the anomaly determination situation, it includes: Constructing a fully connected graph with three discriminators as nodes and the anomaly detection results of the three discriminators as edge weights; Calculating the graph entropy value of the fully connected graph; In the case that the graph entropy value is lower than a preset entropy value threshold, directly outputting the scene anomaly detection result; In the case that the graph entropy value is not lower than the preset entropy value threshold, then marking the scene anomaly detection result as a complex anomaly, and generating a complex anomaly result report based on the anomaly monitoring results of each discriminator and the scene anomaly detection result, and sending the complex anomaly result report to the manual analysis end.

8. A device for detecting scene anomalies, characterized in that, Including: An image processing module, configured to obtain an infrared image and a visible light image corresponding to the to-be-detected scene, and perform temporal alignment processing and spatial registration processing on the infrared image and the visible light image to respectively obtain infrared modal data and visible light modal data; A feature extraction module, configured to use a convolutional network to extract features from the infrared modal data to obtain infrared features corresponding to the infrared modal data; and use a convolutional network to extract features from the visible light modal data to obtain visible light features corresponding to the visible light modal data; A feature fusion module, configured to perform feature fusion processing on the infrared features and the visible light features to obtain multi-modal fusion features; An anomaly detection module, configured to determine a set of generation results based on a preset three-branch generator architecture and the multi-modal fusion features; and perform anomaly detection on the set of generation results through a heterogeneous discriminator group to obtain a scene anomaly detection result regarding the scene to be detected.

9. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

10. A computer device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Infrared and visible light image fusion method based on multi-scale generative adversarial network

    CN111145131A

  • Road condition prediction method and device, equipment and storage medium

    CN114332699A

  • Character action recognition analysis method and system based on infrared laser and deep learning

    CN118747911A

  • Railway anomaly detection method and system based on multi-modal data fusion

    WO2025092018A1

Cited By

  • Channel buoy detection method based on fusion of multi-mode pulse neural network and visual Transform

    CN121616952A