Small target abnormal change detection method and device and storage medium
By combining the Swin Transformer encoder and Pyramid Mamba decoder, the feature loss problem caused by lighting and viewing angle changes in small object detection of power equipment is solved, and the accuracy and robustness of the detection are improved.
Patent Information
- Application Number
- CN202510518817.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-08-01
AI Technical Summary
In the detection of abnormal changes in power equipment, especially in small target detection, there is a problem that the light and viewing angle changes have a large impact, resulting in the disappearance of features or the loss of target positions, and the high misjudgment rate.
Using a combination of Swin Transformer encoder and Pyramid Mamba decoder, the two-time phase images are processed through front-end feature extraction, multi-scale feature extraction, feature fusion and decoding, and differential discrimination modules to improve the detection capability of small objects.
Effectively retain low-frequency information of the image, reduce information redundancy, enhance multi-scale feature representation ability, and improve the detection accuracy of abnormal changes in small targets and local areas.
Smart Images

Figure CN120411037A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of anomaly change detection, and particularly relates to a small target anomaly change detection method, device, and storage medium. Background Art
[0002] In the power scenario, when performing a series of tasks such as maintaining, overhauling, and fault handling of power equipment, it is necessary to conduct multi-scenario and all-round inspections to ensure the stable operation of the power system. When detecting anomaly changes in the state of power equipment based on dual-temporal power images, it is necessary to cope with the changes in lighting conditions and the impacts brought by perspective deviations, and ensure that subtle anomalies of small target equipment can be accurately detected.
[0003] When performing anomaly change detection, existing technologies usually adopt traditional convolutional methods to deeply mine and extract deep features in images, methods based on attention mechanisms to obtain more refined and accurate feature representations during feature reconstruction, and methods based on Transformer to effectively capture global interaction information of context. Although these methods have achieved good results in dealing with lighting and perspective problems, there are still obvious deficiencies in the detection of small targets in multi-scale scenarios: for example, in the process of continuous downsampling, the features of small targets may disappear or the precise spatial positions of targets may be lost, resulting in the uncertainty of edge target pixels and misjudgment of small target detection. Therefore, it is of crucial significance to study a method that can improve the model's ability to detect small targets and local area anomaly changes in multi-scale scenarios. Summary of the Invention
[0004] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a small target anomaly change detection method, device, equipment, and storage medium. On the one hand, the present invention uses a Swin Transformer encoder to extract deep features from shallowly extracted features, and on the other hand, directly maps the shallowly extracted features into a Pyramid Mamba decoder to combine multi-scale semantic features. Finally, a differential discrimination module is used to detect small target anomaly changes between dual-temporal images in the same scene, thereby improving the model's detection ability.
[0005] To achieve the above object, the present invention is implemented by the following technical solutions:
[0006] In a first aspect, the present invention provides a small target anomaly change detection method, and the method includes:
[0007] Obtain dual-temporal images of different target devices in the same scene to be detected;
[0008] Input the dual - temporal image into the trained small - target anomaly change detection model for judgment. Among them, the small - target anomaly change detection model includes a front - end feature extraction module, a multi - scale feature extraction module, a feature fusion and decoding module, and a differential discrimination module;
[0009] Extract shallow - layer features from the dual - temporal image through the 3×3 convolutional layer in the front - end feature extraction module to generate shallow - layer extracted features;
[0010] Input the shallow - layer extracted features into the multi - scale feature extraction module improved based on Swin Transformer to extract deep features and generate multi - scale semantic features;
[0011] Integrate the multi - scale semantic features and the shallow - layer extracted features through the fuser in the feature fusion and decoding module, and then upsample the integrated features through the Pyramid Mamba decoder to generate a difference feature image;
[0012] Input the difference feature image into the differential discrimination module for differential processing to obtain the anomaly change region image of the dual - temporal image;
[0013] Extract the region with non - zero pixel values in the anomaly change region image and output the coordinate information, so as to obtain the anomaly change region of the dual - temporal image.
[0014] Combined with the first aspect, further, the multi - scale feature extraction module improved based on Swin Transformer includes a patch partition module, a first - stage structure, a second - stage structure, and a third - stage structure connected in series in sequence;
[0015] The method of inputting the shallow - layer extracted features into the multi - scale feature extraction module improved based on Swin Transformer to extract deep features and generate multi - scale semantic features includes:
[0016] Input the shallow - layer extracted features of the dual - temporal image into the multi - scale feature extraction module, and perform image block division through the patch partition module to obtain non - overlapping image block features;
[0017] Input each image block feature into the first - stage structure, the second - stage structure, and the third - stage structure in sequence for processing to obtain the first - stage semantic feature, the second - stage semantic feature, and the multi - scale semantic feature respectively.
[0018] In combination with the first aspect, further, the first-stage structure includes a linear embedding module, a first Swin transformation module, and a second Swin transformation module connected in series in sequence. The second-stage structure includes a first patch merging module, a third Swin transformation module, and a fourth Swin transformation module connected in series in sequence. The third-stage structure includes a second patch merging module, a fifth Swin transformation module, a sixth Swin transformation module, a seventh Swin transformation module, an eighth Swin transformation module, a ninth Swin transformation module, and a tenth Swin transformation module connected in series in sequence;
[0019] The first-stage structure, the second-stage structure, and the third-stage structure perform the following processing steps on the input features:
[0020] The linear embedding module in the first-stage structure maps each input image patch feature to a high-dimensional embedding space, and then successively passes through the first Swin transformation module and the second Swin transformation module to realize image feature modeling through the windowed multi-head self-attention mechanism and the shifted window multi-head self-attention mechanism, obtaining the first-stage semantic features;
[0021] The first patch merging module in the second-stage structure performs downsampling and merging operations on the first-stage semantic features, and then successively passes through the third Swin transformation module and the fourth Swin transformation module to realize image feature modeling through the windowed multi-head self-attention mechanism and the shifted window multi-head self-attention mechanism, obtaining the second-stage semantic features;
[0022] The second patch merging module in the third-stage structure performs downsampling and merging operations on the second-stage semantic features, and then successively passes through the fifth Swin transformation module to the tenth Swin transformation module to realize image feature modeling through the windowed multi-head self-attention mechanism and the shifted window multi-head self-attention mechanism, obtaining the multi-scale semantic features.
[0023] In combination with the first aspect, further, the feature fusion and decoding module includes a fuse and a PyramidMamba decoder connected in series; wherein, the Pyramid Mamba decoder includes a dense spatial pyramid pooling module and a pyramid fusion module connected in series;
[0024] The method of integrating the multi-scale semantic features and the shallow extraction features through the fuse in the feature fusion and decoding module, and then performing upsampling on the integrated features by the Pyramid Mamba decoder to generate a difference feature image includes:
[0025] Fusing the multi-scale semantic features of the dual-temporal images and the shallow extraction features through the fuse to obtain fused features;
[0026] The dense spatial pyramid pooling module encodes the fused features using multiple pooling scales and upsamples the encoded features using bilinear interpolation to match the high-level features, obtaining upsampled features;
[0027] The pyramid fusion module reduces the semantic redundancy in the upsampled features through a selective filtering mechanism and further enhances the multi-scale feature representation through a convolutional feed-forward neural network, obtaining a differential feature image.
[0028] Combined with the first aspect, further, the differential discrimination module includes a first fully convolutional neural network, a second fully convolutional neural network, a third fully convolutional neural network, and a differentiator connected in series in sequence;
[0029] The method of inputting the differential feature image into the differential discrimination module for differential processing to obtain the abnormal change region image of the dual-temporal image includes:
[0030] Successively passing the differential feature image of the dual-temporal image through the first fully convolutional neural network, the second fully convolutional neural network, and the third fully convolutional neural network for feature restoration to obtain the restored feature image of the dual-temporal image;
[0031] Calculating the absolute value of the subtraction of the restored feature image of the dual-temporal image by the differentiator to obtain the abnormal change region image of the dual-temporal image.
[0032] Combined with the first aspect, further, the calculation of the absolute value of the subtraction of the restored feature image of the dual-temporal image by the differentiator can be expressed by the following formula:
[0033] ,
[0034] In the formula, is the abnormal change region image; and are the restored feature images of the dual-temporal image, , are the height, width, and number of channels respectively; is the classifier; is the softmax function, which is used to convert the output of the classifier into a probability distribution to more accurately discriminate the abnormal change region.
[0035] Combined with the first aspect, further, the training method of the small target abnormal change detection model includes:
[0036] Obtain multiple groups of dual-temporal sample images of different target devices in the same scene;
[0037] Determine the abnormal change region in the dual-temporal sample images and perform mask marking;
[0038] Construct a dual - temporal sample dataset using the dual - temporal sample images marked with the mask;
[0039] Divide the dual - temporal sample dataset into a training set, a validation set, and a test set;
[0040] Use the training set to train a pre - constructed small - target anomaly change detection model, use the validation set to validate the trained small - target anomaly change detection model, use the test set to test the validated small - target anomaly change detection model, and take the small - target anomaly change detection model with the optimal test result as the finally trained small - target anomaly change detection model.
[0041] Combined with the first aspect, further, the training method of the small - target anomaly change detection model further includes:
[0042] Combine the cross - entropy loss function and the mean - square error loss function to construct a joint loss function;
[0043] Optimize the small - target anomaly change detection model with the goal of minimizing the value of the joint loss function;
[0044] Among them, the expression of the joint loss function is:
[0045] ,
[0046] In the formula, is the cross - entropy loss function; is the mean - square error loss function; is a hyper - parameter, set to 0.7, used to balance the influence of the two losses;
[0047] The expression of the cross - entropy loss function is:
[0048] ,
[0049] In the formula, represents the cross - entropy loss with label at position ; are the height, width of the image and the corresponding label respectively during the model training optimization process, ; and are the height and width of the dual - temporal original image;
[0050] The expression of the mean - square error loss function is:
[0051] ,
[0052] In the formula, is the true value of the th sample; is the predicted value of the model for the th sample; is the total number of samples.
[0053] In a second aspect, the present invention further provides a computer device, including a storage medium and a processor;
[0054] The storage medium is used to store instructions;
[0055] The processor is used to operate according to the instructions to execute the steps of the method according to any one of the first aspect.
[0056] In a third aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method according to any one of the first aspect are implemented.
[0057] Compared with the prior art, the present invention can at least achieve the following beneficial effects:
[0058] The small target abnormal change detection method provided by the present invention processes dual-temporal images in the same scene based on the combination of a Swin Transformer encoder and a Pyramid Mamba decoder. First, the front-end feature extraction module finely captures the shallow extraction features containing rich image details and local information; then, on the one hand, the shallow extraction features are input into the multi-scale feature extraction module, and the Swin Transformer encoder is used to extract the deep features, effectively retaining the low-frequency information of the image; on the other hand, the shallow extraction features are directly mapped into the Pyramid Mamba decoder to combine multi-scale semantic features, which can significantly reduce information redundancy while further enhancing the representation ability of multi-scale features; finally, the differential discrimination module performs differential on the dual-temporal difference feature image to effectively capture and distinguish the multi-scale change information in the dual-temporal images in the same scene. This method not only maintains the accuracy of the model in detecting abnormal changes of large-scale targets, but also further improves the ability of the model to detect abnormal changes of small targets and local regions. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0060] Figure 1It is a flowchart of a small target abnormal change detection method provided by an embodiment of the present invention;
[0061] Figure 2 It is a schematic structural diagram of a small target abnormal change detection model provided by an embodiment of the present invention;
[0062] Figure 3 It is a schematic structural diagram of a multi-scale feature extraction module provided by an embodiment of the present invention;
[0063] Figure 4 It is a schematic structural diagram of a differential discrimination module provided by an embodiment of the present invention;
[0064] Figure 5 It is the detection result of common small targets in the power scene provided by an application example of the present invention;
[0065] Figure 6 It is the detection result of small targets weakened due to the change of shooting angle in the power scene provided by an application example of the present invention;
[0066] Figure 7 It is the detection result of conventional scale targets in the power scene provided by an application example of the present invention;
[0067] Figure 8 It is an internal structure diagram of a computer device provided by an embodiment of the present invention. Specific embodiments
[0068] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.
[0069] Embodiment 1:
[0070] This embodiment provides a small target abnormal change detection method. As Figure 1 shown, it is a flowchart of the method provided in this embodiment, mainly including the following steps:
[0071] Step S1: Obtain dual-temporal images of different target devices in the same scene to be detected;
[0072] Step S2: Input the dual-temporal images into the trained small target abnormal change detection model for judgment;
[0073] Step S3: Extract shallow features from the dual-temporal images through the 3×3 convolutional layer in the front-end feature extraction module to generate shallow extraction features;
[0074] Step S4: Input the shallow extraction features into the multi-scale feature extraction module improved based on Swin Transformer to extract deep features and generate multi-scale semantic features;
[0075] Step S5: Integrate the multi-scale semantic features and the shallow extraction features through the fuser in the feature fusion and decoding module, and then upsample the integrated features through the Pyramid Mamba decoder to generate a difference feature image;
[0076] Step S6: Input the difference feature image into the differential discrimination module for differential processing to obtain an abnormal change region image of the bi-temporal image;
[0077] Step S7: Extract the regions with non-zero pixel values in the abnormal change region image and output the coordinate information, so as to obtain the abnormal change region of the bi-temporal image.
[0078] It should be noted that the bi-temporal image usually refers to images of the same region or object obtained at different time points, which can capture the dynamic information generated by changes in the spectral reflection characteristics of ground objects, vegetation growth status, land use patterns, etc. at different times. And the bi-temporal images obtained in step S1 of this embodiment cover images of target devices at different scales taken from different perspectives and distances.
[0079] Furthermore, as Figure 2 shown, it is a schematic structural diagram of the small target abnormal change detection model provided by this embodiment, mainly including a front-end feature extraction module, a multi-scale feature extraction module, a feature fusion and decoding module, and a differential discrimination module; among them, the output end of the front-end feature extraction module is connected to the input end of the multi-scale feature extraction module, the output ends of the multi-scale feature extraction module and the front-end feature extraction module are jointly connected to the input end of the feature fusion and decoding module, and the output end of the feature fusion and decoding module is connected to the input end of the differential discrimination module.
[0080] Referring to Figure 2 , the front-end feature extraction module mainly includes a 3×3 convolutional layer. Since the convolutional layer is good at initial visual feature processing, it can bring more stable optimization and better output results to the model. And the shallow extraction features generated in step S3 contain the spatial information of multi-scale targets, which can help the model better capture the target details and local information of the image.
[0081] In this embodiment, in order to more effectively retain the image detail features and position information, and at the same time avoid losing the semantic information of dense small targets during the process of feature reduction, this embodiment proposes a multi-scale feature extraction module improved based on Swin Transformer: Since the first three stages contain richer feature information, the fourth stage structure of Swin Transformer is discarded, and the network structures of the first three stages are retained, so as to obtain different levels of semantic information of the bi-temporal image.
[0082] Specifically, as Figure 3 shown, it is a schematic structural diagram of the multi-scale feature extraction module provided in this embodiment, mainly including a patch partition module Patch Partition, a first-stage structure, a second-stage structure, and a third-stage structure connected in series in sequence; among them, the first-stage structure includes a linear embedding module Linear Embedding, a first Swin Transformer Block, and a second Swin Transformer Block connected in series in sequence, the second-stage structure includes a first patch merging module Patch Merging, a third Swin Transformer Block, and a fourth Swin Transformer Block connected in series in sequence, and the third-stage structure includes a second patch merging module, a fifth Swin Transformer Block, a sixth Swin Transformer Block, a seventh Swin Transformer Block, an eighth Swin Transformer Block, a ninth Swin Transformer Block, and a tenth Swin Transformer Block connected in series in sequence.
[0083] Next, in combination with Figure 3 , a further detailed description will be given on how to use the multi-scale feature extraction module to extract deep features from the shallow-layer extracted features and generate multi-scale semantic features in step S4:
[0084] Input the shallow-layer extracted features of the bi-temporal image into the multi-scale feature extraction module, and divide them into non-overlapping image patch features of equal size through the patch partition module;
[0085] Map each image patch feature to a high-dimensional embedding space through the linear embedding module in the first-stage structure to form a learnable feature representation; then sequentially pass through the first Swin Transformer Block and the second Swin Transformer Block to implement image feature modeling through the windowed multi-head self-attention mechanism and the shifted window multi-head self-attention mechanism, and obtain the first-stage semantic features;
[0086] Perform downsampling and merging adjacent image patch features on the first-stage semantic features through the first patch merging module in the second-stage structure to reduce the spatial resolution and increase the number of channels, and obtain the first-level features; then sequentially pass through the third Swin Transformer Block and the fourth Swin Transformer Block to implement image feature modeling through the windowed multi-head self-attention mechanism and the shifted window multi-head self-attention mechanism, and obtain the second-stage semantic features;
[0087] Perform downsampling and merging adjacent image patch features on the second-stage semantic features through the second patch merging module in the third-stage structure to reduce the spatial resolution and increase the number of channels, and obtain the second-level features; then sequentially pass through the fifth Swin Transformer Block to the tenth Swin Transformer Block to implement image feature modeling through the windowed multi-head self-attention mechanism and the shifted window multi-head self-attention mechanism, and obtain the multi-scale semantic features.
[0088] It should be noted that Swin Transformer is a new vision model based on the Transformer architecture. It adopts a hierarchical design and constructs feature maps of different scales by gradually merging the features of adjacent image patches, similar to the downsampling operation in traditional convolutional neural networks. It can efficiently process images of different sizes and capture feature information at different scales. In addition, the first, third, fifth, seventh, and ninth Swin transformation modules all adopt the windowed multi-head self-attention mechanism W-MSA, which confines the attention to a specific window when processing images, thereby reducing the computational and storage overhead; the second, fourth, sixth, eighth, and tenth Swin transformation modules all adopt the shifted window multi-head self-attention mechanism SW-MSA, which moves the window along the width and height directions through local translation operations to obtain different local neighborhood information, ensuring the interaction between the information in different windows and reducing the computational amount. Among them, the windowed multi-head self-attention mechanism W-MSA and the shifted window multi-head self-attention mechanism SW-MSA appear in pairs, and the output end of W-MSA is connected to the input end of SW-MSA.
[0089] The multi-scale feature extraction module provided in this embodiment combines the advantages of CNN and Transformer self-attention mechanisms by retaining the first three-stage structure of Swin Transformer. It not only has strong multi-scale modeling capabilities but also can find a balance between global modeling and local features, which can significantly improve the detection ability for small target abnormal changes.
[0090] As an embodiment, in order to retain the shallow features of the dual-temporal images and supplement high-level semantic information to further improve the accuracy of image processing, this embodiment proposes a feature fusion and decoding module. Refer to Figure 2 , the feature fusion and decoding module includes a concatenated fuser and a Pyramid Mamba decoder; among them, the Pyramid Mamba decoder includes a concatenated dense spatial pyramid pooling module and a pyramid fusion module.
[0091] Refer to Figure 2 , in step S5, the feature fusion and decoding module processes the input features as follows:
[0092] The fuser integrates the multi-scale semantic features output by the multi-scale feature extraction module and the shallow extraction features output by the front-end feature extraction module to obtain fused features;
[0093] The dense spatial pyramid pooling module in the Pyramid Mamba decoder encodes the fused features through multiple pooling scales to capture finer multi-scale information, and uses the bilinear interpolation method to upsample the encoded features to match the high-level features, obtaining the upsampled features;
[0094] The pyramid fusion module in the Pyramid Mamba decoder introduces Mamba to aggregate pyramid semantic features, reduces semantic redundancy in the upsampled features through a unique selective filtering mechanism to effectively represent the core semantic information across scales, and further enhances the multi-scale feature representation through a convolutional feed-forward neural network, obtaining the differential feature image.
[0095] It should be noted that Pyramid Mamba is a further improvement and extension of the Mamba architecture. By introducing a pyramid structure, it further improves the model's ability and efficiency in processing long sequences. First, Pyramid Mamba adopts a pyramid-style hierarchical design. When processing sequences, as the hierarchy increases, the sequence length gradually shrinks, while the dimension of the features gradually increases, allowing the model to abstract and process sequence information at different scales, capturing both local detailed information and global context information. In addition, Pyramid Mamba inherits the characteristics of the Mamba based on the state space model SSM, and each hierarchical module uses SSM to process the sequence, enabling the model to effectively handle long-distance information interaction.
[0096] The feature fusion and decoding module provided in this embodiment can effectively aggregate multi-scale features of the input dual-temporal images, significantly reducing information redundancy while further enhancing the representation ability of multi-scale features.
[0097] As an optional embodiment, in order to determine whether there are abnormal changes in the dual-temporal images, this embodiment provides a differential discrimination module to discriminate the target state by identifying the multi-scale semantic feature differences in the changed regions of the dual-temporal images. As Figure 4 shown, it is a schematic structural diagram of the differential discrimination module provided in this embodiment, mainly including a first fully convolutional neural network, a second fully convolutional neural network, a third fully convolutional neural network, and a differentiator connected in series in sequence.
[0098] The following combines Figure 4 to further elaborate in detail on the method of using the differential discrimination module to perform differential processing on the differential feature image to obtain the abnormal change region image of the dual-temporal image in step S6:
[0099] The difference feature image of the dual-temporal image is successively passed through a first fully convolutional neural network, a second fully convolutional neural network, and a third fully convolutional neural network for feature restoration to obtain a restored feature image of the dual-temporal image;
[0100] The differentiator calculates the absolute value of the subtraction of the restored feature image of the dual-temporal image to obtain an abnormal change region image of the dual-temporal image.
[0101] Specifically, the differentiator calculates the absolute value of the subtraction of the restored feature image of the dual-temporal image, which can be expressed by the following formula:
[0102] ,
[0103] In the formula, is the abnormal change region image; and are the restored feature images of the dual-temporal image, , are the height, width, and number of channels respectively; is the classifier; is the softmax function, which is used to convert the output of the classifier into a probability distribution to more accurately discriminate the abnormal change region.
[0104] The differential discrimination module provided in this embodiment performs differential processing on two groups of feature maps from the feature fusion and decoding module. By performing element-wise absolute value calculation on these two groups of feature maps, it can effectively capture and distinguish multi-scale change information in the dual-temporal image and obtain an abnormal change region image.
[0105] It should be noted that in the abnormal change region image output by the differential discrimination module, the region where the pixel value is zero is the region where there is no abnormal change; while the region where the pixel value is non-zero is the region where there is an abnormal change. Therefore, in step S7, the region where the pixel value is non-zero is extracted and output to obtain the abnormal change region of the dual-temporal image.
[0106] The small target abnormal change detection method provided in this embodiment processes dual-temporal images in the same scene based on the combination of a Swin Transformer encoder and a Pyramid Mamba decoder. First, a front-end feature extraction module is used to finely capture shallow extraction features containing rich image details and local information. Then, on the one hand, the shallow extraction features are input into a multi-scale feature extraction module, and the Swin Transformer encoder is used to extract deep features from them, effectively retaining the low-frequency information of the image. On the other hand, the shallow extraction features are directly mapped into the Pyramid Mamba decoder to combine multi-scale semantic features, which can significantly reduce information redundancy while further enhancing the representation ability of multi-scale features. Finally, a differential discrimination module performs differencing on the dual-temporal difference feature image to effectively capture and distinguish multi-scale change information in the dual-temporal images in the same scene. This method not only maintains the accuracy of the model in detecting abnormal changes in large-scale targets but also further improves the model's ability to detect abnormal changes in small targets and local regions.
[0107] This flowchart only shows the logical order of the method described in this embodiment. On the premise of no conflict, in other possible embodiments of the present invention, the steps shown or described can be completed in a different order from that Figure 1 shown. The small target abnormal change detection method provided in this embodiment can be applied to a terminal and can be executed by a small target abnormal change detection device, which can be implemented in software and / or hardware, and the device can be integrated in the terminal, for example: any smart phone, tablet computer or computer device with communication functions.
[0108] Embodiment 2:
[0109] This embodiment provides a small target abnormal change detection method, which further details the method of training a small target abnormal change detection model in step S2 of Embodiment 1:
[0110] Step S201: Obtain multiple groups of dual-temporal sample images of different target devices in the same scene;
[0111] Step S202: Determine the abnormal change area in the dual-temporal sample images and perform mask marking;
[0112] Step S203: Construct a dual-temporal sample data set using the dual-temporal sample images marked with masks;
[0113] Step S204: Divide the dual-temporal sample data set into a training set, a validation set, and a test set;
[0114] Step S205: Train the pre-constructed small target anomaly change detection model using the training set, verify the trained small target anomaly change detection model using the validation set, test the verified small target anomaly change detection model using the test set, and use the small target anomaly change detection model with the optimal test result as the finally trained small target anomaly change detection model.
[0115] It should be noted that the dual-temporal sample images obtained in step S201 cover images of target devices at different scales taken from different perspectives at different distances.
[0116] In some embodiments, the training method of the small target anomaly change detection model further includes:
[0117] Combine the cross-entropy loss function and the mean squared error loss function to construct a joint loss function;
[0118] Optimize the small target anomaly change detection model with the goal of minimizing the value of the joint loss function.
[0119] Among them, the expression of the joint loss function is:
[0120] ,
[0121] In the formula, is the cross-entropy loss function; is the mean squared error loss function; is a hyperparameter, set to 0.7, used to balance the influence of the two losses.
[0122] The cross-entropy loss function is used to evaluate the difference between the true probability distribution and the predicted probability distribution, and its expression is:
[0123] ,
[0124] In the formula, represents the cross-entropy loss with label at position ; are the height, width of the image and the corresponding label during the model training optimization process respectively, ; and are the height and width of the dual-temporal original image.
[0125] The mean squared error loss function is used to measure the average variance between the predicted image and the true image, and its expression is:
[0126] ,
[0127] In the formula, is the The true value of a sample; For the prediction value of the th sample by the model; is the total number of samples.
[0128] The joint loss function provided in this embodiment can significantly improve the optimization efficiency of the model by jointly using the cross-entropy loss function and the mean squared error loss function, thereby improving the performance of the model in the detection task of micro-change regions.
[0129] In order to comprehensively evaluate the detection ability of the method proposed in the embodiments of the present invention for small target abnormal change regions in dual-temporal images, this application selects small targets in different scenarios and dual-temporal images taken from different perspectives of distance for verification.
[0130] As Figure 5 shown, it is the detection result of common small targets in the power scenario provided by this application. Although the features of these small targets are likely to be lost as the network depth increases in conventional change detection algorithms, the abnormal change regions can be accurately marked in the method proposed in the embodiments of the present invention. At the same time, as Figure 6 shown, it is the detection result of small targets weakened due to the change of shooting perspective in the power scenario provided by this application. It can be found that the method proposed in the embodiments of the present invention can accurately capture the abnormal change regions. In addition, as Figure 7 shown, it is the detection result of conventional-scale targets in the power scenario provided by this application. The method proposed in the embodiments of the present invention can still detect the abnormal change regions well.
[0131] In summary, for the appearance and disappearance of targets, the small target abnormal change detection method provided in the embodiments of the present invention can complete the detection by annotating the target region in the dual-temporal image; for targets with changes in position or state, the regions before and after the change of the target can be detected. In addition, this method is mainly applied to the abnormal change detection task of small targets and local regions in the power scenario, and can also provide a reference for the abnormal detection applications of small targets in other scenarios or fields, such as medical imaging, industrial inspection, and remote sensing image analysis.
[0132] Embodiment 3:
[0133] This embodiment also provides a computer device, which can be a server, and its internal structure diagram can be as Figure 8 shown. This computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface.
[0134] Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the data obtained and generated in the method for the robot to autonomously enter the packaging container. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements the method of the foregoing Embodiment 1 or Embodiment 2.
[0135] Those skilled in the art can understand that Figure 8 the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0136] The computer device provided in this embodiment can execute the small target abnormal change detection method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0137] Embodiment 4:
[0138] This embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps of the method described in Embodiment 1 or Embodiment 2.
[0139] The computer-readable storage medium provided in this embodiment can execute the small target abnormal change detection method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0140] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. In addition, terms such as "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features.
[0141] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce a means for implementing the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 a means for implementing the functions specified in one or more blocks or a plurality of blocks.
[0142] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 a means for implementing the functions specified in one or more blocks or a plurality of blocks.
[0143] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 a means for implementing the functions specified in one or more blocks or a plurality of blocks.
[0144] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope of the present invention as protected by the claims. All of these are within the protection scope of the present invention.
Claims
1. A method for detecting abnormal changes in small targets, characterized in that The method includes: Obtaining dual-temporal images of different target devices in the same scene to be detected; Inputting the dual-temporal images into a trained small-target anomaly change detection model for judgment, where the small-target anomaly change detection model includes a front-end feature extraction module, a multi-scale feature extraction module, a feature fusion and decoding module, and a differential discrimination module; Extracting shallow features of the dual-temporal images through a 3×3 convolutional layer in the front-end feature extraction module to generate shallow extraction features; Inputting the shallow extraction features into a multi-scale feature extraction module improved based on Swin Transformer to extract deep features and generate multi-scale semantic features; Integrating the multi-scale semantic features and the shallow extraction features through a fuser in the feature fusion and decoding module, and then performing upsampling on the integrated features through a Pyramid Mamba decoder to generate a difference feature image; Inputting the difference feature image into the differential discrimination module for differential processing to obtain an anomaly change region image of the dual-temporal images; Extracting the region with non-zero pixel values in the anomaly change region image and outputting coordinate information, thereby obtaining the anomaly change region of the dual-temporal images.
2. The small target abnormal change detection method according to claim 1, characterized in that The multi-scale feature extraction module improved based on Swin Transformer includes a patch partitioning module, a first-stage structure, a second-stage structure, and a third-stage structure connected in series in sequence; The method of inputting the shallow extraction features into a multi-scale feature extraction module improved based on Swin Transformer to extract deep features and generate multi-scale semantic features includes: Inputting the shallow extraction features of the dual-temporal images into the multi-scale feature extraction module, and performing image block division through the patch partitioning module to obtain non-overlapping image block features; Sequentially inputting each image block feature into the first-stage structure, the second-stage structure, and the third-stage structure for processing to respectively obtain first-stage semantic features, second-stage semantic features, and multi-scale semantic features.
3. The small target abnormal change detection method according to claim 2, wherein, The first-stage structure includes a linear embedding module, a first Swin transformation module, and a second Swin transformation module connected in series in sequence, the second-stage structure includes a first patch merging module, a third Swin transformation module, and a fourth Swin transformation module connected in series in sequence, and the third-stage structure includes a second patch merging module, a fifth Swin transformation module, a sixth Swin transformation module, a seventh Swin transformation module, an eighth Swin transformation module, a ninth Swin transformation module, and a tenth Swin transformation module connected in series in sequence; The first-stage structure, the second-stage structure, and the third-stage structure perform the following processing steps on the input features: The linear embedding module in the first-stage structure maps each input image block feature to a high-dimensional embedding space, and then sequentially passes through the first Swin transformation module and the second Swin transformation module to implement image feature modeling through a windowed multi-head self-attention mechanism and a shifted window multi-head self-attention mechanism to obtain the first-stage semantic features; The first patch merging module in the second-stage structure performs downsampling and merging operations on the first-stage semantic features, and then successively passes through the third Swin transformation module and the fourth Swin transformation module to implement image feature modeling through the windowed multi-head self-attention mechanism and the shifted window multi-head self-attention mechanism, obtaining the second-stage semantic features; The second patch merging module in the third-stage structure performs downsampling and merging operations on the second-stage semantic features, and then successively passes through the fifth Swin transformation module to the tenth Swin transformation module to implement image feature modeling through the windowed multi-head self-attention mechanism and the shifted window multi-head self-attention mechanism, obtaining the multi-scale semantic features.
4. The small target abnormal change detection method according to claim 1, characterized in that The feature fusion and decoding module includes a concatenated fuser and a Pyramid Mamba decoder; wherein, the Pyramid Mamba decoder includes a concatenated dense spatial pyramid pooling module and a pyramid fusion module; The method of integrating the multi-scale semantic features and the shallow extraction features through the fuser in the feature fusion and decoding module, and then performing upsampling on the integrated features through the Pyramid Mamba decoder to generate a differential feature image includes: Fusing the multi-scale semantic features and the shallow extraction features of the dual-temporal images through the fuser to obtain fused features; The dense spatial pyramid pooling module encodes the fused features using multiple pooling scales, and performs upsampling on the encoded features using bilinear interpolation to match high-level features, obtaining upsampled features; The pyramid fusion module reduces the semantic redundancy in the upsampled features through a selective filtering mechanism, and further enhances the multi-scale feature representation through a convolutional feed-forward neural network, obtaining a differential feature image.
5. The small target abnormal change detection method according to claim 1, wherein The differential discrimination module includes a first fully convolutional neural network, a second fully convolutional neural network, a third fully convolutional neural network, and a differentiator connected in series in sequence; The method of inputting the differential feature image into the differential discrimination module for differential processing to obtain the abnormal change region image of the dual-temporal image includes: Successively passing the differential feature image of the dual-temporal image through the first fully convolutional neural network, the second fully convolutional neural network, and the third fully convolutional neural network for feature restoration, obtaining the restored feature image of the dual-temporal image; Calculating the absolute value of the subtraction of the restored feature image of the dual-temporal image through the differentiator to obtain the abnormal change region image of the dual-temporal image.
6. The small target abnormal change detection method according to claim 5, wherein Calculating the absolute value of the subtraction of the restored feature image of the dual-temporal image through the differentiator can be represented by the following formula: , In the formula, is the image of the abnormal change area; and are the restored feature images of the dual-temporal images, , are the height, width, and number of channels respectively; is the classifier; is the softmax function, which is used to convert the output of the classifier into a probability distribution to more accurately identify the abnormal change area.
7. The small target abnormal change detection method according to claim 1, wherein The training method of the small target abnormal change detection model includes: Obtaining multiple groups of dual-temporal sample images of different target devices in the same scene; Determining the abnormal change regions in the dual-temporal sample images and performing mask marking; Constructing a dual-temporal sample data set using the dual-temporal sample images marked with the mask; Dividing the dual-temporal sample data set into a training set, a validation set, and a test set; The pre-constructed small target anomaly change detection model is trained using the training set, the trained small target anomaly change detection model is verified using the validation set, and the verified small target anomaly change detection model is tested using the test set. The small target anomaly change detection model with the optimal test result is used as the finally trained small target anomaly change detection model.
8. The small target abnormal change detection method according to claim 7, wherein The training method of the small target anomaly change detection model further includes: Combining the cross-entropy loss function and the mean squared error loss function to construct a joint loss function; Optimizing the small target anomaly change detection model with the goal of minimizing the value of the joint loss function; Among them, the expression of the joint loss function is: , In the formula, is the cross-entropy loss function; is the mean squared error loss function; is a hyperparameter, set to 0.7, used to balance the influence of the two losses; The expression of the cross-entropy loss function is: , In the formula, represents the cross-entropy loss with the label at the position ; are respectively the height, width of the image and the corresponding label during the model training optimization process, ; and are the height and width of the dual-temporal original image; The expression of the mean squared error loss function is: , Wherein, is the true value of the th sample; is the predicted value of the model for the th sample; is the total number of samples.
9. A computer device, characterized in that, Including a storage medium and a processor; The storage medium is used to store instructions; The processor is used to operate according to the instructions to execute the steps of the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Remote sensing image semantic change detection method and device based on Mamba model
CN119580258A