Image difference detection method, device, equipment, storage medium and program product

By using a feature encoding and decoding network of a difference-aware model, global and local features of images are extracted and fused, solving the problem of inaccurate image difference detection in existing technologies and achieving higher-precision image difference detection.

CN116229271BActive Publication Date: 2026-05-05INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INDUSTRIAL AND COMMERCIAL BANK OF CHINA
Filing Date
2023-03-06
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing image difference detection methods typically analyze pixel differences between different images to be detected directly, resulting in inaccurate image difference detection results and low detection accuracy.

Method used

Global and local features of the first and second images to be detected are extracted by the feature encoding network of the difference perception model, and difference perception processing is performed. The target difference region is decoded by the feature decoding network, and feature fusion and difference feature extraction are performed by multi-layer sub-encoding network and difference perception network.

Benefits of technology

It improves the accuracy of image difference detection by extracting and fusing multi-dimensional features, ensuring the accuracy and precision of the detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116229271B_ABST
    Figure CN116229271B_ABST
Patent Text Reader

Abstract

This application relates to an image difference detection method, apparatus, device, storage medium, and program product. It pertains to the field of artificial intelligence technology. The method includes: extracting a first global local feature corresponding to a first image to be detected and a second global local feature corresponding to a second image to be detected through a feature encoding network of a difference-aware model; wherein the first image to be detected and the second image to be detected are images of the same region acquired at different times; performing difference-aware processing on the first and second global local features to determine the difference-aware features between the first and second images to be detected; and decoding the difference-aware features through a feature decoding network of the difference-aware model to obtain the target difference region between the first and second images to be detected. This method can improve the accuracy of image difference detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to an image difference detection method, apparatus, device, storage medium, and program product. Background Technology

[0002] Image difference detection (e.g., remote sensing image difference detection) aims to analyze and detect changes in land features over time by examining images of the same area acquired at different times. Image difference detection has significant importance and value in natural disaster early warning, urban planning, engineering progress surveys, and military activities.

[0003] However, current image difference detection methods typically analyze pixel differences between different images when performing image difference detection. This can lead to insufficient utilization of image features, resulting in inaccurate image difference detection results and low detection accuracy, which urgently need to be addressed. Summary of the Invention

[0004] Therefore, it is necessary to provide an image difference detection method, apparatus, device, storage medium, and program product that can improve the accuracy of image difference detection in response to the above-mentioned technical problems.

[0005] Firstly, this application provides an image difference detection method. The method includes:

[0006] The first global local feature corresponding to the first image to be detected and the second global local feature corresponding to the second image to be detected are extracted through the feature encoding network of the difference perception model; wherein the first image to be detected and the second image to be detected are images of the same area collected at different times;

[0007] Perform difference sensing processing on the first global local features and the second global local features to determine the difference sensing features between the first image to be detected and the second image to be detected.

[0008] The feature decoding network of the difference perception model is used to decode the difference perception features to obtain the target difference region between the first image to be detected and the second image to be detected.

[0009] In one embodiment, the feature encoding network includes at least two end-to-end connected sub-encoding networks; the feature encoding network of the difference-aware model extracts a first global local feature corresponding to a first image to be detected and a second global local feature corresponding to a second image to be detected, including:

[0010] By encoding the input data of each sub-coding network, the first global local features extracted by the sub-coding network for the first image to be detected and the second global local features extracted by the sub-coding network for the second image to be detected are obtained.

[0011] The input data for the first sub-coding network consists of the first image to be detected and the second image to be detected; the input data for each other sub-coding network consists of the first global local feature and the second global local feature output by the previous sub-coding network.

[0012] In one embodiment, each sub-coding network includes: a global feature coding network, a local feature coding network, and a feature fusion network; the input data of each sub-coding network is encoded to obtain the first global and local features extracted by the sub-coding network for the first image to be detected, including:

[0013] The input data of each sub-coding network is parsed and processed by the global feature encoding network of each sub-coding network to obtain the first global feature corresponding to that sub-coding network.

[0014] The input data of the sub-coding network is parsed and processed by the local feature encoding network of the sub-coding network to obtain the first local feature corresponding to the sub-coding network;

[0015] The feature fusion network of the sub-coding network is used to fuse the first global feature and the first local feature corresponding to the sub-coding network to obtain the first global and local features extracted by the sub-coding network from the first image to be detected.

[0016] In one embodiment, the input data of each sub-coding network is parsed and processed by the global feature encoding network of each sub-coding network to obtain the first global feature corresponding to that sub-coding network, including:

[0017] By encoding the global features of each sub-encoding network, and based on channel attention and spatial attention mechanisms, the input data of the sub-encoding network is parsed and processed to obtain the first global feature corresponding to the sub-encoding network.

[0018] In one embodiment, the difference sensing model further includes: a difference sensing network; the difference sensing network includes at least two sub-difference sensing networks connected end-to-end, and each sub-difference sensing network corresponds one-to-one with each sub-coding network in the feature coding network;

[0019] Accordingly, difference-sensing processing is performed on the first global local features and the second global local features to determine the difference-sensing features between the first image to be detected and the second image to be detected, including:

[0020] By using each sub-difference sensing network, the sub-difference features are determined based on the input data corresponding to that sub-difference sensing network.

[0021] The sub-difference features determined by the last sub-difference sensing network are fused with the first global local features and the second global local features extracted by the sub-coding network corresponding to the last sub-difference sensing network to obtain the difference sensing features between the first image to be detected and the second image to be detected.

[0022] The input data for the first sub-differential sensing network consists of the input data for the first sub-coding network, as well as the first global local feature and the second global local feature extracted by the first sub-coding network. The input data for each other sub-differential sensing network consists of the input data of the sub-coding network corresponding to the sub-differential sensing network, the first global local feature and the second global local feature extracted by the sub-coding network corresponding to the sub-differential sensing network, and the sub-differential feature determined by the previous sub-differential sensing network of the sub-differential sensing network.

[0023] In one embodiment, sub-difference features are determined based on the input data corresponding to each sub-difference sensing network, including:

[0024] For each sub-difference sensing network, an initial difference feature is determined based on the first and second global local features extracted by the sub-encoding network corresponding to that sub-difference sensing network. The initial difference feature, the sub-difference feature determined by the previous sub-difference sensing network, and the input data of the sub-encoding network corresponding to that sub-difference sensing network are then fused to obtain the sub-difference feature determined by that sub-difference sensing network.

[0025] In one embodiment, the feature decoding network includes at least two end-to-end connected sub-decoding networks;

[0026] The feature decoding network of the difference-aware model decodes the difference-aware features to obtain the target difference region between the first and second images to be detected, including:

[0027] Each sub-decoding network decodes its input data to obtain the sub-difference region determined by that sub-decoding network.

[0028] The sub-difference region determined by the last sub-decoding network is used as the target difference region between the first image to be detected and the second image to be detected.

[0029] The input data for the first sub-decoding network is the difference-aware feature, and the input data for each other sub-decoding network is the sub-difference region determined by the previous sub-decoding network.

[0030] In one embodiment, the training process of the difference-aware model includes:

[0031] The sample image pairs are input into the difference perception model to obtain the sample difference regions and difference perception features corresponding to the sample images predicted by the difference perception model; wherein, the sample image pairs contain sample images of the same region collected at different times.

[0032] The cross-entropy loss value is determined based on the true difference region and the sample difference region corresponding to the sample image.

[0033] The contrast loss value is determined based on the differential perception features and differential feature labels corresponding to the sample images;

[0034] The difference-aware model is trained based on the cross-entropy loss value and the contrast loss value.

[0035] Secondly, this application also provides an image difference detection device. The device includes:

[0036] The feature extraction module is used to extract the first global local features corresponding to the first image to be detected and the second global local features corresponding to the second image to be detected through the feature encoding network of the difference perception model; wherein the first image to be detected and the second image to be detected are images of the same region collected at different times;

[0037] The difference perception module is used to perform difference perception processing on the first global local features and the second global local features to determine the difference perception features between the first image to be detected and the second image to be detected.

[0038] The difference determination module is used to decode the difference-aware features through the feature decoding network of the difference-aware model to obtain the target difference region between the first image to be detected and the second image to be detected.

[0039] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:

[0040] The first global local feature corresponding to the first image to be detected and the second global local feature corresponding to the second image to be detected are extracted through the feature encoding network of the difference perception model; wherein the first image to be detected and the second image to be detected are images of the same area collected at different times;

[0041] Perform difference sensing processing on the first global local features and the second global local features to determine the difference sensing features between the first image to be detected and the second image to be detected.

[0042] The feature decoding network of the difference perception model is used to decode the difference perception features to obtain the target difference region between the first image to be detected and the second image to be detected.

[0043] Fourthly, this application also provides a computer-readable storage medium. This computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:

[0044] The first global local feature corresponding to the first image to be detected and the second global local feature corresponding to the second image to be detected are extracted through the feature encoding network of the difference perception model; wherein the first image to be detected and the second image to be detected are images of the same area collected at different times;

[0045] Perform difference sensing processing on the first global local features and the second global local features to determine the difference sensing features between the first image to be detected and the second image to be detected.

[0046] The feature decoding network of the difference perception model is used to decode the difference perception features to obtain the target difference region between the first image to be detected and the second image to be detected.

[0047] Fifthly, this application also provides a computer program product. This computer program product includes a computer program that, when executed by a processor, performs the following steps:

[0048] The first global local feature corresponding to the first image to be detected and the second global local feature corresponding to the second image to be detected are extracted through the feature encoding network of the difference perception model; wherein the first image to be detected and the second image to be detected are images of the same area collected at different times;

[0049] Perform difference sensing processing on the first global local features and the second global local features to determine the difference sensing features between the first image to be detected and the second image to be detected.

[0050] The feature decoding network of the difference perception model is used to decode the difference perception features to obtain the target difference region between the first image to be detected and the second image to be detected.

[0051] The aforementioned image difference detection method, apparatus, device, storage medium, and program product extract first and second global local features from a first image to be detected and a second image to be detected. Then, difference-sensing processing is performed on the first and second global local features to obtain difference-sensing features corresponding to the first and second images to be detected. Finally, a feature decoding network is used to decode the difference-sensing features corresponding to the first and second images to obtain the target difference region between the first and second images to be detected. Since the first global local feature is obtained after deep feature extraction of the first image to be detected, it more accurately represents the features corresponding to the first image to be detected. Furthermore, the first global local feature integrates global and local features corresponding to the first image to be detected, thus it can also reflect the features of the first image to be detected from multiple dimensions. Correspondingly, the second global local feature corresponding to the second image to be detected will also more accurately reflect the features of the second image to be detected from multiple dimensions. This process makes the difference-sensing features obtained based on the first and second global local features more accurate. Furthermore, after decoding the difference-sensing features, the target difference regions between the first and second images to be detected are more accurate. In other words, the entire process can improve the accuracy of image difference detection. Attached Figure Description

[0052] Figure 1 This is an application environment diagram of an image difference detection method provided in this embodiment;

[0053] Figure 2 This is a flowchart illustrating the first image difference detection method provided in this embodiment;

[0054] Figure 3 This is an internal structure diagram of the first difference-perceiving model provided in this embodiment;

[0055] Figure 4 This is an internal structure diagram of the second difference-perceiving model provided in this embodiment;

[0056] Figure 5 This is an internal structure diagram of a sub-coding network provided in this embodiment;

[0057] Figure 6 This embodiment provides a flowchart for extracting the first global local features;

[0058] Figure 7 This is an internal structure diagram of a channel attention mechanism provided in this embodiment;

[0059] Figure 8This is an internal structural diagram of a spatial attention mechanism provided in this embodiment;

[0060] Figure 9 This embodiment provides a flowchart illustrating the process of determining the difference-perceived features between a first image to be detected and a second image to be detected.

[0061] Figure 10 This is a schematic diagram of the structure of a differential sensing network outputting differential sensing features, provided in this embodiment.

[0062] Figure 11 This is a schematic diagram illustrating a process for training a difference-aware model, as provided in this embodiment.

[0063] Figure 12 This is a structural block diagram of the first image difference detection device provided in this embodiment;

[0064] Figure 13 This is a structural block diagram of the second image difference detection device provided in this embodiment;

[0065] Figure 14 This is a structural block diagram of the third image difference detection device provided in this embodiment;

[0066] Figure 15 This is a structural block diagram of the fourth image difference detection device provided in this embodiment;

[0067] Figure 16 This is a structural block diagram of the fifth image difference detection device provided in this embodiment;

[0068] Figure 17 This is an internal structural diagram of a computer device provided in this embodiment. Detailed Implementation

[0069] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0070] The image difference detection method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, in one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows. Figure 1As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores relevant data for image difference detection. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements an image difference detection method.

[0071] In one embodiment, such as Figure 2 As shown, an image difference detection method is provided, which can be applied to, for example... Figure 1 The following steps are illustrated using the computer shown as an example:

[0072] S201, using the feature encoding network of the difference perception model, extract the first global local features corresponding to the first image to be detected and the second global local features corresponding to the second image to be detected.

[0073] In this model, the first image to be detected and the second image to be detected are images of the same area collected at different times. For example, both the first image to be detected and the second image to be detected can be remote sensing images. The difference-aware model can be a pre-trained model capable of detecting image differences between two input images. The feature encoding network of the difference-aware model can be a network that extracts features from the input images. It is understood that the first global local feature corresponding to the first image to be detected can be a feature image, and correspondingly, the second global local feature corresponding to the second image to be detected can also be a feature image.

[0074] Optionally, in this embodiment, the difference perception model may contain only one feature encoding network. First, the feature encoding network extracts features from the first image to be detected, outputting the first global local feature corresponding to the first image. Then, the same feature encoding network extracts features from the second image to be detected, outputting the second global local feature corresponding to the second image. Alternatively, the difference perception model may contain two feature encoding networks (e.g., a first feature encoding network and a second feature encoding network). The first image to be detected is input into the first feature encoding network, which extracts features from the first image and outputs the first global local feature corresponding to the first image. The second image to be detected is input into the second feature encoding network, which extracts features from the second image and outputs the second global local feature corresponding to the second image. This is not limited.

[0075] It should be noted that if the difference-aware model contains two feature encoding networks, then these two feature encoding networks can be Siamese neural networks. That is, the two feature encoding networks are composed of two neural networks with the same structure and shared weights, and can perform the same operations on the input data.

[0076] S202, perform difference perception processing on the first global local features and the second global local features to determine the difference perception features between the first image to be detected and the second image to be detected.

[0077] The difference-aware feature can be a feature obtained by processing the first global local feature corresponding to the first image to be detected and the second global local feature corresponding to the second image to be detected, which is used to characterize the difference between the first global local feature and the second global local feature. It can be understood that when the first global local feature corresponding to the first image to be detected is a feature image and the second global local feature corresponding to the second image to be detected is also a feature image, the difference-aware feature can be a difference-aware image.

[0078] Optionally, in this embodiment, there are many ways to perform difference-aware processing on the first global local feature and the second global local feature, and there is no limitation on this. One possible approach is to calculate, for example, by subtraction, the first global local feature and the second global local feature based on a preset difference-aware algorithm to obtain the difference between the first global local feature and the second global local feature. This difference is then characterized, and the characteristic-processed difference is used as the difference-aware feature corresponding to the first and second images to be detected. Another possible approach is that the difference-aware model may include a pre-trained difference-aware network capable of performing difference-aware processing on the first global local feature and the second global local feature. The first global local feature and the second global local feature are input into the difference-aware network, which processes the first global local feature and the second global local feature and outputs the difference-aware feature corresponding to the first global local feature and the second global local feature.

[0079] S203, through the feature decoding network of the difference perception model, the difference perception features are decoded to obtain the target difference region between the first image to be detected and the second image to be detected.

[0080] The feature decoding network can be a network used to decode various features and map them back to the original input size. For example, in this embodiment, the feature decoding network can be a network that decodes difference-aware features and processes them into an image of the same size as the first and second images to be detected. The target difference region can be a region where there is a difference between the first and second images to be detected. In this embodiment, the difference-aware features corresponding to the first and second global local features are input into the feature decoding network of the difference-aware model. The feature decoding network decodes the input difference-aware features and outputs the target difference region between the first and second images to be detected.

[0081] In the aforementioned image difference detection method, global and local features are extracted from the first and second images to be detected, resulting in first and second global and local features. These features are then subjected to difference-aware processing to obtain difference-aware features corresponding to the first and second images. Finally, a feature decoding network is used to decode these difference-aware features, yielding the target difference region between the first and second images. Since the first global and local features are obtained after deep feature extraction from the first image, they more accurately represent the features of the first image. Furthermore, the first global and local features integrate global and local features, allowing them to represent the features of the first image from multiple dimensions. Consequently, the second global and local features corresponding to the second image are also more accurate, representing the features of the second image from multiple dimensions. This makes the difference-aware features obtained from the processing of the first and second global and local features more accurate, and further, it makes the target difference region between the first and second images more accurate after decoding the difference-aware features. In other words, the entire process can improve the accuracy of image difference detection.

[0082] Furthermore, taking a difference-sensing model that includes a first feature encoding network and a second feature encoding network, and that includes a difference-sensing network for difference-sensing processing, as an example, the aforementioned difference-sensing model can be as follows: Figure 3The difference sensing model 1 shown in the figure includes a first feature encoding network 10 for feature extraction of a first image to be detected; a second feature encoding network 11 for feature extraction of a second image to be detected; a difference sensing network 12 for difference sensing processing of a first global local feature and a second global local feature; and a feature decoding network 13 for decoding the difference sensing features. For example, taking both the first and second images to be detected as remote sensing images, the first and second images to be detected are respectively input into the first feature encoding network 10 and the second feature encoding network 11. The first and second feature encoding networks 10 and 11 extract features from the first and second images to be detected, respectively. Then, the first feature encoding network 10 outputs the first global local feature corresponding to the first image to be detected and inputs the first global local feature into the difference sensing network 12. The second feature encoding network 11 outputs the second global local feature corresponding to the second image to be detected and inputs the second global local feature into the difference sensing network 12. The difference perception network 12 performs difference perception processing on the received first global local features and second global local features, outputs the difference perception features corresponding to the first image to be detected and the second image to be detected, and inputs the difference perception features into the feature decoding network 13. The feature decoding network 13 decodes the received difference perception features, outputs the target difference region corresponding to the first image to be detected and the second image to be detected, and determines the image difference between the first image to be detected and the second image to be detected based on the output target difference region.

[0083] Furthermore, in the process of extracting the first global local features corresponding to the first image to be detected and the second global local features corresponding to the second image to be detected using the feature encoding network of the difference perception model, in order to make the extracted first and second global local features more complete, the feature encoding network may include at least two contiguous sub-encoding networks. Each sub-encoding network encodes its input data to obtain the first global local features extracted for the first image to be detected and the second global local features extracted for the second image to be detected. The input data of the first sub-encoding network consists of the first and second images to be detected; the input data of the other sub-encoding networks consists of the first and second global local features output by the previous sub-encoding network. The sub-encoding network can be a network in the feature encoding network that performs feature encoding on the image. Taking a feature encoding network including two contiguous sub-encoding networks as an example, it can be as follows: Figure 4The differential sensing model 1 shown in the figure includes a first feature encoding network 10 comprising a first sub-encoding network 100 and a second sub-encoding network 101; and a second feature encoding network 11 comprising a first sub-encoding network 110 and a second sub-encoding network 111. In this embodiment, after the first image to be detected is input into the first feature encoding network 10, features are extracted from the first image to be detected through the first sub-encoding network (i.e., the first sub-encoding network 100) to obtain the first global local features extracted by the first sub-encoding network. The first sub-encoding network outputs its first global local features to the next sub-encoding network (i.e., the second sub-encoding network 101). The next sub-encoding network further extracts features based on the received first global local features to obtain the first global local features extracted by the second sub-encoding network 101. The process of feature extraction for the second image to be detected is the same as that for the first image to be detected, and will not be described in detail here.

[0084] It should be noted that each sub-encoding network that performs feature extraction on the first image to be detected will output a set of first global local features. In this embodiment, all the first global local features will be output to the difference perception network for difference perception.

[0085] In this embodiment, feature extraction of the first and second images to be detected is completed through multiple sub-coding networks, making the output first and second global local features more comprehensive, thus providing a foundation for improving the accuracy of image difference detection.

[0086] Furthermore, to make the first global-local features and the second global-local features output by the feature encoding network more accurate, in one embodiment, each sub-encoding network includes: a global feature encoding network, a local feature encoding network, and a feature fusion network. Each sub-encoding network can be as follows: Figure 5 As shown in the figure, the first sub-encoding network 100 includes a global feature encoding network 1001 for global feature extraction of the image to be detected, and a local feature encoding network 1002 for local feature extraction of the image to be detected; wherein, the global feature encoding network 1001 includes a normalization layer 10010, a channel attention mechanism 10011, and a spatial attention mechanism 10012. The first global feature output by the global feature encoding network 1001 and the first local feature output by the local feature encoding network are fused to obtain the fused result as the first global-local feature corresponding to the first sub-encoding network. Figure 5 Taking the first sub-coding network 100 as an example, the input data of each sub-coding network is encoded to obtain the first global local features extracted by the sub-coding network from the first image to be detected. This will be described in detail below. Figure 6 The steps shown may include the following:

[0087] S601, the input data of each sub-coding network is parsed and processed by the global feature encoding network of each sub-coding network to obtain the first global feature corresponding to the sub-coding network.

[0088] The input data consists of the data received by each sub-coding network. For example, the input data for the first sub-coding network is the first image to be detected and the second image to be detected. The input data for each of the other sub-coding networks is the first global local feature and the second global local feature output by the previous sub-coding network. Taking the first sub-coding network as an example, the first image to be detected is input into the first sub-coding network. The global feature encoding network in the first sub-coding network extracts global features from the received first image to be detected and processes them to obtain the first global feature.

[0089] Optionally, the global feature encoding network may include a channel attention mechanism and a spatial attention mechanism, used to extract the channel features and spatial features corresponding to the first image to be detected, respectively. In this embodiment, the input data of each sub-encoding network can be parsed and processed based on the channel attention mechanism and the spatial attention mechanism to obtain the first global feature corresponding to the sub-encoding network. For the input data, after tiling it into a one-dimensional token sequence, a global context is modeled between the tokens. For the modeled global context, it is first normalized by layer normalization. Then, the input data after normalization is first processed by the channel attention mechanism, and the processing result is processed by the spatial attention mechanism to extract useful information from the features from different dimensions. For example, taking the channel attention mechanism in the first sub-encoding network as an example, it can be as follows: Figure 7 As shown in the figure, the channel attention mechanism 10011 includes a max pooling layer, an average pooling layer, a multilayer perceptron layer, and an activation layer. For the first image to be detected after normalization, the average pooling layer and the max pooling layer are used to process it to generate two aggregated vectors of size C×1×1. Then, the two aggregated vectors are input into a weight-shared multilayer perceptron layer with a channel reduction ratio r to assign weights to the two aggregated vectors. The two weighted vectors are then output to the activation layer. The two weighted vectors are processed by the activation function to obtain the channel feature map. The channel feature map is multiplied with the first image to be detected to obtain the channel attention feature. In addition, the channel attention mechanism can be calculated by the following formula (1):

[0090] M c (F)=σ(MLP r (AvgPool(F)))+MLP r (MaxPool(F)) (1)

[0091] In the formula, F represents the input data; M c (F) represents the channel attention feature; σ represents the activation function; MLP represents the multilayer perceptron; r represents the reduction weight; AvgPool(F) represents the average pooling feature; MaxPool(F) represents the maximum average pooling feature.

[0092] The channel attention features are vectorized to obtain a feature matrix, which is then input into the spatial attention mechanism. For example, the spatial attention mechanism can be implemented as follows: Figure 8 As shown in the figure, the spatial attention mechanism 10012 includes a max pooling layer, an average pooling layer, a convolutional layer, and an activation layer. For the channel attention features after input normalization, the spatial attention mechanism also uses average pooling and max pooling for the first step, compressing the input channel attention features into two 1×H×W matrices. The two 1×H×W matrices are concatenated and input into a convolutional layer with a kernel size of k for convolution processing. The two 1×H×W matrices are processed into a matrix of the same size as the spatial attention features, and the activation layer is used to process the matrix to obtain the final spatial attention features. In addition, the spatial attention mechanism can be calculated using the following formula (2):

[0093] M s (F')=σ(f (k×k) (AvgPool(F'));MaxPool(F')) (2)

[0094] In the formula, M s (F') represents the spatial attention feature; F' represents the channel attention feature after normalization; f (k×k) This is used for processing convolutional layers with a kernel size of k; AvgPool(F') represents the average pooling feature; MaxPool(F') represents the maximum average pooling feature; σ represents the activation function.

[0095] In this embodiment, the input data is parsed and processed through channel attention mechanism and spatial attention mechanism, and the output channel attention feature and spatial attention feature are fused together as the first global feature, which can make the first global feature more accurate, thereby providing a guarantee for improving the accuracy of image difference detection.

[0096] S602, the input data of the sub-coding network is parsed and processed by the local feature encoding network of the sub-coding network to obtain the first local feature corresponding to the sub-coding network.

[0097] Specifically, in this embodiment, the first image to be detected is input to the local feature encoding network in the first sub-encoding network. The local feature encoding network in the first sub-encoding network extracts local features from the received first image to be detected, and processes them to obtain the first local feature. Optionally, the local feature encoding network in the first sub-encoding network may contain at least two contiguous basic block modules for extracting local features from the first image to be detected. The first basic block module normalizes the input first image to be detected before extracting features. The first basic block module inputs the extracted initial local features to the next basic block module. The next basic block module normalizes the initial local features before extracting features, and so on, until the last basic block module outputs the initial local features. The initial local features output by the last basic block module are used as the first local feature corresponding to the sub-encoding network. It should be noted that the feature extraction process of the basic block module is performed through a channel attention mechanism. The feature extraction process of the channel attention mechanism is the same as the process mentioned in S601, and will not be described again here.

[0098] S603, through the feature fusion network of the sub-coding network, the first global feature and the first local feature corresponding to the sub-coding network are fused to obtain the first global and local features extracted by the sub-coding network for the first image to be detected.

[0099] Specifically, in this embodiment, the first global feature extracted by the global feature coding network and the first local feature extracted by the local feature coding network are fused together to fuse the first global and local features mentioned by the sub-coding network for the first image to be detected.

[0100] In the above embodiments, the first global local feature is obtained by integrating the first global feature and the first local feature. Therefore, the first global local feature can characterize the features of the image to be detected in more dimensions, providing a guarantee for the subsequent difference detection process of the first global local feature and the second global local feature.

[0101] Furthermore, the aforementioned difference-sensing model also includes a difference-sensing network for difference sensing, and this difference-sensing network comprises at least two end-to-end connected sub-difference-sensing networks, with each sub-difference-sensing network corresponding one-to-one with each sub-encoding network in the feature encoding network. Accordingly, the process of determining the difference-sensing features between the first and second images to be detected is described in detail, such as... Figure 9 As shown, it may include the following steps:

[0102] S901, through each sub-difference sensing network, determine the sub-difference features based on the input data corresponding to that sub-difference sensing network.

[0103] The input data of the first sub-differential sensing network consists of the input data of the first sub-encoding network, and the first and second global local features extracted by the first sub-encoding network. The input data of each other sub-differential sensing network consists of the input data of the corresponding sub-encoding network, the first and second global local features extracted by that sub-differential sensing network, and the sub-differential features determined by the previous sub-differential sensing network. Optionally, in this embodiment, each sub-differential sensing network can determine an initial difference feature based on the first and second global local features extracted by its corresponding sub-encoding network. The initial difference feature, the sub-differential features determined by the previous sub-differential sensing network, and the input data of the corresponding sub-encoding network are then fused to obtain the sub-differential feature determined by that sub-differential sensing network. The initial difference feature can be the difference between the first and second global local features, used to characterize the image difference between the first and second images to be detected. By fusing the initial difference features, the sub-difference features determined by the previous sub-difference sensing network of the sub-difference sensing network, and the input data of the sub-coding network corresponding to the sub-difference sensing network, the sub-difference features become more comprehensive and accurate.

[0104] For example, such as Figure 10 As shown, the difference sensing network 12 includes a first sub-difference sensing network 120 and a second sub-difference sensing network 121. First sub-coding networks 100 and 110 correspond to the first sub-difference sensing network 120, and second sub-coding networks 101 and 111 correspond to the second sub-difference sensing network 121. Taking the first sub-difference sensing network 120 as an example, the first image to be detected, the second image to be detected, the first global local feature extracted by the first sub-coding network 100, and the second global local feature extracted by the first sub-coding network 110 are input into the first sub-difference sensing network 120. The first sub-difference sensing network 120 performs difference sensing processing on the received data and outputs the sub-difference features corresponding to the first sub-difference sensing network 120. Optionally, the first image to be detected and the second image to be detected can be subtracted to obtain initial difference features. The initial difference features, the first global local feature, and the second global local feature are then fused to obtain the sub-difference features corresponding to the first sub-difference sensing network.

[0105] In addition, taking the second sub-difference sensing network 121 as an example, the sub-difference features output by the first sub-difference sensing network 120, the first global local features extracted by the second sub-coding network 101 and the second global local features extracted by the second sub-coding network 111, the first global local features extracted by the first sub-coding network 100 and the second global local features extracted by the first sub-coding network 110 are fused to obtain the sub-difference features corresponding to the second sub-difference sensing network.

[0106] S902, the sub-difference features determined by the last sub-difference sensing network are fused with the first global local features and the second global local features extracted by the sub-coding network corresponding to the last sub-difference sensing network to obtain the difference sensing features between the first image to be detected and the second image to be detected.

[0107] Specifically, such as Figure 10 As shown, in this embodiment, the sub-difference feature determined by the last sub-difference perception network is not directly used as the difference perception feature between the first image to be detected and the second image to be detected. Instead, the sub-difference feature determined by the last sub-difference perception network is fused with the first global local feature output by the first feature coding network 10 (i.e., the second sub-coding network 101) and the second global local feature output by the second feature coding network 11 (i.e., the second sub-coding network 111), and the fusion result is used as the difference perception feature.

[0108] In the above embodiments, the image differences between the first image to be detected and the second image to be detected are perceived through multiple sub-difference perception networks to obtain sub-difference features. Instead of directly using the sub-difference features as difference perception features, the sub-difference features are fused with the first global local features and the second global local features, and the fusion result is used as the difference perception features. Therefore, the difference perception features output by this embodiment fuse the features of the first image to be detected and the second image to be detected from multiple aspects, thereby making the difference perception features more complete.

[0109] It should be noted that, in order to decode the output of each sub-difference sensing network and thus make the target difference region more accurate after decoding, in one embodiment, the feature decoding network includes at least two end-to-end connected sub-decoding networks. Accordingly, decoding the difference-aware features includes decoding the input data of each sub-decoding network to obtain the sub-difference region determined by that sub-decoding network; and using the sub-difference region determined by the last sub-decoding network as the target difference region between the first image to be detected and the second image to be detected; wherein, the input data of the first sub-decoding network is the difference-aware feature, and the input data of each of the other sub-decoding networks is the sub-difference region determined by the previous sub-decoding network. Taking a feature decoding network comprising two interconnected sub-decoding networks (a first sub-decoding network and a second sub-decoding network) as an example, in this embodiment, difference-aware features are input into the first sub-decoding network, which decodes the received difference-aware features to obtain the sub-difference region corresponding to the first sub-decoding network. The sub-difference region corresponding to the first sub-decoding network is then input into the second sub-decoding network, which decodes the received sub-difference region again to output the target difference region between the first and second images to be detected. In the above embodiment, the difference-aware features are decoded through at least two sub-decoding networks, ensuring sufficient decoding and making the output target difference region more accurate.

[0110] In addition, this embodiment also provides the process of training the difference-aware model, such as... Figure 11 As shown, it may include the following steps:

[0111] S1101, Input the sample image pair into the difference perception model to obtain the sample difference region and difference perception feature corresponding to the sample image predicted by the difference perception model.

[0112] The sample image pairs contain sample images of the same region collected at different times. Specifically, in this embodiment, the sample image pairs are input into the difference perception model. The difference perception model processes the input sample image pairs and outputs the predicted sample difference region and predicted difference perception feature corresponding to the sample image pairs. The process of outputting the predicted sample difference region and predicted difference perception feature is the same as the process of determining the difference region and difference perception feature, and will not be described in detail here.

[0113] S1102, determine the cross-entropy loss value based on the true difference region and the sample difference region corresponding to the sample image.

[0114] Specifically, in this embodiment, the predicted difference region of the output sample can be compared with the actual difference region, for example, by subtraction, and the subtraction result can be used as the cross-entropy loss value.

[0115] Alternatively, there are many ways to determine the true difference region. For example, it can be determined by using a certain device to identify the difference feature region of the sample image; or it can be determined based on the staff's observation of the sample image. There is no limitation on this.

[0116] S1103, determine the contrast loss value based on the difference perception features and difference feature labels corresponding to the sample images.

[0117] Here, the differential feature labels are pre-determined true differential features. The contrastive loss value can be the loss between the differential perceived features and the differential feature labels, determined by the contrastive loss function.

[0118] In this embodiment, the difference-aware features and difference feature labels can be vectorized first, treating each point in the difference features and difference feature labels as a vector value, and then the contrast loss value can be calculated. For example, the contrast loss value can be determined using the following formula (3):

[0119]

[0120] In the formula, L cont It is the contrast loss value; d ij It is the Euclidean distance between the feature vectors of the first and second images to be detected; w represents the balancing weights; y ij The label at position (i,j) represents the image difference region; m represents the hyperparameter. In this embodiment, the label of the image difference region is set to 1, and the label of the image difference region is set to 0. Therefore, the first term corresponding to the image difference region in the above formula (3) is zero, and the second term corresponding to the image difference region is zero. If the sample image pairs are similar, and the Euclidean distance in the feature space is large, it indicates that the current model is not good, and the comparison loss value is correspondingly large. If the Euclidean distance in the feature space is small when the sample image pairs are similar but dissimilar, the loss value will also be correspondingly large.

[0121] S1104, The difference-aware model is trained based on the cross-entropy loss value and the contrast loss value.

[0122] Optionally, in this embodiment, the determined cross-entropy loss value and contrastive loss value can be fused, for example, by weighted summation, to obtain a summation result, which is then used to train the difference-aware model. Alternatively, in this embodiment, the difference-aware model can first be trained based on the cross-entropy loss value, the parameters of the difference-aware model can be updated, and then the contrastive loss value can be used to train the updated difference-aware model again.

[0123] In the above embodiments, training the difference perception model using cross-entropy loss and contrastive loss can enrich the training process of the difference perception model and make the difference perception results of the trained difference perception model more accurate.

[0124] Furthermore, after training the difference perception model, in this embodiment, the difference regions of the samples can also be evaluated based on the F1-Score, overall accuracy (OA), and Kappa coefficient to determine the accuracy of the output results of the difference perception model.

[0125] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0126] Based on the same inventive concept, this application also provides an image difference detection apparatus for implementing the image difference detection method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more image difference detection apparatus embodiments provided below can be found in the limitations of the image difference detection method described above, and will not be repeated here.

[0127] In one embodiment, such as Figure 12 As shown, an image difference detection device 2 is provided, comprising: a feature extraction module 20, a difference perception module 21, and a difference determination module 22, wherein:

[0128] The feature extraction module 20 is used to extract the first global local features corresponding to the first image to be detected and the second global local features corresponding to the second image to be detected through the feature encoding network of the difference perception model.

[0129] The first image to be detected and the second image to be detected are images of the same area collected at different times.

[0130] The difference perception module 21 is used to perform difference perception processing on the first global local features and the second global local features to determine the difference perception features between the first image to be detected and the second image to be detected.

[0131] The difference determination module 22 is used to decode the difference-aware features through the feature decoding network of the difference-aware model to obtain the target difference region between the first image to be detected and the second image to be detected.

[0132] In one embodiment, the feature encoding network includes at least two end-to-end connected sub-encoding networks. Accordingly, the feature extraction module 20 is specifically used to encode the input data of each sub-encoding network to obtain the first global local features extracted by the sub-encoding network for the first image to be detected, and the second global local features extracted by the sub-encoding network for the second image to be detected.

[0133] The input data for the first sub-coding network consists of the first image to be detected and the second image to be detected; the input data for each other sub-coding network consists of the first global local feature and the second global local feature output by the previous sub-coding network.

[0134] In one embodiment, each sub-coding network includes: a global feature coding network, a local feature coding network, and a feature fusion network, such as... Figure 13 As shown, the feature extraction module 20 includes a first extraction unit 201, a second extraction unit 202, and a third extraction unit 203. Wherein:

[0135] The first extraction unit 201 is used to parse and process the input data of each sub-coding network through the global feature encoding network of each sub-coding network to obtain the first global feature corresponding to the sub-coding network.

[0136] The second extraction unit 202 is used to parse and process the input data of the sub-coding network through the local feature encoding network of the sub-coding network to obtain the first local feature corresponding to the sub-coding network.

[0137] The third extraction unit 203 is used to perform fusion processing on the first global feature and the first local feature corresponding to the sub-coding network through the feature fusion network of the sub-coding network to obtain the first global and local features extracted by the sub-coding network from the first image to be detected.

[0138] In one embodiment, the first extraction unit 201 is specifically used to parse and process the input data of each sub-encoding network by using the global feature encoding network of each sub-encoding network, based on channel attention mechanism and spatial attention mechanism, to obtain the first global feature corresponding to the sub-encoding network.

[0139] In one embodiment, the difference-aware model further includes: a difference-aware network; the difference-aware network includes at least two end-to-end connected sub-difference-aware networks, and each sub-difference-aware network corresponds one-to-one with each sub-encoding network in the feature encoding network; correspondingly, such as Figure 14 As shown, the difference perception module 21 includes a first determining unit 210 and a fusion unit 211. Wherein:

[0140] The first determining unit 210 is used to determine sub-difference features based on the input data corresponding to each sub-difference sensing network.

[0141] The fusion unit 211 is used to fuse the sub-difference features determined by the last sub-difference sensing network with the first global local features and the second global local features extracted by the sub-coding network corresponding to the last sub-difference sensing network to obtain the difference sensing features between the first image to be detected and the second image to be detected.

[0142] The input data for the first sub-differential sensing network consists of the input data for the first sub-coding network, as well as the first global local feature and the second global local feature extracted by the first sub-coding network. The input data for each other sub-differential sensing network consists of the input data of the sub-coding network corresponding to the sub-differential sensing network, the first global local feature and the second global local feature extracted by the sub-coding network corresponding to the sub-differential sensing network, and the sub-differential feature determined by the previous sub-differential sensing network of the sub-differential sensing network.

[0143] In one embodiment, the first determining unit 210 is specifically used to determine initial difference features through each sub-difference sensing network based on the first global local features and the second global local features extracted by the sub-encoding network corresponding to the sub-difference sensing network, and to fuse the initial difference features, the sub-difference features determined by the previous sub-difference sensing network of the sub-difference sensing network, and the input data of the sub-encoding network corresponding to the sub-difference sensing network to obtain the sub-difference features determined by the sub-difference sensing network.

[0144] In one embodiment, the feature decoding network includes at least two end-to-end connected sub-decoding networks; correspondingly, such as Figure 15 As shown, the difference determination module 22 includes a second determination unit 220 and a third determination unit 221. Wherein:

[0145] The second determining unit 220 is used to decode the input data of each sub-decoding network to obtain the sub-difference region determined by the sub-decoding network.

[0146] The third determining unit 221 is used to take the sub-difference region determined by the last sub-decoding network as the target difference region between the first image to be detected and the second image to be detected.

[0147] The input data for the first sub-decoding network is the difference-aware feature, while the input data for each of the other sub-decoding networks is the sub-difference region determined by the previous sub-decoding network.

[0148] In one embodiment, the above Figure 12 The image difference detection device 2 shown also includes a training module 23, such as Figure 16 As shown, the training module 23 includes a sample input unit 230, a fourth determination unit 231, a fifth determination unit 232, and a training unit 233. Wherein:

[0149] The input unit 230 is used to input sample image pairs into the difference perception model to obtain the sample difference region and difference perception feature corresponding to the sample image predicted by the difference perception model.

[0150] Each sample image pair contains sample images of the same area collected at different times.

[0151] The fourth determining unit 231 is used to determine the cross-entropy loss value based on the real difference region and the sample difference region corresponding to the sample image.

[0152] The fifth determining unit 232 is used to determine the contrast loss value based on the difference perception features and difference feature labels corresponding to the sample image.

[0153] Training unit 233 is used to train the difference-aware model based on the cross-entropy loss value and the contrast loss value.

[0154] Each module in the aforementioned image difference detection device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0155] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 17As shown, the computer device includes a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When executed by the processor, the computer program implements an image difference detection method. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the computer device casing, or an external keyboard, touchpad, or mouse.

[0156] Those skilled in the art will understand that Figure 17 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0157] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:

[0158] The first global local feature corresponding to the first image to be detected and the second global local feature corresponding to the second image to be detected are extracted through the feature encoding network of the difference perception model; wherein the first image to be detected and the second image to be detected are images of the same area collected at different times;

[0159] Perform difference sensing processing on the first global local features and the second global local features to determine the difference sensing features between the first image to be detected and the second image to be detected.

[0160] The feature decoding network of the difference perception model is used to decode the difference perception features to obtain the target difference region between the first image to be detected and the second image to be detected.

[0161] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0162] The first global local feature corresponding to the first image to be detected and the second global local feature corresponding to the second image to be detected are extracted through the feature encoding network of the difference perception model; wherein the first image to be detected and the second image to be detected are images of the same area collected at different times;

[0163] Perform difference sensing processing on the first global local features and the second global local features to determine the difference sensing features between the first image to be detected and the second image to be detected.

[0164] The feature decoding network of the difference perception model is used to decode the difference perception features to obtain the target difference region between the first image to be detected and the second image to be detected.

[0165] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, performs the following steps:

[0166] The first global local feature corresponding to the first image to be detected and the second global local feature corresponding to the second image to be detected are extracted through the feature encoding network of the difference perception model; wherein the first image to be detected and the second image to be detected are images of the same area collected at different times;

[0167] Perform difference sensing processing on the first global local features and the second global local features to determine the difference sensing features between the first image to be detected and the second image to be detected.

[0168] The feature decoding network of the difference perception model is used to decode the difference perception features to obtain the target difference region between the first image to be detected and the second image to be detected.

[0169] It should be noted that the user information (including but not limited to the image information and feature information to be detected) and data (including but not limited to the input data and output data) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0170] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0171] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0172] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. An image difference detection method, characterized in that, The method includes: The feature encoding network of the difference perception model is used to extract the first global local features corresponding to the first image to be detected and the second global local features corresponding to the second image to be detected; wherein the first image to be detected and the second image to be detected are images of the same region collected at different times; Perform difference sensing processing on the first global local features and the second global local features to determine the difference sensing features between the first image to be detected and the second image to be detected. The difference-aware features are decoded through the feature decoding network of the difference-aware model to obtain the target difference region between the first image to be detected and the second image to be detected. The feature coding network includes at least two end-to-end connected sub-coding networks and a difference sensing network; the difference sensing network includes at least two end-to-end connected sub-difference sensing networks, and each sub-difference sensing network corresponds one-to-one with each sub-coding network in the feature coding network. Perform difference-sensing processing on the first global local features and the second global local features to determine the difference-sensing features between the first image to be detected and the second image to be detected, including: By using each sub-difference sensing network, the sub-difference features are determined based on the input data corresponding to that sub-difference sensing network. The sub-difference features determined by the last sub-difference sensing network are fused with the first global local features and the second global local features extracted by the sub-coding network corresponding to the last sub-difference sensing network to obtain the difference sensing features between the first image to be detected and the second image to be detected. The input data for the first sub-differential sensing network consists of the input data for the first sub-coding network, as well as the first global local feature and the second global local feature extracted by the first sub-coding network. The input data for each other sub-differential sensing network consists of the input data of the sub-coding network corresponding to the sub-differential sensing network, the first global local feature and the second global local feature extracted by the sub-coding network corresponding to the sub-differential sensing network, and the sub-differential feature determined by the previous sub-differential sensing network of the sub-differential sensing network.

2. The method according to claim 1, characterized in that, The step of extracting the first global local features corresponding to the first image to be detected and the second global local features corresponding to the second image to be detected through the feature encoding network of the difference-aware model includes: By encoding the input data of each sub-coding network, the first global local features extracted by the sub-coding network for the first image to be detected and the second global local features extracted by the sub-coding network for the second image to be detected are obtained. The input data for the first sub-coding network consists of the first image to be detected and the second image to be detected; the input data for each other sub-coding network consists of the first global local feature and the second global local feature output by the previous sub-coding network.

3. The method according to claim 2, characterized in that, Each sub-coding network includes: a global feature coding network, a local feature coding network, and a feature fusion network; the process of encoding the input data of each sub-coding network to obtain the first global and local features extracted by the sub-coding network for the first image to be detected includes: The input data of each sub-coding network is parsed and processed by the global feature encoding network of each sub-coding network to obtain the first global feature corresponding to that sub-coding network. The input data of the sub-coding network is parsed and processed by the local feature encoding network of the sub-coding network to obtain the first local feature corresponding to the sub-coding network; The feature fusion network of the sub-coding network is used to fuse the first global feature and the first local feature corresponding to the sub-coding network to obtain the first global and local features extracted by the sub-coding network from the first image to be detected.

4. The method according to claim 3, characterized in that, The step of parsing and processing the input data of each sub-coding network through the global feature encoding network of each sub-coding network to obtain the first global feature corresponding to that sub-coding network includes: By encoding the global features of each sub-encoding network, and based on channel attention and spatial attention mechanisms, the input data of the sub-encoding network is parsed and processed to obtain the first global feature corresponding to the sub-encoding network.

5. The method according to claim 1, characterized in that, The step of determining sub-difference features based on the input data corresponding to each sub-difference sensing network includes: For each sub-difference sensing network, an initial difference feature is determined based on the first and second global local features extracted by the sub-encoding network corresponding to that sub-difference sensing network. The initial difference feature, the sub-difference feature determined by the previous sub-difference sensing network, and the input data of the sub-encoding network corresponding to that sub-difference sensing network are then fused to obtain the sub-difference feature determined by that sub-difference sensing network.

6. The method according to claim 1, characterized in that, The feature decoding network includes at least two end-to-end connected sub-decoding networks; The feature decoding network of the difference-aware model decodes the difference-aware features to obtain the target difference region between the first image to be detected and the second image to be detected, including: Each sub-decoding network decodes its input data to obtain the sub-difference region determined by that sub-decoding network. The sub-difference region determined by the last sub-decoding network is taken as the target difference region between the first image to be detected and the second image to be detected. The input data for the first sub-decoding network is the difference-aware feature, and the input data for each other sub-decoding network is the sub-difference region determined by the previous sub-decoding network.

7. The method according to any one of claims 1-6, characterized in that, The training process of the difference-aware model includes: The sample image pairs are input into the difference perception model to obtain the sample difference regions and difference perception features corresponding to the sample images predicted by the difference perception model; wherein, the sample image pairs contain sample images of the same region collected at different times; The cross-entropy loss value is determined based on the true difference region corresponding to the sample image and the sample difference region; The contrast loss value is determined based on the difference perception features and difference feature labels corresponding to the sample images; The difference-aware model is trained based on the cross-entropy loss value and the contrast loss value.

8. An image difference detection device, characterized in that, The device includes: The feature extraction module is used to extract the first global local features corresponding to the first image to be detected and the second global local features corresponding to the second image to be detected through the feature encoding network of the difference perception model; wherein the first image to be detected and the second image to be detected are images of the same region collected at different times; The difference perception module is used to perform difference perception processing on the first global local features and the second global local features to determine the difference perception features between the first image to be detected and the second image to be detected. The difference determination module is used to decode the difference-aware features through the feature decoding network of the difference-aware model to obtain the target difference region between the first image to be detected and the second image to be detected. The feature coding network includes at least two end-to-end connected sub-coding networks and a difference sensing network; the difference sensing network includes at least two end-to-end connected sub-difference sensing networks, and each sub-difference sensing network corresponds one-to-one with each sub-coding network in the feature coding network. The difference perception module includes: The first determining unit is used to determine the sub-difference features based on the input data corresponding to each sub-difference sensing network. The fusion unit is used to fuse the sub-difference features determined by the last sub-difference sensing network with the first global local features and the second global local features extracted by the sub-coding network corresponding to the last sub-difference sensing network to obtain the difference sensing features between the first image to be detected and the second image to be detected. The input data for the first sub-differential sensing network consists of the input data for the first sub-coding network, as well as the first global local feature and the second global local feature extracted by the first sub-coding network. The input data for each other sub-differential sensing network consists of the input data of the sub-coding network corresponding to the sub-differential sensing network, the first global local feature and the second global local feature extracted by the sub-coding network corresponding to the sub-differential sensing network, and the sub-differential feature determined by the previous sub-differential sensing network of the sub-differential sensing network.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Image processing method and device, electronic equipment and computer readable storage medium

    CN113887544A

  • Image segmentation method and device, equipment and storage medium

    CN114581462A