Image change detection method and device, electronic equipment and storage medium

Through image change detection models based on preset neighborhood size and multi-scale feature fusion, the problem of inaccurate feature extraction and detection in the prior art is solved, and more efficient image change detection is achieved.

CN120580211APending Publication Date: 2025-09-02AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510709678.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

In the prior art, the pixel-by-pixel change analysis method has problems of inaccurate feature extraction and detection in image change detection, especially inaccuracy caused by fixed feature extraction scales.

Method used

Using image segmentation based on preset neighborhood size and pre-trained image change detection model, multi-scale features are extracted, and image change detection results are determined through multi-scale features fusion and classification.

Benefits of technology

The accuracy and efficiency of image change detection are improved, and the changes between images can be more accurately identified.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120580211A_ABST
    Figure CN120580211A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an image change detection method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring a first image and a second image in a to-be-analyzed Internet of Things video, and performing image processing on the first image and the second image to determine a to-be-detected target image; performing image segmentation on the target image based on a preset neighborhood size, and determining a neighborhood image corresponding to each pixel point in the target image; based on a pre-trained image change detection model, performing image change detection on each neighborhood image, and determining an image change detection result between the first image and the second image; wherein the image change detection model is used for extracting multi-scale features corresponding to neighborhood images, performing feature fusion on the multi-scale features, and performing classification according to the fused features. Through the technical scheme of the embodiment of the invention, the multi-scale features can be accurately and conveniently extracted, so that the accuracy of feature extraction and change detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of image processing technology, and in particular to an image change detection method, device, electronic device, and storage medium. Background Art

[0002] With the development of technology, the Internet of Everything (IoE) is becoming increasingly widespread. In the IoT, sensor devices can collect real-time information about any object that needs to be monitored, connected, and interacted with. These devices can then access the IoT through wired or wireless networks, ultimately connecting objects to one another and to people. Furthermore, IoT video analytics enable intelligent perception, monitoring, and analysis of objects and processes.

[0003] In IoT video analysis, many scenes remain unchanged for extended periods of time. Using video analysis algorithms to continuously analyze every IoT video is a waste of resources. Therefore, a method has emerged to detect changes in the preceding image before performing IoT video analysis.

[0004] Currently, pixel-by-pixel change analysis is commonly used to detect changes in the entire image. However, this approach uses a fixed feature extraction scale and performs pixel-by-pixel single-scale feature comparisons across the entire image, leading to inaccurate feature extraction and change detection. Summary of the Invention

[0005] Embodiments of the present invention provide an image change detection method, apparatus, electronic device, and storage medium to accurately and conveniently extract multi-scale features, thereby improving the accuracy of feature extraction and change detection.

[0006] In a first aspect, an embodiment of the present invention provides an image change detection method, comprising:

[0007] Acquire a first image and a second image in the IoT video to be analyzed, and perform image processing on the first image and the second image to determine a target image to be detected;

[0008] Performing image segmentation on the target image based on a preset neighborhood size to determine a neighborhood image corresponding to each pixel in the target image;

[0009] performing image change detection on each of the neighborhood images based on a pre-trained image change detection model to determine an image change detection result between the first image and the second image;

[0010] The image change detection model is used to extract multi-scale features corresponding to neighborhood images, perform feature fusion on the multi-scale features, and perform classification based on the fused features.

[0011] Optionally, the method further includes: performing difference analysis based on the first image and the second image to determine a difference image between the first image and the second image; and performing image synthesis based on the first image, the second image and the difference image to determine a target image to be detected.

[0012] Optionally, the method also includes: inputting the neighborhood image corresponding to each pixel point into a pre-trained image change detection model to perform image change detection to obtain the change state corresponding to each pixel point; based on a preset display method and the change state corresponding to each pixel point in the target image output by the image change detection model, generating an image change detection result between the first image and the second image.

[0013] Optionally, the method further includes: the image change detection model includes: a multi-scale feature extraction module, a multi-scale feature fusion module and a classification module; for each pixel point, inputting the neighborhood image corresponding to the current pixel point into the multi-scale feature extraction module for multi-scale feature extraction to obtain multiple candidate scale features corresponding to the neighborhood image; inputting the multiple candidate scale features into the multi-scale feature fusion module for multi-scale feature fusion to obtain target scale features corresponding to the neighborhood image; inputting the target scale features into the classification module for classification processing to obtain the change state corresponding to the current pixel point.

[0014] Optionally, the method further includes: the multi-scale feature fusion module includes: a cross-region association unit and a multi-scale feature fusion unit; inputting the multiple candidate scale features into the cross-region association unit for cross-region association to obtain multiple reference scale features corresponding to the neighborhood image; inputting the multiple reference scale features into the multi-scale feature fusion unit for multi-scale feature fusion to obtain target scale features corresponding to the neighborhood image.

[0015] Optionally, the method further includes: determining a comparison result corresponding to the IoT video to be analyzed based on a comparison of the image change detection result and a preset change threshold; if the comparison result indicates that the IoT video needs to be analyzed and identified, performing target recognition based on a pre-trained target recognition model and the IoT video to be analyzed to obtain a target recognition result.

[0016] In a second aspect, an embodiment of the present invention further provides an image change detection device, the device comprising:

[0017] a target image determination module, configured to obtain a first image and a second image in the IoT video to be analyzed, and perform image processing on the first image and the second image to determine a target image to be detected;

[0018] A neighborhood image determination module is used to perform image segmentation on the target image based on a preset neighborhood size, and determine a neighborhood image corresponding to each pixel in the target image;

[0019] an image change detection result determination module, configured to perform image change detection on each of the neighborhood images based on a pre-trained image change detection model, and determine an image change detection result between the first image and the second image;

[0020] The image change detection model is used to extract multi-scale features corresponding to neighborhood images, perform feature fusion on the multi-scale features, and perform classification based on the fused features.

[0021] In a third aspect, an embodiment of the present invention further provides an electronic device, comprising:

[0022] one or more processors;

[0023] a memory for storing one or more programs;

[0024] When the one or more programs are executed by the one or more processors, the one or more processors implement the image change detection method provided by any embodiment of the present invention.

[0025] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the image change detection method provided by any embodiment of the present invention.

[0026] In a fifth aspect, an embodiment of the present invention provides a computer program product, including a computer program, which, when executed by a processor, implements the image change detection method provided by any embodiment of the present invention.

[0027] The technical solution of the embodiment of the present invention is to obtain the first image and the second image in the IoT video to be analyzed, and perform image processing on the first image and the second image to determine the target image to be detected; perform image segmentation on the target image based on a preset neighborhood size to determine the neighborhood image corresponding to each pixel in the target image; perform image change detection on each of the neighborhood images based on a pre-trained image change detection model to determine the image change detection result between the first image and the second image; wherein the image change detection model is used to extract multi-scale features corresponding to the neighborhood images, perform feature fusion on the multi-scale features, and perform classification based on the fused features, so that the multi-scale features can be accurately and conveniently extracted, thereby improving the accuracy of feature extraction and change detection.

[0028] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0030] Figure 1 This is a flowchart of an image change detection method provided by Example 1 of the present invention;

[0031] Figure 2 This is a flow chart of an image change detection method provided by Embodiment 2 of the present invention;

[0032] Figure 3 This is an architectural diagram of a cross-region association unit involved in the second embodiment of the present invention;

[0033] Figure 4 1 is a structural diagram of an image change detection device provided by Embodiment 3 of the present invention;

[0034] Figure 5 It is a structural diagram of an electronic device for implementing the image change detection method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0035] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0036] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0037] Example 1

[0038] Figure 1 A flowchart of an image change detection method is provided for the first embodiment of the present invention. This embodiment is applicable to the case of detecting whether the content of a video frame image at different moments has changed. The method can be performed by an image change detection device, which can be implemented in the form of hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:

[0039] S110: Acquire a first image and a second image in the IoT video to be analyzed, and perform image processing on the first image and the second image to determine a target image to be detected.

[0040] Among them, the IoT video to be analyzed may refer to the video to be subjected to IoT video analysis. All images in the IoT video to be analyzed are images of the same perspective and size. In this embodiment, before performing the IoT video analysis operation on the IoT video to be analyzed, it is necessary to first determine whether there are changes between different frames of the IoT video, that is, to determine the image change detection result. For example, the IoT video to be analyzed may be, but is not limited to, a product quality inspection video or a factory surveillance video. The products in the product quality inspection video may be, but are not limited to, printed clothing or custom-shaped workpieces. Product quality inspection videos can be used to monitor whether the products produced are qualified. Factory surveillance videos can be used to monitor the entry and exit status of people, vehicles, and goods in the factory. The first image may refer to a frame of the IoT video to be analyzed. The second image may refer to another frame of the IoT video to be analyzed that is different from the first image. The target image may refer to an image obtained by fusion of images at different times.

[0041] Specifically, the user periodically sends the IoT video to be analyzed to the server through the communication protocol. A server that is pre-configured to implement the image change detection method in the embodiment of the present invention can obtain the first image and the second image from the IoT video to be analyzed based on a preset interval duration. Detect the image attributes of the first image and / or the second image. If it is detected that the image attribute is a single-channel image (such as a grayscale image), the first image and the second image are fused using a multi-channel fusion technology to obtain a target image. If it is detected that the image attribute is a multi-channel image, the preset channel (focus channel) is extracted from the first image and the second image to obtain a first image of a single channel and a second image of a single channel, and the first image and the second image are fused to obtain a target image.

[0042] On the basis of the above technical solution, "performing image processing on the first image and the second image to determine the target image to be detected" may include: performing difference analysis based on the first image and the second image to determine the difference image between the first image and the second image; performing image synthesis based on the first image, the second image and the difference image to determine the target image to be detected.

[0043] The difference image may refer to an image obtained by preliminarily analyzing the difference between the first image and the second image, and the size of the difference image may be the same as that of the first image and / or the second image.

[0044] Specifically, a logarithmic ratio operator is used to perform difference analysis on the first image and the second image to obtain a difference image between the first image and the second image. In this embodiment, the difference image is calculated as follows: DI =|logI1-logI2|. Where I1 represents the first image, I2 represents the second image, and I DI represents the difference image between images I1 and I2, |·| represents the absolute value operation, and log represents the logarithm operation with base 10. The first image, the second image, and the difference image are multi-channel fused to determine the target image to be detected.

[0045] S120 , performing image segmentation on the target image based on a preset neighborhood size, and determining a neighborhood image corresponding to each pixel in the target image.

[0046] The preset neighborhood size may refer to the preset size of the domain image corresponding to each pixel. For example, the preset domain size may be, but is not limited to, 5×5, 7×7, 9×9, or 11×11. A pixel may refer to the smallest unit of the basic element in an image. A domain image may refer to an image centered on a pixel and containing a specific number of pixels surrounding the central pixel.

[0047] Specifically, in the target image, the r×r image region surrounding each pixel is taken as the neighborhood image corresponding to each pixel. If the obtained neighborhood image meets the preset neighborhood size, no operation is required. If the obtained neighborhood image does not meet the preset neighborhood size, that is, the obtained neighborhood image is incomplete, the incomplete image portion is padded with zeros to obtain a complete neighborhood image. Pixels with incomplete neighborhood images are called edge pixels.

[0048] S130. Based on a pre-trained image change detection model, perform image change detection on each neighborhood image to determine the image change detection result between the first image and the second image; wherein the image change detection model is used to extract multi-scale features corresponding to the neighborhood images, perform feature fusion on the multi-scale features, and perform classification based on the fused features.

[0049] The image change detection model can be used to perform multi-scale feature extraction and multi-scale feature fusion on neighborhood images, and then use the fused features to classify whether the corresponding pixels in the neighborhood images have changed or not. The image change detection result can be an image change detection result between the first image and the second image, formed by combining the classification results of each pixel. In this embodiment, the image change detection result can be used to indicate whether there has been a substantial change between the content of the first image and the content of the second image. For example, in a product quality inspection video of a top that requires a sun pattern, the first image is a top with a sun pattern, but the second image is pants with a sun pattern. The image change detection result between the first image and the second image can be used to indicate whether a change has occurred between the first and second images. For another example, in a factory surveillance video of goods that require monitoring, the first image shows goods awaiting shipment in a certain area, while the second image shows the goods awaiting shipment remaining in the same area and not moving. The image change detection result between the first image and the second image can be used to indicate that no change has occurred between the first and second images.

[0050] Specifically, the neighborhood image is input into a pre-trained image change detection model, which performs multi-scale feature extraction and fusion on the neighborhood image. The fused features are then used to classify whether the corresponding pixels in the neighborhood image have changed or not, resulting in a model output. The model output corresponding to each pixel is then used to determine the image change detection result between the first and second images.

[0051] On the basis of the above technical solution, "based on a pre-trained image change detection model, performing image change detection on each neighborhood image, and determining the image change detection result between the first image and the second image" may include: inputting the neighborhood image corresponding to each pixel point into a pre-trained image change detection model for image change detection, and obtaining the change state corresponding to each pixel point; based on the preset display method and the change state corresponding to each pixel point in the target image output by the image change detection model, generating the image change detection result between the first image and the second image.

[0052] Among them, the image change detection results can have multiple display forms. For example, the image change detection results can be displayed in text form, and can also be displayed in image form. The model output result is the change state corresponding to each pixel point. The change state can be understood as a classification result. The change state can include changes and no changes. The preset display method can refer to a preset display method of the image change detection results. For example, the preset display method can be a method of displaying the image change detection results in text form, or a method of displaying the image change detection results in image form.

[0053] Specifically, the neighborhood image corresponding to each pixel is input into a pre-trained image change detection model for image change detection, obtaining the change state corresponding to each pixel. An image change detection result corresponding to the target image is formed based on the change states corresponding to all pixels in the target image, and this is used as the image change detection result between the first and second images. If the preset display method is a text display method, the position information and change state corresponding to each pixel can be combined into an array, such as (position information, change state), and the data generated by each pixel in the target image can be displayed as the image change detection result. Furthermore, data analysis can be performed on the data generated by each pixel to determine the pixel count percentage for each change state, and this pixel count percentage can also be displayed as the image change detection result. If the preset display method is an image display method, pixels corresponding to one change state (e.g., change) in the target image are rendered using a grayscale value (e.g., 0), and pixels corresponding to another change state (e.g., no change) in the target image are rendered using another grayscale value (e.g., 1), thereby obtaining a grayscale image and displaying this grayscale image as the image change detection result.

[0054] The technical solution of the embodiment of the present invention is to obtain the first image and the second image in the IoT video to be analyzed, and perform image processing on the first image and the second image to determine the target image to be detected; perform image segmentation on the target image based on a preset neighborhood size to determine the neighborhood image corresponding to each pixel in the target image; perform image change detection on each neighborhood image based on a pre-trained image change detection model to determine the image change detection result between the first image and the second image; wherein the image change detection model is used to extract multi-scale features corresponding to the neighborhood images, perform feature fusion on the multi-scale features, and classify them according to the fused features, so as to fully obtain useful multi-scale features from different scales to cope with situations with different image sizes, thereby improving the accuracy of feature extraction and change detection.

[0055] Based on the above technical solution, the method also includes: comparing the image change detection result with a preset change threshold to determine the comparison result corresponding to the IoT video to be analyzed; if the comparison result indicates that the IoT video needs to be analyzed and identified, target recognition is performed based on a pre-trained target recognition model and the IoT video to be analyzed to obtain a target recognition result.

[0056] The preset change threshold may refer to a pre-set threshold for determining whether an IoT video requires analysis and identification. The preset change threshold may include, but is not limited to, a threshold for the percentage of pixels whose change state is changed and a threshold for the area of ​​adjacent pixels whose change state is changed. The target recognition model may refer to a pre-trained neural network model for target recognition. For example, the target recognition model may be used to identify the type of target.

[0057] Specifically, if the image change detection result (such as the number of changed pixels accounts for 70%) is greater than the preset change threshold (the maximum number of changed pixels that do not require target recognition is 50%), the comparison result corresponding to the IoT video to be analyzed is determined to be the IoT video that needs to be analyzed and identified, and target recognition is performed based on the pre-trained target recognition model and the IoT video to be analyzed to obtain the target recognition result.

[0058] The technical solution of the embodiment of the present invention can also be applied to the situation where there are multiple groups of first images and second images in the IoT video to be analyzed. As long as the image change detection result of any group is greater than the preset change threshold, the comparison result of the IoT video to be analyzed to which the regrouped image belongs is that the IoT video needs to be analyzed and identified.

[0059] Example 2

[0060] Figure 2This is a flowchart of an image change detection method provided by the second embodiment of the present invention. Based on the above embodiments, this embodiment describes in detail the process of determining the change state using the image change detection model. The explanations of the terms that are the same or corresponding to the above embodiments are not repeated here. Figure 2 As shown, the method includes:

[0061] S210: Acquire a first image and a second image in the IoT video to be analyzed, and perform image processing on the first image and the second image to determine a target image to be detected.

[0062] S220 , performing image segmentation on the target image based on a preset neighborhood size, and determining a neighborhood image corresponding to each pixel in the target image.

[0063] S230 : For each pixel point, input the neighborhood image corresponding to the current pixel point into a multi-scale feature extraction module for multi-scale feature extraction to obtain a plurality of candidate scale features corresponding to the neighborhood image.

[0064] The image change detection model includes a multi-scale feature extraction module, a multi-scale feature fusion module, and a classification module. Candidate scale features may refer to dynamic scale features extracted using a dynamic convolution kernel. The multiple candidate scale features may include, but are not limited to, first-scale features, second-scale features, and third-scale features. First-scale features correspond to low-level features, second-scale features correspond to mid-level features, and third-scale features correspond to high-level features.

[0065] For example, low-level features are used to capture basic visual elements, such as edge / line detection (e.g., horizontal, vertical, and diagonal lines), color contrast and simple textures (e.g., spots and stripes), and local changes in brightness and darkness. Low-level features identify the overall scene structure (e.g., grass and sky in the background) and the general outline of an object (e.g., the overall shape of a puppy). This is equivalent to the "first glance" of human vision. Mid-level features are used to combine low-level features to form component-level semantics, such as the morphology of object components (e.g., a dog's head, ears, and limbs), moderately complex textures (e.g., hair direction and limb joints), and spatial positional relationships (e.g., "the dog's head is above the body"). Mid-level features identify object category (e.g., the key component differences between dogs and cats) and the basis for pose estimation (e.g., inferring movement from component positions). This is equivalent to recognizing "this is an animal's head and torso." High-level features are used to abstract and refine semantics, such as detailed textures (e.g., hair bifurcations and pupil reflections), category-specific features (e.g., differences in eye shape between different dog breeds), and global contextual associations (e.g., "the relative proportion of the eyes to the nose"). High-level features are used for precise classification (e.g., distinguishing between a golden retriever and a Labrador) and fine-grained recognition (e.g., determining an animal's mood and health status). This is equivalent to expert-level detailed observation.

[0066] Taking an image of a puppy as an example, low-level features mainly focus on the overall scene, such as the background and the outline of the puppy. Mid-level features mainly focus on the puppy's shape, such as the dog's head, body, legs, etc. High-level features mainly focus on the details of the puppy, such as hair, eyes, etc. That is, each layer pays more attention to details than the previous one. Compared with the existing method of extracting fixed features at a scale only to obtain high-level features, and using high-level features, a single-dimensional feature, for change detection, this embodiment uses multiple scales to extract multi-scale features, thereby obtaining more comprehensive features, and integrating multi-scale features for change detection, further improving detection accuracy.

[0067] Specifically, for each pixel point, the neighborhood image corresponding to the current pixel point is input into the multi-scale feature extraction module. The dynamic convolution kernel corresponding to each image is determined according to the input of different images, and the image is convolved using the convolution kernel corresponding to each image to obtain multiple candidate scale features corresponding to the neighborhood image.

[0068] In this embodiment, the determination of the dynamic convolution kernel and multiple candidate scale features can be achieved using three pre-set feature extraction branches. The feature extraction branch corresponding to the first scale feature is composed of a group of global feature encoding submodules, context feature projection submodules and convolution kernel weight generation submodules. The feature extraction branch corresponding to the second scale feature is composed of two groups of global feature encoding submodules, context feature projection submodules and convolution kernel weight generation submodules. The feature extraction branch corresponding to the third scale feature is composed of three groups of global feature encoding submodules, context feature projection submodules and convolution kernel weight generation submodules. Taking the feature extraction branch corresponding to the first scale feature as an example, first, the global feature encoding submodule encodes the global information through the average pooling layer and the linear layer; then the context feature projection submodule projects the encoded features to the output dimension space through the context feature projection; finally, the convolution kernel weight generation submodule generates a convolution kernel containing global context information, and uses the generated new convolution kernel and the group of inputs to perform conventional convolution to obtain the first scale feature. The input of the second group in the feature extraction branch corresponding to the second scale feature is the output of the first group.

[0069] For example, the global feature encoding submodule consists of an average pooling layer, a linear layer, a normalization layer, and an activation layer. The average pooling layer extracts global information from the input data and reduces the input data space size to k×k, where k is the convolution kernel size; then the features in all c channels are projected onto a vector of size m through a linear layer; and then the output features of the global feature encoding submodule are obtained after normalization and ReLU activation function: G∈R C×mThe context feature projection submodule consists of a linear layer, a normalization layer, and an activation layer. The linear layer projects the obtained feature G into a space with an output dimension of n, and then undergoes normalization and ReLU activation function to obtain the output feature of the context feature projection module: C∈R n×m The convolution kernel weight generation submodule modifies the convolution kernel by using different weights. The convolution kernel weight generation submodule first converts the features G and C obtained by two linear layers into the spatial size of the convolution kernel. The first linear layer converts the feature G∈R C×m Convert to M G ∈R c×k×k , the second linear layer transforms the feature C∈R n×m Convert to M C ∈R n×k×k , then M G Expand the dimension to get M' G ∈R n×c×k×k , M C Expand the dimension to get M' C ∈R n×c×k×k Then, M G and M C Add together to generate the convolution kernel weight M, the calculation formula is: M = δ (M' G +M′ C ). Where M∈R n×c×k×k The size of the original convolution kernel W is the same as that of the original convolution kernel. δ represents the Sigmoid activation function. Finally, the generated convolution kernel weight M is element-wise multiplied by the current convolution kernel weight W. Combining global and local information, a new convolution kernel weight W′ that is adaptive to the input data is obtained. The calculation formula is: W′ = M⊙W. ⊙ represents element-wise multiplication.

[0070] S240: Input the multiple candidate scale features into a multi-scale feature fusion module for multi-scale feature fusion to obtain target scale features corresponding to the neighborhood image.

[0071] The target scale feature may refer to a fusion feature that integrates multiple scales. Specifically, the first scale feature, the second scale feature, and the third scale feature are input into a multi-scale feature fusion module for multi-scale feature fusion to obtain the target scale feature corresponding to the neighborhood image.

[0072] Based on the above technical solution, "inputting multiple candidate scale features into a multi-scale feature fusion module for multi-scale feature fusion to obtain target scale features corresponding to the neighborhood image" may include: inputting multiple candidate scale features into a cross-region association unit for cross-region association to obtain multiple reference scale features corresponding to the neighborhood image; and inputting multiple reference scale features into a multi-scale feature fusion unit for multi-scale feature fusion to obtain target scale features corresponding to the neighborhood image.

[0073] The multi-scale feature fusion module includes: a cross-region correlation unit and a multi-scale feature fusion unit. The reference scale feature may include the first scale feature, the second scale feature, and the third scale feature after cross-region correlation. In this embodiment, Figure 3 A diagram of the architecture of a cross-region association unit is given. Figure 3 The cross-region association unit can be composed of a layer normalization layer (LN), a multi-head self-attention layer (MSA), a layer normalization layer (LN), and a multi-layer perceptron (MLP) layer, with residual connections in between. To address the problem that convolutional neural networks cannot achieve long-distance dependencies, this embodiment uses cross-region units to simulate long-distance dependencies, thereby improving the global semantic extraction capability of the image change detection model and enhancing the global understanding of the image.

[0074] For example, in the feature extraction branch corresponding to the first scale feature, the input data is processed by a group and then the first scale feature is extracted. The calculation formula is: Among them, X represents the input data in the input layer, W l ′ represents the dynamic convolution kernel weight obtained by the convolution kernel weight generation submodule, b l Represents the bias term of the dynamic convolution kernel, which is obtained by random initialization and optimized through network iterative training; BN represents the batch normalization operation, σ represents the ReLU activation function, Represents the first-scale feature. On this basis, the first-scale feature after cross-region association can be briefly expressed as ATM *3 Indicates three Figure 3 The cross-region correlation unit shown. Similarly, the second-scale features and third-scale features after cross-region correlation are: and in, represents the second-scale feature of the feature extraction branch corresponding to the second-scale feature, It represents the third-scale feature of the feature extraction branch corresponding to the third-scale feature. The specific calculation process is: and Among them, W m ′ and W h ′ represents the dynamic convolution kernel weights obtained by each branch convolution kernel weight generation submodule, b m and b h They represent the bias terms of the dynamic convolution kernels of each scale, which are obtained by random initialization and optimized through network iterative training. Finally, the results of the three feature extraction branches are added together element by element to obtain the feature fusion result, which is expressed as: F = F l +F m +Fh .

[0075] It should be noted that each feature extraction branch corresponds to at least three cross-region association units. The execution order of each cross-region association unit is: input → LN1 → MSA (residual connection 1) → LN2 → MLP (residual connection 2) → output. Residual connections all use element-by-element addition. The first LN layer can be used to normalize the input data channel dimension to eliminate feature dimensionality differences. The first LN layer improves model training stability. The second LN layer can be used for secondary normalization after MSA to stabilize gradient propagation. The second LN layer prevents feature distribution shift caused by MSA. Multi-head self-attention (MSA) splits features into h heads (e.g., 8 heads) through Q / K / V projection, and each head independently calculates attention weights. The multi-head mechanism facilitates the capture of mixed-domain features (local / global, color / texture). The adaptive mechanism is reflected in the introduction of a learnable scale factor in the calculation of the attention matrix to dynamically adjust the association strength of different spatial locations. The learnable scale parameter enables the network to dynamically adjust the attention range to adapt to the association requirements of objects of different scales. The multi-layer perceptron (MLP) consists of two fully connected layers (with GELU activation in the middle) with an expansion ratio of 4:1. Residual connections can directly transfer original features across layers, alleviating the gradient vanishing problem in deep networks. Taking the input feature X as an example, the first layer LN: X norm1 =LN(X); MSA calculation: X attn =MAS(X norm1 )+X; second layer LN: X norm2 =LN(X attn ); MLP processing: X mlp =MLP(X norm2 )+X attn .

[0076] S250: Input the target scale feature into a classification module for classification processing to obtain a change state corresponding to the current pixel point.

[0077] Specifically, the obtained fusion feature F (i.e., target scale feature) is input into the classification module and obtained through the full connection layer: Y = W fc2 (W fc1 F). Among them, W fc1 Represents the first layer of fully connected operations, W fc2 Represents the second layer of full connection operation. After two full connection operations, the Y obtained is a vector of dimension 2×1 Where a represents the probability that the pixel point corresponding to the input neighborhood image belongs to the unchanged class (no change), b represents the probability that the pixel point corresponding to the input neighborhood image belongs to the changed class (change), according to the vector To output the change state of the pixel points corresponding to each neighborhood image When a>b, Equal to the category to which a belongs, that is When a<b, Equal to the category to which b belongs, that is

[0078] It should be noted that the first fully connected (FC1) operation takes as input the high-dimensional features (e.g., 512-dimensional) from the preceding network (e.g., convolutional or Transformer modules). ReLU is typically used to introduce nonlinearity, mapping the 512-dimensional features to an intermediate dimension (e.g., 64-dimensional) for initial semantic abstraction. The second fully connected (FC2) operation takes as input the 64-dimensional features output by FC1, converts the output to probability values ​​using Softmax, and then maps it to 2D.

[0079] S260: Generate an image change detection result between the first image and the second image based on a preset display method and a change state corresponding to each pixel point in the target image output by the image change detection model.

[0080] The technical solution of the embodiment of the present invention is to input the neighborhood image corresponding to each pixel into a multi-scale feature extraction module for multi-scale feature extraction, thereby obtaining multiple candidate scale features corresponding to the neighborhood image. To address the problem that objects in natural scenes may appear in images of different sizes, a multi-scale network structure is designed to extract the first-scale features, second-scale features, and third-scale features of the image, thereby further fully acquiring useful multi-scale features from different scale perspectives, giving the model stronger multi-scale representation capabilities. The multiple candidate scale features are input into a multi-scale feature fusion module for multi-scale feature fusion, obtaining the target scale features corresponding to the neighborhood image. The target scale features are input into a classification module for classification processing, obtaining the change state corresponding to the current pixel, further improving the accuracy of change detection.

[0081] The following is an embodiment of an image change detection device provided by an embodiment of the present invention. The device and the image change detection methods of the above-mentioned embodiments belong to the same inventive concept. For details not fully described in the embodiment of the image change detection device, please refer to the embodiment of the above-mentioned image change detection method.

[0082] Example 3

[0083] Figure 4 This is a schematic diagram of the structure of an image change detection device provided by the third embodiment of the present invention. Figure 4 As shown, the device includes: a target image determination module 410, a neighborhood image determination module 420 and an image change detection result determination module 430.

[0084] Among them, the target image determination module 410 is used to obtain the first image and the second image in the IoT video to be analyzed, and perform image processing on the first image and the second image to determine the target image to be detected; the neighborhood image determination module 420 is used to perform image segmentation on the target image based on a preset neighborhood size, and determine the neighborhood image corresponding to each pixel in the target image; the image change detection result determination module 430 is used to perform image change detection on each neighborhood image based on a pre-trained image change detection model, and determine the image change detection result between the first image and the second image; wherein, the image change detection model is used to extract multi-scale features corresponding to the neighborhood images, perform feature fusion on the multi-scale features, and perform classification based on the fused features.

[0085] The technical solution of the embodiment of the present invention is to obtain the first image and the second image in the IoT video to be analyzed, and perform image processing on the first image and the second image to determine the target image to be detected; perform image segmentation on the target image based on a preset neighborhood size to determine the neighborhood image corresponding to each pixel in the target image; perform image change detection on each neighborhood image based on a pre-trained image change detection model to determine the image change detection result between the first image and the second image; wherein the image change detection model is used to extract multi-scale features corresponding to the neighborhood images, perform feature fusion on the multi-scale features, and perform classification based on the fused features, so that the multi-scale features can be accurately and conveniently extracted, thereby improving the accuracy of feature extraction and change detection.

[0086] Based on the above technical solution, the target image determination module 410 is specifically used to: perform difference analysis based on the first image and the second image to determine the difference image between the first image and the second image; perform image synthesis based on the first image, the second image and the difference image to determine the target image to be detected.

[0087] Based on the above technical solution, the image change detection result determination module 430 may include:

[0088] The change state determination submodule is used to input the neighborhood image corresponding to each pixel into a pre-trained image change detection model to perform image change detection and obtain the change state corresponding to each pixel;

[0089] The image change detection result determination submodule is used to generate an image change detection result between the first image and the second image based on a preset display method and the change state corresponding to each pixel point in the target image output by the image change detection model.

[0090] Based on the above technical solution, the image change detection model includes: a multi-scale feature extraction module, a multi-scale feature fusion module and a classification module;

[0091] The change state determination submodule may include:

[0092] A candidate scale feature determination unit is configured to input, for each pixel, a neighborhood image corresponding to the current pixel into a multi-scale feature extraction module for multi-scale feature extraction, thereby obtaining a plurality of candidate scale features corresponding to the neighborhood image;

[0093] The target scale feature determination unit is used to input multiple candidate scale features into the multi-scale feature fusion module for multi-scale feature fusion to obtain the target scale feature corresponding to the neighborhood image;

[0094] The change state determination unit is used to input the target scale feature into the classification module for classification processing to obtain the change state corresponding to the current pixel point.

[0095] Based on the above technical solution, the multi-scale feature fusion module includes: a cross-region association unit and a multi-scale feature fusion unit;

[0096] The target scale feature determination unit is specifically configured to: input multiple candidate scale features into a cross-region association unit for cross-region association to obtain multiple reference scale features corresponding to the neighborhood image; and input multiple reference scale features into a multi-scale feature fusion unit for multi-scale feature fusion to obtain target scale features corresponding to the neighborhood image.

[0097] On the basis of the above technical solution, the device further includes:

[0098] A comparison result determination module is used to compare the image change detection result with a preset change threshold to determine the comparison result corresponding to the IoT video to be analyzed;

[0099] The target recognition result determination module is used to obtain the target recognition result by performing target recognition based on the pre-trained target recognition model and the IoT video to be analyzed if the comparison result indicates that the IoT video needs to be analyzed and identified.

[0100] The image change detection device provided by the embodiment of the present invention can execute the image change detection method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the image change detection method.

[0101] It is worth noting that in the above-mentioned embodiment of image change detection, the various units and modules included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present invention.

[0102] Example 4

[0103] Figure 5A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0104] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0105] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0106] The processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the image change detection method.

[0107] In some embodiments, the image change detection method may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the image change detection method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the image change detection method in any other suitable manner (e.g., via firmware).

[0108] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0109] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0110] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0111] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0112] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0113] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.

[0114] An embodiment of the present invention further provides a computer program product, including a computer program, which, when executed by a processor, implements the image change detection method provided in any embodiment of the present application.

[0115] During the implementation of the computer program product, the computer program code for performing the operations of the present invention can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, using an Internet service provider to connect through the Internet). The program product and the image change detection method disclosed in the embodiments of the present application belong to the same inventive concept, so they will not be described here.

[0116] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.

[0117] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A method for detecting image changes, characterized in that: include: Acquire a first image and a second image in the IoT video to be analyzed, and perform image processing on the first image and the second image to determine a target image to be detected; Performing image segmentation on the target image based on a preset neighborhood size to determine a neighborhood image corresponding to each pixel in the target image; performing image change detection on each of the neighborhood images based on a pre-trained image change detection model to determine an image change detection result between the first image and the second image; The image change detection model is used to extract multi-scale features corresponding to neighborhood images, perform feature fusion on the multi-scale features, and perform classification based on the fused features.

2. The method according to claim 1, characterized in that The performing image processing on the first image and the second image to determine the target image to be detected includes: performing a difference analysis based on the first image and the second image to determine a difference image between the first image and the second image; Image synthesis is performed based on the first image, the second image, and the difference image to determine a target image to be detected.

3. The method according to claim 1, characterized in that The performing image change detection on each of the neighborhood images based on the pre-trained image change detection model to determine the image change detection result between the first image and the second image includes: The neighborhood image corresponding to each pixel is input into the pre-trained image change detection model to perform image change detection and obtain the change state corresponding to each pixel; Based on a preset display method and the change state corresponding to each pixel point in the target image output by the image change detection model, an image change detection result between the first image and the second image is generated.

4. The method according to claim 3, characterized in that The image change detection model includes: a multi-scale feature extraction module, a multi-scale feature fusion module and a classification module; The process of inputting the neighborhood image corresponding to each pixel into a pre-trained image change detection model to perform image change detection and obtain the change state corresponding to each pixel includes: For each pixel point, the neighborhood image corresponding to the current pixel point is input into the multi-scale feature extraction module for multi-scale feature extraction to obtain multiple candidate scale features corresponding to the neighborhood image; Inputting the multiple candidate scale features into the multi-scale feature fusion module for multi-scale feature fusion to obtain the target scale feature corresponding to the neighborhood image; The target scale feature is input into the classification module for classification processing to obtain the change state corresponding to the current pixel point.

5. The method according to claim 4, characterized in that The multi-scale feature fusion module includes: a cross-region association unit and a multi-scale feature fusion unit; Inputting the plurality of candidate scale features into the multi-scale feature fusion module for multi-scale feature fusion to obtain the target scale feature corresponding to the neighborhood image includes: Inputting the plurality of candidate scale features into the cross-region association unit for cross-region association to obtain a plurality of reference scale features corresponding to the neighborhood image; The multiple reference scale features are input into the multi-scale feature fusion unit for multi-scale feature fusion to obtain the target scale feature corresponding to the neighborhood image.

6. The method according to claim 1, wherein The method further comprises: Determine a comparison result corresponding to the IoT video to be analyzed based on a comparison between the image change detection result and a preset change threshold; If the comparison result indicates that the IoT video needs to be analyzed and identified, target recognition is performed based on a pre-trained target recognition model and the IoT video to be analyzed to obtain a target recognition result.

7. An image change detection device, characterized in that: The device comprises: a target image determination module, configured to obtain a first image and a second image in the IoT video to be analyzed, and perform image processing on the first image and the second image to determine a target image to be detected; A neighborhood image determination module is used to perform image segmentation on the target image based on a preset neighborhood size, and determine a neighborhood image corresponding to each pixel in the target image; an image change detection result determination module, configured to perform image change detection on each of the neighborhood images based on a pre-trained image change detection model, and determine an image change detection result between the first image and the second image; The image change detection model is used to extract multi-scale features corresponding to neighborhood images, perform feature fusion on the multi-scale features, and perform classification based on the fused features.

8. An electronic device, characterized in that: The electronic device comprises: one or more processors; a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the image change detection method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the image change detection method according to any one of claims 1 to 6 is implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the computer program implements the image change detection method according to any one of claims 1 to 6.