Image difference information detection method and device, and nonvolatile storage medium

By performing serialization and feature fusion on remote sensing images, images containing both global and local information are generated, solving the problem of low accuracy in remote sensing image detection and achieving higher detection accuracy.

CN116994142BActive Publication Date: 2025-12-26BEIJING INFORMATION SCI & TECH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311041597.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-17
Publication Date
2025-12-26
Estimated Expiration
2043-08-17

AI Technical Summary

Technical Problem

Existing technologies for analyzing remote sensing images ignore global information, resulting in low accuracy of detection results.

Method used

By acquiring remote sensing images of the target area at different times, serialization processing is performed to generate a global feature image. Based on the global feature image, multiple local feature images are generated, and feature equalization processing and fusion are performed to generate an image containing both global and local information.

Benefits of technology

This technology enables the simultaneous consideration of global and local information during remote sensing image detection and analysis, thereby improving the accuracy of detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116994142B_ABST
    Figure CN116994142B_ABST
Patent Text Reader

Abstract

The application discloses a kind of detection method and device of image difference information, nonvolatile storage medium.Therein, the method includes: obtaining first target image and second target image;First target image and second target image are serialized, and first feature image is obtained;Multiple second feature images are generated based on first feature image;First feature image and multiple second feature images are processed by feature equalization, and multiple third feature images are obtained;First feature image, multiple second feature images and third feature image are fused, and third target image is obtained, wherein, third target image includes: first type information and second type information, first type information is used to record the partial image of second target image relative to first target image change, and second type information is used to record the partial image of second target image relative to first target image unchanged.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, in particular to a method and device for detecting image difference information and a nonvolatile storage medium. BACKGROUND

[0002] With the rapid development of remote sensing sensors and space technology, a large number of high-resolution remote sensing images in space can be easily obtained; by analyzing the remote sensing images, the details of the region corresponding to the remote sensing images can be captured, which can meet various specific needs. In the related art, when analyzing the remote sensing images, the local information of the remote sensing images is mainly considered, therefore, the global change of the region corresponding to the remote sensing images cannot be captured, resulting in low accuracy of the analysis result.

[0003] At present, no effective solution has been proposed for the above problems. SUMMARY

[0004] Embodiments of the present application provide a method and device for detecting image difference information and a nonvolatile storage medium, to at least solve the technical problem of low accuracy of the detection result caused by ignoring the global information of the remote sensing images when detecting and analyzing the remote sensing images in the related art.

[0005] According to an aspect of an embodiment of the present application, a method for detecting image difference information is provided, comprising: obtaining a first target image and a second target image, wherein the first target image and the second target image are remote sensing images of a target region at different time instants; performing serialization processing on the first target image and the second target image to obtain a first feature image, wherein the first feature image is used to show the global feature information of the target region; generating a plurality of second feature images based on the first feature image, wherein each second feature image in the plurality of second feature images is used to show the change feature distribution information of the target region at different resolution scales; performing feature balancing processing on the first feature image and the plurality of second feature images to obtain a plurality of third feature images; and performing fusion processing on the first feature image, the plurality of second feature images and the third feature images to obtain a third target image, wherein the third target image includes: first type information and second type information, the first type information is used to record the part of the image that changes in the second target image relative to the first target image, and the second type information is used to record the part of the image that does not change in the second target image relative to the first target image.

[0006] Optionally, the serialization processing on the first target image and the second target image to obtain the first feature image comprises: superimposing the first target image and the second target image according to a preset superimposition rule to obtain an initial feature image; cutting the initial feature image into a plurality of pixel blocks of the same size; and performing serialization processing on the plurality of pixel blocks to obtain the first feature image.

[0007] Optionally, the plurality of pixel blocks are serialized to obtain the first feature image, including: determining a position code of each pixel block according to feature information of each pixel block in the plurality of pixel blocks, wherein the position code is used to indicate a spatial position of the pixel block in the initial feature image; arranging the plurality of pixel blocks into a target sequence according to the position code to obtain a plurality of target sequences which are completely same, and generating a plurality of target matrices according to the plurality of target sequences; processing the target matrices and the target sequence by using a first activation function to obtain a feature vector; restoring the feature vector to the first feature image according to the position code.

[0008] Optionally, a plurality of second feature images are generated based on the first feature image, wherein each second feature image in the plurality of second feature images is obtained by: performing a dilated convolution on an input feature image under different dilated parameters to obtain a plurality of fourth feature images with different information granularities, wherein the input feature image includes the first feature image and a target second feature image, and the target second feature image is a second feature image output by a previous position layer of a current position layer; performing weak and small target enhancement processing on the plurality of fourth feature images and the first feature image to obtain a first processing result; simultaneously, performing edge refinement processing on the plurality of fourth feature images to obtain a second processing result; determining a sum of the first processing result and the second processing result, and generating each second feature image according to the sum.

[0009] Optionally, the weak and small target enhancement processing is performed on the plurality of fourth feature images and the first feature image to obtain a first processing result, including: splicing any two fourth feature images in the plurality of fourth feature images into a fifth feature image to obtain a plurality of fifth feature images; performing a normalization operation on each fifth feature image in the plurality of fifth feature images to obtain a plurality of sixth feature images, wherein pixel values corresponding to pixel blocks constituting each sixth feature image belong to a same preset interval; processing the plurality of sixth feature images by using a second activation function to obtain a processed sixth feature image, wherein the processed sixth feature image is a sixth feature image with enhanced weak and small target information, and the processed sixth feature image is the first processing result.

[0010] Optionally, the edge refinement processing is performed on the plurality of fourth feature images to obtain a second processing result, including: obtaining a plurality of target difference values, wherein the plurality of target difference values are difference values of a third pixel value of a third type of pixel block and a fourth pixel value of a fourth type of pixel block at a same spatial position, the third type of pixel block and the fourth type of pixel block belong to any two fourth feature images, and the plurality of target difference values indicate edge information in the first feature image; performing convolution operation on a matrix composed of the plurality of target difference values and a preset matrix to obtain a convolution matrix, wherein an image corresponding to the convolution matrix is same as an image corresponding to the edge information in the first feature image, and the image corresponding to the convolution matrix is the second processing result.

[0011] Optionally, the first feature image and the plurality of second feature images are subjected to feature balancing processing to obtain a plurality of third feature images, including: fusing the first feature image and the plurality of second feature images according to a plurality of sets of preset parameters to obtain the plurality of third feature images, wherein each set of preset parameters in the plurality of sets of preset parameters includes: image width, image height, and image depth.

[0012] Optionally, the first feature image, the plurality of second feature images, and the third feature images are subjected to fusion processing to obtain a third target image, including: determining a plurality of target image groups, wherein each image group in the plurality of target image groups includes: the first feature image and the third feature image corresponding to the first feature image, or each second feature image and the third feature image corresponding to each second feature image; arranging the plurality of target image groups in a preset order to obtain an image group sequence; sequentially convolving the plurality of target image groups in the preset order to obtain a convolution result; and performing visualization processing on an image generated by the convolution result to obtain the third target image.

[0013] According to another aspect of the embodiments of the present application, a device for detecting image difference information is also provided, including: an acquisition module configured to acquire a first target image and a second target image, wherein the first target image and the second target image are remote sensing images of a target region at different time instants; a first processing module configured to perform serialization processing on the first target image and the second target image to obtain a first feature image, wherein the first feature image is used to show global feature information of the target region; a generation module configured to generate a plurality of second feature images based on the first feature image, wherein each second feature image in the plurality of second feature images is used to show change feature distribution information of the target region at different resolution scales; a second processing module configured to perform feature balancing processing on the first feature image and the plurality of second feature images to obtain a plurality of third feature images; and a third processing module configured to perform fusion processing on the first feature image, the plurality of second feature images, and the third feature images to obtain a third target image, wherein the third target image includes: first type information and second type information, the first type information is used to record a part of image of the second target image that has changed relative to the first target image, and the second type information is used to record a part of image of the second target image that has not changed relative to the first target image.

[0014] According to another aspect of the embodiments of the present application, a non-volatile storage medium is also provided, and the non-volatile storage medium stores a computer program, wherein a device in which the non-volatile storage medium is located executes the image difference information detection method described above by running the computer program.

[0015] According to another aspect of the embodiments of the present application, an electronic device is also provided, which comprises a memory and a processor, the memory stores a computer program, and the processor is configured to execute the image difference information detection method described above by using the computer program.

[0016] In the embodiments of the present application, the first target image and the second target image are obtained, wherein the first target image and the second target image are remote sensing images of the target region at different time points; the first target image and the second target image are sequentially processed to obtain a first feature image, wherein the first feature image is used to show the global feature information of the target region; a plurality of second feature images are generated based on the first feature image, wherein each second feature image in the plurality of second feature images is used to show the change feature distribution information of the target region at different resolution scales; the first feature image and the plurality of second feature images are processed by feature balancing to obtain a plurality of third feature images; the first feature image, the plurality of second feature images and the third feature images are fused to obtain a third target image, wherein the third target image includes first type information and second type information, the first type information is used to record the part of the image that changes in the second target image relative to the first target image, and the second type information is used to record the way of the part of the image that does not change in the second target image relative to the first target image. By constructing the global structural feature image and the local feature image of the remote sensing image sample, adjusting the weight to perform feature balancing on the global structural feature image and the local feature image, and fusing the global structural feature image, the local feature image and the feature balanced image at different scale features to obtain the image recording the change information of the remote sensing image, the purpose of considering the global information and the local information of the remote sensing image at the same time when detecting and analyzing the remote sensing image is achieved, thereby realizing the technical effect of improving the accuracy of identifying the change information of the remote sensing image, and further solving the technical problem of low detection result accuracy caused by ignoring the global information of the remote sensing image when the related technology detects and analyzes the remote sensing image. BRIEF DESCRIPTION OF DRAWINGS

[0017] The drawings described herein are used to provide further understanding of the present application, and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute improper limitations on the present application. In the drawings:

[0018] Figure 1 is a hardware structure block diagram of a computer terminal (or mobile device) for implementing the image difference information detection method according to the embodiments of the present application;

[0019] Figure 2 is a step flowchart of the image difference information detection method according to the embodiments of the present application;

[0020] Figure 3is a schematic diagram of generating a global feature image according to an embodiment of the present application;

[0021] Figure 4 is a schematic diagram of generating a local feature image according to an embodiment of the present application;

[0022] Figure 5 is a structural diagram of an image difference information detection device according to an embodiment of the present application;

[0023] Figure 6 is a working flow diagram of an image difference information detection device according to an embodiment of the present application. DETAILED DESCRIPTION

[0024] In order to make the personnel in the technical field better understand the scheme of the present application, the technical scheme in the embodiments of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by the person skilled in the art without creative labor should belong to the scope of protection of the present application.

[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0026] In order to better understand the embodiments of the present application, the technical terms involved in the embodiments of the present application are explained as follows:

[0027] Remote sensing image: an image of the earth's surface or atmosphere obtained by remote sensing technology. Remote sensing technology uses sensors on satellites, aircraft or other carriers to sense electromagnetic wave radiation of the earth's surface or atmosphere, and converts it into visual images or digital data. Remote sensing images can provide information about the earth's surface features, topography, vegetation, water bodies, urban buildings, etc. and information about atmospheric composition, cloud layers, temperature, etc.

[0028] Information granularity: the level of detail of the information contained in an image. The finer the information granularity, the more details the image contains.

[0029] Weak target enhancement processing: the weak target enhancement processing in the embodiments of the present application is an image processing technology for improving the visibility and recognizability of weak targets in an image; a weak target is a target in an image that is smaller in size, lower in contrast, and contains less image details than a background or other targets.

[0030] Edge thinning: edge information of an image is extracted and processed to obtain clearer edge information.

[0031] In the related art, a remote sensing image is detected and analyzed by using a convolutional neural network or a generative adversarial network (GAN network). When the remote sensing image is detected and analyzed by using the convolutional neural network, the local features of the remote sensing image are extracted for analysis, and the global features of the remote sensing image are ignored. Therefore, the accuracy of the detection result obtained by detecting the remote sensing image by using the convolutional neural network is low. The training process of the GAN network is relatively complex, and includes a game process between a generator and a discriminator. Such a game often leads to instability of the training process, and problems such as mode collapse or mode collapse are prone to occur. Therefore, the detection and analysis of the remote sensing image by using the GAN network has the problem of instability. In order to solve the above problems, the present application provides related solutions, which are described in detail below.

[0032] According to the embodiments of the present application, a method embodiment of an image difference information detection method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0033] The method embodiment provided by the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing the image difference information detection method is shown. As shown in Figure 1 The computer terminal 10 (or mobile device 10) can include one or more processors 102 (the processor 102 can include but is not limited to a microprocessor MCU or a programmable logic device FPGA processing device), a memory 104 for storing data, and a transmission device 106 for communication function. In addition, it can also include a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. Those skilled in the art can understand that Figure 1The illustrated architecture is merely an example and does not impose a limitation on the architecture of the electronic device. For example, the computer terminal 10 can further include more or fewer components than those shown, or have a different configuration of components than those shown. Figure 1 Figure 1

[0034] It should be noted that the one or more processors 102 and / or other data processing circuitry described above can be generally referred to herein as "data processing circuitry". The data processing circuitry can be embodied in whole or in part as software, hardware, firmware, or any combination thereof. In addition, the data processing circuitry can be a single standalone processing module, or incorporated in whole or in part within any of the other elements of the computer terminal 10 (or mobile device). As referred to in embodiments of the present application, the data processing circuitry serves as a processor to control, for example, the selection of the variable resistance terminal path connected to the interface.

[0035] The memory 104 can be used to store software programs of application software and modules, such as program instructions / data storage means corresponding to the image difference information detection method in embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, i.e. implements the image difference information detection method described above. The memory 104 can include a high-speed random access memory, and can further include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory 104 can further include a memory remotely arranged with respect to the processor 102, which can be connected to the computer terminal 10 through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0036] The transmission device 106 is configured to receive or send data via a network. Examples of the network include, but are not limited to, a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network interface controller (NIC) which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module configured to communicate with the Internet in a wireless manner.

[0037] The display can be, for example, a touch screen type liquid crystal display (LCD) which enables a user to interact with the user interface of the computer terminal 10 (or mobile device).

[0038] In the above-described operating environment, embodiments of the present application provide an image difference information detection method,​​Figure 2 is a step flow chart of a method for detecting image difference information according to an embodiment of the present application, as shown in the figure, the method comprises the following steps: Figure 2

[0039] In step S202, a first target image and a second target image are acquired, wherein the first target image and the second target image are remote sensing images of a target region at different time instants.

[0040] The embodiments of the present application provide a network model for detecting change information of remote sensing images. The network model is used for detecting and analyzing any two remote sensing images and finally outputting difference information of the two remote sensing images. In step S202, the network model for detecting change information of remote sensing images acquires two remote sensing images (i.e., the first target image and the second target image) of the target region at different time instants.

[0041] In step S204, the first target image and the second target image are subjected to serialization processing to obtain a first feature image, wherein the first feature image is used to show global feature information of the target region.

[0042] In step S204, after the two remote sensing images (i.e., the first target image and the second target image) of the target region at different time instants are acquired in step S202, the two remote sensing images (i.e., the first target image and the second target image) are subjected to serialization processing: the two remote sensing images (i.e., the first target image and the second target image) at different time instants are cut into a plurality of pixel blocks of the same size, and the plurality of pixel blocks of the same size are combined into a (first) feature image showing global feature information of the target region.

[0043] According to an optional embodiment of the present application, the serialization processing of the first target image and the second target image to obtain the first feature image comprises: superimposing the first target image and the second target image according to a preset superimposition rule to obtain an initial feature image; cutting the initial feature image into a plurality of pixel blocks of the same size; and subjecting the plurality of pixel blocks to serialization processing to obtain the first feature image.

[0044] Figure 3 is a schematic diagram of generating a global feature image, as shown in the figure, Figure 3 ​As shown, generating a global feature image (i.e., the first feature image) includes two steps: slicing and combining. In this embodiment, two remote sensing images from different times (i.e., the first target image and the second target image) are sliced ​​into multiple pixel blocks of the same size. The slicing method is as follows: Before slicing the image, the two remote sensing images from different times (i.e., the first target image and the second target image) are first superimposed into one image according to a preset superposition rule. In this embodiment, the superposition rule is to superimpose according to the channel dimension of the image. The two images are superimposed into one (initial) feature image according to the superposition order of top and bottom. The superimposed (initial) feature image is then sliced ​​into multiple pixel blocks of the same size. The multiple pixel blocks are then combined and serialized to generate a global feature image (i.e., the first feature image).

[0045] Optionally, the serialization process of multiple pixel blocks to obtain a first feature image includes: determining the position code of each pixel block based on the feature information of each pixel block in the multiple pixel blocks, wherein the position code is used to indicate the spatial position of the pixel block in the initial feature image; arranging the multiple pixel blocks into a target sequence according to the position code to obtain multiple identical target sequences, and generating multiple target matrices based on the multiple target sequences; processing the target matrices and target sequences using a first activation function to obtain feature vectors; and restoring the feature vectors to the first feature image based on the position code.

[0046] like Figure 3 As shown, the method for combining multiple pixel blocks into a global feature image (i.e., the first feature image) that displays global feature information of the target region is as follows: First, the input sequence composed of multiple pixel blocks is input into the embedding layer of a network model used to detect change information in remote sensing images. The embedding layer converts each pixel block in the input sequence into a high-dimensional vector. Simultaneously, the semantic tagger in the network model extracts feature information F from the high-dimensional vector of each pixel block. in ; where the feature information F of each pixel block in This includes the location information of the region corresponding to the pixel block within the target region, as well as the size information of the region corresponding to the pixel block; a dynamic encoding method is used, based on formula P. ε =MLP(GELU(g) (3*3) (MLP(F in The process represented by )))) represents the feature information F in The location information in the network model is extracted and enhanced sequentially through a multi-layer perceptron (MLP) and a convolutional kernel (g). (3*3) The convolution operator and Gaussian error linear unit activation function (GELU) are used to process the data to obtain the position code P of each pixel block. ε. Each pixel block is arranged according to its corresponding position encoding P ε The pixel blocks are arranged according to the position encoding P The multiple target matrices and the (target) sequence (for example, sequence V) that does not generate the target matrix are processed by a normalization exponential function (σ) to output a global feature image (i.e., a first feature image). The purpose of processing the multiple target matrices by the normalization exponential function (σ) is to use an attention mechanism for the network model, so that the network model only focuses on the distance relationship between two pixel blocks when processing the pixel blocks at this step, thereby improving the accuracy of the output result. For example, according to the position encoding P ε After the pixel blocks are arranged to generate three identical (target) sequences Q, K, and V, the activation function σ takes the Q matrix as the query vector, the K matrix as the key vector, and the V matrix as the value vector, and processes the matrices Q, K, and V in the manner shown in formula V, and finally outputs a global feature image (i.e., a first feature image). Before outputting the global feature image (i.e., the first feature image), the query vector Q, the key vector K, and the value vector V are arranged according to the position encoding P T is the bias matrix of the matrix K, d k is the dimension of the matrix K, and the finally output image is an image that shows the global feature information of the target region (i.e., the first feature image).

[0047] Step S206, generating a plurality of second feature images based on the first feature image, wherein each second feature image in the plurality of second feature images is used to show the change feature distribution information of the target region at different resolution scales.

[0048] In step S206, a plurality of local feature images (i.e., second feature images) are generated based on the image (i.e., the first feature image) obtained in step S204, which shows the global feature information of the target region, wherein the plurality of local feature images (i.e., the second feature images) respectively show the distribution information of the local features of the target region at different resolutions, and the local feature information of the target region includes the information of the region (i.e., the edge region) with large gray value change in the remote sensing image of the target region, the information of the corner (such as the corner of an object or the intersection) region in the remote sensing image of the target region, and the information of the spot region (a region with small area and obvious difference in gray value from the surrounding) in the remote sensing image of the target region, etc. The clarity of the local feature information shown at different resolutions is different, so in step S206, the plurality of local feature images (i.e., the second feature images) are used to respectively show the plurality of local feature information of the target region.

[0049] According to an optional embodiment of the present application, the plurality of second feature images are generated based on the first feature image, wherein each of the plurality of second feature images is obtained by the following method: performing dilated convolution on the input feature image at different dilated parameters to obtain a plurality of fourth feature images with different information granularities, wherein the input feature image includes the first feature image and a target second feature image, and the target second feature image is the second feature image output by the previous position layer of the current position layer; performing weak and small target reinforcement processing on the plurality of fourth feature images and the first feature image to obtain a first processing result; at the same time, performing edge thinning processing on the plurality of fourth feature images to obtain a second processing result; determining the sum of the first processing result and the second processing result, and generating each second feature image according to the sum.

[0050] Figure 4 is a schematic diagram of generating a local feature image, and the flow of generating a local feature image based on a global feature image (i.e., a first feature image) is as follows Figure 4As shown, in the embodiment, firstly, the feature image as input is dilated convolutional using different dilated convolutional parameters (Dilated) to obtain the local feature images (i.e., the fourth feature images) generated based on different dilated convolutional parameters (Dilated) convolution, wherein the length, width and information granularity of the plurality of local feature images (i.e., the fourth feature images) generated based on different dilated convolutional parameters (Dilated) convolution are different. Wherein, if the current position layer is the first layer for generating the local feature image, the last position layer is the position layer for generating the global feature image (i.e., the first feature image), at this time, the input feature image of the current position layer is the global feature image (i.e., the first feature image); if the current position layer is the second layer or other position layer above the second layer for generating the local feature image, the last position layer is the position layer for generating the local feature image (i.e., the second feature image), at this time, the input feature image of the current position layer is the local feature image (i.e., the second feature image) output by the last position layer. For example Figure 4 As shown, the global feature image (i.e., the first feature image) is dilated convolutional when the dilated convolutional parameter is Dilated0 to obtain the (fourth) feature image f c The global feature image (i.e., the first feature image) is dilated convolutional when the dilated convolutional parameter is Dilated1 to obtain the (fourth) feature image f w Next, the plurality of local feature images (i.e., the fourth feature images) with different information granularities are subjected to weak and small target enhancement processing and edge refinement processing to obtain the feature image after weak and small target enhancement processing (i.e., the first processing result) and the feature image after edge refinement processing (i.e., the second processing result); the pixel values of the corresponding pixel points in the feature image after weak and small target enhancement processing (i.e., the first processing result) and the feature image after edge refinement processing (i.e., the second processing result) are added, and the image generated according to the addition result is determined as the image (i.e., the second feature image) showing the local features of the target region.

[0051] According to another optional embodiment of the present application, the plurality of fourth feature images and the first feature image are subjected to weak and small target enhancement processing to obtain the first processing result, comprising: splicing any two fourth feature images in the plurality of fourth feature images into a fifth feature image to obtain a plurality of fifth feature images; performing normalization operation processing on each fifth feature image in the plurality of fifth feature images to obtain a plurality of sixth feature images, wherein the pixel values corresponding to the pixel blocks constituting each sixth feature image belong to the same preset interval; using a second activation function to process the plurality of sixth feature images to obtain a processed sixth feature image, wherein the processed sixth feature image is a sixth feature image with enhanced weak and small target information, and the processed sixth feature image is the first processing result.

[0052] In the embodiment, the small target enhancement processing of the local feature image includes the following steps: as shown in the formula, the feature images f Figure 4 and f c are processed by small target enhancement processing. w When the feature images f c and f w are processed by small target enhancement processing, the two feature images f c and f w are first merged into one feature image (i.e., the fifth feature image), and then the merged image (i.e., the fifth feature image) is normalized by the normalization layer (Batch Normalization, BN) of the network model used for detecting the change information of the remote sensing image, so that the pixel values of each pixel block in the merged feature image belong to the same preset pixel value interval, such as the interval [0, 1], so as to enhance the feature information of the merged image. Further, the network model uses a parameterized non-saturated function (i.e., the second activation function) (Parametric Rectified Linear Unit, PRELU) to extract the (first) pixel value of the pixel block (i.e., the first type of pixel block) of the small target image in the fifth feature image (i.e., the sixth feature image) after normalization processing; and the (first) pixel value of the pixel block corresponding to the small target image is multiplied by the (second) pixel value of the pixel block (i.e., the second type of pixel block) in the global feature image (i.e., the first feature image). Through the above steps, the features of the small target image in the global feature image are enhanced, and the technical effect of small target enhancement is realized. Finally, the global feature image after enhancing the small target information is generated as the (first) processing result output by the small target enhancement processing. The above process of small target enhancement based on the global feature image can be summarized as the formula F m = PRELU(BN(Cat(conv(Y n-1 ,W c ), conv(Y n-1 ,W w ))), wherein F m is used to extract the feature information of the edge image, Cat represents the connection operation, which is used to merge the two feature images generated by the convolution of the input global feature image Y n-1 into one image, conv represents the convolution operation, W c and W w are two different parameters of the dilated convolution.

[0053] According to some optional embodiments of the present application, the plurality of fourth feature images are subjected to edge thinning processing to obtain a second processing result, including: obtaining a plurality of target difference values, wherein the plurality of target difference values are difference values of third pixel values of third pixel blocks and fourth pixel values of fourth pixel blocks at the same spatial position, the third pixel blocks and the fourth pixel blocks belong to any two fourth feature images respectively, and the plurality of target difference values indicate edge information in the first feature image; and performing convolution operation on a matrix composed of the plurality of target difference values and a preset matrix to obtain a convolution matrix, wherein the image corresponding to the convolution matrix is the same as the image corresponding to the edge information in the first feature image, and the image corresponding to the convolution matrix is the second processing result.

[0054] As shown in Figure 4 , when the feature images f c and f w are subjected to edge thinning processing, first, pixel difference values (i.e., target difference values) of the feature images f c and f w are calculated, wherein when the pixel difference values of the feature images f c and f w are calculated, the pixel values of two pixel blocks (i.e., third pixel blocks and fourth pixel blocks) belonging to f c and f w at the same spatial position of the two feature images are subtracted. When the pixel difference values of the feature images f c and f w are calculated, the pixel difference values (i.e., target difference values) of the same images in f c and f w are 0, indicating the part that has not changed; and the pixel difference values (i.e., target difference values) of different images in f c and f w are not 0, and the image corresponding to the part of the pixel difference values (i.e., target difference values) that are not 0 is the edge information to be extracted. The edge information in the global feature image is extracted by subtracting f c and f w , and the matrix composed of the extracted edge information is subjected to convolution processing, and finally the enhanced edge feature information is output. When the matrix composed of the extracted edge information is subjected to convolution processing, it includes twice convolution processing, once normalization processing, and once activation function processing. For example, the matrix composed of the (third) pixel values of the (third type) pixel blocks in the feature image f c is F c , and the matrix composed of the (fourth) pixel values of the (fourth type) pixel blocks in the feature image f w is F w ; then according to the formula F edg = g (3*3) (σ(BN(g (3*3) (|FC -F w |))))indicated processing process, the matrix F c and the difference value matrix |F w of the matrix F C -F w |is convolved with the convolution kernel matrix (i.e., the preset matrix) g (3*3) , and the convolution result is normalized in the normalization layer (BN); the edge feature information is extracted again by using the normalization exponential function (i.e., the first activation function) σ, and the extracted feature information is convolved again with the convolution kernel matrix (i.e., the preset matrix) g (3*3) , to obtain the feature information F edg of the enhanced edge image.

[0055] In step S208, the first feature image and the plurality of second feature images are subjected to feature balancing processing to obtain a plurality of third feature images.

[0056] In step S208, after the global feature information image (i.e., the first feature image) and the plurality of local feature information images (i.e., the second feature images) of the target region are obtained, the global feature information image (i.e., the first feature image) and the plurality of local feature information images (i.e., the second feature images) are fused into one (third) feature image which includes both the global feature information and the local feature information of the target region.

[0057] According to an optional embodiment of the present application, the first feature image and the plurality of second feature images are subjected to feature balancing processing to obtain a plurality of third feature images, which includes: fusing the first feature image and the plurality of second feature images according to a plurality of preset parameters to obtain a plurality of third feature images, wherein each of the plurality of preset parameters includes: image width, image height, and image depth.

[0058] In this embodiment, the global feature information image (i.e., the first feature image) and the plurality of local feature information images (i.e., the second feature images) are fused into one (third) feature image according to different weights according to the pre-defined weight parameters (i.e., the preset parameters); wherein the pre-defined weight parameters (i.e., the preset parameters) are determined by the width (W) of the image, the height (H) of the image, and the depth of the image in the image, which are three dimensional information of the image, and the preset weight parameter is the product of the width (W) of the image, the height (H) of the image, and the depth of the image. For example, there is one global feature image (i.e., the first feature image) and four local feature images (i.e., the second feature images), and when generating the fused image (i.e., the third feature image) corresponding to the global feature image (i.e., the first feature image), the weight of the global feature image (i.e., the first feature image) is set to and the remaining four local feature images (i.e., the second feature images) are set to The global feature image (i.e., the first feature image) and the local feature image (i.e., the second feature image) are fused into a global feature image (i.e., the first feature image) corresponding to the fusion image (i.e., the third feature image) according to the set weight. It should be noted that the weight parameters set when generating the fusion image corresponding to different feature images (including global feature images and local feature images) are different.

[0059] In step S210, the first feature image, the plurality of second feature images and the third feature image are fused to obtain a third target image, wherein the third target image includes: first type information and second type information, the first type information is used to record the part of the second target image that changes relative to the first target image, and the second type information is used to record the part of the second target image that does not change relative to the first target image.

[0060] In step S210, the global feature image (i.e., the first feature image), the local feature image (i.e., the second feature image) and the feature balanced fusion image (i.e., the third feature image) are subjected to multi-scale feature fusion processing, and an image (third target) including both change information (i.e., first type information) and unchanged information (i.e., second type information) is output, wherein the change information (i.e., first type information) is information of a target region that changes in two remote sensing images at different times, for example, changes in artificial objects (changes in human activities), changes in natural objects (vegetation, seasons, etc.), changes in mixed objects (regional changes), etc. The unchanged information (i.e., second type information) is information of a target region that does not change in two remote sensing images at different times.

[0061] According to another optional embodiment of the present application, the first feature image, the plurality of second feature images and the third feature image are fused to obtain a third target image, including: determining a plurality of target image groups, wherein each image group in the plurality of target image groups includes: the first feature image and the third feature image corresponding to the first feature image, or each second feature image and the third feature image corresponding to each second feature image; arranging the plurality of target image groups in a predetermined order to obtain an image group sequence; sequentially convolving the plurality of target image groups according to the predetermined order to obtain a convolution result; and visualizing the image generated by the convolution result to obtain the third target image.

[0062] The global feature image (i.e., the first feature image) is a low-level feature map containing coarser-grained global information, the local feature image (i.e., the second feature image) is a high-level feature map containing finer-grained local information, and the fused feature image (i.e., the third feature image) is a feature map with balanced granularity. Therefore, in some embodiments, the network model used to detect changes in remote sensing images fuses the three types of images with different scales—the global feature image (i.e., the first feature image), the local feature image (i.e., the second feature image), and the feature-balanced image (the third feature image)—through convolution according to a pre-set fusion order, and finally outputs the fused result (i.e., the third target image). For example, the feature images to be fused include: the global feature image (i.e., the first feature image) D1, the local feature images (i.e., the second feature images) D2 to D5, and the (third) feature images corresponding to D1 to D5 respectively, obtained by performing feature balance processing on D1 to D5. Then proceed in the following order to process D1 through D5 and Merge: Combine D4, D5 and Perform convolution and output the convolution result. Then with D3 and Perform convolution and output the convolution result. Will With D2 and Perform convolution and output the convolution result. Finally With D1 and Output the final convolution result. The corresponding image (i.e., the third target image) is the detection result of detecting changes in the remote sensing image of the target area at different times. It should be noted that the method provided in this application embodiment also applies to the image generated based on the detection result (i.e.... The corresponding image is visualized so that the image generated based on the detection results (i.e. In the corresponding image, changed and unchanged information are displayed in different colors. For example, for the image generated based on the detection results (i.e. The corresponding image is binarized so that changed information is displayed in white and unchanged information is displayed in black.

[0063] By following the steps described above, both global and local feature information of the target region in the remote sensing image can be considered as influencing factors when performing change detection on the remote sensing image. This avoids the problem of ignoring global feature information when performing change detection on remote sensing images in existing technologies, and improves the accuracy of the detection results.

[0064] Figure 5This is a structural diagram of an image difference information detection device according to an embodiment of this application, such as... Figure 5 As shown, the detection device includes: an acquisition module 50, used to acquire a first target image and a second target image, wherein the first target image and the second target image are remote sensing images of the target area at different times; a first processing module 52, used to perform serialization processing on the first target image and the second target image to obtain a first feature image, wherein the first feature image is used to display global feature information of the target area; a generation module 54, used to generate multiple second feature images based on the first feature image, wherein each of the multiple second feature images is used to display the distribution information of the changing features of the target area at different resolution scales; a second processing module 56, used to perform feature equalization processing on the first feature image and the multiple second feature images to obtain multiple third feature images; and a third processing module 58, used to perform fusion processing on the first feature image, the multiple second feature images, and the third feature images to obtain a third target image, wherein the third target image includes: a first type of information and a second type of information, wherein the first type of information is used to record the part of the second target image that has changed relative to the first target image, and the second type of information is used to record the part of the second target image that has not changed relative to the first target image.

[0065] Figure 6 This is a schematic diagram of the workflow of an image difference information detection device, such as... Figure 6 As shown, the acquisition module 50 acquires two remote sensing images (remote sensing image 1 and remote sensing image 2) of the target area at different times. These two images are then superimposed and input into the first processing module 52 for serialization processing to generate a global feature image D1. The first processing module 52 saves the global feature image D1 and inputs copies of D1 into four generation modules 54 to obtain local feature images D2, D3, D4, and D5. The second processing module 56 performs feature equalization processing on images D1 to D5 according to preset weight parameters, obtaining multiple images that fuse global and local feature information (i.e., the third feature image). Finally, the third processing module 58 processes the global feature image D1, the local feature images D2 to D5, and the feature equalization output image (i.e., the third feature image). Perform fusion processing, following the order of D1 to D5 and Merge: Combine D4, D5 and Perform convolution and output the convolution result. Then with D3 and Perform convolution and output the convolution result. Will With D2 and Perform convolution and output the convolution result. Finally With D1 and Output the final convolution result. The corresponding image (i.e., the third target image); at the same time, the third processing module 58 will... The corresponding image (i.e., the third target image) is binarized, and the detection result is output, such as... Figure 6 As shown in the image corresponding to the detection results, the white part represents information that has changed after comparing two remote sensing images, and the black part represents information that has not changed after comparing two remote sensing images.

[0066] It should be noted that, Figure 5 Preferred embodiments of the shown examples can be found in [reference needed]. Figure 2 The relevant descriptions of the embodiments shown will not be repeated here.

[0067] This application also provides a non-volatile storage medium storing a computer program, which is used to execute the above-described image difference information detection method on the device where the non-volatile storage medium is located by running the computer program.

[0068] The aforementioned non-volatile storage medium is used to store a program that performs the following functions: acquiring a first target image and a second target image, wherein the first target image and the second target image are remote sensing images of a target region at different times; performing serialization processing on the first target image and the second target image to obtain a first feature image, wherein the first feature image is used to display global feature information of the target region; generating multiple second feature images based on the first feature image, wherein each of the multiple second feature images is used to display the distribution information of the changing features of the target region at different resolution scales; performing feature equalization processing on the first feature image and the multiple second feature images to obtain multiple third feature images; and performing fusion processing on the first feature image, the multiple second feature images, and the third feature images to obtain a third target image, wherein the third target image includes: a first type of information and a second type of information, wherein the first type of information is used to record the part of the second target image that has changed relative to the first target image, and the second type of information is used to record the part of the second target image that has not changed relative to the first target image.

[0069] This application also provides an electronic device, including a memory and a processor. The memory stores a computer program, and the processor is configured to execute the above-described method for detecting image difference information through the computer program.

[0070] The processor in the electronic device is configured to run a program for performing the following functions: obtaining a first target image and a second target image, wherein the first target image and the second target image are remote sensing images of a target region at different time instants; performing serialization processing on the first target image and the second target image to obtain a first feature image, wherein the first feature image is used to show global feature information of the target region; generating a plurality of second feature images based on the first feature image, wherein each second feature image in the plurality of second feature images is used to show change feature distribution information of the target region at different resolution scales; performing feature balancing processing on the first feature image and the plurality of second feature images to obtain a plurality of third feature images; and performing fusion processing on the first feature image, the plurality of second feature images, and the third feature images to obtain a third target image, wherein the third target image includes first type information and second type information, the first type information is used to record a part of the second target image that has changed relative to the first target image, and the second type information is used to record a part of the second target image that has not changed relative to the first target image.

[0071] It should be noted that each module in the image difference information detection device described above can be a program module (for example, a set of program instructions that implement a certain specific function) or a hardware module. For the latter, it can be in the following form, but is not limited to this: the form of each module described above is a processor, or the functions of each module described above are implemented by a processor.

[0072] The serial numbers of the embodiments of the present application described above are only for description, and do not represent the advantages or disadvantages of the embodiments.

[0073] In the above embodiments of the present application, the description of each embodiment has its own emphasis, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0074] In the several embodiments provided by the present application, it should be understood that the disclosed technology can be implemented in other ways. Of course, the device embodiment described above is only schematic. For example, the division of the units can be a logical function division, and there can be another division manner in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, and can be electrical or other forms.

[0075] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, may be located in one place, or may be distributed to multiple units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0076] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present alone, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0077] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the part that contributes to the related art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various program code storage media.

[0078] The above is only the preferred embodiment of the present application, and it should be pointed out that for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, which should be considered as the protection scope of the present application.

Claims

1. A method of detecting image difference information, characterized by, The method comprises the following steps: obtaining a first target image and a second target image, wherein the first target image and the second target image are remote sensing images of a target region at different time instants; performing serialization processing on the first target image and the second target image to obtain a first feature image, comprising: superimposing the first target image and the second target image according to a preset superimposition rule to obtain an initial feature image; cutting the initial feature image into a plurality of pixel blocks of the same size; determining the position encoding of each pixel block in the plurality of pixel blocks according to the feature information of the pixel block, wherein the position encoding is used to indicate the spatial position of the pixel block in the initial feature image; arranging the plurality of pixel blocks into a target sequence according to the position encoding to obtain a plurality of target sequences which are completely the same, and generating a plurality of target matrices according to the plurality of target sequences; processing the target matrices and the target sequences by using a first activation function to obtain a feature vector; restoring the feature vector to the first feature image according to the position encoding, wherein the first feature image is used to show the global feature information of the target region; generating a plurality of second feature images based on the first feature image, wherein each second feature image in the plurality of second feature images is used to show the change feature distribution information of the target region at different resolution scales; performing feature equalization processing on the first feature image and the plurality of second feature images to obtain a plurality of third feature images; performing fusion processing on the first feature image, the plurality of second feature images and the third feature images to obtain a third target image, wherein the third target image comprises first information and second information, the first information is used to record the part of the second target image that has changed relative to the first target image, and the second information is used to record the part of the second target image that has not changed relative to the first target image.

2. The method of claim 1, wherein, generating a plurality of second feature images based on the first feature image, wherein each second feature image in the plurality of second feature images is obtained by the following method: performing dilated convolution on an input feature image under different dilated parameters to obtain a plurality of fourth feature images with different information granularities, wherein the input feature image comprises the first feature image and a target second feature image, and the target second feature image is a second feature image output by a previous position layer of a current position layer; performing weak and small target enhancement processing on the plurality of fourth feature images and the first feature image to obtain a first processing result; at the same time, performing edge refinement processing on the plurality of fourth feature images to obtain a second processing result; determining the sum of the first processing result and the second processing result, and generating each second feature image according to the sum.

3. The method of claim 2, wherein, performing weak and small target enhancement processing on the plurality of fourth feature images and the first feature image to obtain a first processing result, comprising: splicing any two fourth feature images in the plurality of fourth feature images into a fifth feature image to obtain a plurality of fifth feature images; The normalization operation is performed on each of the plurality of fifth feature images to obtain a plurality of sixth feature images, wherein pixel values of pixel blocks constituting each of the plurality of sixth feature images belong to a same preset interval; The second activation function is used to process the plurality of sixth feature images to obtain a processed sixth feature image, wherein the processed sixth feature image is a sixth feature image in which weak target information is enhanced, and the processed sixth feature image is the first processing result.

4. The method of claim 2, wherein, The plurality of fourth feature images are subjected to edge thinning processing to obtain a second processing result, including: A plurality of target difference values are obtained, wherein the plurality of target difference values are difference values of third pixel values of third pixel blocks and fourth pixel values of fourth pixel blocks at a same spatial position, the third pixel blocks and the fourth pixel blocks belong to any two of the fourth feature images respectively, and the plurality of target difference values indicate edge information in the first feature image; A matrix composed of the plurality of target difference values is subjected to convolution operation with a preset matrix to obtain a convolution matrix, wherein an image corresponding to the convolution matrix is the same as an image corresponding to the edge information in the first feature image, and the image corresponding to the convolution matrix is the second processing result.

5. The method of claim 1, wherein, The first feature image and the plurality of second feature images are subjected to feature equalization processing to obtain a plurality of third feature images, including: The first feature image and the plurality of second feature images are fused according to a plurality of sets of preset parameters to obtain the plurality of third feature images, wherein each set of preset parameters in the plurality of sets of preset parameters includes an image width, an image height, and an image depth.

6. The method of claim 2, wherein, The first feature image, the plurality of second feature images, and the third feature images are fused to obtain a third target image, including: A plurality of target image groups are determined, wherein each image group in the plurality of target image groups includes the first feature image and a third feature image corresponding to the first feature image, or each second feature image and a third feature image corresponding to each second feature image; The plurality of target image groups are arranged in a preset order to obtain an image group sequence; The plurality of target image groups are sequentially convolved in the preset order to obtain a convolution result; An image generated by the convolution result is subjected to visual processing to obtain the third target image.

7. An image difference information detection device, characterized in that, including: An acquisition module is configured to acquire a first target image and a second target image, wherein the first target image and the second target image are remote sensing images of a target region at different time instants. The first processing module is configured to perform serialization processing on the first target image and the second target image to obtain a first feature image, including: superimposing the first target image and the second target image according to a preset superimposition rule to obtain an initial feature image; cutting the initial feature image into a plurality of pixel blocks of the same size; determining a position code of each pixel block in the plurality of pixel blocks according to feature information of the each pixel block, wherein the position code is used to indicate a spatial position of the pixel block in the initial feature image; arranging the plurality of pixel blocks into a target sequence according to the position code to obtain a plurality of target sequences which are completely same, and generating a plurality of target matrices according to the plurality of target sequences; processing the target matrices and the target sequences by using a first activation function to obtain a feature vector; and restoring the feature vector to the first feature image according to the position code, wherein the first feature image is used to show global feature information of the target region. The generation module is configured to generate a plurality of second feature images based on the first feature image, wherein each second feature image in the plurality of second feature images is used to show change feature distribution information of the target region at different resolution scales. The second processing module is configured to perform feature equalization processing on the first feature image and the plurality of second feature images to obtain a plurality of third feature images. The third processing module is configured to perform fusion processing on the first feature image, the plurality of second feature images and the third feature images to obtain a third target image, wherein the third target image includes first type information and second type information, the first type information is used to record a partial image of the second target image which changes relative to the first target image, and the second type information is used to record a partial image of the second target image which does not change relative to the first target image.

8. A non-volatile storage medium, comprising: The non-volatile storage medium stores a computer program, wherein a device in which the non-volatile storage medium is located performs the image difference information detection method in any one of claims 1 to 6 by running the computer program.

Citation Information

Patent Citations

  • Image processing method and device, electronic equipment and storage medium

    CN109829501A

  • Image change detection method, image change training method, image change detection device, image change training device and computer equipment

    CN114419406A