Method for change detection of images and related device
By extracting the feature of the pictures in the change detection and building the attention matrix combination, the problems of large amount of calculation and low inference efficiency in the deep learning model in the prior art are solved, and efficient change detection is achieved.
Patent Information
- Application Number
- CN202210403181.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-18
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-04-18
AI Technical Summary
Among the existing change detection methods, the deep learning model has a complex structure and high computational volume, which leads to low inference efficiency and is difficult to implement and apply.
By extracting the first picture and the second picture respectively, multi-level feature pairs are constructed, and the feature pairs at each level are convolutional to generate attention matrix combinations, spatial and temporal feature combinations are constructed, and fusion operations are performed to obtain multi-level fusion features, and a change map is obtained through downsampling and upsampling.
On the premise of ensuring high precision, the calculation amount of the model is reduced, the inference efficiency of the model is improved, and the model can be easily implemented and applied to their respective scenarios.
Smart Images

Figure CN115131284B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology. Specifically, it relates to a method for detecting changes in pictures and related devices. Background Art
[0002] Change detection is one of the important tasks in the field of computer vision. The purpose of this task of change detection is to detect and segment the semantic changes existing between two pictures taken at different times.
[0003] Currently, the proposed change detection solutions mainly adopt deep learning models. However, the model structures adopted by these solutions are complex and the computational cost is high, which seriously restricts the implementation and application of the solutions. Summary of the Invention
[0004] Embodiments of this application provide a method for detecting changes in pictures and related devices, which can, at least to a certain extent, reduce the computational cost of the model while ensuring high accuracy, thereby improving the inference efficiency of the model.
[0005] Other features and advantages of this application will become apparent through the following detailed description, or be learned in part through the practice of this application.
[0006] According to one aspect of the embodiments of this application, a method for detecting changes in pictures is provided. The method includes: respectively performing feature extraction on a first picture and a second picture to obtain multi-level feature pairs, where the feature pairs include a first feature extracted from the first picture and a second feature extracted from the second picture; for each level of feature pairs, respectively performing a convolution operation on the first feature and the second feature in the feature pairs to obtain a first attention matrix combination corresponding to the first feature and a second attention matrix combination corresponding to the second feature, and constructing a spatio-temporal feature combination corresponding to the feature pairs according to the first attention matrix combination and the second attention matrix combination corresponding to each level of feature pairs, where the spatio-temporal feature combination includes a first spatio-temporal feature corresponding to the first attention matrix combination and a second spatio-temporal feature corresponding to the second attention matrix combination; respectively performing a fusion operation on the spatio-temporal feature combinations corresponding to each level of feature pairs to obtain multi-level fusion features; performing a downsampling operation on the multi-level fusion features to obtain a downsampling result, and performing an upsampling operation on the downsampling result to obtain a change map.
[0007] According to one aspect of the embodiments of the present application, there is provided a device for detecting changes in pictures. The device includes: a feature extraction unit configured to respectively extract features from a first picture and a second picture to obtain multi-level feature pairs, where the feature pairs include a first feature extracted from the first picture and a second feature extracted from the second picture; a spatio-temporal feature construction unit configured to, for each level of feature pairs, respectively perform convolution operations on the first feature and the second feature in the feature pairs to obtain a first attention matrix combination corresponding to the first feature and a second attention matrix combination corresponding to the second feature, and construct a spatio-temporal feature combination corresponding to the feature pairs according to the first attention matrix combination and the second attention matrix combination corresponding to each level of feature pairs, where the spatio-temporal feature combination includes a first spatio-temporal feature corresponding to the first attention matrix combination and a second spatio-temporal feature corresponding to the second attention matrix combination; a fusion unit configured to respectively perform fusion operations on the spatio-temporal feature combinations corresponding to each level of feature pairs to obtain multi-level fusion features; and a sampling unit configured to perform downsampling operations on the multi-level fusion features to obtain a downsampling result, and perform upsampling operations on the downsampling result to obtain a change map.
[0008] In some embodiments of the present application, based on the foregoing solution, the fusion unit is configured to: for the spatio-temporal feature combination corresponding to each level of feature pairs, respectively perform two convolution operations on the spatio-temporal feature combination to respectively obtain a third attention matrix combination and a fourth attention matrix combination, where both the third attention matrix combination and the fourth attention matrix combination include multiple attention matrices, and the same attention matrix in the third attention matrix combination and the fourth attention matrix combination is obtained by performing convolution operations on different spatio-temporal features in the spatio-temporal feature combination; construct a first cross-interaction feature corresponding to each level of feature pairs according to the third attention matrix combination corresponding to each spatio-temporal feature combination, and construct a second cross-interaction feature corresponding to each level of feature pairs according to the fourth attention matrix combination corresponding to each spatio-temporal feature combination; determine semantic change information between the first cross-interaction feature and the second cross-interaction feature corresponding to each level of feature pairs, and use the semantic change information corresponding to each level of feature pairs as the fusion features of each level.
[0009] In some embodiments of the present application, based on the aforementioned scheme, the change map is generated by a main branch of a change detection model, the change detection model also includes an auxiliary branch for auxiliary training of the main branch, the change detection model is trained by generating a multi-level affine map in the auxiliary branch, and the device also includes an auxiliary training unit, and the affine map of the target level in the multi-level affine map is generated by the auxiliary training unit performing the following process: embedding the first feature and the second feature in the feature pair of the target level respectively to obtain a first feature embedding result and a second feature embedding result, and generating a channel correlation matrix based on the first feature embedding result and the second feature embedding result; shielding some elements in the channel correlation matrix by masking to obtain a shielded matrix; extracting the maximum value of each row in the shielded matrix to obtain a first maximum similarity map; extracting the maximum value of each column in the shielded matrix to obtain a second maximum similarity map; fusing the first maximum similarity map and the second maximum similarity map to obtain an affine map of the target level.
[0010] In some embodiments of the present application, based on the aforementioned scheme, the auxiliary training unit is configured to: extract the maximum value of the elements at each position in the first maximum similarity graph and the second maximum similarity graph as the value of the element at the position in the target maximum similarity graph to obtain the target maximum similarity graph; perform a correction operation on the target maximum similarity graph to obtain an affine graph of the target level.
[0011] In some embodiments of the present application, based on the aforementioned scheme, the auxiliary training unit is configured to: perform a mapping operation on the elements in the target maximum similarity graph to obtain a mapping graph; and perform a resizing operation on the mapping graph to obtain an affine graph of the target level.
[0012] In some embodiments of the present application, based on the aforementioned scheme, the change detection model is trained according to a loss function, the loss function includes a main branch loss corresponding to the main branch and an auxiliary branch loss corresponding to the auxiliary branch, and the main branch loss includes a cross entropy loss and a Diess loss.
[0013] In some embodiments of the present application, based on the aforementioned scheme, the feature extraction unit is configured to: perform feature extraction on the first image and the second image respectively to obtain feature pairs of a predetermined number of levels; eliminate feature pairs of a specified level from the feature pairs of the predetermined number of levels to obtain multi-level feature pairs, wherein the specified level is lower than the level corresponding to the multi-level feature pairs.
[0014] In some embodiments of the present application, based on the foregoing solution, the spatio-temporal feature construction unit is configured to: perform multiple groups of convolution operations on the first feature in the feature pair respectively to obtain multiple first attention matrix combinations corresponding to the first feature; perform multiple groups of convolution operations on the second feature in the feature pair respectively to obtain multiple second attention matrix combinations corresponding to the second feature.
[0015] In some embodiments of the present application, based on the foregoing solution, the feature extraction unit is configured to: perform feature extraction on the first picture and the second picture respectively to obtain multi-level original feature pairs; superimpose position encoding features on the first original feature and the second original feature in each level of the original feature pairs respectively to obtain multi-level feature pairs.
[0016] According to one aspect of the embodiments of the present application, there is provided a computer-readable medium having a computer program stored thereon, and when the computer program is executed by a processor, it implements the method for change detection of pictures as described in the above embodiments.
[0017] According to one aspect of the embodiments of the present application, there is provided an electronic device, including: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method for change detection of pictures as described in the above embodiments.
[0018] According to one aspect of the embodiments of the present application, there is provided a computer program product, the computer program product includes computer instructions, the computer instructions are stored in a computer-readable storage medium, a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method for change detection of pictures as described in the above embodiments.
[0019] In the technical solutions provided by some embodiments of the present application, by first extracting features from a pair of pictures to obtain multi-level feature pairs, and then performing convolution operations on the first feature and the second feature in the feature pairs of each level respectively, a first attention matrix combination corresponding to the first feature and a second attention matrix combination corresponding to the second feature are obtained. A spatio-temporal feature combination corresponding to each feature pair is constructed according to the first attention matrix combination and the second attention matrix combination corresponding to the feature pairs of each level; on this basis, by performing fusion operations on the spatio-temporal feature combinations respectively, multi-level fusion features are obtained, and downsampling operations and upsampling operations are sequentially performed on the multi-level fusion features, thereby obtaining a change map. Therefore, the solution of the embodiment of the present application constructs a first attention matrix combination corresponding to the first feature and a second attention matrix combination corresponding to the second feature in each feature pair, and constructs a spatio-temporal feature combination corresponding to each feature pair according to the first attention matrix combination and the second attention matrix combination corresponding to the feature pairs of each level. Since the spatio-temporal feature combination corresponding to each feature pair is constructed based on the attention matrix combinations corresponding to the feature pairs of each level, the perception of global context information is realized, and feature encoding can be performed more efficiently. Thus, on the premise of ensuring high accuracy, the computational amount of the model is reduced, and the inference efficiency of the model is significantly improved, enabling the model to be conveniently applied to respective scenarios.
[0020] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the accompanying drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts. In the drawings:
[0022] Figure 1 An example diagram showing a change detection task is shown;
[0023] Figure 2 An example diagram showing the difficulties of change detection is shown;
[0024] Figure 3 A schematic diagram showing the model architecture of a change detection model with shallow fusion in the related art is shown;
[0025] Figure 4 A schematic diagram showing the model architecture of a change detection model with deep fusion in the related art is shown;
[0026] Figure 5Shows a schematic diagram of the model architecture of the change detection model with multi-layer fusion in the related art;
[0027] Figure 6 Shows a schematic diagram of the architecture of the HPCFNet model in the related art;
[0028] Figure 7 Shows a schematic diagram of the architecture of the change detection model according to an embodiment of the present application;
[0029] Figure 8 Shows a schematic diagram of an exemplary system architecture to which the technical solution of the embodiment of the present application can be applied;
[0030] Figure 9 Shows a flowchart of the change detection method for pictures according to an embodiment of the present application;
[0031] Figure 10 Shows a schematic diagram of the overall framework of the change detection model according to an embodiment of the present application;
[0032] Figure 11 Shows according to an embodiment of the present application Figure 9 The flowchart of the details of step 910;
[0033] Figure 12 Shows according to another embodiment of the present application Figure 9 The flowchart of the details of step 910;
[0034] Figure 13 Shows a schematic diagram of the structure of the self-interaction sub-module according to an embodiment of the present application;
[0035] Figure 14 Shows a flowchart of obtaining multi-level fusion features through a fusion operation according to an embodiment of the present application;
[0036] Figure 15 Shows a schematic diagram of the structure of the cross-interaction sub-module according to an embodiment of the present application;
[0037] Figure 16 Shows a flowchart of generating the affine graph of the target level in the multi-level affine graph according to an embodiment of the present application;
[0038] Figure 17 Shows according to an embodiment of the present application Figure 16 The flowchart of the details of step 1650;
[0039] Figure 18 Shows a schematic diagram of the effect comparison between the method of the embodiment of the present application and the related art;
[0040] Figure 19 The block diagram of a change detection device for a picture according to an embodiment of the present application is shown;
[0041] Figure 20 The structural schematic diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown. Detailed implementation manners
[0042] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.
[0043] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be used. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.
[0044] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0045] The flowcharts shown in the drawings are only illustrative and do not necessarily include all the contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.
[0046] The purpose of the change detection task is to detect and segment the semantic changes existing between two pictures taken at different times.
[0047] Figure 1 An example diagram of the change detection task is shown. Please refer to Figure 1 As shown, the pictures taken at time T1 and the pictures taken at time T2 are analyzed and processed using the change detection method, and finally a change diagram is obtained.
[0048] Since the changes in images taken at different times usually include semantic changes and various non-semantic changes, such as non-semantic changes including illumination changes, shadow changes, seasonal changes, and perspective changes, etc. Therefore, how to distinguish semantic changes from non-semantic changes and measure semantic changes is crucial for change detection. Figure 2 The example diagram showing the difficulties of change detection can be seen that non-semantic changes such as sunlight, shadow, and season are very crucial in change detection.
[0049] Change detection has many application scenarios, such as land cover monitoring, medical diagnosis, urbanization rate analysis, intelligent transportation, and industrial quality inspection, etc. In recent years, many deep learning methods have emerged to solve the change detection problem. Most of the methods of related technologies focus on using Convolutional Neural Network (CNN) to automatically extract and fuse spatio-temporal features at multiple levels.
[0050] Dividing the existing change detection methods from the perspective of fusion strategies, they can generally be divided into three categories of solutions: shallow fusion, deep fusion, and multi-layer fusion, as Figures 3 - 5 shown.
[0051] Figure 3 The schematic diagram of the model architecture of the change detection model with shallow fusion in related technologies is shown. Through Figure 3 it can be seen that the shallow fusion solution fuses features at a shallow position (such as the input position).
[0052] Figure 4 The schematic diagram of the model architecture of the change detection model with deep fusion in related technologies is shown. Through Figure 4 it can be seen that the deep fusion solution fuses features at a deep position (such as the high level of CNN).
[0053] Figure 5 The schematic diagram of the model architecture of the change detection model with multi-layer fusion in related technologies is shown. Through Figure 5 it can be seen that the multi-layer fusion solution fuses features at multiple levels (multiple levels of CNN) in order to make full use of multi-level feature information.
[0054] These solutions all have certain defects, such as low accuracy, large computational amount, etc.
[0055] In addition, related technologies also proposed an HPCFNet model that can complete the change detection task. This model is also a change detection method based on a multi-level fusion strategy. Figure 6 The schematic diagram of the architecture of the HPCFNet model in related technologies is shown.
[0056] Please refer toFigure 6 As shown in the figure, HPCFNet adopts a siamese network structure. Multiple convolutional layers (conv1_2, conv2_2, conv3_3, conv4_3, conv5_3) are used to extract multi-layer semantic features from picture t0 and picture t1 respectively. Then, the PCF (Paired Channel Fusion) module is used for feature fusion at the channel level. At the same time, the MPFL (Multi-Part Feature Learning) module is used to integrate features from the whole to the local and adapt to various change distributions. Finally, the output change segmentation map is obtained.
[0057] However, the design of the dual-channel fusion module in this method is too complex, using a large number of parallel dilated convolutions and cross-feature stacking operations, which makes the method inefficient, lacks the perception of global context information, has insufficient robustness to non-semantic changes, and cannot be widely applied to actual scenarios.
[0058] Therefore, this application first provides a method for detecting changes in pictures. The method for detecting changes in pictures provided based on the embodiments of this application can overcome the above defects, strengthen the algorithm's perception of global context information, reduce the computational amount of the model, significantly improve the inference efficiency of the model, and enable the model to be conveniently applied to respective scenarios; at the same time, it can also improve the robustness to non-semantic changes, and the model has high accuracy.
[0059] Figure 7 The schematic diagram of the architecture of the change detection model according to an embodiment of this application is shown. Please refer to Figure 7 As shown in the figure, the change detection model proposed in this application has an SSI module, which is a semantic interaction module based on the spatial level. It can be seen that each convolutional layer is correspondingly provided with an SSI module. When each SSI module is processing, it can interact with other SSI modules to capture long-range dependencies and strengthen the algorithm's perception of global information. The SSI module will be introduced in the subsequent part.
[0060] Figure 8 The schematic diagram of an exemplary system architecture to which the technical solution of the embodiments of this application can be applied is shown. As Figure 8 shown, the system architecture 800 may include a server 801, a network 802, a camera 803, and a display screen 804. The camera 803 continuously takes pictures of a certain location in the city. The camera 803 can communicate with the server 801 through the network 802. The server 801 has established a communication connection with the display screen 804. The server 801 is the implementation terminal of the embodiments of this application, and the change detection model provided by the embodiments of this application is deployed on it. When the method for detecting changes in pictures provided by this application is applied toFigure 8 In the system architecture shown, a process can be as follows: First, the camera 803 continuously takes pictures of the city and uploads the pictures taken of the city to the server 801 through the network 802; then, after the server 801 obtains a pair of pictures taken by the camera 803 at different times, it can analyze and process this pair of pictures by using a change detection model to obtain a change map; finally, the server 801 sends the change map to the display screen 804, and the change map is displayed on the display screen 804. In addition, the pictures based on which the change map is generated, the analysis results of the changed areas in the change map, etc. can also be displayed on the display screen 804.
[0061] In some embodiments of the present application, the server 801 will regularly perform statistical analysis on the change maps generated within a period of time and generate an analysis report.
[0062] It should be understood that Figure 8 the numbers of the server, network, camera, and display screen in are merely illustrative. According to the implementation requirements, there can be any number of servers, networks, cameras, and display screens. For example, the server 801 can be a server cluster composed of multiple servers, etc.
[0063] It should be noted that Figure 8 only one embodiment of the present application is shown. Although in the Figure 8 solution of the embodiment, a pair of pictures provided to the change detection model are taken by the same camera, but in other embodiments of the present application, a pair of pictures provided to the change detection model can be taken by two cameras respectively; although Figure 8 the solution of the embodiment is applied to the urban management scenario, but in fact, the solution of the embodiment of the present application can be applied to various scenarios. For example, it can be applied to the intelligent transportation scenario, and even can be applied to the industrial scenario; although Figure 8 in the solution of the embodiment, a pair of pictures provided to the change detection model are taken by the camera and transmitted to the server, but in other embodiments of the present application, the sources of the pictures provided to the change detection model can be various. For example, they can be uploaded by users, or can also be computer-generated. The embodiments of the present application do not make any limitations in this regard, and the protection scope of the present application should not be limited thereby.
[0064] It is easy to understand that the method for detecting changes in pictures provided by the embodiments of the present application is generally executed by the server. Correspondingly, the device for detecting changes in pictures is generally set in the server. However, in other embodiments of the present application, the terminal device can also have a similar function as the server, so as to execute the solution for detecting changes in pictures provided by the embodiments of the present application.
[0065] Therefore, the embodiments of the present application can be applied to a terminal or a server. The server can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and the present application does not make any restrictions in this regard.
[0066] The implementation details of the technical solutions of the embodiments of the present application are elaborated in detail below:
[0067] Figure 9 The flowchart of the picture change detection method according to an embodiment of the present application is shown. The picture change detection method can be executed by various devices capable of computing and processing, such as a user terminal or a cloud server. The user terminal includes, but is not limited to, a mobile phone, a computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, a wearable device, etc. The embodiments of the present application can be applied to various scenarios, including, but not limited to, cloud technology, artificial intelligence, intelligent transportation, assisted driving, etc. Please refer to Figure 9 As shown, the picture change detection method at least includes the following steps:
[0068] In step 910, feature extraction is respectively performed on the first picture and the second picture to obtain multi-level feature pairs. The feature pair includes a first feature extracted from the first picture and a second feature extracted from the second picture.
[0069] By performing feature extraction on the first picture, multi-level first features can be obtained; by performing feature extraction on the second picture, multi-level second features can be obtained. The first feature and the second feature at the same level are combined into a feature pair at one level, so that multi-level feature pairs can be obtained. The number of levels of the first feature and the second feature is the same, for example, both can be 5. It is easy to understand that other numbers of levels can be set according to actual needs.
[0070] The first picture and the second picture can be pictures generated or taken at different times. The first picture and the second picture are usually pictures taken in the same scene, and the shooting angles of the first picture and the second picture can be different.
[0071] The first image and the second image can be images in various formats such as.jpg,.bmp,.png, etc. The first image and the second image can be photos taken, or images obtained by computer processing and output, and can even be video frames in a video file. The first image and the second image can contain various things such as buildings, vehicles, people, landscapes, items, etc.
[0072] Figure 10 The overall framework schematic diagram of the change detection model according to an embodiment of the present application is shown. Next, the solution of the embodiment of the present application will be described in conjunction with Figure 10 Please refer to Figure 10 As shown, the change detection model includes a Siamese CNN Extractor, a Spatial-wise Semantic Interaction Module (SSI), a Change Perception Head (CPH), a Channel-wise Semantic Interaction Module (CSI), and Joint Learning.
[0073] The Siamese CNN Extractor includes a group of CNN feature extraction networks. The network structures of this group of CNN feature extraction networks can be the same, and all CNN feature extraction networks can share parameters. The Siamese CNN Extractor can be implemented by various networks such as the VGG network, the Inception network, and the ResNet network.
[0074] Given a pair of images collected at time t0 and time t1 Using the Siamese CNN Extractor, multi-level feature pairs can be extracted and where is the first image collected at time t0, is the second image collected at time t1, 3 is the number of channels of the image, H is the height of the image, W is the width of the image, is the multi-level first feature corresponding to the first image, is the multi-level second feature corresponding to the second image, i is the level, and 5 represents that the total number of levels is 5.
[0075] Figure 11 The details of step 910 according to an embodiment of the present application are shown in Figure 9 Please refer to Figure 11 As shown, step 910 can specifically include the following steps:
[0076] In step 911, feature extraction is performed on the first picture and the second picture respectively to obtain feature pairs at a predetermined number of levels.
[0077] The predetermined number can correspond to the number of convolutional layers in the Siamese CNN feature extractor.
[0078] In step 912, the feature pairs at a specified level among the feature pairs at a predetermined number of levels are removed to obtain multi-level feature pairs, where the specified level is lower than the levels corresponding to the multi-level feature pairs.
[0079] The specified level can be one or more levels. For example, the specified level can be the lowest level, or the lowest level and the second lowest level.
[0080] Please refer to Figure 10 As shown, the Siamese CNN feature extractor outputs feature pairs at 4 levels and This represents the first feature and the second feature at the lowest level being removed.
[0081] The semantic information of low-level features is weak and the computational load is large. In the embodiments of the present application, by removing the feature pairs at the specified level with the lowest level, the computational load can be greatly reduced while ensuring the accuracy of the model.
[0082] Figure 12 shows the details of step 910 according to another embodiment of the present application Figure 9 Please refer to Figure 12 As shown, step 910 may specifically further include the following steps:
[0083] In step 911', feature extraction is performed on the first picture and the second picture respectively to obtain multi-level original feature pairs.
[0084] By performing feature extraction on the first picture and the second picture respectively to obtain original feature pairs for subsequent operations.
[0085] In step 912', position encoding features are respectively superimposed on the first original feature and the second original feature in each level of the original feature pairs to obtain multi-level feature pairs.
[0086] The position encoding features are used to encode the position information of the features.
[0087] Please continue to refer to Figure 10As shown, the semantic interaction module based on the spatial level is the core module of the change detection model, which is used to explore the global context information and enhance the interaction between two spatio-temporal features. The semantic interaction module based on the spatial level includes a self-interaction sub-module (self), a cross-interaction sub-module (cross), and a spatio-temporal feature fusion sub-module. Among them, the input of the semantic interaction module based on the spatial level is a feature pair at a certain level. and Then, the first feature in this feature pair and the second feature need to be respectively superimposed with PE (Positional Encoding), and then the superimposed results are input into the self-interaction sub-module. PE is the positional encoding feature.
[0088] In the embodiment of the present application, by superimposing the positional encoding feature on the basis of the original feature, the accuracy of the model is improved.
[0089] Please continue to refer to Figure 9 , in step 920, for each level of feature pair, a convolution operation is respectively performed on the first feature and the second feature in the feature pair to obtain a first attention matrix combination corresponding to the first feature and a second attention matrix combination corresponding to the second feature, and a spatio-temporal feature combination corresponding to the feature pair is constructed according to the first attention matrix combination and the second attention matrix combination corresponding to each level of feature pair. The spatio-temporal feature combination includes a first spatio-temporal feature corresponding to the first attention matrix combination and a second spatio-temporal feature corresponding to the second attention matrix combination.
[0090] In this step, each level of feature pair needs to be processed to obtain the corresponding spatio-temporal feature combination. By performing a convolution operation on the first feature in the feature pair, a first attention matrix combination corresponding to the first feature is obtained; by performing a convolution operation on the second feature in the feature pair, a second attention matrix combination corresponding to the second feature is obtained. When constructing the spatio-temporal feature combination corresponding to each feature pair, it is necessary to according to the first attention matrix combination and the second attention matrix combination corresponding to all levels of feature pairs. Specifically, when constructing the first spatio-temporal feature in the spatio-temporal feature combination corresponding to each feature pair, it is necessary to according to the first attention matrix combination corresponding to all levels of feature pairs; when constructing the second spatio-temporal feature in the spatio-temporal feature combination corresponding to each feature pair, it is necessary to according to the second attention matrix combination corresponding to all levels of feature pairs. Therefore, when constructing the spatio-temporal feature at a certain level, it is necessary to pay attention to the attention matrix combinations at all levels. By constructing the spatio-temporal feature combination in this way, the perception of global context information can be realized.
[0091] The first attention matrix combination and the second attention matrix combination may both include a query matrix, a key matrix, and a value matrix. Among them, the first attention matrix combination may include a first query matrix, a first key matrix, and a first value matrix, and the second attention matrix combination may include a second query matrix, a second key matrix, and a second value matrix.
[0092] In one embodiment of the present application, performing convolution operations on the first feature and the second feature in the feature pair respectively to obtain the first attention matrix combination corresponding to the first feature and the second attention matrix combination corresponding to the second feature includes: performing multiple groups of convolution operations on the first feature in the feature pair respectively to obtain multiple first attention matrix combinations corresponding to the first feature; performing multiple groups of convolution operations on the second feature in the feature pair respectively to obtain multiple second attention matrix combinations corresponding to the second feature.
[0093] Each group of convolution operations in the multiple groups of convolution operations performed on the first feature or the second feature is carried out with different convolution parameters.
[0094] Therefore, in the embodiment of the present application, by performing multiple groups of convolution operations on the first feature and the second feature respectively, the multi-head self-attention mechanism is realized, and the ability of the model to focus on features at different levels is improved.
[0095] This step can be implemented by Figure 10 the self-interaction sub-module in the semantic interaction module based on the spatial level in Figure 13 as shown. Figure 13 FIG. shows a schematic structural diagram of the self-interaction sub-module according to an embodiment of the present application. In Figure 13 the self-interaction sub-module specifically performs the following operations on the first feature in the feature pair:
[0096] First, initialize and generate an initial query matrix through the following expression an initial key matrix and an initial value matrix
[0097]
[0098] The above expression represents that the initial query matrix the initial key matrix and the initial value matrix are initialized to the first feature at a certain level that is, the first feature at a certain level is respectively assigned to the initial query matrix the initial key matrix and the initial value matrix
[0099] Then, the initial query matrix is respectively convolved through the following expressions the initial key matrix and the initial value matrix as follows:
[0100]
[0101] where and are the parameters of the 1x1 convolution for the i-th head, is the parameter of the 1x1 convolution corresponding to the initial query matrix is the parameter of the 1x1 convolution corresponding to the initial key matrix is the parameter of the 1x1 convolution corresponding to the initial value matrix is the query matrix, is the key matrix, is the value matrix, which together form the first attention matrix combination.
[0102] Next, according to Figure 13 the attention map is calculated in the following way, which is equivalent to the affinity and softmax in
[0103]
[0104] where the softmax function is calculated along the channel dimension, and d is a preset value, such as the number of dimensions of the feature.
[0105] Subsequently, an aggregation operation is performed based on the following function to calculate the first spatio-temporal feature
[0106]
[0107] where is the first spatio-temporal feature, is the attention map,
[0108] is the value matrix.
[0109] Similarly, by performing the above operations on the second feature, the second spatio-temporal feature can be obtained.
[0110] In step 930, a fusion operation is respectively performed on the spatio-temporal feature combinations corresponding to the feature pairs at each level to obtain multi-level fusion features.
[0110] By performing a fusion operation on the spatio-temporal feature combination corresponding to a feature pair at one level, a fusion feature at one level can be obtained. Thus, since throughFigure 13 The self-interaction sub-module shown can directly obtain spatio-temporal features, so the cross-interaction sub-module can be not set.
[0111] Figure 14 The flowchart of obtaining multi-level fusion features through a fusion operation according to an embodiment of the present application is shown. Please refer to Figure 14 As shown, the following steps can be included:
[0112] In step 1410, for the spatio-temporal feature combination corresponding to each level of feature pairs, two convolution operations are respectively performed on the spatio-temporal feature combination to obtain a third attention matrix combination and a fourth attention matrix combination respectively, where both the third attention matrix combination and the fourth attention matrix combination include multiple attention matrices, and the same attention matrices in the third attention matrix combination and the fourth attention matrix combination are obtained by performing convolution operations on different spatio-temporal features in the spatio-temporal feature combination respectively.
[0113] Performing one convolution operation on the spatio-temporal feature combination can obtain the third attention matrix combination; performing another convolution operation on the spatio-temporal feature combination can obtain the fourth attention matrix combination. The third attention matrix combination can include a third query matrix, a third key matrix, and a third value matrix, and the fourth attention matrix combination can include a fourth query matrix, a fourth key matrix, and a fourth value matrix. The same attention matrices in the third attention matrix combination and the fourth attention matrix combination are the query matrix, the key matrix, or the value matrix.
[0114] In step 1420, the first cross-interaction features corresponding to each level of feature pairs are constructed according to the third attention matrix combination corresponding to each spatio-temporal feature combination, and the second cross-interaction features corresponding to each level of feature pairs are constructed according to the fourth attention matrix combination corresponding to each spatio-temporal feature combination.
[0115] Step 1410 and step 1420 can be implemented through a cross-interaction sub-module. Figure 15 The structural schematic diagram of the cross-interaction sub-module according to an embodiment of the present application is shown. Please refer to Figure 15 As shown, by respectively performing two 1x1 convolution processes on the first spatio-temporal feature the corresponding query matrix is obtained and the key matrix Performing one 1x1 convolution process on the second spatio-temporal feature the corresponding value matrix is obtained Next, based on the query matrix the key matrix and the value matrix the same operation is performed to obtain the second cross-interaction feature Similarly, for the second spatio-temporal feature obtained by Perform two 1x1 convolution operations respectively to obtain the corresponding query matrix and key matrix Perform a 1x1 convolution operation on the first spatio-temporal feature to obtain the corresponding value matrix Next, based on the query matrix key matrix and value matrix perform the same operation to obtain the first cross-interaction feature
[0116] Through Figure 13 and Figure 15 It can be seen that after the aggregation operation, the aggregation result is also superimposed with the first feature and the first spatio-temporal feature respectively for cross-layer connection. The role of this cross-layer connection is to facilitate model training and gradient backpropagation, so as to converge quickly.
[0117] In step 1430, determine the semantic change information between the first cross-interaction feature and the second cross-interaction feature corresponding to the feature pairs at each level, and use the semantic change information corresponding to the feature pairs at each level as the fusion feature at each level.
[0118] After obtaining a pair of cross-interaction features, the fusion feature can be obtained through the following fusion strategy:
[0119]
[0120] Among them, is the first cross-interaction feature, is the second cross-interaction feature, is the fusion feature.
[0121] It can be seen that in the above embodiment of the present application, the absolute value of the subtraction of the two cross-interaction features is used to measure the semantic change. This method can not only intuitively reflect the semantic change, is relatively symmetric, but also has low computational complexity.
[0122] Please continue to refer to Figure 10 , in the semantic interaction module based on the spatial level, there is also a spatio-temporal feature fusion sub-module located after the cross-interaction sub-module. The spatio-temporal feature fusion sub-module performs corresponding operations according to the following expression: In this expression, by calculating the difference between the first cross-interaction feature of the i-th level and the second cross-interaction feature of the i-th level, the fusion feature of the i-th level is obtained, so as to obtain the fusion features of multiple levels
[0123] In the embodiments of the present application, by performing convolution operations on different spatio-temporal features in the spatio-temporal feature combination to obtain the same attention matrix in the third attention matrix combination and the fourth attention matrix combination, information interaction between different spatio-temporal features is achieved, and feature encoding can be performed more precisely, thereby further improving the model efficiency.
[0124] In step 940, downsampling operations are performed on the multi-level fusion features to obtain a downsampling result, and upsampling operations are performed on the downsampling result to obtain a change map.
[0125] Both the downsampling operation and the upsampling operation may include multiple sub-operations, that is, layer-by-layer downsampling and upsampling are performed.
[0126] In an embodiment of the present application, performing downsampling operations on the multi-level fusion features to obtain a downsampling result includes: starting from the fusion features of the lowest level, performing downsampling operations on the fusion features of each level respectively, and when performing downsampling operations on the fusion features of each level, fusing the downsampling intermediate results corresponding to the level one lower than the current level to obtain the downsampling result.
[0127] Step 940 is implemented by Figure 10 the change perception head in. The change perception head adopts a pyramid structure. Specifically, starting from the shallowest fusion features beginning, with a stride of 2, a convolutional layer with a convolution kernel size of 3x3 is used for downsampling to obtain the corresponding downsampling intermediate result; then, the downsampling intermediate result is merged with and continue to perform downsampling on the merged result to obtain the corresponding downsampling intermediate result, and so on, until the downsampling result is obtained.
[0128] In order to obtain a high-resolution feature map, the upsampling operation adopts an operation opposite to the downsampling operation. Specifically, assuming the downsampling result is perform 1x1 convolution on the downsampling result to obtain the deepest upsampling intermediate result; perform Rectified Linear Unit (Relu) and batch normalization processing on the deepest upsampling intermediate result, merge the processed result with the downsampling intermediate result of the corresponding level to obtain a merged result, and perform upsampling operations on the merged result to obtain the second deepest upsampling intermediate result, and so on, until the upsampling result is obtained.
[0129] Through Figure 10 It can be seen that two 1x1 convolution operations are performed on the upsampling result to obtain a convolution result. In addition, the convolution result can be upsampled once again to finally obtain the change map.
[0130] In one embodiment of the present application, the change map is generated by the main branch of the change detection model. The change detection model further includes an auxiliary branch for assisting in training the main branch. The change detection model is trained by generating multi-level affine maps in the auxiliary branch. The affine map of the target level in the multi-level affine maps is generated by Figure 16 the process shown.
[0131] The main branch of the change detection model includes a semantic interaction module based on the spatial level and a change perception head. The auxiliary branch of the change detection model includes a semantic interaction module based on the channel level.
[0132] Figure 16 The flowchart of generating the affine map of the target level in the multi-level affine maps according to one embodiment of the present application is shown. Please refer to Figure 16 shown, and may include the following steps:
[0133] In step 1610, embedding operations are respectively performed on the first feature and the second feature in the feature pair of the target level to obtain a first feature embedding result and a second feature embedding result, and a channel correlation matrix is generated according to the first feature embedding result and the second feature embedding result.
[0134] Specifically, the first feature embedding result and the second feature embedding result can be implemented through a fully connected layer. For any level of feature pair ( and ), the channel correlation matrix is calculated by using a fully connected layer in the following manner:
[0135]
[0136] Among them, is the first feature, is the second feature, ψ and φ are linear embedding functions provided by the fully connected layer, is the channel correlation matrix, and the size of the channel correlation matrix is hw×hw, where h is the height of the first feature or the second feature, and w is the width of the first feature or the second feature.
[0137] The channel correlation matrix effectively focuses on the and commonality, representing the correlation between each pixel in the first feature and each pixel in the second feature. Figure 10 The shown in i w i ×h i w i .
[0138] In step 1620, some elements in the channel correlation matrix are masked through a mask to obtain a masked matrix.
[0139] In an embodiment of the present application, masking some elements in the channel correlation matrix through a mask to obtain a masked matrix includes: masking the elements on the diagonal of the channel correlation matrix and multiple pairs of elements that are symmetric about the diagonal and adjacent to the elements on the diagonal through a mask to obtain a masked matrix.
[0140] Specifically, the elements on the diagonal and the elements within the neighborhood of the diagonal can be masked.
[0141] The diagonal represents the one-to-one correlation at the same position in a pair of features. Considering the presence of perspective changes in change detection, the mask E is used for restriction, so that only the diagonal and its neighborhood need to be concerned about instead of the entire This can improve the efficiency of model training.
[0142] Specifically, the following operations are performed:
[0143]
[0144] Among them, is the channel correlation matrix, E is the diagonal unit matrix for amplifying Δ pixels, ⊙ represents pixel-level dot multiplication, S is the masked matrix, and the i-th row in S represents the correlation between the i-th pixel in and the i-th pixel in and its Δ neighborhood. Figure 10 The S shown in i is the masked matrix at the i-th level.
[0145] In step 1630, the maximum value of each row in the masked matrix is extracted to obtain the first maximum similarity map.
[0146] The first maximum similarity map is obtained by performing a maximum operation in the horizontal direction as follows:
[0147]
[0148] Among them, is the first maximum similarity map, and its size is h i w i ×1, and i and j are the positions of the elements in the masked matrix.
[0149] In step 1640, the maximum value of each column in the masked matrix is extracted to obtain the second maximum similarity map.
[0150] The second largest similar graph is obtained by performing a maximum operation in the vertical direction as follows:
[0151]
[0152] Among them, is the second largest similar graph, and i and j are the positions of the elements in the masked matrix.
[0153] In step 1650, the first largest similar graph and the second largest similar graph are fused to obtain the affine graph of the target level.
[0154] The affine graph of the target level is obtained by combining the first largest similar graph and the second largest similar graph.
[0155] Although high-level features contain rich semantic features, they lack detailed information. Using multi-scale features can alleviate this problem, but low-level features often contain a large amount of irrelevant noise.
[0156] The inventors found that non-changing regions often have a certain similarity in appearance and features regardless of the influence of illumination or other noises. In the embodiments of the present application, when performing model training, two largest similar graphs are first obtained and then fused to obtain the affine graph. Since only the similar parts are considered, it is possible to compactly represent features while suppressing the influence of noise.
[0157] Figure 17 Shows the flowchart of the details of step 1650 according to an embodiment of the present application Figure 16 in.
[0158] Please refer to Figure 17 shown. Step 1650 may specifically include the following steps:
[0159] In step 1651, the maximum value of the elements at each position in the first largest similar graph and the second largest similar graph is extracted as the value of the element at the corresponding position in the target largest similar graph to obtain the target largest similar graph.
[0160] Specifically, the operation of this step can be implemented by the following expression:
[0161]
[0162] Among them, is the first largest similar graph, is the second largest similar graph, and A s is the target largest similar graph.
[0163] It is easy to understand that although the embodiments of the present application obtain the target maximum similarity graph by taking the maximum value of elements, in other embodiments of the present application, the target maximum similarity graph can also be obtained by other means such as the average value of elements.
[0164] Figure 10 The Maximize module shown in is used to perform the operations of this step. That is the target maximum similarity graph.
[0165] In step 1652, a correction operation is performed on the target maximum similarity graph to obtain the affine graph of the target level.
[0166] In one embodiment of the present application, performing a correction operation on the target maximum similarity graph to obtain the affine graph of the target level includes: performing a mapping operation on the elements in the target maximum similarity graph to obtain a mapping graph; performing a size adjustment operation on the mapping graph to obtain the affine graph of the target level.
[0167] Figure 10 The SR module shown in Figure 10 is used to perform the operations of this step. The SR module performs a Softmax operation and a Reshape operation. The Softmax operation is the mapping operation, and the Reshape operation is the size adjustment operation.
[0168] For example, if the size of A s is hw×1, then the size of the mapping result obtained through the mapping operation is also hw×1. Through the Reshape operation, the mapping result with a size of hw×1 can be adjusted to a size of h×w.
[0169] By generating affine graphs for the feature pairs of 4 levels and respectively, 4-level affine graphs can be obtained
[0170] In one embodiment of the present application, the change detection model is trained according to a loss function. The loss function includes a main branch loss corresponding to the main branch and an auxiliary branch loss corresponding to the auxiliary branch. The main branch loss includes a cross-entropy loss and a Dice loss.
[0171] The change detection model is trained by minimizing the loss function. The loss function of the change detection model is shown in the joint learning part of Figure 10 . The loss function of the change detection model can specifically be: Figure 10 wherein,
[0172]
[0173] where is the main branch loss, is the auxiliary branch loss, and λ is a hyperparameter used to balance the main branch loss and the auxiliary branch loss.
[0174] The main branch loss is specifically:
[0175]
[0176] Among them, is the Dice loss, is the cross-entropy loss. A hyperparameter for balancing the Dice loss and the cross-entropy loss can also be added in this formula.
[0177] The Dice loss and the cross-entropy loss are respectively calculated through the following formulas:
[0178]
[0179]
[0180] Among them, is the Dice loss, is the cross-entropy loss, is the output of the change-aware head, and y i is the annotation information of the changed area.
[0181] For the affine graphs of the 4 levels in the auxiliary branch the auxiliary branch loss can be calculated through the following formula:
[0182]
[0183] Among them, is the auxiliary branch loss, and y i is the annotation information of the changed area, and resize represents the scaling operation.
[0184] The inventor conducted experiments on the method provided by this application using the public academic dataset VL-CMU-CD, using the F1 score as the performance metric, and the results are shown in Table 1:
[0185]
[0186]
[0187] Table 1
[0188] The inventor also conducted experiments on the method provided by this application based on the public academic dataset PCD. The dataset PCD includes GSV and TSUNAMI, using the F1 score as the performance metric, and the results are shown in Table 2:
[0189]
[0190] Table 2
[0191] The inventors also compared with existing methods based on the performance metric of mIoU, and the results are shown in Table 3 as follows:
[0192] Method Backbone network Change Static mIoU ChangeNet ResNet - 18 17.6 73.3 45.4 CSCDNet ResNet - 18 22.9 87.3 55.1 This method ResNet - 18 23.6 91.5 57.5
[0193] Table 3
[0194] As can be seen from Tables 1 - 3, the method provided by this application is significantly better than the existing methods.
[0195] Figure 18 shows a schematic diagram of the effect comparison between the method of the embodiment of this application and the related technology. Through Figure 18 it can be seen that the effect of the method of the embodiment of this application is better than that of the related technology.
[0196] In summary, according to the change detection method of the pictures provided by the embodiments of this application, it pays more attention to target changes than the existing methods, and is more robust to changes such as illumination, shadow, season, weather, and perspective; on a V100 GPU, the inference time of this method is 15 ms, with a small amount of calculation and high accuracy, which is convenient for practical application; this method can be used as one of the alternative options for the existing industrial AI quality inspection algorithms and the new electronic eye detection algorithms in the existing intelligent transportation.
[0197] The following introduces the device embodiments of this application, which can be used to execute the change detection method of the pictures in the above embodiments of this application. For the details not disclosed in the device embodiments of this application, please refer to the embodiments of the change detection method of the pictures above in this application.
[0198] Figure 19 shows a block diagram of a picture change detection device according to an embodiment of this application.
[0199] Refer to Figure 19As shown in the figure, a change detection device 1900 for pictures according to an embodiment of the present application includes: a feature extraction unit 1910, a spatio-temporal feature construction unit 1920, a fusion unit 1930, and a sampling unit 1940. Among them, the feature extraction unit 1910 is configured to perform feature extraction on the first picture and the second picture respectively to obtain multi-level feature pairs, where the feature pairs include a first feature extracted from the first picture and a second feature extracted from the second picture; the spatio-temporal feature construction unit 1920 is configured to perform convolution operations on the first feature and the second feature in the feature pair respectively for each level of the feature pair to obtain a first attention matrix combination corresponding to the first feature and a second attention matrix combination corresponding to the second feature, and construct a spatio-temporal feature combination corresponding to the feature pair according to the first attention matrix combination and the second attention matrix combination corresponding to the feature pairs at each level, where the spatio-temporal feature combination includes a first spatio-temporal feature corresponding to the first attention matrix combination and a second spatio-temporal feature corresponding to the second attention matrix combination; the fusion unit 1930 is configured to perform fusion operations on the spatio-temporal feature combinations corresponding to the feature pairs at each level respectively to obtain multi-level fusion features; the sampling unit 1940 is configured to perform downsampling operations on the multi-level fusion features to obtain a downsampling result, and perform upsampling operations on the downsampling result to obtain a change map.
[0200] In some embodiments of the present application, based on the foregoing solution, the fusion unit 1930 is configured to: for the spatio-temporal feature combination corresponding to the feature pair at each level, perform two convolution operations on the spatio-temporal feature combination respectively to obtain a third attention matrix combination and a fourth attention matrix combination, where both the third attention matrix combination and the fourth attention matrix combination include multiple attention matrices, and the same attention matrices in the third attention matrix combination and the fourth attention matrix combination are obtained by performing convolution operations on different spatio-temporal features in the spatio-temporal feature combination respectively; construct a first cross-interaction feature corresponding to the feature pair at each level according to the third attention matrix combination corresponding to each spatio-temporal feature combination, and construct a second cross-interaction feature corresponding to the feature pair at each level according to the fourth attention matrix combination corresponding to each spatio-temporal feature combination; determine the semantic change information between the first cross-interaction feature and the second cross-interaction feature corresponding to the feature pair at each level, and use the semantic change information corresponding to the feature pair at each level as the fusion feature at each level.
[0201] In some embodiments of the present application, based on the foregoing solution, the change map is generated by the main branch of the change detection model. The change detection model further includes an auxiliary branch for assisting in training the main branch. The change detection model is trained by generating multi-level affine maps in the auxiliary branch. The apparatus further includes an auxiliary training unit. The affine map of the target level in the multi-level affine maps is generated by the auxiliary training unit by performing the following process: respectively performing embedding operations on the first feature and the second feature in the feature pair of the target level to obtain a first feature embedding result and a second feature embedding result, and generating a channel correlation matrix according to the first feature embedding result and the second feature embedding result; masking some elements in the channel correlation matrix through a mask to obtain a masked matrix; extracting the maximum value of each row in the masked matrix to obtain a first maximum similarity map; extracting the maximum value of each column in the masked matrix to obtain a second maximum similarity map; and fusing the first maximum similarity map and the second maximum similarity map to obtain the affine map of the target level.
[0202] In some embodiments of the present application, based on the foregoing solution, the auxiliary training unit is configured to: extract the maximum value of the elements at each position in the first maximum similarity map and the second maximum similarity map as the value of the element at the corresponding position in the target maximum similarity map to obtain the target maximum similarity map; and perform a correction operation on the target maximum similarity map to obtain the affine map of the target level.
[0203] In some embodiments of the present application, based on the foregoing solution, the auxiliary training unit is configured to: perform a mapping operation on the elements in the target maximum similarity map to obtain a mapping map; and perform a size adjustment operation on the mapping map to obtain the affine map of the target level.
[0204] In some embodiments of the present application, based on the foregoing solution, the change detection model is trained according to a loss function. The loss function includes a main branch loss corresponding to the main branch and an auxiliary branch loss corresponding to the auxiliary branch. The main branch loss includes a cross-entropy loss and a Dice loss.
[0205] In some embodiments of the present application, based on the foregoing solution, the feature extraction unit 1910 is configured to: respectively perform feature extraction on the first picture and the second picture to obtain feature pairs of a predetermined number of levels; and remove the feature pair of a specified level from the feature pairs of the predetermined number of levels to obtain multi-level feature pairs, where the specified level is lower than the level corresponding to the multi-level feature pairs.
[0206] In some embodiments of the present application, based on the foregoing solution, the spatio-temporal feature construction unit 1920 is configured to: perform multiple groups of convolution operations on the first feature in the feature pair respectively to obtain multiple first attention matrix combinations corresponding to the first feature; perform multiple groups of convolution operations on the second feature in the feature pair respectively to obtain multiple second attention matrix combinations corresponding to the second feature.
[0207] In some embodiments of the present application, based on the foregoing solution, the feature extraction unit 1910 is configured to: perform feature extraction on the first picture and the second picture respectively to obtain multi-level original feature pairs; superimpose position encoding features on the first original feature and the second original feature in each level of the original feature pairs respectively to obtain multi-level feature pairs.
[0208] Figure 20 The structural schematic diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown.
[0209] It should be noted that Figure 20 The computer system 2000 of the electronic device shown is only an example, and should not impose any limitation on the functions and usage scopes of the embodiments of the present application.
[0210] As Figure 20 shown, the computer system 2000 includes a central processing unit (CPU) 2001, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 2002 or the program loaded from the storage section 2008 into the random access memory (RAM) 2003, such as executing the method described in the above embodiments. In the RAM 2003, various programs and data required for system operation are also stored. The CPU 2001, the ROM 2002, and the RAM 2003 are connected to each other through a bus 2004. The input / output (I / O) interface 2005 is also connected to the bus 2004.
[0211] The following components are connected to the I / O interface 2005: an input section 2006 including a keyboard, a mouse, etc.; an output section 2007 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 2008 including a hard disk, etc.; and a communication section 2009 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 2009 performs communication processing via a network such as the Internet. A drive 2010 is also connected to the I / O interface 2005 as needed. A removable medium 2011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 2010 as needed so that a computer program read from it can be installed into the storage section 2008 as needed.
[0212] Specifically, according to an embodiment of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 2009, and / or installed from the removable medium 2011. When the computer program is executed by a central processing unit (CPU) 2001, various functions defined in the system of the present application are executed.
[0213] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0214] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0215] The units involved in the embodiments described in this application can be implemented in software or in hardware, and the described units can also be provided in a processor. In some cases, the names of these units do not constitute a limitation on the units themselves.
[0216] As one aspect, the present application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or may exist separately without being assembled into the electronic device. When one or more of the above programs are executed by an electronic device, the electronic device is caused to implement the method described in the above embodiments.
[0217] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more of the above-described modules or units may be embodied in one module or unit. Conversely, the features and functions of one module or unit described above may be further divided and embodied by a plurality of modules or units.
[0218] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or by software in combination with necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.
[0219] It can be understood that in the specific embodiments of the present application, data related to games is involved. When the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards in relevant countries and regions.
[0220] After considering the specification and practicing the disclosed embodiments herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present application.
[0221] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A method for change detection of pictures, characterized in that, The method includes: Performing feature extraction on the first picture and the second picture respectively to obtain multi-level feature pairs, where the feature pairs include first features extracted from the first picture and second features extracted from the second picture; For each level of feature pairs, performing convolution operations on the first feature and the second feature in the feature pairs respectively to obtain a first attention matrix combination corresponding to the first feature and a second attention matrix combination corresponding to the second feature, and constructing a spatio-temporal feature combination corresponding to the feature pairs according to the first attention matrix combination and the second attention matrix combination corresponding to each level of feature pairs, where the spatio-temporal feature combination includes a first spatio-temporal feature corresponding to the first attention matrix combination and a second spatio-temporal feature corresponding to the second attention matrix combination; Performing fusion operations on the spatio-temporal feature combinations corresponding to each level of feature pairs respectively in the following manner to obtain multi-level fused features: For the spatio-temporal feature combination corresponding to each level of feature pairs, performing two convolution operations on the spatio-temporal feature combination respectively to obtain a third attention matrix combination and a fourth attention matrix combination, where both the third attention matrix combination and the fourth attention matrix combination include multiple attention matrices, and the same attention matrix in the third attention matrix combination and the fourth attention matrix combination is obtained by performing convolution operations on different spatio-temporal features in the spatio-temporal feature combination; constructing a first cross-interaction feature corresponding to each level of feature pairs according to the third attention matrix combination corresponding to each spatio-temporal feature combination, and constructing a second cross-interaction feature corresponding to each level of feature pairs according to the fourth attention matrix combination corresponding to each spatio-temporal feature combination; determining the semantic change information between the first cross-interaction feature and the second cross-interaction feature corresponding to each level of feature pairs, and using the semantic change information corresponding to each level of feature pairs as the fused features at each level; Performing downsampling operation on the multi-level fused features to obtain a downsampling result, and performing upsampling operation on the downsampling result to obtain a change map.
2. The method for change detection of pictures according to claim 1, wherein The change map is generated through the main branch of a change detection model, the change detection model further includes an auxiliary branch for assisting in training the main branch, and the change detection model is trained by generating multi-level affine maps in the auxiliary branch, and the affine map of the target level in the multi-level affine maps is generated through the following process: Performing embedding operations on the first feature and the second feature in the feature pair of the target level respectively to obtain a first feature embedding result and a second feature embedding result, and generating a channel correlation matrix according to the first feature embedding result and the second feature embedding result; Masking some elements in the channel correlation matrix through a mask to obtain a masked matrix; Extracting the maximum value of each row in the masked matrix to obtain a first maximum similarity map; Extracting the maximum value of each column in the masked matrix to obtain a second maximum similarity map; Fuse the first maximum similarity graph and the second maximum similarity graph to obtain an affine graph of the target level.
3. The method for change detection of the picture according to claim 2, wherein The step of fusing the first maximum similarity graph and the second maximum similarity graph to obtain an affine graph of the target level includes: Extract the maximum value of the elements at each position in the first maximum similarity graph and the second maximum similarity graph as the value of the element at the corresponding position in the target maximum similarity graph to obtain the target maximum similarity graph; Perform a correction operation on the target maximum similarity graph to obtain an affine graph of the target level.
4. The method for change detection of a picture according to claim 3, characterized in that, The step of performing a correction operation on the target maximum similarity graph to obtain an affine graph of the target level includes: Perform a mapping operation on the elements in the target maximum similarity graph to obtain a mapping graph; Perform a size adjustment operation on the mapping graph to obtain an affine graph of the target level.
5. The method for change detection of pictures according to claim 2, wherein The change detection model is trained according to a loss function, and the loss function includes a main branch loss corresponding to the main branch and an auxiliary branch loss corresponding to the auxiliary branch. The main branch loss includes a cross-entropy loss and a Dice loss.
6. The method for change detection of a picture according to claim 1, characterized in that The step of respectively performing feature extraction on the first picture and the second picture to obtain multi-level feature pairs includes: Respectively perform feature extraction on the first picture and the second picture to obtain feature pairs of a predetermined number of levels; Remove the feature pairs of the specified level from the feature pairs of the predetermined number of levels to obtain multi-level feature pairs, where the specified level is lower than the level corresponding to the multi-level feature pairs.
7. The method for change detection of pictures according to any one of claims 1-6, characterized in that The step of respectively performing convolution operations on the first feature and the second feature in the feature pairs to obtain a first attention matrix combination corresponding to the first feature and a second attention matrix combination corresponding to the second feature includes: Respectively perform multiple groups of convolution operations on the first feature in the feature pairs to obtain multiple first attention matrix combinations corresponding to the first feature; Respectively perform multiple groups of convolution operations on the second feature in the feature pairs to obtain multiple second attention matrix combinations corresponding to the second feature.
8. The method for change detection of pictures according to any one of claims 1-6, wherein the step of respectively performing feature extraction on the first picture and the second picture to obtain multi-level feature pairs includes: Respectively perform feature extraction on the first picture and the second picture to obtain multi-level original feature pairs; Superimpose position encoding features on the first original feature and the second original feature in the original feature pairs of each level respectively to obtain multi-level feature pairs.
9. An apparatus for change detection of pictures, characterized in that, The device includes: A feature extraction unit, configured to respectively perform feature extraction on the first picture and the second picture to obtain multi-level feature pairs, where the feature pairs include a first feature extracted from the first picture and a second feature extracted from the second picture; A spatio-temporal feature construction unit, configured to perform convolution operations on the first feature and the second feature in each pair of features at each level respectively, to obtain a first attention matrix combination corresponding to the first feature and a second attention matrix combination corresponding to the second feature, and construct a spatio-temporal feature combination corresponding to the pair of features according to the first attention matrix combination and the second attention matrix combination corresponding to each pair of features at each level, where the spatio-temporal feature combination includes a first spatio-temporal feature corresponding to the first attention matrix combination and a second spatio-temporal feature corresponding to the second attention matrix combination; A fusion unit, configured to perform fusion operations on the spatio-temporal feature combinations corresponding to each pair of features at each level respectively, to obtain multi-level fusion features; A sampling unit, configured to perform downsampling operations on the multi-level fusion features to obtain a downsampling result, and perform upsampling operations on the downsampling result to obtain a change map; Wherein, the fusion unit is configured to: for the spatio-temporal feature combination corresponding to each pair of features at each level, perform two convolution operations on the spatio-temporal feature combination respectively to obtain a third attention matrix combination and a fourth attention matrix combination, where both the third attention matrix combination and the fourth attention matrix combination include multiple attention matrices, and the same attention matrix in the third attention matrix combination and the fourth attention matrix combination is obtained by performing convolution operations on different spatio-temporal features in the spatio-temporal feature combination; construct a first cross-interaction feature corresponding to each pair of features at each level according to the third attention matrix combination corresponding to each spatio-temporal feature combination, and construct a second cross-interaction feature corresponding to each pair of features at each level according to the fourth attention matrix combination corresponding to each spatio-temporal feature combination; determine semantic change information between the first cross-interaction feature and the second cross-interaction feature corresponding to each pair of features at each level, and use the semantic change information corresponding to each pair of features at each level as the fusion features at each level.
10. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method for detecting changes in a picture according to any one of claims 1 to 8.
11. An electronic device, characterized in that, Including: One or more processors; A storage device, configured to store one or more programs, when the one or more programs are executed by the one or more processors, enable the one or more processors to implement the method for detecting changes in a picture according to any one of claims 1 to 8.
12. A computer program product, characterized in that, The computer program product includes computer instructions, the computer instructions are stored in a computer-readable storage medium, a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, enabling the computer device to execute the method for detecting changes in a picture according to any one of claims 1 to 8.
Citation Information
Patent Citations
Novel ultra-high-definition remote sensing image change detection method based on AFFPN
CN112818818A