Image Change Segmentation Method Based on Perspective Alignment
Through perspective alignment and feature fusion technology, the challenge of image change detection in computer vision is solved, accurate recognition of image changes and bounding box positioning are achieved, and the reliability and accuracy of detection are improved.
Patent Information
- Application Number
- CN202411754952.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-03
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-12-03
AI Technical Summary
In computer vision, it is difficult to accurately detect and position changes between two images of the same scene taken from different viewpoints at different moments, especially while ignoring the photometric differences and camera posture changes, while capturing physically different changes in the scene.
The image change segmentation method based on view angle alignment is adopted, and the images are aligned through the 2D scene alignment module and the 3D scene image registration module, combined with feature extraction, preliminary change detection network and feature fusion module, change information is generated, and the bounding box of the change area is identified through the positioning network detected by the border.
It is realized that the changes between images are accurately detected without luminance differences and camera posture changes, and the unnecessary changes are ignored, and only physically different changes are captured, which improves the accuracy and reliability of image change detection.
Smart Images

Figure CN119251507B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image recognition, and particularly relates to an image change segmentation method based on view alignment. Background Art
[0002] In computer vision, accurately detecting and localizing changes, whether in a 3D or 2D scene, remains a challenging task. Imagine being able to identify the changes between two images of the same scene captured from different viewpoints at different times, which is the challenge to be addressed in this work. With applications in areas such as robotics, facility monitoring, forensics, and augmented reality, our work has the potential to provide new possibilities for understanding and interacting with the dynamic world. We formulate the problem we will study as follows: Given a pair of RGB images of any scene (3D or 2D), captured with significant changes in camera position and at different times, we want to localize the changes between them. In particular, we want to capture everything physically different in the regions visible in the two images, while ignoring regions that appear or disappear from the view due to changes in camera pose or occlusion. This includes objects that may have been added or removed from the scene, as well as text or decorations that may have been added to the objects, while ignoring photometric differences such as lighting changes. Additionally, in order to operate only on RGB images without considering additional information such as camera parameters and camera pose, and be able to detect changes between image pairs regardless of the scene, it is a challenging problem to provide reliable and accurate detection of changes regardless of various scenes and environments. Summary of the Invention
[0003] In order to overcome the above deficiencies in technology, the present invention provides an image change segmentation method based on view alignment.
[0004] To achieve the above object, the present invention is realized through the following technical solutions:
[0005] In a first aspect, the present invention provides an image change segmentation method based on view alignment, including the following steps:
[0006] S1. Obtain the original image pair and preprocess the image pair to obtain two preprocessed images;
[0007] S2. The two preprocessed images pass through a 2D scene alignment module to obtain the aligned images corresponding to the two images;
[0008] S3. Construct a feature extraction module, input the two preprocessed images and the aligned images corresponding to the two images into the feature extraction module for feature extraction to obtain the feature information corresponding to the two images;
[0009] S4. Construct a preliminary change detection network, which includes a 3D scene image registration module and a difference module. The feature information corresponding to two aligned images is input into the preliminary change detection network to obtain the difference information corresponding to the two images;
[0010] S5. Construct a feature fusion module, input the difference information corresponding to the two images into the feature fusion module for feature fusion, and obtain the change information corresponding to the two images;
[0011] S6. Construct a positioning network for border detection, input the change information corresponding to the two images into the positioning network for border detection, and obtain the bounding boxes of the change regions of the two images respectively.
[0012] Further, step S1 is specifically as follows:
[0013] Perform geometric transformation operations on the two original images to obtain the preprocessed left image with a size of and the right image , , , The data type of the matrix elements is real, is the height of the image, is the width of the image, 3 is the number of channels of the image, indicates that the image is composed of a real number matrix with a shape and size of , Similarly.
[0014] Further, the 2D scene alignment module in step S2 is composed of an image feature point matching unit and a homography registration unit, and specifically:
[0015] S21. The preprocessed left image and the right image pass through the image feature point matching unit to extract feature points using the feature matching extractor to obtain the set of matching feature points of the left image and the set of matching feature points of the right image . At the same time, the preprocessed left image and the right image use the monocular depth estimator to obtain the corresponding depth maps respectively;
[0016] S22. The sets of feature points and pass through the homography registration unit to calculate the homography transformation matrix and obtain the transformation matrix of , and transformation matrix , the transformation matrix and are applied to the corresponding images to achieve image alignment, obtaining the left image aligned image in the 2D scenario based on the right image R and the right image aligned image in the 2D scenario based on the left image L , .
[0017] Furthermore, the feature extraction module in step S3 includes a reference feature extraction model and a U-Net encoder, specifically:
[0018] S31. Input the preprocessed left image , the preprocessed right image , the left image aligned image and the right image aligned image into the reference feature extraction model for feature extraction, obtaining the corresponding preliminary feature information , and merge the preliminary feature information corresponding to the preprocessed left image with the preliminary feature information corresponding to the aligned image of the right image in a way of merging the number of channels to obtain the merged feature information of the left image L , and merge the preliminary feature information corresponding to the preprocessed right image with the preliminary feature information corresponding to the aligned image of the left image in a way of merging the number of channels to obtain the merged feature information of the right image R ; ;
[0019] ; S32. Input and into their respective U-Net compilers respectively, and output and corresponding five different-scale intermediate feature maps and , , when s = 1, the scale size of the feature map is 64×64; when s = 2, the scale size of the feature map is 32×32; when s = 3, the scale size of the feature map is 16×16; when s = 4, the scale size of the feature map is 8×8; when s = 5, the scale size of the feature map is 4×4.
[0020] Furthermore, the preliminary change detection network described in step S4 includes a 3D scene image registration module and a difference module. The difference module is composed of a subtraction operation, specifically:
[0021] S41. Divide the corresponding intermediate feature maps and and and into two equal parts in the channel dimension respectively to obtain the feature maps corresponding to the left image , the feature maps corresponding to the image , the feature maps corresponding to the right image , and the feature maps corresponding to the image , where ; ; ; . ;
[0022] S42. In the 3D scene image registration module, according to the feature point set corresponding to the preprocessed left image , the feature point set corresponding to the right image , the depth map corresponding to the preprocessed left image , and the depth map corresponding to the right image , first back-project each point in the feature point set to obtain the projection vector , and then obtain the sparse three-dimensional point clouds corresponding to the homogeneous coordinate feature point sets respectively. According to the sparse three-dimensional point clouds, estimate the 3D linear transformation matrix for aligning to using the calculation formula . At the same time, estimate the 3D linear transformation matrix for aligning to using the calculation formula . The formula is as follows:
[0023] .
[0024] ,
[0025] ,
[0026] wherein, represents 's depth value, represents 's matrix generalized inverse, represents 's matrix generalized inverse;
[0027] S43. Use the 3D linear transformation matrix obtained in step S42 to transform the image coordinates of the feature map corresponding to the left image into image coordinates in the three-dimensional coordinate system , and use the point cloud tool to encapsulate the transformed image coordinates and the feature map into a point cloud object, and finally render the point cloud object into a feature map based on the aligned image features corresponding to the 3D scene , ; similarly, obtain the feature map based on the aligned image features corresponding to the 3D scene , ;
[0028] S44. Input the feature map corresponding to the left image , the feature map corresponding to the image and the feature map based on the aligned image features corresponding to the 3D scene into the difference module, and through subtraction operation, obtain the difference information of the left image under the 2D scene and the difference information of the left image under the 3D scene. The formula is as follows:
[0029]
[0030] ,
[0031] Similarly, obtain the difference information of the right image R under the 2D scene and the difference information of the right image R under the 3D scene.
[0032] Further, the feature fusion module described in step S5 is specifically as follows:
[0033] Concatenate the difference information and to obtain a concatenated feature, and input the concatenated feature into a linear layer to obtain an intermediate feature , then perform a convolution operation to obtain a weight . Multiply the weight element-wise with the feature map corresponding to the left image to finally obtain the change information of the left image . The formula is as follows:
[0034] ,
[0035] ,
[0036] ,
[0037] wherein, represents channel merging of two features, represents a non-linear activation function, represents the first convolutional layer, represents the second convolutional layer, represents the Sigmoid activation function, represents an element-wise multiplication operation; similarly, obtain the change information of the right image .
[0038] Further, step S6 is specifically as follows:
[0039] The change information and are upsampled and decoded through a U-Net decoder to obtain the feature map corresponding to the left image and the feature map corresponding to the right image . Then, input the feature maps and into the CenterNet network head for bounding box detection and localization to obtain the bounding boxes of the changed regions of the two images respectively.
[0040] In a second aspect, an image change segmentation device based on view alignment includes:
[0041] A data acquisition unit: configured to acquire an original image pair and preprocess the image pair to obtain two preprocessed images;
[0042] 2D Scene Alignment Unit: It is used to obtain the aligned images corresponding to the two pre - processed images through the 2D scene alignment module;
[0043] Feature Extraction Unit: It is used to input the two pre - processed images and the aligned images corresponding to the two images into the feature extraction module for feature extraction, and obtain the feature information corresponding to the two images;
[0044] First Construction Unit: It is used to construct a preliminary change detection network. The preliminary change detection network includes a 3D scene image registration module and a difference module. The feature information corresponding to the two aligned images is input into the preliminary change detection network to obtain the difference information corresponding to the two images;
[0045] Second Construction Unit: It is used to construct a feature fusion module, and input the difference information corresponding to the two images into the feature fusion module for feature fusion to obtain the change information corresponding to the two images;
[0046] Third Construction Unit: It is used to construct a positioning network for border detection, and input the change information corresponding to the two images into the positioning network for border detection to obtain the bounding boxes of the changed regions of the two images respectively.
[0047] In a third aspect, an electronic device includes: a processor, a memory, and a bus. The memory stores machine - readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine - readable instructions are executed by the processor, the steps of the image change segmentation method based on perspective alignment described in the first aspect are executed.
[0048] In a fourth aspect, a computer - readable storage medium is characterized in that a computer program is stored on the computer - readable storage medium. When the computer program is run by a processor, the steps of the image change segmentation method based on perspective alignment described in the first aspect are executed.
[0049] The advantages of the present invention are as follows:
[0050] The present invention adopts a network composed of a combination of an image registration difference structure and a fusion structure. By two image registration methods, RGB image pairs are aligned respectively to obtain two kinds of difference information, one is the difference information tending to the 2D scene, and the other is the difference information tending to the 3D scene. Then, the two kinds of difference information are effectively fused through the fusion structure, making up for the problem of insufficient change difference information. At the same time, the network adopts a Siamese neural network architecture, so that two images can be operated simultaneously, and the recognition of the changed region can be better completed. Description of the Drawings
[0051] The accompanying drawings are used to provide a further understanding of the present invention and form a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention.
[0052] Figure 1 is the method flowchart of the present invention;
[0053] Figure 2 is the 3D-based image registration operation diagram of the present invention;
[0054] Figure 3 is the operation diagram of the difference module and the feature fusion module of the present invention;
[0055] Figure 4 is the network structure diagram of the image change segmentation method based on view alignment of the present invention;
[0056] Figure 5 is the segmentation result diagram of the method of the present invention. Detailed implementation manners
[0057] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0058] Embodiment 1
[0059] In this embodiment, as Figure 1 shown, the present invention provides an image change segmentation method based on view alignment. The specific steps include:
[0060] S1. Obtain the original image pair and preprocess the image pair to obtain two preprocessed images;
[0061] Specifically, perform geometric transformation operations on the two original images to obtain the preprocessed left image with a size of and the right image . . The data type of the matrix elements is real numbers. is the height of the image. is the width of the image, and 3 is the number of channels of the image. indicates that the image is composed of a real number matrix with a shape and size of . Similarly.
[0062] S2. The two preprocessed images are input into the 2D scene alignment module to obtain the aligned images corresponding to the two images;
[0063] Specifically, the 2D scene alignment module consists of an image feature point matching unit and a homography registration unit;
[0064] Specifically, S21. The preprocessed left image and the right image pass through the image feature point matching unit, and the feature matching extractor is used to extract feature points, obtaining the set of matched feature points of the left image and the set of matched feature points of the right image Meanwhile, the preprocessed left image and the right image .
[0065] Specifically, S22. The set of and pass through the homography registration unit to calculate the homography transformation matrix, obtaining the transformation matrix , and the transformation matrix . The transformation matrix and are applied to the corresponding images to achieve image alignment, obtaining the aligned image of the left image under the 2D scene based on the right image R and the aligned image of the right image under the 2D scene based on the left image L
[0066] S3. Construct a feature extraction module, input the two preprocessed images and the aligned images corresponding to the two images into the feature extraction module for feature extraction, obtaining the feature information corresponding to the two images;
[0067] Specifically, the feature extraction module includes a reference feature extraction model and a U-Net encoder;
[0068] Specifically, S31. The preprocessed left image , the preprocessed right image , the aligned image of the left image and the aligned image of the right image Input it into the benchmark feature extraction model for feature extraction to obtain the corresponding preliminary feature information , the preprocessed left image The corresponding preliminary feature information and the right image The aligned image of The corresponding preliminary feature information are merged in the way of merging the number of channels to obtain the merged feature information of the left image L , the preprocessed right image The corresponding preliminary feature information and the left image The aligned image of The corresponding preliminary feature information are merged in the way of merging the number of channels to obtain the merged feature information of the right image R ;
[0069] Specifically, S32. Input and into their respective U-Net compilers respectively, and output and obtain and The corresponding five intermediate feature maps of different scales and , , when s = 1, the scale size of the feature map is 64×64, when s = 2, the scale size of the feature map is 32×32, when s = 3, the scale size of the feature map is 16×16, when s = 4, the scale size of the feature map is 8×8, when s = 5, the scale size of the feature map is 4×4.
[0070] S4. Construct a preliminary change detection network, the preliminary change detection network includes a 3D scene image registration module and a difference module, and the feature information corresponding to the two aligned images is input into the preliminary change detection network to obtain the difference information corresponding to the two images;
[0071] Specifically, S41. Perform a channel half-division operation on the corresponding intermediate feature map and corresponding intermediate feature map respectively to obtain the feature map corresponding to the left image 、the feature map corresponding to the image 、the feature map corresponding to the right image 、the feature map corresponding to the image and the feature map corresponding to the image 、the feature map corresponding to the image and the feature map corresponding to the image 、the feature map corresponding to the image , where ;
[0072] Specifically, in step S42, in the 3D scene image registration module, according to the preprocessed left image obtained in step S21 and the corresponding feature point set , the right image and its corresponding feature point set , the preprocessed left image and its corresponding depth map , as well as the right image and its corresponding depth map , first, for each point in the feature point set perform back-projection to obtain the projection vector , and then obtain the feature point set in homogeneous coordinates and their corresponding sparse 3D point clouds . According to the sparse 3D point clouds, through the calculation formula , estimate the 3D linear transformation matrix aligned with . At the same time, according to the sparse 3D point clouds, through the calculation formula , estimate the 3D linear transformation matrix aligned with . The formula is as follows:
[0073] ,
[0074] ,
[0075] ,
[0076] wherein, represents the depth value of , represents the matrix generalized inverse of , represents the matrix generalized inverse of
[0077] Specifically, in step S43, use the 3D linear transformation matrix obtained in step S42 to transform the image coordinates of the feature map corresponding to the left image into the image coordinates in the 3D coordinate system . Use the point cloud tool to encapsulate the transformed image coordinates and the feature map into a point cloud object, and finally render the point cloud object into a feature map based on the aligned image features corresponding to the 3D scene , ; Similarly, the feature map is obtained Based on the aligned image features corresponding to the 3D scene , .
[0078] Specifically, in S44, the left image The corresponding feature map , the image The corresponding feature map And the feature map Based on the aligned image features corresponding to the 3D scene Are input into the difference module. Through subtraction operation, the difference information of the left image Based on the 2D scene And the left image Based on the difference information of the 3D scene are obtained. The formula is as follows:
[0079]
[0080] ,
[0081] Similarly, the difference information of the right image R based on the 2D scene And the difference information of the right image R based on the 3D scene are obtained.
[0082] S5. Construct a feature fusion module, input the difference information corresponding to the two images into the feature fusion module for feature fusion, and obtain the change information corresponding to the two images;
[0083] Specifically, as Figure 3 shown, in the feature fusion module, the difference information and are concatenated to obtain a concatenated feature. The concatenated feature is input into a linear layer to obtain an intermediate feature , and then a convolution operation is performed to obtain a weight . The weight is multiplied element-wise with the feature map corresponding to the left image , and finally the change information of the left image is obtained. The formula is as follows:
[0084] ,
[0085] ,
[0086] ,
[0087] where, Indicates the channel merging of two features, Indicates a non-linear activation function, Indicates the first convolutional layer, Indicates the second convolutional layer, Indicates the Sigmoid activation function, Indicates an element-wise multiplication operation; similarly, the change information of the right image is obtained of .
[0088] S6. Construct a localization network for bounding box detection, input the change information corresponding to the two images into the localization network for bounding box detection, and obtain the bounding boxes of the changed regions of the two images respectively.
[0089] Specifically, the change information and are upsampled and decoded by a U-Net decoder to obtain the feature maps corresponding to the left image and the feature maps corresponding to the right image and respectively. Then, the feature maps and and are input into the CenterNet network head for bounding box detection and localization to obtain the bounding boxes of the changed regions of the two images respectively.
[0090] Embodiment 2
[0091] In this embodiment, in order to verify the effectiveness of the method of the present invention, evaluations are carried out on the KC-3D dataset, RC-3D dataset, COCO-Inpainted dataset, Synthtext-Change dataset, VIRAT-STD dataset, and Kubric-Change dataset. The COCO-Inpainted dataset is a test set based on changes organized by us from the COCO test subset. We divide this test set into three categories according to the size of the changed objects, namely small, medium, and large. All represents the integration of the test sets of the three categories. We organized 1655 pairs of image pairs for small objects, 1747 pairs of image pairs for medium objects, and 1006 pairs of image pairs for large objects, totaling 4408 pairs of image pairs for the COCO-Inpainted test set.
[0092] The KC-3D dataset uses the Kubric dataset generator to manage 86,407 pairs of 3D scene images with controllable variations (4,548 pairs of images are used for validation). The scenes consist of randomly selected 3D objects that are generated at random positions on randomly textured planes. We iteratively remove these objects and capture "before" and "after" image pairs. We capture images of various camera poses within the cylindrical space around the object. An additional 4,548 pairs of 3D scene images are extracted from Kubric for testing.
[0093] The RC-3D dataset quantifies the ability of the model to generalize to real-world images. A small test set consisting of 100 pairs of images was manually collected and labeled, capturing various common objects found in daily places such as offices, kitchens, and lounges. These images were taken with a handheld Apple iPad Pro (4th generation) and an Apple iPhone 14 Pro, which have built-in lidar that can provide aligned RGB-D images.
[0094] The Synthtext-Change dataset adds random text to "background" images through synthetic techniques and generates 5,000 pairs of images in a way consistent with their geometry. To detect changes in outdoor scenes, we randomly selected 1,000 pairs of images from the STD dataset. Since STD does not provide the basic GroundTruth for changes, automated tools were used to obtain the basic GroundTruth. Since the camera is static, there is an identical geometric transformation between the images, but the photometric conditions may change due to the time of day, weather conditions, etc. The Kubric-Change dataset consists of 1,605 realistic image pairs of changes. The scenes consist of a randomly selected set of 3D objects located on a randomly textured ground plane. For a given scene, objects are iteratively removed from it and "before" and "after" image pairs are captured.
[0095] For quantitative evaluation, we calculate the average precision AP based on the predicted bounding boxes and the ground truth bounding boxes as the evaluation metric according to previous related methods.
[0096] Figure 4 This is the network structure diagram of the method of the present invention. The performance comparison between the classical image change detection algorithm and the method of the present invention is shown in Table 1. The experiment was set with 12 epochs, using the optimization method Adam. The default learning rate is 0.00001, and the weight decay is 0.0005. To enhance the model's fitting ability to the data, we adopted random affine transformation, contrast enhancement, lighting enhancement, and saturation enhancement.
[0097] Table 1 Comparison of the performance of the currently optimal change detection model and the method of the present invention on different datasets:
[0098]
[0099] The CYWS-3D model is the optimal change detection model in the current research field. It can be found from this table that the performance of our model is excellent compared with the CYWS-3D model in the COCO-Inpainted, KC-3D dataset, and Synthtext-Change datasets, and the performance in other datasets is stable.
[0100] Example 3
[0101] In this embodiment, as Figure 5 shown, the initial images L and R are taken from different perspectives. Due to the different perspectives, there will be occlusion between objects in the images. The purpose of this model is to capture under significant changes in the camera position and at different times. We hope to locate the changes between them. In particular, we hope to capture everything physically different in the regions visible in the two images, while ignoring the regions that appear or disappear from the view due to changes in camera pose or occlusion. This includes objects that may have been added or removed from the scene, as well as text or decorations that may have been added to the objects, while ignoring photometric differences such as lighting changes. In addition, in order not to consider additional information such as camera parameters and camera pose, and only operate on RGB images, the changes between image pairs in any scene can be detected. Figure 5 shown, the object that changes between the two images is the shoe. The shoe exists in Figure L but does not exist in Figure R. Through the model, the different change regions of this image pair are finally detected, and the detection boxes are used for annotation to obtain the segmentation regions.
[0102] Example 4
[0103] An image change segmentation device based on view alignment, comprising:
[0104] A data acquisition unit: used to acquire the original image pair and preprocess the image pair to obtain two preprocessed images;
[0105] A 2D scene alignment unit: used to obtain the aligned images corresponding to the two images by passing the two preprocessed images through a 2D scene alignment module;
[0106] A feature extraction unit: used to input the two preprocessed images and the aligned images corresponding to the two images into a feature extraction module for feature extraction to obtain the feature information corresponding to the two images;
[0107] The first construction unit: used to construct a preliminary change detection network, the preliminary change detection network includes a 3D scene image registration module and a difference module, and the feature information corresponding to two aligned images is input into the preliminary change detection network to obtain the difference information corresponding to the two images;
[0108] The second construction unit: used to construct a feature fusion module, and input the difference information corresponding to the two images into the feature fusion module for feature fusion to obtain the change information corresponding to the two images;
[0109] The third construction unit: used to construct a positioning network for border detection, and input the change information corresponding to the two images into the positioning network for border detection to obtain the bounding boxes of the change regions of the two images respectively.
[0110] Embodiment 5
[0111] An electronic device includes: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the above-mentioned image change segmentation method based on view alignment are executed.
[0112] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely set relative to the processor, and these remote memories may be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0113] Embodiment 6
[0114] A computer-readable storage medium, characterized in that a computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, the steps of the above-mentioned image change segmentation method based on view alignment are executed.
[0115] The storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.
[0116] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for image change segmentation based on perspective alignment, characterized in that: The following steps are involved: S1. Get the original image pair and preprocess the image pair to obtain two preprocessed images: left image , right image ; S2. The two preprocessed images are passed through a 2D scene alignment module to obtain aligned images corresponding to the two images: left image Aligned images in 2D scenes based on the right image R , right image Aligned images in 2D scenes based on the left image L ; S3. Construct a feature extraction module, input the two preprocessed images and the aligned images corresponding to the two images into the feature extraction module for feature extraction, and obtain the merged feature information of the left image L And the combined feature information of the right image R , and then get The corresponding intermediate feature map and The corresponding intermediate feature map ; S4. constructing a preliminary change detection network, wherein the preliminary change detection network includes a 3D scene image registration module and a difference module, and inputting feature information corresponding to the two aligned images into the preliminary change detection network to obtain difference information corresponding to the two images; The preliminary change detection network includes a 3D scene image registration module and a difference module, and the difference module is composed of a subtraction operation, specifically: S41. The corresponding intermediate feature map and The corresponding intermediate feature map Perform the channel half-division operation respectively to obtain the left image The corresponding feature map ,image The corresponding feature map , right image The corresponding feature map and images The corresponding feature map ,in, ; S42. In the 3D scene image registration module, the preprocessed left image is obtained according to step S21. The corresponding feature point set , right image The corresponding feature point set And the preprocessed left image The corresponding depth map And the right image The corresponding depth map , first for each point in the feature point set Perform back projection to get the projection vector , and then get the feature point set in homogeneous coordinates The corresponding sparse 3D point clouds , according to the sparse 3D point cloud through the calculation formula , it is estimated Towards Aligned 3D linear transformation matrix , and according to the sparse three-dimensional point cloud, the calculation formula , it is estimated Towards Aligned 3D linear transformation matrix , the formula is as follows: , , , in, express The depth value of express The matrix generalized inverse of express The matrix generalized inverse of ; S43. The left image The corresponding feature map The image coordinates are used to obtain the 3D linear transformation matrix in step S42 Transform it into Image coordinates in three-dimensional coordinate system , use the point cloud tool to convert the converted image coordinates and feature map Encapsulate into a point cloud object, and finally render the point cloud object into a feature map through a differentiable renderer Aligning image features based on 3D scene correspondence , ; Similarly, we get the feature map Aligning image features based on 3D scene correspondence , ; S44. The left image The corresponding feature map ,image The corresponding feature map And feature map Aligning image features based on 3D scene correspondence Input into the difference module, and through the subtraction operation, the left image is obtained Based on the difference information in 2D scene and left image Based on the difference information in 3D scene , the formula is as follows: , Similarly, the right image R is obtained based on the difference information in the 2D scene And the right image R based on the difference information in the 3D scene ; S5. Construct a feature fusion module, input the difference information corresponding to the two images into the feature fusion module for feature fusion, and obtain the change information corresponding to the two images; The feature fusion module is specifically: The difference information and Perform splicing to obtain splicing features, and input the splicing features into the linear layer to obtain intermediate features , and then perform convolution operation to get the weight , the weight With the left image The corresponding feature map Multiply element by element and finally get the left image Change information , the formula is as follows: , , , in, Indicates that two features are merged into one channel. represents a nonlinear activation function, represents the first convolutional layer, represents the second convolutional layer, represents the Sigmoid activation function, Represents an element-by-element multiplication operation; similarly, we get the right image Change information ; S6. Construct a positioning network for border detection, input the change information corresponding to the two images into the positioning network for border detection, and obtain the bounding boxes of the changed areas of the two images respectively.
2. The image change segmentation method based on view alignment according to claim 1, characterized in that: Step S1 is specifically as follows: Perform geometric transformation on the two original images and obtain the preprocessed size as Left image and right image , , Indicates that the data type of the matrix elements is real number, is the height of the image, is the width of the image, 3 is the number of channels of the image, Representing images It is a shape with a size of The real matrix of Same reason.
3. The image change segmentation method based on view alignment according to claim 2, characterized in that: The 2D scene alignment module in step S2 is composed of an image feature point matching unit and a homography registration unit, specifically: S21. Preprocessed left image and right image The image feature point matching unit uses the feature matching extractor Extract feature points and get the left image Matching feature point set and right image Matching feature point set , while the preprocessed left image and right image Using a monocular depth estimator Get the corresponding depth map respectively ; S22. The feature point set and After the homography registration unit calculates the homography transformation matrix, we get The transformation matrix ,and The transformation matrix , the transformation matrix and Apply to the corresponding image to achieve image alignment and get the left image Aligned images in 2D scenes based on the right image R and right image Aligned images in 2D scenes based on the left image L , .
4. The image change segmentation method based on view alignment according to claim 3, characterized in that: The feature extraction module in step S3 includes a baseline feature extraction model and a U-Net encoder, specifically: S31. The preprocessed left image , right image after preprocessing , left image Aligned images of and right image Aligned images of Input into the benchmark feature extraction model for feature extraction, and obtain the corresponding preliminary feature information , the preprocessed left image Corresponding preliminary feature information With the right image Aligned images of Corresponding preliminary feature information Merge by merging the number of channels to obtain the merged feature information of the left image L , the preprocessed right image Corresponding preliminary feature information With the left image Aligned images of Corresponding preliminary feature information Merge by merging the number of channels to obtain the merged feature information of the right image R ; S32. and Input them into their respective U-Net compilers and output them respectively and Corresponding five intermediate feature maps of different scales and , When s=1, the scale of the feature map is 64×64, when s=2, the scale of the feature map is 32×32, when s=3, the scale of the feature map is 16×16, when s=4, the scale of the feature map is 8×8, and when s=5, the scale of the feature map is 4×4.
5. The image change segmentation method based on view alignment according to claim 4, characterized in that: Step S6 is specifically as follows: The change information and After upsampling and decoding by U-Net decoder, the left image is obtained The corresponding feature map and right image The corresponding feature map , and then the feature map and The data is input into the CenterNet network head for border detection and positioning to obtain the bounding boxes of the changed areas of the two images.
6. A device for image change segmentation based on view alignment, executing the image change segmentation method based on view alignment as claimed in claim 1, characterized in that: include: Data acquisition unit: used to acquire the original image pair and preprocess the image pair to obtain two preprocessed images; 2D scene alignment unit: used for passing the two pre-processed images through the 2D scene alignment module to obtain an aligned image corresponding to the two images; Feature extraction unit: used to input the two preprocessed images and the aligned images corresponding to the two images into the feature extraction module for feature extraction to obtain feature information corresponding to the two images; A first construction unit is used to construct a preliminary change detection network, wherein the preliminary change detection network includes a 3D scene image registration module and a difference module, and feature information corresponding to two aligned images is input into the preliminary change detection network to obtain difference information corresponding to the two images; The second construction unit is used to construct a feature fusion module, and input the difference information corresponding to the two images into the feature fusion module for feature fusion to obtain the change information corresponding to the two images; The third construction unit is used to construct a positioning network for border detection, and input the change information corresponding to the two images into the positioning network for border detection to obtain the boundary boxes of the change areas of the two images respectively.
7. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate via the bus, and when the machine-readable instructions are executed by the processor, the image change segmentation method based on perspective alignment as described in any one of claims 1 to 5 is performed.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the image change segmentation method based on viewpoint alignment as claimed in any one of claims 1 to 5 is executed.
Citation Information
Patent Citations
Scene change detection method and device based on deep learning, medium and equipment
CN118097566A