Open panorama segmentation method based on deformation enhancement and distortion comparison

Through random Gaussian deformation enhancement and distortion perception contrast learning technology, the problems of insufficient distortion perception and feature misalignment in open panoramic image segmentation are solved, and the segmentation performance of the model on panoramic images is significantly improved.

CN120147328APending Publication Date: 2025-06-13HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510211032.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing open panoramic image segmentation method lacks supervision perception of panoramic images during the training process, resulting in insufficient distortion perception and feature misalignment, which in turn affects the generalization performance of the model on panoramic images.

Method used

Random Gaussian deformation enhancement (RGDA) and distortion-aware contrast learning (DaCL) technology are used to enhance pinhole images through random parameterized Gaussian deformation, distorted images are generated, and contrast learning is performed at the patch level and pixel level to extract features related to distortion invariant and class.

Benefits of technology

It effectively alleviates the problems of insufficient distortion perception and feature misalignment, improves the segmentation performance of the model on panoramic images, which is specifically manifested as the improvement of segmentation mIoU performance on indoor and outdoor reference datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147328A_ABST
    Figure CN120147328A_ABST
Patent Text Reader

Abstract

The invention discloses an open panorama segmentation method based on deformation enhancement and distortion comparison, and the method comprises the steps: enabling RGDA to enhance a training set of a pinhole image with a label through the random parameterized Gaussian deformation, and providing a distorted image for a segmentation network. Secondly, the DaCL performs patch-level and pixel-level comparative learning between the pinhole and the image after deformation enhancement so as to extract distortion-invariant and class-related features, and the feature space difference between the pinhole image and the panorama is reduced; wide experiments carried out on four disclosed indoor and outdoor references prove the effectiveness of the open panorama segmentation method based on deformation enhancement and distortion comparison provided by the invention. According to the method, the segmentation mIoU performance of Stanford2D3D and Matterport3D in an indoor environment and the segmentation mIoU performance of DensePASS and WildPASS in an outdoor environment are improved to 3.2%, 1.4%, 2.4% and 1.7% respectively, and the segmentation mIoU performance of the DensePASS and the WildPASS in the indoor environment is improved to 3.2%, 1.4%, 2.4% and 1.7% respectively.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a data augmentation technique, particularly a random Gaussian deformation augmentation technique; and a contrast learning technique, particularly a deformation contrast learning technique at the patch level and pixel level. Background Art

[0002] Recently, Open Panoramic Segmentation (OPS) has received extensive attention because it can be trained only on pinhole images within a closed vocabulary and can be effectively generalized to panoramic images in an open vocabulary setting.

[0003] Due to the ability to provide a 360° comprehensive view, panoramic semantic segmentation has attracted increasing attention in environmental perception, drone remote sensing, and autonomous driving scenarios. Although these methods have achieved impressive results, they require a large number of annotated panoramic images during the training process, which is time-consuming. Many Domain Adaptation (DA) methods have solved this problem without the need for panoramic image annotation. They transfer the knowledge learned by the segmentation network from pinhole images to panoramic images.

[0004] However, these methods require predefined categories, and when new categories need to be recognized, the segmentation network needs to be retrained. Recently, Open-vocabulary Semantic Segmentation (OVS) has emerged in the field, which can recognize any input category during the test process without predefined categories. These methods apply the powerful zero-shot image-level classification capabilities of Vision Language Models (VLMs), such as CLIP or ALIGN, to pixel-level classification to achieve OVS.

[0005] These open-vocabulary semantic segmentation methods are designed specifically for pinhole images and may produce suboptimal performance when directly applied to panoramic images because they have never been exposed to panoramic images before. The introduction of Open Panoramic Segmentation (OPS) adapts the OVS task to panoramic images to alleviate this problem. OOOPS is the pioneering work of OPS, which proposed Random Equiangular Projection (RERP) and Deformable Adaptation Network (DAN) to enhance the perception of panoramic scenes. RERP alleviates the problem of insufficient distortion perception. However, it can only produce a limited range of deformation styles, which are neither flexible nor effective enough. DAN mainly focuses on expanding the receptive field of the panoramic image segmentation network but ignores the feature misalignment caused by distortion, resulting in poor generalization of panoramic images.

[0006] Despite significant progress in open-vocabulary semantic segmentation and open panoramic segmentation methods, especially in terms of open vocabulary, they have not fully addressed two key issues in open panoramic segmentation: (1) insufficient distortion perception. The segmentation network only receives annotated pinhole images during training and lacks supervised perception of panoramic images. (2) Feature misalignment caused by distortion. The open panoramic segmentation model trained on pinhole images performs well on pinhole images (such as the Cityscapes dataset), showing high confidence in the "road" category. However, its performance on panoramic images (such as the DensePASS dataset) is not ideal, showing low confidence in large areas of the "road" category. This phenomenon indicates that when the model is only trained on pinhole images and then tested on panoramic images, the feature space changes.

[0007] Previous open panoramic segmentation methods have mainly focused on open vocabulary and paid less attention to open field of view (FoV) and open domain. This has led to two key issues: (1) insufficient distortion perception, as the segmentation model is only trained on pinhole images and lacks perception of panoramic distortion; and (2) distortion-induced feature misalignment, where the distortion difference between pinhole images and panoramic images results in domain shift in the feature space. Summary of the Invention

[0008] The purpose of the present invention is to solve the above problems existing in the background technology and provide an open panoramic segmentation method based on deformation enhancement and distortion contrast, which consists of random Gaussian deformation augmentation (RGDA) and distortion-aware contrastive learning (DaCL).

[0009] To achieve the above purpose, the technical solutions adopted by the present invention are as follows:

[0010] An open panoramic segmentation method based on deformation enhancement and distortion contrast, the method is as follows:

[0011] The open panoramic segmentation (OPS) task aims to train an open vocabulary segmentation model only on pinhole samples (x pin , y pin ) ∈ S and test it on panoramic samples (x pan , y pan ) ∈ T, where the label spaces do not overlap, that is, new categories may be encountered during the test;

[0012] Step 1: Random Gaussian Deformation Augmentation (RGDA)

[0013] Random Gaussian deformation augmentation provides distorted images for the segmentation network S, alleviating the lack of distortion perception in the following ways: (1) constructing 2D Gaussian kernels with random parameters at multiple scales; (2) applying first-order horizontal and vertical differences to these kernels in the 2D plane; (3) combining the results to form a deformation field, which is then applied to the pinhole image x pin to obtain the deformed image x def , as Figure 2 shown;

[0014] For an input image x pin ∈R 1×H×W with label y pin ∈R 3×H×W apply the deformation transformation φ ∈ R 2×H×W , to obtain a distorted image x def ∈R 1×H×W with the corresponding label y def ∈R 3×H×W , using the grid sampling function as follows:

[0015]

[0016] Use a 2D independent Gaussian function to construct the deformation transformation matrix φ:

[0017]

[0018] where x and y are the horizontal and vertical coordinates of a point (x, y) in the 2D plane, and A, μ, and σ are the amplitude, mean, and standard deviation of the Gaussian function respectively; μ x is the mean in the horizontal direction in the plane; μ y is the mean in the vertical direction in the plane; σ x is the standard deviation in the horizontal direction in the plane; σ y is the standard deviation in the vertical direction in the plane;

[0019] Apply horizontal and vertical first-order differences to G to construct φ:

[0020] φ = [diff x (G), diff y (G)] (3)

[0021] where φ ∈ R 2×R×W , diff x and diff y represent the horizontal and vertical first-order difference operators;

[0022] Uniformly initialize several Gaussian functions on the 2D plane to construct multiple deformation transformation matrices at different positions Then sum these matrices to form a composite deformation transformation matrix Represents multiple distortion regions; to introduce deformations at multiple scales, a hierarchical structure is constructed where the influence range of the deformation at each layer increases with the level, and finally a deformation transformation matrix is obtained

[0023] Step 2: Distortion-Aware Contrastive Learning (DaCL)

[0024] The input pinhole image x pin and the deformed image x augmented by random Gaussian deformation def are input into the visual encoder V of the open-vocabulary segmentation network S to obtain visual features V pin and V def :

[0025] V pin = V(x pin ), V def = V(x def ) (4)

[0026] Among them, C is the number of feature channels;

[0027] The projection networks h pat and h pix are used to project these features:

[0028]

[0029] For the features of the image x def after deformation and will be obtained in the same way, and they are respectively shared with the features of the pinhole image x pin and and the same projection networks h pat and h pix ; among them, and

[0030] First, patch-level contrastive learning is applied to and

[0031]

[0032] Among them, N pat = H 2 × W 2 , p i , q j , q k are respectively The i-th, j-th, and k-th features, all with a dimension size of C 2 ; d represents the similarity metric, specifically, the exponential cosine similarity d(a, b) = exp(cos(a, b) / τ), where cos is the cosine similarity and τ is the temperature parameter;

[0033] Secondly, pixel-level contrastive learning is defined as:

[0034]

[0035] where N pix = H 1 × W 1 , m i , n j , n k are respectively the i-th, j-th, and k-th features, all with a dimension size of C 1 ; C(i) = C(j) indicates that i and j belong to the same category;

[0036] Step 3: The learning process of an open panoramic image segmentation method based on deformation enhancement and distortion contrast

[0037] An open panoramic image segmentation method based on deformation enhancement and distortion contrast does not rely on any specific open segmentation network S during training or testing, making it compatible with other off-the-shelf methods; the total loss function is defined as:

[0038] L = L seg + α(L pat + L pix ) (8)

[0039] where L seg represents the segmentation loss related to the selected segmentation network, and α represents a hyperparameter.

[0040] The beneficial effects of the present invention compared with the prior art are as follows: RGDA applies randomly parameterized Gaussian deformation to enhance the training set of labeled pinhole images, providing distorted images for the segmentation network. Secondly, DaCL performs patch-level and pixel-level contrastive learning between the images after pinhole and deformation enhancement to extract distortion-invariant and class-related features, reducing the feature space gap between pinhole images and panoramic images.

[0041] Extensive experiments conducted on four publicly available indoor and outdoor benchmarks have demonstrated the effectiveness of an open panoramic segmentation method based on deformation enhancement and distortion contrast proposed by the present invention. The improvements in the segmentation mIoU performance for Stanford2D3D and Matterport3D in indoor environments, and DensePASS and WildPASS in outdoor environments are 3.2%, 1.4%, 2.4%, and 1.7% respectively. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 It is the overall structure diagram of an open panoramic segmentation method based on deformation enhancement and distortion contrast;

[0043] Figure 2 It is the overall structure diagram of distortion-aware contrastive learning. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] The technical solutions of the present invention will be further described below in conjunction with the accompanying drawings and embodiments, but are not limited thereto. Any modifications or equivalent replacements to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention shall be covered within the protection scope of the present invention.

[0045] Random Gaussian Deformation Augmentation (RGDA) and Distortion-aware Contrastive Learning (DaCL) are the most important implementation manners. The rest can have alternatives, such as the selection of a specific segmentation network (for example, representative SAN and OOOPS can both be used as alternatives).

[0046] The present invention consists of two main modules: (1) Random Gaussian Deformation Augmentation (RGDA) enhances the pinhole image with a randomly parameterized Gaussian deformation network to provide distorted images for the segmentation network, thereby alleviating the problem of insufficient distortion perception. Specifically, RGDA constructs several Gaussian kernels with random parameters (i.e., amplitude, mean, and standard deviation), performs a first-order difference on these Gaussian kernels to construct a deformation field in the 2D plane, and then applies the deformation field to the pinhole image and its annotation to obtain the deformed image and the corresponding annotation. (2) Distortion-aware Contrastive Learning (DaCL) performs contrastive learning at the patch and pixel levels between the pinhole and deformed images to extract distortion-invariant and class-related features, thereby reducing the feature misalignment caused by distortion. Specifically, patch-level contrastive learning is applied to the visual features of the pinhole image and its deformed version, considering the patches at the same position as positive examples and the patches at different positions as negative examples. Pixel-level contrastive learning is performed on the visual features of the pinhole image and its deformed counterpart, using pixels of the same category as positive examples and pixels of different categories as negative examples.

[0047] An open panoramic image segmentation method based on deformation enhancement and distortion contrast proposed by the present invention includes two main components: Random Gaussian Deformation Augmentation (RGDA) and Distortion-Aware Contrastive Learning (DaCL), as Figure 1 shown. First, Random Gaussian Deformation Augmentation constructs multiple deformation fields with random Gaussian parameters (no learnable parameters) and applies them to the pinhole image x pin and the corresponding label y pin to obtain its deformed image x def and label y def . Then, the pinhole image and the deformed image are fed into the segmentation network S together to calculate the segmentation loss L seg , enhancing the segmentation network's perception of distorted images. Secondly, Distortion-Aware Contrastive Learning calculates the patch-level and pixel-level contrast losses L pat and L pix to adjust the visual features V pin and V def of x pin and x def to mitigate the feature shift caused by distortion. The proposed method does not rely on a specific open vocabulary segmentation network and can use any OVS or OPS method as the segmentation network S, such as SAN and OOOPS.

[0048] Example 1:

[0049] An open panoramic image segmentation method based on deformation enhancement and distortion contrast, the method is as follows:

[0050] The open panoramic segmentation (OPS) task aims to train an open vocabulary segmentation model only on the pinhole samples (x pin , y pin ) ∈ S and test it on the panoramic samples (x pan , y pan ) ∈ T, where the label spaces do not overlap, that is, new categories may be encountered during the test;

[0051] Step 1: Random Gaussian Deformation Augmentation (RGDA)

[0052] Random Gaussian Deformation Augmentation provides distorted images for the segmentation network S and alleviates the insufficient perception of distortion by the following means: (1) constructing 2D Gaussian kernels with random parameters at multiple scales; (2) applying first-order horizontal and vertical differences to these kernels in the 2D plane; (3) combining the results to form a deformation field, and then applying it to the pinhole image x pin to obtain the deformed image x def , as Figure 2 shown;

[0053] For an input image x with a label y pin ∈R 1×H×W apply a deformation transformation φ ∈ R pin ∈R 3×H×W to obtain a distorted image x with the corresponding label y 2×H×W ∈R def ∈R 1×H×W using a grid sampling function def ∈R 3×H×W as follows:

[0054]

[0055] The current challenge is how to construct the deformation transformation matrix φ to simulate the image distortion in panoramic images. This matrix φ needs to meet three main conditions: (1) Most importantly, it should produce image distortion similar to the appearance of panoramic images; (2) It should have good mathematical properties such as smoothness and differentiability to reduce the impact of image quality degradation caused by the deformation transformation; (3) It needs to meet some constraints and can be determined with a small number of randomly initialized parameters because there are no paired planar-panoramic images for backpropagation training.

[0056]

[0056] Fortunately, the 2D Gaussian function (or Gaussian kernel) meets the above three conditions: (1) Its shape is similar to the distorted appearance in fisheye images, simulating the distortion in panoramic images to a certain extent, as Figure 2 shown; (2) The Gaussian function inherently has good mathematical properties; (3) Adjacent points in the two-dimensional plane meet the constraints of the Gaussian function and only require the amplitude, mean, and standard deviation to completely determine the Gaussian function. Therefore, the present invention uses a 2D independent Gaussian function to construct the deformation transformation matrix φ:

[0057]

[0058] where x and y are the horizontal and vertical coordinates of a point (x, y) in the 2D plane, and A, μ, and σ are the amplitude, mean, and standard deviation of the Gaussian function respectively; μ x is the mean in the horizontal direction of the plane; μ y is the mean in the vertical direction of the plane; σ x is the standard deviation in the horizontal direction of the plane; σ y is the standard deviation in the vertical direction of the plane;

[0059] The matrix G obviously cannot be directly used as the deformation transformation matrix φ because its shape is R 1×H×W ; to maintain its good mathematical properties while constructing φ, horizontal and vertical first-order differences are applied to G to construct φ:

[0060] φ = [diff x(G), diff y (G)] (3)

[0061] where φ ∈ R 2×R×W , diff x and diff y represent the horizontal and vertical first-order difference operators; in the present invention, the first-order difference is selected instead of the first-order derivative because the first-order difference is more efficient, produces smaller numerical changes, and helps to avoid drastic local deformations.

[0062] The matrix φ obtained above can only simulate the distorted appearance of a single region in the 2D plane, while a panoramic image usually contains multiple distortions. To solve this problem, several Gaussian functions are uniformly initialized on the 2D plane to construct multiple deformation transformation matrices at different positions Then these matrices are summed to form a composite deformation transformation matrix represent multiple distorted regions; to introduce deformations at multiple scales, a hierarchical structure is constructed, where the influence range of the deformation at each layer increases with the level, as Figure 2 shown. Finally, the deformation transformation matrix

[0063] Step 2: Distortion-Aware Contrastive Learning (DaCL)

[0064] Distortion-Aware Contrastive Learning aims to help the segmentation network learn distortion invariance and class-related features, and mitigate the feature shift caused by the distortion between the pinhole and panoramic images. It performs patch-level and pixel-level contrastive learning on the visual features of a batch of pinhole and deformed images, as Figure 1 shown on the right. Specifically, patch-level contrastive learning regards the patch features at the same position in the pinhole and distorted images as positive examples, and the features at different positions as negative examples, enabling the segmentation network to learn features independent of the distortion. Pixel-level contrastive learning regards the pixel features of the same class in the pinhole and distorted images as positive examples, and the pixel features of different classes as negative examples, helping the network learn class-discriminative features.

[0065] The input pinhole image x pin and the deformed image x def augmented by random Gaussian deformation are input into the visual encoder V of the open-vocabulary segmentation network S to obtain the visual features V pin and V def :

[0066] V pin = V(x pin ), V def = V(x def ) (4)

[0067] where, C is the number of feature channels; the non - linear projection head is crucial for contrastive learning, and the present invention adopts the projection networks h pat and h pix to project these features:

[0068]

[0069] For the features def of the warped image x and they will be obtained in the same way, and they share the same projection networks h pin and h and with the features pat and h pix of the pinhole image x respectively; where, and

[0070] First, patch - level contrastive learning is applied toand

[0071]

[0072] where N pat = H 2 × W 2 , p i , q j , q k are respectively the i - th, j - th, and k - th features of with the dimension size of C 2 ; d represents the similarity metric, specifically, the exponential cosine similarity d(a, b)=exp(cos(a, b) / τ), where cos is the cosine similarity and τ is the temperature parameter, which is set to 0.1 in this present invention;

[0073] Secondly, pixel - level contrastive learning is defined as:

[0074]

[0075] where N pix = H 1 × W 1 , m i , n j , n k are respectively the i - th, j - th, and k - th features of with the dimension size of C 1 ; C(i)=C(j) indicates that i and j belong to the same category;

[0076] Step 3: The learning process of an open panoramic image segmentation method based on deformation enhancement and distortion contrast

[0077] An open panoramic image segmentation method based on deformation enhancement and distortion contrast does not rely on any specific open-word segmentation network S during training or testing, making it compatible with other off-the-shelf methods; the total loss function is defined as:

[0078] L = L seg + α(L pat + L pix ) (8)

[0079] where L seg represents the segmentation loss related to the selected segmentation network, α represents a hyperparameter, which is set to 0.1 in the present invention; the pinhole image x pin and the deformed image x def both need to calculate L seg . The specific calculation is well-known, and L seg is an ordinary image segmentation loss.

Claims

1. An open panorama segmentation method based on deformation enhancement and distortion contrast, characterized by: The method is as follows: first, random Gaussian deformation enhancement constructs multiple deformation fields with random Gaussian parameters (no learnable parameters) and applies them to the pinhole image x pin and the corresponding label y pin , to obtain its deformed image x def and label y def ; Then, the pinhole image and the deformed image are fed into the segmentation network S together to calculate the segmentation loss L seg , enhance the segmentation network’s perception of distorted images; Second, distortion-aware contrastive learning computes the patch-level and pixel-level contrastive losses L pat and L pix , to x pin and x def Visual features V pin and V def Adjustments are made to mitigate feature shifts caused by distortion.

2. The open panorama segmentation method based on deformation enhancement and distortion contrast according to claim 1, characterized in that: The method is specifically as follows: The Open Panoptic Segmentation (OPS) task aims to segment only the pinhole samples (x pin ,y pin )∈S, and train an open vocabulary segmentation model on the panoramic samples (x pan ,y pan )∈T, where the label space does not overlap, that is, new categories may be encountered during the test; Step 1: Random Gaussian Deformation Augmentation (RGDA) Random Gaussian deformation augmentation provides a distorted image to the segmentation network S to alleviate the distortion perception problem by: (1) constructing 2D Gaussian kernels with random parameters at multiple scales; (2) applying first-order horizontal and vertical differences to these kernels in the 2D plane; (3) combining the results to form a deformation field, which is then applied to the pinhole image x pin To obtain the deformed image x def ; To label y pin ∈R 1×H×W The input image x pin ∈R 3×H×W Apply the deformation transformation φ∈R 2×H×W , to obtain the corresponding label y def ∈R 1×H×W The distorted image x def ∈R 3×H×W , using the grid sampling function as follows: Use 2D independent Gaussian functions to construct the deformation transformation matrix φ: Where x, y are the horizontal and vertical coordinates of a point (x, y) in the 2D plane, A, μ and σ are the amplitude, mean and standard deviation of the Gaussian function respectively; μ x is the mean value in the transverse direction of the plane; μ y is the mean value in the longitudinal direction of the plane; σ x is the standard deviation in the transverse direction of the plane; σ y is the standard deviation in the longitudinal direction of the plane; Apply horizontal and vertical first-order differences on G to construct φ: φ=[diff x (G),diff y (G)] (3) where φ∈R 2×R×W , diff x and diff y represents the horizontal and vertical first-order difference operators; Initialize several Gaussian functions uniformly on the 2D plane to construct multiple deformation transformation matrices at different positions These matrices are then summed to form the composite deformation transformation matrix Represents multiple distortion regions; in order to introduce deformation at multiple scales, a hierarchical structure is constructed, in which the deformation influence range of each layer increases with the increase of the level, and finally the deformation transformation matrix is ​​obtained Step 2: Distortion-aware Contrastive Learning (DaCL) Input pinhole image x pin and the deformed image x augmented by random Gaussian deformation def is input into the visual encoder V of the open vocabulary segmentation network S to obtain the visual features V pin and V def : V pin =V(x pin ),V def =V(x def ) (4) in, C is the number of feature channels; Using projection network h pat and h pix To project these features: For the deformed image x def Features and will be obtained in the same way, and they are respectively pin Features and Share the same projection network h pat and h pix ;in, as well as First, patch-level contrastive learning is applied to and Among them, N pat =H2×W2, p i ,q j ,q k They are The i, j, and kth features of , all have a dimension size of C2; d represents the similarity measure, specifically, the exponential cosine similarity d(a, b) = exp(cos(a, b) / τ), where cos is the cosine similarity and τ is the temperature parameter; Second, pixel-level contrastive learning is defined as: Among them, N pix =H1×W1, m i ,,n j , n k They are The i-th, j-th, and k-th features have dimensions of C1; C(i) = C(j) means that i and j belong to the same category; Step 3: Learning process of an open panorama segmentation method based on deformation enhancement and distortion comparison A deformation-enhanced and distortion-contrast-based open panorama segmentation method does not rely on any specific open segmentation network S during training or testing, making it compatible with other off-the-shelf methods; the total loss function is defined as: L=L seg +α(L pat +L pix ) (8) in, L seg represents the segmentation loss associated with the chosen segmentation network and α represents a hyperparameter.