Night semantic segmentation method and device based on wavelet transform detail enhancement and text prompt

By combining wavelet transformation and text prompts deep learning models, the problems of lighting inhomogeneity and object boundary blur in night scenes are solved, and the accuracy and robustness of semantic segmentation of night scenes are improved.

CN120236080APending Publication Date: 2025-07-01QUZHOU UNIV
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510319727.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The lighting conditions of night scenes are complex and uneven, which makes the semantic segmentation task of night scenes difficult, the object boundaries are blurred, the texture is significantly weakened, and the target area identification is inaccurate.

Method used

A deep learning model using a three-modal feature extractor, a two-branch cross-modal feature interaction module and a multi-scale feature segmentation decoder is used to combine wavelet transformation and text cues to enhance the details of low-light areas and the target edges, and enhance the positioning and understanding of the target areas through natural language priors.

Benefits of technology

It improves the accuracy and robustness of semantic segmentation at night scenes, overcomes the problems of blurred object boundaries and significantly weakened textures, and achieves more accurate target area recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236080A_ABST
    Figure CN120236080A_ABST
Patent Text Reader

Abstract

The invention discloses a night semantic segmentation method and device based on wavelet transform detail enhancement and text prompt, and the method comprises the steps: obtaining a night image, carrying out the preprocessing of the night image, and carrying out the reconstruction of a wavelet image; and inputting the preprocessed night image and the image after wavelet transform reconstruction into a deep learning model for semantic segmentation to obtain a segmentation result of the night scene object. A new three-stage network structure is designed and formed, in the first stage, a three-mode feature extractor composed of an image encoder, a night semantic category encoder and a wavelet image encoder is used for extracting features, in the second stage, a double-branch cross-mode feature interaction module is designed, and the feature extraction is carried out through the image encoder. In the first stage, features of different spatial resolutions and semantic hierarchies and natural language priori of a target object are integrated, all-directional semantic information from coarse granularity to fine granularity is captured, in the third stage, a multi-scale feature segmentation decoder is introduced, details of a low-light area are enhanced, fine texture edges and target contours are captured, and the target object is obtained. Through positioning and understanding of the target area by the natural language prior enhancement model, the precision of night scene semantic segmentation can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image segmentation, and particularly relates to a nighttime semantic segmentation method and device based on wavelet transform detail enhancement and text prompts. Background Art

[0002] Nighttime scene semantic segmentation, as an important branch of image processing tasks, is of extremely important significance to multiple fields such as intelligent transportation, security monitoring, and autonomous driving. However, due to the complex lighting conditions in nighttime scenes, ranging from low-intensity natural light to diverse artificial light sources, the lighting distribution often shows non-uniformity, and the intensity and direction also vary greatly. These inherent complexities make the nighttime scene semantic segmentation task particularly difficult. In addition, external environmental factors further increase the difficulty of segmentation. Reflections and shadows on the object surface often interfere with the target area, affecting its clarity in the image. Various light sources in nighttime scenes, such as street lights, car lights, and billboards, are closely intertwined with the target objects, forming complex light and shadow relationships that blur the boundaries of the objects. Therefore, there is an urgent need for an effective nighttime scene semantic segmentation method to specifically address the specific challenges in nighttime image semantic segmentation and improve the accuracy and robustness of segmentation. Summary of the Invention

[0003] The present invention aims to overcome the above-mentioned drawbacks of the prior art and provides a nighttime semantic segmentation method and device based on wavelet transform detail enhancement and text prompts.

[0004] The present invention takes a three-modal feature extractor composed of an image encoder, a nighttime semantic category encoder, and a wavelet image encoder as the first stage, a newly designed dual-branch cross-modal feature interaction module as the second stage, integrates features of different spatial resolutions and semantic levels with the natural language prior of the target object, captures all-round semantic information from coarse-grained to fine-grained, and takes a segmentation decoder for multi-scale features as the third stage to enhance the details of low-light regions, capture fine texture edges and the edge contours of the target, and strengthen the model's positioning and understanding of the target area through the natural language prior, thus overcoming the difficult-to-segment problems such as blurred object boundaries, significantly weakened textures, and inaccurate recognition of target areas in nighttime scenes.

[0005] To achieve the above object, the technical solutions adopted by the present invention are as follows:

[0006] The first aspect of the present invention relates to a nighttime semantic segmentation method based on wavelet transform detail enhancement and text prompts, including the following steps:

[0007] Collect a nighttime image dataset, and perform preprocessing on the nighttime images and reconstruction of the images after wavelet transform;

[0008] Input the preprocessed night image into a deep learning model for semantic segmentation to obtain the segmentation result of night objects;

[0009] Among them, the deep learning model has a three-stage network structure, including a three-modal feature extractor composed of an image encoder, a night semantic category encoder, and a wavelet image encoder, a dual-branch cross-modal feature interaction module, and a segmentation decoder for multi-scale features. Among them:

[0010] The three-modal feature extractor receives the preprocessed night image, the wavelet reconstructed image, and the night semantic category hint, and outputs four sizes of image feature maps {F1, F2, F3, F4}, wavelet feature f, and category text embedding t through the feature extraction structures of each modality;

[0011] The dual-branch cross-modal feature interaction module has two branches: a text-image feature interaction branch and a wavelet-guided detail enhancement branch. The text-image feature interaction branch performs feature interaction on the feature map F4 and t to obtain the visual text matching score Score. The wavelet-guided detail enhancement branch fuses the image features {F1, F2, F3, F4} and the wavelet feature f at different scales to obtain the feature maps {Ff1, Ff2, Ff3, Ff4};

[0012] The segmentation decoder for multi-scale features decodes the feature maps {Ff1, Ff2, Ff3, Ff4} and the visual text matching score Score to obtain the final segmentation map F seg , and use the final segmentation map as the segmentation result of night objects.

[0013] The following also provides several optional methods, which are not additional limitations to the above overall solution, but only further supplements or optimizations. Without technical or logical contradictions, each optional method can be combined with the above overall solution separately, or multiple optional methods can be combined with each other.

[0014] Preferably, the preprocessing includes scaling the night image to an image size of 512×1024 and performing data augmentation to expand the dataset, including using horizontal and vertical flipping and rotation.

[0015] Preferably, the reconstruction of the image after wavelet transform includes performing wavelet transform on the preprocessed night image to extract the low-frequency component and the horizontal, vertical, and diagonal high-frequency components, increasing the contribution of the high-frequency components through a weight factor, and then reconstructing the image through inverse wavelet transform.

[0016] Preferably, the structure of the three-modal feature extractor is defined as an image encoder, a wavelet image encoder, and a night semantic category encoder according to the input modality;

[0017] The image encoder is composed of a Swin Transformer network, the wavelet image encoder is composed of ResNet18, and the night semantic category encoder is composed of the text encoder of the CLIP model;

[0018] Among them, the image encoder outputs multi-scale image feature maps {F1, F2, F3, F4}, the wavelet image encoder outputs a wavelet feature map f, and the night semantic category encoder outputs a category text embedding t.

[0019] Preferably, the dual-branch cross-modal feature interaction module fuses the multi-scale image feature maps {F1, F2, F3, F4}, the wavelet feature map f, and the category text embedding t and then outputs feature maps {Ff1, Ff2, Ff3, Ff4}, including:

[0020] For the image feature map F4 and the category text embedding t, in the corresponding text-image feature interaction branch, on the one hand, global average pooling is performed on F4 to obtain a global feature Then the global feature is concatenated with the original feature in the spatial dimension to obtain a concatenated feature Subsequently, it is fed into a multi-head attention layer (MHSA) for processing to obtain On the other hand, in order to more effectively use the input text to guide visual object detection and localization, through a language-guided query selection mechanism, and t are queried to select visual information I more relevant to the input text Nq ; Subsequently, the cross-attention mechanism in the Transformer decoder is used to simulate the interaction between vision and language, and I Nq and t are processed to obtain v post , and finally, the residual connection is used to update the text embedding t;

[0021] For the image features {F1, F2, F3, F4} and the wavelet feature f at different scales, in the corresponding wavelet-guided detail enhancement branch, starting from the lowest resolution, the two input feature maps F4 and f are mapped to the same channel dimension through convolution to obtain new feature maps F′4 and f′4; Next, the image feature map F′4 is used as the key and value, and the wavelet image f′4 is used as the query, and deep information exchange and complementation are performed through the multi-head cross-attention mechanism to generate a feature map Ff4 with richer texture details. Then, after the feature map Ff4 is added to the original image feature F4, an upsampling operation is performed to obtain a wavelet feature f3 with the same resolution as the previous layer. The above process is repeated until the highest resolution to obtain output features {Ff1, Ff2, Ff3, Ff4} at different resolutions.

[0022] Preferably, the multi-scale feature segmentation decoder decodes the multi-scale feature maps {Ff1, Ff2, Ff3, Ff4} and the visual text matching score Score to obtain the final segmentation map F seg , and uses the final segmentation map as the night object segmentation result, including:

[0023] First, the lowest-resolution feature map Ff4 and the visual text matching score Score are concatenated in the spatial dimension to utilize the visual language prior. Secondly, the multi-scale features {Ff1, Ff2, Ff3, [Ff4, Score]} are aligned and normalized through the Feature Pyramid Network (FPN) to obtain {Ff′1, Ff′2, Ff′3, Ff′4}. Then, C learnable query embeddings Q = {q1, q2,..., q C} (C is the number of categories) interact with the transformed multi-scale features through a cross-scale attention mechanism, and after passing through L layers of the transformer decoder and the classification prediction head, the final semantic segmentation prediction map F of the night object is obtained seg .

[0024] The second aspect of the present invention relates to a night semantic segmentation device based on wavelet transform detail enhancement and text prompts, including a memory and one or more processors. Executable code is stored in the memory, and when the one or more processors execute the executable code, it is used to implement the night semantic segmentation method based on wavelet transform detail enhancement and text prompts of the present invention

[0025] The third aspect of the present invention relates to a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the night semantic segmentation method based on wavelet transform detail enhancement and text prompts of the present invention

[0026] A night semantic segmentation method based on wavelet transform detail enhancement and text prompts provided by the present invention performs feature extraction through a three-modal feature extractor composed of an image encoder, a night semantic category encoder, and a wavelet image encoder in the first stage. In the second stage, it integrates features of different spatial resolutions and semantic levels with the natural language prior of the target object to capture all-round semantic information from coarse-grained to fine-grained. In the third stage, a multi-scale feature segmentation decoder is introduced to enhance the details in low-light regions, capture subtle texture edges and target contours, and strengthen the model's localization and understanding of the target region through the natural language prior, which can effectively improve the accuracy of night scene semantic segmentation

[0027] The advantages of the present invention are: improving the accuracy and robustness of night scene semantic segmentation BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1It is a flowchart of the method of the present invention;

[0029] Figure 2 It is a model diagram of the method of the present invention;

[0030] Figure 3 It is a schematic diagram of the device of the present invention. Detailed implementation manners

[0031] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the description of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0033] Embodiment 1

[0034] This embodiment provides a nighttime semantic segmentation method based on wavelet transform detail enhancement and text prompts to improve the accuracy of nighttime scene semantic segmentation.

[0035] As Figure 1 shown, the nighttime semantic segmentation method based on wavelet transform detail enhancement and text prompts in this embodiment includes the following steps:

[0036] Step 1: Obtain a nighttime image, and perform preprocessing on the nighttime image and reconstruct the wavelet image.

[0037] To meet the input requirements of the neural network, in this embodiment, the nighttime image needs to be scaled to an image size of 512×1024. The preprocessing in this embodiment lies in scaling the image size and expanding the data set through data enhancement methods such as horizontal and vertical flipping and rotation. To meet the input requirements of the neural network, in this embodiment, it is necessary to reconstruct the wavelet image of the nighttime image. The reconstruction of the wavelet image in this embodiment lies in performing wavelet transform on the preprocessed nighttime image to extract low-frequency components and extract horizontal, vertical, and diagonal high-frequency components, and improving the contribution of the high-frequency components through a weight factor, and then reconstructing the image through inverse wavelet transform.

[0038] Step 2: Input the preprocessed nighttime image into a deep learning model for semantic segmentation to obtain a nighttime scene segmentation result.

[0039] In this embodiment, a deep learning model is introduced for semantic segmentation, which not only reduces the complexity of the segmentation operation, but also improves the accuracy and stability of the segmentation operation. As Figure 2 shown, the deep learning model of this embodiment is a three-stage network structure, including a three-modal feature extractor composed of an image encoder, a night semantic category encoder, and a wavelet image encoder, a dual-branch cross-modal feature interaction module, and a segmentation decoder for multi-scale features. The deep learning model adopted in this embodiment will be described below through a detailed introduction of each module.

[0040] 1) The three-modal feature extractor is located in the first stage of the network structure, receives the preprocessed night image, wavelet reconstruction image, and night semantic category prompt, and outputs four-sized image feature maps {F1, F2, F3, F4}, wavelet feature f, and category text embedding t through the feature extraction structures of each modality.

[0041] The structure of the three-modal feature extractor in this embodiment is defined as an image encoder, a wavelet image encoder, and a night semantic category encoder according to the input modality.

[0042] 2) The dual-branch cross-modal feature interaction module is located in the second stage of the network structure. The image feature map F4 and the category text embedding t output in the first stage are subjected to feature interaction through the text-image feature interaction branch to obtain the visual text matching score Score. The different-scale image features {F1, F2, F3, F4} and the wavelet feature f output in the first stage are fused through the wavelet-guided detail enhancement branch to obtain the feature maps {Ff1, Ff2, Ff3, Ff4}.

[0043] For the image feature map F4 and the category text embedding t, in the corresponding text-image feature interaction branch, on the one hand, global average pooling is performed on F4 to obtain a global feature Then the global feature is concatenated with the original feature in the spatial dimension to obtain the concatenated feature Subsequently, it is sent to the multi-head attention layer (MHSA) for processing to obtain On the other hand, in order to more effectively use the input text to guide the visual object detection and localization, through the language-guided query selection mechanism, and t are queried to select the visual information I more relevant to the input text Nq ; Subsequently, the cross-attention mechanism in the Transformer decoder is used to simulate the interaction between vision and language, and I Nq and t are processed to obtain v post , and finally, the residual connection is used to update the text embedding t;

[0044] For image features {F1, F2, F3, F4} and wavelet features f at different scales, the corresponding wavelet-guided detail enhancement branches start from the lowest resolution. The two input feature maps F4 and f are mapped to the same channel dimension through convolution to obtain new feature maps F′4 and f′4. Next, taking the image feature map F′4 as the key and value, and the wavelet image f′4 as the query, deep information exchange and complementation are performed through the multi-head cross-attention mechanism to generate a feature map Ff4 with richer texture details. Then, after adding the feature map Ff4 to the original image feature F4, through the upsampling operation, the wavelet feature f3 with the same resolution as the previous layer is obtained. Repeat the above process until the highest resolution to obtain output features {Ff1, Ff2, Ff3, Ff4} at different resolutions.

[0045] 4) Preferably, the segmentation decoder of the multi-scale features decodes the multi-scale feature maps {Ff1, Ff2, Ff3, Ff4} and the visual text matching score Score to obtain the final segmentation map F seg and takes the final segmentation map as the night object segmentation result.

[0046] First, the lowest-resolution feature map Ff4 and the visual text matching score Score are concatenated in the spatial dimension to utilize the visual language prior. Secondly, the multi-scale features {Ff1, Ff2, Ff3, [Ff4, Score]} are aligned and normalized through the Feature Pyramid Network (FPN) to obtain {Ff′1, Ff′2, Ff′3, Ff′4}. Then, C learnable query embeddings Q = {q1, q2,..., q C}(C is the number of classes) interact with the transformed multi-scale features through the cross-scale attention mechanism, and after passing through L layers of the transformer decoder and the classification prediction head, the semantic segmentation prediction map F of the final night object is obtained. seg 。

[0047] To ensure the application effect of the deep learning model, the deep learning model needs to be pre-trained. This embodiment provides a training process as follows:

[0048] Step S1: Obtain a night image training dataset with annotated street scene segmentation masks. First, adjust the image size of the night image training dataset, that is, adjust it to a size of 512×1024. Secondly, perform data augmentation such as horizontal and vertical flipping and rotation. Then, reconstruct the wavelet image for the preprocessed image, and then start training in batches.

[0049] This embodiment performs data augmentation such as horizontal and vertical flipping and rotation on the training dataset. Among them, horizontal and vertical flipping means that the pictures trained in each batch are flipped horizontally or vertically with a probability of 50%.

[0050] The reconstruction of the wavelet image in this embodiment lies in performing wavelet transform on the preprocessed nighttime images in the training dataset to extract low-frequency components and extract high-frequency components in the horizontal, vertical, and diagonal directions, and enhancing the contribution of the high-frequency components through weight factors, and then reconstructing the image through inverse wavelet transform.

[0051] Step S2: Input the preprocessed nighttime image, the wavelet-reconstructed image, and the nighttime semantic category prompt into the three-modal feature extractor, and obtain four sizes of image feature maps {F1, F2, F3, F4}, wavelet feature f, and category text embedding t through the feature extraction structures of each modality.

[0052] This application uses a three-modal feature extractor composed of an image encoder, a wavelet image encoder, and a nighttime semantic category encoder for feature extraction. Training is carried out in batches. During the training process, the batch size is 16 (i.e., 16 images are processed in each batch), the optimizer is AdamW, the initial learning rate is 0.0001, and the weight decay is 0.05. The learning rate adjustment strategy uses polynomial decay (PolyLR). From the start of training to the 90,000th iteration, the learning rate decays according to a power function with a decay coefficient of 0.9.

[0053] After putting a 512×1024 nighttime image into the image encoder of the three-modal feature extractor, four sizes of feature maps {F1, F2, F3, F4} of 256×512, 128×256, 64×128, and 32×64 are output; after putting a 512×1024 wavelet-reconstructed image into the wavelet image encoder in the three-modal feature extractor, a wavelet feature f of size 32×64 is output; after inputting the category names of 19 semantic classes as text prompts into the nighttime semantic category encoder of the three-modal feature extractor, a category text embedding t of 19×512 is output.

[0054] Step S3: In the dual-branch cross-modal feature interaction module, the text-image feature interaction branch performs feature interaction on the feature map F4 and t to obtain the visual-text matching score Score, and the wavelet-guided detail enhancement branch fuses the image features {F1, F2, F3, F4} at different scales and the wavelet feature f to obtain the feature maps {Ff1, Ff2, Ff3, Ff4}.

[0055] Step S4: Input the processed feature maps {Ff1, Ff2, Ff3, Ff4} and the visual-text matching score Score into the segmentation decoder of the multi-scale features for feature decoding to obtain the final segmentation map F segCalculate the loss and perform backpropagation to update the network parameters to complete the training of the network. Additionally, since the night semantic segmentation task targets common night-time urban road scenes, there are 19 categories of labels, including: road, sidewalk, building, wall, fence, pole, traffic light, traffic sign, vegetation, terrain, sky, person, cyclist, car, truck, bus, train, motorcycle, and bicycle. Therefore, the final number of channels is only 19 channels.

[0056] Calculate the main loss of the segmentation prediction map and the ground truth label through the Focal Loss function, and calculate the auxiliary loss of the ground truth label and the visual text matching score through the Cross-Entropy Loss function to implicitly utilize the natural language prior. The specific formulas are as follows:

[0057]

[0058] where y is the label value, is the predicted value, Score is the visual text matching score, γ = 2.0, λ1 = 0.8, λ2 = 0.2.

[0059] It should be noted that the calculations of Cross-Entropy Loss and Focal Loss are already relatively mature technologies in this field and will not be elaborated here.

[0060] Thus, the loss between the predicted value and the ground truth value is obtained. Before the end of each batch, backpropagation is performed to reduce the loss. At the same time, the network parameters are updated, and the training of the next batch starts until all the training data of all batches are trained. Finally, the trained weights are obtained, and all the updated parameters will be saved in the Works weight file.

[0061] Embodiment 2

[0062] Referring to Figure 3 , this embodiment relates to a night semantic segmentation device based on wavelet transform detail enhancement and text prompt, including a memory and one or more processors. An executable code is stored in the memory. When the one or more processors execute the executable code, it is used to implement the night semantic segmentation method based on wavelet transform detail enhancement and text prompt in Embodiment 1.

[0063] Embodiment 3

[0064] This embodiment relates to a computer-readable storage medium on which a program is stored. When the program is executed by a processor, it implements the night semantic segmentation method based on wavelet transform detail enhancement and text prompt in Embodiment 1.

[0065] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as falling within the scope described in this specification.

[0066] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the appended claims.

Claims

1. A nighttime semantic segmentation method based on wavelet transform detail enhancement and text prompt, characterized in that: The steps include: Collect night image data sets, preprocess night images, and reconstruct images after wavelet transformation; The preprocessed nighttime images and the images reconstructed by wavelet transform are input into the deep learning model for semantic segmentation to obtain the segmentation results of nighttime objects; The deep learning model is a three-stage network structure, including a trimodal feature extractor consisting of an image encoder, a night semantic category encoder and a wavelet image encoder, a dual-branch cross-modal feature interaction module, and a multi-scale feature segmentation decoder, wherein: The trimodal feature extractor receives the preprocessed nighttime image, the wavelet reconstructed image, and the nighttime semantic category prompt, and outputs four-sized image feature maps {F1, F2, F3, F4}, wavelet features f, and category text embedding t through the feature extraction structure of each modality; The dual-branch cross-modal feature interaction module has two branches: a text-image feature interaction branch and a wavelet-guided detail enhancement branch. The text-image feature interaction branch performs feature interaction on feature maps F4 and t to obtain a visual text matching score Score. The wavelet-guided detail enhancement branch fuses image features {F1, F2, F3, F4} of different scales with wavelet features f to obtain feature maps {Ff1, Ff2, Ff3, Ff4}. The multi-scale feature segmentation decoder decodes the feature map {Ff1, Ff2, Ff3, Ff4} and the visual text matching score Score to obtain the final segmentation map F seg , and the final segmentation map is used as the segmentation result of night objects.

2. The nighttime semantic segmentation method based on wavelet transform detail enhancement and text prompting as claimed in claim 1, characterized in that: The preprocessing includes scaling the nighttime images to an image size of 512×1024, and performing data enhancement to expand the data set, including using horizontal and vertical flipping and rotation.

3. The nighttime semantic segmentation method based on wavelet transform detail enhancement and text prompting as claimed in claim 1, characterized in that: The reconstruction of the wavelet transformed image includes performing wavelet transform on the preprocessed nighttime image to extract low-frequency components and extracting horizontal, vertical and diagonal high-frequency components, and increasing the contribution of the high-frequency components by weight factors, and then reconstructing the image by inverse wavelet transform.

4. The nighttime semantic segmentation method based on wavelet transform detail enhancement and text prompting as claimed in claim 1, characterized in that: The structure of the tri-modal feature extractor is defined as an image encoder, a wavelet image encoder, and a night semantic category encoder according to the input modality; The image encoder is composed of the SwinTransformer network, the wavelet image encoder is composed of ResNet18, and the night semantic category encoder is composed of the text encoder of the clip model; The image encoder outputs multi-scale image feature maps {F1, F2, F3, F4}, the wavelet image encoder outputs wavelet feature maps f, and the night semantic category encoder outputs category text embedding t.

5. The nighttime semantic segmentation method based on wavelet transform detail enhancement and text prompting as claimed in claim 1, characterized in that: The dual-branch cross-modal feature interaction module performs feature interaction on the image feature map F4 and the category text embedding t through the text-image feature interaction branch to obtain the visual text matching score Score, and fuses the image features {F1, F2, F3, F4} of different scales with the wavelet feature f through the wavelet-guided detail enhancement branch to obtain the feature map {Ff1, Ff2, Ff3, Ff4}, including: For the image feature map F4 and the category text embedding t, the corresponding text-image feature interaction branch performs global average pooling on F4 to obtain a global feature Then the global features Splice it with the original feature in the spatial dimension to get the spliced ​​feature It is then sent to the multi-head attention layer (MHSA) for processing On the other hand, in order to more effectively use the input text to guide visual object detection and positioning, a language-guided query selection mechanism is used to and t to select visual information I that is more relevant to the input text Nq ; Then use the cross-attention mechanism in the Transformer decoder to simulate the interaction between vision and language. I Nq and t to obtain v post ,Finally, the text embedding t is updated using the residual connection; For image features {F1, F2, F3, F4} and wavelet features f of different scales, the corresponding wavelet-guided detail enhancement branch starts from the lowest resolution and maps the two input feature maps F4 and f to the same channel dimension through convolution to obtain new feature maps F′4 and f′4; next, the image feature map F′4 is used as the key and value, and the wavelet image f′4 is used as the query, and deep information exchange and complementation are performed through the multi-head cross-attention mechanism to generate a feature map Ff4 with richer texture detail information. Then, the feature map Ff4 is added to the original image feature F4 and then up-sampled to obtain the wavelet feature f3 with the same resolution as the previous layer. The above process is repeated until the highest resolution, and output features {Ff1, Ff2, Ff3, Ff4} of different resolutions are obtained.

6. The nighttime semantic segmentation method based on wavelet transform detail enhancement and text prompting as claimed in claim 1, characterized in that: The multi-scale feature segmentation decoder decodes the multi-scale feature map {Ff1, Ff2, Ff3, Ff4} and the visual text matching score Score to obtain the final segmentation map F seg , the final segmentation map is used as the night object segmentation result, including: First, the lowest resolution feature map Ff4 is concatenated with the visual text matching score Score in the spatial dimension to explicitly utilize the visual language prior. Second, the multi-scale features {Ff1, Ff2, Ff3, [Ff4, Score]} are aligned and normalized through the feature pyramid network FPN to obtain {Ff′1, Ff′2, Ff′3, Ff′4}. Then, C learnable query embeddings Q = {q1, q2, ..., q C } interacts with the transformed multi-scale features through a cross-scale attention mechanism. C is the number of categories. After L layers of transformer decoders and classification prediction heads, the final semantic segmentation prediction map F of night objects is obtained. seg .

7. A nighttime semantic segmentation device based on wavelet transform detail enhancement and text prompt, characterized in that: It comprises a memory and one or more processors, wherein the memory stores executable codes, and when the one or more processors execute the executable codes, they are used to implement the nighttime semantic segmentation method based on wavelet transform detail enhancement and text prompt as described in any one of claims 1 to 6.

8. A computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, the night semantic segmentation method based on wavelet transform detail enhancement and text prompting as described in any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Fine-grained image classification method based on deep wavelet feature fusion

    CN120673186A

  • Fine-grained image classification method based on deep wavelet feature fusion

    CN120673186B

  • Scene text super-resolution method and system based on text prior and stationary wavelet domain transformation

    CN120807290A

  • A Scene Text Super-Resolution Method and System Based on Text Prior and Stationary Wavelet Domain Transform

    CN120807290B

  • Unmanned aerial vehicle image enhancement method and system

    CN120953059A