Open-Vocabulary Semantic Segmentation Method for Aerial Drone Images with Vision-Language Fusion

By constructing a visual language fusion segmentation model, the open vocabulary semantic segmentation problem of drone aerial images in complex backgrounds is solved, high-precision and robust segmentation of known and unknown categories is achieved, and the processing capability of drone aerial images is improved.

CN120014280BActive Publication Date: 2025-07-04CIVIL AVIATION FLIGHT UNIV OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510470102.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-04
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

In the open vocabulary semantic segmentation of drone aerial images, the prior art faces the problems of insufficient complex background processing capabilities, high computing resource consumption, low accuracy of new category recognition and poor stability in harsh environments. Especially when processing drone aerial images, the traditional method has insufficient generalization capabilities and high cost of labeling data.

Method used

A visual language fusion segmentation model is constructed, and aerial photography mixed images under different environmental conditions are collected by drones to generate language description data. The visual-language feature extraction model, heterogeneous cross-modal graph fusion model and semantic segmentation model are used, and a variety of attention mechanisms and multi-level fusion modules are combined to achieve high-precision and robust segmentation of aerial images.

Benefits of technology

It realizes high-precision segmentation of aerial images of known and unknown categories in complex scenarios, improves the dynamic balance of local details and global features, enhances the ability to fusion across modal semantic relationships, and improves the accuracy and stability of segmentation effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014280B_ABST
    Figure CN120014280B_ABST
Patent Text Reader

Abstract

The present invention discloses an open-vocabulary semantic segmentation method for UAV aerial images integrating vision and language, which relates to the field of multi-modal artificial intelligence technology. Based on multiple attention mechanisms, multi-level fusion modules, and dynamic adjustment mechanisms, this method constructs a vision-language fusion segmentation model to ensure high-precision and robust segmentation effects for aerial images of known and unknown categories in complex scenarios; uses VIT and Mamba models to extract global image information and local image details, and adopts adaptive weighted fusion to achieve dynamic balance between global and local features, and uses deformable convolution to strengthen local structures to ensure accurate expression of the overall scene semantics; uses a heterogeneous cross-modal graph fusion model to integrate cross-modal semantic relationships at greater distances, and continuously fuses multi-dimensional information from vision, text, and domain knowledge.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimodal artificial intelligence, and particularly to an open-vocabulary semantic segmentation method for drone aerial images with visual-language fusion. Background Art

[0002] In recent years, drone aerial images have played an increasingly important role in remote sensing applications such as disaster assessment, precision agriculture, and urban planning. However, traditional semantic segmentation techniques mainly rely on fully supervised deep learning methods, which usually require a large amount of manually annotated data and are only trained for predefined closed-set categories. This closed-set recognition method often shows insufficient generalization ability and inaccurate recognition when facing unknown categories frequently appearing in practical applications. In addition, drone aerial images usually present unique geographical and structural features, such as sparse distribution of target objects, large scale variations, and complex backgrounds, making traditional models trained on natural images have obvious limitations in processing such images. At the same time, obtaining comprehensively covered annotated data also faces high costs and operational difficulties, further restricting the effectiveness and practicality of traditional methods in practical applications.

[0003] Multimodal fusion technology has shown great potential in the field of image understanding. By introducing text description information, the model can obtain semantic information that is difficult to directly extract from visual signals in the image, thereby making up for the deficiencies of traditional visual models in open-vocabulary recognition. However, current multimodal methods mainly target natural images, and there are still certain difficulties in dealing with the complex geographical and structural features in drone aerial images, and a mature and efficient solution has not yet been formed.

[0004] In summary, in view of the deficiencies of the existing technology in dealing with complex backgrounds, computational resource consumption, accuracy of new category recognition, and stability in harsh environments, there is an urgent need for a new open-vocabulary semantic segmentation method for drone aerial images based on visual and language fusion. Summary of the Invention

[0005] The purpose of the present invention is to provide an open-vocabulary semantic segmentation method for drone aerial images with visual-language fusion to improve the above technical problems.

[0006] To achieve the above object of the invention, the embodiments of the present invention provide the following technical solutions:

[0007] An open-vocabulary semantic segmentation method for drone aerial images with visual-language fusion includes:

[0008] Using a drone to collect unclassified aerial mixed images under different environmental conditions and generating corresponding language description data;

[0009] Construct a visual-language fusion segmentation model; the visual-language fusion segmentation model includes a vision-language feature extraction model, a heterogeneous cross-modal graph fusion model, and a semantic segmentation model;

[0010] Input the unclassified aerial mixed images and language description data into the vision-language feature extraction model, and output multi-scale spatio-temporal visual features and language features;

[0011] Input the language features and multi-scale spatio-temporal visual features into the heterogeneous cross-modal graph fusion model, and output vision-language matching features;

[0012] Input the vision-language matching features into the semantic segmentation model to complete the semantic segmentation of the aerial images.

[0013] Furthermore, the method for using a drone to collect unclassified aerial mixed images under different environmental conditions and generate corresponding language description data includes:

[0014] Use a drone platform to collect initial unclassified aerial images and perform preprocessing to obtain corresponding unclassified aerial images;

[0015] Select unclassified aerial images that meet the selection conditions and perform inspection and stitching to obtain unclassified aerial mixed images; the selection conditions are different shooting angles or shooting environments;

[0016] Use the GPT model to generate description texts for each unclassified aerial mixed image to obtain language description data.

[0017] Furthermore, the vision-language feature extraction model includes a parallel vision feature extraction sub-model and a language feature extraction sub-model; the vision feature extraction sub-model includes a series-connected vision feature extraction module and a multi-scale spatio-temporal visual feature fusion module; the vision feature extraction module includes a parallel VIT global feature extraction module and a local detail feature extraction module; the VIT global feature extraction module includes a Patch-Embedding layer and a transformer layer based on the self-attention mechanism; the multi-scale spatio-temporal visual feature fusion module includes a series-connected dynamic feature fusion layer and a deformable convolutional layer; the local detail feature extraction module uses the Mamba model; the language feature extraction sub-model is the BERT model;

[0018] The heterogeneous cross-modal graph fusion model uses a heterogeneous cross-modal graph network based on the multi-head attention mechanism;

[0019] The semantic segmentation model uses a lightweight U-Net++ model; the lightweight U-Net++ model includes an encoder connected in series and a decoder based on channel-spatial dual attention; the encoder includes N decoding layers; the decoder based on channel-spatial dual attention includes N decoding modules, and every two decoding modules are connected through a multi-level fusion module; each decoding module includes a CSDA layer and a decoding layer connected in series.

[0020] Further, the training process of the vision-language feature extraction model includes:

[0021] Obtain unclassified aerial mixed training images and corresponding language description training data;

[0022] Input the unclassified aerial mixed training images into the Patch-Embedding layer to obtain unclassified aerial mixed training segmentation images;

[0023] Input the unclassified aerial mixed training segmentation images into the transformer layer based on the self-attention mechanism to obtain unclassified aerial mixed training global feature maps;

[0024] Input the unclassified aerial mixed training images into the Mamba model to obtain unclassified aerial mixed training local feature maps;

[0025] Input the unclassified aerial mixed training local feature maps and the unclassified aerial mixed training global feature maps into the multi-scale spatio-temporal visual feature fusion module to obtain multi-scale spatio-temporal visual training features;

[0026] Input the language description training data into the BERT model to obtain language training features;

[0027] Based on the multi-scale spatio-temporal visual training features and the language training features, adjust the weight parameters of the visual feature extraction sub-model.

[0028] Further, the step of inputting the language description training data into the BERT model to obtain language training features includes:

[0029] Generate a corresponding continuous vector text representation based on the language description training data;

[0030] Based on the unclassified aerial mixed training local feature maps and the unclassified aerial mixed training all feature maps, set a group of visual prototypes;

[0031] Calculate the CSA loss function based on the positive / negative similarity of the continuous vector text representation;

[0032] Based on the CSA loss function, match and align the visual prototypes and the continuous vector text representation to obtain a matching result;

[0033] Generate corresponding language training features based on the matching results and continuous vector text representations.

[0034] Furthermore, the training process of the heterogeneous cross-modal graph fusion model includes:

[0035] Construct an initial heterogeneous cross-modal graph; obtain language training features and multi-scale spatio-temporal visual training features and use them as text nodes and visual nodes respectively, and map them into the initial heterogeneous cross-modal graph;

[0036] Calculate the node similarity between the text nodes and the visual nodes; set the weights of the edges and the knowledge graph based on each node similarity, construct the corresponding edges, and obtain the heterogeneous cross-modal graph;

[0037] According to the multi-head graph attention mechanism, obtain the aggregation information of each node and its adjacent nodes;

[0038] Based on each aggregation information, update the representation of each node through multi-hop aggregation of the multi-head graph attention mechanism; the representation of the node is the visual-language matching training feature, which includes language features and multi-scale spatio-temporal visual features;

[0039] Based on the updated representation of the node, optimize the heterogeneous cross-modal graph, that is, optimize the heterogeneous cross-modal graph fusion model.

[0040] Furthermore, the training process of the semantic segmentation model includes:

[0041] Obtain the visual-language matching training features and input them into the encoder to obtain the i-th visual-language matching training encoding feature corresponding to each encoding layer; except that the input of the first encoding layer is the visual-language matching training features, the input of the i-th encoding layer is the output of the i-1-th encoding layer;

[0042] Input the N-th visual-language matching training feature into the N-th decoding module, and output to obtain the N-th visual-language matching decoding result;

[0043] Input the N-th visual-language matching decoding result and the N-1-th visual-language matching training feature into the multi-level fusion module, and output to obtain the N-1-th visual-language matching training fusion feature;

[0044] Input the N-1-th visual-language matching training fusion feature into the N-1-th decoding module, and output to obtain the N-1-th visual-language matching decoding result;

[0045] Repeat the processing process of the multi-level fusion module and the N-1-th decoding module until the output of the first decoding module is obtained, that is, obtain the semantic segmentation training result;

[0046] Based on the semantic segmentation training results, calculate the dynamic weight triplet loss function and the generalized intersection over union loss function;

[0047] Based on the dynamic weight triplet loss function and the generalized intersection over union loss function, adjust the weight parameters of the semantic segmentation model.

[0048] Furthermore, the formula corresponding to the dynamic weight triplet loss function is:

[0049] ;

[0050] where and represent the triplet dynamic weight parameter and the predetermined hyperparameter respectively, represents the expectation operation, represents the probability distribution predicted by the semantic segmentation model for each category in the visual - language matching training feature for each category in it, represents the maximum value function.

[0051] Furthermore, the visual - language fusion segmentation model also includes an optimization process, and the corresponding steps are:

[0052] Use the NAS algorithm to process the initially optimized visual - language fusion segmentation model to obtain the once - optimized visual - language fusion segmentation model;

[0053] Perform parameter pruning and refinement on the once - optimized visual - language fusion segmentation model to obtain the twice - optimized visual - language fusion segmentation model;

[0054] Use the quantization - aware training method to process the twice - optimized visual - language fusion segmentation model to obtain the optimized visual - language fusion segmentation model.

[0055] The beneficial effects of the present invention are as follows:

[0056] Based on multiple attention mechanisms, multi - level fusion modules, and dynamic adjustment mechanisms, this method constructs a visual - language fusion segmentation model, ensuring high - precision and robust segmentation effects for aerial images of known and unknown categories in complex scenarios; uses the VIT and Mamba models to extract global image information and local image details, and adopts adaptive weighted fusion to achieve dynamic balance between global and local features, uses deformable convolution to strengthen local structures, ensures accurate expression of the overall scene semantics, and enhances the efficiency and accuracy of visual - language fusion features; uses a heterogeneous cross - modal graph fusion model to integrate cross - modal semantic relationships at a greater distance, and continuously fuses multi - dimensional information from vision, text, and domain knowledge. Description of the Drawings

[0057] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0058] Figure 1 It is the method flowchart in the embodiment of the present invention;

[0059] Figure 2 It is the structural diagram of the vision-language feature extraction model in the embodiment of the present invention;

[0060] Figure 3 It is the structural diagram of the semantic segmentation model in the embodiment of the present invention. Detailed implementation manners

[0061] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Usually, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but only represents the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.

[0062] Please refer to Figure 1 , an open-vocabulary semantic segmentation method for UAV aerial images with vision-language fusion provided in this embodiment, which includes:

[0063] S1. Use a UAV to collect unclassified aerial mixed images under different environmental conditions, and generate corresponding language description data; different environmental conditions refer to environmental conditions such as different weather, time, or / and scenes.

[0064] The S1 includes:

[0065] S1-1. Collect high-resolution initial unclassified aerial images using a drone platform; the high resolution is 1280×720 or higher; the initial unclassified aerial images are usually RGB images. During the collection process, various scenarios need to be covered, including different lighting conditions (sunny, cloudy, night), diverse weather (rainy, foggy), and complex environments (urban, rural, industrial areas, etc.), further ensuring the diversity and representativeness of the initial unclassified aerial images, which is beneficial to the training process of the visual language fusion segmentation model and improves the accuracy of semantic segmentation corresponding to aerial images under various environmental conditions.

[0066] S1-2. Preprocess the initial unclassified aerial images to obtain unclassified aerial images; the preprocessing includes filtering and noise reduction, normalization, size adjustment, and format conversion.

[0067] For filtering and noise reduction, Gaussian filtering or median filtering is adopted. When using Gaussian filtering for processing, set the convolution kernel size to smooth the image noise while retaining the corresponding image edge information; the convolution kernel size can be set to 3×3 or 5×5. When using median filtering for processing, set the corresponding window size to determine the filtering effect and avoid losing image details due to excessive smoothing.

[0068] The formula for normalization is:

[0069] ;

[0070] where represents the normalized unclassified aerial image, represents the pixel value of the initial unclassified aerial image, and represent the minimum and maximum pixel values of the initial unclassified aerial image respectively.

[0071] For size adjustment, bilinear interpolation or other resampling techniques are used to adjust the normalized unclassified aerial images to a unified size, obtaining the adjusted unclassified aerial images, ensuring that all images meet the input requirements of the visual language fusion segmentation model or GPT model. When performing size adjustment, it is necessary to pay attention to maintaining the aspect ratio of the image and, if necessary, use padding techniques to prevent deformation.

[0072] The format conversion is to the tensor format required by the visual language fusion segmentation model or GPT model; the input tensor form The corresponding formula is:

[0073] ;

[0074] where represents the unclassified aerial image, and respectively represent the height and width of the unclassified aerial images, represent the number of channels. The number of channels can be set to 3.

[0075] S1-3. Select the unclassified aerial images that meet the selection criteria, check and splice them to obtain an unclassified aerial hybrid image.

[0076] In order to utilize the rich information of images from different perspectives, unclassified aerial images from different shooting angles are mixed and generated, and the unclassified aerial images after mixed generation that meet the selection criteria are selected. The selection criteria are images with representativeness, different shooting angles and / or shooting environments, and these images should cover different parts of the same scene or similar scenes to ensure that each image provides unique spatial details.

[0077] Perform non-overlapping splicing on the unclassified aerial images after mixed generation that meet the selection criteria to ensure that the unique features of each image are retained, and obtain an unclassified aerial hybrid image. Non-overlapping splicing arranges and splices the images according to a predetermined rule to form a composite image, and the corresponding formula is:

[0078] ;

[0079] where represents the unclassified aerial hybrid image, represents the splicing operation, 、 、 respectively represent the first, second and th unclassified aerial images after mixed generation that meet the selection criteria.

[0080] S1-4. Use the GPT model to generate descriptive texts for each unclassified aerial hybrid image to obtain language description data, providing accurate semantic information for cross-modal feature fusion. Format each unclassified aerial hybrid image or its local area according to the requirements of the GPT model. The formatting process includes adjusting the image size, ensuring that the image channels and format match the model input, and making full preparations for the model input.

[0081] S2. Construct a visual-language fusion segmentation model; the visual-language fusion segmentation model includes a visual-language feature extraction model, a heterogeneous cross-modal graph fusion model and a semantic segmentation model.

[0082] S3. Input the unclassified aerial hybrid image and the language description data into the visual-language feature extraction model, and output multi-scale spatio-temporal visual features.

[0083] As Figure 2As shown, the vision - language feature extraction model includes a parallel vision feature extraction sub - model and a language feature extraction sub - model; the vision feature extraction sub - model includes a series - connected vision feature extraction module and a multi - scale spatio - temporal vision feature fusion module; the vision feature extraction module includes a parallel VIT global feature extraction module (Vision Transformer) and a local detail feature extraction module; the VIT global feature extraction module includes a Patch - Embedding layer and a transformer layer based on the self - attention mechanism; the multi - scale spatio - temporal vision feature fusion module includes a series - connected dynamic feature fusion layer and a deformable convolutional layer; the local detail feature extraction module uses the Mamba model; the language feature extraction sub - model includes a series - connected BERT layer and a CSA - cross - modal alignment layer.

[0084] Thus, the S3 includes:

[0085] S3 - 1: Obtain unclassified aerial mixed training images and corresponding language description training data;

[0086] S3 - 2: Input the unclassified aerial mixed training images into the Patch - Embedding layer to obtain unclassified aerial mixed training segmented images; the unclassified aerial mixed training images are segmented into several small blocks through Patch Embedding technology, and each small block is regarded as an independent Token.

[0087] S3 - 3: Input the unclassified aerial mixed training segmented images into the transformer layer based on the self - attention mechanism to obtain unclassified aerial mixed training global feature maps; capture the long - distance dependence relationships and global semantic connections between each patch through the self - attention mechanism, so as to extract the context information of the overall scene. For example, the VIT global feature extraction module can identify building clusters, large - area roads, and open areas in the aerial images, providing a solid foundation for modeling the overall scene.

[0088] The transformer layer based on the self - attention mechanism includes a corresponding first encoder and a first decoder, and the first encoder uses an encoder based on the self - attention mechanism. Thus, the S3 - 3 includes:

[0089] S3 - 3 - 1: Perform linear transformation on each Token and append the corresponding position information to obtain the corresponding Token vector representation;

[0090] S3 - 3 - 2: Input each Token vector representation into the encoder based on the self - attention mechanism, calculate the correlation between each Token vector representation and other Token vector representations to obtain the corresponding attention weights; based on the attention weights, determine the corresponding Token key vector representation;

[0091] In each layer of the encoder, the self-attention mechanism calculates the correlation between each Token vector representation and all other Token vector representations. Each patch "attends" to the information of all other patches in the image. Three vectors, namely query, key, and value, are generated for each Token vector representation. By calculating the dot product of the query and all key vectors and then performing Softmax normalization, the attention weights of each Token vector representation with respect to other Token vector representations are obtained. Then, these weights are used to perform weighted summation on the values corresponding to all Tokens, thereby generating a new representation. This fusion of global information can capture the relationships and context semantics between distant regions. For example, in an aerial image, the self-attention mechanism can associate the scattered information such as buildings and roads everywhere to construct a comprehensive scene understanding. Thus, after multiple layers of self-attention processing, the global key training features output by the encoder contain rich context information and long-range dependencies, which can provide more accurate feature support for subsequent segmentation or other visual tasks.

[0092] S3-3-3. Perform weighted summation on all Token key vector representations to generate global key training features;

[0093] S3-3-4. Input the global key training features into the first decoder to output an unclassified aerial mixed training global feature map.

[0094] S3-4. Input the unclassified aerial mixed training image into the Mamba model to obtain an unclassified aerial mixed training local feature map; The Mamba model is constructed based on the state space model (SSM) and is used for feature extraction of local regions of aerial images. It can effectively capture long-range spatial dependence relationships and efficiently model details such as edges and textures in the image, thereby compensating for the weakening problem of the VIT global feature extraction module in fine-grained target extraction and ensuring the fine-grained characterization of small targets and detail regions in complex scenes.

[0095] S3-5. Input the unclassified aerial mixed training local feature map and the unclassified aerial mixed training global feature map into the multi-scale spatio-temporal visual feature fusion module to obtain multi-scale spatio-temporal visual training features;

[0096] The above S3-5 includes:

[0097] S3-5-1. Input the unclassified aerial mixed training local feature map and the unclassified aerial mixed training global feature map into the dynamic feature fusion layer to obtain multi-scale spatio-temporal visual initial training features;

[0098] The dynamic feature fusion layer adopts an adaptive weighted fusion mechanism, which dynamically calculates weights to balance the contributions of the two (classless aerial mixed training local feature maps, classless aerial mixed training global feature maps). The corresponding formula is:

[0099] ;

[0100] ;

[0101] Among them, 、 respectively represent the classless aerial mixed training global feature map and the classless aerial mixed training local feature map, represents the activation function, represents the adaptive weight parameter, represents the adaptive weight matrix, represents the multi-scale spatio-temporal visual initial training feature, represents the concatenated feature map of the classless aerial mixed training global feature map and the classless aerial mixed training local feature map. The adaptive weight parameter ranges from (0, 1).

[0102] The multi-scale spatio-temporal visual features are obtained through the adaptive weighted fusion mechanism, which can dynamically adjust the weights according to the importance of different image regions. It relies more on the Mamba model in regions with richer local details, while emphasizing the VIT global feature extraction module in regions with clear global structures, so as to obtain a more balanced and accurate visual representation.

[0103] S3-5-2. Input the multi-scale spatio-temporal visual initial training feature into the deformable convolutional layer to obtain the multi-scale spatio-temporal visual training feature, which can further improve the ability to capture local structures; the deformable convolution allows the convolutional kernel to adaptively adjust the sampling position according to the image content, so as to better adapt to target deformation and complex occlusion situations; it can perform fine-grained modeling on the edges and local shapes in the image, improving the accuracy of the segmentation boundary and the expression ability of the overall fine-grained information.

[0104] The VIT global feature extraction module and the Mamba model are respectively used to extract global image information and local image details, and adaptive weighted fusion is used to achieve the dynamic balance between global and local features. The deformable convolution is used to strengthen the local structure, thus constituting a set of multi-scale spatio-temporal visual feature extraction schemes, which can not only ensure the accurate expression of the overall scene semantics, but also capture the local details in the image in detail.

[0105] S3-6. Input the language description training data into the BERT model to obtain the language training feature;

[0106] The S3-6 includes:

[0107] S3-6-1. Language-description-based training data Generate the corresponding continuous vector text representation , and the corresponding formula is:

[0108] ;

[0109] Among them, represents the BERT model, represents the word segmentation operation.

[0110] S3-6-2. Based on the unlabeled aerial photography mixed training local feature map and the unlabeled aerial photography mixed training global feature map, set a group of visual prototypes (Prototype); in order to perform cross-modal alignment with the multi-scale spatio-temporal visual training features, a group of visual prototypes need to be preset. These prototypes usually come from the regional features (unlabeled aerial photography mixed training local / global feature maps) extracted by the visual feature extraction module, and the continuous vector text representation is in the same latent space as the visual prototype, but it is difficult to directly align. Therefore, it is necessary to further optimize the text features so that they can be more closely aligned with the positive sample visual features.

[0111] When the visual feature extraction sub-model extracts regional features, it actually combines the advantages of the VIT global feature extraction module and the local detail feature extraction module (Mamba model): the VIT global feature extraction module extracts global information through the self-attention mechanism, while the Mamba model focuses on capturing local details. Therefore, the corresponding regional features are used as visual prototypes, and then aligned with the multi-scale spatio-temporal visual training features to assist in the further optimization of the text features (continuous vector text representation), making them more closely correspond to the positive sample visual features. Finally, cross-modal alignment is achieved through the contrastive semantic alignment (CSA) loss function.

[0112] S3-6-3. Calculate the CSA loss function based on the positive / negative similarity of the continuous vector text representation;

[0113] To ensure the accurate alignment of the continuous vector text representation and the visual prototype in the same latent space, a contrastive semantic alignment (CSA) pre-training task is designed. The goal of this task is to make the positive similarity between the continuous vector text representation and its corresponding positive sample visual features as high as possible, while the negative similarity between the continuous vector text representation and the negative sample visual features is as low as possible. The positive sample visual features and the negative sample visual features respectively represent the multi-scale spatio-temporal visual training features that match and do not match the continuous vector text representation.

[0114] Thus, the CSA loss function has the corresponding formula as follows:

[0115] ;

[0116] wherein, represents the similarity metric function, represents the exponential function with the base of the constant e, represents the logarithmic function with a constant base, represents the summation function, represents the temperature parameter. The similarity metric function can adopt the cosine similarity. The temperature parameter is used to adjust the smoothness of the distribution.

[0117] The CSA loss function forces the text encoder (BERT model) to learn more discriminative semantic representations, enabling the generated continuous vector text representations to better reflect the key information in the unclassified aerial images and achieve precise alignment with the corresponding multi-scale spatio-temporal visual training features.

[0118] S3-6-4. Based on the CSA loss function, match and align the visual prototype and the continuous vector text representation to obtain the matching result;

[0119] S3-6-5. Based on the matching result and the continuous vector text representation , generate the corresponding language training features.

[0120] Driven by the CSA pre-training task, the text encoder continuously updates its parameters, thereby shortening the distance between the matching pair of continuous vector text representations and the positive sample visual features in the latent space, while lengthening the distance between the continuous vector text representation and all negative sample visual features . Through this contrastive learning mechanism, the language training features not only retain the natural expression ability of the language but also embed richer cross-modal semantic information, improving the matching accuracy with the multi-scale spatio-temporal visual training features. The feature representation output by the text encoder will have stronger discriminability, providing a solid semantic foundation for subsequent cross-modal graph fusion and segmentation tasks.

[0121] S3-7. Based on the multi-scale spatio-temporal visual training features and the language training features, adjust the weight parameters of the visual feature extraction sub-model.

[0122] S4. Input the language features and the multi-scale spatio-temporal visual features into the heterogeneous cross-modal graph fusion model, and output the visual-language matching features.

[0123] The heterogeneous cross-modal graph fusion model adopts a heterogeneous cross-modal graph network based on the multi-head attention mechanism.

[0124] The training process of the heterogeneous cross-modal graph fusion model includes:

[0125] S4-1. Construct an initial heterogeneous cross-modal graph; obtain language training features and multi-scale spatio-temporal visual training features, and use them as text nodes and visual nodes respectively, and map them into the initial heterogeneous cross-modal graph;

[0126] S4-2. Calculate the node similarity between text nodes and visual nodes; set the weights of edges and knowledge graphs based on each node similarity, and construct corresponding edges to obtain a heterogeneous cross-modal graph; the edges include visual-text edges and text-knowledge edges; the node similarity can adopt cosine similarity. Cosine similarity is used to measure the angular similarity between two vectors. The value range of cosine similarity is [-1, 1], and the larger the value, the more similar the two vectors are. Between text nodes and visual nodes, cosine similarity can be used to calculate their similarity in the feature space, usually by normalizing the text feature and visual feature vectors and then calculating. The corresponding formula is:

[0127] ;

[0128] Among them, 、 represent text nodes and visual nodes respectively, represents the text node and the visual node the cosine similarity between them, represents the Euclidean norm.

[0129] If the "damaged house" in the language feature and the multi-scale spatio-temporal visual feature match highly, a higher-weighted edge is established between the visual node and the corresponding text node. Secondly, construct text-knowledge edges to connect the text nodes with relevant nodes in the external domain knowledge graph, such as the "farmland" node and the "crop type" node. Through this connection, external professional knowledge is introduced to enhance the semantic expression in the graph. This step ensures that the graph not only contains direct visual-text correspondence relationships but also can incorporate domain knowledge to expand the depth of cross-modal information.

[0130] S4-3. According to the multi-head graph attention mechanism (GAT), obtain the aggregated information of each node and its adjacent nodes;

[0131] S4-4. Based on each aggregated information, update the representation of each node through multi-hop aggregation of the multi-head graph attention mechanism; the representation of the node is visual-language matching training features, which include language features and multi-scale spatio-temporal visual features; the update formula is:

[0132] ;

[0133] ;

[0134] Among them, represents concatenating the outputs of all graph attention heads, represents the number of multi-head graph attention, , , , respectively represent the output data corresponding to the first graph attention head, the second graph attention head, the th graph attention head, and the last graph attention head, represents the set of neighbor nodes of node , represents the th graph attention head's normalized attention weight of node to neighbor node , represents the learnable weight matrix of the th graph attention head, , respectively represent the representation of node at the th layer, and the representation of neighbor node at the th layer.

[0135] Through this multi-head attention mechanism, each node can capture cross-modal semantic information from different angles and scales, ensuring the comprehensiveness and depth of information aggregation. Through the multi-hop aggregation of multi-head graph attention, it is ensured that each node not only considers the information of direct neighbors during update, but also can integrate cross-modal semantic relationships at greater distances. This aggregation mechanism enables the representations of each node in the graph to continuously fuse multi-dimensional information from vision, text, and domain knowledge during the layer-by-layer propagation process, forming richer and more discriminative node representations. Through the multi-layer stacking of graph neural networks, the entire heterogeneous graph can fully capture and transmit cross-modal fine-grained semantic information, providing strong semantic support for subsequent segmentation tasks.

[0136] S4-5. Optimize the heterogeneous cross-modal graph based on the updated node representations, that is, optimize the heterogeneous cross-modal graph fusion model.

[0137] By constructing visual nodes, text nodes, visual-text edge relationships, and text-knowledge edge relationships, introducing a multi-head graph attention mechanism to achieve weighted aggregation of node information, and continuously optimizing the graph structure through multi-hop propagation, an efficient heterogeneous cross-modal graph is constructed, enabling fine-grained interaction between visual features and text features in the same graph structure and providing rich cross-modal semantic expressions for subsequent tasks.

[0138] S5. Input the visual-language matching features into the semantic segmentation model to complete the semantic segmentation of the aerial images.

[0139] As Figure 3 shown, the semantic segmentation model uses a lightweight U-Net++ model; the lightweight U-Net++ model includes an encoder connected in series and a decoder based on channel-spatial dual attention; multi-level skip connections are constructed between the encoder and the decoder to ensure that fine-grained image information can be fully retained during the information transmission process. The encoder includes N decoding layers; the decoder based on channel-spatial dual attention includes N decoding modules, and every two decoding modules are connected through a multi-level fusion module; each decoding module includes a CSDA layer (channel-spatial dual attention layer) and a decoding layer connected in series.

[0140] Thus, the S5 includes:

[0141] S5-1. Obtain the visual-language matching training features and input them into the encoder to obtain the i-th visual-language matching training coding feature corresponding to each coding layer; the input of the first coding layer is the visual-language matching training features, and the input of the remaining coding layers is the output of the previous coding layer, that is, the input of the i-th coding layer is the output of the i-1-th coding layer.

[0142] S5-2. Input the N-th visual-language matching training feature into the N-th decoding module and output the N-th visual-language matching decoding result. The corresponding process is:

[0143] S5-2-1. Input the Nth visual-language matching training feature into the Nth layer of the CSDA layer to obtain the corresponding Nth visual-language matching training key feature. The lightweight U-Net++ model embeds a channel-spatial dual attention module (CSDA). The CSDA module further enhances the discriminative ability of feature expression by introducing attention mechanisms in both the channel and spatial dimensions respectively. The channel attention module can adaptively assign weights to different feature channels, thereby emphasizing the semantic information that is more critical for the segmentation task. The spatial attention module focuses on important regions in the spatial dimension to ensure that regions with rich details (such as small targets or blurred edges) are fully expressed. Through the synergistic effect of these two attention mechanisms, the CSDA module significantly improves the sensitivity and expression ability of the segmentation head to fine-grained information, providing a more robust feature basis for subsequent fine segmentation.

[0144] S5-2-2. Input the Nth visual-language matching training key feature into the Nth layer of the decoding layer to obtain the Nth visual-language matching decoding result.

[0145] The processing process of each decoding module is the same as that of the Nth layer decoding module.

[0146] S5-3. Input the Nth visual-language matching decoding result and the (N - 1)th visual-language matching training feature into the multi-level fusion module, and output the (N - 1)th visual-language matching training fusion feature. By introducing the multi-level fusion module, the lightweight U-Net++ model enables the features from the low layer to the high layer to be gradually integrated, while strengthening the capture of small target and complex edge information, enabling the segmentation head to obtain accurate feature expressions at different scales and providing a solid foundation for the final segmentation task.

[0147] S5-4. Input the (N - 1)th visual-language matching training fusion feature into the (N - 1)th layer decoding module, and output the (N - 1)th visual-language matching decoding result.

[0148] S5-5. Repeat the processing processes of the multi-level fusion module and the (N - 1)th layer decoding module until the output of the first layer decoding module is obtained, that is, the semantic segmentation training result is obtained.

[0149] S5-6. Based on the semantic segmentation training result, calculate the dynamic weight triplet loss function and the generalized intersection over union loss function.

[0150] The dynamic weight triplet loss function dynamically calculates the triplet dynamic weight parameters through the uncertainty estimation mechanism , thus, the formula corresponding to the dynamic weight triplet loss function is:

[0151] ;

[0152] Among them, and respectively represent the three - element dynamic weight parameter and the predetermined hyperparameter, represents the desired operation, represents the probability distribution predicted by the semantic segmentation model for each category in the visual - language matching training feature and represents the maximum value function. When the uncertainty of the prediction of a certain sample (visual - language matching training feature) by the semantic segmentation model is relatively high (low confidence), the three - element dynamic weight parameter is dynamically adjusted through the dynamic weight triple loss function, so that the semantic segmentation model can pay more attention to unknown categories during the training process, thereby achieving balanced optimization of known and unknown categories.

[0153] When the semantic segmentation model has a relatively high prediction uncertainty (low confidence) for a certain sample (visual - language matching training feature), the three - element dynamic weight parameter is dynamically adjusted through the dynamic weight triple loss function, enabling the semantic segmentation model to pay more attention to unknown categories during the training process, thus realizing the balanced optimization of known and unknown categories.

[0154] The generalized intersection over union (GIoU) loss function is adopted to replace the traditional cross - entropy loss function. The traditional cross - entropy loss function may have the problem of insufficient fitting when dealing with segmentation boundaries, while the GIoU loss function further improves the boundary accuracy and overall robustness of the segmentation result by comprehensively considering the overlap, difference, and enclosing situation between the predicted segmentation region and the real region, and can more finely regulate the prediction performance of the semantic segmentation model at each pixel level to ensure that the segmentation map can better fit the real boundary.

[0155] S5 - 7. Based on the dynamic weight triple loss function and the generalized intersection over union loss function, adjust the weight parameters of the semantic segmentation model.

[0156] The U - Net++ model is improved based on the channel - spatial dual - attention mechanism and the multi - level fusion module to ensure that the fused multi - modal features can be fully utilized; the dynamic weight triple loss function is introduced, and the loss weight is dynamically adjusted through uncertainty estimation to solve the problem of unbalanced optimization between known and unknown categories; the GIoU loss is used to further improve the fitting accuracy of the segmentation boundary and the overall robustness; through the efficient design of the segmentation head and the fine regulation of the adaptive loss function, high - precision segmentation of multi - modal fusion features is achieved.

[0157] The visual - language fusion segmentation model also includes an optimization process, and the corresponding steps are:

[0158] The initial optimized vision-language fusion segmentation model is processed using the NAS algorithm to obtain a once-optimized vision-language fusion segmentation model; perform neural architecture search (NAS) on the heterogeneous cross-modal graph fusion model. Through an automated search algorithm, find the optimal network structure in the predefined architecture space, and eliminate redundant structures and parameters, thereby obtaining a heterogeneous cross-modal graph fusion model that not only meets the accuracy requirements but also has a lower computational complexity. The NAS algorithm can evaluate the performance of different network configurations under specific tasks and finally select those structures that can operate efficiently on edge devices to achieve the goal of overall lightweighting.

[0159] Perform parameter pruning and refinement on the once-optimized vision-language fusion segmentation model to obtain a twice-optimized vision-language fusion segmentation model; perform parameter pruning on the vision-language fusion segmentation model. By analyzing the weight contributions of each module in the vision-language fusion segmentation model, eliminate those parameters that have little impact on the output or are redundant. The pruning operation can not only reduce the number of model parameters but also reduce memory occupancy and computational requirements. After fine pruning and parameter refinement, the model will be more compact, facilitating deployment on resource-constrained edge devices while maintaining a high prediction accuracy.

[0160] Process the twice-optimized vision-language fusion segmentation model using the quantization-aware training method to obtain the trained vision-language fusion segmentation model. Adopt the quantization-aware training (QAT) method to achieve 8-bit quantization deployment of the model. During the training process, the QAT method will simulate a low-precision computing environment, enabling the model to gradually adapt to 8-bit integer operations. After quantization training, the weights and activation functions of the model are represented in low precision, thus significantly reducing memory requirements and computational resource consumption. This quantization not only reduces the model size but also significantly improves the inference speed while maintaining the model accuracy, with an expected acceleration effect of about 3 times.

[0161] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. An open-vocabulary semantic segmentation method for UAV aerial images with visual language fusion, characterized in that, Including: Using a drone to collect unclassified aerial mixed images under different environmental conditions and generating corresponding language description data; Constructing a vision-language fusion segmentation model; The vision-language fusion segmentation model includes a vision-language feature extraction model, a heterogeneous cross-modal graph fusion model, and a semantic segmentation model; Inputting the unclassified aerial mixed images and language description data into the vision-language feature extraction model to output multi-scale spatio-temporal vision features and language features; Inputting the language features and multi-scale spatio-temporal vision features into the heterogeneous cross-modal graph fusion model to output vision-language matching features; Inputting the vision-language matching features into the semantic segmentation model to complete the semantic segmentation of the aerial images; The semantic segmentation model uses a lightweight U-Net++ model; the training process of the semantic segmentation model includes: Obtaining vision-language matching training features and inputting them into the lightweight U-Net++ model to output semantic segmentation training results; based on the semantic segmentation training results, calculating a dynamic weight triplet loss function and a generalized intersection over union loss function; based on the dynamic weight triplet loss function and the generalized intersection over union loss function, adjusting the weight parameters of the semantic segmentation model; The formula corresponding to the dynamic weight triplet loss function is: ; Among them, and respectively represent the ternary dynamic weight parameter and the predetermined hyperparameter, represents the desired operation, represents the probability distribution predicted by the semantic segmentation model for each category in the visual - language matching training feature among, and represents the maximum value function.

2. The method for open-vocabulary semantic segmentation of UAV aerial images with visual language fusion according to claim 1, wherein The using a drone to collect unclassified aerial mixed images under different environmental conditions and generating corresponding language description data includes: Using a drone platform to collect initial unclassified aerial images and preprocessing them to obtain corresponding unclassified aerial images; Selecting unclassified aerial images that meet the selection conditions and performing inspection and stitching to obtain unclassified aerial mixed images; the selection conditions are different shooting angles or shooting environments; Using a GPT model to generate description texts for each unclassified aerial mixed image to obtain language description data.

3. The open-vocabulary semantic segmentation method for UAV aerial images with visual-language fusion according to claim 1, characterized in that, The vision-language feature extraction model includes a parallel vision feature extraction sub-model and a language feature extraction sub-model; The vision feature extraction sub-model includes a series-connected vision feature extraction module and a multi-scale spatio-temporal vision feature fusion module; the vision feature extraction module includes a parallel VIT global feature extraction module and a local detail feature extraction module; the VIT global feature extraction module includes a Patch-Embedding layer and a transformer layer based on a self-attention mechanism; the multi-scale spatio-temporal vision feature fusion module includes a series-connected dynamic feature fusion layer and a deformable convolutional layer; the local detail feature extraction module uses a Mamba model; the language feature extraction sub-model is a BERT model; The heterogeneous cross-modal graph fusion model uses a heterogeneous cross-modal graph network based on a multi-head attention mechanism; The lightweight U-Net++ model includes a series-connected encoder and a decoder based on channel-spatial dual attention; the encoder includes N decoding layers; the decoder based on channel-spatial dual attention includes N decoding modules, and every two decoding modules are connected through a multi-level fusion module; each decoding module includes a series-connected CSDA layer and a decoding layer.

4. The method for open-vocabulary semantic segmentation of UAV aerial images with visual language fusion according to claim 3, characterized in that The training process of the vision-language feature extraction model includes: Obtaining unclassified aerial mixed training images and corresponding language description training data; Input the unclassified aerial mixed training images into the Patch-Embedding layer to obtain unclassified aerial mixed training segmentation images; Input the unclassified aerial mixed training segmentation images into the transformer layer based on the self-attention mechanism to obtain unclassified aerial mixed training global feature maps; Input the unclassified aerial mixed training images into the Mamba model to obtain unclassified aerial mixed training local feature maps; Input the unclassified aerial mixed training local feature maps and the unclassified aerial mixed training global feature maps into the multi-scale spatio-temporal visual feature fusion module to obtain multi-scale spatio-temporal visual training features; Input the language description training data into the BERT model to obtain language training features; Based on the multi-scale spatio-temporal visual training features and the language training features, adjust the weight parameters of the visual feature extraction sub-model.

5. The method for open-vocabulary semantic segmentation of UAV aerial images with visual-language fusion according to claim 4, characterized in that, The step of inputting the language description training data into the BERT model to obtain language training features includes: Generate a corresponding continuous vector text representation based on the language description training data; Based on the unclassified aerial mixed training local feature maps and the unclassified aerial mixed training all feature maps, set a group of visual prototypes; Based on the positive / negative similarity of continuous vector text representations, calculate the CSA loss function; the CSA loss function The corresponding formula is: ; Among them, represents a similarity metric function, represents an exponential function with the constant e as the base, represents a logarithmic function with a constant as the base, represents a summation function, represents a temperature parameter, represents a visual prototype and a continuous vector text representation, 、 represent positive sample visual features and sample visual features respectively; Based on the CSA loss function, match and align the visual prototypes and the continuous vector text representation to obtain a matching result; Based on the matching result and the continuous vector text representation, generate corresponding language training features.

6. The method for open-vocabulary semantic segmentation of UAV aerial images with visual language fusion according to claim 3, wherein The training process of the heterogeneous cross-modal graph fusion model includes: Construct an initial heterogeneous cross-modal graph; obtain the language training features and the multi-scale spatio-temporal visual training features and use them as text nodes and visual nodes respectively, and map them into the initial heterogeneous cross-modal graph; Calculate the node similarity between the text nodes and the visual nodes; set the weights of the edges based on each node similarity, construct the corresponding edges and update the external domain knowledge graph to obtain a heterogeneous cross-modal graph; According to the multi-head graph attention mechanism, obtain the aggregation information of each node and its adjacent nodes; Based on each aggregation information, update the representation of each node through multi-hop aggregation of the multi-head graph attention mechanism; the representation of the node is visual-language matching training features, which include language features and multi-scale spatio-temporal visual features; Based on the updated representation of the node, optimize the heterogeneous cross-modal graph, that is, optimize the heterogeneous cross-modal graph fusion model.

7. The method for open-vocabulary semantic segmentation of UAV aerial images with visual language fusion according to claim 3, characterized in that The training process of the semantic segmentation model includes: Obtain the visual-language matching training features and input them into the encoder to obtain the corresponding i-th visual-language matching training coding features of each coding layer; the input of the first coding layer is the visual-language matching training features, and the input of the i-th coding layer is the output of the (i - 1)-th coding layer; Input the N-th visual-language matching training features into the N-th decoding module, and output the N-th visual-language matching decoding result; Input the N-th visual-language matching decoding result and the (N - 1)-th visual-language matching training features into the multi-level fusion module, and output the (N - 1)-th visual-language matching training fusion features; Input the (N - 1)-th visual-language matching training fusion features into the (N - 1)-th decoding module, and output the (N - 1)-th visual-language matching decoding result; Repeat the processing of the repeated multi-level fusion module and the (N-1)th layer decoding module until the output of the first layer decoding module is obtained, that is, the semantic segmentation training result is obtained; Based on the semantic segmentation training result, calculate the dynamic weight triplet loss function and the generalized intersection over union loss function; Based on the dynamic weight triplet loss function and the generalized intersection over union loss function, adjust the weight parameters of the semantic segmentation model.

8. The open-vocabulary semantic segmentation method for UAV aerial images with visual-language fusion according to claim 1, characterized in that The visual language fusion segmentation model further includes an optimization process, and the corresponding steps are: Process the initially optimized visual language fusion segmentation model using the NAS algorithm to obtain the once-optimized visual language fusion segmentation model; Perform parameter pruning and simplification on the once-optimized visual language fusion segmentation model to obtain the twice-optimized visual language fusion segmentation model; Process the twice-optimized visual language fusion segmentation model using the quantization-aware training method to obtain the optimized visual language fusion segmentation model.