Remote sensing image building semantic segmentation system based on visual language model
By constructing a diverse query text set and an adaptive cross-modal attention mechanism, the initial building mask is generated, and the pseudo-label quality is optimized through pseudo-label screening and model iterative training, the problem that natural scene image pre-trained VLMs cannot be directly applied to remote sensing images, improving the building segmentation accuracy of remote sensing images, and reducing manual annotation dependence.
Patent Information
- Application Number
- CN202510873921.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-06-27
AI Technical Summary
The visual language model pre-trained by natural scene images cannot be directly applied to the building segmentation task of remote sensing images, resulting in insufficient segmentation accuracy and relying on manual annotation.
A diverse set of query texts is constructed, text features are generated using the alignment prompt encoder, and visual features are fused through an adaptive cross-modal attention mechanism to generate initial building masks, combining pseudo-label screening, model iterative training and dynamic parameter adjustment to optimize pseudo-label quality.
It improves the segmentation accuracy of buildings in remote sensing images, reduces the dependence on manual annotation, and enhances the model's adaptability in remote sensing scenarios.
Smart Images

Figure CN120388178A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a remote sensing image building semantic segmentation system based on a vision - language model. Background Art
[0002] Building semantic segmentation, as a key task in the analysis of remote sensing images in the field of computer vision, is committed to accurately classifying each pixel in a remote sensing image into building or non - building categories, laying a solid foundation for subsequent geographical information analysis, and having indispensable applications in many fields such as urban planning, disaster monitoring, and environmental assessment. In urban planning, accurate building segmentation results can help planners clarify the urban spatial layout and reasonably plan land use and infrastructure construction; in the disaster monitoring scenario, it can quickly identify the damage status of buildings in the affected area and provide key information for rescue operations; in environmental assessment, it helps analyze the impact of building distribution on the ecological environment.
[0003] Traditional building semantic segmentation methods highly rely on large - scale manually annotated data. However, manual annotation not only requires a large amount of human, material, and time costs, but also the annotation process is easily interfered by subjective factors, resulting in inconsistent annotation results. At the same time, in the face of complex and variable real - world scenarios, such as different architectural styles, scale differences, complex backgrounds, and occlusion situations, the segmentation accuracy of traditional methods is insufficient.
[0004] With the breakthrough of vision - language models (VLMs) in cross - modal semantic alignment and zero - shot inference capabilities, they have shown great potential in image semantic understanding. Currently, several high - performance VLMs have been developed, and these VLMs are usually pre - trained using natural scene images. Due to the significant domain differences between natural images and remote sensing images (such as imaging perspective, resolution, and distribution of ground object features), VLMs pre - trained with natural scene images cannot be directly applied to the building segmentation task of remote sensing images. Summary of the Invention
[0005] In view of this, the present invention provides a remote sensing image building semantic segmentation system based on a vision - language model to solve the problem that VLMs pre - trained with natural scene images cannot be directly applied to the building segmentation task of remote sensing images, improve the segmentation accuracy of the model for buildings in remote sensing images, and reduce the dependence on manual annotation of remote sensing images.
[0006] A remote sensing image building semantic segmentation system based on a vision-language model, comprising an initial building mask generation module and a pseudo-label optimization and model iterative enhancement module; the initial building mask generation module includes a query text set construction unit, a feature extraction unit, a cross-modal attention interaction unit, an initial building instance generation unit, and an instance fusion and mask generation unit; the pseudo-label optimization and model iterative enhancement module includes a pseudo-label screening unit, a model iterative training unit, and a dynamic parameter adjustment unit; The query text set construction unit is used to construct a diverse query text set composed of multiple prompt words; The feature extraction unit is used to extract the text features of the query text set using an aligned prompt encoder, and to extract the multi-scale visual features of the remote sensing image; The cross-modal attention interaction unit is used to interact the text features and the multi-scale visual features by adopting an adaptive cross-modal attention mechanism to obtain visual features integrating text information; The initial building instance generation unit is used to perform threshold processing, connected region analysis, and overlap removal on the visual features integrating text information to generate initial building instances; The instance fusion and mask generation unit is used to fuse the initial building instances obtained based on different query texts to obtain an initial building mask, and to obtain an initial building segmentation pseudo-label based on the initial building mask; The pseudo-label screening unit is used to screen out reliable pseudo-labels from the initial building segmentation pseudo-labels based on uncertainty quantification, cross-validation consistency, and geometric constraint filtering; The model iterative training unit is used to gradually enhance the generalization ability of the system by adopting a two-stage training method of weak supervision training and fine-tuning optimization; The dynamic parameter adjustment unit is used to perform dynamic parameter adjustment by adopting threshold adaptive update and query text weight reallocation.
[0007] According to the remote sensing image building semantic segmentation system based on the vision-language model provided by the present invention, a diverse query text set is constructed. The text encoder of the Aligning and Prompting Everything all at once for Universal (APE) is used to generate text features. At the same time, the multi-scale features of the dual-temporal remote sensing images are extracted through the vision encoder. Then, an adaptive cross-modal attention mechanism is adopted to deeply fuse the text features and the vision features, enabling precise alignment of the features. Threshold processing, connected region analysis, and overlapping removal are performed on the fused features to generate initial building instances. The instance confidence of different query texts is fused through weighted summation to obtain an initial building mask, which is then binary-processed to obtain an initial building segmentation pseudo-label. Then, through pseudo-label screening, model iterative training, and dynamic parameter adjustment, the reliability of the pseudo-label and the adaptability of the system to the remote sensing scene are gradually improved. It can effectively utilize the initial pseudo-label and reduce the influence of noise, improving the segmentation accuracy of buildings in remote sensing images, solving the problem that the building labels generated by pre-training vision-language models (VLMs) for natural scene images cannot be directly applied to remote sensing images, and reducing the dependence on manual annotation of remote sensing images. Brief Description of the Drawings
[0008] Figure 1 It is a structural block diagram of the remote sensing image building semantic segmentation system based on the vision-language model provided by the embodiment of the present invention. Detailed Embodiment
[0009] The embodiments of the present invention are described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the drawings are exemplary and are intended to explain the embodiments of the present invention and should not be construed as a limitation to the present invention.
[0010] Please refer to Figure 1 , the embodiment of the present invention provides a remote sensing image building semantic segmentation system based on the vision-language model, including an initial building mask generation module and a pseudo-label optimization and model iterative enhancement module; the initial building mask generation module includes a query text set construction unit, a feature extraction unit, a cross-modal attention interaction unit, an initial building instance generation unit, and an instance fusion and mask generation unit; the pseudo-label optimization and model iterative enhancement module includes a pseudo-label screening unit, a model iterative training unit, and a dynamic parameter adjustment unit.
[0011] Among them, in order to generate the initial building mask more accurately, the initial building mask generation module processes the th remote sensing image from the training set ( ) Perform in-depth processing to generate a building mask.
[0012] The query text set construction unit is used to construct a diverse query text set composed of multiple prompt words.
[0013] Among them, the present application first constructs a diverse query text set composed of k prompt words , where , , are the 1st, 2nd, and rd query texts respectively. The prompt words can be words or phrases, such as various building-related words and phrases like "building", "residential building", "commercial building", "skyscraper", "villa", etc. Compared with using only a single word (such as "house" or "building", etc.) as a prompt word for query in the prior art, the diverse query text set of the present application can make full use of the powerful text understanding ability of the vision foundation model to capture the building features in the image from different perspectives. Different query texts can cover buildings of different types, scales, and uses, thereby improving the recognition ability of various buildings.
[0014] The feature extraction unit is used to extract the text features of the query text set using the aligned prompt encoder, and extract the multi-scale visual features of the remote sensing image.
[0015] Among them, the feature extraction unit is specifically used to input each query text in the query text set into the text encoder in the aligned prompt encoder, and the text encoder converts the query text into the corresponding text features. For the rd query text , the present application inputs it into the text encoder of the APE. The text encoder will convert the query text into the corresponding text features , and the text features contain the semantic information of the query text.
[0016] The feature extraction unit is also used to process the remote sensing image using the visual encoder in the aligned prompt encoder to extract multi-scale visual features. For the <00> nd remote sensing image , extract multi-scale visual features , , represents the number of visual feature scales extracted, , 、 They are the feature maps of the first scale, the feature maps of the second scale, and the feature maps of the Feature maps of different scales.
[0017] The cross-modal attention interaction unit is used to interact text features and multi-scale visual features using an adaptive cross-modal attention mechanism to obtain visual features that integrate text information.
[0018] In order to better adapt to the matching of different query texts and image features in the query text set, this application adopts an adaptive cross-modal attention mechanism, which includes four main processes: similarity calculation, attention weight calculation, adaptive adjustment of attention matrix weights and feature fusion.
[0019] Among them, the cross-modal attention interaction unit is specifically used for: Based on the cosine similarity, the similarity between the text features of each query text in the query text set and the feature maps of each scale in the multi-scale visual features is calculated. Remote sensing images No. Feature maps of different scales ,Will Flattened into a sequence of vectors , then calculate the Text features and The dot product of each element in gets the similarity score matrix ; Then use the softmax function to normalize the similarity score matrix to obtain the normalized attention weight matrix, which is expressed as:
[0020] in, express The elements, express The elements, is the sequence length of the similarity score matrix, Represents the normalized attention weight matrix The elements; In order to further enhance the adaptive capability, this application then introduces the adaptive adjustment factor , this adaptive adjustment factor is dynamically adjusted according to the recognition accuracy of the query text for the feature map at this scale in historical experiments. If a certain query text shows a high recognition accuracy on the feature map at a certain scale, the corresponding adaptive adjustment factor will be larger; otherwise, it will be smaller.
[0021] By adjusting the normalized attention weight matrix, the expression is:
[0022] Among them, represents the element at the th position in the final attention weight matrix ; Finally, use to perform weighted summation to obtain the visual feature that fuses text information , and the expression is:
[0023] Among them, represents the element at the th position in
[0024] Through this adaptive cross-modal attention mechanism, the system can automatically focus on the image regions that are more semantically relevant to the query text, so as to more accurately capture the features of the building.
[0025] The initial building instance generation unit is used to perform threshold processing, connected component analysis, and overlapping removal on the visual feature that fuses text information to generate initial building instances.
[0026] Among them, the initial building instance generation unit is specifically used for: Perform threshold processing on , set the threshold , and use the elements in that are greater than as the parts belonging to the building, and use the elements in that are less than or equal to as the background, so as to obtain a binary feature map ; Use a connected component analysis algorithm (such as the flood fill algorithm) to process the binary feature map , and this algorithm will find out the set of foreground pixels that are connected to each other in To ensure that building instances do not overlap, post - process the detected connected regions. When two or more connected regions overlap, screen them according to the confidence of the connected regions (such as the sum of pixel values in the region, average similarity score, etc.), retain the connected region with the highest confidence, and remove other connected regions.
[0027] After the above - mentioned processing, a set of non - overlapping initial building instances can be obtained.
[0028] The instance fusion and mask generation unit is used to fuse the initial building instances obtained based on different query texts to obtain an initial building mask, and obtain an initial building segmentation pseudo - label based on the initial building mask.
[0029] Among them, the way the instance fusion and mask generation unit fuses the initial building instances obtained based on different query texts is: perform a weighted sum of the instance confidences at each pixel position, and the weights are dynamically allocated according to the effectiveness of different query texts.
[0030] Specifically, this application will maintain a historical accuracy record, recording the recognition accuracy of each query text for buildings in previous training or testing. For the th query text weight , the calculation formula is:
[0031] Among them, is the total number of query texts in the query text set, is the historical accuracy of the th query text , is the historical accuracy of the th query text ; For the pixel position , is the abscissa of the pixel, is the ordinate of the pixel, and perform a weighted sum of the building instance confidences obtained based on the query text to obtain the fused confidence , and the expression is:
[0032] After fusion, this application obtains a more accurate and comprehensive set of building instances, and then combines the building instances in the set of building instances into an initial building mask , , is a set of real numbers, and are respectively the height and width of. is a pixel-by-pixel probability map, where the pixel position corresponding value , reflects the image in the pixel position the possibility of belonging to a building.
[0033] Due to the domain difference between the pre-trained data (natural images) of APE and remote sensing images, the confidence score of the system for buildings is usually conservatively calibrated. To address this domain difference problem, this application uses a loose binarization threshold (e.g., taking 0.25) to binarize the probability map to obtain an initial building segmentation pseudo-label, and the expression is:
[0034] where, is the value corresponding to the pixel position in the initial building mask ; is the value corresponding to the pixel position in the initial building segmentation pseudo-label ; is the binarization threshold.
[0035] This binarization threshold achieves a balance between recall and precision, while accepting controllable noise, retains areas with medium building features (e.g., roofs with unclear partial structures or spectral features). The obtained as the initial building segmentation pseudo-label, will subsequently be further optimized through uncertainty quantification to filter out false activation areas.
[0036] After obtaining the initial building segmentation pseudo-label , to further improve the building semantic segmentation accuracy of the system on the optical remote sensing image dataset, this application constructs a self-training optimization framework through three core links: pseudo-label screening, model iterative training, and dynamic parameter adjustment.
[0037] Since the initial pseudo-labels are affected by domain differences and model uncertainties, reliable pseudo-labels need to be screened out through multi-dimensional evaluation. The pseudo-label screening unit is used to filter out reliable pseudo-labels from the initial building segmentation pseudo-labels based on uncertainty quantification, cross-validation consistency, and geometric constraints.
[0038] Among them, the process of the pseudo-label screening unit screening based on uncertainty quantification is: Calculate the pixel-level entropy value of the initial building mask. The calculation formula is as follows:
[0039] where is the pixel position corresponding to the pixel-level entropy value; For the pixel regions with pixel-level entropy values higher than the entropy threshold lower the pseudo-label confidence or mark them as to be corrected; The process of filtering by the pseudo-label filtering unit based on cross-validation consistency is as follows: Use the alignment prompt encoder with different initializations to generate multiple sets of pseudo-labels for the same remote sensing image, calculate the pixel-level intersection over union (IoU). If the average pixel-level IoU of multiple sets of pseudo-labels for the target region is lower than the IoU threshold , then remove the pseudo-labels of this target region from the training set. The process of filtering by the pseudo-label filtering unit based on geometric constraint filtering is as follows: Based on the geometric priors of buildings in the remote sensing image (such as minimum area, aspect ratio range, etc.), remove the pseudo-label regions that do not conform to the preset shape features. For example, connected regions with an area less than 10 pixels or an aspect ratio exceeding 10:1 will be filtered.
[0040] The model iterative training unit is used to gradually enhance the generalization ability of the system by adopting a two-stage training method of weak supervision training and fine-tuning optimization.
[0041] In weak supervision training, merge the filtered high-quality pseudo-labels with a small amount of labeled data as the pre-training dataset . During training, the expression of the weighted cross-entropy loss function is as follows:
[0042] where is the number of samples in the pre-training dataset , is the pseudo-label quality weight, is the cross-entropy function, is the true label (pseudo-labels can be used when there is no true label).
[0043] On the basis of pre-training, use the labeled data for fine-tuning. Introduce a contrastive learning module, input different enhanced versions (such as rotation, scaling) of the same image into the model, and further strengthen the system's discriminative ability for building features by maximizing the similarity of homogeneous pixel features and minimizing the similarity of heterogeneous pixels. Finally, the total loss function of the remote sensing image building semantic segmentation system is: ; Among them, is the contrastive learning loss function, is the balance coefficient.
[0044] To adapt to the scene differences of remote sensing images, the dynamic parameter adjustment unit is used to perform dynamic parameter adjustment by threshold adaptive update and query text weight redistribution.
[0045] Among them, the process of the dynamic parameter adjustment unit performing dynamic parameter adjustment by threshold adaptive update is as follows: According to the precision-recall curve of the system on the validation set during the training process, the binarization threshold is dynamically adjusted , when the precision increases and the recall decreases, reduce ; when the precision decreases and the recall increases, increase to balance the segmentation accuracy and integrity.
[0046] After each training iteration, the historical accuracy corresponding to the query text is statistically calculated, and its weight in the instance fusion stage is dynamically updated. Specifically, during the process of the dynamic parameter adjustment unit performing dynamic parameter adjustment by query text weight redistribution, The update formula of
[0047] Among them, is the exponential decay factor.
[0048] Through threshold adaptive update and query text weight redistribution, the system gradually focuses on the query texts with excellent performance, enhancing the recognition ability of complex building forms. Through the above self-training process, the system can effectively utilize the valid information in the initial pseudo-labels, and at the same time reduce the influence of noise through iterative optimization, ultimately achieving a significant improvement in the semantic segmentation accuracy of buildings in remote sensing images.
[0049] In summary, the remote sensing image building semantic segmentation system based on the vision-language model according to the above embodiments constructs a diverse query text set, generates text features using the text encoder of the alignment prompt encoder, and at the same time extracts multi-scale features of the dual-temporal remote sensing images through the vision encoder. Then, an adaptive cross-modal attention mechanism is adopted to deeply fuse the text features and the visual features, enabling precise alignment of the features. Threshold processing, connected region analysis, and overlapping removal are performed on the fused features to generate initial building instances. The instance confidence of different query texts is fused by weighted summation to obtain an initial building mask, which is then binarized to obtain an initial building segmentation pseudo-label. Then, through pseudo-label screening, model iterative training, and dynamic parameter adjustment, the reliability of the pseudo-label and the adaptability of the system to the remote sensing scene are gradually improved. It can effectively utilize the initial pseudo-label and reduce the influence of noise, improving the segmentation accuracy of buildings in remote sensing images, solving the problem that the building labels generated by pre-training VLMs on natural scene images cannot be directly applied to remote sensing images, and reducing the dependence on manual annotation of remote sensing images.
[0050] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.
Claims
1. A remote sensing image building semantic segmentation system based on a vision language model, characterized in that, It includes an initial building mask generation module, a pseudo-label optimization and model iterative enhancement module; the initial building mask generation module includes a query text set construction unit, a feature extraction unit, a cross-modal attention interaction unit, an initial building instance generation unit, and an instance fusion and mask generation unit; the pseudo-label optimization and model iterative enhancement module includes a pseudo-label screening unit, a model iterative training unit, and a dynamic parameter adjustment unit; The query text set construction unit is used to construct a diverse query text set composed of multiple prompt words; The feature extraction unit is used to extract the text features of the query text set using an aligned prompt encoder, and extract the multi-scale visual features of the remote sensing image; The cross-modal attention interaction unit is used to interact the text features and multi-scale visual features using an adaptive cross-modal attention mechanism to obtain visual features integrating text information; The initial building instance generation unit is used to perform threshold processing, connected region analysis, and overlapping removal on the visual features integrating text information to generate initial building instances; The instance fusion and mask generation unit is used to fuse the initial building instances obtained based on different query texts to obtain an initial building mask, and obtain an initial building segmentation pseudo-label based on the initial building mask; The pseudo-label screening unit is used to screen out reliable pseudo-labels from the initial building segmentation pseudo-labels based on uncertainty quantification, cross-validation consistency, and geometric constraint filtering; The model iterative training unit is used to gradually enhance the generalization ability of the system using a two-stage training method of weak supervision training and fine-tuning optimization; The dynamic parameter adjustment unit is used to perform dynamic parameter adjustment using threshold adaptive update and query text weight redistribution.
2. The remote sensing image building semantic segmentation system based on a vision-language model according to claim 1, wherein Specifically, the feature extraction unit is used to input each query text in the query text set into the text encoder in the aligned prompt encoder, and the text encoder converts the query text into corresponding text features; The feature extraction unit is also used to process the remote sensing image using the visual encoder in the aligned prompt encoder to extract multi-scale visual features.
3. The remote sensing image building semantic segmentation system based on a vision-language model according to claim 2, wherein Specifically, the cross-modal attention interaction unit is used to: Calculate the similarity between the text features of each query text in the query text set and the feature maps of each scale in the multi-scale visual features based on cosine similarity. For the th remote sensing image at the th scale of the feature map , flatten into a vector sequence , and then calculate the dot product of the th text feature and each element in to obtain a similarity score matrix ; Then use the softmax function to normalize the similarity score matrix to obtain a normalized attention weight matrix, and the expression is: Among them, represents the th element in represents the th element in is the sequence length of the similarity score matrix, represents the th element in the normalized attention weight matrix; Next, an adaptive adjustment factor is introduced The normalized attention weight matrix is adjusted, and the expression is as follows: Among them, represents the -th element in the final attention weight matrix; Finally, use to perform weighted summation to obtain the visual features that fuse text information , and the expression is: Among them, denotes the th element.
4. The remote sensing image building semantic segmentation system based on a vision-language model according to claim 3, wherein Specifically, the initial building instance generation unit is used to: Perform threshold processing, and set the threshold . Consider the elements in that are greater than as the part belonging to the building, and consider the elements in that are less than or equal to as the background, thereby obtaining a binary feature map ; Process the binarized feature map using the connected component analysis algorithm to find the set of foreground pixels that are connected to each other. Each connected component represents a potential building instance; Perform post-processing on the detected connected regions. When there is an overlap between two or more connected regions, screen according to the confidence of the connected regions, retain the connected region with the highest confidence, and remove other connected regions, so as to obtain a group of non-overlapping initial building instances.
5. The remote sensing image building semantic segmentation system based on a vision-language model according to claim 4, wherein The way the instance fusion and mask generation unit fuses the initial building instances obtained based on different query texts is as follows: perform a weighted sum of the instance confidences at each pixel position, and the weights are dynamically assigned according to the effectiveness of different query texts. Among them, for the th query text weight , the calculation formula is: Among them, is the total number of query texts in the query text set, is the th query text 's historical accuracy rate, is the th query text 's historical accuracy rate; For the pixel position , is the abscissa of the pixel, is the ordinate of the pixel, and the confidence of the building instance obtained based on the query text is weighted and summed to obtain the fused confidence , and the expression is: After fusion, a set of building instances is obtained, and then the building instances in the set of building instances are merged into an initial building mask , , is a set of real numbers, and are respectively the height and width of; Then, based on the initial building mask an initial building segmentation pseudo-label is obtained, and the expression is: Among them, is the initial building mask the pixel position the corresponding value, is the initial building segmentation pseudo-label the pixel position the corresponding value, is the binarization threshold.
6. The remote sensing image building semantic segmentation system based on a vision-language model according to claim 5, characterized in that, The process of the pseudo-label screening unit screening based on uncertainty quantification is: Calculate the pixel-level entropy value of the initial building mask, and the calculation formula is: Among them, is the pixel position corresponding to the pixel-level entropy value; For pixel regions where the pixel-level entropy value is higher than the entropy threshold the pseudo-label confidence is downscaled or marked for correction; The process of the pseudo-label screening unit screening based on cross-validation consistency is: Generate multiple sets of pseudo-labels for the same remote sensing image using alignment hint encoders with different initializations, calculate the pixel-level intersection over union (IoU). If the average pixel-level IoU of multiple sets of pseudo-labels for the target area is lower than the IoU threshold , then remove the pseudo-labels of this target area from the training set The process of the pseudo-label screening unit screening based on geometric constraint filtering is: Based on the geometric prior of the buildings in the remote sensing image, eliminate the pseudo-label regions that do not conform to the preset shape features.
7. The remote sensing image building semantic segmentation system based on a vision-language model according to claim 6, wherein The model iterative training unit satisfies the following formula: ; Among them, is the total loss function of the remote sensing image building semantic segmentation system, is the weighted cross-entropy loss function, is the contrastive learning loss function, is the balance coefficient, is the pre-training dataset is the number of samples in it, is the pseudo-label quality weight, is the cross-entropy function, is the true label.
8. The remote sensing image building semantic segmentation system based on a vision-language model according to claim 7, characterized in that The process of the dynamic parameter adjustment unit performing dynamic parameter adjustment using threshold adaptive update is: Dynamically adjust the binarization threshold according to the precision-recall curve of the system on the validation set during the training process , when the precision increases and the recall decreases, lower ; when the precision decreases and the recall increases, raise , to balance the segmentation accuracy and integrity; During the process of the dynamic parameter adjustment unit performing dynamic parameter adjustment by querying text weight redistribution, The update formula is as follows: Among them, is the exponential decay factor.
Citation Information
Patent Citations
Remote sensing image cross-modal retrieval method based on language and visual detail feature fusion
CN116775922A
Semi-supervised change detection method and device based on visual language model
CN118447344A
Remote sensing image semantic segmentation method and device based on point annotation expansion network
CN118968046A
Weak supervision semantic segmentation method and system based on text mask collaborative prompt
CN119919663A
Building semantic segmentation method for correcting unsupervised domain adaptive pseudo tag by using segmentation large model
CN120014643A
Cited By
Building programmed modeling method and system based on unmanned aerial vehicle image
CN120635338A
Remote sensing image building change detection system based on self-training and consistency learning
CN121147220A
Remote sensing image building change detection system based on self-training and consistency learning
CN121147220B
Waste slag field functional area zero sample identification method based on image processing and product
CN121170608A
Semantic-based building design image hierarchical data enhancement method and device, and medium
CN121236530A