Vision-language collaborative heterogeneous agricultural greenhouse remote sensing extraction method and system
Patent Information
- Application Number
- CN202610957064.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-25
AI Technical Summary
对于异构大棚识别而言,虽然不同类型大棚在材料、形态和纹理上存在差异,但其视觉表征仍应与大棚语义概念保持较高一致性;而与大棚在局部响应上相似的干扰地物,通常难以与目标语义形成稳定匹配
[0069]本发明实施例所提供的基于视觉-语言协同的异构农业大棚遥感提取方法及系统,针对异构大棚形态、材质、颜色差异大且易与复杂背景混淆的问题,构建了正负提示词约束的CLIP文本词汇概率先验生成方法,通过正向目标提示增强大棚响应、负向背景提示抑制道路、地膜、屋顶、光伏板等相似背景响应;同时引入DINOv3稠密视觉特征获得的结构引导信息,补充大棚区域的结构一致性、边界纹理和上下文关系表达,并结合全局局部特征生成模块(Global Local Feature Generation, GLFG)和双向交互注意力融合模块,实现RGB主干特征、文本词汇概率先验和结构引导信息的协同建模,提升了模型对跨场景、跨材质异构大棚的识别稳定性和复杂背景下的分割精度。
Smart Images

Figure CN122821384A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of agricultural remote sensing image recognition, specifically relating to a method and system for remote sensing extraction of heterogeneous agricultural greenhouses based on vision-language collaboration. Background Technology
[0002] Land use change is a crucial factor influencing surface processes and regional climate, with the expansion of artificial land cover impacting regional surface temperature and energy exchange. Agricultural greenhouses, as a rapidly expanding typical artificial land cover, alter original surface characteristics, affecting surface albedo, evapotranspiration, and heat exchange processes, thus disrupting regional surface energy balance. Therefore, accurately acquiring surface information about agricultural greenhouses is essential for regional surface process analysis and climate simulation research. However, agricultural greenhouses exhibit significant spatiotemporal heterogeneity. Spatially, the diversity of covering materials, structural forms, and distribution patterns among different greenhouses leads to significant differences in their spectral response, texture characteristics, and geometric shapes. Temporally, seasonal changes, crop growth, and production management processes cause continuous variations in greenhouse covering conditions and surface reflectance characteristics, making these changes difficult to consistently capture and increasing the difficulty of greenhouse identification. Furthermore, similar backgrounds such as bare land, roads, plastic film mulch, and farmland are often interspersed with greenhouses, exacerbating background interference and spatial confusion in greenhouse identification, posing significant challenges to regional surface energy balance assessment and climate simulation parameterization. With the continuous development of the application of multi-source heterogeneous data in remote sensing monitoring, the method based on multi-source data collaboration has become an important technical means to support the accurate extraction of agricultural greenhouse data.
[0003] In existing technologies, remote sensing extraction methods for agricultural plastic greenhouses can be categorized into two types. The first type is based on prior rules and index / threshold methods, typically relying on artificial features such as spectral indices, reflectance differences, and texture. Rapid identification is achieved through threshold discrimination or rule constraints. Some scholars have proposed the Plastic Greenhouse Index (PGI) based on the sensitivity and discriminative power of the spectral characteristics of plastic greenhouses to accurately map and estimate the distribution of plastic greenhouses (PGs). Others have emphasized the importance of texture features in reflecting image grayscale changes and repetitive patterns, extracting key texture features from Landsat 8 imagery using the GLCM method and combining them with other features to achieve efficient classification. However, while these methods have low computational overhead, they are sensitive to changes in imaging conditions and background interference, and have limited adaptability across regions and seasons. The second type comprises novel segmentation methods represented by deep learning. These methods often employ a convolutional neural network encoder-decoder framework to learn multi-scale representations in end-to-end training, significantly improving extraction capabilities in complex scenes and promoting refined mapping applications. However, under conditions where heterogeneous greenhouses exhibit significant differences in appearance and coexist closely with similar ground features, they are still susceptible to the combined effects of heterogeneous spectra within the same category and heterogeneous spectra between different categories, resulting in prominent issues of false detection and missed detection.
[0004] Agricultural greenhouse extraction methods based on multi-source data collaboration mainly utilize various data sources such as optical images, lidar, temperature data, and textual semantic information to extract greenhouse data. Based on data modal attributes, these methods can be further categorized into two types: homomodal data collaboration and multimodal data collaboration. Homomodal data collaboration primarily involves introducing information from different spatial resolutions, sensors, or feature levels within the same data type to enhance the spatial constraints and spectral discrimination capabilities of the target. For example, González-Yebra et al. introduced detailed plot boundaries provided by high-resolution aerial imagery to constrain the spatial units involved in classification using medium-resolution pixels from Landsat. They extracted and combined multispectral band features, as well as Greenhouse Detection Index (GDI) and Plastic-Mulched Landcover Index (PMLI) features, thereby enhancing the spectral discrimination capability of greenhouse targets, reducing the interference of mixed pixels from medium-resolution images on classification results, and improving the accuracy of random forest models in identifying greenhouses. Multimodal data collaboration methods emphasize the complementarity of different modalities in information representation. By combining heterogeneous information such as optical imagery, SAR data, LiDAR data, and textual information, the representation of greenhouse targets is enhanced from multiple dimensions, including spectral, scattering, structural, and semantic dimensions. For example, some researchers extract SAR scattering features from optical spectra and texture features separately, and construct a joint feature vector by concatenating the two sets of feature vectors. This joint feature vector is then input into a machine learning classifier for extracting agricultural greenhouses covered by plastic film. Employing multi-source data collaboration methods can supplement greenhouse target information from multiple dimensions, including spectral, structural, and semantic dimensions, thereby enhancing the comprehensive representation capability of agricultural greenhouses in complex scenarios.
[0005] Existing multimodal collaborative methods can generally be summarized into three paradigms: one-way transfer, fusion, and interactive methods. One-way transfer methods rely on the model results of upstream modalities to guide subsequent branches. If there are misclassifications, omissions, or boundary shifts in the upstream results, these biases may continue to propagate during subsequent learning. Fusion methods typically perform channel splicing, weighted overlay, or unified encoding of spectral, texture, structural, or high-level semantic features extracted from different modalities. These methods can integrate multi-source information, but they often remain at the modal feature level, lacking sufficient constraints on the correspondence between modal information and target semantics. Interactive collaborative methods establish information transfer and feature association mechanisms between different modalities, fully exploring complementary information in material properties, structural morphology, and spatial distribution, thereby enhancing the ability to represent ground features. For example, some researchers have constructed feature association relationships between SAR data and optical data, using this relationship to guide the SAR branch to accurately locate and extract edge structures, and feeding back the calibrated association information to the optical branch to strengthen the texture feature expression in the optical branch, achieving iterative optimization of the features of both modalities. These methods lack an effective information filtering mechanism during the interaction process, which can easily lead to a large amount of low-contribution or even non-contribution information being repeatedly propagated between modalities through association weights, thereby interfering with the extraction and expression of effective features.
[0006] To alleviate these problems, some studies have further introduced feature selection and constraint mechanisms into multimodal interaction processes to reduce the propagation of redundant information. For example, some researchers have combined dot product similarity measurement with clustering selection mechanisms to calculate the statistical correlation of features from two different modalities, selecting effective features that are highly correlated with both modalities. However, the statistical correlation weight only reflects the similarity of cross-modal features in numerical responses and cannot truly represent their semantic consistency. Therefore, this type of method is prone to misjudging features that are numerically highly similar but have different category attributes as effective information and transmitting them back to the interaction process, thereby interfering with the effective expression of target ground features and weakening the separability between the target and similar interfering ground features.
[0007] The aforementioned research indicates that while existing multimodal interaction methods can improve information utilization efficiency through feature association and filtering mechanisms, their discrimination criteria primarily remain at the level of mathematical and statistical responses between modalities, lacking further constraints on the semantic attributes of cross-modal information. In complex scenarios, this lack of discrimination makes it difficult for interaction mechanisms to distinguish between complementary information that truly contributes to the target expression and interfering information that is only similar in local responses, thus leading to the retention or even amplification of semantic mismatches during the interaction process. Therefore, the key deficiency of existing multimodal collaborative interaction methods is not merely insufficient information filtering, but rather the widespread lack of semantic consistency discrimination capability, which is also a significant root cause of information mismatches and category confusion.
[0008] This type of semantic mismatch caused by statistical similarity can be identified using semantic constraints provided by visual-language models, but existing cross-modal interaction mechanisms have not yet fully incorporated this semantic constraint capability. Visual-language pre-trained models align image representations with textual semantics into a unified embedding space, allowing target categories to be directly defined by semantic concepts such as "color greenhouse" and "glass greenhouse." Based on this alignment, whether an image region belongs to the target category no longer depends solely on the numerical response similarity between modalities, but can also be determined by the degree of matching between its visual representation and the target semantic concept. For heterogeneous greenhouse identification, although different types of greenhouses differ in materials, shapes, and textures, their visual representations should still maintain a high degree of consistency with the greenhouse semantic concept; while interfering features that are similar to greenhouses in local responses are usually difficult to form a stable match with the target semantics. Therefore, this type of statistically similar but semantically inconsistent information can be identified through visual-language semantic constraints and suppressed during the interaction process.
[0009] The present invention aims to design a multimodal interaction model that combines effective information filtering and semantic consistency discrimination capabilities to suppress the erroneous propagation of semantically mismatched information during cross-modal information transmission, thereby improving the accurate identification capability of heterogeneous greenhouses. Summary of the Invention
[0010] In view of the above-mentioned defects or deficiencies in the prior art, the present invention aims to provide a remote sensing extraction method and system for heterogeneous agricultural greenhouses based on vision-language collaboration. It adopts a vision-language collaboration strategy that combines interaction and fusion. Through interaction, semantic priors provided by textual vocabulary are injected into visual representations. Through fusion, local discriminative information and scene-level contextual constraints are uniformly integrated in multi-scale spatial features, thereby achieving stable identification and fine segmentation of heterogeneous greenhouses.
[0011] To achieve the above objectives, the embodiments of the present invention adopt the following technical solutions:
[0012] In a first aspect, embodiments of the present invention provide a remote sensing extraction method for heterogeneous agricultural greenhouses based on vision-language collaboration, the method comprising the following steps:
[0013] Step S1: Obtain RGB remote sensing images with artificial labels for heterogeneous greenhouses and divide them into training set, validation set and test set;
[0014] Step S2: Based on RGB remote sensing images, generate corresponding dense visual features using a visual basic model as priors for dense visual features;
[0015] Step S3: Based on RGB remote sensing imagery, generate CLIP text vocabulary probability priors guided by text semantics using a visual-language model, and construct corresponding soft labels;
[0016] Step S4: Construct a heterogeneous agricultural greenhouse target remote sensing extraction model based on a vision-language collaborative network, and output the greenhouse segmentation results;
[0017] Step S5: Construct the multi-source prior consistency loss function for the heterogeneous agricultural greenhouse target remote sensing extraction model;
[0018] Step S6: Train and optimize the parameters of the remote sensing extraction model for heterogeneous agricultural greenhouse targets based on the training set and validation set;
[0019] Step S7: Input the remote sensing images and corresponding dense visual feature priors and CLIP text word probability priors from the dataset of the area to be tested into the trained heterogeneous agricultural greenhouse target remote sensing extraction model to obtain the category prediction probability, and then obtain the binary segmentation result of the heterogeneous greenhouse area.
[0020] As a preferred embodiment of the present invention, the heterogeneous greenhouse includes greenhouse targets with different covering materials, different greenhouse colors, different structural forms, and different spatial distribution methods.
[0021] In a preferred embodiment of the present invention, in step S2, the visual base model is DINOv3; RGB remote sensing images are input into DINOv3 to obtain block-level dense visual features.
[0022] In a preferred embodiment of the present invention, the visual-language model in step S3 employs the contrastive language-image pre-trained model CLIP; generating a text semantically guided CLIP text vocabulary probability prior includes:
[0023] Step S31: Construct a dynamic semantic prompt pool consisting of basic greenhouse vocabulary, expert experience vocabulary, and regional network retrieval information, and divide the prompt words into positive target prompt words and negative background prompt words;
[0024] Step S32: Input each positive target cue word and negative background cue word into the CLIP text encoder to obtain the text embedding, and then perform normalization processing:
[0025] (2)
[0026] (3)
[0027] In equations (2) and (3), For the i-th positive prompt text, For CLIP text encoder, This is the normalized text embedding vector; For the j-th negative prompt text, This is the normalized negative text embedding vector;
[0028] Step S33: Input RGB remote sensing image According to multi-scale spatial grid The image is divided into regions at different scales; for each scale... The next Image regions Normalized image embeddings are obtained using the CLIP image encoder:
[0029] (4)
[0030] In equation (4), This indicates the CLIP image encoder. This represents the normalized image region embedding;
[0031] Aggregate similarity response between image regions and positive target cues and negative background cues , They are represented as follows:
[0032] (5)
[0033] In equation (5), The cosine similarity between the image region embedding and the positive target cue text embedding is represented. The cosine similarity between the image region embedding and the negative background cue word text embedding is used to represent the cosine similarity between the image region embedding and the negative background cue word text embedding. This represents an aggregation operation that evaluates the similarity responses to multiple positive target prompts. This represents an aggregation operation that evaluates the similarity responses to multiple negative background cue words.
[0034] The scale is obtained by the difference between positive and negative responses. Next The target semantic prior response of each image region:
[0035] (6)
[0036] In equation (6), The text lexical probabilistic prior response of the image region to the greenhouse target;
[0037] Step S34 involves spatial reconstruction, weighted fusion, and normalization of the target semantic prior responses at different scales to obtain text semantic guidance target semantic probability prior information corresponding to the original image space. .
[0038] As a preferred embodiment of the present invention, step S34 further includes: constructing soft labels based on the prior response, which are used to jointly constrain the model with artificial labels in the subsequent training stage, so that the model can enhance its ability to suppress similar backgrounds while learning the target features of the greenhouse.
[0039] In a preferred embodiment of the present invention, the heterogeneous agricultural greenhouse target remote sensing extraction model in step S4 takes RGB remote sensing imagery, dense visual feature priors, and text semantic-guided CLIP text vocabulary probability priors as input, and outputs greenhouse segmentation results. The model obtains greenhouse segmentation results by extracting the backbone visual features of RGB remote sensing imagery, dense visual feature priors, and text semantic-guided CLIP text vocabulary probability priors, performing structure-guided mapping, semantic prior mapping, global and local feature generation, and bidirectional interactive attention fusion.
[0040] In a preferred embodiment of the present invention, the loss function in step S5 includes manual label supervision loss, semantic prior soft supervision loss, prior consistency loss, and distillation consistency loss.
[0041] As a preferred embodiment of the present invention, the manual label supervision loss The formula is as follows:
[0042] (7)
[0043] In equation (7), Represents cross-entropy loss, Indicates pixel position, and These represent the image height and width, respectively. Indicates the segmentation model at pixel location The output is the category prediction score; Indicates pixel position The corresponding manually labeled category tag;
[0044] Define the high confidence region as:
[0045] (8)
[0046] In equation (8), This indicates a priori high-confidence regions. Indicates the confidence margin threshold; Indicates pixel position The prior probability value generated by the similarity response of positive target cue words and negative background cue words is used to characterize the prior confidence level that the pixel belongs to the target greenhouse;
[0047] Introduce a prior consistency loss within the high-confidence prior region:
[0048] (9)
[0049] In equation (9), This represents the unnormalized prediction score of the binary cross-entropy. Represents pixels The unnormalized prediction score of the greenhouse prospect at the main branch. This indicates the number of pixels within the high confidence region;
[0050] In high confidence region Internally, based on the prior probability response of text words. As a soft objective, the semantic prior soft supervision loss is defined as follows:
[0051] (10)
[0052] In equation (10), Indicates text semantic auxiliary branches at the pixel level The unnormalized prediction score of the corresponding greenhouse prompt;
[0053] The formula for the distillation consistency loss is as follows:
[0054] (11)
[0055] In equation (11), This represents the Sigmoid function. Indicates the consistency threshold. The set of pixels representing the auxiliary branch that is consistent with the prior probability response of the text words;
[0056] The total loss function is expressed as:
[0057] (12)
[0058] In equation (12), Losses due to manual labeling supervision The prior consistency loss between the main segmentation branch and the prior probability of the text words. The semantic prior soft-supervised loss is used to bridge the semantic auxiliary branch of the text and the text lexical probability prior. To account for the distillation consistency loss from the auxiliary branch to the main split branch, , and These are the weighting coefficients for the corresponding loss terms.
[0059] In a preferred embodiment of the present invention, during the training process in step S6, the manual labeling supervision loss is used to ensure that the model learns the manually labeled greenhouse areas, the text vocabulary probability prior consistency loss is used to enhance the model's response to high-confidence greenhouse prior areas, the auxiliary consistency loss is used to constrain the output of the text semantic auxiliary branch, and the distillation consistency loss is used to pass the text semantic auxiliary information to the main segmentation branch within the reliable area. Through the joint optimization of multiple losses, the model can simultaneously utilize manual labeling information and external text vocabulary probability prior information to improve the ability to extract greenhouse targets under different materials, different shapes, and complex background conditions.
[0060] Secondly, embodiments of the present invention also provide a heterogeneous agricultural greenhouse remote sensing extraction system based on vision-language collaboration. The system includes: a tag data acquisition module, a dense visual feature prior generation module, a CLIP text vocabulary probability prior generation module, a model building module, a loss function construction module, a model training module, and a greenhouse image extraction module; wherein...
[0061] The tag data acquisition module is used to acquire RGB remote sensing images with heterogeneous greenhouse artificial tags and divide them into training set, validation set and test set;
[0062] The dense visual feature prior generation module is used to generate corresponding dense visual features based on RGB remote sensing images using a visual basic model, as dense visual feature priors.
[0063] The CLIP text vocabulary probability prior generation module is used to generate CLIP text vocabulary probability priors guided by text semantics based on RGB remote sensing images using a visual-language model, and to construct corresponding soft tags.
[0064] The model building module is used to construct a heterogeneous agricultural greenhouse target remote sensing extraction model based on a vision-language collaborative network and output the greenhouse segmentation results;
[0065] The loss function construction module is used to construct a multi-source prior consistency loss function for a heterogeneous agricultural greenhouse target remote sensing extraction model.
[0066] The model training module is used to train and optimize the parameters of the remote sensing extraction model for heterogeneous agricultural greenhouse targets based on the training set and validation set.
[0067] The greenhouse image extraction module is used to input the remote sensing images in the dataset of the area to be tested, along with the corresponding dense visual feature priors and CLIP text word probability priors, into the trained heterogeneous agricultural greenhouse target remote sensing extraction model to obtain the category prediction probability, and thus obtain the binary segmentation result of the heterogeneous greenhouse area.
[0068] The technical solutions provided in the embodiments of the present invention have the following beneficial effects:
[0069] The remote sensing extraction method and system for heterogeneous agricultural greenhouses based on vision-language collaboration provided in this invention addresses the problem that heterogeneous greenhouses have large differences in shape, material, and color, and are easily confused with complex backgrounds. It constructs a CLIP text vocabulary probability prior generation method with positive and negative cue word constraints. Positive target cues enhance the greenhouse response, while negative background cues suppress responses from similar backgrounds such as roads, plastic film, roofs, and photovoltaic panels. Simultaneously, it introduces structural guidance information obtained from DINOv3 dense visual features to supplement the structural consistency, boundary texture, and contextual relationship expression of the greenhouse area. Combined with a Global Local Feature Generation (GLFG) module and a bidirectional interactive attention fusion module, it achieves collaborative modeling of RGB backbone features, text vocabulary probability priors, and structural guidance information, improving the model's recognition stability for heterogeneous greenhouses across scenes and materials, and its segmentation accuracy in complex backgrounds.
[0070] Of course, implementing any product or method of the present invention does not necessarily require achieving all of the advantages described above at the same time. Attached Figure Description
[0071] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0072] Figure 1 This is a flowchart of the remote sensing extraction method for heterogeneous agricultural greenhouses based on vision-language collaboration as described in an embodiment of the present invention;
[0073] Figure 2 This is a schematic diagram of the remote sensing extraction method for heterogeneous agricultural greenhouses based on vision-language collaboration as described in an embodiment of the present invention.
[0074] Figure 3 This is a visual comparison chart of the method described in the embodiments of the present invention on the second region test dataset and the classic network;
[0075] Figure 4 This is a visualization comparing the method described in this embodiment of the invention with other advanced greenhouse extraction methods on test data in the second region;
[0076] Figure 5 This is an ablation visualization effect based on different prior information on the test dataset in the second region using the method described in the embodiments of the present invention;
[0077] Figure 6This is an ablation visualization effect diagram based on different key modules on the second region test dataset using the method described in the embodiments of the present invention. Detailed Implementation
[0078] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. It should be noted that, without conflict, the embodiments and features in the embodiments of the present invention can also be combined with each other.
[0079] It should be noted that similar reference numerals and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In the description of the embodiments of the present invention, the terms "first," "second," "third," "fourth," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance. In addition, sometimes a subscript such as W1 may be written in a non-subscript form such as W1, and their meanings are consistent unless the distinction is emphasized.
[0080] It should be understood that the term "and / or" in this embodiment is merely a description of the relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. A and B can be singular or plural. Additionally, the character " / " in this embodiment generally indicates an "or" relationship between the preceding and following associated objects, but it may also indicate an "and / or" relationship. Please refer to the preceding and following text for a more detailed understanding.
[0081] This invention provides a method and system for remote sensing extraction of heterogeneous agricultural greenhouses based on vision-language collaboration. It employs a vision-language collaboration strategy combining interaction and fusion to construct a model with semantic awareness capabilities that integrates multimodal information. Through interaction, semantic priors provided by textual vocabulary are injected into the visual representation. Fusion integrates local discriminative information and scene-level contextual constraints across multi-scale spatial features, thereby achieving stable identification and fine segmentation of heterogeneous greenhouses. The method follows a process of "backbone representation construction—multi-source prior generation—collaborative fusion optimization—result output." First, multi-scale feature encoding is completed using RGB remote sensing images as the main input, forming the basic semantic representation required for subsequent fusion. Second, two types of external priors are introduced in parallel to enhance cross-scene stability: one is the dense structural feature prior provided by DINOv3, used to supplement more robust spatial structural guidance information; the other is the probabilistic prior and training-period soft labels generated by CLIP based on text prompts, used to provide spatial indication and semantic consistency constraints for the target area. Subsequently, the core features and multi-source priors are aligned and complementarily modeled in the collaborative fusion module. The interaction of global and local information enhances the balance between structural consistency and boundary details. Finally, the fused features are decoded to restore spatial resolution and output pixel-level segmentation results. Meanwhile, during the training phase, prior weights and soft supervision are used to constrain the optimization process, thereby improving the method's generalization ability and segmentation robustness under cross-scene and cross-material conditions.
[0082] On the one hand, this invention proposes a unified framework consisting of CLIP textual lexical probabilistic semantic priors, DINOv3 dense visual structure guidance, GLFG global / local structured feature generation, and subsequent heterogeneous fusion. This framework achieves the organic organization and collaborative modeling of multi-source prior information and backbone visual features. DINOv3 is used to extract dense visual representations with strong spatial consistency and structural stability, providing structural guidance for the boundary textures, area continuity, and contextual relationships of the greenhouse region, thus compensating for the insufficient structural expression of a single backbone network under heterogeneous materials, complex backgrounds, and cross-scene conditions. On the other hand, this invention proposes a method for constructing positive and negative cue words, a differentiated aggregation method for Agg⁺ / Agg⁻, and a probabilistic prior allocation mechanism oriented towards global / local branches. Through the joint constraint of positive target cues and negative background cues, the model's responsiveness to heterogeneous greenhouses and its ability to distinguish complex backgrounds are enhanced, allowing external priors to participate more effectively in feature expression, semantic guidance, and fusion decision-making.
[0083] like Figure 1 and Figure 2 As shown, the vision-language collaborative remote sensing extraction method for heterogeneous agricultural greenhouses can be implemented by an electronic device, which can be a terminal or a server. The method includes the following steps:
[0084] Step S1: Obtain RGB remote sensing images with artificial labels for heterogeneous greenhouses and divide them into training set, validation set and test set.
[0085] In this step, RGB remote sensing images of the study area are acquired as the main input data, and heterogeneous greenhouse labels are created based on the results of manual interpretation. These heterogeneous greenhouses include targets with different covering materials, greenhouse colors, structural forms, and spatial distribution patterns, such as white plastic film greenhouses, green shade net greenhouses, strip-shaped arched greenhouses, and greenhouse areas distributed in clusters. The manual labels are used to identify greenhouse areas in the RGB remote sensing images and serve as real-world supervisory information for subsequent model training and accuracy evaluation.
[0086] After obtaining RGB remote sensing images and their corresponding artificial labels, spatial registration, cropping, and format standardization were performed on the images and labels to obtain paired image samples and label samples. Subsequently, the samples were divided into training, validation, and test sets. The training set was used for model parameter learning, the validation set for monitoring the training process and parameter selection, and the test set for evaluating the model's extraction performance on heterogeneous greenhouse targets.
[0087] After the above processing, the basic data for subsequent DINOv3 spatial structure guidance information generation, CLIP text vocabulary probability prior generation, soft label construction, and visual-language heterogeneous fusion network training are obtained.
[0088] Step S2: Based on RGB remote sensing imagery, generate corresponding dense visual features using a visual basic model as priors for dense visual features.
[0089] In this step, the visual base model uses DINOv3. RGB remote sensing images are input into DINOv3 to obtain block-level dense visual representations. Since the output features of DINOv3 have high dimensionality and their spatial scale is not entirely consistent with the backbone features of the segmentation network, this embodiment further converts them into spatial structure guidance information through channel compression, scale alignment, and convolutional mapping. This information is primarily used to supplement the structural consistency, boundary texture, and contextual relationships of the greenhouse area, enhancing the model's structural perception capabilities across scenes, materials, and complex backgrounds. Specifically, RGB images are input into DINOv3 to obtain mid-to-high-level feature maps, which are then aligned to the network fusion resolution.
[0090] , = (1)
[0091] In equation (1), For DINOv3 feature extractor, It is a multi-channel dense visual representation extracted by DINOv3; This indicates interpolation and scale alignment operations used to adjust the DINOv3 output features to the spatial resolution required for subsequent network feature fusion. This represents the aligned visual structural features of DINOv3. These features serve as spatial structure guidance information in subsequent network feature fusion, supplementing the structural consistency, boundary texture, and contextual relationships of the greenhouse area.
[0092] Step S3: Based on RGB remote sensing imagery, generate CLIP text vocabulary probability priors guided by text semantics using a visual-language model, and construct corresponding soft labels.
[0093] In this step, the visual-language model employs a Contrastive Language-Image Pre-training (CLIP) model. CLIP is not used as the final segmentation network, but rather to calculate the similarity response between image regions and text prompts, thereby generating a CLIP text vocabulary probability prior guided by text semantics. This CLIP text vocabulary probability prior represents the confidence that an image region belongs to the greenhouse target in the form of a continuous response map, and is further used to construct soft labels during the training phase.
[0094] The CLIP text lexical probability prior, which generates semantically guided text, includes:
[0095] Step S31 involves constructing a dynamic semantic cue pool composed of basic greenhouse vocabulary, expert experience vocabulary, and regional network retrieval information, and dividing the cue words into positive target cue words and negative background cue words. Positive target cue words are used to describe greenhouse targets of different shapes, materials, colors, and spatial distributions, such as white plastic film greenhouses, green shade net greenhouses, black mesh greenhouses, multi-span greenhouses, and strip-shaped arched greenhouses. Negative background cue words are used to describe background features easily confused with greenhouses, such as bare land, roads, plastic mulch, rural roofs, photovoltaic panels, and farmland.
[0096] Step S32: Input each positive target cue word and negative background cue word into the CLIP text encoder to obtain the text embedding, and then perform normalization processing:
[0097] (2)
[0098] (3)
[0099] In equations (2) and (3), For the i-th positive prompt text, For CLIP text encoder, This is the normalized text embedding vector; For the j-th negative prompt text, This is the normalized negative text embedding vector.
[0100] Step S33: To accommodate greenhouse shapes at different scales, input RGB remote sensing images... According to multi-scale spatial grid The image is divided into regions at different scales. (Regarding scale...) The next Image regions Normalized image embeddings are obtained using the CLIP image encoder:
[0101] (4)
[0102] In equation (4), This indicates the CLIP image encoder. This represents the normalized image region embedding. Since both the image embedding and the text embedding have been normalized, the cosine similarity between them can be calculated using the vector dot product. Therefore, the aggregated similarity response between the image region and the positive target cue and the negative background cue is... , They are represented as follows:
[0103] (5)
[0104] In equation (5), The cosine similarity between the image region embedding and the positive target cue text embedding is represented. The cosine similarity between the image region embedding and the negative background cue word text embedding is used to represent the cosine similarity between the image region embedding and the negative background cue word text embedding. This represents an aggregation operation that evaluates the similarity responses to multiple positive target prompts. This represents an aggregation operation that evaluates the similarity responses to multiple negative background cue words.
[0105] The scale is obtained by the difference between positive and negative responses. Next The target semantic prior response of each image region:
[0106] (6)
[0107] In equation (6), This represents the textual lexical probabilistic prior response of an image region to the greenhouse target.
[0108] Positive target cue words are used to enhance the response in the greenhouse area, while negative background cue words are used to suppress the response to similar backgrounds.
[0109] Step S34 involves spatial reconstruction, weighted fusion, and normalization of the target semantic prior responses at different scales to obtain text semantic guidance target semantic prior information corresponding to the original image space. Simultaneously, soft labels are constructed based on this prior response and used in subsequent training phases to jointly constrain the model with artificial labels, enabling the model to enhance its ability to suppress similar backgrounds while learning the features of the greenhouse target.
[0110] Step S4: Construct a remote sensing extraction model for heterogeneous agricultural greenhouse targets based on a vision-language collaborative network, and output the greenhouse segmentation results.
[0111] In this step, a remote sensing extraction model for heterogeneous agricultural greenhouse targets is constructed based on a vision-language collaborative network to perform pixel-level extraction of heterogeneous greenhouse areas. The model uses RGB remote sensing imagery as the main input, while also incorporating DINOv3 spatial structure guidance information (dense visual feature prior) obtained in step S2 and CLIP textual semantic guidance target semantic prior information obtained in step S3. Through backbone visual feature extraction, structure guidance mapping, semantic prior mapping, global-local feature generation, and bidirectional interactive attention fusion, the greenhouse segmentation result is obtained.
[0112] Specifically, RGB remote sensing images are first input into the ResNet-50 backbone network to extract visual features at different levels. Then, channel projection and spatial scale uniformity processing are performed on the mid-to-high-level visual features to transform them to a unified feature dimension and spatial resolution. Simultaneously, the block-level dense visual representations extracted from DINOv3 are further mined through continuous convolution operations to obtain spatial structure guidance information. After channel projection and spatial resolution transformation, the spatial structure guidance information is converted to a feature space consistent with the backbone visual features, used to supplement the structural consistency, boundary texture, and contextual relationships of the greenhouse area. The CLIP textual semantic guidance target semantic prior information is converted into semantic prior features through convolutional mapping, used to provide spatial cues and semantic constraints for the greenhouse target.
[0113] Subsequently, the main visual features and spatial structure guidance information, after the channel projection and spatial scale consistency processing are completed, are input into the Global Local Feature Generation (GLFG) module to generate global and local representations, respectively. The global representation is used to describe the overall distribution, long-range spatial relationships, and regional consistency of the greenhouse area; the local representation is used to enhance the greenhouse boundaries, strip textures, and local geometric details.
[0114] In the feature interaction and fusion stage, global representations, local representations, semantic prior features, and structurally guided features are input into the Bidirectional Interaction Attention Fusion (BDIAF) module. This module primarily interacts with global and local representations, using a bidirectional attention mechanism to guide local detail expression with global information, while simultaneously allowing local structural information to supplement the global contextual representation. It also combines semantic prior features and structurally guided features to enhance the target area of the greenhouse and suppress similar backgrounds such as roads, bare land, mulch film, roofs, and photovoltaic panels. Through this process, the model establishes complementary relationships between global spatial relationships, local boundary structures, textual lexical probability priors, and spatial structural guidance, thereby improving its ability to distinguish heterogeneous greenhouses from complex backgrounds.
[0115] After bidirectional interactive attention fusion, the obtained fused features are further refined through convolution and linear interpolation to restore the original image resolution, outputting pixel-level prediction results for the greenhouse area. The model output is a greenhouse segmentation probability map, which can be converted into a binary segmentation result according to a set threshold, where the foreground represents the greenhouse area and the background represents the non-greenhouse area. Through the above process, the heterogeneous agricultural greenhouse target remote sensing extraction model can comprehensively utilize RGB visual features, DINOv3 spatial structure guidance information, and CLIP text vocabulary probability prior information to achieve fine extraction of agricultural greenhouses with different materials, shapes, and complex backgrounds.
[0116] Step S5: Construct the multi-source prior consistency loss function for the heterogeneous agricultural greenhouse target remote sensing extraction model.
[0117] In this step, based on the greenhouse segmentation prediction results output in step S4, the manual labels in step S1, and the text vocabulary probability prior information and soft labels generated in step S3, a loss function for model training is constructed. The loss function includes manual label supervision loss, semantic prior soft supervision loss, prior consistency loss, and distillation consistency loss, used to simultaneously constrain the model's prediction results for manually labeled regions, text vocabulary probability prior response regions, and high-confidence soft label regions.
[0118] Specifically, let the output of the main split branch of the model be... The foreground passage corresponds to the greenhouse category, denoted as Where logit represents the prediction score output by the neural network before Softmax or Sigmoid normalization. Let the manual label be... Where 1 represents the greenhouse area and 0 represents the non-greenhouse area. A system of artificially labeled loss monitoring is constructed. As the primary supervisory loss, it is used to ensure that the model output remains consistent with the manually labeled results.
[0119] (7)
[0120] In equation (7), Represents cross-entropy loss, Indicates pixel position, and These represent the image height and width, respectively. Indicates the segmentation model at pixel location The output is the category prediction score; Indicates pixel position The corresponding manually labeled category.
[0121] To utilize the prior information of text word probabilities generated in step S3, let its corresponding continuous prior response be... ,in Represents pixels The prior confidence level of belonging to the greenhouse target. Considering that the prior probability of text words may produce false responses in some similar background regions, this embodiment only introduces prior consistency loss in high-confidence prior regions. The high-confidence region is defined as:
[0122] (8)
[0123] In equation (8), This indicates a priori high-confidence regions. Indicates the confidence margin threshold; Indicates pixel position The prior probability value generated by the similarity responses of positive target cue words and negative background cue words is used to characterize the prior confidence level that the pixel belongs to the target greenhouse. This region selection condition is used to determine which pixels participate in the prior consistency loss calculation, and is not a separate loss function. In the high-confidence region... Internally, based on the prior probability response of text words. As a soft objective, the greenhouse prospect prediction results of the main segmentation branch are constrained, resulting in the prior consistency loss:
[0124] (9)
[0125] In equation (9), This represents the unnormalized prediction score of the binary cross-entropy. Represents pixels The unnormalized prediction score of the greenhouse prospect at the main branch. This represents the number of pixels within the high-confidence region. This loss is used to ensure that the model aligns with the text's lexical probability prior in the high-confidence prior region, thereby enhancing its response to the target region of the greenhouse.
[0126] Furthermore, if the model includes a text semantic auxiliary branch, this auxiliary branch outputs a pixel-level response corresponding to the prompt word, denoted as... In the high confidence region Internally, the prior response is also based on the probability of the text's vocabulary. As a soft target, the semantic prior soft supervision loss formula is as follows:
[0127] (10)
[0128] In equation (10), Indicates text semantic auxiliary branches at the pixel level The unnormalized prediction score for the corresponding greenhouse prompt. This item is used to improve the consistency between the text semantic auxiliary branch and the text prior response.
[0129] To prevent auxiliary branches from interfering with the main split branch in error response areas, this embodiment further sets up a consistency region. A pixel is only used for distillation consistency constraints when the output probability of the text semantic auxiliary branch is sufficiently close to the prior response of the text lexical probability; in consistent regions... Internally, the output probability of the text semantic auxiliary branch is used as a soft objective to constrain the foreground probability of the greenhouse in the main segmentation branch, resulting in the distillation consistency loss:
[0130] (11)
[0131] In equation (11), This represents the Sigmoid function. Indicates the consistency threshold. This represents the set of pixels whose auxiliary branch matches the prior probabilistic responses of the text words. This set is used to filter reliable pixels and avoid passing erroneous auxiliary responses to the main segmentation branch. This represents the distillation consistency loss, used to maintain consistency between the main segmentation branch and the textual semantic auxiliary branch within a reliable region.
[0132] Finally, the total loss function is expressed as:
[0133] (12)
[0134] In equation (12), Losses due to manual labeling supervision The prior consistency loss between the main segmentation branch and the prior probability of the text words. The semantic prior soft-supervised loss is used to bridge the semantic auxiliary branch of the text and the text lexical probability prior. To account for the distillation consistency loss from the auxiliary branch to the main split branch, , and These are the weighting coefficients for the corresponding loss terms.
[0135] Step S6: Train and optimize the parameters of the remote sensing extraction model for heterogeneous agricultural greenhouse targets based on the training set and validation set.
[0136] In this step, the training set partitioned in step S1 is used to optimize the parameters of the heterogeneous agricultural greenhouse target remote sensing extraction model. During training, RGB remote sensing images are input into the model. After backbone visual feature extraction, DINOv3 spatial structure guidance, CLIP text vocabulary probability prior guidance, global and local feature generation, and bidirectional interactive attention fusion, a greenhouse segmentation probability map is output. Subsequently, the multi-source prior consistency loss function constructed in step S5 is used to calculate the error between the model prediction results and the manual labels, text vocabulary probability prior responses, and soft labels, and the model parameters are updated through backpropagation.
[0137] During training, the manual labeling supervision loss ensures the model learns manually labeled greenhouse regions, the text lexical probability prior consistency loss enhances the model's response to high-confidence greenhouse prior regions, the auxiliary consistency loss constrains the output of the text semantic auxiliary branch, and the distillation consistency loss transmits text semantic auxiliary information to the main segmentation branch within reliable regions. Through joint optimization of multiple losses, the model can simultaneously utilize manually labeled information and external text lexical probability prior information, improving its ability to extract greenhouse targets under different materials, shapes, and complex background conditions.
[0138] Step S7: Input the RGB remote sensing images and corresponding dense visual feature priors and CLIP text word probability priors from the dataset of the area to be tested into the trained heterogeneous agricultural greenhouse target remote sensing extraction model to obtain the category prediction probability, and then obtain the binary segmentation result of the heterogeneous greenhouse area.
[0139] Based on the same approach, this invention also provides a heterogeneous agricultural greenhouse remote sensing extraction system based on vision-language collaboration. The system includes: a label data acquisition module, a dense visual feature prior generation module, a CLIP text vocabulary probability prior generation module, a model building module, a loss function construction module, a model training module, and a greenhouse image extraction module; wherein...
[0140] The tag data acquisition module is used to acquire RGB remote sensing images with heterogeneous greenhouse artificial tags and divide them into training set, validation set and test set;
[0141] The dense visual feature prior generation module is used to generate corresponding dense visual features based on RGB remote sensing images using a visual basic model, as dense visual feature priors.
[0142] The CLIP text vocabulary probability prior generation module is used to generate CLIP text vocabulary probability priors guided by text semantics based on RGB remote sensing images using a visual-language model, and to construct corresponding soft tags.
[0143] The model building module is used to construct a heterogeneous agricultural greenhouse target remote sensing extraction model based on a vision-language collaborative network and output the greenhouse segmentation results;
[0144] The loss function construction module is used to construct a multi-source prior consistency loss function for a heterogeneous agricultural greenhouse target remote sensing extraction model.
[0145] The model training module is used to train and optimize the parameters of the remote sensing extraction model for heterogeneous agricultural greenhouse targets based on the training set and validation set.
[0146] The greenhouse image extraction module is used to input the remote sensing images in the dataset of the area to be tested, along with the corresponding dense visual feature priors and CLIP text word probability priors, into the trained heterogeneous agricultural greenhouse target remote sensing extraction model to obtain the category prediction probability, and thus obtain the binary segmentation result of the heterogeneous greenhouse area.
[0147] The system or device for executing the method in this embodiment of the invention can be a terminal or a server. The system includes a processor, a memory, and / or a transceiver, etc., and is connected via a communication bus. Each module can be implemented by a processor, a memory, and / or a transceiver, etc. The processor can be, but is not limited to, one or more microprocessors (MPUs), central processing units (CPUs), network processors (NPs), digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), and other programmable logic devices, discrete gates, transistor logic devices, discrete hardware components, etc., or can be configured to implement one or more integrated circuits of this invention. The processor can perform various functions by running or executing software programs in the memory and calling data in the memory. The memory includes Random Access Memory (RAM), Read-Only Memory (ROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Compact Disc Read-Only Memory (CD-ROM), and / or Non-Volatile Memory (NVM), etc. The transceiver is used to communicate with network devices or with terminal devices, and includes a receiver and a transmitter. The memory and transceiver can be integrated with the processor or exist independently.
[0148] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.
[0149] It should also be noted that the heterogeneous agricultural greenhouse remote sensing extraction system based on vision-language collaboration described in this embodiment corresponds to the heterogeneous agricultural greenhouse remote sensing extraction method based on vision-language collaboration. The description and limitations of the method also apply to the system, and will not be repeated here.
[0150] The remote sensing extraction method and system for heterogeneous agricultural greenhouses based on vision-language collaboration described in this invention were applied to the remote sensing extraction process of heterogeneous agricultural greenhouse targets in a certain region. Two typical plastic greenhouse distribution areas were selected as study areas to evaluate the applicability of the model under different greenhouse types and background conditions in different regions. The first region has a high degree of greenhouse scale, with significant differences in greenhouse shape and covering materials, and diverse background features. To test the model's cross-regional generalization performance, data from the first region was used as the test set, consisting of 1663 images, all of which were manually annotated to obtain pixel-level ground truth. The second region has a dense distribution of plastic greenhouses with a wide coverage area, including various greenhouse types with different covering materials and structural forms. These greenhouses often coexist closely with roads, buildings, bare land, and plastic film, making the scene complex and representative. Data from the second region was used as the validation set, consisting of 1615 images, all of which were finely annotated manually.
[0151] The data source was Google Earth Level 18 imagery with a spatial resolution of 0.5m. The parameters were set as follows: The experiment was conducted on a Windows platform, implemented using Python 3.11.7 (Anaconda) and PyTorch 2.2.2 (CUDA 11.8), and trained on an NVIDIA A100 GPU. The task was set as binary semantic segmentation (2 classes). Training lasted 60 epochs, using AdamW as the optimizer, with an initial learning rate of 3×10⁻⁻⁻⁶. 4 The weight decays to 0.05.
[0152] To comprehensively evaluate the performance of the proposed model, overall precision (OA), recall, F1 score, and mean intersection-union ratio (MIOU) were used to evaluate the experimental results (MIOU was used to evaluate the VLHF model's ability to extract greenhouses in complex scenes, while also considering the segmentation accuracy of the background and greenhouses). OA represents the proportion of correctly predicted samples out of the total number of samples; recall represents the proportion of all positive samples that were correctly predicted as positive; the F1 score is the harmonic mean of precision and recall; and MIOU represents the ratio of the intersection and union of the actual segmented regions and the model-predicted segmented regions. It is the number of categories. Indicates the first There are several categories. The calculation formulas for each evaluation indicator are as follows:
[0153] (13)
[0154] (14)
[0155] (15)
[0156] (16)
[0157] In equations (13) to (16), P is the number of positive samples; N is the number of negative samples; TP is the number of correctly predicted positive (true positive) samples; FP is the number of incorrectly predicted positive (false positive) samples; TN is the number of correctly predicted negative (true negative) samples; and FN is the number of incorrectly predicted negative (false negative) samples.
[0158] The comparison results with classic networks are as follows:
[0159] To verify the effectiveness of the proposed VLHF-Net in the task of extracting heterogeneous plastic greenhouses, three classic and representative semantic segmentation networks were selected as baselines for comparison: U-Net, ResUNet, and DeepLabv3 (all of which are widely used general segmentation frameworks in remote sensing scenarios). These methods represent three typical technical approaches: "encoder-decoder," "residual-enhanced encoder-decoder," and "dilated convolutional multi-scale contextual modeling," respectively, and can comprehensively reflect the capabilities of traditional segmentation models.
[0160] U-Net employs a symmetrical encoder-decoder structure and fuses shallow details with deep semantics through skip connections, offering advantages such as simplicity, stable training, and strong preservation of boundary details. In scenarios with limited sample size or relatively regular target shapes, U-Net often provides reliable baseline performance. ResUNet introduces residual units into the U-Net framework to enhance feature representation and gradient propagation capabilities, thereby supporting deeper networks or stronger representation learning and improving convergence stability in complex backgrounds. This structure is generally more robust to texture variations and local occlusion, making it suitable as a "structure-enhanced" traditional baseline. DeepLabv3, with dilated convolutions and ASPP (Atrous Spatial Pyramid Pooling) at its core, models multi-scale contexts without significantly reducing resolution, exhibiting strong long-range receptive field and scale adaptation capabilities. In remote sensing imagery, DeepLabv3 often demonstrates good overall consistency for large-scale continuous features or targets with significant scale variations.
[0161] Table 1. Comparison results of VLHFNet and classic networks on the second test set.
[0162]
[0163] As shown in Table 1, VLHFNet achieves the best results in all four metrics: OA, Recall, F1, and mIoU (99.11%, 99.07%, 98.98%, and 98.20%, respectively). Compared with U-Net, ResUNet, and DeepLabv3, overall accuracy and region consistency are improved simultaneously, especially the mIoU improvement, indicating a higher degree of spatial overlap between the predicted region and the ground truth annotation, and effective suppression of intra- and inter-class confusion. Compared with DeepLabv3, VLHFNet improves OA, Recall, F1, and mIoU by 2.71, 2.62, 3.10, and 5.27 percentage points, respectively, demonstrating that it further improves the reliability and overall consistency of predictions while maintaining detection capabilities.
[0164] Combination Figure 3The visualization results show that VLHFNet demonstrates more stable target representation capabilities in typical difficult regions. Its predictions are more consistent with the labeled boundary locations, exhibiting higher integrity of target regions and fewer responses in non-target regions. Compared to U-Net and ResUNet, which are more prone to local false responses in complex backgrounds, VLHFNet can more effectively distinguish targets from similar background textures. Compared to DeepLabv3, which is more prone to spatial expansion or contraction near boundaries, VLHFNet's boundary localization is more accurate, and the region shape is preserved more reasonably. This is consistent with the simultaneous improvement in Recall and F1 in Table 1, indicating that the model reduces false positives while also controlling false negatives, resulting in higher mIoU and overall accuracy.
[0165] The results of comparison with other advanced greenhouse extraction networks are as follows:
[0166] To further verify the effectiveness of the proposed method, this embodiment compares VLHFNet with representative advanced greenhouse extraction networks developed in recent years. Addressing the issues of similar spectral characteristics between plastic greenhouses and mulched farmland, easily confused boundaries, and insufficient cross-regional generalization ability, some researchers have proposed the APC-Net semantic segmentation model based on ultra-high resolution remote sensing imagery. This method combines a DNCNN encoder, a multi-task edge learning decoder, and an L2 regularized transfer learning strategy based on source priors, achieving good results in greenhouse boundary characterization and category differentiation. Other researchers have proposed the VFM-UNet / VFM-ASPP network to address the problems of remote sensing greenhouse segmentation relying on a large number of labeled samples and the difficulty of directly adapting the visual base model to remote sensing scenes. This method extracts multi-level embedded features by freezing the visual base model and combines it with a lightweight convolutional network for feature fusion, reducing the dependence on labeled samples while improving segmentation accuracy, boundary regularity, and model generalization ability.
[0167] Table 2. Ablation Experiment Results for Different Prior Methods
[0168]
[0169] As can be seen from the comparison results in Table 2, the proposed VLHFNet achieved the best performance in all evaluation metrics. Compared with APC-Net, VLHFNet improved OA, Recall, F1, and mIoU by 3.22, 6.61, 3.84, and 6.27 percentage points, respectively; compared with VFM-Net, VLHFNet improved OA, Recall, F1, and mIoU by 1.09, 1.17, 1.24, and 2.63 percentage points, respectively. The improvement in mIoU was particularly significant, indicating that VLHFNet not only improves overall classification accuracy but also more effectively improves the cross-union ratio performance in greenhouse areas, reducing misclassification with similar features such as plastic-covered farmland, bare land, and building roofs.
[0170] from Figure 4 As can be seen, APC-Net mainly enhances greenhouse boundary representation through edge learning and transfer regularization, while VFM-Net relies on a visual base model to improve feature generalization under few-sample conditions. In contrast, the proposed VLHFNet can more fully utilize the complementary relationship between optical visual features and heterogeneous auxiliary information, exhibiting more stable performance in target recognition, boundary preservation, and region generalization. Therefore, VLHFNet has a stronger ability to accurately extract greenhouse features in complex agricultural scenarios.
[0171] Table 3. Comparison of ablation methods with different prior techniques
[0172]
[0173] As shown in Table 3, the impact of different prior methods on performance varies significantly. The base network achieves 96.79% OA, 96.04% Recall, 96.31% F1, and 96.69% mIoU, respectively. After adding the DINOv3 prior, OA increases to 97.39%, F1 increases to 96.78%, but Recall decreases to 95.75% and mIoU decreases to 94.74%. This indicates that the DINOv3 prior provides some gain in overall discrimination, but negatively impacts target region consistency and boundary localization, leading to a decrease in spatial overlap with the annotation. After adding the CLIP prior (VLHFNet), all four metrics simultaneously improve to 99.11%, 99.07%, 98.98%, and 98.20%. Compared to the basic network, OA, Recall, F1, and mIoU were improved by 2.32, 3.03, 2.67, and 1.51 percentage points, respectively; compared to the DINOv3 prior scheme, mIoU was improved by 3.46 percentage points, indicating that CLIP prior can simultaneously improve detection capability and regional consistency.
[0174] In challenging regions, the predictions of the base network are more susceptible to interference from background structure and local texture, manifesting as false responses in non-target regions and deviations in target boundary localization. Even with the addition of the DINOv3 prior, some samples still retain a significant number of non-target responses, while the consistency within the target region remains insufficient, consistent with the situation where DINOv3 primarily provides features without offering semantic information. In contrast, with the addition of CLIP semantic prior features, VLHFNet's predictions are more consistent with the annotations, the target region is more complete, there are fewer non-target responses, and the boundary localization is more stable, especially in complex backgrounds and samples with strong similar interference. This aligns with the simultaneous improvement in Recall and mIoU shown in Table 2, indicating that the CLIP prior not only reduces missed detections but also effectively suppresses false detections, resulting in higher overall consistency. Figure 5 As shown.
[0175] The results of ablation comparison of different modules are as follows:
[0176] Table 4 Ablation Results of Key Modules
[0177]
[0178] As shown in Table 4, the ablation results indicate a progressive relationship in performance improvement among the modules. To ensure fair comparison, the base network incorporates the same prior information at the input, removing only key modules. When no interaction mechanism is introduced, prior information participates in feature extraction as a common input. The base network's OA, Recall, F1, and mIoU are 96.79%, 96.05%, 96.31%, and 93.69%, respectively. After adding the global-local module, Recall increases to 97.06%, and mIoU increases to 94.74%, indicating that this module enhances the detection capability of target regions and improves regional consistency. Further addition of the fusion module increases OA to 97.74%, Recall to 98.02%, F1 to 97.43%, and mIoU to 94.98%, demonstrating that the fusion mechanism helps align and complement prior information and visual representations, thereby reducing class confusion and stabilizing the spatial distribution of predicted regions. After introducing the dynamic loss block, the four metrics were significantly improved to 99.11%, 99.07%, 98.98%, and 98.20%, respectively. Among them, mIoU was improved by 4.51% compared to the basic network, which shows that the optimization of difficult-to-distinguish regions and boundary neighborhoods was the most complete.
[0179] Figure 6 The visualization results are consistent with the trends described above. For example... Figure 6As shown, in the typical difficult regions indicated by the yellow box, the base network is more prone to non-target response remnants and insufficient target region characterization. Adding the global-local module enhances the continuity and local consistency of the target region, alleviating missed detections. Adding the fusion module further improves the spatial consistency between the predicted region and the annotation, and non-target responses are more effectively suppressed. Introducing the dynamic loss block results in more stable boundary localization, more reasonable region shapes, and a significant reduction in erroneous responses in small-scale structures and complex backgrounds. Corresponding to the improvement in Recall and mIoU in Table 3, this indicates that the module is more effective in constraining training for difficult samples and key regions.
[0180] This embodiment addresses the problem of large intra-class differences and strong inter-class confusion in heterogeneous plastic greenhouses under conditions of cross-regional, cross-material, and complex backgrounds. It proposes a visual-language collaborative segmentation framework, VLHF-Net, which integrates backbone visual representations and semantic prior constraints in a unified network to improve segmentation robustness and generalization ability. Experiments were conducted in two typical greenhouse distribution areas, using pixel-level manually labeled data constructed from 0.5m resolution Google Earth imagery to evaluate the model's applicability under different regional backgrounds and greenhouse morphological differences.
[0181] Comparative experiments show that VLHF-Net achieves superior overall performance compared to the classic baselines U-Net, ResUNet, and DeepLabv3, with optimal levels in OA, Recall, F1, and mIoU. This indicates that the method effectively suppresses false detections while reducing missed detections and significantly improves the spatial consistency between the predicted region and the ground truth annotation. Visualization results further validate this conclusion, demonstrating that the model maintains more stable target characterization capabilities in similar interference neighborhoods such as roads, buildings, bare land, and plastic film, with more accurate boundary localization and a more complete target region.
[0182] Ablation experiments further explain the source of performance gains. Comparison of different prior methods shows that the DINOV3 prior helps with overall discrimination but weakens region consistency. Introducing the CLIP semantic prior, however, leads to simultaneous improvements in all four metrics, indicating that semantic-level prior constraints are more crucial for mitigating similarity interference. Module ablation results show that global-local representation, fusion modules, and dynamic loss blocks contribute progressively, with dynamic loss being the most significant for optimizing difficult-to-distinguish regions and boundary neighborhoods, thus resulting in higher mIoU and overall stability.
[0183] In summary, VLHF-Net achieves high-precision and robust segmentation of heterogeneous plastic greenhouses through a vision-language semantic prior guidance and interactive fusion mechanism, providing effective technical support for rapid extraction and refined monitoring of greenhouses under high-resolution remote sensing imagery. Future validation can be expanded to more regions and with a wider range of material types, and text prompts and prior calibration strategies can be further optimized to improve generalization capabilities in larger-scale applications.
[0184] Therefore, the remote sensing extraction method and system for heterogeneous agricultural greenhouses based on vision-language collaboration provided in this invention enhances generalization and robustness under cross-domain conditions by jointly modeling local discriminative features and scene-level contextual relationships in a unified network. A two-level text lexical mechanism of spatial prior constraints and semantic conditional discrimination is constructed: firstly, a soft prior map obtained through text guidance provides regional constraints to reduce background interference and search uncertainty; then, a pixel-level similarity discrimination module with text conditionalization is introduced, using text embedding as a category prototype to complete pixel-level semantic decision-making and improve inter-class separability; a dynamic bidirectional interactive fusion module is designed, achieving adaptive collaborative fusion of cross-branch features through bidirectional attention interaction and gating weighting, thereby improving representation consistency and discrimination stability.
[0185] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0186] The above description is merely a preferred embodiment of the present invention and an explanation of the technical principles employed, and is not intended to limit the scope of the claimed invention, but merely to illustrate preferred embodiments of the invention. Those skilled in the art should understand that the scope of the invention is not limited to the specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the inventive concept. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
Claims
1. A remote sensing extraction method for heterogeneous agricultural greenhouses based on vision-language collaboration, characterized in that, The method includes the following steps: Step S1: Obtain RGB remote sensing images with artificial labels for heterogeneous greenhouses and divide them into training set, validation set and test set; Step S2: Based on RGB remote sensing images, generate corresponding dense visual features using a visual basic model as priors for dense visual features; Step S3: Based on RGB remote sensing imagery, generate CLIP text vocabulary probability priors guided by text semantics using a visual-language model, and construct corresponding soft labels; Step S4: Construct a heterogeneous agricultural greenhouse target remote sensing extraction model based on a vision-language collaborative network, and output the greenhouse segmentation results; Step S5: Construct the multi-source prior consistency loss function for the heterogeneous agricultural greenhouse target remote sensing extraction model; Step S6: Train and optimize the parameters of the remote sensing extraction model for heterogeneous agricultural greenhouse targets based on the training set and validation set; Step S7: Input the remote sensing images and corresponding dense visual feature priors and CLIP text word probability priors from the dataset of the area to be tested into the trained heterogeneous agricultural greenhouse target remote sensing extraction model to obtain the category prediction probability, and then obtain the binary segmentation result of the heterogeneous greenhouse area.
2. The method according to claim 1, characterized in that, The heterogeneous greenhouses include greenhouse targets with different covering materials, different greenhouse colors, different structural forms, and different spatial distribution methods.
3. The method according to claim 1, characterized in that, In step S2, the visual base model adopts the DINOv3 model; RGB remote sensing images are input into the DINOv3 model to obtain block-level dense visual features.
4. The method according to claim 1, characterized in that, The visual-language model described in step S3 uses the contrastive language-image pre-trained model CLIP. Generate CLIP text lexical probabilistic priors guided by text semantics, including: Step S31: Construct a dynamic semantic prompt pool consisting of basic greenhouse vocabulary, expert experience vocabulary, and regional network retrieval information, and divide the prompt words into positive target prompt words and negative background prompt words; Step S32: Input each positive target cue word and negative background cue word into the CLIP text encoder to obtain the text embedding, and then perform normalization processing: ,(2) ,(3) In equations (2) and (3), For the i-th positive prompt text, For CLIP text encoder, This is the normalized text embedding vector; For the j-th negative prompt text, This is the normalized negative text embedding vector; Step S33: Input RGB remote sensing image According to multi-scale spatial grid The image is divided into regions at different scales; for each scale... The next Image regions Normalized image embeddings are obtained using the CLIP image encoder: ,(4) In equation (4), This indicates the CLIP image encoder. This represents the normalized image region embedding; Aggregate similarity response between image regions and positive target cues and negative background cues , They are represented as follows: ,(5) In equation (5), The cosine similarity between the image region embedding and the positive target cue text embedding is represented. The cosine similarity between the image region embedding and the negative background cue word text embedding is used to represent the cosine similarity between the image region embedding and the negative background cue word text embedding. This represents an aggregation operation that evaluates the similarity responses to multiple positive target prompts. This represents an aggregation operation that evaluates the similarity responses to multiple negative background cue words. The scale is obtained by the difference between positive and negative responses. Next The target semantic prior response of each image region: ,(6) In equation (6), The text lexical probabilistic prior response of the image region to the greenhouse target; Step S34 involves spatial reconstruction, weighted fusion, and normalization of the target semantic prior responses at different scales to obtain text semantic guidance target semantic probability prior information corresponding to the original image space. .
5. The method according to claim 4, characterized in that, Step S34 further includes: constructing soft labels based on the prior response, which are used in subsequent training phases to jointly constrain the model with artificial labels, so that the model can enhance its ability to suppress similar backgrounds while learning the features of the greenhouse target.
6. The method according to claim 1, characterized in that, The heterogeneous agricultural greenhouse target remote sensing extraction model described in step S4 takes RGB remote sensing imagery, dense visual feature priors, and text semantic-guided CLIP text vocabulary probability priors as input, and outputs greenhouse segmentation results. The model obtains greenhouse segmentation results by extracting the backbone visual features of RGB remote sensing imagery, dense visual feature priors, and text semantic-guided CLIP text vocabulary probability priors, performing structure-guided mapping, semantic prior mapping, global and local feature generation, and bidirectional interactive attention fusion.
7. The method according to claim 1, characterized in that, The loss function described in step S5 includes manual label supervision loss, semantic prior soft supervision loss, prior consistency loss, and distillation consistency loss.
8. The method according to claim 7, characterized in that, The loss of manual label supervision The formula is as follows: ,(7) In equation (7), Represents cross-entropy loss, Indicates pixel position, and These represent the image height and width, respectively. Indicates the segmentation model at pixel location The output is the category prediction score; Indicates pixel position The corresponding manually labeled category tag; Define the high confidence region as: ,(8) In equation (8), This indicates a priori high-confidence regions. Indicates the confidence margin threshold; Indicates pixel position The prior probability value generated by the similarity response of positive target cue words and negative background cue words is used to characterize the prior confidence level that the pixel belongs to the target greenhouse; Introduce a prior consistency loss within the high-confidence prior region: ,(9) In equation (9), This represents the unnormalized prediction score of the binary cross-entropy. Represents pixels The unnormalized prediction score of the greenhouse prospect at the main branch. This indicates the number of pixels within the high confidence region; In high confidence region Internally, based on the prior probability of textual vocabulary responses. As a soft objective, the semantic prior soft supervision loss is defined as follows: ,(10) In equation (10), Indicates text semantic auxiliary branches at the pixel level The unnormalized prediction score of the corresponding greenhouse prompt; The formula for the distillation consistency loss is as follows: ,(11) In equation (11), This represents the Sigmoid function. Indicates the consistency threshold. The set of pixels representing the auxiliary branch that is consistent with the prior probability response of the text words; The total loss function is expressed as: ,(12) In equation (12), Losses due to manual labeling supervision The prior consistency loss between the main segmentation branch and the prior probability of the text words. The semantic prior soft-supervised loss is used to bridge the semantic auxiliary branch of the text and the text lexical probability prior. To account for the distillation consistency loss from the auxiliary branch to the main split branch, , and These are the weighting coefficients for the corresponding loss terms.
9. The method according to claim 1, characterized in that, In step S6, during the training process, the manual labeling supervision loss is used to ensure that the model learns the manually labeled greenhouse areas, the text vocabulary probability prior consistency loss is used to enhance the model's response to high-confidence greenhouse prior areas, the auxiliary consistency loss is used to constrain the output of the text semantic auxiliary branch, and the distillation consistency loss is used to pass the text semantic auxiliary information to the main segmentation branch within the reliable region. Through the joint optimization of multiple losses, the model can simultaneously utilize manual labeling information and external text vocabulary probability prior information to improve the ability to extract greenhouse targets under different materials, different shapes, and complex background conditions.
10. A heterogeneous agricultural greenhouse remote sensing extraction system based on vision-language collaboration, characterized in that, The system includes: a label data acquisition module, a dense visual feature prior generation module, a CLIP text vocabulary probability prior generation module, a model building module, a loss function construction module, a model training module, and a greenhouse image extraction module; among which... The tag data acquisition module is used to acquire RGB remote sensing images with heterogeneous greenhouse artificial tags and divide them into training set, validation set and test set; The dense visual feature prior generation module is used to generate corresponding dense visual features based on RGB remote sensing images using a visual basic model, as dense visual feature priors. The CLIP text vocabulary probability prior generation module is used to generate CLIP text vocabulary probability priors guided by text semantics based on RGB remote sensing images using a visual-language model, and to construct corresponding soft tags. The model building module is used to construct a heterogeneous agricultural greenhouse target remote sensing extraction model based on a vision-language collaborative network and output the greenhouse segmentation results; The loss function construction module is used to construct a multi-source prior consistency loss function for a heterogeneous agricultural greenhouse target remote sensing extraction model. The model training module is used to train and optimize the parameters of the remote sensing extraction model for heterogeneous agricultural greenhouse targets based on the training set and validation set. The greenhouse image extraction module is used to input the remote sensing images in the dataset of the area to be tested, along with the corresponding dense visual feature priors and CLIP text word probability priors, into the trained heterogeneous agricultural greenhouse target remote sensing extraction model to obtain the category prediction probability, and thus obtain the binary segmentation result of the heterogeneous greenhouse area.