A method and system for extracting images of plastic greenhouses based on multi-feature deep fusion

Through the multi-feature deep fusion method and boundary learning neural network, the problems of low extraction accuracy and unclear boundaries in the existing technology are solved, and a higher precision and stable extraction of plastic greenhouses is achieved.

CN119991471BActive Publication Date: 2025-07-11YUNNAN NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510457619.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-11
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

The existing plastic greenhouse extraction method is difficult to accurately extract the boundary information of plastic greenhouses when the spectral characteristics are small and the remote sensing image background is complex. The deep learning network leads to boundary blur and contour sticking, which affects the extraction accuracy and stability.

Method used

The multi-feature deep fusion method is adopted to build a multi-branch architecture, combining multi-scale feature learning and boundary learning neural networks, to guide boundary information fusion and improve the recognition ability of plastic greenhouses.

Benefits of technology

It significantly improves the accuracy and stability of plastic greenhouse extraction, enhances the generalization ability of the model, solves the problems of boundary blur and contour sticking, and improves the accuracy and consistency of the extraction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991471B_ABST
    Figure CN119991471B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for extracting plastic greenhouse images based on multi-feature deep fusion, belonging to the field of machine vision and image extraction. The method determines the research area and year and obtains remote sensing images and multi-spectral remote sensing data, and extracts remote sensing image samples and index samples with the same resolution; then constructs a greenhouse image extraction model and a boundary learning model; downsamples the index samples and inputs them into the multi-spectral branch of the greenhouse image extraction model to obtain multi-spectral features; downsamples the remote sensing image samples and inputs them into the five convolutional layers of the semantic branch; respectively obtains the first and second semantic features based on the second and fifth convolutional layers; inputs the remote sensing image samples into the detail branch to extract the detail features of the plastic greenhouse; fuses the fused edge map, multi-spectral features, first semantic feature, second semantic feature and detail features to extract the plastic greenhouse image. The present invention solves the problems of blurred boundaries and contour adhesion, and improves the identification ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of machine vision and image processing, and particularly relates to a method and system for extracting plastic greenhouse images based on multi-feature depth fusion. Background Art

[0002] A plastic greenhouse is an agricultural shed mainly made of plastic materials, usually composed of plastic films, plastic pipes, steel frames and other materials. It has the characteristics of being light, durable, and having good light transmittance, providing good environmental conditions for the growth of crops and playing an important role in the development of modern agriculture. Although plastic greenhouses have improved the efficiency of agricultural production, the plastic films used in plastic greenhouses usually need to be replaced every 1-2 years. If the aged plastic films are not properly recycled and treated, they will be discarded in farmland or the nearby environment, causing white pollution. In addition, due to their extremely long degradation period, plastic films will persist in the soil and water bodies, posing a serious threat to the ecosystem. These non-degradable plastic films will gradually split and degrade into microplastic particles, which can penetrate into the soil, interfere with the normal growth and development of crop roots, and then enter the food chain through the absorption and accumulation of crops, posing a potential threat to the health of humans and animals. Therefore, accurately obtaining the area and spatial layout of agricultural plastic greenhouses is crucial for agricultural pollution prevention and efficient management, and helps to formulate scientific and reasonable agricultural policies to ensure the safety of food supply.

[0003] Relevant data such as the area and spatial layout of agricultural plastic greenhouses are generally obtained through manual on-site surveys. However, such methods are time-consuming and laborious, and have low timeliness. In recent years, with the continuous development and maturity of satellite remote sensing technology, some researchers have also extracted plastic greenhouses based on satellite remote sensing data, which has the advantages of wide coverage, short revisit period, and low data acquisition cost.

[0004] In the prior art, the methods for extracting plastic greenhouses based on remote sensing data include three categories: the first is the method for extracting plastic greenhouses based on spectral index threshold segmentation, the second is the method for extracting plastic greenhouses based on traditional machine learning, and the third is the method for mapping plastic greenhouse extraction based on deep learning.

[0005] Among them, the first type of method uses mathematical methods to expand the difference between plastic greenhouses and other ground objects based on the difference in spectral reflectance curves between plastic greenhouses and other ground objects, so that the plastic greenhouse to be studied can obtain the maximum brightness enhancement on the generated index image, while other ground objects are generally suppressed, and the extraction of plastic greenhouses is completed by setting appropriate segmentation thresholds. For example, common plastic greenhouse extraction indexes include greenhouse vegetable land extraction index (greenhousevegetable land index, VI), plastic-mulched land cover index (plastic-mulched landcover index, PMLI), advanced plastic greenhouse index (advanced plastic greenhouse index, APGI), etc. Although the principle of plastic greenhouse extraction method based on spectral index threshold segmentation is simple, easy to understand and calculate. However, this type of method only relies on spectral feature differences to extract plastic greenhouses. Plastic greenhouses are made of various materials, and the spectral features within the class vary greatly, making it difficult to construct a spectral index with strong universality.

[0006] The second type of method uses traditional machine learning algorithms to extract plastic greenhouses based on the differences in texture, geometry, spectrum and other features between plastic greenhouses and other landforms. For example, based on 37 features such as brightness and density, the random forest method of sample optimization selection is used to extract plastic greenhouses. Another example is the classification effects of three machine learning methods: random forest, CART decision tree and support vector machine for the greenhouse extraction task of GF-2 remote sensing image, and the conclusion is that random forest classification has the best effect. Compared with the plastic greenhouse extraction method based on spectral index threshold segmentation, this method can learn the feature differences between the machine learning target and the background without a fixed threshold. However, this type of method can only use shallow feature differences such as spatial texture and spectrum to extract plastic greenhouses, and lacks in-depth analysis of feature engineering.

[0007] The third method uses neural networks to mine the deep feature differences between plastic greenhouses and other landforms, and uses the deep features to extract plastic greenhouses. For example, five fully convolutional neural networks of different scales are constructed through different feature depth fusion methods, and greenhouses and mulch fields are extracted based on these five models; another example is the use of the ENVINet5 deep learning architecture to extract sparsely distributed plastic greenhouses in high-resolution remote sensing images; another example is the use of the SSD network to make greenhouse labels based on GF-2 satellite images, and the overall accuracy of the extracted random sample points is 84.5%, and the Kappa coefficient is 0.831. However, due to multiple convolution and downsampling operations, the neural network based on the codec structure will cause the spatial resolution of remote sensing data to decrease, resulting in the loss of boundary information of the extracted agricultural plastic greenhouses, and the phenomenon of adhesion and fusion between greenhouse boundaries is prone to occur.

[0008] It can be seen that the existing spectral index threshold segmentation method faces challenges such as the diversity, complexity of plastic greenhouses and small differences in spectral characteristics when extracting plastic greenhouses, making it difficult to obtain a universal threshold to completely extract the detailed information of the greenhouse contour and difficult to meet the actual needs. In addition, the background of remote sensing images is complex, and traditional machine learning methods rely on artificial feature engineering, making it difficult to mine their deep features and vulnerable to interference from similar ground objects, resulting in unstable extraction results and the "salt and pepper" phenomenon. Summary of the Invention

[0009] In view of the above defects or deficiencies in the prior art, the present invention aims to provide a method and system for extracting plastic greenhouse images based on multi-feature deep fusion. By constructing a multi-branch structure, deep fusion of multi-scale features is carried out, and an independent boundary learning neural network is introduced to guide the fusion mechanism of boundary information, so as to solve the problems such as blurred boundaries and contour adhesion often encountered in the plastic greenhouse extraction task by deep learning convolutional neural networks, and improve the recognition ability of plastic greenhouses.

[0010] To achieve the above object, the embodiments of the present invention adopt the following technical solutions:

[0011] In the first aspect, the embodiments of the present invention provide a method for extracting plastic greenhouse images based on multi-feature deep fusion, including the following steps:

[0012] Step S1, determine the research area and year, and obtain high-resolution remote sensing image data and multi-spectral remote sensing data corresponding to the time in the research area;

[0013] Step S2, based on the high-resolution remote sensing image data, extract remote sensing image samples; resample the multi-spectral remote sensing data and calculate index samples so that the resolution of the calculated index samples after resampling is the same as that of the remote sensing image samples;

[0014] Step S3, build a backbone network based on the deep convolutional neural network VGG-16; based on the backbone network, create a multi-spectral branch, a semantic branch and a detail branch using an encoder-decoder structure to build a greenhouse image extraction model;

[0015] Step S4, build a boundary learning model using an encoder-decoder structure based on a deep learning neural network, where the boundary learning model includes M convolutional blocks, an upsampling module and a splicing module;

[0016] Step S5, input the remote sensing image samples of the original size into the boundary learning model. The remote sensing image samples introduce a model training mechanism guided by boundary information through M convolutional blocks, extract M feature maps with different scales, and then through upsampling and splicing, extract edge information from the M feature maps with different scales, perform deep feature fusion on the edge information, and finally generate a fused edge map;

[0017] Step S6: Input the index sample into the first downsampling module of the greenhouse image extraction model, output the index sample with 1 / 64 low resolution and input it into the multispectral branch to obtain multispectral features; input the remote sensing image sample into the second downsampling module of the greenhouse image extraction model, output the image sample with 1 / 4 low resolution, and sequentially input it into the five convolutional layers of the semantic branch; obtain the first semantic feature based on the second convolutional layer and the second semantic feature based on the fifth convolutional layer; input the remote sensing image sample into the detail branch to extract the detail features of the plastic greenhouse.

[0018] Step S7: Use the attention perception mechanism to fuse the fused edge map, multispectral features, first semantic feature, second semantic feature and detail features to extract the plastic greenhouse image.

[0019] As a preferred embodiment of the present invention, step S7 specifically includes:

[0020] Step S71: Input the fused edge map, multispectral features and second semantic feature into the first attention fusion module at the same time to fuse the fused edge map, multispectral features and second semantic feature to obtain the first fused image feature map;

[0021] Step S72: After the first fused image feature map is upsampled 8 times by the first upsampling module, input it into the second attention fusion module together with the first semantic feature to obtain the second fused image feature map;

[0022] Step S73: After the second fused image feature map is upsampled 4 times by the second upsampling module, input it into the third attention fusion module together with the detail features to obtain the third fused image feature map;

[0023] Step S74: After the third fused image feature map is upsampled 2 times by the third upsampling module, output the fused plastic greenhouse image map.

[0024] As a preferred embodiment of the present invention, the high-resolution remote sensing image data and multispectral remote sensing data in step S1 come from the Google Earth platform high-resolution RGB image - level 18 and Sentinel-2 data.

[0025] As a preferred embodiment of the present invention, in step S2, the size of the extracted remote sensing image sample is 256×256 or 512×512.

[0026] As a preferred embodiment of the present invention, when resampling the multispectral remote sensing data in step S2, the exponential threshold method is used; the index uses the new greenhouse index APGI , and the index calculation formula is as follows:

[0027] (1)

[0028] In formula (1), is the wavelength of the aerosol band, is the wavelength of the red band, is the wavelength of the near-infrared band, is the wavelength of the short-wave infrared band, APGI represents the new plastic greenhouse index;

[0029] When APGI is greater than the preset index threshold, it is confirmed as a plastic greenhouse sample.

[0030] As a preferred embodiment of the present invention, in step S3, the greenhouse image extraction model further includes: a first downsampling module before the multispectral branch and a second downsampling module before the semantic branch; after the three branches, it includes a first attention module, a first upsampling module, a second attention module, a second upsampling module, a third attention module and a third upsampling module.

[0031] As a preferred embodiment of the present invention, the multispectral branch includes a convolutional layer, a batch normalization layer and a ReLU activation function; the semantic branch includes several convolutional layers; the detail branch includes two convolutional layers.

[0032] As a preferred embodiment of the present invention, in the boundary learning model constructed in step S5, the encoder is composed of five convolutional blocks; among them, the first convolutional block includes a 1x1 convolutional layer, two 3x3 convolutional layers and a pooling layer; the second convolutional block includes a 1x1 convolutional layer, two 3x3 convolutional layers and a pooling layer; the third convolutional block includes a 1x1 convolutional layer, three 3x3 convolutional layers and a pooling layer; the fourth convolutional block includes a 1x1 convolutional layer, three 3x3 convolutional layers, and the fifth convolutional block includes a 1x1 convolutional layer, three 3x3 convolutional layers; each convolutional block obtains feature maps of different scales.

[0033] As a preferred embodiment of the present invention, the decoder is composed of an upsampling module and a splicing module, which are used to fuse M different scales of feature maps. At this time, the output of the current convolutional block is connected to the output of the previous convolutional block through the splicing module to form a feature stack, and finally a fused edge map is generated.

[0034] In a second aspect, an embodiment of the present invention further provides a plastic greenhouse image extraction system based on multi-feature deep fusion. The system includes: a data acquisition module, a sample generation module, a greenhouse extraction model construction module, a boundary learning model construction module, a feature deep fusion module and a result output module; among them,

[0035] The data acquisition module is used to determine the research area and year, and obtain high-resolution remote sensing image data and multispectral remote sensing data corresponding to the time within the research area;

[0036] The sample generation module is used to extract remote sensing image samples based on the high-resolution remote sensing image data; resample the multispectral remote sensing data and calculate index samples, so that the resolution of the calculated index samples after resampling is consistent with that of the remote sensing image samples;

[0037] The greenhouse extraction model construction module is used to build a backbone network based on the deep convolutional neural network VGG-16; based on the backbone network, create a detail branch, a multispectral branch and a semantic branch using an encoder-decoder structure to construct a greenhouse image extraction model; among them, the multispectral branch is used to obtain multispectral features based on the index samples; the semantic branch is used to obtain semantic features based on the remote sensing image samples after 1 / 4 downsampling; the detail branch is used to extract the detail features of the plastic greenhouse based on the remote sensing image samples of the original size;

[0038] The greenhouse image extraction module further includes: a first downsampling module is provided before the multispectral branch, and a second downsampling module is provided before the semantic branch; after the three branches, a first attention module, a first upsampling module, a second attention module, a second upsampling module, a third attention module and a third upsampling module are provided; the semantic branch includes five convolutional layers, where the second convolutional layer is used to obtain the first semantic feature, and the fifth convolutional layer is used to obtain the second semantic feature;

[0039] The boundary learning model construction module is used to construct a boundary learning model using an encoder-decoder structure based on a deep learning neural network. The boundary learning model includes M convolutional blocks, which are used to introduce a model training mechanism guided by boundary information, and each convolutional block can obtain M different-scale feature maps; the boundary learning module further includes an upsampling module and a splicing module, which are used to fuse the M different-scale feature maps to generate a fused edge map;

[0040] The feature depth fusion module is used to fuse the fused edge map, multispectral features, first semantic features, second semantic features and detail features using an attention fusion mechanism, extract the plastic greenhouse image, and send it to the result output module;

[0041] The result output module is used to output the plastic greenhouse image.

[0042] The technical solution provided by the embodiment of the present invention has the following beneficial effects:

[0043] The plastic greenhouse image extraction method and system based on multi-feature deep fusion provided by the embodiments of the present invention, based on the multi-feature deep fusion strategy, deeply excavates the reliable spectral prior knowledge hidden in the spectral index, and through the advanced attention mechanism, deeply integrates this spectral information with the high-resolution image, thereby fusing the rich spectral features of the multi-spectral remote sensing image and the fine details of the high-resolution remote sensing image; and by constructing a multi-branch structure, integrates the multi-scale data processing ability; and through an independent boundary neural network, proposes a boundary information-guided fusion mechanism to accurately guide the image segmentation and extraction process, thereby significantly improving the accuracy of plastic greenhouse image extraction, while enhancing the generalization ability and practicability of the model.

[0044] Of course, it is not necessary for any product or method implementing the present invention to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0046] Figure 1 It is a schematic diagram of the principle of the plastic greenhouse image extraction method based on multi-feature deep fusion described in the embodiments of the present invention;

[0047] Figure 2 It is a flowchart of the plastic greenhouse image extraction method based on multi-feature deep fusion described in the embodiments of the present invention;

[0048] Figure 3 It is a detailed comparison diagram of the plastic greenhouse image extraction of the first area using the extraction method described in the embodiments of the present invention;

[0049] Figure 4 It is a comparison diagram of the UA, OA, and F1score results of the plastic greenhouse image extraction of the first area using the extraction method described in the embodiments of the present invention;

[0050] Figure 5 It is a detailed comparison diagram of the plastic greenhouse image extraction of the second area using the extraction method described in the embodiments of the present invention;

[0051] Figure 6 It is a comparison diagram of the UA, OA, and F1score results of the plastic greenhouse image extraction of the second area using the extraction method described in the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. The components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can also be combined with each other.

[0053] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. In the description of the present invention, the terms "first", "second", "third", "fourth", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0054] Based on satellite remote sensing data, how to extract plastic greenhouses more precisely. The embodiments of the present invention propose a method and system for extracting plastic greenhouse images based on multi-feature deep fusion. By constructing a multi-branch architecture that integrates a detail branch, a multi-spectral branch, and a semantic branch, multi-scale feature learning is performed to take into account the local details and global features of plastic greenhouses. In this architecture, an independent boundary learning neural network is introduced to guide the boundary information fusion mechanism, so as to solve the problems often encountered in the plastic greenhouse extraction task by deep learning convolutional neural networks, such as blurred boundaries and contour adhesion, improve the extraction accuracy and precision of plastic greenhouses, and at the same time enhance the generalization ability and practicality of the model.

[0055] Such as Figure 1As shown in the figure, specifically, each branch in the multi-branch structure is responsible for extracting features in different dimensions: the detail branch uses high-resolution images to capture the color and texture features of plastic greenhouses, which are crucial for distinguishing greenhouses from other surface objects; the multi-spectral branch delves into spectral index information to reveal the physical properties of surface objects, further enhancing the model's ability to identify plastic greenhouses; the semantic branch provides global context information to help the model understand the position and relationship of greenhouses in the overall scene. However, relying solely on these features is still difficult to completely solve the problem of incomplete boundaries. Therefore, the present invention innovatively introduces a boundary information-guided fusion mechanism. This mechanism constructs a boundary learning neural network to accurately extract the boundary information of the greenhouse and incorporates it as an auxiliary supervision signal into the deep learning segmentation and extraction process. The boundary learning neural network can capture more refined boundary details, thereby guiding the main network to more accurately extract the contour of the plastic greenhouse during segmentation. In terms of the fusion strategy, a multi-feature deep fusion method is adopted to efficiently fuse the spectral features in multi-spectral remote sensing data, the fine detail information and semantic information in high-resolution images, and the boundary information extracted by the boundary information-guided fusion mechanism. This deep fusion not only significantly improves the extraction accuracy of plastic greenhouses but also greatly enriches the feature dimensions of training data, enabling the model to exhibit stronger adaptability and generalization ability when facing complex scenes.

[0056] As Figure 2 shown, the method for extracting plastic greenhouse images based on multi-feature deep fusion includes the following steps:

[0057] Step S1, determine the study area and year, and obtain high-resolution remote sensing image data and multi-spectral remote sensing data corresponding to the time within the study area.

[0058] In this step, the high-resolution remote sensing image data and multi-spectral remote sensing data come from satellite remote sensing data. Among them, the remote sensing image data generally refers to RGB data, such as high-resolution RGB images (level 18) of the Google Earth platform, Sentinel-2 data, etc.

[0059] Step S2, based on the high-resolution remote sensing image data, extract remote sensing image samples; resample the multi-spectral remote sensing data and calculate index samples so that the resolution of the calculated index samples after resampling is the same as that of the remote sensing image samples.

[0060] In this step, the size of the extracted remote sensing image samples can be 256×256 or 512×512, etc.

[0061] When resampling the multi-spectral remote sensing data, the exponential threshold method is adopted. The exponential threshold method here uses the downsampling method. Through downsampling, the resolution is made consistent with the remote sensing image samples, which better meets the actual application requirements and at the same time reduces the computational pressure on the corresponding branch; moreover, the result of the exponential threshold processing itself is an effective identifier for identifying plastic greenhouses.

[0062] When the index adopts the new greenhouse index APGI the index calculation formula is as follows::

[0063] (1)

[0064] In formula (1), is the wavelength of the aerosol band, is the wavelength of the red band, is the wavelength of the near-infrared band, is the wavelength of the short-wave infrared band, APGI is the new plastic greenhouse index.

[0065] When using APGI the index to determine the plastic greenhouse samples, the value of the index threshold is determined specifically according to the actual situation. When APGI is greater than the preset index threshold, it is confirmed as a plastic greenhouse sample.

[0066] Step S3: Build a backbone network based on the deep convolutional neural network VGG-16; based on the backbone network, use an encoder-decoder structure to create a multi-spectral branch, a semantic branch, and a detail branch to construct a greenhouse image extraction model.

[0067] In this step, the greenhouse image extraction model further includes: a first downsampling module before the multi-spectral branch, and a second downsampling module before the semantic branch; after the three branches, there are a first attention module, a first upsampling module, a second attention module, a second upsampling module, a third attention module, and a third upsampling module.

[0068] Among them, the detail branch processes the remote sensing image samples of the original size and learns the detail features of the plastic greenhouses. Therefore, there is no need to set a downsampling module before the detail branch; the multi-spectral branch learns the multi-spectral features of the downsampled 1 / 64 index samples; the semantic branch processes the remote sensing image samples downsampled by 1 / 4 and learns the semantic features of the images.

[0069] In this step, preferably, the encoder-decoder structure is adopted and the stacked autoencoder algorithm is used. Among them, the semantic branch uses five convolutional layers of VGG-16 as the core of the feature encoder.

[0070] Among them, the encoder of the multi-spectral branch includes a convolutional layer, a batch normalization layer, and a ReLU activation function. The semantic branch includes several convolutional layers, generally five convolutional layers, and is a complete encoder to deeply extract the semantic features of remote sensing image samples. However, since the semantic branch performs multiple convolutional layers, it may cause the loss of spatial detail information of the samples. The detail branch is used to process the remote sensing image samples of the original size and includes two convolutional layers to finely learn the detail features of the plastic greenhouse, ensuring that the model can capture key detail information while maintaining high-efficiency operation.

[0071] Preferably, the greenhouse image extraction model is completed using Python code and implemented based on the Pytorch 3.6 framework. When training the model, the Adam optimizer with an initial learning rate of 0.0001 is selected for training, and the weight decay is set to the recommended default value. For example, all models for comparison are trained from scratch for 100 epochs until convergence. The batch size is set to 8, and the same parameter settings are ensured to equally evaluate the performance of different methods. For the dataset used in training, considering the risk of overfitting, some common data augmentation methods, such as vertical-horizontal flipping and random rotation, are applied to each 256×256 pixel image to expand the dataset.

[0072] Step S4, construct a boundary learning model based on the encoder-decoder structure of the deep learning neural network. The boundary learning model includes M convolutional blocks, an upsampling module, and a splicing module.

[0073] Construct a boundary learning model based on the encoder-decoder structure of the deep learning neural network. The boundary learning model includes M convolutional blocks.

[0074] In this step, the boundary learning model adopts a DexiNed edge detection model optimized based on the encoder-decoder structure. Among them, the encoder consists of M convolutional blocks, which are used to introduce a model training mechanism guided by boundary information. Each convolutional block can obtain M feature maps of different scales. In this embodiment, M = 5, that is, the encoder part is composed of five core modules, including five convolutional blocks, and these modules all adopt an optimized DenseNet structure. Through a series of densely connected convolutional layers, the encoder can gradually and effectively extract multi-scale features. Specifically, there are a total of five convolutional blocks, each of which is embedded with a 1x1 convolutional layer for reducing the feature dimension and a 3x3 convolutional layer for focusing on feature extraction; among them, the first convolutional block contains a 1x1 convolutional layer, two 3x3 convolutional layers and a pooling layer; the second convolutional block contains a 1x1 convolutional layer, two 3x3 convolutional layers and a pooling layer; the third convolutional block contains a 1x1 convolutional layer, three 3x3 convolutional layers and a pooling layer; the fourth convolutional block contains a 1x1 convolutional layer, three 3x3 convolutional layers, and the fifth convolutional block contains a 1x1 convolutional layer, three 3x3 convolutional layers; each convolutional block obtains feature maps of different scales. As the processing stage deepens from the first convolutional block to the fifth convolutional block, the number of feature channels of the convolutional block gradually increases, thus enhancing the feature expression ability; at the same time, after the feature extraction of the first three convolutional blocks, a max pooling operation is applied for downsampling, aiming to reduce the size of the feature map while retaining key information. In addition, the output of the current convolutional block will be connected to the output of the previous convolutional block through a splicing block to form a feature stack, and this mechanism greatly promotes the capture of detailed information.

[0075] The decoder consists of an upsampling module and a splicing module, which are used to fuse M feature maps of different scales to generate a fine fused edge map. The decoder part is responsible for layer-by-layer restoration of the features extracted by the encoder. Through the upsampling technique and the layer-by-layer feature depth fusion strategy, the decoder can generate more fine-grained edge features. Deconvolution operations are used to restore the size of the feature map, ensure the consistency of information in terms of scale, and output the final edge map. During the decoding process, the output of each convolutional block will be spliced with the corresponding encoder feature. This step allows the extraction of edge information from feature maps of different scales and the fusion of these features to finally generate a fused edge map. The fused edge map carries a boundary information guidance mechanism and is deeply fused with the feature extraction results of the multi-spectral branch, thereby guiding the semantic branch of the model to learn more precisely. Such a design not only improves the accuracy of edge detection but also enhances the sensitivity of the model to detailed information.

[0076] Step S5: Input the remote sensing image sample of the original size into the boundary learning model. The remote sensing image sample is extracted by M convolutional blocks to obtain M feature maps with different scales, and then through upsampling and splicing, edge information is extracted from the M feature maps with different scales, and the edge information is subjected to feature depth fusion to finally generate a fused edge map;

[0077] Step S6: Input the index sample into the first downsampling module of the greenhouse image extraction model, and output the index sample with a low resolution of 1 / 64 and input it into the multispectral branch to obtain multispectral features;

[0078] Input the remote sensing image sample into the second downsampling module of the greenhouse image extraction model, output the image sample with a low resolution of 1 / 4, and input it into the five convolutional layers of the semantic branch in sequence; obtain the first semantic feature based on the second convolutional layer, and obtain the second semantic feature based on the fifth convolutional layer;

[0079] Input the remote sensing image sample into the detail branch to extract the detail features of the plastic greenhouse.

[0080] In this step, among the five convolutional layers of the semantic branch, the number of features extracted by each convolutional layer depends on the number of output channels of each convolutional layer. For example, in a preferred embodiment, convolutional block 1: the number of input channels is 3, and the number of output channels is 64; convolutional block 2: the number of input channels is 64, and the number of output channels is 128; convolutional block 3: the number of input channels is 128, and the number of output channels is 256; convolutional block 4: the number of input channels is 256, and the number of output channels is 512; convolutional block 5: the number of input channels is 512, and the number of output channels is 512. The total number of features extracted by the semantic branch is the sum of the number of features extracted by all convolutional blocks: 64 + 128 + 256 + 512 + 512 = 1456. The semantic branch has extracted a total of 1456 kinds of features.

[0081] Step S7: Use the attention perception mechanism to fuse the fused edge map, multispectral features, first semantic feature, second semantic feature and detail features to extract the plastic greenhouse image.

[0082] This step specifically includes:

[0083] Step S71: Input the fused edge map, multispectral features and second semantic feature into the first attention fusion module at the same time, fuse the fused edge map, multispectral features and second semantic feature to obtain the first fused image feature map;

[0084] Step S72: After the first fused image feature map is upsampled 8 times by the first upsampling module, input it into the second attention fusion module together with the first semantic feature to obtain the second fused image feature map;

[0085] Step S73: After the second fused image feature map is upsampled 4 times by the second upsampling module, it is input into the third attention fusion module together with the detail features to obtain the third fused image feature map;

[0086] Step S74: After the third fused image feature map is upsampled 2 times by the third upsampling module, the fused plastic greenhouse image map is output.

[0087] Among them, the attention mechanism in the first to third attention fusion modules can automatically capture the non-linear correlation between features, thereby strengthening important features and weakening the influence of irrelevant features. The formula of the attention mechanism is as follows:

[0088] (2)

[0089] In formula (2), X 1 ∈R H×W×C1 and X 2 ∈R H×W×C2 are two input feature maps, both belonging to the shape of R H×W×C . R refers to the resolution of the feature map, indicating the total number of pixels contained in the feature map. H and W respectively represent the height and width of the input image, while C1 and C2 respectively represent the number of channels of the two feature maps. After stacking X 1 and X 2 in the channel dimension, the new feature set X 1 , X 2 is obtained. θ is a vector belonging to the shape of R 1×1×(C1+C2) , obtained by AF learning, representing the weight of each channel.

[0090] To determine the channel correlation between different feature maps, first encode the spatial information of each channel, and then model its channel correlation. Since the receptive field of traditional convolution operations is limited, only limited spatial local features can be learned. To obtain spatial global features, first obtain the global distribution of different channels through global average pooling (GAP) operation. The GAP operation can be expressed in the following form:

[0091] (3)

[0092] In formula (3), X C (i,j) is the feature map of the pixel value (i,j) in the C-th channel, and P C ∈R 1×1×(C1+C2)It is the result after the GAP operation on the C-th channel, and (i, j) represents the pixel value at the i-th row and j-th column.

[0093] After obtaining the spatial global information of each channel, the relationships between different channels of different input features are modeled. To learn the non-linear correlation relationships between channels, a gating mechanism is adopted, which is expressed in the following form:

[0094] (4)

[0095] In Equation (4), and represent the Sigmoid and ReLU activation functions respectively, and are the weight parameters of two fully connected layers, and r is an empirical value used to adjust the computational cost and complexity of the model, usually set to 1 / 16; θ represents the weight. P represents the set of GAP operation results of all channels, representing the combination of global pooling results of all channels.

[0096] After completing the modeling of the relationships between channels, each channel will automatically obtain its corresponding weight θ . Then, multiplying this weight θ by its corresponding channel feature, a re-weighted feature set can be obtained, and the process is expressed as follows:

[0097] (5)

[0098] In Equation (5), X F represents the re-weighted feature set.

[0099] To verify the effectiveness of this method, the plastic greenhouse extraction method based on multi-feature deep fusion in this embodiment is applied in the first region and the second region respectively to extract the plastic greenhouse images of the two regions, and the extraction effect is evaluated for accuracy. At the same time, the mapping accuracy and robustness are compared with the existing plastic greenhouse extraction methods based on deep learning. The first region is Luliang County, Yunnan Province, and the second region is Shouguang City, Shandong Province. The selected time is November 2022.

[0100] Based on the high-resolution RGB images (level 18) of the Google Earth platform and Sentinel-2 data in the two regions, steps S1 to S7 in this embodiment are executed to extract the plastic greenhouse images of the two regions, and the accuracy of the extracted images is quantitatively evaluated.

[0101] The quantitative evaluation indicators include Overall Accuracy (OA), Producer’s Accuracy (PA), User’s Accuracy (UA), and F1 score, and their calculation formulas are as follows:

[0102] (6)

[0103] (7)

[0104] (8)

[0105] (9)

[0106] In formulas (6)-(9), TP (True Positive): refers to the number of samples that the model correctly predicts as the positive class. That is, the samples that are actually the positive class are predicted as the positive class; FP (False Positive): refers to the number of negative class samples that the model incorrectly predicts as the positive class. That is, the samples that are actually the negative class are predicted as the positive class; FN (False Negative): refers to the number of positive class samples that the model incorrectly predicts as the negative class, that is, the samples that are actually the positive class are predicted as the negative class; TN (True Negative): refers to the number of samples that the model correctly predicts as the negative class. That is, the samples that are actually the negative class are predicted as the negative class.

[0107] Convert UA, OA, and F1 score into scores, compare the accuracy of the method of this application with that of the existing model, and draw Figure 4 and Figure 6 . The score conversion process is as follows:

[0108] Normalize the UA, OA, and F1 score results expressed as percentages to the range of 0-1, and then multiply by 3 to expand to the range of 0-3. The abscissa represents different methods (Unet++, DeepLab V3+, PSPNet, Linknet, MFF-GHNet). For example, taking the Luliang area as an example, for the F1 score of MFF-GHNet (96.40%): score = 96.40 / 100 × 3 = 2.892. This is consistent with Figure 4 the height of the F1 score bar of MFF-GHNet in

[0109] As Figure 3 and Figure 4 shown, compare the accuracy of the predicted extraction results of the Luliang dataset with the extraction results of other deep learning network models. The results are as follows:

[0110] Table 1 Comparison of extraction accuracies of plastic greenhouses using different extraction methods (Luliang dataset)

[0111]

[0112] As can be seen from Table 1, the model proposed in this embodiment (Multi-FeatureFusion GreenhouseNetwork, MFF-GHNet) achieved better extraction results compared to the highest accuracy (UA: 60.41%, OA: 72.71%, F1 score: 69.27%) achieved by directly applying other networks such as Linknet. The remote sensing classification extraction performance of the plastic greenhouse image extraction method based on multi-feature deep fusion in this embodiment on the Luliang dataset was more excellent (MFF-GHNet) (UA: 95.65%, OA: 96.18%, F1 score: 96.40%). The main reason for obtaining the above results may be that the plastic greenhouses in Luliang are affected by government department management, and the distribution of plastic greenhouses is very uniform and the types are relatively single. In addition, the edge detection model algorithm based on deep learning used in this embodiment, that is, the constructed boundary learning model, effectively filters out the noise information on the boundary, optimizes the quality of the extraction results to a certain extent, and thus improves the classification extraction accuracy of plastic greenhouses. Therefore, compared with the existing deep learning neural network model extraction methods, the method proposed in this embodiment can effectively filter out the edge adhesion phenomenon of the extraction results, improve the quality of the extraction results, and improve the interpretation accuracy of plastic greenhouse images.

[0113] As Figure 5 and Figure 6 shown, the accuracy of the predicted extraction results of the Shouguang dataset was compared with the accuracy of the extraction results of other deep learning network models, and the results are as follows:

[0114] Table 2 Comparison of extraction accuracies of plastic greenhouses using different extraction methods (Shouguang dataset)

[0115]

[0116] As can be seen from Table 2, the extraction method (MFF-GHNet) proposed in this embodiment is overall more effective than existing deep learning models in the remote sensing classification of plastic greenhouses. However, there are significant differences in the extraction accuracy of plastic greenhouses among different deep learning model extraction methods. For example, the extraction result of the Unet++ model has the worst UA, OA, and F1score accuracies among all the comparison models. Compared with the other four comparison network model methods, the plastic greenhouse extraction method with multi-feature deep fusion proposed in this embodiment has significantly improved the extraction accuracy of plastic greenhouses. The PA, UA, and F1score of the plastic greenhouse extraction on the Shouguang dataset reached 92.10%, 94.33%, and 91.92% respectively. Compared with the accuracy results of the previous Luliang dataset, the overall PA, UA, and F1score of Shouguang have decreased. This is because the distribution of plastic greenhouses in Shouguang City is diverse, with different types and sizes. There are significant differences in styles, lighting, greenhouse types, etc. between the two datasets. The model may be difficult to adapt to the categories, resulting in a relative decrease in overall indicators. However, compared with other models, it still shows good learning and extraction accuracy.

[0117] The above research results show that compared with directly applying the deep learning neural network model to the extraction of plastic greenhouses in the target area, using the method of multi-feature branch fusion to extract plastic greenhouses effectively improves the reliability of plastic greenhouse interpretation and the classification accuracy under high-resolution images. However, due to the limitations of the number and quality of plastic greenhouse samples, and the characteristic differences of plastic greenhouses in different regions and different years, the detection accuracy of plastic greenhouses detected by this method also varies slightly in different regions, but still maintains a relatively high extraction accuracy. The proposed method effectively retains the spectral information on the multi-spectral image and combines the detail information of the high-resolution image, improves the quality of the extraction result, and effectively improves the extraction accuracy of plastic greenhouse images.

[0118] Based on the same idea, the embodiment of the present invention also provides a plastic greenhouse image extraction system based on multi-feature deep fusion. The system includes: a data acquisition module, a sample generation module, a greenhouse extraction model construction module, a boundary learning model construction module, a feature deep fusion module, and a result output module.

[0119] Among them, the data acquisition module is used to determine the research area and year, and obtain high-resolution remote sensing image data and multi-spectral remote sensing data corresponding to the time within the research area.

[0120] The sample generation module is used to extract remote sensing image samples based on the high-resolution remote sensing image data; resample the multi-spectral remote sensing data and calculate index samples to make the resolution of the calculated index samples after resampling consistent with that of the remote sensing image samples.

[0121] The greenhouse extraction model construction module is used to build a backbone network based on the deep convolutional neural network VGG-16; based on the backbone network, an encoder-decoder structure is adopted to create a detail branch, a multispectral branch and a semantic branch, and a greenhouse image extraction model is constructed; among them, the multispectral branch is used to obtain multispectral features based on exponential samples; the semantic branch is used to obtain semantic features based on remote sensing image samples after 1 / 4 downsampling; the detail branch is used to extract the detail features of plastic greenhouses based on remote sensing image samples of the original size;

[0122] The greenhouse image extraction module further includes: a first downsampling module is arranged before the multispectral branch, and a second downsampling module is arranged before the semantic branch; after the three branches, a first attention module, a first upsampling module, a second attention module, a second upsampling module, a third attention module and a third upsampling module are arranged; the semantic branch includes five convolutional layers, wherein the second convolutional layer is used to obtain the first semantic feature, and the fifth convolutional layer is used to obtain the second semantic feature;

[0123] The boundary learning model construction module is used to construct a boundary learning model based on a deep learning neural network using an encoder-decoder structure. The boundary learning model includes M convolutional blocks, which are used to introduce a model training mechanism guided by boundary information, and each convolutional block can obtain M feature maps of different scales; the boundary learning module further includes an upsampling module and a splicing module, which are used to fuse the M feature maps of different scales to generate a fused edge map;

[0124] The feature depth fusion module is used to fuse the fused edge map, multispectral features, first semantic features, second semantic features and detail features by using an attention fusion mechanism, extract and obtain a plastic greenhouse image, and send it to the result output module;

[0125] The result output module is used to output the plastic greenhouse image.

[0126] In this embodiment, each module is implemented by a processor, and a memory is appropriately added when storage is required. Among them, the processor may be, but is not limited to, a microprocessor MPU, a central processing unit (CPU), a network processor (NP), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), other programmable logic devices, discrete gate, transistor logic devices, discrete hardware components, etc. The memory may include a random access memory (RAM), and may also include a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0127] In the above embodiment, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.).

[0128] In addition, it should be noted that the plastic greenhouse image extraction system based on multi-feature deep fusion described in this embodiment corresponds to the plastic greenhouse image extraction method based on multi-feature deep fusion. The description and limitation of the method also apply to the system, and will not be repeated here.

[0129] As can be seen from the above technical solutions, the method and system for extracting plastic greenhouses with multi-feature deep fusion provided by the embodiments of the present invention can effectively utilize the spectral information of multi-source remote sensing images and the ground object detail information of high-resolution images to fuse and extract the plastic greenhouse mapping of the target area. Case studies in the monitoring area show that the present invention can effectively improve the interpretation accuracy of plastic greenhouses, especially for phenomena such as boundary adhesion and unclear boundaries, effectively increasing the number of true pixels in the classification results and improving the classification accuracy of plastic greenhouses by deep learning algorithms under different regions and different types of conditions. Compared with the existing deep learning methods, the accuracy and quality of the plastic greenhouse images extracted by the present invention are higher. There is a higher consistency between the results of this embodiment and the field label results, and it has higher robustness for extracting plastic greenhouses in different regions and complex terrain regions.

[0130] The above description is only the preferred embodiments of the present invention and the description of the applied technical principles, and is not intended to limit the scope of the present invention claimed, but only represents the preferred embodiments of the present invention. Those skilled in the art should understand that the scope of the invention involved in the present invention is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the inventive concept. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.

Claims

1. A method for extracting plastic greenhouse images based on multi-feature fusion, characterized in that, It includes the following steps: Step S1: Determine the research area and year, and obtain high-resolution remote sensing image data and multispectral remote sensing data corresponding to the time within the research area; Step S2: Based on the high-resolution remote sensing image data, extract remote sensing image samples; resample the multispectral remote sensing data and calculate index samples to make the resolution of the calculated index samples after resampling consistent with that of the remote sensing image samples; Step S3: Build a backbone network based on the deep convolutional neural network VGG-16; Based on the backbone network, create a multispectral branch, a semantic branch, and a detail branch using an encoder-decoder structure to construct a greenhouse image extraction model; Step S4: Build a boundary learning model using an encoder-decoder structure based on a deep learning neural network. The boundary learning model includes M convolutional blocks, an upsampling module, and a splicing module; Step S5: Input the remote sensing image samples of the original size into the boundary learning model. The remote sensing image samples are introduced into the model training mechanism guided by boundary information through M convolutional blocks, and M feature maps with different scales are extracted. Then, through upsampling and splicing, edge information is extracted from the M feature maps with different scales, and the edge information is feature-fused to finally generate a fused edge map; In the constructed boundary learning model, the encoder consists of five convolutional blocks; among them, the first convolutional block contains a 1x1 convolutional layer, two 3x3 convolutional layers, and a pooling layer; the second convolutional block contains a 1x1 convolutional layer, two 3x3 convolutional layers, and a pooling layer; the third convolutional block contains a 1x1 convolutional layer, three 3x3 convolutional layers, and a pooling layer; the fourth convolutional block contains a 1x1 convolutional layer, three 3x3 convolutional layers, and the fifth convolutional block contains a 1x1 convolutional layer, three 3x3 convolutional layers; each convolutional block obtains feature maps with different scales; the decoder consists of an upsampling module and a splicing module, which are used to fuse the M feature maps with different scales. At this time, the output of the current convolutional block is connected to the output of the previous convolutional block through the splicing module to form a feature stack, and finally a fused edge map is generated; Step S6: Input the index samples into the first downsampling module of the greenhouse image extraction model, and output the index samples with a low resolution of 1 / 64 and input them into the multispectral branch to obtain multispectral features; input the remote sensing image samples into the second downsampling module of the greenhouse image extraction model, output the image samples with a low resolution of 1 / 4, and input them into the five convolutional layers of the semantic branch in sequence; obtain the first semantic feature based on the second convolutional layer, and obtain the second semantic feature based on the fifth convolutional layer; input the remote sensing image samples into the detail branch to extract the detail features of the plastic greenhouse; Step S7: Use an attention perception mechanism to fuse the fused edge map, multispectral features, first semantic feature, second semantic feature, and detail features to extract the plastic greenhouse image.

2. The method for extracting plastic greenhouse images based on multi-feature fusion according to claim 1, wherein Step S7 specifically includes: Step S71: Input the fused edge map, multispectral features, and second semantic feature into the first attention fusion module at the same time, fuse the fused edge map, multispectral features, and second semantic feature to obtain the first fused image feature map; Step S72: After the first fused image feature map is upsampled 8 times by the first upsampling module, it is input into the second attention fusion module together with the first semantic feature to obtain the second fused image feature map; Step S73: After the second fused image feature map is upsampled 4 times by the second upsampling module, it is input into the third attention fusion module together with the detail feature to obtain the third fused image feature map; Step S74: After the third fused image feature map is upsampled 2 times by the third upsampling module, the fused plastic greenhouse image is output.

3. The method for extracting plastic greenhouse images based on multi-feature fusion according to claim 1, characterized in that, The high-resolution remote sensing image data and multi-spectral remote sensing data in Step S1 are from the Google Earth platform high-resolution RGB image - level 18 and Sentinel-2 data.

4. The method for extracting plastic greenhouse images based on multi-feature fusion according to claim 1, wherein In Step S2, the size of the extracted remote sensing image samples is 256×256 or 512×512.

5. The method for extracting plastic greenhouse images based on multi-feature fusion according to claim 1, characterized in that, When resampling the multispectral remote sensing data in step S2, the exponential threshold method is adopted; the exponent adopts the new greenhouse index APGI , and the exponent calculation formula is as follows: (1), In formula (1), is the wavelength of the aerosol band, is the wavelength of the red band, is the wavelength of the near-infrared band, is the wavelength of the short-wave infrared band, APGI represents the new plastic greenhouse index; When APGI is greater than a preset exponential threshold, it is confirmed as a plastic greenhouse sample.

6. The method for extracting plastic greenhouse images based on multi-feature fusion according to claim 1, characterized in that, In Step S3, the greenhouse image extraction model further includes: a first downsampling module before the multi-spectral branch, and a second downsampling module before the semantic branch; after the three branches, there are a first attention module, a first upsampling module, a second attention module, a second upsampling module, a third attention module and a third upsampling module.

7. The method for extracting plastic greenhouse images based on multi-feature fusion according to any one of claims 1-6, characterized in that, The multi-spectral branch includes a convolutional layer, a batch normalization layer and a ReLU activation function; the semantic branch includes several convolutional layers; the detail branch includes two convolutional layers.

8. An image extraction system for plastic greenhouses based on multi-feature fusion, characterized in that, The system includes: a data acquisition module, a sample generation module, a greenhouse extraction model construction module, a boundary learning model construction module, a feature fusion module and a result output module; among them, The data acquisition module is used to determine the research area and year, and obtain the high-resolution remote sensing image data and multi-spectral remote sensing data corresponding to the time in the research area; The sample generation module is used to extract remote sensing image samples based on the high-resolution remote sensing image data; resample the multi-spectral remote sensing data and calculate the index samples, so that the calculated index samples after resampling are consistent with the resolution of the remote sensing image samples; The greenhouse extraction model construction module is used to build the backbone network based on the deep convolutional neural network VGG-16; based on the backbone network, create a detail branch, a multi-spectral branch and a semantic branch using the encoder-decoder structure to build the greenhouse image extraction model; among them, the multi-spectral branch is used to obtain multi-spectral features based on the index samples after 1 / 64 downsampling; the semantic branch is used to obtain semantic features based on the remote sensing image samples after 1 / 4 downsampling; the detail branch is used to extract the detail features of the plastic greenhouse based on the remote sensing image samples of the original size; The greenhouse image extraction module further includes: a first downsampling module is set before the multi-spectral branch, and a second downsampling module is set before the semantic branch; after the three branches, a first attention module, a first upsampling module, a second attention module, a second upsampling module, a third attention module and a third upsampling module are set; the semantic branch includes five convolutional layers, where the second convolutional layer is used to obtain the first semantic feature, and the fifth convolutional layer is used to obtain the second semantic feature; The boundary learning model construction module is used to construct a boundary learning model based on a deep learning neural network with an encoder-decoder structure. The boundary learning model includes M convolutional blocks, which are used to introduce a model training mechanism guided by boundary information. Each convolutional block can obtain M feature maps of different scales. The boundary learning module also includes an upsampling module and a splicing module, which are used to fuse the M feature maps of different scales to generate a fused edge map. In the constructed boundary learning model, the encoder consists of five convolutional blocks. Among them, the first convolutional block contains a 1x1 convolutional layer, two 3x3 convolutional layers, and a pooling layer; the second convolutional block contains a 1x1 convolutional layer, two 3x3 convolutional layers, and a pooling layer; the third convolutional block contains a 1x1 convolutional layer, three 3x3 convolutional layers, and a pooling layer; the fourth convolutional block contains a 1x1 convolutional layer, three 3x3 convolutional layers, and the fifth convolutional block contains a 1x1 convolutional layer, three 3x3 convolutional layers. Each convolutional block obtains feature maps of different scales. The decoder consists of an upsampling module and a splicing module, which are used to fuse the M feature maps of different scales. At this time, the output of the current convolutional block is connected to the output of the previous convolutional block through the splicing module to form a feature stack, and finally a fused edge map is generated. The feature fusion module is used to fuse the fused edge map, multi-spectral features, first semantic features, second semantic features, and detail features by using an attention fusion mechanism, extract the plastic greenhouse image, and send it to the result output module. The result output module is used to output the plastic greenhouse image.

Citation Information

Patent Citations

  • Forest unstructured scene segmentation method based on multispectral image fusion

    CN114821064A

  • Hyperspectral and multispectral image fusion method and system based on improved Transform

    CN119559066A