Plastic greenhouse image extraction method and system based on multi-feature deep fusion
Through the combination of multi-feature deep fusion and boundary learning neural network, the problems of boundary blur and contour sticking in plastic greenhouse extraction are solved, and the extraction accuracy and model generalization ability are improved.
Patent Information
- Application Number
- CN202510457619.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The existing technology faces problems such as blurred boundary and contour sticking in plastic greenhouse extraction tasks. In addition, traditional machine learning methods rely on artificial feature engineering, making it difficult to dig deep features and are susceptible to interference from similar objects, resulting in unstable extraction results.
A method based on multi-feature deep fusion is adopted to deeply fusion of multi-scale features by building a multi-branch architecture, and an independent boundary learning neural network is introduced to guide the fusion mechanism of boundary information to solve the problems of boundary blur and contour sticking.
It significantly improves the accuracy of image extraction in plastic greenhouses, enhances the generalization ability and practicality of the model, improves the quality of the extraction results, and reduces boundary adhesion.
Smart Images

Figure CN119991471A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of machine vision and image processing, and specifically relates to a plastic greenhouse image extraction method and system based on multi-feature deep fusion. Background Art
[0002] Plastic greenhouses are agricultural sheds with plastic materials as the main structure. They are usually composed of plastic film, plastic pipes and steel frames. They are light, durable and have good light transmittance. They provide good environmental conditions for the growth of crops and play an important role in the development of modern agriculture. Although plastic greenhouses have improved the efficiency of agricultural production, the plastic films used in plastic greenhouses usually need to be replaced every 1-2 years. If the aged plastic films are not properly recycled, they will be discarded in farmland or nearby environments, causing white pollution. In addition, due to their extremely long degradation cycle, plastic films will persist in soil and water bodies, posing a serious threat to the ecosystem. These plastic films that are not easy to decompose will gradually split and degrade into microplastic particles, which can penetrate into the soil, interfere with the normal growth and development of crop roots, and then enter the food chain through crop absorption and accumulation, posing a potential threat to human and animal health. Therefore, accurately obtaining the area and spatial layout of agricultural plastic greenhouses is crucial for agricultural pollution prevention and control and efficient management, and helps to formulate scientific and reasonable agricultural policies and ensure the safety of food supply.
[0003] The area and spatial layout of agricultural plastic greenhouses and other related data are generally obtained through manual field surveys. However, this method is time-consuming and labor-intensive, and has low timeliness. In recent years, with the continuous development and maturity of satellite remote sensing technology, some researchers have also extracted plastic greenhouses based on satellite remote sensing data, which has the advantages of wide coverage, short revisit cycle and low data collection cost.
[0004] In the prior art, there are three types of plastic greenhouse extraction methods based on remote sensing data: the first type is a plastic greenhouse extraction method based on spectral index threshold segmentation, the second type is a plastic greenhouse extraction method based on traditional machine learning, and the third type is plastic greenhouse extraction mapping based on deep learning.
[0005] Among them, the first type of method uses mathematical methods to expand the difference between plastic greenhouses and other ground objects based on the difference in spectral reflectance curves between plastic greenhouses and other ground objects, so that the plastic greenhouse to be studied can obtain the maximum brightness enhancement on the generated index image, while other ground objects are generally suppressed, and the extraction of plastic greenhouses is completed by setting appropriate segmentation thresholds. For example, common plastic greenhouse extraction indexes include greenhouse vegetable land extraction index (greenhousevegetable land index, VI), plastic-mulched land cover index (plastic-mulched landcover index, PMLI), advanced plastic greenhouse index (advanced plastic greenhouse index, APGI), etc. Although the principle of plastic greenhouse extraction method based on spectral index threshold segmentation is simple, easy to understand and calculate. However, this type of method only relies on spectral feature differences to extract plastic greenhouses. Plastic greenhouses are made of various materials, and the spectral features within the class vary greatly, making it difficult to construct a spectral index with strong universality.
[0006] The second type of method uses traditional machine learning algorithms to extract plastic greenhouses based on the differences in texture, geometry, spectrum and other features between plastic greenhouses and other landforms. For example, based on 37 features such as brightness and density, the random forest method of sample optimization selection is used to extract plastic greenhouses. Another example is the classification effects of three machine learning methods: random forest, CART decision tree and support vector machine for the greenhouse extraction task of GF-2 remote sensing image, and the conclusion is that random forest classification has the best effect. Compared with the plastic greenhouse extraction method based on spectral index threshold segmentation, this method can learn the feature differences between the machine learning target and the background without a fixed threshold. However, this type of method can only use shallow feature differences such as spatial texture and spectrum to extract plastic greenhouses, and lacks in-depth analysis of feature engineering.
[0007] The third method uses neural networks to mine the deep feature differences between plastic greenhouses and other landforms, and uses the deep features to extract plastic greenhouses. For example, five fully convolutional neural networks of different scales are constructed through different feature depth fusion methods, and greenhouses and mulch fields are extracted based on these five models; another example is the use of the ENVINet5 deep learning architecture to extract sparsely distributed plastic greenhouses in high-resolution remote sensing images; another example is the use of the SSD network to make greenhouse labels based on GF-2 satellite images, and the overall accuracy of the extracted random sample points is 84.5%, and the Kappa coefficient is 0.831. However, due to multiple convolution and downsampling operations, the neural network based on the codec structure will cause the spatial resolution of remote sensing data to decrease, resulting in the loss of boundary information of the extracted agricultural plastic greenhouses, and the phenomenon of adhesion and fusion between greenhouse boundaries is prone to occur.
[0008] It can be seen that the existing spectral index threshold segmentation method faces the challenges of diversity, complexity and small differences in spectral characteristics of plastic greenhouses when extracting plastic greenhouses. It is difficult to obtain a universal threshold to fully extract detailed information on the greenhouse contour, which is difficult to meet actual needs. In addition, the background of remote sensing images is complex, and traditional machine learning methods rely on artificial feature engineering, which makes it difficult to mine its deep features and is easily disturbed by similar objects, resulting in unstable extraction results and "salt and pepper" phenomenon. Summary of the invention
[0009] In view of the above-mentioned defects or deficiencies in the prior art, the present invention aims to provide a method and system for plastic greenhouse image extraction based on multi-feature deep fusion. By constructing a multi-branch architecture, deep fusion of multi-scale features is performed, and an independent boundary learning neural network is introduced at the same time to guide the fusion mechanism of boundary information. The problems of blurred boundaries and contour adhesion often encountered by deep learning convolutional neural networks in plastic greenhouse extraction tasks are solved, thereby improving the recognition ability of plastic greenhouses.
[0010] In order to achieve the above purpose, the embodiment of the present invention adopts the following technical solution: In a first aspect, an embodiment of the present invention provides a method for extracting plastic greenhouse images based on multi-feature deep fusion, comprising the following steps: Step S1, determine the study area and year, and obtain high-resolution remote sensing image data and multispectral remote sensing data of the corresponding time in the study area; Step S2, extracting remote sensing image samples based on high-resolution remote sensing image data; resampling the multispectral remote sensing data and calculating index samples, so that the index samples calculated after resampling are consistent with the resolution of the remote sensing image samples; Step S3, building a backbone network based on the deep convolutional neural network VGG-16; based on the backbone network, using the encoder-decoder structure to create multispectral branches, semantic branches and detail branches, and constructing a greenhouse image extraction model; Step S4, constructing a boundary learning model based on a deep learning neural network using an encoder-decoder structure, wherein the boundary learning model includes M convolution blocks, an upsampling module, and a splicing module; Step S5, inputting the remote sensing image samples of the original size into the boundary learning model, introducing the model training mechanism guided by boundary information through M convolution blocks, extracting M feature maps with different scales, and then extracting edge information from the feature maps of M different scales through upsampling and splicing, performing feature deep fusion on the edge information, and finally generating a fused edge map; Step S6, input the index sample into the first downsampling module of the greenhouse image extraction model, output the 1 / 64 low-resolution index sample and input it into the multispectral branch to obtain the multispectral feature; input the remote sensing image sample into the second downsampling module of the greenhouse image extraction model, output the 1 / 4 low-resolution image sample, and input it into the five convolutional layers of the semantic branch in sequence; obtain the first semantic feature based on the second convolutional layer, and obtain the second semantic feature based on the fifth convolutional layer; input the remote sensing image sample into the detail branch to extract the detail feature of the plastic greenhouse; Step S7, using the attention perception mechanism to fuse the fused edge map, the multispectral feature, the first semantic feature, the second semantic feature and the detail feature to extract the plastic greenhouse image.
[0011] As a preferred embodiment of the present invention, step S7 specifically includes: Step S71, simultaneously inputting the fused edge map, the multispectral feature and the second semantic feature into a first attention fusion module, fusing the fused edge map, the multispectral feature and the second semantic feature to obtain a first fused image feature map; Step S72, after the first fused image feature map is upsampled by 8 times by the first upsampling module, the first fused image feature map is input into the second attention fusion module together with the first semantic feature to obtain a second fused image feature map; Step S73, after the second fused image feature map is upsampled by 4 times by the second upsampling module, it is input into the third attention fusion module together with the detail features to obtain the third fused image feature map; Step S74, after the third fused image feature map is upsampled by 2 times by the third upsampling module, the fused plastic greenhouse image is output.
[0012] As a preferred embodiment of the present invention, the high-resolution remote sensing image data and multispectral remote sensing data in step S1 come from Google Earth platform high-resolution RGB image-level 18 and Sentinel-2 data.
[0013] As a preferred embodiment of the present invention, in step S2, the size of the extracted remote sensing image sample is 256×256 or 512×512.
[0014] As a preferred embodiment of the present invention, when resampling the multispectral remote sensing data in step S2, an index threshold method is used; the index adopts a new greenhouse index APGI , and the index calculation formula is as follows: (1) In formula (1), is the wavelength of the aerosol band, is the wavelength of the red band, is the wavelength in the near-infrared band, is the wavelength of the short-wave infrared band, APGI It represents the new plastic greenhouse index; when APGI When it is greater than the preset index threshold, it is confirmed as a plastic greenhouse sample.
[0015] As a preferred embodiment of the present invention, in step S3, the greenhouse image extraction model also includes: a first downsampling module before the multispectral branch, and a second downsampling module before the semantic branch; after the three branches, it includes a first attention module, a first upsampling module, a second attention module, a second upsampling module, a third attention module and a third upsampling module.
[0016] As a preferred embodiment of the present invention, the multispectral branch includes a convolution layer, a batch normalization layer and a ReLU activation function; the semantic branch includes several convolution layers; and the detail branch includes two convolution layers.
[0017] As a preferred embodiment of the present invention, in the boundary learning model constructed in step S5, the encoder consists of five convolution blocks; wherein the first convolution block includes a 1x1 convolution layer, two 3x3 convolution layers and a pooling layer; the second convolution block includes a 1x1 convolution layer, two 3x3 convolution layers and a pooling layer; the third convolution block includes a 1x1 convolution layer, three 3x3 convolution layers and a pooling layer; the fourth convolution block includes a 1x1 convolution layer, three 3x3 convolution layers, and the fifth convolution block includes a 1x1 convolution layer, three 3x3 convolution layers; each convolution block obtains feature maps of different scales.
[0018] As a preferred embodiment of the present invention, the decoder is composed of an upsampling module and a splicing module, which is used to fuse M feature maps of different scales. At this time, the output of the current convolution block is connected with the output of the previous convolution block through the splicing module to form a feature stack, and finally generate a fused edge map.
[0019] In a second aspect, an embodiment of the present invention further provides a plastic greenhouse image extraction system based on multi-feature deep fusion, the system comprising: a data acquisition module, a sample generation module, a greenhouse extraction model construction module, a boundary learning model construction module, a feature deep fusion module and a result output module; wherein, The data acquisition module is used to determine the study area and year, and obtain high-resolution remote sensing image data and multispectral remote sensing data of the corresponding time in the study area; The sample generation module is used to extract remote sensing image samples based on high-resolution remote sensing image data; resample the multispectral remote sensing data and calculate index samples, so that the index samples calculated after resampling are consistent with the resolution of the remote sensing image samples; The greenhouse extraction model construction module is used to build a backbone network based on the deep convolutional neural network VGG-16; based on the backbone network, an encoder-decoder structure is used to create detail branches, multispectral branches and semantic branches to build a greenhouse image extraction model; wherein the multispectral branch is used to obtain multispectral features based on index samples; the semantic branch is used to obtain semantic features based on 1 / 4 down-sampled remote sensing image samples; the detail branch is used to extract detail features of plastic greenhouses based on remote sensing image samples of original size; The greenhouse image extraction module also includes: a first downsampling module is arranged before the multispectral branch, and a second downsampling module is arranged before the semantic branch; after the three branches, a first attention module, a first upsampling module, a second attention module, a second upsampling module, a third attention module and a third upsampling module are arranged; the semantic branch includes five convolution layers, wherein the second convolution layer is used to obtain the first semantic feature, and the fifth convolution layer is used to obtain the second semantic feature; The boundary learning model construction module is used to construct a boundary learning model based on a deep learning neural network using an encoder-decoder structure. The boundary learning model includes M convolution blocks for introducing a model training mechanism guided by boundary information. Each convolution block can obtain feature maps of M different scales. The boundary learning module also includes an upsampling module and a splicing module for fusing the feature maps of M different scales to generate a fused edge map. The feature depth fusion module is used to fuse the fused edge map, multispectral features, first semantic features, second semantic features and detail features using an attention fusion mechanism, extract the plastic greenhouse image, and send it to the result output module; The result output module is used to output the plastic greenhouse image.
[0020] The technical solution provided by the embodiment of the present invention has the following beneficial effects: The plastic greenhouse image extraction method and system based on multi-feature deep fusion provided by the embodiment of the present invention, based on the multi-feature deep fusion strategy, deeply mines the reliable spectral prior knowledge implicit in the spectral index, and deeply integrates these spectral information with the high-resolution image through an advanced attention mechanism, thereby integrating the rich spectral features of the multi-spectral remote sensing image and the fine details of the high-resolution remote sensing image; and by constructing a multi-branch architecture, it integrates multi-scale data processing capabilities; and through an independent boundary neural network, a boundary information guided fusion mechanism is proposed to accurately guide the image segmentation and extraction process, thereby significantly improving the accuracy of plastic greenhouse image extraction, while enhancing the generalization ability and practicality of the model.
[0021] Of course, it is not necessary to achieve all of the advantages described above at the same time to implement any product or method of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0023] Figure 1 It is a schematic diagram of the principle of the plastic greenhouse image extraction method based on multi-feature deep fusion according to an embodiment of the present invention; Figure 2 It is a flow chart of a method for extracting plastic greenhouse images based on multi-feature deep fusion according to an embodiment of the present invention; Figure 3 is a detailed comparison diagram of the plastic greenhouse image extraction in the first area using the extraction method in the embodiment of the present invention; Figure 4 is a comparison chart of UA, OA, and F1score results of applying the extraction method to extract plastic greenhouse images in the first region in an embodiment of the present invention; Figure 5 is a detailed comparison diagram of applying the extraction method to extract plastic greenhouse images in the second region in an embodiment of the present invention; Figure 6 It is a comparison chart of UA, OA, and F1score results of applying the extraction method to extract plastic greenhouse images in the second area in an embodiment of the present invention. DETAILED DESCRIPTION
[0024] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. The components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations. It should be noted that the embodiments of the present invention and the features in the embodiments can also be combined with each other without conflict.
[0025] It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. In the description of the present invention, the terms "first", "second", "third", "fourth", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.
[0026] In order to extract plastic greenhouses more accurately based on satellite remote sensing data, the embodiment of the present invention proposes a method and system for extracting plastic greenhouse images based on multi-feature deep fusion, which constructs a multi-branch architecture that integrates detail branches, multi-spectral branches, and semantic branches, and performs multi-scale feature learning to take into account both local details and global features of plastic greenhouses. In this architecture, an independent boundary learning neural network is introduced to guide the boundary information fusion mechanism to solve the problems of blurred boundaries and contour adhesion that deep learning convolutional neural networks often encounter in plastic greenhouse extraction tasks, improve the extraction precision and accuracy of plastic greenhouses, and enhance the generalization ability and practicality of the model.
[0027] like Figure 1 As shown in the figure, specifically, each branch in the multi-branch architecture is responsible for extracting features of different dimensions: the detail branch uses high-resolution images to capture the color and texture features of the plastic greenhouse, which are crucial for distinguishing the greenhouse from other surface objects; the multispectral branch deeply mines the spectral index information, reveals the physical properties of the surface objects, and further enhances the model's recognition ability of the plastic greenhouse; the semantic branch provides global context information to help the model understand the position and relationship of the greenhouse in the overall scene. However, it is still difficult to completely solve the problem of incomplete boundaries by relying solely on these features. Therefore, the present invention innovatively introduces a boundary information guided fusion mechanism. This mechanism accurately extracts the boundary information of the greenhouse by constructing a boundary learning neural network, and integrates it into the segmentation and extraction process of deep learning as an auxiliary supervision signal. The boundary learning neural network can capture finer boundary details, thereby guiding the main network to more accurately extract the outline of the plastic greenhouse during segmentation. In terms of fusion strategy, a multi-feature deep fusion method is adopted to efficiently fuse the spectral features in multispectral remote sensing data, the fine detail information and semantic information of high-resolution images, and the boundary information extracted by the boundary information guided fusion mechanism. This deep fusion not only significantly improves the extraction accuracy of plastic greenhouses, but also greatly enriches the feature dimensions of the training data, enabling the model to show stronger adaptability and generalization capabilities when facing complex scenarios.
[0028] like Figure 2 As shown, the plastic greenhouse image extraction method based on multi-feature deep fusion includes the following steps: Step S1, determine the study area and year, and obtain high-resolution remote sensing image data and multispectral remote sensing data of the corresponding time in the study area.
[0029] In this step, the high-resolution remote sensing image data and multispectral remote sensing data come from satellite remote sensing data, wherein the remote sensing image data generally refers to RGB data, such as Google Earth platform high-resolution RGB images (level 18), Sentinel-2 data, etc.
[0030] Step S2, extracting remote sensing image samples based on high-resolution remote sensing image data; resampling the multispectral remote sensing data and calculating index samples, so that the index samples calculated after resampling are consistent with the resolution of the remote sensing image samples.
[0031] In this step, the size of the extracted remote sensing image sample can be 256×256 or 512×512, etc.
[0032] When resampling the multispectral remote sensing data, the exponential threshold method is used. The exponential threshold method here adopts a downsampling method. Through downsampling, the resolution is kept consistent with the remote sensing image sample, which is more in line with the actual application requirements and reduces the calculation pressure of the corresponding branch; and the exponential threshold processing result itself is an effective identification for identifying plastic greenhouses.
[0033] When the index adopts the new greenhouse index APGI When , the index calculation formula is as follows: (1) In formula (1), is the wavelength of the aerosol band, is the wavelength of the red band, is the wavelength in the near-infrared band, is the wavelength of the short-wave infrared band, APGI It is the new plastic greenhouse index.
[0034] When using APGI When determining the plastic greenhouse samples based on the index, the value of the index threshold is determined based on the actual situation. APGI When it is greater than the preset index threshold, it is confirmed as a plastic greenhouse sample.
[0035] Step S3, building a backbone network based on the deep convolutional neural network VGG-16; based on the backbone network, using the encoder-decoder structure to create multispectral branches, semantic branches and detail branches, and constructing a greenhouse image extraction model.
[0036] In this step, the greenhouse image extraction model also includes: a first downsampling module before the multispectral branch, and a second downsampling module before the semantic branch; after the three branches, it includes a first attention module, a first upsampling module, a second attention module, a second upsampling module, a third attention module and a third upsampling module.
[0037] Among them, the detail branch processes remote sensing image samples of the original size and learns the detail features of the plastic greenhouse, so there is no need to set up a downsampling module before the detail branch; the multispectral branch learns the multispectral features of downsampled 1 / 64 index samples; the semantic branch processes remote sensing image samples downsampled 1 / 4 and learns the semantic features of the images.
[0038] In this step, preferably, the encoder-decoder structure adopts a stacked autoencoder algorithm, wherein the semantic branch adopts five convolutional layers of VGG-16 as the feature encoder core.
[0039] Among them, the encoder of the multispectral branch includes a convolution layer, a batch normalization layer and a ReLU activation function. The semantic branch includes several layers of convolution layers, generally including five convolution layers, and is a structurally complete encoder to deeply mine the semantic features of remote sensing image samples. However, due to the execution of multiple layers of convolution in the semantic branch, the spatial detail information of the sample may be lost. The detail branch is used to process remote sensing image samples of original size, including two layers of convolution layers to finely learn the detail features of the plastic greenhouse, ensuring that the model can capture key detail information while maintaining efficient operation.
[0040] Preferably, the greenhouse image extraction model is completed using Python code and implemented based on the Pytorch3.6 framework. When training the model, the Adam optimizer with an initial learning rate of 0.0001 is selected for training, and the weight decay is set to the recommended default value. For example, all models used for comparison are trained from scratch for 100 iterations until convergence. The batch size is set to 8, and the same parameter settings are guaranteed to evaluate the performance of different methods equally. For the data set used for training, considering the risk of overfitting, some common data enhancement methods are applied to each 256×256 pixel image, such as vertical-horizontal flipping and random rotation to expand the data set.
[0041] Step S4, constructing a boundary learning model based on a deep learning neural network using an encoder-decoder structure, wherein the boundary learning model includes M convolution blocks, an upsampling module, and a splicing module.
[0042] Based on the deep learning neural network, an encoder-decoder structure is used to construct a boundary learning model, which includes M convolution blocks. In this step, the boundary learning model adopts the DexiNed edge detection model based on the encoder-decoder structure optimization. Among them, the encoder is composed of M convolution blocks, which is used to introduce a model training mechanism guided by boundary information, and each convolution block can obtain M feature maps of different scales. In this embodiment, let M=5, that is, the encoder part is composed of five core modules, including five convolution blocks, and these modules all adopt the optimized DenseNet structure. Through a series of densely connected convolutional layers, the encoder can gradually and effectively extract multi-scale features. Specifically, there are five convolution blocks in total, each of which is embedded with a 1x1 convolution layer to reduce the feature dimension, and a 3x3 convolution layer to focus on feature extraction; among them, the first convolution block contains a 1x1 convolution layer, two 3x3 convolution layers and a pooling layer; the second convolution block contains a 1x1 convolution layer, two 3x3 convolution layers and a pooling layer; the third convolution block contains a 1x1 convolution layer, three 3x3 convolution layers and a pooling layer; the fourth convolution block contains a 1x1 convolution layer, three 3x3 convolution layers, and the fifth convolution block contains a 1x1 convolution layer, three 3x3 convolution layers; each convolution block obtains feature maps of different scales. As the processing stage from the first convolution block to the fifth convolution block deepens, the number of feature channels of the convolution block gradually increases, thereby enhancing the expressive power of the features; at the same time, after the feature extraction of the first three convolution blocks, the maximum pooling operation is applied for downsampling, aiming to reduce the size of the feature map while retaining key information. In addition, the output of the current convolution block will be connected with the output of the previous convolution block through the splicing block to form a feature stack. This mechanism greatly facilitates the capture of detail information.
[0043] The decoder consists of an upsampling module and a splicing module, which are used to fuse M feature maps of different scales to generate a fine fused edge map. The decoder part is responsible for restoring the features extracted by the encoder layer by layer. Through upsampling technology and layer-by-layer feature deep fusion strategy, the decoder can generate more refined edge features. Deconvolution operation is used to restore the size of the feature map, ensure the consistency of information in scale, and output the final edge map. During the decoding process, the output of each convolution block is spliced with the corresponding encoder feature. This step allows edge information to be extracted from feature maps of different scales and these features are fused to finally generate a fused edge map. The fused edge map carries a boundary information guidance mechanism and is deeply fused with the feature extraction results of the multispectral branch, thereby guiding the semantic branch of the model to learn more accurately. This design not only improves the accuracy of edge detection, but also enhances the model's sensitivity to detail information.
[0044] Step S5, inputting the remote sensing image samples of original size into the boundary learning model, extracting the remote sensing image samples into M feature maps with different scales through M convolution blocks, then upsampling and splicing, extracting edge information from the M feature maps with different scales, performing feature deep fusion on the edge information, and finally generating a fused edge map; Step S6, input the index sample into the first downsampling module of the greenhouse image extraction model, output the 1 / 64 low-resolution index sample into the multispectral branch, and obtain the multispectral feature; The remote sensing image samples are input into the second downsampling module of the greenhouse image extraction model, and 1 / 4 low-resolution image samples are output, which are sequentially input into the five convolutional layers of the semantic branch; the first semantic feature is obtained based on the second convolutional layer, and the second semantic feature is obtained based on the fifth convolutional layer; Input the remote sensing image samples into the detail branch to extract the detail features of the plastic greenhouse.
[0045] In this step, in the five convolutional layers of the semantic branch, the number of features extracted by each convolutional layer depends on the number of output channels of each convolutional layer. For example, in a preferred embodiment, convolution block 1: the number of input channels is 3, and the number of output channels is 64; convolution block 2: the number of input channels is 64, and the number of output channels is 128; convolution block 3: the number of input channels is 128, and the number of output channels is 256; convolution block 4: the number of input channels is 256, and the number of output channels is 512; convolution block 5: the number of input channels is 512, and the number of output channels is 512. The total number of features extracted by the semantic branch is the sum of the number of features extracted by all convolutional blocks: 64+128+256+512+512=1456. The semantic branch extracts a total of 1456 features.
[0046] Step S7, using the attention perception mechanism to fuse the fused edge map, the multispectral feature, the first semantic feature, the second semantic feature and the detail feature to extract the plastic greenhouse image.
[0047] This step specifically includes: Step S71, simultaneously inputting the fused edge map, the multispectral feature and the second semantic feature into a first attention fusion module, fusing the fused edge map, the multispectral feature and the second semantic feature to obtain a first fused image feature map; Step S72, after the first fused image feature map is upsampled by 8 times by the first upsampling module, the first fused image feature map is input into the second attention fusion module together with the first semantic feature to obtain a second fused image feature map; Step S73, after the second fused image feature map is upsampled by 4 times by the second upsampling module, it is input into the third attention fusion module together with the detail features to obtain the third fused image feature map; Step S74, after the third fused image feature map is upsampled by 2 times by the third upsampling module, the fused plastic greenhouse image is output.
[0048] Among them, the attention mechanism in the first to third attention fusion modules can automatically capture the nonlinear correlation between features, thereby strengthening important features and weakening the influence of irrelevant features. The attention mechanism formula is as follows: (2) In formula (2), X 1 ∈R H×W×C1 and X 2 ∈R H×W×C2 are two input feature maps, both belonging to R H×W×C The shape of the feature map is R, which refers to the resolution of the feature map, indicating the total number of pixels contained in the feature map. H and W represent the height and width of the input image, respectively, while C1 and C2 represent the number of channels of the two feature maps. X 1 and X 2 After superposition in the channel dimension, a new feature set is obtained [ X 1 , X 2 ]. θ It is a R 1×1×(C1+C2) A vector of the shape of , obtained through AF learning, represents the weight of each channel.
[0049] In order to determine the channel correlation between different feature maps, the spatial information of each channel is first encoded, and then its channel correlation is modeled. Since the traditional convolution operation has a limited receptive field, only limited spatial local features can be learned. In order to obtain spatial global features, the global average pooling (GAP) operation is first performed to obtain the global distribution of different channels. The GAP operation can be expressed as follows: (3) In formula (3), X C (i,j) is the feature map of the pixel value (i, j) in the Cth channel, P C ∈R 1×1×(C1+C2) It is the result of the GAP operation on the Cth channel, and (i, j) represents the pixel value of the i-th row and j-th column.
[0050] After obtaining the spatial global information of each channel, the relationship between different channels of different input features is modeled. In order to learn the nonlinear correlation between channels, a gating mechanism is used, which is expressed as follows: (4) In formula (4), and Represent the Sigmoid and ReLU activation functions respectively, and are the weight parameters of the two fully connected layers, and r is an empirical value used to adjust the computational cost and complexity of the model, usually set to 1 / 16; θ Represents weight. P Represents the collection of GAP operation results of all channels, representing the combination of global pooling results of all channels.
[0051] After the modeling of the channel relationships is completed, each channel will automatically obtain its corresponding weight θ Then, this weight θ By multiplying the corresponding channel features, we can get the re-weighted feature set. The process is as follows: (5) In formula (5), X F Represents the re-weighted feature set.
[0052] In order to verify the effectiveness of this method, the plastic greenhouse extraction method based on multi-feature deep fusion of this embodiment is applied in the first region and the second region respectively to extract the plastic greenhouse images of the two regions, and the accuracy of the extraction effect is evaluated. At the same time, the mapping accuracy and robustness are compared with the existing plastic greenhouse extraction method based on deep learning. The first region is Luliang County, Yunnan Province, and the second region is Shouguang City, Shandong Province. The selected time is November 2022.
[0053] Based on the Google Earth platform high-resolution RGB images (level 18) and Sentinel-2 data of the two regions, steps S1 to S7 of this embodiment are executed to extract the plastic greenhouse images of the two regions, and the accuracy of the extracted images is quantitatively evaluated.
[0054] The quantitative evaluation indicators include Overall Accuracy (OA), Producer's Accuracy (PA), User's Accuracy (UA) and F1 score (F1score), and the calculation formula is as follows: (6) (7) (8) (9) In formulas (6)-(9), TP (True Positive): refers to the number of samples correctly predicted as positive by the model. That is, samples that are actually positive are predicted as positive; FP (False Positive): refers to the number of negative samples that are incorrectly predicted as positive by the model. That is, samples that are actually negative are predicted as positive; FN (False Negative): refers to the number of positive samples that are incorrectly predicted as negative by the model, that is, samples that are actually positive are predicted as negative; TN (True Negative): refers to the number of samples correctly predicted as negative by the model. That is, samples that are actually negative are predicted as negative.
[0055] Convert UA, OA and F1score into scores, compare the accuracy of this method with the existing model, and draw Figure 4 and Figure 6 The score conversion process is as follows: The UA, OA and F1score results expressed as percentages are normalized to the range of 0-1, and then multiplied by 3 to expand to the range of 0-3. The horizontal axis represents different methods (Unet++, DeepLab V3+, PSPNet, Linknet, MFF-GHNet). For example, taking Luliang area as an example, for the F1score of MFF-GHNet (96.40%): score = 96.40 / 100 × 3 =2.892. This is consistent with Figure 4 The F1score columns of MFF-GHNet are highly consistent.
[0056] like Figure 3 and Figure 4 As shown in the figure, the accuracy of the inferred Luliang dataset extraction results is compared with the accuracy of the extraction results of other deep learning network models. The results are as follows: Table 1 Comparison of plastic greenhouse extraction accuracy of different extraction methods (Luliang dataset)
[0057] As can be seen from Table 1, the model (Multi-FeatureFusion GreenhouseNetwork, MFF-GHNet) proposed in this embodiment has a better extraction effect than the highest accuracy achieved by directly applying other networks (such as Linknet) (UA: 60.41%, OA: 72.71%, F1 score: 69.27%). The plastic greenhouse image extraction method based on multi-feature deep fusion in this embodiment has better remote sensing classification extraction performance on the Luliang data set (MFF-GHNet) (UA: 95.65%, OA: 96.18%, F1 score: 96.40%). The main reason for obtaining the above results may be that the plastic greenhouses in Luliang are affected by the management of government departments, and the distribution of plastic greenhouses is very uniform and the types are relatively single; in addition, the edge detection model algorithm based on deep learning adopted in this embodiment, that is, the constructed boundary learning model, optimizes the quality of the extraction results to a certain extent by effectively filtering out the noise information on the boundary, thereby improving the classification extraction accuracy of plastic greenhouses. Therefore, compared with the existing deep learning neural network model extraction method, the method proposed in this embodiment can effectively filter out the edge adhesion phenomenon of the extraction results, improve the quality of the extraction results, and improve the accuracy of plastic greenhouse image interpretation.
[0058] like Figure 5 and Figure 6 As shown in the figure, the accuracy of the inferred Shouguang dataset extraction results is compared with the accuracy of the extraction results of other deep learning network models. The results are as follows: Table 2 Comparison of plastic greenhouse extraction accuracy of different extraction methods (Shouguang dataset)
[0059] As can be seen from Table 2, the effectiveness of the extraction method (MFF-GHNet) proposed in this embodiment in the remote sensing classification of plastic greenhouses is generally better than the existing deep learning model, but different deep learning model extraction methods have significant differences in the extraction accuracy of plastic greenhouses; for example, the extraction result of the Unet++ model is the model with the worst UA, OA, and F1score accuracy among all the compared models. Compared with the other four comparative network model methods, the multi-feature deep fusion plastic greenhouse extraction method proposed in this embodiment has significantly improved the accuracy of plastic greenhouse extraction; the plastic greenhouse extraction accuracy PA, UA and F1score on the Shouguang dataset reached 92.10%, 94.33% and 91.92% respectively. Compared with the accuracy results of the previous Luliang dataset, Shouguang's overall PA, UA and F1score have all decreased. This is because the plastic greenhouses in Shouguang City are distributed differently, of different types and uneven sizes. The two types of datasets have significant differences in style, lighting, greenhouse type, etc. The model may be difficult to adapt in terms of category, resulting in a relative decrease in the overall index, but compared with other models, it still shows good learning extraction accuracy.
[0060] The above research results show that compared with directly applying the deep learning neural network model to the plastic greenhouse extraction in the target area, the multi-feature branch fusion method effectively improves the reliability of plastic greenhouse interpretation and the classification accuracy under high-resolution images; however, due to the limitations of the number and quality of plastic greenhouse samples, and the differences in the characteristics of plastic greenhouses in different regions and years, the detection accuracy of plastic greenhouses detected by this method is slightly different in different regions, but still maintains a high extraction accuracy. The proposed method effectively retains the spectral information on the multispectral image and combines the detailed information of the high-resolution image, improves the quality of the extraction results, and effectively improves the extraction accuracy of the plastic greenhouse image.
[0061] Based on the same idea, an embodiment of the present invention also provides a plastic greenhouse image extraction system based on multi-feature deep fusion, the system including: a data acquisition module, a sample generation module, a greenhouse extraction model construction module, a boundary learning model construction module, a feature deep fusion module and a result output module.
[0062] The data acquisition module is used to determine the research area and year, and obtain high-resolution remote sensing image data and multispectral remote sensing data of the corresponding time in the research area; The sample generation module is used to extract remote sensing image samples based on high-resolution remote sensing image data; resample the multispectral remote sensing data and calculate index samples, so that the index samples calculated after resampling are consistent with the resolution of the remote sensing image samples; The greenhouse extraction model construction module is used to build a backbone network based on the deep convolutional neural network VGG-16; based on the backbone network, an encoder-decoder structure is used to create detail branches, multispectral branches and semantic branches to build a greenhouse image extraction model; wherein the multispectral branch is used to obtain multispectral features based on index samples; the semantic branch is used to obtain semantic features based on 1 / 4 down-sampled remote sensing image samples; the detail branch is used to extract detail features of plastic greenhouses based on remote sensing image samples of original size; The greenhouse image extraction module also includes: a first downsampling module is arranged before the multispectral branch, and a second downsampling module is arranged before the semantic branch; after the three branches, a first attention module, a first upsampling module, a second attention module, a second upsampling module, a third attention module and a third upsampling module are arranged; the semantic branch includes five convolution layers, wherein the second convolution layer is used to obtain the first semantic feature, and the fifth convolution layer is used to obtain the second semantic feature; The boundary learning model construction module is used to construct a boundary learning model based on a deep learning neural network using an encoder-decoder structure. The boundary learning model includes M convolution blocks for introducing a model training mechanism guided by boundary information. Each convolution block can obtain feature maps of M different scales. The boundary learning module also includes an upsampling module and a splicing module for fusing the feature maps of M different scales to generate a fused edge map. The feature depth fusion module is used to fuse the fused edge map, multispectral features, first semantic features, second semantic features and detail features using an attention fusion mechanism, extract the plastic greenhouse image, and send it to the result output module; The result output module is used to output the plastic greenhouse image.
[0063] In this embodiment, each module is implemented by a processor, and a memory is appropriately added when storage is required. Among them, the processor can be but is not limited to a microprocessor MPU, a central processing unit (CPU), a network processor (NP), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), other programmable logic devices, discrete gates, transistor logic devices, discrete hardware components, etc. The memory may include a random access memory (RAM) and may also include a non-volatile memory (NVM), such as at least one disk storage. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.
[0064] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode.
[0065] It should also be noted that the plastic greenhouse image extraction system based on multi-feature deep fusion described in this embodiment corresponds to the plastic greenhouse image extraction method based on multi-feature deep fusion. The description and limitation of the method are also applicable to the system and will not be repeated here.
[0066] It can be seen from the above technical solutions that the multi-feature deep fusion plastic greenhouse extraction method and system provided by the embodiment of the present invention can effectively utilize the spectral information of multi-source remote sensing images and the ground object detail information of high-resolution images to fuse and extract plastic greenhouse mapping in the target area. Case studies in the monitoring area have shown that the present invention can effectively improve the interpretation accuracy of plastic greenhouses, especially for phenomena such as boundary adhesion and unclear boundaries, effectively increase the number of real pixels in the classification results, and improve the classification accuracy of deep learning algorithms for plastic greenhouses in different regions and under different types of conditions. Compared with existing deep learning methods, the plastic greenhouse images extracted by the present invention have higher accuracy and quality, and the results of this embodiment have higher consistency with the field label results, and have higher robustness for monitoring plastic greenhouse extraction in different areas and complex terrain areas.
[0067] The above description is only a preferred embodiment of the present invention and an explanation of the technical principles used. It is not intended to limit the scope of the invention claimed for protection, but only represents the preferred embodiment of the present invention. Those skilled in the art should understand that the scope of the invention involved in the present invention is not limited to the technical solution formed by a specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the inventive concept. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present invention.
Claims
1. A method for extracting plastic greenhouse images based on multi-feature deep fusion, characterized in that: The steps include: Step S1, determine the study area and year, and obtain high-resolution remote sensing image data and multispectral remote sensing data of the corresponding time in the study area; Step S2, extracting remote sensing image samples based on high-resolution remote sensing image data; resampling the multispectral remote sensing data and calculating index samples, so that the index samples calculated after resampling are consistent with the resolution of the remote sensing image samples; Step S3, building a backbone network based on the deep convolutional neural network VGG-16; Based on the backbone network, the encoder-decoder structure is used to create multispectral branches, semantic branches and detail branches to build a greenhouse image extraction model. Step S4, constructing a boundary learning model based on a deep learning neural network using an encoder-decoder structure, wherein the boundary learning model includes M convolution blocks, an upsampling module, and a splicing module; Step S5, inputting the remote sensing image samples of the original size into the boundary learning model, introducing the model training mechanism guided by boundary information through M convolution blocks, extracting M feature maps with different scales, and then extracting edge information from the feature maps of M different scales through upsampling and splicing, performing feature deep fusion on the edge information, and finally generating a fused edge map; Step S6, input the index sample into the first downsampling module of the greenhouse image extraction model, output the 1 / 64 low-resolution index sample into the multispectral branch, and obtain the multispectral feature; Input the remote sensing image sample into the second downsampling module of the greenhouse image extraction model, output 1 / 4 low-resolution image samples, and input them into the five convolutional layers of the semantic branch in sequence; obtain the first semantic feature based on the second convolutional layer, and obtain the second semantic feature based on the fifth convolutional layer; input the remote sensing image sample into the detail branch to extract the detail features of the plastic greenhouse; Step S7, using the attention perception mechanism to fuse the fused edge map, the multispectral feature, the first semantic feature, the second semantic feature and the detail feature to extract the plastic greenhouse image.
2. The method for extracting plastic greenhouse images based on multi-feature deep fusion according to claim 1 is characterized in that: Step S7 specifically includes: Step S71, simultaneously inputting the fused edge map, the multispectral feature and the second semantic feature into a first attention fusion module, fusing the fused edge map, the multispectral feature and the second semantic feature to obtain a first fused image feature map; Step S72, after the first fused image feature map is upsampled by 8 times by the first upsampling module, the first fused image feature map is input into the second attention fusion module together with the first semantic feature to obtain a second fused image feature map; Step S73, after the second fused image feature map is upsampled by 4 times by the second upsampling module, it is input into the third attention fusion module together with the detail features to obtain the third fused image feature map; Step S74, after the third fused image feature map is upsampled by 2 times by the third upsampling module, the fused plastic greenhouse image is output.
3. The method for extracting plastic greenhouse images based on multi-feature deep fusion according to claim 1, characterized in that: The high-resolution remote sensing image data and multispectral remote sensing data in step S1 come from the Google Earth platform high-resolution RGB image-level 18 and Sentinel-2 data.
4. The method for extracting plastic greenhouse images based on multi-feature deep fusion according to claim 1, characterized in that: In step S2, the size of the extracted remote sensing image sample is 256×256 or 512×512.
5. The method for extracting plastic greenhouse images based on multi-feature deep fusion according to claim 1, characterized in that: In step S2, when resampling the multispectral remote sensing data, an index threshold method is used; the index adopts a new greenhouse index APGI , and the index calculation formula is as follows: (1); In formula (1), is the wavelength of the aerosol band, is the wavelength of the red band, is the wavelength in the near-infrared band, is the wavelength of the short-wave infrared band, APGI It represents the new plastic greenhouse index; when APGI When it is greater than the preset index threshold, it is confirmed as a plastic greenhouse sample.
6. The method for extracting plastic greenhouse images based on multi-feature deep fusion according to claim 1, characterized in that: In step S3, the greenhouse image extraction model also includes: a first downsampling module before the multispectral branch, and a second downsampling module before the semantic branch; after the three branches, it includes a first attention module, a first upsampling module, a second attention module, a second upsampling module, a third attention module and a third upsampling module.
7. The method for extracting plastic greenhouse images based on multi-feature deep fusion according to any one of claims 1 to 6, characterized in that: The multispectral branch includes a convolution layer, a batch normalization layer and a ReLU activation function; the semantic branch includes several convolution layers; and the detail branch includes two convolution layers.
8. The method for extracting plastic greenhouse images based on multi-feature deep fusion according to claim 1, characterized in that: In the boundary learning model constructed in step S5, the encoder consists of five convolution blocks; among them, the first convolution block contains a 1x1 convolution layer, two 3x3 convolution layers and a pooling layer; the second convolution block contains a 1x1 convolution layer, two 3x3 convolution layers and a pooling layer; the third convolution block contains a 1x1 convolution layer, three 3x3 convolution layers and a pooling layer; the fourth convolution block contains a 1x1 convolution layer, three 3x3 convolution layers, and the fifth convolution block contains a 1x1 convolution layer and three 3x3 convolution layers; each convolution block obtains feature maps of different scales.
9. The method for extracting plastic greenhouse images based on multi-feature deep fusion according to claim 1, characterized in that: The decoder consists of an upsampling module and a splicing module, which is used to fuse M feature maps of different scales. At this time, the output of the current convolution block is connected with the output of the previous convolution block through the splicing module to form a feature stack and finally generate a fused edge map.
10. A plastic greenhouse image extraction system based on multi-feature deep fusion, characterized in that: The system includes: a data acquisition module, a sample generation module, a greenhouse extraction model construction module, a boundary learning model construction module, a feature depth fusion module and a result output module; wherein, The data acquisition module is used to determine the study area and year, and obtain high-resolution remote sensing image data and multispectral remote sensing data of the corresponding time in the study area; The sample generation module is used to extract remote sensing image samples based on high-resolution remote sensing image data; resample the multispectral remote sensing data and calculate index samples, so that the index samples calculated after resampling are consistent with the resolution of the remote sensing image samples; The greenhouse extraction model construction module is used to build a backbone network based on the deep convolutional neural network VGG-16; based on the backbone network, an encoder-decoder structure is used to create detail branches, multispectral branches and semantic branches to build a greenhouse image extraction model; wherein the multispectral branch is used to obtain multispectral features based on index samples; the semantic branch is used to obtain semantic features based on 1 / 4 down-sampled remote sensing image samples; the detail branch is used to extract detail features of plastic greenhouses based on remote sensing image samples of original size; The greenhouse image extraction module also includes: a first downsampling module is arranged before the multispectral branch, and a second downsampling module is arranged before the semantic branch; after the three branches, a first attention module, a first upsampling module, a second attention module, a second upsampling module, a third attention module and a third upsampling module are arranged; the semantic branch includes five convolution layers, wherein the second convolution layer is used to obtain the first semantic feature, and the fifth convolution layer is used to obtain the second semantic feature; The boundary learning model construction module is used to construct a boundary learning model based on a deep learning neural network using an encoder-decoder structure. The boundary learning model includes M convolution blocks for introducing a model training mechanism guided by boundary information. Each convolution block can obtain feature maps of M different scales. The boundary learning module also includes an upsampling module and a splicing module for fusing the feature maps of M different scales to generate a fused edge map. The feature depth fusion module is used to fuse the fused edge map, multispectral features, first semantic features, second semantic features and detail features using an attention fusion mechanism, extract the plastic greenhouse image, and send it to the result output module; The result output module is used to output the plastic greenhouse image.
Citation Information
Patent Citations
Remote sensing image fusion method based on knowledge guidance
CN113887619A
Forest unstructured scene segmentation method based on multispectral image fusion
CN114821064A
Homestead identification method and system based on multi-branch learning
CN115082778A
Multi-modal remote sensing data fused impervious surface extraction method and device
CN117593664A
Remote sensing extraction method and system for agricultural plastic greenhouse
CN117612021A
Cited By
Full-biological mulching film degradation intelligent monitoring method and system based on multi-source data acquisition
CN120279493A
Intelligent monitoring method and system for full biological film degradation based on multi-source data acquisition
CN120279493B