Target area extraction device and extraction method
By combining three-dimensional and one-dimensional feature extraction networks to process multispectral remote sensing images, the time-consuming, labor-intensive and poor-precision problems of rice-growing area extraction were solved, and low-cost and high-accuracy target area extraction was achieved.
Patent Information
- Application Number
- CN202211220145.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-08
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2042-10-08
AI Technical Summary
Existing technologies for rice-growing area extraction are time-consuming and labor-intensive, with poor precision, high cost, and poor generalization capabilities. In particular, they are unable to meet the requirements of real-time and accurate extraction in satellite imagery and multispectral imagery applications.
A target area extraction device is adopted, including an image input module, a trained three-dimensional feature extraction network and a one-dimensional feature extraction network. Through multispectral remote sensing image processing, three-dimensional and one-dimensional feature extraction are combined, and feature vectors are fused and classified using a data fusion module to output the target area image.
It achieves low-cost, high-accuracy and fast rice planting area extraction, can accurately distinguish target areas from non-target areas, and improves the real-time and accuracy of extraction.
Smart Images

Figure CN115731391B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image target extraction, and more particularly to a target region extraction device and method. Background Art
[0002] Rice is my country's most important food crop. Accurately and in real time, extracting rice-growing areas is crucial for rice yield estimation, precision agriculture implementation, and disaster assessment, providing valuable information for the government's formulation of national food security policies. Traditional methods for extracting rice areas rely on manual field measurements, which are then reported and aggregated through administrative systems. This is time-consuming, labor-intensive, and suffers from poor accuracy.
[0003] Satellite imagery-based extraction methods often use vegetation index thresholding or feature classification to identify large areas of rice-growing areas. However, due to limitations such as satellite revisit cycles, image resolution, and weather conditions, satellite imagery is expensive and cannot meet the requirements for accurate, real-time planted area extraction. Furthermore, the selection of vegetation indices and thresholds is influenced by factors such as rice variety and growing season, resulting in poor generalization and low accuracy when crop types are complex. Furthermore, existing feature classification methods are often designed for optical and hyperspectral remote sensing imagery, making them less applicable to multispectral imagery and requiring further improvement in accuracy. Summary of the Invention
[0004] An object of the present invention is to provide a target region extraction device and an extraction method to solve at least one of the problems existing in the prior art.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions:
[0006] A first aspect of the present invention provides a target region extraction device, comprising: an image input module, a trained three-dimensional feature extraction network connected in series with the image input module, a trained one-dimensional feature extraction network connected in series with the image input module and in parallel with the three-dimensional feature extraction network, a data fusion module, and an image output module, wherein:
[0007] The image input module is used to process a plurality of multispectral remote sensing images to generate first image data input to the three-dimensional feature extraction network and second image data input to the one-dimensional feature extraction network:
[0008] The three-dimensional feature extraction network is used to perform three-dimensional feature extraction on the first input image data and then output a first feature vector associated with the multispectral remote sensing image;
[0009] The one-dimensional feature extraction network is used to perform one-dimensional feature extraction on the second input image data and then output a second feature vector that distinguishes the target area from the non-target area:
[0010] a data fusion module, configured to fuse and classify the second eigenvector and the first eigenvector using a preset weight parameter to obtain image category probabilities corresponding to the second eigenvector and the first eigenvector;
[0011] An image output module is configured to output a target area image including a target area according to the image category probability.
[0012] Furthermore, the image input module includes:
[0013] a first image data generating unit, configured to slice the multispectral remote sensing image at preset slice values to generate three-dimensional slice data, wherein the three-dimensional slice data serves as the first image data;
[0014] a second image data generating unit, configured to generate first one-dimensional data, second one-dimensional data, and third one-dimensional data respectively according to the band parameters of the spectral remote sensing image, and to combine the first one-dimensional data, the second one-dimensional data, and the third one-dimensional data in a spectral dimension to generate the second image data;
[0015] The band parameters include the normalized vegetation index, the ratio vegetation index and the number of spectral bands, wherein:
[0016] The normalized vegetation index is obtained according to the following formula:
[0017]
[0018] The ratio vegetation index is obtained according to the following formula:
[0019]
[0020] Among them, Band q is the reflectance value of band q in the spectral remote sensing image, Band p is the reflectance value of band p in the spectral remote sensing image;
[0021] The first one-dimensional data includes a normalized vegetation index determined by any two bands;
[0022] The second one-dimensional data includes a ratio vegetation index determined by any two bands;
[0023] The third one-dimensional data includes spectral data of each wavelength band.
[0024] Furthermore, the second image data generating unit is further used to normalize and convolve the first one-dimensional data, normalize and convolve the second one-dimensional data, and normalize and convolve the third one-dimensional data, and combine the normalized and convoluted first one-dimensional data, second one-dimensional data and third one-dimensional data in the spectral dimension to obtain the second image data.
[0025] Furthermore, the three-dimensional feature extraction network includes:
[0026] a three-dimensional convolution module, at least one three-dimensional feature extraction module connected in series with the three-dimensional convolution module, and a first fully connected layer connected in series with the three-dimensional feature extraction module,
[0027] in,
[0028] The three-dimensional convolution module is used to perform three-dimensional convolution on the input first image data to output a three-dimensional convolution feature map.
[0029] The three-dimensional feature extraction module is used to take the three-dimensional convolution feature map as input, perform feature association, calibration and pooling on the three-dimensional convolution feature map, and then output a three-dimensional pooling feature map;
[0030] The first fully connected layer is used to take the three-dimensional pooling feature map as input, connect the three-dimensional pooling feature map, and output the first feature vector.
[0031] Furthermore, the three-dimensional feature extraction module includes a spatial self-attention unit, a three-dimensional residual unit, a first feature calibration unit, and a first pooling layer connected in series.
[0032] The other end of the spatial self-attention unit is connected in series with the three-dimensional convolution module.
[0033] The other end of the first pooling layer is connected in series with the first fully connected layer.
[0034] The spatial self-attention unit is used to take the three-dimensional convolution feature map as input, calculate the correlation features between the central pixel and other pixels of the three-dimensional convolution feature map, and output a spatial correlation feature map;
[0035] A three-dimensional residual unit, configured to perform residual calculation based on the spatial correlation feature map to obtain a three-dimensional residual feature map output;
[0036] The first feature calibration unit is configured to take the three-dimensional residual feature map as input, determine the correlation between adjacent channels in the three-dimensional residual feature map, and output a three-dimensional calibration feature map;
[0037] The first pooling layer is used to perform pooling compression on the three-dimensional calibration feature map, thereby outputting a three-dimensional pooling feature map.
[0038] Furthermore, the spatial self-attention unit includes:
[0039] The sub-feature map generation unit is configured to take the three-dimensional convolutional feature map as input and output a first sub-feature map, a second sub-feature map, and a third sub-feature map, respectively, where each sub-feature map satisfies the following formula:
[0040] F 1,2,3 =ReLU(f bn (F*W 3×3×1 +b));
[0041]
[0042] Among them, F 1,2,3 Corresponding to the first, second and third sub-feature maps respectively, W and b represent the weight and bias, X is the input three-dimensional convolution feature map, μ and σ 2 are the mean and variance of the input 3D convolution feature map, ∈, γ and β are training parameters;
[0043] a spatial self-attention weight map generating unit, configured to reshape the first sub-feature map and the second sub-feature map to generate corresponding first sub-feature matrices and second sub-feature matrices, and perform product calculation and normalization on the transposed matrices of the first sub-feature matrix and the second sub-feature matrix to generate a spatial self-attention weight map;
[0044] The original feature size map generation unit is used to multiply the third sub-feature map with the transposed matrix of the spatial self-attention weight map and perform feature reconstruction to obtain the original feature size map;
[0045] The spatial correlation feature map generation unit is used to multiply the original feature size map by the scale coefficient and add the product result to the input three-dimensional convolution feature map to obtain and output the spatial correlation feature map.
[0046] Furthermore, the first feature calibration unit includes:
[0047] The first descriptor determination unit is configured to take the three-dimensional residual feature map as input and determine the first descriptor y corresponding to each channel of the three-dimensional residual feature map by the following calculation formula:
[0048]
[0049] Among them, X l is the input three-dimensional residual feature map, S is the height or width of the input three-dimensional residual feature map, x i,j,c is the position of the pixel in the input three-dimensional residual feature map;
[0050] A first feature recalibration vector unit is used to determine a first correlation vector ω between multiple adjacent channels of the three-dimensional residual feature map according to the first descriptor and using the following calculation formula c :
[0051]
[0052] Where M is an integer, β m is the training parameter of the first feature recalibration vector unit in the mth three-dimensional feature extraction module in the M three-dimensional feature extraction modules, is the descriptor of the c-th channel of the first feature recalibration vector unit in the m-th three-dimensional feature extraction module in the M three-dimensional feature extraction modules;
[0053] A three-dimensional calibration feature map output unit is configured to determine the correlation of all channels of the three-dimensional residual feature map according to the first correlation vector and the following formula, thereby outputting the three-dimensional calibration feature map.
[0054]
[0055] Among them, x c is the pixel point of the c-th channel in the three-dimensional calibration feature map,
[0056] Furthermore, the one-dimensional feature extraction network includes: a one-dimensional convolution module, at least one one-dimensional feature extraction module connected in series with the one-dimensional convolution module, and a second fully connected layer connected in series with the one-dimensional feature extraction module.
[0057] The one-dimensional convolution module is used to perform one-dimensional convolution on the second input image data to output a one-dimensional convolution feature map.
[0058] The one-dimensional feature extraction module is used to perform residual calculation, feature calibration and pooling on the one-dimensional convolution feature map and output a one-dimensional pooling feature map.
[0059] The second fully connected layer is used to take the one-dimensional pooled feature map as input, connect the one-dimensional pooled feature map, and output the second feature vector.
[0060] Furthermore, the one-dimensional feature extraction module includes a one-dimensional residual unit, a second feature calibration unit, and a second pooling layer connected in series.
[0061] The other end of the one-dimensional residual unit is connected in series with the one-dimensional convolution module.
[0062] The other end of the second pooling layer is connected in series with the second fully connected layer.
[0063] A one-dimensional residual unit, configured to perform residual calculation based on the one-dimensional convolution feature map to obtain a one-dimensional residual feature map output;
[0064] The second feature calibration unit is configured to take the one-dimensional residual feature map as input, determine the correlation between adjacent channels in the one-dimensional residual feature map, and output a one-dimensional calibration feature map;
[0065] The second pooling layer is used to perform pooling compression on the one-dimensional calibration feature map, thereby outputting a one-dimensional pooling feature map.
[0066] A second aspect of the present invention provides a method for extracting a target area image using the target extraction device according to the first aspect of the present invention, the method comprising:
[0067] Processing a plurality of multispectral remote sensing images to generate first image data input to the three-dimensional feature extraction network and second image data input to the one-dimensional feature extraction network:
[0068] Performing three-dimensional feature extraction on the first input image data and outputting a first feature vector associated with the multispectral remote sensing image;
[0069] After performing one-dimensional feature extraction on the second input image data, a second feature vector is output to distinguish the target area from the non-target area:
[0070] fusing and classifying the second eigenvector and the first eigenvector using a preset weight parameter to obtain image category probabilities corresponding one-to-one to the second eigenvector and the first eigenvector;
[0071] A target area image including the target area is output according to the image category probability.
[0072] The beneficial effects of the present invention are as follows:
[0073] The target region extraction device of the embodiment of the present invention utilizes the three-dimensional feature extraction network to extract deep spatial spectral combination features of multispectral images, and utilizes a one-dimensional feature extraction network to mine spectral information and extract deep spectral features used to distinguish target regions from non-target regions, forming a dual-branch network architecture. The target region extraction device of the embodiment of the present invention has the advantages of low cost, high accuracy, and fast extraction speed. The target region image extracted by the target extraction device of the embodiment of the present invention can accurately distinguish target regions from non-target regions, and the target region extraction is accurate and complete. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] The specific embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0075] Figure 1A schematic diagram showing the network architecture of a target region extraction device according to an embodiment of the present invention;
[0076] Figure 2 A schematic diagram illustrating the architecture of an image input module according to an embodiment of the present invention;
[0077] Figure 3 A schematic diagram showing first image data according to an embodiment of the present invention;
[0078] Figure 4 A diagram showing the purpose of second image data according to an embodiment of the present invention;
[0079] Figure 5 A schematic diagram showing a framework of a three-dimensional feature extraction network according to an embodiment of the present invention;
[0080] Figure 6 A schematic diagram illustrating two-dimensional convolution and three-dimensional convolution according to an embodiment of the present invention is shown;
[0081] Figure 7 A schematic diagram illustrating a framework of a three-dimensional feature extraction module according to an embodiment of the present invention;
[0082] Figure 8 A schematic diagram illustrating a framework of a spatial self-attention unit according to an embodiment of the present invention is shown;
[0083] Figure 9 The feature extraction process of the spatial self-attention unit according to one embodiment of the present invention is shown;
[0084] Figure 10 A schematic diagram showing a framework of a three-dimensional residual unit according to an embodiment of the present invention;
[0085] Figure 11 A schematic diagram showing a framework of a first feature calibration unit according to an embodiment of the present invention;
[0086] Figure 12 The feature extraction process of the first feature calibration unit according to one embodiment of the present invention is shown;
[0087] Figure 13 A schematic diagram showing a framework of a three-dimensional feature extraction network and a one-dimensional feature extraction network according to an embodiment of the present invention;
[0088] Figures 14a to 14e It shows the target area images obtained by different methods using a multispectral remote sensing image as input;
[0089] Figures 15a to 15e It shows the target area images obtained in different ways using another multispectral remote sensing image as input;
[0090] Figures 16a to 16eIt shows the target area images obtained in different ways using another multispectral remote sensing image as input. DETAILED DESCRIPTION
[0091] In order to more clearly illustrate the present invention, the present invention will be further described below in conjunction with the embodiments and drawings. Similar components in the drawings are represented by the same reference numerals. It should be understood by those skilled in the art that the following specific description is illustrative rather than restrictive and should not be used to limit the scope of protection of the present invention.
[0092] Based on the above discussion, the first embodiment of the present invention proposes a target area extraction device, such as Figure 1 As shown, its network architecture includes: an image input module, a trained three-dimensional feature extraction network connected in series with the image input module, a trained one-dimensional feature extraction network connected in series with the image input module and in parallel with the three-dimensional feature extraction network, a data fusion module, and an image output module, wherein:
[0093] The image input module is used to process a plurality of multispectral remote sensing images to generate first image data input to the three-dimensional feature extraction network and second image data input to the one-dimensional feature extraction network:
[0094] The three-dimensional feature extraction network is used to perform three-dimensional feature extraction on the first input image data and then output a first feature vector associated with the multispectral remote sensing image;
[0095] The one-dimensional feature extraction network is used to perform one-dimensional feature extraction on the second input image data and then output a second feature vector that distinguishes the target area from the non-target area:
[0096] a data fusion module, configured to fuse and classify the second eigenvector and the first eigenvector using a preset weight parameter to obtain image category probabilities corresponding one-to-one to the second eigenvector and the first eigenvector;
[0097] An image output module is configured to output a target area image including a target area according to the image category probability, the second feature vector, and the first feature vector.
[0098] The target region extraction device of the embodiment of the present invention utilizes the three-dimensional feature extraction network to extract deep spatial spectral combination features of multispectral images, and utilizes a one-dimensional feature extraction network to mine spectral information and extract deep spectral features used to distinguish target regions from non-target regions, forming a dual-branch network architecture. The target region extraction device of the embodiment of the present invention has the advantages of low cost, high accuracy, and fast extraction speed. The target region image extracted by the target extraction device of the embodiment of the present invention can accurately distinguish target regions from non-target regions, and the target region extraction is accurate and complete.
[0099] In an optional embodiment, the multispectral remote sensing image is collected by an unmanned aerial vehicle remote sensing platform.
[0100] The multispectral remote sensing images collected by the UAV remote sensing platform have the characteristics of multispectral, multi-spatial and high temporal resolution, which can realize fixed-point multi-angle observation. By frequently shooting multispectral images of the current farmland, it can accurately guide the operation area of crops and effectively improve the real-time and accuracy of the target extraction area.
[0101] In a specific example, the target extraction device of an embodiment of the present invention is applied to extract rice areas in a multispectral remote sensing image. The workflow and principle of the target extraction device of the embodiment of the present invention will now be described based on this specific application.
[0102] In an optional embodiment, if Figure 2 As shown, the image input module includes a first image data generating unit and a second image data generating unit.
[0103] The first image data generating unit is configured to slice the multispectral remote sensing image at preset slice values to generate three-dimensional slice data, wherein the three-dimensional slice data serves as the first image data;
[0104] The second image data generating unit is used to generate first one-dimensional data, second one-dimensional data and third one-dimensional data according to the band parameters of the spectral remote sensing image, and to generate the second image data after combining the first one-dimensional data, the second one-dimensional data and the third one-dimensional data in the spectral dimension.
[0105] That is to say, the first image data generation unit and the second image data unit of the embodiment of the present invention generate three-dimensional data and one-dimensional data, respectively, which serve as inputs of the three-dimensional feature extraction network and the one-dimensional feature extraction network, respectively, to realize dual-branch feature extraction of three-dimensional spatial features and one-dimensional spectral features, thereby improving the accuracy of feature extraction.
[0106] In the embodiment of the present invention, the first image data generating unit performs pixel-by-pixel label assignment and efficient feature extraction on the entire multispectral remote sensing image by slicing processing, such as Figure 3 As shown, the entire multispectral image is sliced using a preset slice size. In one optional embodiment, the 3D slice data uses the preset slice size as width and height, and the number of channels (i.e., depth) is the number of bands in the multispectral remote sensing image. For example, if the multispectral remote sensing image has six bands and the slice size is set to 11×11, the 3D slice data size is 11×11×6, which serves as the input to the 3D feature extraction network. In this embodiment, the slice size should not be too large, otherwise it will introduce more negative information and increase the computational complexity.
[0107] In this embodiment, the second image data generation unit is used to extract spectral information from the multispectral remote sensing image, calculate the normalized difference vegetation index (NDVI) and ratio vegetation index (RVI) between the six bands respectively, and combine them with the original spectral data to form a one-dimensional second image data as the input of the one-dimensional feature extraction network.
[0108] In an optional embodiment, the band parameters include the normalized vegetation index, the ratio vegetation index and the number of spectral bands, wherein:
[0109] The normalized difference vegetation index NDVI is obtained according to the following formula:
[0110]
[0111] The ratio vegetation index RVI is obtained according to the following formula:
[0112]
[0113] Among them, Band q is the reflectance value of band q in the spectral remote sensing image, Band p is the reflectance value of band p in the spectral remote sensing image.
[0114] Based on the above formula, it can be seen that two different bands have different normalized vegetation indexes NDVI and ratio vegetation indices RVI. Therefore, the embodiment of the present invention calculates the normalized vegetation index NDVI and the ratio vegetation index RVI for each of two bands in all the bands, thereby generating multiple first one-dimensional data and multiple second one-dimensional data, respectively.
[0115] In an optional embodiment, the first one-dimensional data includes the normalized vegetation index determined by any two bands; the second one-dimensional data includes the ratio vegetation index determined by any two bands; and the third one-dimensional data includes spectral data for each band. Furthermore, to improve the accuracy of feature extraction for each spectral information, the first one-dimensional data in this embodiment of the present invention includes the normalized vegetation index corresponding to all two-band combinations for all bands, and the third one-dimensional data includes the normalized vegetation index corresponding to all two-band combinations for all bands.
[0116] For example, a multispectral remote sensing image includes 6 bands, and any two bands can form 15 different band combinations of normalized vegetation index and ratio vegetation index. Therefore, the depth of the first one-dimensional data is the same as the number of combinations of any two bands in all bands, and the depth of the third one-dimensional data is the same as the number of combinations of any two bands in all bands, so as to extract accurate spectral information. Figure 4 As shown, the first one-dimensional data is 1×1×15 three-dimensional data, which includes the normalized vegetation index, the third one-dimensional data is 1×1×15 three-dimensional data, which includes the ratio vegetation index, and the second one-dimensional data is 1×1×6 three-dimensional data, which includes the original spectral data of all bands.
[0117] In an optional embodiment, the second image data generating unit is further used to normalize and convolve the first one-dimensional data, normalize and convolve the second one-dimensional data, and normalize and convolve the third one-dimensional data, and combine the normalized and convoluted first one-dimensional data, second one-dimensional data and third one-dimensional data in the spectral dimension to obtain the second image data.
[0118] like Figure 4 As shown, after the second image data generation unit generates first one-dimensional data, second one-dimensional data and third one-dimensional data according to the multispectral remote sensing image, each one-dimensional data is normalized and convolved for feature extraction, and then combined in the spectral dimension to form a 1×1×36 second image data as the input of the one-dimensional feature extraction network.
[0119] In an optional embodiment, if Figure 5 As shown, the three-dimensional feature extraction network includes:
[0120] a three-dimensional convolution module, at least one three-dimensional feature extraction module connected in series with the three-dimensional convolution module, and a first fully connected layer connected in series with the three-dimensional feature extraction module,
[0121] in,
[0122] The three-dimensional convolution module is used to perform three-dimensional convolution on the input first image data to output a three-dimensional convolution feature map.
[0123] The three-dimensional feature extraction module is used to take the three-dimensional convolution feature map as input, perform feature association, calibration and pooling on the three-dimensional convolution feature map, and then output a three-dimensional pooling feature map;
[0124] The first fully connected layer is used to take the three-dimensional pooling feature map as input, connect the three-dimensional pooling feature map, and output the first feature vector, wherein the number of dimensions of the first feature vector is determined according to the target category of the multispectral remote sensing image.
[0125] The three-dimensional feature extraction network of the embodiment of the present invention takes the spatial 11×11 neighborhood centered on the sample point as input, that is, the first image data as input. The three-dimensional feature extraction network slides in both spatial and spectral dimensions to perform feature extraction, which can better preserve the spectral information inherent in multispectral remote sensing images.
[0126] In an embodiment of the present invention, the three-dimensional convolution module is used to perform three-dimensional convolution on the input first image data to output a three-dimensional convolution feature map. Figure 6 As shown in (a) on the left, the standard two-dimensional convolution operation is as follows: the learned two-dimensional convolution kernel is moved along the spatial dimension, and the kernel weights are multiplied and summed with the pixel values in the corresponding area to capture spatial features and generate a feature map for a channel. However, two-dimensional convolution is less effective for multispectral images and can easily cause spectral distortion because it destroys the correlation between bands at the beginning of the convolution. Therefore, the embodiment of the present invention uses three-dimensional convolution for feature extraction to fully utilize the spectral information of multispectral remote sensing images.
[0127] like Figure 6 As shown in (b) on the right, unlike two-dimensional convolution, the three-dimensional convolution kernel is three-dimensional and slides in both the spatial and spectral dimensions. The output of the three-dimensional convolution module is a feature map composed of many stacked cubes, each of which is generated by summing the weighted values of the input multi-band. In this way, three-dimensional convolution can better preserve the spectral information inherent in multispectral remote sensing images, thereby effectively extracting deep spatial-spectral combination features without any pre-processing or post-processing.
[0128] In an optional embodiment, if Figure 7 As shown, the three-dimensional feature extraction module includes a spatial self-attention unit, a three-dimensional residual unit, a first feature calibration unit, and a first pooling layer connected in series.
[0129] The other end of the spatial self-attention unit is connected in series with the three-dimensional convolution module.
[0130] The other end of the first pooling layer is connected in series with the first fully connected layer.
[0131] The spatial self-attention unit is used to take the three-dimensional convolution feature map as input, calculate the correlation features between the central pixel and other pixels of the three-dimensional convolution feature map, and output a spatial correlation feature map;
[0132] A three-dimensional residual unit, configured to perform residual calculation based on the spatial correlation feature map to obtain a three-dimensional residual feature map output;
[0133] The first feature calibration unit is configured to take the three-dimensional residual feature map as input, determine the correlation between adjacent channels in the three-dimensional residual feature map, and output a three-dimensional calibration feature map;
[0134] The first pooling layer is used to perform pooling compression on the three-dimensional calibration feature map, thereby outputting a three-dimensional pooling feature map.
[0135] In an optional embodiment, if Figure 8 As shown, the spatial self-attention unit includes:
[0136] The sub-feature map generation unit is configured to take the three-dimensional convolutional feature map as input and output a first sub-feature map, a second sub-feature map, and a third sub-feature map, respectively, where each sub-feature map satisfies the following formula:
[0137] F 1,2,3 =ReLU(f bn (F*W 3×3×1 +b))(2-1);
[0138]
[0139] Among them, the ReLU activation function, f bn (X)BN layer function with respect to X, * represents 3D convolution, F 1,2,3 Corresponding to the first, second and third sub-feature maps respectively, W and b represent the weight and bias, X is the input three-dimensional convolution feature map, μ and σ 2 are the mean and variance of the input 3D convolution feature map, ∈, γ and β are training parameters;
[0140] a spatial self-attention weight map generating unit, configured to reshape the first sub-feature map and the second sub-feature map to generate corresponding first sub-feature matrices and second sub-feature matrices, and perform product calculation and normalization on the transposed matrices of the first sub-feature matrix and the second sub-feature matrix to generate a spatial self-attention weight map;
[0141] The original feature size map generation unit is used to multiply the third sub-feature map with the transposed matrix of the spatial self-attention weight map and perform feature reconstruction to obtain the original feature size map;
[0142] The spatial correlation feature map generation unit is used to multiply the original feature size map by the scale coefficient and add the product result to the input three-dimensional convolution feature map to obtain and output the spatial correlation feature map.
[0143] Considering that the neighborhood spatial support around class boundaries is usually invalid because these neighborhood pixels sometimes have different categories from the center pixel, these neighboring pixels will have a negative impact on feature learning during the 3D convolution operation.
[0144] Therefore, after the 3D convolution module performs 3D convolution, in order to more effectively perform object classification, it is necessary to consider the potential correlation between the center pixel and its surrounding environment. In this embodiment of the present invention, a spatial self-attention unit (PAM) is introduced into the 3D feature calibration module connected in series with the 3D convolution module. The spatial self-attention unit can calculate the similarity between the features of the center pixel and the surrounding pixels and assign weights to different features. Therefore, the spatial self-attention unit can adaptively enhance the long-range features of the surrounding pixels related to the center pixel, while suppressing unnecessary features, so as to improve the spatial feature representation when predicting the center pixel and improve the accuracy of the 3D feature extraction network.
[0145] The three-dimensional feature extraction module of an embodiment of the present invention takes the three-dimensional convolution feature map generated by the three-dimensional convolution module as input, and outputs a three-dimensional pooling feature map after processing by each unit of the three-dimensional feature extraction module. The three-dimensional feature extraction module of this embodiment introduces a spatial self-attention unit and a first feature calibration unit on the basis of the three-dimensional residual unit to improve the effectiveness of feature information.
[0146] The feature extraction process of the spatial self-attention unit is now illustrated by way of example. Figure 9 As shown:
[0147] The feature map F passes through three different CBR units to obtain sub-feature maps F1, F2, and F3, which are expressed as follows:
[0148] F 1,2,3 =ReLU(f bn (F*W 3×3×1 +b))(2-1);
[0149]
[0150] ReLU activation function, f bn (X)BN layer function with respect to X, * represents 3D convolution, F1,2,3 Corresponding to the first sub-feature map F1, the second sub-feature map F2 and the third sub-feature map F3, W and b represent the weight and bias, X is the input three-dimensional convolution feature map, μ and σ 2 are the mean and variance of the input 3D convolution feature map, ∈, γ and β are training parameters.
[0151] In this embodiment, the CBR unit consists of a convolutional layer (Conv), a normalization layer (BN), and an activation function layer (ReLU). The CBR units that generate different sub-feature maps have different weight parameters, thereby obtaining sub-feature maps F1, F2, and F3 with different weight parameters. The inputs of sub-feature maps F1, F2, and F3 all come from the same input, namely the three-dimensional convolution feature map output by the three-dimensional convolution module.
[0152] The spatial self-attention weight map generation unit is used to calculate the similarity between the first sub-feature map F1 and the second sub-feature map F2, that is, the first sub-feature map F1 is converted from three dimensions to a two-dimensional sub-feature matrix. For example, the size of the first sub-feature map F1 is S×S×B, then the converted first sub-feature matrix is B×(S×S), and the size of the second sub-feature map F2 is S×S×B. After the conversion, the second sub-feature matrix is obtained and the transposed second sub-feature matrix F2 is obtained after the transposition. T , the first sub-feature matrix F1 and the transposed second sub-feature matrix F2 T Multiply them together and normalize the result using the softmax function to generate a two-dimensional spatial self-attention weight map A. spa , spatial self-attention weight map A spa The size of (S×S)×(S×S) is calculated as follows:
[0153]
[0154]
[0155] In this embodiment, the multiplication of the two sub-feature matrices is used to calculate the product of any two pixel values, and the softmax function normalizes all eigenvalues at each position, so A spa (c1, c2) is the influence weight of the c2th channel on the c1th channel in the original feature map.
[0156] Subsequently, the original feature size map generation unit combines the third sub-feature map F3 with the spatial self-attention weight map A spaThe original feature size map F4 is obtained by multiplying the transposed matrix of the original feature map F4 and performing feature reconstruction, mapping the two-dimensional attention weights to the three-dimensional spatial feature domain. Finally, the spatial correlation feature map generation unit multiplies the original feature size map F4 by the scale factor α and adds it to the input feature (three-dimensional convolution feature map), thereby obtaining a spatial correlation feature map with long-range context representation.
[0157] Spatially correlated feature maps can effectively enhance global feature representation using fewer parameters. By calculating the similarity between the central feature and surrounding features of the feature map, weights are assigned to different features, adaptively enhancing the features of other pixels related to the central pixel while suppressing unnecessary features, effectively improving the spatial feature representation when predicting the central pixel. In a specific example, α is initialized to 0 and gradually learned throughout the training process.
[0158] In an optional embodiment, if Figure 10 As shown, ResNet (residual network) can handle the vanishing gradient problem well. Therefore, the embodiment of the present invention uses a three-dimensional residual network to construct a three-dimensional residual unit, such as Figure 10 As shown in the figure, the 3D residual unit consists of multiple 3D residual blocks (ResBlocks) connected in series. The features output by the last 3D residual block are added to the input of the 3D residual unit, such as a spatial correlation feature map. After passing through the activation function layer, a 3D residual feature map is output. In a specific example, the 3D residual block consists of a concatenated convolutional layer (Conv), a normalization layer (BN), an activation function layer (ReLU), a convolutional layer (Conv), and a normalization layer (BN).
[0159] In an optional embodiment, if Figure 11 As shown, the first feature calibration unit includes:
[0160] The first descriptor determination unit is configured to take the three-dimensional residual feature map as input and determine the first descriptor y corresponding to each channel of the three-dimensional residual feature map by the following calculation formula:
[0161]
[0162] Among them, X l is the input three-dimensional residual feature map, S is the height or width of the input three-dimensional residual feature map, x i,j,c is the position of the pixel in the input three-dimensional residual feature map;
[0163] A first feature recalibration vector unit is used to determine a first correlation vector ω between multiple adjacent channels of the three-dimensional residual feature map according to the first descriptor and using the following calculation formula c :
[0164]
[0165] Where M is an integer, β m is the training parameter of the first feature recalibration vector unit in the mth three-dimensional feature extraction module in the M three-dimensional feature extraction modules, is the descriptor of the c-th channel of the first feature recalibration vector unit in the m-th three-dimensional feature extraction module in the M three-dimensional feature extraction modules;
[0166] A three-dimensional calibration feature map output unit is configured to determine the correlation of all channels of the three-dimensional residual feature map according to the first correlation vector and the following formula, thereby outputting the three-dimensional calibration feature map.
[0167]
[0168] Among them, x c is the pixel point of the cth channel in the three-dimensional calibration feature map.
[0169] Existing convolution operations basically fuse all channels of the input feature map by default, including spatial (H and W) and inter-channel (C) feature fusion. To improve model accuracy and reduce computational complexity, the present invention utilizes a first feature recalibration module (FRM) to automatically learn the importance of features in different channels, thereby improving classification performance.
[0170] Now, let's take the feature extraction process of the first feature calibration unit as an example. Specifically, Figure 12 As shown,
[0171] The first feature calibration unit takes the feature map output by the previous unit as input, such as the 3D residual feature map output by the 3D residual unit Any feature in Figure X l is the input, that is, the input three-dimensional residual feature map Output 3D calibration feature map after recalibration Specifically, we first calculate a descriptor y, which is used to characterize each channel by performing a feature map compression operation using global average pooling:
[0172]
[0173] Among them, X l is the input three-dimensional residual feature map, S is the height or width of the input three-dimensional residual feature map, x i,j,c is the position of the pixel in the input 3D residual feature map.
[0174] For example, Figure 12 As shown, the original size is S×S×B three-dimensional residual features Figure X l , after calculation, the first descriptor y corresponding to each channel is determined, and the feature size of the first descriptor y of all channels is 1×1×B.
[0175] Furthermore, in order to find useful feature maps, the first feature recalibration vector unit reweights the descriptor y along the channel dimension by considering a preset number of local neighborhoods, e.g. Figure 12 The four adjacent channels shown in the figure use one-dimensional convolution to capture the dependencies between all channels along the channel dimension, forming a network as shown in the figure. Figure 12 The feature map shown in FIG1 is captured with a preset number of local adjacent channels as the capture size and all channels as the capture end point. Further, the first feature recalibration vector unit sets the first correlation vector ω c The descriptors of all channels are associated, and the variance function is satisfied between them, thus forming Figure 12 The first correlation vector ω represented by the variance function is shown as a feature size of 1×1×B c , the first correlation vector ω c The descriptor y satisfies the following formula:
[0176]
[0177] Where M is an integer, β m is the training parameter of the first feature recalibration vector unit in the mth three-dimensional feature extraction module in the M three-dimensional feature extraction modules, It is the descriptor of the c-th channel of the first feature recalibration vector unit in the m-th three-dimensional feature extraction module in the M three-dimensional feature extraction modules.
[0178] In this embodiment, ω c Is a characteristic recalibration vector, after obtaining the first characteristic recalibration vector ω c After that, the first feature recalibration vector ω is further c Residual features with input Figure X l Perform product calculation to convert the input residual features Figure X l Convert to 3D calibration feature map It emphasizes the features of multiple convolution channels, has low computational complexity, and has significant improvements in multispectral remote sensing image classification.
[0179] 3D calibration feature map Produced by performing the following calculation:
[0180]
[0181] Among them, x c is the pixel point of the c-th channel in the three-dimensional calibration feature map,
[0182] Based on the spatial correlation between the central pixel and other pixels achieved by the spatial self-attention unit, the embodiment of the present invention introduces a first feature calibration unit to capture the local correlation of the channels in the spatial correlation feature map, and selectively enhance useful features while suppressing useless features. Unlike the correlation of all pixels in the spatial self-attention unit, the first feature calibration unit focuses on the interdependence between multiple adjacent channels. This significantly reduces the amount of computation while maintaining the feature extraction effect, thereby improving feature extraction efficiency.
[0183] In an optional embodiment, if Figure 13 As shown, the one-dimensional feature extraction network includes: a one-dimensional convolution module, at least one one-dimensional feature extraction module connected in series with the one-dimensional convolution module, and a second fully connected layer connected in series with the one-dimensional feature extraction module.
[0184] The one-dimensional convolution module is used to perform one-dimensional convolution on the second input image data to output a one-dimensional convolution feature map.
[0185] The one-dimensional feature extraction module is used to perform residual calculation, feature calibration and pooling on the one-dimensional convolution feature map and output a one-dimensional pooling feature map.
[0186] The second fully connected layer is used to take the one-dimensional pooled feature map as input, connect the one-dimensional pooled feature map, and output the second feature vector.
[0187] Based on the use of the three-dimensional feature extraction network to extract deep spatial-spectral combination features of multispectral images, the present invention also provides a one-dimensional feature extraction network connected in parallel with the three-dimensional feature extraction network. The one-dimensional feature extraction network in this embodiment of the present invention takes image data generated based on band parameters as input, that is, the second image data serves as the input for the one-dimensional feature extraction image. Because the first image data is a one-dimensional vector, each module in the one-dimensional feature extraction network performs one-dimensional operations, further utilizing the one-dimensional feature extraction network to mine one-dimensional spectral information from the spectral image.
[0188] In an optional embodiment, if Figure 13 As shown, the one-dimensional feature extraction module includes a one-dimensional residual unit, a second feature calibration unit, and a second pooling layer connected in series.
[0189] The other end of the one-dimensional residual unit is connected in series with the one-dimensional convolution module.
[0190] The other end of the second pooling layer is connected in series with the second fully connected layer.
[0191] A one-dimensional residual unit, configured to perform residual calculation based on the one-dimensional convolution feature map to obtain a one-dimensional residual feature map output;
[0192] The second feature calibration unit is configured to take the one-dimensional residual feature map as input, determine the correlation between adjacent channels in the one-dimensional residual feature map, and output a one-dimensional calibration feature map;
[0193] The second pooling layer is used to perform pooling compression on the one-dimensional calibration feature map, thereby outputting a one-dimensional pooling feature map.
[0194] In this embodiment, the main network structure of the one-dimensional residual unit is similar to the network structure of the three-dimensional residual unit, including multiple one-dimensional residual blocks connected in series. The features output by the last one-dimensional residual block are added to the input of the one-dimensional residual unit, such as a one-dimensional convolution feature map, and a one-dimensional residual feature map is output after the activation function layer. In a specific example, the three-dimensional residual block is composed of a convolution layer (Conv), a normalization layer (BN), an activation function layer (Relu), a convolution layer (Conv), and a normalization layer (BN) connected in series. For the specific network architecture, please refer to Figure 10 The three-dimensional residual unit shown is different from the one-dimensional residual unit in terms of input, operation, and output.
[0195] Based on the same principle, the structure, formula and feature calibration process of the second feature calibration unit of the embodiment of the present invention can also refer to the second feature calibration unit of the above embodiment. For example, with the one-dimensional residual feature map as input, the second descriptor determination unit determines the second descriptor y corresponding to each channel of the one-dimensional residual feature map by the following calculation formula 一维 :
[0196]
[0197] Among them, X 一维 is the input one-dimensional residual feature map, S is the height or width of the input one-dimensional residual feature map, x i,j,c is the position of the pixel in the input three-dimensional residual feature map;
[0198] Then, the second feature recalibration vector unit is based on the second descriptor y 一维 The second correlation vector ω between multiple adjacent channels of the one-dimensional residual feature map is determined by the following calculation formula: c一维
[0199]
[0200] Where M is an integer, β 一维 mis the training parameter of the second feature recalibration vector unit in the mth one-dimensional feature extraction module in the M one-dimensional feature extraction modules, is the descriptor of the c-th channel of the second feature recalibration vector unit in the m-th three-dimensional feature extraction module in the M three-dimensional feature extraction modules;
[0201] The one-dimensional calibration feature map output unit determines the correlation of all channels of the one-dimensional residual feature map according to the second correlation vector by the following formula, thereby outputting the one-dimensional calibration feature map
[0202] Among them, x c is the pixel point of the cth channel in the one-dimensional calibration feature map.
[0203] Furthermore, in an embodiment of the present invention, the one-dimensional calibration feature map is output to the second pooling layer for feature compression, so that the second fully connected layer can quickly connect nodes to output a second feature vector.
[0204] In an optional embodiment, the number of dimensions of the first eigenvector and the second eigenvector is determined according to the target category of the multispectral remote sensing image. For example, the final extraction goal is to classify the samples into 10 categories, so the output dimension needs to be 10 dimensions, each dimension is a probability, and by comparing the sizes of the 10 probabilities and selecting the largest probability, the classification of each pixel in the input multispectral image can be obtained.
[0205] Furthermore, in an optional embodiment, the number of neurons in the first fully connected layer and the second fully connected layer is equal to the number of target categories, the first fully connected layer outputs multiple first feature vectors F-B1, and the second fully connected layer outputs multiple second feature vectors F-B2, and the data fusion module fuses and classifies the second feature vectors and the first feature vectors with a preset weight parameter. For example, Figure 13 As shown, the weight parameter of the first eigenvector is θ, and the weight parameter of the second eigenvector is 1-θ, where θ is a learnable parameter in the range [0,1]. The first eigenvector and the second eigenvector are added to the softmax function for normalization to obtain the image category probability corresponding to the second eigenvector and the first eigenvector. That is, the final image category probability is a weighted vector obtained by the weighted sum of the scores of the first eigenvector F-B1 and the second eigenvector F-B2, and the value of each weighted vector is the category probability.
[0206] The image output module generates a target area image according to the category probability output by the data fusion module. The extracted target area image only includes the target area, and the non-target area is eliminated after extraction by the target extraction device of the embodiment of the present invention.
[0207] Exemplarily, the target area image of the embodiment of the present invention and the target area image of the related art are evaluated using the running time and the extraction accuracy.
[0208] The experimental data are 6-band UAV multispectral images in tif format with a size of 1280×960.
[0209] Figure 14a A multispectral remote sensing image captured by a drone according to an embodiment is shown. Figure 14b Shown with Figure 14a The image is taken as input, and the target area image is generated using the vegetation index threshold method (threshold 0.2); Figure 14c Shown with Figure 14a The image is taken as input, and the target area image is generated using the vegetation index threshold method (threshold 0.5); Figure 14d Shown with Figure 14a The image of is taken as input, and the target area image is generated by the SVM-based object classification method; Figure 14e Shown with Figure 14a The image is taken as input, and the target area image is extracted by the target extraction device according to the embodiment of the present invention.
[0210] Figure 15a A multispectral remote sensing image taken by a drone according to another embodiment is shown. Figure 15b Shown with Figure 15a The image is taken as input, and the target area image is generated using the vegetation index threshold method (threshold 0.2); Figure 15c Shown with Figure 15a The image is taken as input, and the target area image is generated using the vegetation index threshold method (threshold 0.5); Figure 15d Shown with Figure 15a The image of is taken as input, and the target area image is generated by the SVM-based object classification method; Figure 15e Shown with Figure 15a The image is taken as input, and the target area image is extracted by the target extraction device according to the embodiment of the present invention.
[0211] Figure 16a A multispectral remote sensing image taken by a drone according to another embodiment is shown. Figure 16b Shown with Figure 16a The image is taken as input, and the target area image is generated using the vegetation index threshold method (threshold 0.2); Figure 16c Shown with Figure 16a The image is taken as input, and the target area image is generated using the vegetation index threshold method (threshold 0.5); Figure 16d Shown with Figure 16a The image of is taken as input, and the target area image is generated by the SVM-based object classification method; Figure 16eShown with Figure 16a The image is taken as input, and the target area image is extracted by the target extraction device according to the embodiment of the present invention.
[0212] Experimental results from three different sets of multispectral remote sensing imagery indicate that the vegetation index threshold method extracts incomplete rice regions and suffers from numerous false positives, making it unable to accurately distinguish rice from other vegetation. Two thresholds, 0.2 and 0.5, were selected during the experiment. Comparison of the extraction results using these two thresholds reveals that the extraction performance of the vegetation index threshold method is closely related to the threshold selection: excessively high thresholds lead to incomplete extraction results, while excessively low thresholds produce numerous false positives. Comparison of extraction results across different scenes reveals that the extraction performance is better for scenes with simple feature types and distributions, but poorer for complex scenes. Furthermore, the optimal threshold selection varies with the scene, resulting in poor generalization. This is due to the high spatial resolution of UAV remote sensing imagery and the continuous enrichment of spatial features, particularly texture information, in rice. Consequently, the vegetation index threshold method, which currently relies solely on spectral information, no longer meets the accuracy requirements for rice region extraction, resulting in incomplete extraction results and numerous false positives.
[0213] While the SVM-based feature classification method utilizes the spatial and spectral features of multispectral images, achieving better extraction results than the traditional vegetation index threshold method, it still suffers from false alarms and incomplete rice area extraction. This is because the SVM-based feature classification method primarily consists of two steps: feature extraction and classifier design. The spectral and spatial features of the image are manually constructed, and then a trained classifier is used for classification. Manually constructed features often only yield the underlying features of the image, resulting in inadequate feature extraction. Furthermore, due to the irregular structure of the features and the presence of mixed pixels, this method struggles to correctly classify the edges of different features, leading to a series of issues such as incomplete rice area extraction.
[0214] Compared with the above two methods, the target extraction device of the embodiment of the present invention ensures accurate and complete extraction of the rice area by obtaining deep features of the image. Judging from the experimental results of the three sets of data, the target area image obtained by the embodiment of the present invention has no false alarms and the extraction results are relatively complete. There are only a small number of holes in the extracted rice area, which is caused by the occlusion of ground objects in the image. In other words, the embodiment of the present invention can accurately restore the target area in the multispectral remote sensing image.
[0215] Therefore, in summary, the traditional vegetation index threshold method and the SVM-based ground feature classification method have incomplete rice area extraction results and have certain false alarms. The extraction accuracy of the target area image of the channel of the present invention is high. Figures 14a to 14eTaking the image data of as an example, the extraction results of the existing method and the proposed algorithm are compared quantitatively, as shown in Table 1-1. The experimental environment is an Intel i5-10400 CPU, 16G RAM, and Nvidia GeForce RTX 2080 GPU.
[0216]
[0217] Table 1-1 Comparison of rice region extraction accuracy
[0218] It can be seen from the above table data that the target extraction device of the embodiment of the present invention has a faster running time while giving priority to ensuring the extraction accuracy, and has improved real-time and accuracy compared to the existing methods. The method of the embodiment of the present invention utilizes the three-dimensional feature extraction network to extract the deep spatial spectral combination features of the multi-spectral image, and utilizes the one-dimensional feature extraction network to mine spectral information, extract deep spectral features for distinguishing target areas from non-target areas, and forms a dual-branch network architecture. The target extraction device of the embodiment of the present invention has the advantages of low cost, high accuracy and fast extraction speed. The target area image extracted by the target extraction device of the embodiment of the present invention can accurately distinguish the target area from the non-target area, and extract the target area accurately and completely, laying the foundation for subsequent growth monitoring, yield estimation and other applications.
[0219] Another embodiment of the present invention provides a target extraction method, which includes the following steps:
[0220] Processing a plurality of multispectral remote sensing images to generate first image data input to the three-dimensional feature extraction network and second image data input to the one-dimensional feature extraction network:
[0221] Performing three-dimensional feature extraction on the first input image data and outputting a first feature vector associated with the multispectral remote sensing image;
[0222] After performing one-dimensional feature extraction on the second input image data, a second feature vector is output to distinguish the target area from the non-target area:
[0223] fusing and classifying the second eigenvector and the first eigenvector using a preset weight parameter to obtain image category probabilities corresponding one-to-one to the second eigenvector and the first eigenvector;
[0224] A target area image including the target area is output according to the image category probability.
[0225] It is worth noting that the method of the embodiment of the present invention can refer to the principle and process of the aforementioned target extraction device, which will not be repeated here. Similarly, the specific process of each step of the method can also refer to the aforementioned principle, which will not be repeated here.
[0226] Based on the method of the embodiment of the present invention, the three-dimensional feature extraction network is used to extract the deep spatial spectral combination features of the multispectral image, and the one-dimensional feature extraction network is used to mine spectral information to extract deep spectral features for distinguishing target areas from non-target areas. A dual-branch network architecture is formed, which can accurately distinguish target areas from non-target areas and extract target areas accurately and completely, laying the foundation for subsequent applications such as growth monitoring and yield estimation.
[0227] In the description of the present invention, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.
[0228] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not limitations on the implementation methods of the present invention. For ordinary technicians in this field, other different forms of changes or modifications can be made based on the above description. It is impossible to list all the implementation methods here. All obvious changes or modifications derived from the technical solution of the present invention are still within the scope of protection of the present invention.
Claims
1. A target area extraction device, characterized in that: include: An image input module, a trained three-dimensional feature extraction network connected in series with the image input module, a trained one-dimensional feature extraction network connected in series with the image input module and in parallel with the three-dimensional feature extraction network, a data fusion module, and an image output module, wherein: The image input module is used to process a plurality of multispectral remote sensing images to generate first image data input to the three-dimensional feature extraction network and second image data input to the one-dimensional feature extraction network: The three-dimensional feature extraction network is used to perform three-dimensional feature extraction on the first input image data and then output a first feature vector associated with the multispectral remote sensing image; The one-dimensional feature extraction network is used to perform one-dimensional feature extraction on the second input image data and then output a second feature vector that distinguishes the target area from the non-target area: The data fusion module is used to fuse and classify the second eigenvector and the first eigenvector using a preset weight parameter to obtain image category probabilities corresponding to the second eigenvector and the first eigenvector; The image output module is configured to output a target area image including a target area according to the image category probability.
2. The device according to claim 1, characterized in that The image input module includes: The first image data generating unit is configured to slice the multispectral remote sensing image at preset slice values to generate three-dimensional slice data, wherein the three-dimensional slice data serves as the first image data; The second image data generating unit is configured to generate first one-dimensional data, second one-dimensional data, and third one-dimensional data according to the band parameters of the spectral remote sensing image, and to combine the first one-dimensional data, the second one-dimensional data, and the third one-dimensional data in the dimension of the spectrum to generate the second image data; The band parameters include normalized vegetation index, ratio vegetation index and spectral band number, where: The normalized vegetation index is obtained according to the following formula: ; The ratio vegetation index is obtained according to the following formula: ; in, is the reflectance value of band q in the spectral remote sensing image, is the reflectance value of band p in the spectral remote sensing image; The first one-dimensional data includes a normalized vegetation index determined based on any two bands; The second one-dimensional data includes a ratio vegetation index determined based on any two bands; The third one-dimensional data includes spectral data of each wavelength band.
3. The device according to claim 2, characterized in that The second image data generating unit is further used to normalize and convolve the first one-dimensional data, normalize and convolve the second one-dimensional data, and normalize and convolve the third one-dimensional data, and combine the normalized and convoluted first one-dimensional data, the second one-dimensional data, and the third one-dimensional data in the spectral dimension to obtain the second image data.
4. The device according to claim 1, characterized in that The three-dimensional feature extraction network includes: a three-dimensional convolution module, at least one three-dimensional feature extraction module connected in series with the three-dimensional convolution module, and a first fully connected layer connected in series with the three-dimensional feature extraction module, in, The three-dimensional convolution module is used to perform three-dimensional convolution on the input first image data to output a three-dimensional convolution feature map. The three-dimensional feature extraction module is used to take the three-dimensional convolution feature map as input, perform feature association, calibration and pooling on the three-dimensional convolution feature map, and then output a three-dimensional pooling feature map; The first fully connected layer is used to take the three-dimensional pooled feature map as input, connect the three-dimensional pooled feature map, and output the first feature vector.
5. The device according to claim 4, characterized in that The three-dimensional feature extraction module includes a spatial self-attention unit, a three-dimensional residual unit, a first feature calibration unit, and a first pooling layer connected in series. The other end of the spatial self-attention unit is connected in series with the three-dimensional convolution module. The other end of the first pooling layer is connected in series with the first fully connected layer. The spatial self-attention unit is used to take the three-dimensional convolution feature map as input, calculate the correlation features between the central pixel and other pixels of the three-dimensional convolution feature map, and output a spatial correlation feature map; The three-dimensional residual unit is used to perform residual calculation based on the spatial correlation feature map to obtain a three-dimensional residual feature map output; The first feature calibration unit is configured to take the three-dimensional residual feature map as input, determine the correlation between adjacent channels in the three-dimensional residual feature map, and output a three-dimensional calibration feature map; The first pooling layer is used to perform pooling compression on the three-dimensional calibration feature map, thereby outputting a three-dimensional pooling feature map.
6. The device according to claim 5, characterized in that The spatial self-attention unit includes: The sub-feature map generation unit is configured to take the three-dimensional convolutional feature map as input and output a first sub-feature map, a second sub-feature map, and a third sub-feature map, respectively, where each sub-feature map satisfies the following formula: ; ; in, Corresponding to the first sub-feature graph, the second sub-feature graph and the third sub-feature graph respectively, is the activation function, is the functional relationship of the BN layer function with respect to X, F is the feature map, represents 3D convolution, and Represents the weights and biases, X is the input three-dimensional convolution feature map, and are the mean and variance of the input three-dimensional convolution feature map, 、 and is the training parameter; The spatial self-attention weight map generating unit is used to reshape the first sub-feature map and the second sub-feature map to generate corresponding first sub-feature matrices and second sub-feature matrices, and perform product calculation and normalization on the transposed matrix of the first sub-feature matrix and the second sub-feature matrix to generate a spatial self-attention weight map; The original feature size map generating unit is used to multiply the third sub-feature map by the transposed matrix of the spatial self-attention weight map and perform feature reconstruction to obtain an original feature size map; The spatial correlation feature map generating unit is used to multiply the original feature size map by a scale coefficient, and add the product result to the input three-dimensional convolution feature map to obtain a spatial correlation feature map and output it.
7. The device according to claim 5, characterized in that The first feature calibration unit includes: The first descriptor determination unit is configured to take the three-dimensional residual feature map as input and determine the first descriptor corresponding to each channel of the three-dimensional residual feature map by the following calculation formula: : ; in, is the input three-dimensional residual feature map, is the height or width of the input three-dimensional residual feature map, is the position of the pixel in the input three-dimensional residual feature map; The first feature recalibration vector unit is used to determine a first correlation vector between multiple adjacent channels of the three-dimensional residual feature map according to the first descriptor and the following calculation formula: : , Where M is an integer, is the training parameter of the first feature recalibration vector unit in the mth three-dimensional feature extraction module in the M three-dimensional feature extraction modules, is the descriptor of the c-th channel of the first feature recalibration vector unit in the m-th three-dimensional feature extraction module in the M three-dimensional feature extraction modules; The three-dimensional calibration feature map output unit is used to determine the correlation of all channels of the three-dimensional residual feature map according to the first correlation vector by the following formula, thereby outputting the three-dimensional calibration feature map ; ; in, is the pixel point of the cth channel in the three-dimensional calibration feature map.
8. The device according to claim 1, characterized in that The one-dimensional feature extraction network includes: a one-dimensional convolution module, at least one one-dimensional feature extraction module connected in series with the one-dimensional convolution module, and a second fully connected layer connected in series with the one-dimensional feature extraction module. The one-dimensional convolution module is used to perform one-dimensional convolution on the input second image data to output a one-dimensional convolution feature map. The one-dimensional feature extraction module is used to perform residual calculation, feature calibration and pooling on the one-dimensional convolution feature map and output a one-dimensional pooling feature map. The second fully connected layer is used to take the one-dimensional pooled feature map as input, connect the one-dimensional pooled feature map, and output the second feature vector.
9. The device according to claim 8, characterized in that The one-dimensional feature extraction module includes a one-dimensional residual unit, a second feature calibration unit, and a second pooling layer connected in series. The other end of the one-dimensional residual unit is connected in series with the one-dimensional convolution module. The other end of the second pooling layer is connected in series with the second fully connected layer. The one-dimensional residual unit is used to perform residual calculation according to the one-dimensional convolution feature map to obtain a one-dimensional residual feature map output; The second feature calibration unit is configured to take the one-dimensional residual feature map as input, determine the correlation between adjacent channels in the one-dimensional residual feature map, and output a one-dimensional calibration feature map; The second pooling layer is used to perform pooling compression on the one-dimensional calibration feature map, thereby outputting a one-dimensional pooling feature map.
10. A method for extracting a target area image using the target extraction device according to any one of claims 1 to 9, characterized in that: The method comprises: Processing a plurality of multispectral remote sensing images to generate first image data input to the three-dimensional feature extraction network and second image data input to the one-dimensional feature extraction network: Performing three-dimensional feature extraction on the first input image data and outputting a first feature vector associated with the multispectral remote sensing image; After performing one-dimensional feature extraction on the second input image data, a second feature vector is output to distinguish the target area from the non-target area: fusing and classifying the second eigenvector and the first eigenvector using a preset weight parameter to obtain image category probabilities corresponding one-to-one to the second eigenvector and the first eigenvector; A target area image including the target area is output according to the image category probability.
Citation Information
Patent Citations
Hyper-object information based remote sensing image target extraction method and device, and medium
CN110287962A
Hyperspectral remote sensing image recognition method and device, electronic equipment and storage medium
CN113822207A