An orchard land space information identification method and device
Patent Information
- Application Number
- CN202410151838.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-02
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-02-02
AI Technical Summary
[0041]本发明还提供一种计算机程序产品,包括计算机程序,所述计算机程序被处理器执行时实现如上述任一种所述苹果园地空间信息识别方法。
Smart Images

Figure CN118072162B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a method and apparatus for spatial information recognition in apple orchards. Background Technology
[0002] Due to the influence of numerous factors such as climate and region on fruit tree cultivation, there are few studies on orchard remote sensing monitoring and the methods are limited. For example, data such as Landsat TM / ETM+, SPOT, GF-2, Sentinel satellite remote sensing data and UAV imagery data are used to extract orchard information such as fruit tree growth and planting area.
[0003] Currently, orchards can usually be extracted directly by using single-temporal high-resolution imagery through the linear relationship between ground features and their characteristics. Combining machine learning methods with single-temporal remote sensing imagery for fruit tree remote sensing monitoring can significantly improve the efficiency and accuracy of fruit tree identification and monitoring.
[0004] However, compared to other orchard spatial information extraction, there is relatively little research on remote sensing spatial information extraction for apple orchards. The methods mentioned above mostly use low-to-medium spatial resolution remote sensing data and conventional supervised classification methods, which have limited recognition accuracy and are difficult to extract spatiotemporal information, and are not effective for extracting small-target apple orchards. Summary of the Invention
[0005] This invention provides a method and apparatus for identifying spatial information of apple orchards, which addresses the shortcomings of existing technologies that mostly use low-to-medium spatial resolution remote sensing data and conventional supervised classification methods, resulting in limited identification accuracy, difficulty in extracting spatiotemporal information, and poor extraction effect for small-target apple orchards, thereby improving the extraction effect of small-target apple orchards.
[0006] This invention provides a method for identifying spatial information of apple orchards, comprising:
[0007] Acquire single-temporal and multi-temporal remote sensing images of the apple orchard to be identified, and extract the single-temporal feature representation of the single-temporal remote sensing image and the multi-temporal feature representation of the multi-temporal remote sensing image of the apple orchard to be identified;
[0008] The single-temporal feature representation is fused with the multi-temporal feature representation to obtain the fused feature representation of the remote sensing image of the apple orchard to be identified;
[0009] The fusion feature representation of the remote sensing image of the apple orchard to be identified is input into the spatial information recognition model to obtain the spatial information recognition result corresponding to the remote sensing image of the apple orchard to be identified, which is output by the spatial information recognition model.
[0010] The spatial information recognition model is trained based on the sample fusion feature representation of the sample apple orchard remote sensing image and the spatial information identifier of the sample apple orchard remote sensing image.
[0011] According to the method for identifying spatial information of apple orchards provided by the present invention, the spatial information identification model includes a feature extraction layer and a loss calculation layer;
[0012] Correspondingly, the step of inputting the fused feature representation of the remote sensing image of the apple orchard to be identified into the spatial information recognition model, and obtaining the spatial information recognition result corresponding to the remote sensing image of the apple orchard to be identified output by the spatial information recognition model, specifically includes:
[0013] The fused feature representation of the remote sensing image of the apple orchard to be identified is input into the feature extraction layer to obtain the semantic features of the remote sensing image of the apple orchard to be identified output by the feature extraction layer.
[0014] The semantic features are input into the loss calculation layer to obtain the spatial information recognition result of the multispectral remote sensing image of the apple orchard to be identified, output by the loss calculation layer.
[0015] According to the apple orchard spatial information identification method provided by the present invention, the feature extraction layer includes a feature encoding layer, a weight allocation layer, and a feature decoding layer;
[0016] Correspondingly, the step of inputting the fused feature representation of the remote sensing image of the apple orchard to be identified into the feature extraction layer to obtain the semantic features of the remote sensing image of the apple orchard to be identified output by the feature extraction layer specifically includes:
[0017] The fusion feature representation of the remote sensing image of the apple orchard to be identified is input into the feature coding layer to obtain the multi-channel feature coding information of the multispectral remote sensing image of the apple orchard to be identified output by the feature coding layer.
[0018] The feature encoding information of each channel is input into the weight allocation layer to obtain the weight allocation value of the feature encoding information of each channel output by the weight allocation layer;
[0019] The feature encoding information of each channel and the weight allocation value of the feature encoding information of each channel are input to the feature decoding layer to obtain the semantic features of the remote sensing image of the apple orchard to be identified output by the feature decoding layer.
[0020] According to the apple orchard spatial information identification method provided by the present invention, the weight allocation layer includes a channel weight allocation layer, a spatial weight allocation layer and a weight fusion layer;
[0021] Correspondingly, the step of inputting the feature encoding information of each channel into the weight allocation layer to obtain the weight allocation value of the feature encoding information of each channel output by the weight allocation layer includes:
[0022] The feature encoding information of each channel is input into the channel weight allocation layer to obtain the channel weight allocation value of the feature encoding information of each channel output by the channel weight allocation layer;
[0023] The feature encoding information of each channel is input into the spatial weight allocation layer to obtain the spatial weight allocation value of the feature encoding information of each channel output by the spatial weight allocation layer;
[0024] The channel weight allocation value and the spatial weight allocation value of the feature encoding information of each channel are input into the weight fusion layer to obtain the weight allocation value of the feature encoding information of each channel.
[0025] According to the spatial information identification method for apple orchards provided by the present invention, the step of inputting the feature encoding information of each channel and the weight allocation value of the feature encoding information of each channel into the feature decoding layer to obtain the semantic features of the multispectral remote sensing image of the apple orchard to be identified output by the feature decoding layer includes:
[0026] Based on the weight allocation values of the feature encoding information of each channel, the weights of the semantic features of each channel are determined, wherein the semantic features of each channel are obtained by upsampling, convolution and batch normalization of the feature encoding information of each channel;
[0027] Based on the weights of the semantic features in each channel, the semantic features in each channel are weighted to obtain the semantic features of the remote sensing image of the apple orchard to be identified.
[0028] According to the spatial information recognition method for apple orchards provided by the present invention, the step of inputting the semantic features into the loss calculation layer to obtain the spatial information recognition result of the remote sensing image of the apple orchard to be identified output by the loss calculation layer includes:
[0029] Based on the semantic features, the loss value of the spatial information recognition model is determined;
[0030] Based on the loss value and the preset loss function, the spatial information recognition result is determined.
[0031] According to the spatial information identification method for apple orchards provided by the present invention, the method extracts multi-temporal feature representations of multi-temporal remote sensing images of the apple orchard to be identified, including:
[0032] Extract the temporal and spectral index features of the multi-temporal remote sensing images of the apple orchard to be identified;
[0033] By fusing the temporal features and spectral index features, a multi-temporal feature representation of the apple orchard to be identified is obtained.
[0034] The present invention also provides an apple orchard spatial information identification device, comprising:
[0035] The feature extraction module is used to acquire single-temporal and multi-temporal remote sensing images of the apple orchard to be identified, and to extract the single-temporal feature representation of the single-temporal remote sensing image of the apple orchard to be identified and the multi-temporal feature representation of the multi-temporal remote sensing image.
[0036] The feature fusion module is used to fuse the single-temporal feature representation with the multi-temporal feature representation to obtain the fused feature representation of the remote sensing image of the apple orchard to be identified;
[0037] The spatial information recognition module is used to input the fusion feature representation of the remote sensing image of the apple orchard to be identified into the spatial information recognition model, and obtain the spatial information recognition result corresponding to the remote sensing image of the apple orchard to be identified output by the spatial information recognition model.
[0038] The spatial information recognition model is trained based on the sample fusion feature representation of the sample apple orchard remote sensing image and the spatial information identifier of the sample apple orchard remote sensing image.
[0039] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the apple orchard spatial information recognition method as described above.
[0040] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the apple orchard spatial information identification method as described above.
[0041] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the apple orchard spatial information recognition method as described above.
[0042] The present invention provides a method and apparatus for spatial information identification of apple orchards. Based on the feature representation of multispectral remote sensing images of apple orchards, it performs spatial information identification, which solves the problems of limited identification accuracy and difficulty in mining spatiotemporal information in existing technologies that use low-to-medium spatial resolution remote sensing data and conventional supervised classification methods for spatial information identification, and the poor extraction effect on small target apple orchards. It fully mines the deep features of multispectral remote sensing images of apple orchards, thereby improving the effect of spatial information identification for small target apple orchards. Attached Figure Description
[0043] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0044] Figure 1 This is a flowchart illustrating the method for identifying spatial information in apple orchards provided in an embodiment of the present invention;
[0045] Figure 2 This is a flowchart illustrating the remote sensing image preprocessing method provided in an embodiment of the present invention;
[0046] Figure 3 This is a graph showing the variation of the EVI (Electrospectral Index) of different land features over time, provided in an embodiment of the present invention.
[0047] Figure 4 A flowchart of a spatial information recognition method based on a spatial information recognition model provided in an embodiment of the present invention;
[0048] Figure 5 This is a diagram illustrating the training process of the spatial information recognition model provided in an embodiment of the present invention.
[0049] Figure 6 A flowchart illustrating the construction of a dataset for identifying spatial information of apple orchards;
[0050] Figure 7 This is a flowchart of label labeling provided in an embodiment of the present invention;
[0051] Figure 8 A flowchart of an apple orchard extraction method based on an improved semantic segmentation model provided in an embodiment of the present invention;
[0052] Figure 9 This is a schematic diagram of the improved SegNet model structure provided in an embodiment of the present invention;
[0053] Figure 10 This is a schematic diagram of the structure of the apple orchard spatial information recognition device provided in an embodiment of the present invention;
[0054] Figure 11 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0055] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0056] In practical applications, monitoring apple-producing areas has traditionally relied on door-to-door surveys and on-site sampling to obtain statistical information on orchard planting. This approach is time-consuming, costly, and cumbersome, lacking spatial information and hindering continuous and dynamic monitoring. Currently, large-scale apple orchard extraction largely utilizes low-to-medium spatial resolution remote sensing data such as Landsat TM / ETM+, often employing conventional supervised classification methods with limited accuracy. While employing high-accuracy deep learning methods presents challenges such as the significant time and cost of manually labeling datasets and poor extraction of small features from satellite imagery. Therefore, further research and exploration are needed to develop an easily implementable apple orchard extraction method that leverages high-accuracy remote sensing image classification techniques to extract rich information from remote sensing images.
[0057] To address this issue, embodiments of the present invention provide a method for identifying spatial information of apple orchards. Figure 1 This is a flowchart illustrating the method for identifying spatial information in apple orchards provided in an embodiment of the present invention, as shown below. Figure 1 As shown, the method includes:
[0058] Step 110: Obtain single-temporal remote sensing images and multi-temporal remote sensing images of the apple orchard to be identified, and extract the single-temporal feature representation of the single-temporal remote sensing image of the apple orchard to be identified and the multi-temporal feature representation of the multi-temporal remote sensing image.
[0059] Step 120: The single-temporal feature representation and the multi-temporal feature representation are fused to obtain the fused feature representation of the remote sensing image of the apple orchard to be identified;
[0060] Step 130: Input the fusion feature representation of the remote sensing image of the apple orchard to be identified into the spatial information recognition model to obtain the spatial information recognition result corresponding to the remote sensing image of the apple orchard to be identified, output by the spatial information recognition model; wherein, the spatial information recognition model is trained based on the sample fusion feature representation of the sample apple orchard remote sensing image and the spatial information identifier of the sample apple orchard remote sensing image.
[0061] The above steps will be explained in detail below with reference to specific embodiments.
[0062] Step 110: Obtain single-temporal remote sensing images and multi-temporal remote sensing images of the apple orchard to be identified, and extract the single-temporal feature representation of the single-temporal remote sensing image of the apple orchard to be identified and the multi-temporal feature representation of the multi-temporal remote sensing image.
[0063] In this step, "apple orchard to be identified" refers to identifying the spatial information of an apple orchard. The multispectral remote sensing image of the apple orchard to be identified is the remote sensing image for which spatial information identification is required. This multispectral remote sensing image can be acquired by a remote sensing satellite, which can be a high-resolution Sentinel-2, QuickBird, SPOT, or other remote sensing satellites. After acquiring the multispectral remote sensing image data of the apple orchard to be identified, the data needs to be preprocessed. The preprocessing process varies depending on the remote sensing satellite used to acquire the data. For example, Sentinel-2 has five product levels: Level-0, Level-1A, Level-1B, Level-1C, and Level-2A. Level-1C is orthorectified image data with geometric correction, while Level-2A is surface reflectance data obtained after atmospheric correction and radiometric calibration based on Level-1C. If Google Earth ultra-high-resolution imagery is used, the image data includes red, green, and blue bands and uses the WGS84 Web Mercator projection. In other words, if the selected data is the Sentinel-2L2A product, which has already undergone preprocessing such as atmospheric correction and radiometric calibration by the data provider platform, no further preprocessing is required. However, issues such as high cloud cover in some areas, missing data, and data unsuitable for semantic segmentation still exist, necessitating satellite image preprocessing to convert the image data into standard data suitable for image analysis. If data from other remote sensing satellites is selected, atmospheric correction and radiometric calibration may be required.
[0064] Based on this, this embodiment takes the Sentinel-2L2A product as an example. Figure 2 This is a schematic flowchart of the remote sensing image preprocessing method provided in an embodiment of the present invention, as shown below. Figure 2 As shown, the preprocessing method for remote sensing images acquired in this embodiment may include the following steps:
[0065] Step 210, Cloud removal and image fusion:
[0066] Cloud cover percentages were filtered based on the Sentinel-2 image metadata, removing images with cloud cover greater than 10%. Then, cloud masking was applied to the filtered images using the QA60 quality assessment band. Since some areas experienced image gaps during certain time periods after cloud removal, all cloud-removed images within the selected time periods were fused using median composite to generate a complete image.
[0067] Step 220, missing value completion and image cropping:
[0068] For areas where a small number of gaps remain after image fusion, the same cloud removal and image fusion processes are performed on historical images of the same area and time period. These images are then mosaicked with the images obtained in this embodiment to fill in the missing values. Simultaneously, the required portions of the image are cropped using vector boundaries.
[0069] Step 230, Band Normalization and Resampling:
[0070] If this embodiment utilizes different spectral index features from different time phases, the calculated results of the selected spectral index features often have inconsistent or even significantly different orders of magnitude. Directly inputting these features into the model would lead to different weights between channels, affecting the training of the spatial information recognition model. Therefore, the preferred features are pre-normalized. Furthermore, because Sentinel-2 data only has a spatial resolution of 10m for the R, G, B, and NIR bands, the spatial resolution of the selected spectral index features may not all be 10m. Therefore, some features need to be resampled to 10m to ensure consistency in spatial resolution across bands.
[0071] Step 240, Band Combining and Format Conversion:
[0072] The original bands and selected features were combined using the band synthesis tool in ENVI to obtain a multi-band file. Due to the large data volume of the entire Sentinel-2 remote sensing image of the Yantai area and its inconvenience for post-processing, it was necessary to convert it into a data format that could be input into the spatial information recognition model. This study used Python tools, combined with the characteristics of Sentinel-2 imagery, to convert the image from signed 32-bit floating-point numbers to unsigned 8-bit integers using percentage truncation and exponential transformation, and then performed pixel depth conversion.
[0073] It is worth noting that the above preprocessing process and steps are only a specific embodiment, and the preprocessing process and steps are not limited to the above embodiment. They may vary depending on the data acquisition method, and the embodiments of the present invention do not specifically limit them.
[0074] After preprocessing, feature information of the multispectral remote sensing image of the apple orchard to be identified is extracted to obtain the optimal classification time phase and feature representation of the apple orchard.
[0075] Optionally, based on the above embodiments, single-temporal original band images and multi-temporal optimized feature images with multi-temporal features can be generated based on the obtained optimal classification temporal phase and feature representation. Three classic semantic segmentation models, FCN-8s, U-Net, and SegNet, are constructed and trained on the two different images respectively. The accuracy is evaluated using a test set, and then a comparative analysis is performed. Based on the comparison results, it is determined that when performing spatial information identification of apple orchards, the classification effect of using multi-temporal Sentinel-2 images and SegNet network models simultaneously is optimal.
[0076] Specifically, extracting multi-temporal feature representations from the multi-temporal remote sensing images of the apple orchard to be identified includes:
[0077] The temporal features and spectral index features of the multi-temporal remote sensing image of the apple orchard to be identified are extracted; the temporal features and spectral index features are fused to obtain the multi-temporal feature representation of the multi-temporal remote sensing image of the apple orchard to be identified.
[0078] In this step, existing research often uses data augmentation methods to address the small sample size problem in remote sensing image feature extraction. However, data augmentation does not significantly increase the effective feature information of the dataset. Utilizing feature information from bands other than the RGB bands in multispectral satellite imagery can compensate for the insufficient spectral information in the original RGB data, enhancing the model's generalization ability with limited training data. Since spectral indices are parameters calculated by combining reflectance at different wavelengths, the influence of background conditions on spectral reflectance can be reduced, making it more sensitive than the original band data in vegetation extraction and identification. Some researchers have found that using typical temporal features, spectral index features, topographic features, and texture features that distinguish specific land cover categories from other land cover categories as input data for the network model improves or matches the classification performance compared to directly using the original bands of multispectral data.
[0079] Current research mostly focuses on expanding feature channels based on the original bands of single-temporal images, with limited research utilizing different features from multi-temporal images for feature channel expansion. However, when using multi-temporal remote sensing images for deep learning research, directly overlaying all bands from different temporal images into the network model results in a massive amount of image data. This leads to a significant increase in the number of model parameters, requiring more computation time and storage space, and placing high demands on computer hardware and software configurations.
[0080] To address the aforementioned issues, this invention comprehensively utilizes multi-temporal information from multispectral images by introducing different temporal feature variables. This provides the network model with usable prior knowledge from different temporal phases, thereby improving the model's classification performance while maintaining a low parameter computation load. This avoids the problem of low computational efficiency due to massive data volume, saving computation time and storage space. Therefore, this embodiment comprehensively utilizes the temporal features and spectral index features of Sentinel-2 multispectral data. By extracting different features from different temporal phases and using them as the original input of the model, it explores the effect of introducing multi-temporal features on improving the classification accuracy of the network model.
[0081] For example, based on the phenological information of apple trees and other vegetation, and considering the characteristics of apple growth, characteristic values of typical land features can be calculated using six images from different time periods. Then, a statistical method based on mean squared error is used for feature optimization. The difference between the characteristic values of apple orchards and the characteristic values of the most easily confused land features is used as the feature optimization evaluation index to select characteristic variables that are beneficial for extracting apple orchard features from a predetermined area. Table 1 shows the nine characteristic variables selected in this embodiment of the invention.
[0082] Table 1
[0083]
[0084] In Table 1, B2, B3, B4, B5, B6, B7, and B8 represent the reflectance of the corresponding bands in the Sentinel-2 data. If the center wavelength of the Sentinel-2 data does not meet the wavelength used in the characteristic calculation formula, the reflectance of the band at the nearest neighbor wavelength is used instead.
[0085] Figure 3 This is a graph showing the change of EVI (Electrospectral Index) of different land features over time, provided in an embodiment of the present invention. Figure 3 As shown, the Enhanced Vegetation Index (EVI) can reflect changes in vegetation canopy structure while suppressing the influence of background conditions such as soil and atmosphere, and is often used to monitor vegetation density. The EVI of apple orchards at flowering stage (T2) integrates the spectral characteristics of flowers and leaves, while cherry trees are in the post-flowering stage. The EVI of apple trees differs from that of cherry trees and woodlands, which are easily confused with other land cover features, which is helpful in distinguishing apple orchards from other land cover features.
[0086] Similarly, the ratio vegetation index (RVI) reflects the difference in refractive index of vegetation in the near-infrared and red light bands, and enhances the difference in radiation between vegetation and the soil background. It is often used to monitor vegetation changes or reflect land use changes. Likewise, during the apple flowering period (T2), the spectral values of apples, cherries, and woodlands show significant differences in normalized RVI compared to nearby land features, which is helpful in distinguishing apple orchards from other land features.
[0087] The modified chlorophyll uptake ratio (MCARI) can reflect changes in vegetation chlorophyll content while suppressing the influence of canopy non-photosynthetic substances and soil reflectance. Similarly, during the apple flowering period (T2), apples show significant differences in MCARI compared to similar ground features such as cherries and woodlands, which helps distinguish apple orchards from other ground features.
[0088] The improved red-edge simple ratio index (MRESR) can reflect the chlorophyll content of vegetation and is relatively insensitive to changes in vegetation species and leaf structure, making it applicable to a wide range of vegetation chlorophyll monitoring. During the apple fruit coloring stage (T5), vegetation easily confused with apples, such as peanuts and pear trees, is already in harvest time, and their MRESRs show significant differences from those of apple trees. This is helpful in distinguishing apple orchards from other land cover.
[0089] Ground chlorophyll index (MTCI) can reflect changes in vegetation chlorophyll content. It performs better in areas with higher chlorophyll concentration. During the apple budding stage (T1), cherry and pear trees are in the flowering stage, and their MTCI is significantly different from that of apple trees. This is helpful in distinguishing apple orchards from other ground features.
[0090] The Normalized Difference Red Edge Vegetation Index (NDre3) can reflect the degree of canopy cover and is particularly suitable for high-density vegetation areas. It is often used to monitor crops in the ripening stage. During the apple juvenile stage (T3), the easily confused cherry has entered the ripening and harvesting stage. The NDre3 of apples and other land cover in the study area has obvious differences, which is helpful to distinguish apple orchards from other land cover.
[0091] The Normalized Difference Red Edge Vegetation Index (NDVIre32) can reflect the health status of vegetation and is particularly suitable for mature crops with high chlorophyll concentration. Therefore, there are obvious differences in NDVIre32 between apples and other land cover during the young fruit stage (T3) and the fruit ripening stage (T6), which is helpful in distinguishing apple orchards from other land cover.
[0092] Near-infrared reflectance (NIRv) of vegetation, derived from the normalized difference vegetation index (NDVI), can mitigate the influence of soil background in high biomass areas and is often used to estimate the total primary productivity (GPP) of an ecosystem. Similarly, during the apple flowering period (T2), apple trees show significant differences in NIRv compared to nearby ground features such as cherry trees and woodlands, which is helpful in distinguishing apple orchards from other ground features.
[0093] The triangular vegetation index (TVI) has higher saturation resistance than NDVI and can better reflect the chlorophyll content of crops. During the apple budding stage (T1), easily confused ground cover such as peanuts is usually not yet sown. The TVI of apples is significantly different from that of ground cover with similar spectral values such as peanuts and grasslands, which is helpful in distinguishing apple orchards from other ground cover.
[0094] Step 120: The single-temporal feature representation and the multi-temporal feature representation are fused to obtain the fused feature representation of the remote sensing image of the apple orchard to be identified;
[0095] Based on the above embodiments, the multi-temporal preferred feature image is based on the original single-temporal band images B2(T6), B3(T6), and B4(T6), with the addition of the spectral index features of 10 different temporal phases selected above: EVI(T2), RVI(T2), MCARI(T2), MRESR(T5), MTCI(T1), NDre3(T3), NDVIre32(T3)(T6), NIRv(T2), and TVI(T1).
[0096] It is worth noting that, compared with single-temporal raw band image data, multi-temporal image data has rich phenological time series information. Based on the phenological characteristics of typical land cover in the study area, multi-temporal spectral indices and other characteristic information are constructed and added to the classification model, which can highlight the changing land cover, distinguish the change categories, and improve the classification accuracy.
[0097] Step 130: Input the fusion feature representation of the remote sensing image of the apple orchard to be identified into the spatial information recognition model to obtain the spatial information recognition result corresponding to the remote sensing image of the apple orchard to be identified, output by the spatial information recognition model; the spatial information recognition model is trained based on the sample fusion feature representation of the sample apple orchard remote sensing image and the spatial information identifier of the sample apple orchard remote sensing image.
[0098] In existing technologies, when performing spatial information identification based on feature representations of multispectral remote sensing images, the commonly used low-to-medium spatial resolution remote sensing data employs conventional supervised classification methods. Here, low-to-medium spatial resolution remote sensing data is typically single-temporal remote sensing data. Vegetation information in remote sensing images is largely reflected by spectral differences in the vegetation canopy. Single-temporal remote sensing images are significantly affected by topography, shadows, and atmospheric conditions. Furthermore, due to phenomena such as different land cover types exhibiting different spectral features, conventional supervised classification methods struggle to achieve satisfactory extraction results for apple orchards from single-temporal Sentinel-2 remote sensing images.
[0099] In this embodiment of the invention, multi-temporal features are added to eliminate the influence of some environmental factors, so as to highlight the differences in the growth characteristics of apple orchards and other land features over time. Based on the above-mentioned optimal classification temporal and feature optimization research results, a multi-temporal remote sensing image composed of single-temporal original band image and multiple optimized features is constructed.
[0100] Furthermore, the spatial information recognition model is a pre-trained model used to determine whether the area to be identified in the multispectral remote sensing image is the spatial information of an apple orchard, based on the sample feature representation of the input apple orchard multispectral remote sensing image, and output the spatial information recognition result. In this step, the spatial information recognition result can be "apple area" or "non-apple area," where "apple area" indicates that the area to be identified in the multispectral remote sensing image is an apple orchard, and "non-apple area" indicates that the area to be identified in the multispectral remote sensing image is not an apple orchard. Alternatively, different colors can be used to distinguish between apple orchards and non-apple orchards; this embodiment of the invention does not specifically limit this.
[0101] Optionally, the spatial information recognition model includes a feature extraction layer and a loss calculation layer. Correspondingly, Figure 4 A flowchart of the spatial information recognition method based on a spatial information recognition model provided in the embodiments of the present invention is shown below. Figure 4 As shown, step 130 specifically includes:
[0102] Step 131: Input the fused feature representation of the remote sensing image of the apple orchard to be identified into the feature extraction layer to obtain the semantic features of the remote sensing image of the apple orchard to be identified output by the feature extraction layer.
[0103] Step 132: Input the semantic features into the loss calculation layer to obtain the spatial information recognition result of the multispectral remote sensing image of the apple orchard to be identified, output by the loss calculation layer.
[0104] Based on any of the above embodiments, a spatial information recognition model can be pre-trained before performing step 130. Figure 5 This is a diagram illustrating the training process of the spatial information recognition model provided in an embodiment of the present invention, as shown below. Figure 5 As shown, the spatial information recognition model can be trained in the following way:
[0105] Step 510, Sample set construction:
[0106] In this step, the construction of the apple orchard spatial information recognition dataset mainly consists of three parts: sample area delineation, label annotation, and dataset partitioning and data augmentation. Figure 6 A flowchart for constructing a dataset for spatial information identification of apple orchards, such as... Figure 6 As shown, step 510 may specifically include:
[0107] Step 511, Delineation of the sample region:
[0108] Specifically, to ensure the typicality and accuracy of the apple orchard spatial information identification dataset, this embodiment of the invention selects parts of Qixia City, Laiyang City, and Muping District under the jurisdiction of Yantai City as sample areas for the spatial information identification dataset, based on two field surveys conducted in 2021 and 2022 and the apple planting situation in various counties and districts under the jurisdiction of Yantai City. The selected areas cover most of the field survey sample point data, have a large number of apple planting areas, and also include grassland, woodland, cultivated land, other orchards, water areas, residential areas, and other building land categories. The rich variety of background features can effectively improve the extraction accuracy of apple orchard information.
[0109] Step 512, Labeling:
[0110] This invention proposes a method for constructing a spatial information identification dataset for apple orchards based on machine learning and multi-source data fusion. This method comprehensively utilizes the multi-source information features of Sentinel-2 multispectral imagery and Google Earth ultra-high resolution imagery. By establishing two classic machine learning classification models, apple orchards in the sample area are extracted respectively, and the classification results of different models are fused. The fused classification results are visually inspected using Google Earth ultra-high resolution imagery to correct unreasonable classification areas. Finally, the corrected vector apple orchard classification results are converted into satellite imagery data label maps in RGB channel unsigned bit raster format.
[0111] Based on the above embodiments, Figure 7 This is a flowchart of label annotation provided in an embodiment of the present invention, such as... Figure 7 As shown, step 512 can specifically include:
[0112] Step 512-1, Pre-extraction of random forest in apple orchards from Sentinel-2 images:
[0113] This embodiment uses the Random Forest classification method for pre-extraction of apple orchards from Sentinel-2 imagery. The Random Forest algorithm is an ensemble learning algorithm based on the divide-and-conquer principle, which can fully utilize the multispectral information of satellite remote sensing imagery and has advantages such as high classification accuracy and high training efficiency in the field of land use classification. The Random Forest classification method requires randomly and uniformly distributed samples. Therefore, in the sample selection process, a 10km grid was established to uniformly select samples. Simultaneously, high-resolution Google Earth imagery was used as a reference base map, and sample points were selected in combination with field-measured sample point data. In addition, the Normalized Difference Vegetation Index (NDVI) curves of various localities in the sample area were also used for point selection. The NDVI curve of the apple-growing area in Yantai shows a gradual increase starting in March, typically reaching above 0.5 in May, peaking around September, and then rapidly declining. An area sampling method was used to select random forest classification samples. Area sampling integrates the mean values of all pixels within the area, reducing the uncertainty of point sampling. Ultimately, 500 apple samples and 500 non-apple land samples, including woodland, cultivated land, other orchards, residential land and other building land, were selected.
[0114] Due to heavy cloud cover in some areas of Yantai City, resulting in missing image data, Sentinel-2 images from April 1, 2022 to May 31, 2022 were merged into a single image using median processing for random forest classification. Thirty-three different spectral index features and topographic features were selected and added to the random forest classification model for classification. After classification, post-processing such as small patch merging and decomposition was performed to obtain the final random forest pre-extraction results.
[0115] Step 512-2, K-means pre-extraction of apple orchard from Google Images:
[0116] Considering that the apple orchards in the Sentinel-2 image have already been extracted using the random forest supervised classification method, and that unsupervised classification methods have the advantage of less influence from human subjective factors and can avoid human classification errors caused by inaccurate sample selection, the K-means classification method in unsupervised classification is used here to pre-extract the sample area of apple orchards from the Google Earth image.
[0117] Using the K-means unsupervised classification tool in ENVI, all land features in the Google Earth imagery sample area were classified into 20 classes, with three iterations to increase accuracy. After classification, the actual distribution of apple orchards in the Google Earth imagery was considered to merge the classes, resulting in two categories: apple orchards and background features. Post-processing was then performed on the classification results, including principal and secondary analysis of small patches, clustering, and filtering. Finally, the post-processed results were resampled to 10m in ArcGIS to match the Sentinel-2 data resolution.
[0118] Step 512-3: Fusion of multi-source distribution results for apple orchards in the sample area:
[0119] Multi-source remote sensing data, due to differences in imaging mechanisms and observation characteristics, can provide observational information from multiple perspectives. While both classification methods mentioned above can generally extract apple orchards, the pre-extraction results from Sentinel-2 and Google imagery exhibit some misclassification. Therefore, this embodiment of the invention performs raster intersection analysis on the two classification results in ArcGIS, thereby fusing the multi-source distribution results. Raster intersection can cross-validate the two apple orchard extraction results, improving the reliability of the extraction results, while simultaneously removing a large number of misclassified areas and reducing the workload of subsequent data correction.
[0120] Step 512-4, Correction of multi-source distribution classification results in apple orchards:
[0121] Even after fusing the classification results from the two methods, some unreasonable areas still exist. The classification results were converted into vector data, and the multi-source distribution classification fusion results were visually inspected in Google Earth Pro, referencing ultra-high resolution orthophotos of the Yantai area, and the vector boundaries of the unreasonable areas were corrected.
[0122] After converting the generated label data into raster format, perform RGB coloring in ArcGIS. Then, use Python to convert the label data into unsigned 8-bit data suitable for deep learning models to obtain the final label file.
[0123] Step 513, Dataset Partitioning and Data Augmentation:
[0124] Remote sensing images are typically large in size, making it impossible to directly train a network using the entire satellite image. This embodiment of the invention uses a Python sliding window cropping method to perform batch cropping of the images. To maximize the use of data from the selected sample regions, the image data and label data of three sample regions are simultaneously and batch-cropped to a size of 512*512 pixels, resulting in 91 sets of labeled image pairs.
[0125] The 91 datasets were divided into training, validation, and test sets in a 15:4:4 ratio. The training, validation, and test sets were then cropped to 256x256 pixels with a 0.4 overlap. Smaller datasets improve training efficiency, resulting in 540 training pairs, 144 validation pairs, and 135 test pairs. Training deep learning models often requires a large amount of sample data and multiple iterations to fully learn sample features and improve model accuracy. However, the amount of data after cropping and partitioning is still insufficient for good training of the semantic segmentation model. Data augmentation is needed to increase data breadth and improve the model's generalization ability.
[0126] It's worth noting that traditional dataset augmentation methods include image flipping, rotation, scaling, translation, mirroring, cropping, contrast and brightness adjustment, blurring, and adding noise. However, for remote sensing imagery, these methods not only fail to improve model training accuracy but can also lead to the loss of necessary local information during the augmentation process, resulting in a decline in model learning ability. Therefore, this study uses three data augmentation methods from geometric transformations—horizontal flipping, vertical flipping, and diagonal mirroring—in Python for data augmentation, ultimately obtaining 2160 training pairs, 576 validation pairs, and 540 test pairs.
[0127] Based on the above embodiments, sample feature representations are extracted from the multispectral remote sensing images of sample apple orchards, and spatial information identifiers corresponding to each sample apple orchard multispectral remote sensing image are obtained. The spatial information identifiers corresponding to each sample apple orchard multispectral remote sensing image are used to indicate whether the sample multispectral remote sensing image represents the spatial information of an apple orchard. The spatial information identifiers can be obtained through the methods described in the above embodiments.
[0128] Step 520, Model Training:
[0129] The initial model is trained based on the sample feature representation of the multispectral remote sensing image of the apple orchard and the corresponding spatial information label of the multispectral remote sensing image of the apple orchard, thereby obtaining a spatial information recognition model. The initial model can be a single classic semantic segmentation model, such as FCN, U-Net, or SegNet.
[0130] In this embodiment, based on the optimal classification time phase and features obtained in the above embodiments, a single-time phase original band image and a multi-time phase optimized feature image with multi-time phase features are generated. Three classic semantic segmentation models, FCN-8s, U-Net and SegNet, are constructed and trained on the two different images respectively. The accuracy is evaluated using a test set, and then a comparative analysis is performed to determine that the classification effect of the SegNet network model is optimal. Figure 8The flowchart of the apple orchard extraction method based on the improved semantic segmentation model provided in the embodiments of the present invention is as follows: Figure 8 As shown, the embodiments of the present invention can use the SegNet network structure as the basic network model. Based on this, in order to address the problem of small target recognition in apple orchard extraction from satellite remote sensing images, the original SegNet network structure is improved by introducing the CBAM dual attention mechanism module and using a novel custom loss function, Focal loss, to improve the information extraction capability of small target segmentation in apple orchards in remote sensing images, thereby improving the classification effect.
[0131] It is worth noting that in the identification and extraction of ground features from remote sensing images, the problem of small targets often arises. This means that remote sensing images contain a large number of tiny targets, but these small targets carry limited information, leading to inaccurate localization of small targets by existing semantic segmentation methods, which directly affects segmentation performance. In this embodiment, the apple orchard extraction is considered a small target within the entire remote sensing image, with a relatively small pixel ratio and a lack of sufficient feature information. This makes network training difficult, hindering the differentiation of the apple orchard from the background or similar ground features. To address this issue, this embodiment of the invention employs a CBAM dual attention mechanism and a Focal Loss loss function to improve the model.
[0132] In this embodiment of the invention, the feature extraction layer includes a feature encoding layer, a weight allocation layer, and a feature decoding layer;
[0133] Correspondingly, the fused feature representation of the remote sensing image of the apple orchard to be identified is input into the feature extraction layer to obtain the semantic features of the remote sensing image of the apple orchard to be identified output by the feature extraction layer, specifically including:
[0134] The fusion feature representation of the remote sensing image of the apple orchard to be identified is input into the feature coding layer to obtain the multi-channel feature coding information of the multispectral remote sensing image of the apple orchard to be identified output by the feature coding layer.
[0135] The feature encoding information of each channel is input into the weight allocation layer to obtain the weight allocation value of the feature encoding information of each channel output by the weight allocation layer;
[0136] It's worth noting that in commonly used semantic segmentation models, as the network depth increases, the semantic information of small targets weakens, leading to the inability to recover small target pixels during subsequent upsampling operations. While shallow networks contain less semantic information than deep networks, they carry a significant amount of semantic information about small targets. Introducing attention mechanisms into shallower layers of the network can help improve the accuracy of small target recognition.
[0137] Based on the above considerations, the weight allocation layer mentioned in the embodiments of the present invention includes a channel weight allocation layer, a spatial weight allocation layer, and a weight fusion layer.
[0138] Specifically, inputting the feature encoding information of each channel into the weight allocation layer to obtain the weight allocation value of the feature encoding information of each channel output by the weight allocation layer includes: inputting the feature encoding information of each channel into the channel weight allocation layer to obtain the channel weight allocation value of the feature encoding information of each channel output by the channel weight allocation layer; inputting the feature encoding information of each channel into the spatial weight allocation layer to obtain the spatial weight allocation value of the feature encoding information of each channel output by the spatial weight allocation layer; and inputting the channel weight allocation value and the spatial weight allocation value of the feature encoding information of each channel into the weight fusion layer to obtain the weight allocation value of the feature encoding information of each channel.
[0139] For example, the dual attention mechanism CBAM used in this embodiment is a method that combines channel and spatial attention mechanisms. Channel and spatial attention mechanisms are applied sequentially to the input feature layer, causing the model to focus more on the objects to be identified. Therefore, the CBAM dual attention mechanism is highly feasible for extracting small target objects. The CBAM dual attention mechanism can enhance low-level feature information useful for classification without overusing low-level features, effectively eliminating noise. Simultaneously, by strengthening the focus on small target objects, it can improve model learning efficiency and achieve better classification results.
[0140] The feature encoding information of each channel and the weight allocation value of the feature encoding information of each channel are input to the feature decoding layer to obtain the semantic features of the remote sensing image of the apple orchard to be identified output by the feature decoding layer.
[0141] Specifically, after obtaining the feature encoding information of each channel and the weight allocation value of the feature encoding information of each channel, the weight of the semantic features of each channel is determined based on the weight allocation value of the feature encoding information of each channel. The semantic features of each channel are obtained by upsampling, convolution and batch normalization of the feature encoding information of each channel. Based on the weight of the semantic features of each channel, the semantic features of each channel are weighted to obtain the semantic features of the remote sensing image of the apple orchard to be identified.
[0142] For example, Figure 9 This is a schematic diagram of the improved SegNet model structure provided in an embodiment of the present invention, as shown below. Figure 9As shown, the improved SegNet model provided in this application uses the SegNet network structure as its basic structure and mainly consists of three parts. The first part is the encoder, which uses 2D convolutions with a kernel size of 3*3 to extract image features. The number of filters gradually increases as the network deepens, and a total of 4 downsampling layers are used, with the Corrected Linear Unit (ReLU) as the activation function. The second part is the decoder, whose convolutional layers are divided into 4 parts. Upsampling is performed using the indices stored in the encoder's pooling. The last convolutional layer of the model is connected to a non-linear function, Sigmoid, as the activation function. The third part is the CBAM dual attention mechanism module, which is added between the encoder convolutional layers and the decoder upsampling. The output of the encoder is used as the input of the attention module, which uses low-level features to help high-level features recover pixel positions, thereby enhancing the attention to feature information and reducing the loss of important information.
[0143] Based on any of the above embodiments, the loss calculation layer in this embodiment of the invention can be implemented by setting a loss function. Correspondingly, the semantic features are input to the loss calculation layer to obtain the spatial information recognition result of the remote sensing image of the apple orchard to be identified, which is output by the loss calculation layer. Specifically, this may include:
[0144] Based on the semantic features, the loss value of the spatial information recognition model is determined;
[0145] Based on the loss value and the preset loss function, the spatial information recognition result is determined.
[0146] For example, see still Figure 9 The improved SegNet model provided in this application also includes a fourth part: loss function calculation part. When calculating the loss function during the forward and backward propagation of the model, the original cross-entropy loss function is improved into a Focal Loss loss function with weighting factors and modulation factors to alleviate the problem of imbalance between apple orchard and background land cover samples.
[0147] Specifically, the Focal loss function used in this embodiment of the invention is a function that calculates the difference between the predicted value and the true value. The sample is forward propagated through the model to obtain the predicted value. The difference between the predicted value and the true value is called the loss. The model updates the model parameters based on the loss value through backpropagation, thereby reducing the loss value. The smaller the loss value, the stronger the model's performance becomes. The Binary Cross Entropy Loss function is a commonly used loss function in binary classification problems, as shown in formula (1):
[0148]
[0149] Where y takes the values 1 (positive sample) and -1 (negative sample), and p is the probability that the model predicts that the sample belongs to 1 (positive sample).
[0150] Formula (1) can also be simplified by the following formula (2):
[0151]
[0152] From formula (2), we obtain the following formula:
[0153] CE(p, y) = CE(P) t ) = -log(P t (3).
[0154] It can be seen that the binary cross-entropy loss function performs poorly in handling class imbalance. When the number of samples from a certain class is large, it will cause that class to dominate the loss function, and random errors will more easily affect the smaller number of small object classes, thus degrading model performance. In small object segmentation problems, it is often necessary to optimize the loss function to achieve accurate segmentation of small objects.
[0155] To address the aforementioned issues, the Focal loss method employed in this invention is a novel loss function derived from the field of object detection, which is highly sensitive to the segmentation of small objects. Focal loss controls the balance of samples in two ways: by controlling the weights of positive and negative samples, and by controlling the weights of easily classified and difficult-to-classify samples.
[0156] Specifically, Focal loss can first assign different weights to each category by using a weighting factor α before the regular loss function. t The coefficient reduces the impact of a large number of samples. As shown in formulas (4) and (5), the range of α is 0 to 1. When the true label is a positive sample, α... t =ɑ, when the true label is a negative sample. t =1-a, and the contribution of positive and negative samples to the loss function can be controlled by setting the value of a. When a is between 0 and 0.5, the weight of the loss function for positive samples is reduced, while the weight of the loss function for negative samples is increased, and vice versa.
[0157]
[0158]
[0159] It is worth noting that although the weighting factor α tWhile this solves the problem of imbalanced positive and negative samples, when there are many easily distinguishable negative samples, the model training tends to focus only on learning from negative samples, resulting in large backpropagation gradients for negative samples. To address this issue, Focal loss uses a modulation factor (1-p) t ) γ The weights of easily classified and difficult-to-classify samples are controlled as shown in formula (6). Whether it is a positive or negative sample, when p... t When the value is close to 1, it indicates that the sample is easily distinguishable. At this point, the modulation factor (1-p) t ) γ The modulator γ will be smaller, meaning it contributes less to the loss value. This reduces the proportion of loss contribution from easily distinguishable samples, thereby increasing attention to samples that are difficult to distinguish. The range of the modulator γ is usually between 0 and 5. When γ = 0, the difficulty of distinguishing samples is not considered.
[0160]
[0161] By combining the weighting factor and the modulation factor, the final form of Focal Loss can be obtained, as shown in Equation (7).
[0162]
[0163] Based on the above embodiments, it can be seen that Focal Loss is very effective in balancing positive and negative samples, which can effectively improve the model's ability to extract small target features, thereby improving the model's segmentation effect. Therefore, this embodiment alleviates the problem of small targets in apple orchard extraction by replacing the model's cross-entropy loss function with Focal Loss.
[0164] The spatial information identification device for apple orchards provided by the present invention is described below. The spatial information identification device for apple orchards described below can be referred to in correspondence with the spatial information identification method for apple orchards described above.
[0165] Based on any of the above embodiments Figure 10 This is a schematic diagram of the structure of the apple orchard spatial information recognition device provided in an embodiment of the present invention, as shown below. Figure 10 As shown, the apple orchard spatial information recognition device includes a feature extraction module 1010, a feature fusion module 1020, and a spatial information recognition module 1030;
[0166] The feature extraction module 1010 is used to acquire single-temporal remote sensing images and multi-temporal remote sensing images of the apple orchard to be identified, and to extract the single-temporal feature representation of the single-temporal remote sensing image of the apple orchard to be identified and the multi-temporal feature representation of the multi-temporal remote sensing image.
[0167] The feature fusion module 1020 is used to fuse the single-temporal feature representation with the multi-temporal feature representation to obtain the fused feature representation of the remote sensing image of the apple orchard to be identified.
[0168] The spatial information recognition module 1030 is used to input the fusion feature representation of the remote sensing image of the apple orchard to be identified into the spatial information recognition model, and obtain the spatial information recognition result corresponding to the remote sensing image of the apple orchard to be identified output by the spatial information recognition model; the spatial information recognition model is trained based on the sample fusion feature representation of the sample apple orchard remote sensing image and the spatial information identifier of the sample apple orchard remote sensing image.
[0169] The device provided in this invention performs spatial information recognition based on the feature representation of multispectral remote sensing images of apple orchards. This solves the problems of limited recognition accuracy and difficulty in extracting spatiotemporal information in existing technologies that use low-to-medium spatial resolution remote sensing data and conventional supervised classification methods for spatial information recognition, resulting in poor extraction of small-target apple orchards. The device fully explores the deep features of multispectral remote sensing images of apple orchards, thereby improving the effectiveness of spatial information recognition for small-target apple orchards.
[0170] Figure 11 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 11 As shown, the electronic device may include: a processor 1110, a communications interface 1120, a memory 1130, and a communications bus 1140, wherein the processor 1110, the communications interface 1120, and the memory 1130 communicate with each other through the communications bus 1140. The processor 1110 can call logic instructions in the memory 1130 to execute a spatial information identification method for apple orchards. This method includes: acquiring single-temporal and multi-temporal remote sensing images of the apple orchard to be identified, and extracting single-temporal feature representations from the single-temporal remote sensing image and multi-temporal feature representations from the multi-temporal remote sensing image; fusing the single-temporal and multi-temporal feature representations to obtain a fused feature representation of the remote sensing image of the apple orchard to be identified; inputting the fused feature representation of the remote sensing image of the apple orchard to be identified into a spatial information identification model to obtain a spatial information identification result corresponding to the remote sensing image of the apple orchard to be identified, output by the spatial information identification model; wherein the spatial information identification model is trained based on sample fused feature representations of sample apple orchard remote sensing images and spatial information identifiers of the sample apple orchard remote sensing images.
[0171] Furthermore, the logical instructions in the aforementioned memory 1130 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0172] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the apple orchard spatial information identification method provided by the above methods. The method includes: acquiring single-temporal remote sensing images and multi-temporal remote sensing images of the apple orchard to be identified, and extracting single-temporal feature representations of the single-temporal remote sensing images and multi-temporal feature representations of the multi-temporal remote sensing images of the apple orchard to be identified; fusing the single-temporal feature representations and multi-temporal feature representations to obtain a fused feature representation of the remote sensing images of the apple orchard to be identified; inputting the fused feature representation of the remote sensing images of the apple orchard to be identified into a spatial information identification model to obtain a spatial information identification result corresponding to the remote sensing images of the apple orchard to be identified output by the spatial information identification model; wherein, the spatial information identification model is trained based on the sample fused feature representation of sample apple orchard remote sensing images and the spatial information identifier of the sample apple orchard remote sensing images.
[0173] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the method for identifying spatial information of an apple orchard provided by the above methods. The method includes: acquiring single-temporal remote sensing images and multi-temporal remote sensing images of an apple orchard to be identified, and extracting single-temporal feature representations of the single-temporal remote sensing images of the apple orchard to be identified and multi-temporal feature representations of the multi-temporal remote sensing images; fusing the single-temporal feature representations and multi-temporal feature representations to obtain a fused feature representation of the remote sensing images of the apple orchard to be identified; inputting the fused feature representation of the remote sensing images of the apple orchard to be identified into a spatial information identification model to obtain a spatial information identification result corresponding to the remote sensing images of the apple orchard to be identified, output by the spatial information identification model; wherein the spatial information identification model is trained based on sample fused feature representations of sample apple orchard remote sensing images and spatial information identifiers of the sample apple orchard remote sensing images.
[0174] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0175] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0176] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for spatial information identification of apple orchards, characterized in that, include: Acquire single-temporal and multi-temporal remote sensing images of the apple orchard to be identified, and extract the single-temporal feature representation of the single-temporal remote sensing image and the multi-temporal feature representation of the multi-temporal remote sensing image of the apple orchard to be identified; The single-temporal feature representation is fused with the multi-temporal feature representation to obtain the fused feature representation of the remote sensing image of the apple orchard to be identified; The fusion feature representation of the remote sensing image of the apple orchard to be identified is input into the spatial information recognition model to obtain the spatial information recognition result corresponding to the remote sensing image of the apple orchard to be identified, which is output by the spatial information recognition model. The spatial information recognition model is trained based on the sample fusion feature representation of the sample apple orchard remote sensing image and the spatial information identifier of the sample apple orchard remote sensing image. The spatial information recognition model includes a feature extraction layer and a loss calculation layer; Correspondingly, the step of inputting the fused feature representation of the remote sensing image of the apple orchard to be identified into the spatial information recognition model, and obtaining the spatial information recognition result corresponding to the remote sensing image of the apple orchard to be identified output by the spatial information recognition model, specifically includes: The fused feature representation of the remote sensing image of the apple orchard to be identified is input into the feature extraction layer to obtain the semantic features of the remote sensing image of the apple orchard to be identified output by the feature extraction layer. The semantic features are input into the loss calculation layer to obtain the spatial information recognition result of the multispectral remote sensing image of the apple orchard to be identified, output by the loss calculation layer. The feature extraction layer includes a feature encoding layer, a weight allocation layer, and a feature decoding layer; Correspondingly, the step of inputting the fused feature representation of the remote sensing image of the apple orchard to be identified into the feature extraction layer to obtain the semantic features of the remote sensing image of the apple orchard to be identified output by the feature extraction layer specifically includes: The fusion feature representation of the remote sensing image of the apple orchard to be identified is input into the feature coding layer to obtain the multi-channel feature coding information of the multispectral remote sensing image of the apple orchard to be identified output by the feature coding layer. The feature encoding information of each channel is input into the weight allocation layer to obtain the weight allocation value of the feature encoding information of each channel output by the weight allocation layer; The feature encoding information of each channel and the weight allocation value of the feature encoding information of each channel are input to the feature decoding layer to obtain the semantic features of the remote sensing image of the apple orchard to be identified output by the feature decoding layer.
2. The method for identifying spatial information of apple orchards according to claim 1, characterized in that, The weight allocation layer includes a channel weight allocation layer, a spatial weight allocation layer, and a weight fusion layer; Correspondingly, the step of inputting the feature encoding information of each channel into the weight allocation layer to obtain the weight allocation value of the feature encoding information of each channel output by the weight allocation layer includes: The feature encoding information of each channel is input into the channel weight allocation layer to obtain the channel weight allocation value of the feature encoding information of each channel output by the channel weight allocation layer; The feature encoding information of each channel is input into the spatial weight allocation layer to obtain the spatial weight allocation value of the feature encoding information of each channel output by the spatial weight allocation layer; The channel weight allocation value and the spatial weight allocation value of the feature encoding information of each channel are input into the weight fusion layer to obtain the weight allocation value of the feature encoding information of each channel.
3. The method for identifying spatial information of apple orchards according to claim 1, characterized in that, The step of inputting the feature encoding information of each channel and the weight allocation value of the feature encoding information of each channel into the feature decoding layer to obtain the semantic features of the multispectral remote sensing image of the apple orchard to be identified output by the feature decoding layer includes: Based on the weight allocation values of the feature encoding information of each channel, the weights of the semantic features of each channel are determined, wherein the semantic features of each channel are obtained by upsampling, convolution and batch normalization of the feature encoding information of each channel; Based on the weights of the semantic features in each channel, the semantic features in each channel are weighted to obtain the semantic features of the remote sensing image of the apple orchard to be identified.
4. The method for identifying spatial information of apple orchards according to claim 1, characterized in that, The step of inputting the semantic features into the loss calculation layer to obtain the spatial information recognition result of the remote sensing image of the apple orchard to be identified, output by the loss calculation layer, includes: Based on the semantic features, the loss value of the spatial information recognition model is determined; Based on the loss value and the preset loss function, the spatial information recognition result is determined.
5. The method for identifying spatial information of apple orchards according to any one of claims 1-4, characterized in that, Extracting multi-temporal feature representations from the multi-temporal remote sensing images of the apple orchard to be identified includes: Extract the temporal and spectral index features of the multi-temporal remote sensing images of the apple orchard to be identified; By fusing the temporal features and spectral index features, a multi-temporal feature representation of the apple orchard to be identified is obtained.
6. A spatial information recognition device for apple orchards, characterized in that, include: The feature extraction module is used to acquire single-temporal and multi-temporal remote sensing images of the apple orchard to be identified, and to extract the single-temporal feature representation of the single-temporal remote sensing image of the apple orchard to be identified and the multi-temporal feature representation of the multi-temporal remote sensing image. The feature fusion module is used to fuse the single-temporal feature representation with the multi-temporal feature representation to obtain the fused feature representation of the remote sensing image of the apple orchard to be identified; The spatial information recognition module is used to input the fusion feature representation of the remote sensing image of the apple orchard to be identified into the spatial information recognition model, and obtain the spatial information recognition result corresponding to the remote sensing image of the apple orchard to be identified output by the spatial information recognition model. The spatial information recognition model is trained based on the sample fusion feature representation of the sample apple orchard remote sensing image and the spatial information identifier of the sample apple orchard remote sensing image. The spatial information recognition model includes a feature extraction layer and a loss calculation layer; Correspondingly, the spatial information recognition module is specifically used for: The fused feature representation of the remote sensing image of the apple orchard to be identified is input into the feature extraction layer to obtain the semantic features of the remote sensing image of the apple orchard to be identified output by the feature extraction layer. The semantic features are input into the loss calculation layer to obtain the spatial information recognition result of the multispectral remote sensing image of the apple orchard to be identified, output by the loss calculation layer. The feature extraction layer includes a feature encoding layer, a weight allocation layer, and a feature decoding layer; Correspondingly, the spatial information recognition module is specifically used for: The fusion feature representation of the remote sensing image of the apple orchard to be identified is input into the feature coding layer to obtain the multi-channel feature coding information of the multispectral remote sensing image of the apple orchard to be identified output by the feature coding layer. The feature encoding information of each channel is input into the weight allocation layer to obtain the weight allocation value of the feature encoding information of each channel output by the weight allocation layer; The feature encoding information of each channel and the weight allocation value of the feature encoding information of each channel are input to the feature decoding layer to obtain the semantic features of the remote sensing image of the apple orchard to be identified output by the feature decoding layer.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the apple orchard spatial information recognition method as described in any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the apple orchard spatial information identification method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Remote sensing vegetation extraction method and device based on deep learning semantic segmentation
CN116385875A
Few-shot urban remote sensing image information extraction method based on meta learning and attention
US20230215166A1