A method for single-tree segmentation based on infrared thermal images and visible light images
Through the feature fusion of infrared thermal images and visible light images and the main curve rate and watershed algorithm, the problems of high cost and low segmentation accuracy of point cloud data acquisition equipment are solved, and efficient and low-cost single-wood segmentation is achieved.
Patent Information
- Application Number
- CN202411591624.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-08
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-11-08
AI Technical Summary
The existing technology relies on point cloud data acquisition equipment with high cost and low segmentation accuracy, and the quality of traditional point cloud data segmentation is poor, resulting in low efficiency of single wood segmentation.
A feature fusion method based on infrared thermal images and visible light images is adopted, and a single wood segmentation model pre-trained with the main curve rate and watershed algorithm is used to replace the point cloud data acquisition equipment to achieve rich information and accurate segmentation.
The cost of single wood segmentation is reduced, segmentation accuracy and efficiency are improved, and accurate segmentation of feature fusion images is achieved.
Smart Images

Figure CN119169019B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of individual tree segmentation, and particularly to an individual tree segmentation method based on infrared thermal images and visible light images. Background Art
[0002] Accurate individual tree segmentation can help scientists and forestry managers obtain detailed information about each tree, such as health status, growth rate, and environmental adaptability. These data are crucial for pest and disease monitoring, resource management, and forest planning. Through efficient and accurate individual tree segmentation technology, sustainable management of forest resources can be achieved, forestry productivity can be improved, and the health and stability of the ecosystem can be ensured. In the past, individual tree segmentation tasks were basically manual sampling and marking after delimiting quadrats. This method is time-consuming and laborious, with low efficiency and cannot fully cover the responsible area. In recent years, due to the booming development of the drone industry, many methods for using drone-borne radar to achieve individual tree segmentation tasks have been proposed in the field. Currently, the algorithms for individual tree extraction can generally be divided into two ideas: detection methods based on canopy height models and detection methods based on point clouds. The mainstream algorithms are region growing method and watershed method. In the existing technologies, the acquisition of data is generally divided into two categories. One is to directly extract point cloud data in the target area using lidar, and the other is to calculate point cloud data using high-resolution visible light images.
[0003] The existing technologies are extremely dependent on point cloud data, which will lead to: the acquisition of point cloud data generally requires large-scale drone equipment and supporting lidar, with high costs; using oblique photography technology to obtain visible light images and calculating point cloud data by subsequent point cloud reconstruction methods has quality problems. Summary of the Invention
[0004] In order to overcome the deficiencies of the prior art, the purpose of the present invention is to provide an individual tree segmentation method based on infrared thermal images and visible light images. By fusing the features of visible light images and infrared thermal images, the richness of input information is realized, replacing the point cloud data acquisition equipment and reducing costs; through an individual tree segmentation model pre-trained based on the principal curvature rate and watershed algorithm, accurate segmentation of the feature fusion image is achieved, avoiding the problems of poor data quality and low segmentation accuracy in a large range of traditional point cloud data segmentation.
[0005] To achieve the above purpose, the present invention provides the following solution:
[0006] An individual tree segmentation method based on infrared thermal images and visible light images, comprising:
[0007] Using a drone equipped with a visible light camera and a multispectral camera to collect images of the target recognition area, obtaining an original visible light image set and an original thermal imaging image set;
[0008] Stitch the original visible light image set and the original thermal imaging image set respectively to obtain an overall visible light image and an overall thermal imaging image;
[0009] Extract features from and fuse the overall visible light image and the overall thermal imaging image to obtain a feature fusion image;
[0010] Input the feature fusion image into a single-tree segmentation model pre-trained based on the principal curvature and watershed algorithm to obtain a target segmentation image.
[0011] Preferably, extracting features from and fusing the overall visible light image and the overall thermal imaging image to obtain a feature fusion image includes:
[0012] Use a preset visible light feature extraction network to extract the overall visible light image to obtain a visible light feature matrix;
[0013] Use a preset thermal imaging feature extraction network to extract the overall thermal imaging image to obtain a thermal imaging feature matrix;
[0014] Stitch the visible light feature matrix and the thermal imaging feature matrix to obtain a multi-modal stitched feature;
[0015] Perform first global average pooling and first max pooling on the multi-modal stitched feature in the channel domain respectively to obtain a global channel description and a maximum channel description;
[0016] Input the global channel description and the maximum channel description into a basic neural network for calculation to obtain a channel weight coefficient; the basic neural network includes: a Relu activation function, a Sigmoid activation function, and at least two neural network layers;
[0017] Multiply the channel weight coefficient by the multi-modal stitched feature to obtain a channel-corrected multi-modal feature;
[0018] Perform second global average pooling and second max pooling on the channel-corrected multi-modal feature in the spatial domain respectively to obtain a global spatial description and a maximum spatial description;
[0019] Input the global spatial description and the maximum spatial description into a basic neural network for calculation to obtain a spatial weight coefficient;
[0020] Multiply the channel-corrected multi-modal feature by the spatial weight coefficient to obtain a spatially corrected multi-modal feature;
[0021] Perform third global average pooling and Sigmoid calculation on the spatially corrected multi-modal feature to obtain a comprehensive image feature;
[0022] Modify the overall visible light image according to the comprehensive image features to obtain the feature fusion image.
[0023] Preferably, the construction process of the single-tree segmentation model includes:
[0024] Convert the pre-collected forest farm image set into grayscale format to obtain a grayscale image set;
[0025] Extract the grayscale value of each pixel point in the grayscale image set, and define the grayscale value as the height to obtain a three-dimensional point distribution set;
[0026] Use the radial basis function interpolation method to fit the three-dimensional point distribution set into a surface to obtain a grayscale surface set;
[0027] Calculate the principal curvature of the grayscale surface set, and obtain a grayscale surface edge image set according to the principal curvature;
[0028] Use the watershed algorithm to perform an initial segmentation on the grayscale image set to obtain an initial segmentation image set;
[0029] Use the forest farm image set, the initial segmentation image set, and the grayscale surface edge image set as the input of a preset convolutional neural network, and use the manually segmented forest farm single-tree segmentation image set as the output of the convolutional neural network to train the convolutional neural network to obtain the trained single-tree segmentation model.
[0030] Preferably, the tool for stitching the original visible light image set and the original thermal imaging image set is: the Stitcher library file.
[0031] Preferably, the acquisition parameter settings of the drone include: the height range is 60 to 80 meters, the heading overlap is 60%, and the side overlap is 60%.
[0032] Preferably, calculating the principal curvature of the grayscale surface set and obtaining a grayscale surface edge image set according to the principal curvature includes:
[0033] Randomly select a grayscale surface image according to the grayscale surface set;
[0034] Construct a Hessian matrix according to the grayscale surface image;
[0035] Calculate the maximum principal curvature and the minimum principal curvature of each pixel point in the grayscale surface image according to the Hessian matrix;
[0036] Screen the maximum principal curvature and the minimum principal curvature according to a preset edge point determination formula to obtain an original edge point set; the edge point determination formula is: EP is the set of original edge points; k1 is the maximum principal curvature; k2 is the minimum principal curvature; σ1 is the adaptive first threshold; σ2 is the adaptive second threshold;
[0037] Noise reduction is performed on the set of original edge points using the noise reduction formula for speech jamming to obtain the set of main edge points; the noise reduction formula is: EP′ is the set of main edge points; p i is the number of the original edge points within the 8 nearest pixels to the i-th original edge point in the set of original edge points; l i is the length of the curve formed by the edge points where the i-th original edge point in the set of original edge points is located; S i is the area of the region formed by the edge points where the i-th original edge point in the set of original edge points is located; η is the adaptive length threshold; δ is the adaptive area threshold;
[0038] The set of main edge points is matched to the grayscale surface image to obtain the surface image of the fused edge points;
[0039] All the images in the grayscale surface set are traversed to obtain several surface images of the fused edge points, and all the surface images of the fused edge points are integrated to obtain the grayscale surface edge image set.
[0040] Preferably, the watershed algorithm is used to perform an initial segmentation on the grayscale image set to obtain an initial segmentation image set, including:
[0041] Noise reduction is performed on any grayscale image in the grayscale image set to obtain a denoised image;
[0042] According to the preset number of single-image contours, the findContours function is used to find the contours of the denoised image to obtain several contour distributions;
[0043] The contour distributions are used to segment the images in the forest farm image set corresponding to the grayscale image to obtain several contour segmentation images;
[0044] All the images in the grayscale image set are traversed and integrated to obtain the initial segmentation image set.
[0045] Preferably, the forest farm image set, the initial segmentation image set, and the grayscale surface edge image set are used as the input of a preset convolutional neural network, and the manually segmented forest farm single-tree segmentation image set is used as the output of the convolutional neural network. The convolutional neural network is trained to obtain the trained single-tree segmentation model, including:
[0046] Input the forest farm image set and the grayscale surface edge image set into a 1×1 convolution respectively for local detail feature extraction to obtain a first feature representation and a second feature representation;
[0047] Input the first feature representation and the second feature representation into a BN layer for feature normalization to obtain a first normalized feature representation and a second normalized feature representation;
[0048] Fuse the first normalized feature representation and the second normalized feature representation to obtain a fused feature representation;
[0049] Input the fused feature representation into a preset deep learning network for prediction to obtain a predicted segmentation image;
[0050] According to the forest farm single-tree segmentation image set, calculate the loss value of the predicted segmentation image by using a preset loss function; the calculation formula of the loss function is: Loss b is the loss value of the b-th iteration; is the predicted area of the a-th single-tree segmentation region in the b-th iteration; S ab is the overlapping area between the predicted area and the actual area of the a-th single-tree segmentation region in the b-th iteration; m is the total number of actual single-tree segmentation regions;
[0051] Based on the minimum gradient descent method, optimize the parameters of the deep learning network, the edge point determination formula, and the noise reduction formula according to the loss value;
[0052] Iteratively train the deep learning network, the edge point determination formula, and the noise reduction formula to obtain the trained single-tree segmentation model.
[0053] The present invention discloses the following technical effects:
[0054] The present invention provides a single-tree segmentation method based on infrared thermal images and visible light images. By fusing the features of visible light images and infrared thermal images, the problem of high cost of traditional point cloud data acquisition equipment is solved, and the enrichment of input information and the replacement of point cloud data acquisition equipment are realized; through a single-tree segmentation model pre-trained based on the principal curvature rate and the watershed algorithm, the problems of poor data quality and low large-scale segmentation accuracy of traditional point cloud data segmentation are solved, and the accurate segmentation of the feature fusion image is realized. Description of the Drawings
[0055] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for use in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0056] Figure 1 Schematic diagram of the single-tree segmentation process based on infrared thermal images and visible light images provided by the embodiments of the present invention;
[0057] Figure 2 Schematic diagram of the feature fusion image acquisition process provided by the embodiments of the present invention;
[0058] Figure 3 Schematic diagram of the single-tree segmentation model construction process provided by the embodiments of the present invention. Detailed implementation manners
[0059] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0060] The purpose of the present invention is to provide a single-tree segmentation method based on infrared thermal images and visible light images. By fusing the features of visible light images and infrared thermal images, the input information is enriched, replacing the point cloud data acquisition device and reducing costs; through a single-tree segmentation model pre-trained based on the principal curvature and watershed algorithm, accurate segmentation of the feature fusion image is achieved, avoiding the problems of poor data quality and low large-scale segmentation accuracy in traditional point cloud data segmentation.
[0061] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will further describe the present invention in detail in conjunction with the drawings and specific implementation manners.
[0062] Figure 1 Schematic diagram of the single-tree segmentation process based on infrared thermal images and visible light images provided by the embodiments of the present invention, as Figure 1 shown, the present invention provides a single-tree segmentation method based on infrared thermal images and visible light images, including:
[0063] Step 100: Use a drone equipped with a visible light camera and a multispectral camera to collect images of the target recognition area, obtaining an original visible light image set and an original thermal imaging image set;
[0064] Step 200: Stitch the original visible light image set and the original thermal imaging image set respectively to obtain an overall visible light image and an overall thermal imaging image;
[0065] Step 300: Extract and fuse features from the overall visible light image and the overall thermal imaging image to obtain a feature fusion image;
[0066] Step 400: Input the feature fusion image into a single tree segmentation model pre-trained based on the principal curvature rate and the watershed algorithm to obtain a target segmentation image.
[0067] Specifically, as Figure 2 shown, the step 300 specifically includes:
[0068] Step 301: Use a preset visible light feature extraction network to extract the overall visible light image to obtain a visible light feature matrix;
[0069] Step 302: Use a preset thermal imaging feature extraction network to extract the overall thermal imaging image to obtain a thermal imaging feature matrix;
[0070] Step 303: Stitch the visible light feature matrix and the thermal imaging feature matrix to obtain a multi-modal stitched feature;
[0071] Step 304: Perform the first global average pooling and the first max pooling on the multi-modal stitched feature in the channel domain respectively to obtain a global channel description and a max channel description;
[0072] Step 305: Input the global channel description and the max channel description into a basic neural network for calculation to obtain a channel weight coefficient; the basic neural network includes: a Relu activation function, a Sigmoid activation function, and at least two neural network layers;
[0073] Step 306: Multiply the channel weight coefficient by the multi-modal stitched feature to obtain a channel-corrected multi-modal feature;
[0074] Step 307: Perform the second global average pooling and the second max pooling on the channel-corrected multi-modal feature in the spatial domain respectively to obtain a global spatial description and a max spatial description;
[0075] Step 308: Input the global spatial description and the max spatial description into a basic neural network for calculation to obtain a spatial weight coefficient;
[0076] Step 309: Multiply the channel-corrected multi-modal feature by the spatial weight coefficient to obtain a spatially corrected multi-modal feature;
[0077] Step 310: After performing third global average pooling and Sigmoid calculation on the spatially corrected multi-modal features, comprehensive image features are obtained;
[0078] Step 311: The overall visible light image is corrected according to the comprehensive image features to obtain the feature fusion image.
[0079] Further, as Figure 3 shown, the construction process of the single-tree segmentation model in step 400 includes:
[0080] Step 401: Convert the pre-collected forest farm image set into grayscale format to obtain a grayscale image set;
[0081] Step 402: Extract the grayscale value of each pixel point in the grayscale image set, and define the grayscale value as the height to obtain a three-dimensional point distribution set;
[0082] Step 403: Use the radial basis function interpolation method to fit the three-dimensional point distribution set into a surface to obtain a grayscale surface set;
[0083] Step 404: Calculate the principal curvature of the grayscale surface set, and obtain a grayscale surface edge image set according to the principal curvature;
[0084] Step 405: Use the watershed algorithm to perform an initial segmentation on the grayscale image set to obtain an initial segmentation image set;
[0085] Step 406: Use the forest farm image set, the initial segmentation image set, and the grayscale surface edge image set as the input of a preset convolutional neural network, and use the manually segmented forest farm single-tree segmentation image set as the output of the convolutional neural network to train the convolutional neural network to obtain the trained single-tree segmentation model.
[0086] Optionally, the tool for stitching and using the original visible light image set and the original thermal imaging image set is: the Stitcher library file.
[0087] Preferably, the acquisition parameter settings of the drone include: the height range is 60 to 80 meters, the heading overlap is 60%, and the side overlap is 60%.
[0088] Specifically, calculating the principal curvature of the grayscale surface set and obtaining a grayscale surface edge image set according to the principal curvature includes:
[0089] Randomly select a grayscale surface image according to the grayscale surface set;
[0090] Construct a Hessian matrix according to the grayscale surface image;
[0091] Calculate the maximum principal curvature and the minimum principal curvature of each pixel point in the grayscale surface image according to the Hessian matrix;
[0092] Screen the maximum principal curvature and the minimum principal curvature according to a preset edge point determination formula to obtain an original edge point set; the edge point determination formula is: EP is the original edge point set; k1 is the maximum principal curvature; k2 is the minimum principal curvature; σ1 is an adaptive first threshold; σ2 is an adaptive second threshold;
[0093] Denoise the original edge point set by using a denoising formula for speech jamming to obtain a main edge point set; the denoising formula is: EP′ is the main edge point set; p i is the number of the original edge points within the 8 nearest pixels to the i-th original edge point in the original edge point set; l i is the length of the curve formed by the edge points where the i-th original edge point in the original edge point set is located; S i is the area of the region formed by the edge points where the i-th original edge point in the original edge point set is located; η is an adaptive length threshold; δ is an adaptive area threshold;
[0094] Match the main edge point set to the grayscale surface image to obtain a surface image with fused edge points;
[0095] Traverse all the images in the grayscale surface set to obtain several surface images with fused edge points, and integrate all the surface images with fused edge points to obtain the grayscale surface edge image set.
[0096] Preferably, use the watershed algorithm to perform an initial segmentation on the grayscale image set to obtain an initial segmentation image set, including:
[0097] Denoise any grayscale image in the grayscale image set to obtain a denoised image;
[0098] According to a preset number of single-image contours, use the findContours function to find the contours of the denoised image to obtain several contour distributions;
[0099] Use the contour distribution to segment the images in the forest farm image set corresponding to the grayscale image to obtain several contour segmentation images;
[0100] Traverse and integrate all the images in the grayscale image set to obtain the initial segmentation image set.
[0101] Further, taking the forest farm image set, the initial segmentation image set, and the gray surface edge image set as the inputs of a preset convolutional neural network, and taking the manually segmented forest farm single-tree segmentation image set as the output of the convolutional neural network, training the convolutional neural network to obtain the trained single-tree segmentation model, including:
[0102] Inputting the forest farm image set and the gray surface edge image set into a 1×1 convolution respectively for local detail feature extraction to obtain a first feature representation and a second feature representation;
[0103] Inputting the first feature representation and the second feature representation into a BN layer for feature normalization to obtain a first normalized feature representation and a second normalized feature representation;
[0104] Fusing the first normalized feature representation and the second normalized feature representation to obtain a fused feature representation;
[0105] Inputting the fused feature representation into a preset deep learning network for prediction to obtain a predicted segmentation image;
[0106] According to the forest farm single-tree segmentation image set, calculating the loss value of the predicted segmentation image by using a preset loss function; the calculation formula of the loss function is: Loss b is the loss value of the b-th iteration; is the predicted area of the a-th single-tree segmentation region in the b-th iteration; S ab is the overlapping area between the predicted area and the actual area of the a-th single-tree segmentation region in the b-th iteration; m is the total number of actual single-tree segmentation regions;
[0107] Based on the minimum gradient descent method, optimizing the parameters of the deep learning network, the edge point determination formula, and the noise reduction formula according to the loss value;
[0108] Performing iterative training on the deep learning network, the edge point determination formula, and the noise reduction formula to obtain the trained single-tree segmentation model.
[0109] Specifically, using a drone equipped with both a visible light camera and a multispectral camera for data collection to obtain a front view of infrared thermal imaging and visible light; using a drone with a fixed flight path to collect spectral data under permitted conditions to obtain multiple photos with overlapping parts. The drone settings are: flight altitude from 60 meters to 80 meters, forward overlap 60%, and side overlap 60%.
[0110] Further, preprocess the image data: Combine the flight path and the corresponding photos, and use the Stitcher library file to splice to obtain the corresponding front view; Combine the drone GPS data with the scope of the project-contracted forest land, and cut out the orthophoto image of the forest land to be used when expanding 30m outward from the actual contracted area, and use the noise reduction algorithm to improve the quality of the visible light image.
[0111] Optionally, this embodiment provides another neural network design idea, including: Design a neural network with an encoder-decoder structure and 4M parameters, use the different regularization functions and texture differences between the output image and the input image as guidance, and use the Adam optimizer to optimize the network parameters. Use the multi-scale features extracted in the encoder for image feature-level fusion. In the model feature extraction, we use the multi-scale features with different dimensions and sizes generated at different stages of the encoder to participate in the fusion task. The fusion module fuses the global features and texture features of the visible light and thermal imaging at the same scale respectively, and then integrates them. Finally, the integrated multi-scale features enter the decoder for image reconstruction. Use real data in the actual tree crown project, and fine-tune the model with texture information as the index to obtain an image fusion neural network suitable for the project scenario.
[0112] Optionally, the segmentation model output of this embodiment is a json file for each tree crown range. After visualizing the segmentation results according to the json file, data such as the number of green plants, area, and percentage of the area in the region can be statistically obtained based on the visualized segmentation image.
[0113] The beneficial effects of the present invention are as follows:
[0114] By fusing the features of the visible light image and the infrared thermal image, the present invention enriches the input information, successfully replaces the point cloud data acquisition device, and reduces the cost; Through the single-tree segmentation model pre-trained based on the principal curvature and watershed algorithm, the accurate segmentation of the feature fusion image is realized, avoiding the problems of poor data quality and low large-scale segmentation accuracy in the traditional point cloud data segmentation.
[0115] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same and similar parts among the various embodiments can be referred to each other.
[0116] Specific examples are used in this article to elaborate on the principles and implementation methods of the present invention. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present invention; At the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A single-tree segmentation method based on infrared thermal images and visible light images, characterized in that, Including: Converting the pre - collected forest farm image set into grayscale format to obtain a grayscale image set; Extracting the grayscale value of each pixel point in the grayscale image set and defining the grayscale value as height to obtain a three - dimensional point distribution set; Using the radial basis function interpolation method to fit the three - dimensional point distribution set into a surface to obtain a grayscale surface set; Randomly selecting a grayscale surface image according to the grayscale surface set; Constructing a Hessian matrix according to the grayscale surface image; Calculating the maximum principal curvature and the minimum principal curvature of each pixel point in the grayscale surface image according to the Hessian matrix; Screening the maximum principal curvature and the minimum principal curvature according to a preset edge point determination formula to obtain an original edge point set; The edge point determination formula is as follows: EP is the set of original edge points; k1 is the maximum principal curvature; k2 is the minimum principal curvature; σ1 is an adaptive first threshold; σ2 is an adaptive second threshold; Noise reduction is performed on the original edge point set by using a noise reduction formula for speech hesitation to obtain a main edge point set; the noise reduction formula is: EP′ is the main edge point set; p i is the number of the original edge points within the nearest 8 pixels of the i-th original edge point in the original edge point set; l i is the length of the curve formed by the edge points where the i-th original edge point in the original edge point set is located; S i is the area of the region formed by the edge points where the i-th original edge point in the original edge point set is located; η is an adaptive length threshold; δ is an adaptive area threshold; Matching the main edge point set to the grayscale surface image to obtain a surface image with fused edge points; Traversing all the images in the grayscale surface set to obtain several surface images with fused edge points, and integrating all the surface images with fused edge points to obtain the grayscale surface edge image set; Using the watershed algorithm to perform an initial segmentation on the grayscale image set to obtain an initial segmentation image set; Taking the forest farm image set, the initial segmentation image set, and the grayscale surface edge image set as the input of a preset convolutional neural network, and taking the manually segmented forest farm single - tree segmentation image set as the output of the convolutional neural network, and training the convolutional neural network to obtain a trained single - tree segmentation model; Using a drone equipped with a visible - light camera and a multi - spectral camera to collect images of the target recognition area to obtain an original visible - light image set and an original thermal - imaging image set; Performing stitching on the original visible - light image set and the original thermal - imaging image set respectively to obtain an overall visible - light image and an overall thermal - imaging image; Performing feature extraction and fusion on the overall visible - light image and the overall thermal - imaging image to obtain a feature - fused image; Inputting the feature - fused image into the single - tree segmentation model pre - trained based on the principal curvature and the watershed algorithm to obtain a target segmentation image.
2. The single-tree segmentation method based on infrared thermal images and visible light images according to claim 1, wherein Performing feature extraction and fusion on the overall visible - light image and the overall thermal - imaging image to obtain a feature - fused image, including: Using a preset visible - light feature extraction network to extract the overall visible - light image to obtain a visible - light feature matrix; Using a preset thermal - imaging feature extraction network to extract the overall thermal - imaging image to obtain a thermal - imaging feature matrix; Stitching the visible - light feature matrix and the thermal - imaging feature matrix to obtain a multi - modal stitched feature; Performing the first global average pooling and the first max - pooling in the channel domain on the multi - modal stitched feature respectively to obtain a global channel description and a max - channel description; Inputting the global channel description and the max - channel description into a basic neural network for calculation to obtain a channel weight coefficient; the basic neural network includes: a Relu activation function, a Sigmoid activation function, and at least two neural network layers; Multiplying the channel weight coefficient by the multi - modal stitched feature to obtain a channel - corrected multi - modal feature; Perform the second global average pooling and the second max pooling in the spatial domain on the multimodal features corrected for the channels respectively to obtain a global spatial description and a max spatial description; Input the global spatial description and the max spatial description into a basic neural network for calculation to obtain spatial weight coefficients; Multiply the multimodal features corrected for the channels by the spatial weight coefficients to obtain multimodal features corrected in space; After performing the third global average pooling and Sigmoid calculation on the multimodal features corrected in space, obtain comprehensive image features; Correct the overall visible light image according to the comprehensive image features to obtain the feature fusion image.
3. The single-tree segmentation method based on infrared thermal images and visible light images according to claim 1, wherein The tool used for stitching the original visible light image set and the original thermal imaging image set is: the Stitcher library file.
4. The single-tree segmentation method based on infrared thermal images and visible light images according to claim 1, wherein The acquisition parameter settings of the UAV include: the altitude range is 60 to 80 meters, the forward overlap is 60%, and the side overlap is 60%.
5. A single-tree segmentation method based on infrared thermal images and visible light images according to claim 1, characterized in that, Use the watershed algorithm to perform an initial segmentation on the grayscale image set to obtain an initial segmentation image set, including: Denoise any grayscale image in the grayscale image set to obtain a denoised image; According to the preset number of single-image contours, use the findContours function to find the contours of the denoised image to obtain a number of contour distributions; Use the contour distributions to segment the images in the forest farm image set corresponding to the grayscale image to obtain a number of contour segmentation images; Traverse and integrate all the images in the grayscale image set to obtain the initial segmentation image set.
6. The single-tree segmentation method based on infrared thermal images and visible light images according to claim 1, wherein Use the forest farm image set, the initial segmentation image set, and the grayscale surface edge image set as the inputs of a preset convolutional neural network, and use the manually segmented forest farm single-tree segmentation image set as the output of the convolutional neural network to train the convolutional neural network to obtain the trained single-tree segmentation model, including: Input the forest farm image set and the grayscale surface edge image set into a 1×1 convolution respectively for local detail feature extraction to obtain a first feature representation and a second feature representation; Input the first feature representation and the second feature representation into a BN layer for feature normalization to obtain a first normalized feature representation and a second normalized feature representation; Fuse the first normalized feature representation and the second normalized feature representation to obtain a fused feature representation; Input the fused feature representation into a preset deep learning network for prediction to obtain a predicted segmentation image; According to the single-tree segmentation image set of the forest farm, calculate the loss value of the predicted segmentation image by using a preset loss function; the calculation formula of the loss function is: Loss b is the loss value of the b-th iteration; is the predicted area of the a-th single-tree segmentation region of the b-th iteration; S ab is the overlapping area between the predicted area and the actual area of the a-th individual tree segmentation region in the b-th iteration; m is the total number of actual individual tree segmentation regions; Based on the minimum gradient descent method, optimize the parameters of the deep learning network, the edge point determination formula, and the denoising formula according to the loss value; Perform iterative training on the deep learning network, the edge point determination formula, and the denoising formula to obtain the trained single-tree segmentation model.
Citation Information
Patent Citations
High-canopy-density wetland forest parameter extraction method based on consumer-level unmanned aerial vehicle images
CN115760885A
Individual tree biomass estimation method based on hyperspectrum and air-ground collaborative LiDAR
CN118570677A