Automatic Extraction Method of Tree Factors Based on AIM-ASPP

By applying deep learning semantic segmentation technology and AIM-ASPP network in tree factor measurement, the problems of complex operation, high cost and low accuracy of tree factor measurement in the prior art are solved, and efficient and automated tree height and breast diameter measurement are achieved.

CN115375710BActive Publication Date: 2025-05-30NORTHEAST FORESTRY UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211130286.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-16
Publication Date
2025-05-30
Estimated Expiration
2042-09-16

AI Technical Summary

Technical Problem

The prior art has problems in the tree factor measurement, such as inconvenience in carrying, complex operation, high cost, and measurement accuracy is susceptible to camera performance and baseline length.

Method used

Using semantic segmentation methods based on deep learning, the integrated hollow space convolution pooling pyramid (AIM-ASPP) network is used to automate, batch and high-precision segmentation of trees and standards, thereby achieving efficient and automated measurement of tree height and breast diameter.

Benefits of technology

It realizes high-precision tree factor extraction without the need for a large number of preprocessing images and without the need to obtain the camera shooting angle and height in advance. It is low cost, high efficiency, and supports batch processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115375710B_ABST
    Figure CN115375710B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for automatically extracting tree factors based on AIM-ASPP. The extraction method includes: obtaining the internal parameters of a measurement camera through calibration; performing distortion correction on the captured street tree images according to the internal parameters; segmenting the corrected images into trees, reference objects, and backgrounds through a feature extraction network model containing an integrated Atrous Spatial Pyramid Pooling (AIM-ASPP); and obtaining tree factors by calculating parameters including the pixel parameters of the tree area and the reference object area in the segmented image and the measured data of the reference object. The extraction method of the present invention has low cost and high efficiency, and can achieve batch extraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of tree factor extraction methods, and particularly to the technical field of tree factor extraction methods based on AIM-ASPP. Background Art

[0002] Sustainable street tree management requires the measurement of a large number of tree factors. Among common tree factors, tree height and diameter at breast height are often used to evaluate the age of trees and are basic parameters for volume and biomass, etc., which are particularly important. Currently, traditional measurement tools for tree factors include total stations, electronic theodolites, Blume-type hypsometers, etc. However, these instruments have problems such as inconvenient carrying, complex operation, and high cost. Therefore, some technologies use lidar scanning methods. By processing the acquired point cloud data, rapid and high-precision extraction of tree height and diameter at breast height is achieved. However, lidar equipment is expensive, and the point cloud data volume is huge, requiring a large amount of preprocessing work.

[0003] Other existing tree factor extraction methods also include, for example: directly operating on the captured two-dimensional images using multi-view vision in close-range photography to achieve measurement by obtaining three-dimensional coordinates; or according to binocular stereo vision technology, first using a calibration matrix to achieve sparse projection reconstruction, and then using a sphere rotation algorithm for surface modeling to achieve dense reconstruction, and finally restoring the three-dimensional information of the tree canopy; or using algorithms of SGBM (semi-global stereo matching algorithm) and BM (string matching algorithm) for binocular stereo matching to obtain the tree disparity value to calculate the coordinates of key points to achieve the purpose of measuring trees. However, these binocular measurement methods have problems such as difficult stereo matching and the measurement accuracy being easily affected by factors such as camera performance and baseline length. And some monocular measurement methods, such as: using the built-in high-precision gyroscope of a smartphone to obtain the depression angle and the position of the shooting instrument to obtain the height of the tree; obtaining the height of the tree by obtaining the tilt angle of the shooting instrument and the distance from the actual target; obtaining the height of the tree by setting a fixed height of the shooting instrument and combining the internal sensors of the mobile phone, etc., all require manual interaction to determine the tree area and obtain parameters such as the tilt angle, height, or depth of field of the shooting instrument, with low measurement efficiency and relatively poor extraction accuracy. Summary of the Invention

[0004] Aiming at the defects of the prior art, the purpose of the present invention is to propose a method for efficiently and automatically measuring and extracting the height and diameter at breast height (DBH) of street trees using monocular vision. Based on the semantic segmentation method in deep learning, the Aggregate Interaction Module-Atrous Spatial Pyramid Pooling (AIM-ASPP) can be used to automatically, batch-wise, and accurately segment trees and reference objects, and then accurately extract tree factors according to the segmented images.

[0005] The technical solution of the present invention is as follows:

[0006] An automatic tree factor extraction method based on AIM-ASPP, which includes:

[0007] S1 Calibrate the measurement camera that obtains the street tree image to obtain its internal parameters;

[0008] S2 Perform distortion correction on the street tree image captured by the measurement camera according to the internal parameters obtained by calibration to obtain a corrected image;

[0009] S3 Segment the corrected image into trees, reference objects, and the background through a feature extraction network model with image segmentation capabilities to obtain a segmented image divided into tree regions, standard regions, and background regions, where the reference object is a fixed object in the same plane as the tree in the captured image;

[0010] S4 Calculate the tree factors by calculating parameters, where the calculation parameters include: pixel parameters of the tree region and the reference object region in the segmented image and the measured data of the reference object; the pixel parameters include the number of pixels and / or pixel coordinates;

[0011] Among them, the feature extraction network model contains an integrated atrous spatial pyramid pooling network, that is, the AIM-ASPP network.

[0012] According to some preferred embodiments of the present invention, the reference object is a lime white coating on the tree, and the tree factors include the tree height and the DBH of the tree.

[0013] According to some preferred embodiments of the present invention, the Zhang-Zhengyou calibration method is used for calibration, and the internal parameters f x 、f y 、u 0 、v 0 in the transformation formula from the following world coordinate system to the pixel coordinate system, the external parameters R, T, and the tangential distortion coefficients p 1 、p 2 and the radial distortion parameter k1 , k 2 :

[0014]

[0015] Among them, X w , Y w , Z w are coordinates in the world coordinate system; R and T are respectively the rotation matrix and the translation matrix for rigid body transformation from the world coordinate system to the camera pixel coordinate system; f is the focal length; d x , d y are the physical sizes of a pixel point in the x and y axis directions respectively; f x is the ratio of f / d x ; f y is the ratio of f / d y ; Zc is the scale factor; (u, v) are the coordinate values in the pixel coordinate system; (u 0 , v 0 ) is the center point coordinate when the coordinate in the image coordinate system is converted to the pixel coordinate system.

[0016] According to some preferred embodiments of the present invention, the distortion correction uses the following correction model:

[0017]

[0018] r 2 = x 2 + y 2 (3)

[0019] Among them, x and y are the image coordinates without removing the radial distortion and the tangential distortion; k 1 , k 2 are the radial distortion parameters obtained during camera calibration; p 1 , p 2 are the tangential distortion coefficients obtained during camera calibration; X, Y are the image coordinates after removing the radial distortion and the tangential distortion; r is the radius with the point at the center of the optical axis as the origin.

[0020] According to some preferred embodiments of the present invention, in the S3, the obtaining of the feature extraction network model includes:

[0021] S31 Obtain the labeled data set of street and / or tree images;

[0022] S32 Construct a feature extraction network model with image segmentation function;

[0023] S33 Perform image segmentation training or training and testing on the constructed feature extraction network model through the labeled data set to obtain the trained feature extraction network model;

[0024] Among them, the feature extraction network model constructed by S32 includes an encoding end and a decoding end;

[0025] The encoding end includes: a Dense-DRN-D-54 network containing the first to eighth convolutional module groups connected in sequence, an AIM-ASPP network connected to the output of the eighth convolutional module group, first and second convolutional layers respectively connected to the output of the third convolutional module group and the output of the eighth convolutional module group, a third convolutional layer connected to the output of the AIM-ASPP network, a first quadruple upsampling layer connected to the second convolutional layer, and a second quadruple upsampling layer connected to the third convolutional layer; among them, the first, second, and third convolutional layers are all 1×1 convolutional layers;

[0026] The Dense-DRN-D-54 network is obtained by adding dense operations to the 3rd, 4th, 5th, and 6th convolutional module groups in the DRN-D-54 network and removing the average pooling layer and fully connected layer in the 8 convolutional module groups of the DRN-D-54 network;

[0027] The AIM-ASPP network is obtained by adding an integrated interaction structure to the Atrous Spatial Pyramid Pooling (ASPP) network, and the integrated interaction structure refers to a structure that fuses information of different feature maps and integrates the extracted different features into one feature map;

[0028] The decoding end includes: a fusion layer that performs cross-layer splicing and fusion on multiple feature maps obtained by the decoding end, a fourth convolutional layer and a third quadruple upsampling layer connected to the fusion layer in sequence, and the fourth convolutional layer is a 3×3 convolutional layer.

[0029] According to some preferred embodiments of the present invention, the image processing performed by the encoding end includes:

[0030] Performing residual module operations on the input image 1 time, 1 time, 3 times, 4 times, 6 times, 3 times, 1 time, and 1 time respectively through the first to eighth convolutional module groups; after being processed by the first to third convolutional module groups, a low-order feature map is obtained; performing atrous convolutions with dilation rates of 2, 4, and 2 on the low-order feature map through the fourth to sixth convolutional module groups, and then obtaining a middle-order feature map after being processed by the seventh to eighth convolutional module groups;

[0031] After the middle-order feature map is input into the AIM-ASPP network:

[0032] Sequentially passing through 1×1 convolution, 3×3 convolution, 3×3 convolution, 3×3 convolution, and pooling processing to generate feature maps with dilation rates of 1, 6, 12, and 18 and a global pooling feature map, that is, A1 to A5 feature maps;

[0033] The obtained feature map Ai, where i = 2 to 4, is fused with its two adjacent feature maps Ai-1 and Ai+1 to obtain a fused feature map Bi; the feature map A1 and A2 are fused to obtain the feature map B1, and the feature map A5 and the feature map A4 are fused to obtain the feature map B5;

[0034] The obtained feature maps B1 to B5 are respectively fused with the feature maps A1 to A5 one by one to obtain feature maps C1 to C5;

[0035] The feature maps C1 to C5 are fused through channel concatenation to obtain a high-order feature map;

[0036] Subsequently, the low-order feature map and the middle-order feature map respectively pass through the first and second convolutional layers to obtain a post-convolution low-order feature map and a post-convolution middle-order feature map with reduced and equal numbers of channels;

[0037] The post-convolution middle-order feature map and the high-order feature map are subjected to 4-fold bilinear upsampling through the first four-fold upsampling layer and the second four-fold upsampling layer to obtain an upsampled middle-order feature map and an upsampled high-order feature map with the same size as the post-convolution low-order feature map.

[0038] According to some preferred embodiments of the present invention, the image processing performed by the decoding end includes:

[0039] Fusing the post-convolution low-order feature map, the upsampled middle-order feature map, and the upsampled high-order feature map in the fusion layer to obtain a fused image;

[0040] Performing feature refinement on the fused image through the fourth convolutional layer and adjusting the number of channels to 3 to obtain a refined feature map;

[0041] Performing 4-fold bilinear upsampling on the refined feature map through the third four-fold upsampling layer to output a predicted segmentation image.

[0042] According to some preferred embodiments of the present invention, the extraction of the tree factor includes:

[0043] S410 Set the tree area, the standard object area, and the background in the segmentation image to different colors, and obtain the number of longitudinal pixel points in the tree area and the standard object area;

[0044] S411 According to the true height of the standard object and the number of longitudinal pixel points in the tree area and the standard object area, calculate the tree height through the following formula:

[0045]

[0046] where H is the tree height, h is the true height of the standard object; P standard_ysum is the number of longitudinal pixel points in the standard object area; Ptree_ysum is the number of vertical pixels in the tree area.

[0047] According to some preferred embodiments of the present invention, the extraction of the tree factor includes:

[0048] The diameter at breast height of the tree is extracted through the following calculation formula:

[0049]

[0050] where P DBH is the diameter at breast height of the tree, h is the true vertical height of the reference object; f x , f y is the ratio of the internal parameters f / d x and f / d y of the camera, where f is the focal length; d x , d y is the physical size of a pixel in the x and y axis directions respectively; P 1.3h_xsum is the number of horizontal pixels corresponding to the measured diameter at breast height of the tree at a tree height of 1.3 m in the tree area, and P standard_ysum is the number of vertical pixels in the reference object area.

[0051] The present invention has the following beneficial effects:

[0052] (1) The extraction method of the present invention does not require a large amount of preprocessing of the image, nor does it need to obtain the camera shooting angle and the camera placement height in advance, and has a high accuracy.

[0053] (2) The extraction method of the present invention can directly use street view images to realize the automatic extraction of tree factors, and the extraction process is convenient.

[0054] (3) The extraction method of the present invention has low cost, high efficiency, and can also achieve batch processing. Description of the Drawings

[0055] Figure 1 is a schematic structural diagram of a feature extraction network model with an image segmentation function involved in the specific implementation manner.

[0056] Figure 2 is a schematic structural diagram of the original DRN-D-54 network structure involved in the feature extraction network model.

[0057] Figure 3 is a schematic structural diagram of the AIM-ASPP network structure involved in the feature extraction network model.

[0058] Figure 4 is a display of the image segmentation and tree parameter measurement methods in the embodiment.

[0059] Figure 5Schematic diagram of the Dense-DRN-D-54 network structure involved in the feature extraction network model. Detailed implementation manners

[0060] The present invention will be described in detail below in conjunction with embodiments and the accompanying drawings. However, it should be understood that the embodiments and the drawings are only used for exemplary description of the present invention, and do not constitute any limitation to the protection scope of the present invention. All reasonable transformations and combinations within the scope of the inventive concept of the present invention fall within the protection scope of the present invention.

[0061] According to the technical solution of the present invention, some specific implementation manners of the automatic tree factor extraction method based on AIM-ASPP include the following steps:

[0062] S1 Calibrate the measurement camera to obtain its internal parameters.

[0063] Among them, the calibration preferably uses the Zhang Zhengyou calibration method with non-linear distortion terms.

[0064] In some specific embodiments, a black and white checkerboard with a side length of 45 mm and arranged in 12×9 is selected as the calibration board.

[0065] In some specific embodiments, the measurement camera can be a mobile phone camera. First, a group of 15 calibration board checkerboards in different poses can be photographed by the mobile phone camera.

[0066] In some specific embodiments, during calibration, the world coordinate system is converted into the pixel coordinate system through the following calculation formula:

[0067]

[0068] In formula (1), X w , Y w , Z w are the coordinates in the world coordinate system; R and T are respectively the rotation matrix and the translation matrix for rigid body transformation from the world coordinate system to the camera pixel coordinate system; f is the focal length; d x , d y are the physical sizes of a pixel point in the x and y axis directions respectively; f x represents the ratio of f / d x ; f y represents the ratio of f / d y ; Zc is the scale factor; (u, v) are the coordinate values in the pixel coordinate system; (u 0 , v 0 ) are the center point coordinates when the coordinates in the image coordinate system are converted into the pixel coordinate system; among them, the internal parameters include f x , f y , u 0 , v0 , the external parameters include R and T, is a zero vector.

[0069] The calibration process is to obtain the internal parameter f in Equation (1) x , f y , u 0 , v 0 , the external parameters R, T and the distortion coefficients k 1 , k 2 , p 1 , p 2 .

[0070] In some specific embodiments, the calibration includes: performing corner detection on each collected image, solving to obtain the internal parameters and the external parameters; then using the least squares method to solve for the distortion coefficients; and finally using the maximum likelihood estimation method to optimize the solutions of the external parameters, internal parameters and distortion coefficients.

[0071] S2 performs distortion correction on the street tree image captured by the measurement camera according to the internal parameters obtained by calibration.

[0072] Among them, the distortion correction preferably uses the following correction model:

[0073]

[0074] r 2 = x 2 + y 2 (3)

[0075] Among them, x and y are the image coordinates without removing radial distortion and tangential distortion; k 1 , k 2 are the radial distortion parameters obtained during camera calibration; p 1 , p 2 are the tangential distortion coefficients obtained during camera calibration; X, Y are the image coordinates after removing radial distortion and tangential distortion; r represents the radius with the point at the center of the optical axis as the origin.

[0076] S3 segments the distorted-corrected image into trees, standard objects and background through a feature extraction network model with image segmentation function, where the standard object is the white lime coating on the tree.

[0077] In some specific embodiments, the construction of the segmentation model includes:

[0078] S31 obtains the labeled dataset of the image

[0079] In some specific embodiments, the present invention uses labelme to manually annotate the collected street and / or tree images, and the annotation types include three categories: trees, lime white coatings, and backgrounds.

[0080] The embodiment of the present invention using the lime white coating on the tree as a fixed reference object fully considers the characteristics that the lime white coating and the tree are in the same plane, the coating height is approximately the same, and it is easy to obtain.

[0081] In a specific embodiment, the collected images are images of different streets containing trees taken by a mobile phone camera in May under sufficient light, at long and short distances. The images contain environmental noises such as roads, people, cars, buildings, and grass, with a total of 382 images and a resolution of 1440×1920; it also includes 130 street view images with a resolution of 2048×1024 randomly selected from the publicly available dataset Cityscapes; and further uses methods such as contrast change, noise, translation, and mirroring to expand the dataset to 2560 images.

[0082] S32 Construct a feature extraction network model with image segmentation function

[0083] Preferably, referring to the appendix Figure 1 , the feature extraction network model includes an encoding end and a decoding end.

[0084] Furthermore, referring to the appendix Figure 1 , the encoding end includes a connected Dense-DRN-D-54 network and an AIM-ASPP network, two 1×1 convolutional layers connected to the Dense-DRN-D-54 network, namely the first and second convolutional layers, one 1×1 convolutional layer connected to the AIM-ASPP network, namely the third convolutional layer, a first quadruple upsampling layer connected to the second convolutional layer, and a second quadruple upsampling layer connected to the third convolutional layer.

[0085] Among them, referring to the appendix Figure 2 and the appendix Figure 5 , the Dense-DRN-D-54 network (appendix Figure 5 ) is obtained by adding a dense network structure to the DRN-D-54 network (appendix Figure 2 ).

[0086] Referring to the appendix Figure 2, the DRN-D-54 network is a simplified version of the third version DRN-C network in the DRN network (Chen Liang-Chieh, Zhu Yu-kun, Papandreou George, et al. Encoder-decoder with atrous separable convolution for semantic image segmentation [C] / / Proceedings of the European Conference on Computer Vision (ECCV), 2018: 801-818). It includes 54 layers of network, divided into 8 groups from 1 to 8. This group structure is cycled 3 times, 4 times, 6 times, and 3 times in the 3rd, 4th, 5th, and 6th groups respectively. Each group includes 2, 1, 9, 12, 18, 9, 1, and 1 convolutional module (CBR) groups respectively. Each convolutional module group contains 1 convolutional layer (C), 1 normalization layer (B), and 1 activation layer (R). (The convolutional kernel size of each convolutional layer is as shown by the numbers within the rectangular box), and the 1st to 8th convolutional module groups have 16, 32, 256, 512, 1024, 2048, 512, and 512 output channels respectively.

[0087] Based on the above DRN-D-54 network, referring to the appendix Figure 5 , in the Dense-DRN-D-54 network of the present invention, a dense network structure is used in the 3rd to 6th groups, and residual operations are performed 1 time, 2 times, 3 times, and 4 times respectively. The residual operation is to splice the feature map passed from the previous layer in front of the current layer with the feature map output by the current layer. In the dense network structure, it is to receive the feature maps passed from multiple previous layers and splice them with the feature map of the current layer. To avoid the problem of network degradation caused by gradient dispersion, the residual processing and dense network structure are not used in the 7th and 8th groups, which can prevent the aliasing effect generated by the previous dilated convolution from being transmitted through the residual or dense network structure, and the average pooling layer and fully connected layer after the 8 module groups in the DRN-D-54 network are removed. Among them, the feature map output by the 3rd convolutional module group in the Dense-DRN-D-54 network is 1 / 4 of the original image, denoted as the low-order feature map; dilated convolutions with dilation rates of 2, 4, and 2 are used in the 4th, 5th, and 6th convolutional module groups respectively. The last convolutional layer in the 3rd, 4th, 5th, and 6th convolutional module groups needs to receive the residual values passed from the previous convolutional layers before the activation operation. Finally, the feature map output after the CBR operation with a 3×3 convolutional kernel in the 8th convolutional module group has a size of 1 / 16 of the original image, denoted as the middle-order feature map, and the obtained middle-order feature map is further input into the AIM-ASPP network.

[0088] Further, referring to the appendixFigure 3 , the AIM-ASPP network adds an integrated interaction structure on the basis of the Atrous Spatial Pyramid Pooling (ASPP) network (Yu F, Koltun V, Funkhouser T. Dilated residual network[C] / / Proceddings of the IEEE conference on computer vision and pattern recognition. 2017: 472-480.) in DeepLab V3+. The integrated interaction structure refers to a structure that fuses information of different feature maps and integrates the extracted different features into one feature map, which can achieve interactive information transmission, increase the elements participating in effective calculation and the receptive field, and solve the discontinuity of information over long distances.

[0089] Among them, the way of channel splicing and fusion is used between each feature map. The number of channels of the feature maps in column A is 512. For the channels added after fusion, the CBR operation with a 1×1 convolution kernel is used to reduce the number of channels, and the number of channels after fusion is reduced to the same as that before fusion. The same is true for columns B and C. And the dropout function is used after the fusion in column C to prevent overfitting.

[0090] Specifically, the intermediate feature maps first go through the CBR processing of 1×1 convolution, 3×3 convolution, 3×3 convolution, and 3×3 convolution respectively (the dilation rates of the three 3×3 convolutions are 6, 12, and 18 respectively) to output 5 feature maps of the same size, namely the A1-A5 feature maps. Then, for the feature map Ai (i = 2-4) in A1-A5, it is fused with its adjacent two feature maps Ai-1 and Ai+1 to obtain the fused feature map Bi. The feature map A1 and A2 are fused to obtain the feature map B1, and the feature map A5 and A4 are fused to obtain the feature map B5. Then, the feature maps B1-B5 are fused with the feature maps A1-A5 one by one to obtain the feature maps C1-C5. The feature maps C1-C5 are subjected to channel splicing and fusion to obtain the high-order feature map, and the size of this feature map is 1 / 16 of the original image.

[0091] Further, referring to the appendix Figure 1, the number of channels of the low-level feature map and the middle-level feature map are 256 and 512 respectively. After passing through the first and second convolutional layers respectively, they are reduced to 48, so that a higher accuracy can be achieved when performing cross-layer splicing and fusion with the high-level feature map. Further, in order to better restore the edge detail information, the middle-level feature map after passing through the second convolutional layer and the high-level feature map are simultaneously bilinearly upsampled by 4 times through the first four-fold upsampling layer and the second four-fold upsampling layer to restore to the same size as the low-level feature map. Finally, the low-level feature map after passing through the first convolutional layer, the middle-level feature map after passing through the first four-fold upsampling layer, and the high-level feature map after passing through the second four-fold upsampling layer are cross-layer spliced and fused inside the decoding end.

[0092] Further, referring to the attached Figure 1 , the decoding end includes a fusion layer for performing the cross-layer splicing and fusion, a 3×3 convolutional layer connected to the fusion layer, and a third four-fold upsampling layer connected to the 3×3 convolutional layer. The 3×3 convolutional layer therein can refine the features of the fused image and adjust the number of channels to the number of categories to be segmented, which is 3. Then, through bilinear upsampling with a multiple of 4, the fused image with refined features is restored to the resolution of the original image, and the predicted segmentation image is output.

[0093] S33 Use the labeled data set to perform image segmentation training on the feature extraction network model to obtain a trained feature extraction network model

[0094] In some specific embodiments, the present invention randomly distributes the obtained labeled data set into a training set and a test set in a ratio of 8:2 to train and test the feature extraction network model.

[0095] Further, in a specific embodiment, the present invention uses the deep learning framework pytorch to implement the above training and testing in the CUDA 11.1 parallel computing environment, Ubuntu16.04 operating system, and python3.7 programming language. Stochastic Gradient Descent (SGD) method is used for training, the initial learning rate is set to 0.01, the batchsize is set to 4, the number of iterations per round is 516 times, and the epochs are set to 200.

[0096] S4 Extract tree factors from the segmentation image obtained from S3. The tree factors include tree height and tree diameter at breast height.

[0097] In some specific embodiments, the extraction of the tree factors includes:

[0098] S41 Extract tree height, including:

[0099] S410 sets the segmented image background, tree area, and reference object area to different colors, reads the colors of each pixel point in the tree area and the reference object area of the segmented image through the Python language, and compares the pixel coordinates in the tree and reference object areas respectively to obtain the number of pixels vertically included in the tree and reference object areas;

[0100] S411 obtains the true vertical pixel size of a single pixel in the image according to the true height of the reference object: the ratio of the true height of the reference object to the number of pixels vertically included in the reference object, and calculates the tree height through the following formula:

[0101]

[0102] Among them, the true height of the reference object is preferably the average value of the heights of multiple lime white coatings measured on site.

[0103] In formula (4), H is the tree height, h is the true vertical height of the reference object; P standard_ysum is the number of pixels vertically included in the reference object; P tree_ysum is the number of pixels vertically included in the tree; P tree_ysum and P white_ysum are in pixels, and the units of h and H are meters.

[0104] S42 extracts the tree diameter at breast height

[0105] Furthermore, the tree diameter at breast height is obtained through the following formula:

[0106]

[0107] In formula (5), P DBH is the tree diameter at breast height, h is the true vertical height of the reference object; f x 、f y are the aforementioned internal parameters of the camera; P 1.3h_xsum is the number of horizontal pixels occupied by the tree diameter at breast height at a tree height of 1.3 meters in the tree area, P standard_ysum is the number of pixels vertically included in the reference object.

[0108] Among them, preferably, the calculation method of P 1.3h_xsum is: calculate the vertical pixel coordinate value at 1.3 meters according to the true size mapped by each pixel, and then calculate the number of pixels included in the tree at this horizontal position..

[0109] In a specific embodiment, the segmentation of the tree area, reference object area, and background and the measurement method of tree parameters are as shown in the appendix Figure 4 as shown.

[0110] The above embodiments are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the concept of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, improvements and refinements made without departing from the principle of the present invention should also be regarded as within the protection scope of the present invention.

Claims

1. Automatic extraction method of tree factors based on AIM-ASPP, characterized in that, it includes: S1 Calibrate the measurement camera that obtains the street tree image to obtain its internal parameters; S2 Perform distortion correction on the street tree image captured by the measurement camera according to the internal parameters obtained by calibration to obtain a corrected image; S3 Segment the corrected image into trees, standard objects and background through a feature extraction network model with image segmentation function to obtain a segmented image divided into tree area, standard area and background area, wherein the standard object is a fixed object in the captured image that is on the same plane as the tree; S4 Calculate the tree factors by calculating parameters, and the calculation parameters include: pixel parameters of the tree area and the standard object area in the segmented image and the measured data of the standard object; the pixel parameters include the number of pixels and / or pixel coordinates; wherein, the feature extraction network model contains an integrated Atrous Spatial Pyramid Pooling network, namely AIM-ASPP network; In S3, the obtaining of the feature extraction network model includes: S31 Obtain an annotated data set of streets and / or tree images; S32 Construct a feature extraction network model with image segmentation function; S33 Perform image segmentation training or training and testing on the constructed feature extraction network model through the annotated data set to obtain a trained feature extraction network model; wherein, the feature extraction network model constructed in S32 includes an encoding end and a decoding end; The encoding end includes: a Dense-DRN-D-54 network containing the first to eighth convolutional module groups connected in sequence, an AIM-ASPP network connected to the output of the eighth convolutional module group, first and second convolutional layers respectively connected to the output of the third convolutional module group and the output of the eighth convolutional module group, a third convolutional layer connected to the output of the AIM-ASPP network, a first four-fold upsampling layer connected to the second convolutional layer, and a second four-fold upsampling layer connected to the third convolutional layer; wherein, the first, second, and third convolutional layers are all 1×1 convolutional layers; The Dense-DRN-D-54 network is obtained by adding dense operations to the 3rd, 4th, 5th, and 6th convolutional module groups in the DRN-D-54 network and removing the average pooling layer and fully connected layer in the 8 convolutional module groups of the DRN-D-54 network; The AIM-ASPP network is obtained by adding an integrated interaction structure to the Atrous Spatial Pyramid Pooling ASPP network, and the integrated interaction structure refers to a structure that fuses information of different feature maps and integrates the extracted different features into one feature map; The decoding end includes: a fusion layer that performs cross-layer splicing and fusion on multiple feature maps obtained by the decoding end, a fourth convolutional layer and a third four-fold upsampling layer connected to the fusion layer in sequence, and the fourth convolutional layer is a 3×3 convolutional layer.

2. The automatic extraction method of tree factors based on AIM-ASPP according to claim 1, characterized in that, the standard object is the lime white coating on the tree, and the tree factors include tree height and tree diameter at breast height.

3. The automatic extraction method of tree factors based on AIM-ASPP according to claim 1, characterized in that, The calibration uses the Zhang-Zhengyou calibration method, and the internal parameter f in the following conversion formula for converting the world coordinate system to the pixel coordinate system is obtained through the calibration x 、f y 、u 0 、v 0 , the external parameters R, T, and the tangential distortion coefficients p 1 、p 2 and the radial distortion parameters k 1 、k 2 : Among them, X w , Y w , Z w are coordinates in the world coordinate system; R and T are respectively the rotation matrix and the translation matrix for rigid body transformation from the world coordinate system to the camera pixel coordinate system; f is the focal length; d x , d y are the physical sizes of a pixel point in the x and y axis directions respectively; f x is the ratio of f / d x , f y is the ratio of f / d y ; Zc is the scale factor; (u, v) are the coordinate values in the pixel coordinate system; (u 0 , v 0 ) is the center point coordinate when the coordinate in the image coordinate system is transformed to the pixel coordinate system; is the zero vector.

4. The automatic extraction method of tree factors based on AIM-ASPP according to claim 1, characterized in that, the distortion correction uses the following correction model: r 2 = x 2 + y 2 (3) Among them, x and y are the image coordinates before removing radial distortion and tangential distortion; k 1 , k 2 are the radial distortion parameters obtained during camera calibration; p 1 , p 2 are the tangential distortion coefficients obtained during camera calibration; X and Y are the image coordinates after removing radial distortion and tangential distortion; r is the radius with the point at the center of the optical axis as the origin.

5. The automatic extraction method of tree factors based on AIM-ASPP according to claim 1, characterized in that, the image processing performed by the encoding end includes: performing residual module operations on the input image 1 time, 1 time, 3 times, 4 times, 6 times, 3 times, 1 time, and 1 time respectively through the first to eighth convolutional module groups; obtaining low-order feature maps after being processed by the first to third convolutional module groups; performing dilated convolutions with dilation rates of 2, 4, and 2 on the low-order feature maps through the fourth to sixth convolutional module groups, and then obtaining intermediate-order feature maps after being processed by the seventh to eighth convolutional module groups; after the intermediate-order feature maps are input into the AIM-ASPP network: successively performing 1×1 convolution, 3×3 convolution, 3×3 convolution, 3×3 convolution, and pooling processing to generate feature maps with dilation rates of 1, 6, 12, and 18 and a global pooling feature map, namely A1 to A5 feature maps; the obtained feature maps Ai, where i = 2 to 4 are fused with the two adjacent feature maps Ai-1 and Ai+1 to obtain fused feature maps Bi; the feature map A1 is fused with A2 to obtain the feature map B1, and the feature map A5 is fused with the feature map A4 to obtain the feature map B5; the obtained feature maps B1 to B5 are respectively fused with the feature maps A1 to A5 one by one to obtain feature maps C1 to C5; the feature maps C1 to C5 are fused through channel splicing to obtain high-order feature maps; subsequently, the low-order feature maps and the intermediate-order feature maps respectively pass through the first and second convolutional layers to obtain post-convolution low-order feature maps and post-convolution intermediate-order feature maps with reduced and equal numbers of channels; the post-convolution intermediate-order feature maps and the high-order feature maps are bilinearly upsampled 4 times through the first quadruple upsampling layer and the second quadruple upsampling layer to obtain an upsampled intermediate-order feature map and an upsampled high-order feature map with the same size as the post-convolution low-order feature maps.

6. The automatic extraction method of tree factors based on AIM-ASPP according to claim 5, characterized in that, the image processing performed by the decoding end includes: fusing the post-convolution low-order feature maps, the upsampled intermediate-order feature maps, and the upsampled high-order feature maps in the fusion layer to obtain a fused image; performing feature refinement on the fused image through the fourth convolutional layer and adjusting the number of channels to 3 to obtain a refined feature map; performing 4-fold bilinear upsampling on the refined feature map through the third quadruple upsampling layer to output a predicted segmentation image.

7. The automatic extraction method of tree factors based on AIM-ASPP according to claim 1, characterized in that, the extraction of the tree factors includes: S410 Set the tree regions, standard object regions, and background in the segmentation image to different colors, and obtain the number of vertical pixel points in the tree regions and standard object regions; S411 calculates the tree height according to the true height of the standard object and the number of longitudinal pixel points in the tree area and the standard object area through the following formula: Among them, H is the height of the tree, and h is the true longitudinal height of the standard object; P standard_ysum is the number of longitudinal pixel points in the standard object area; P tree_ysum is the number of longitudinal pixel points in the tree area.

8. The automatic extraction method of tree factors based on AIM-ASPP according to claim 1, characterized in that the extraction of the tree factors includes: extracting the diameter at breast height of the tree through the following calculation formula: Among them, P DBH is the tree diameter at breast height, h is the true longitudinal height of the standard object; f x , f y are f / d x and f / d y respectively, where f is the focal length; d x , d y are the physical sizes of a pixel point in the x and y axis directions respectively; P 1.3h_xsum is the number of horizontal pixel points corresponding to the measured tree diameter at breast height at a tree height of 1.3 m in the tree area, and P standard_ysum is the number of longitudinal pixel points in the standard object area.

Citation Information

Patent Citations

  • Standing tree determination method based on binocular vision

    CN111179335A