A method for instant cross-modal matching of mine single tree health diagnosis
Patent Information
- Application Number
- CN202610995578.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-06
- Publication Date
- 2026-09-25
AI Technical Summary
其一,矿区植被尺度差异大,现有网络对不同尺度目标的特征提取能力不足,小尺寸树木的细节信息在深层网络中丢失严重,导致边缘分割不准、小目标漏检率高,难以全面覆盖矿区各类单株树木
[0016]有益效果:本发明提出一种即时跨模态匹配的矿区单株树木健康诊断方法,通过构建深度残差双向多层特征提取网络对正射影像执行多尺度卷积与跨层加权融合,解决了矿区植被尺度差异大导致的小目标细节信息在网络深层传递中丢失的问题,显著提升了不同尺度单株树木的边缘分割精度和识别召回率;同时,本发明采用共享参数的并联多任务检测头同步完成像素级类别判定、空间位置回归和实例掩膜分离,在减少模型参数总量和降低过拟合风险的前提下,实现了矿区复杂环境下不同类型树木的精准识别。本发明通过引入跨模态匹配机制,依据识别阶段输出的定位框与轮廓掩膜从点云影像中精准裁剪单株树木点云并分离树冠点与非树冠点,以高度差分统计获取树高,同时对正射影像中的冠层区域差异化提取颜色、冠幅和纹理等多维健康表征向量,再将上述跨模态特征输入具有压缩-扩张结构的全连接度量网络进行多层次关联映射,最终通过级联识别与诊断模块将树木空间位置、类型及健康程度集成于同一输出结果中,从而解决了现有技术中视觉特征与形态参量分属不同模态难以高效融合、识别与诊断环节相互独立无法端到端一体化处理的缺陷,实现了矿区单株树木的即时精准识别与健康状态同步度量。
Smart Images

Figure CN122821157A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of health diagnosis technology for individual trees in mining areas, and in particular to a real-time cross-modal matching method for health diagnosis of individual trees in mining areas. Background Technology
[0002] Large-scale mining in mining areas leads to nutrient depletion and loose soil structure, making it difficult for planted vegetation to obtain the elements and nutrients needed for growth. This results in low survival rates and a high risk of degradation. Therefore, precise identification and health diagnosis of individual trees in mining areas are urgently needed. Traditional vegetation monitoring mainly relies on manual field surveys, which are costly and inefficient. While remote sensing imagery methods have improved upon these approaches, they struggle to effectively extract contextual features given the complex and diverse morphologies of vegetation in mining areas, leading to errors such as misidentification of similar vegetation with different spectral types. Machine learning methods based on UAV imagery, although improving accuracy, suffer from limitations such as subjective adjustment of model parameters and insufficient generalization ability.
[0003] In existing technologies, deep learning methods have been gradually introduced into vegetation identification in mining areas. By automatically extracting deep, detailed features of vegetation through multi-layer networks, the identification accuracy has been improved to some extent. In specific implementation, a network model is usually constructed using single-modal image data (such as orthophotos or point cloud data). Convolutional operations are used to extract visual features such as the spatial structure, edge texture, and color of trees. Then, a classifier or regression network is used to output tree type and location information. In the health diagnosis stage, indicators such as crown width, height, and color are manually extracted and then input into a traditional classification model for discrimination. The identification and diagnosis stages are relatively independent and do not form an end-to-end integrated processing flow.
[0004] However, existing technologies still have two shortcomings. First, the vegetation scale varies greatly in mining areas, and existing networks are insufficient in their ability to extract features from targets at different scales. Detailed information about small trees is severely lost in deep networks, leading to inaccurate edge segmentation, a high rate of missed detection for small targets, and difficulty in comprehensively covering all types of individual trees in mining areas. Second, the visual features required for tree recognition and the morphological parameters required for health diagnosis belong to different data modalities. Existing methods lack efficient matching and fusion mechanisms for cross-modal information, making it impossible to effectively integrate visual recognition results with diagnostic features. This results in a separation of recognition and diagnosis processes, low computational efficiency, and difficulty in achieving real-time integrated health assessment. Summary of the Invention
[0005] In order to overcome the shortcomings and deficiencies of the existing technology, the present invention provides a method for real-time cross-modal matching of the health diagnosis of a single tree in a mining area.
[0006] The technical solution adopted in this invention is a real-time cross-modal matching method for the health diagnosis of single trees in mining areas, comprising the following steps: S1, acquiring UAV images covering the mining area and generating orthophotos and point cloud images, performing image segmentation of both at the same scale, and constructing a tree sample dataset with one-to-one correspondence between the orthophotos and point cloud images; S2, based on a deep residual bidirectional multi-layer feature extraction network, performing multi-scale convolution and cross-layer weighted fusion operations on the input orthophotos to obtain a high-dimensional feature map set that fuses multi-level contextual responses; S3, using a parallel multi-task detection branch with shared parameters to simultaneously perform pixel-level class determination, spatial location regression, and instance mask separation on each layer of the high-dimensional feature map set. S4. Output the type label, bounding box, and contour mask of a single tree; S5. Based on the bounding box and contour mask, crop the point cloud of a single tree from the corresponding point cloud image, perform height difference statistics on the crown point and non-crown point to obtain the tree height, and simultaneously perform differential feature mapping on the crown layer of the single tree in the orthophoto image for color, crown width, and texture to obtain a multidimensional health representation vector; S6. Input the multidimensional health representation vector into a fully connected metric network with a compression-expansion structure, perform multi-level association mapping, and output a continuous quantitative index representing the health level of a single tree; S7. Cascade the outputs of S3 and S5 to generate an integrated monitoring result including the spatial location, type, and health level of a single tree.
[0007] Furthermore, the deep residual bidirectional multilayer feature extraction network in S2 performs a weighted fusion operation on the feature maps, adjusting the fusion weights of each layer's feature maps using the following calculation formula: , ,in, Indicates the first Layer feature map for the first The fusion weight coefficients of the layer feature maps For learnable log-scalar parameters, This indicates a bilinear upsampling operation. This indicates a max-pooling downsampling operation. For the first The output feature map after bidirectional fusion of the layers. and These are the original feature maps of the upper and lower layers that participated in the fusion.
[0008] Furthermore, the aforementioned The parallel multi-task detection branch in the model uses bounding box loss constraints for bounding box regression, and its loss function is: , ,in, For prediction boxes With real frame The intersection and union ratio, and These are the center coordinates of the predicted bounding box and the ground truth bounding box, respectively. It is a Euclidean distance metric. and These are the width and height of the prediction box, respectively. and These are the width and height of the actual bounding box, respectively. This is the diagonal length of the smallest bounding rectangle between the predicted bounding box and the ground truth bounding box. and These are the width and height of the smallest bounding rectangle, respectively.
[0009] Furthermore, the parallel multi-task detection branch in S3 uses a focus loss function to constrain the category determination, and its calculation formula is as follows: ,in, For the first The true category label for each pixel. The model predicts the first The probability value of each pixel belonging to the target category. To adjust the focusing parameters for the weights of easy and difficult samples, The loss value for class discrimination.
[0010] Furthermore, in step S4, when performing height difference statistics on canopy points and non-canopy points, the average of the top 5% of height values in the canopy point set is extracted as the canopy elevation parameter, and the average of the bottom 5% of height values in the non-canopy point set is extracted as the ground surface elevation parameter. The difference between the two is used as a measure of the height of a single tree, calculated as follows: ,in, The height of a single tree. This is the set of indices of the top 5% of tree crown points, sorted by height. The total number of points in this set. The first point in the set of tree canopy points Elevation values of each point This is the set of indexes for 5% of the non-crown point sets after height sorting. The total number of points in this set. The first point in the set of non-canopy points Elevation values of each point.
[0011] Furthermore, the fully connected metric network in S5 uses the cross-entropy loss function for parameter optimization, the expression of which is: ,in, The total number of samples used in loss calculation. For the first The true health status label of each sample The first output of the fully connected metric network The predicted probability value of the health status of each sample This represents the loss value for the health diagnosis branch.
[0012] Further, S2 includes the following sub-steps: S21, compressing the spatial dimension of the input orthophoto image through initial convolution and pooling operations to extract primary edge and texture response maps; S22, progressively downsampling the primary response maps through five residual modules to output five deep feature maps with different spatial resolutions; S23, performing channel unification mapping on each deep feature map to transform the number of channels of each feature map to the same dimension; S24, performing bidirectional cross-scale weighted aggregation operation on each layer of feature maps after channel unification, wherein the top-down path upsamples high-level features and superimposes them with low-level features, and the bottom-up path downsamples low-level features and superimposes them with high-level features, finally outputting a set of six high-dimensional feature maps fused with multi-scale context.
[0013] Further, step S3 includes the following sub-steps: S31, inputting the six-layer high-dimensional feature maps into a parallel detection head with shared parameters, each detection head including four serial convolutional modules and an end convolutional layer; S32, the segmentation branch in each detection head performs foreground and background binary classification on the feature map pixel by pixel, generating an instance mask probability map of a single tree; S33, the classification branch in each detection head performs multi-class probability prediction of tree type on the feature map pixel by pixel, outputting confidence scores for each category; S34, the regression branch in each detection head performs offset regression of the predicted bounding box center coordinates and width and height parameters on the feature map pixel by pixel, outputting the spatial positioning box of a single tree.
[0014] Further, S4 includes the following sub-steps: S41, cropping the minimum bounding rectangle subset of the point cloud of a single tree from the point cloud image based on the positioning box; S42, separating the canopy point set and the non-canopy point set from the point cloud subset based on the spatial distribution of the contour mask; S43, performing height sorting on the canopy point set and the non-canopy point set respectively, and taking the difference between the mean values of the two ends to obtain the height parameter of a single tree; S44, performing color mean statistics, canopy diameter calculation, and texture encoding extraction through multiple concatenated convolutional layers on the canopy region defined by the contour mask in the orthophoto image, and combining them to form a multidimensional health representation vector.
[0015] Further, S5 includes the following sub-steps: S51, inputting the multidimensional health representation vector into the first hidden layer, which includes five neurons, and performing linear transformation and nonlinear activation on the input vector to expand the feature dimension; S52, inputting the output of the first hidden layer into the second hidden layer, which includes three neurons, and performing dimensionality reduction compression on the features to remove redundant correlation information; S53, inputting the output of the second hidden layer into the third hidden layer, which includes five neurons, and performing expansion mapping on the compressed features to restore the comprehensiveness of the correlation expression; S54, mapping the output of the third hidden layer to a preset numerical range through an activation function, and outputting a continuous quantitative index representing the health level of a single tree.
[0016] Beneficial effects: This invention proposes a real-time cross-modal matching method for the health diagnosis of individual trees in mining areas. By constructing a deep residual bidirectional multi-layer feature extraction network to perform multi-scale convolution and cross-layer weighted fusion on orthophotos, it solves the problem of loss of small target detail information in deep network transmission caused by large differences in vegetation scale in mining areas, and significantly improves the edge segmentation accuracy and recognition recall rate of individual trees at different scales. At the same time, this invention uses a parallel multi-task detection head with shared parameters to simultaneously complete pixel-level category determination, spatial location regression and instance mask separation. Under the premise of reducing the total number of model parameters and reducing the risk of overfitting, it achieves accurate identification of different types of trees in the complex environment of mining areas. This invention introduces a cross-modal matching mechanism, accurately cropping the point cloud of a single tree from the point cloud image based on the positioning box and contour mask output in the recognition stage, and separating the canopy points from the non-canopy points. The tree height is obtained by height difference statistics. At the same time, multi-dimensional health representation vectors such as color, canopy width, and texture are extracted differentially from the canopy region in the orthophoto. The above cross-modal features are then input into a fully connected metric network with a compression-expansion structure for multi-level correlation mapping. Finally, the spatial location, type, and health status of the tree are integrated into the same output result through a cascaded recognition and diagnosis module. This solves the defects of the prior art, where visual features and morphological parameters belong to different modalities and are difficult to integrate efficiently, and the recognition and diagnosis links are independent and cannot be integrated into an end-to-end process. It realizes the real-time and accurate recognition of single trees in the mining area and the synchronous measurement of their health status. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating the overall process of the method of the present invention. Figure 2 This is a framework diagram of the single-plant health diagnosis model in the mining area of the present invention; Figure 3 This is a diagram of the bidirectional multilayer feature extraction network of the present invention; Figure 4 This is the multi-task detection diagram of the present invention. Detailed Implementation
[0018] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0019] like Figure 1 As shown, a real-time cross-modal matching method for single-tree health diagnosis in mining areas includes the following steps: S1, acquiring UAV images covering the mining area and generating orthophotos and point cloud images, performing image segmentation of both at the same scale, and constructing a tree sample dataset with one-to-one correspondence between the orthophotos and point cloud images; S2, based on a deep residual bidirectional multi-layer feature extraction network, performing multi-scale convolution and cross-layer weighted fusion operations on the input orthophotos to obtain a high-dimensional feature map set that fuses multi-level contextual responses; S3, using a parallel multi-task detection branch with shared parameters to simultaneously perform pixel-level category determination, spatial location regression, and instance mask separation on each layer of the high-dimensional feature map set, outputting a single-tree image. S4. Based on the location box and contour mask, crop the point cloud of a single tree from the corresponding point cloud image, perform height difference statistics on the crown point and non-crown point to obtain the tree height, and simultaneously perform differential feature mapping of color, crown width and texture on the crown layer of the single tree in the orthophoto to obtain a multidimensional health representation vector; S5. Input the multidimensional health representation vector into a fully connected metric network with a compression-expansion structure, perform multi-level association mapping, and output a continuous quantitative index representing the health level of a single tree; S6. Cascade the output of S3 and the output of S5 to generate an integrated monitoring result including the spatial location, type and health level of a single tree.
[0020] Step S1 involves acquiring UAV images covering the mining area and generating orthophotos and point cloud images. Both images are then segmented at the same scale to construct a tree sample dataset with a one-to-one correspondence between the orthophotos and point cloud images. Specifically, under clear and windless weather conditions, the mining area is selected for UAV aerial photography. Before takeoff, flight direction, flight overlap, and flight altitude are set using flight path planning software to ensure the images cover the entire research area. After acquisition, UAV aerial image processing software is used to perform geographic information calibration, geometric calibration, and image stitching operations on the original images, generating orthophotos and point cloud images of the mining area in 3D map mode. The generated orthophotos and point cloud images are simultaneously divided into basic image units of 512 pixels by 512 pixels, maintaining a one-to-one spatial correspondence between the orthophotos and point cloud images within each unit. All basic image units are randomly divided into three subsets according to a ratio of 60% for the training set, 30% for the validation set, and 10% for the test set. Using annotation tools, the bounding boxes and true canopy extents of each tree are drawn on the orthophoto. Tree type labels are assigned to each annotated object, generating files that include annotation information and converting the annotation information into a standard format to form a complete sample dataset for model training, validation, and testing.
[0021] Step S2 involves using a deep residual bidirectional multi-layer feature extraction network to perform multi-scale convolution and cross-layer weighted fusion operations on the input orthophoto image, obtaining a high-dimensional feature map set that fuses multi-level contextual responses. Specifically, a tree sample image with dimensions of 1024 pixels by 1024 pixels by three channels is input into the network backbone. It first passes through an initial convolutional layer and a max-pooling layer to suppress background information, extracting low-level image features of edges and textures, and outputting a feature map with dimensions of 256 pixels by 256 pixels by 64 channels. This feature map is then progressively downsampled through five residual modules, outputting five deep feature maps with dimensions of 128 pixels by 128 pixels by 256 channels, 64 pixels by 64 pixels by 512 channels, 32 pixels by 32 pixels by 1024 channels, 16 pixels by 16 pixels by 2048 channels, and 8 pixels by 8 pixels by 4096 channels, respectively. Each of the five feature maps is convolved with a kernel size of 1 pixel by 1 pixel, transforming the number of channels in each feature map to a uniform 256 channels. The five channel-unified feature maps are then input into a bidirectional cross-scale aggregation structure. This structure performs upsampling on the high-level feature maps and weighted superposition with the low-level feature maps, while simultaneously performing downsampling on the low-level feature maps and weighted superposition with the high-level feature maps. The fusion weights of each layer are dynamically adjusted using learnable logarithmic scalar parameters. The final output is a set of six high-dimensional feature maps that integrate multi-scale contextual information.
[0022] Step S3 involves using a parallel multi-task detection branch with shared parameters to simultaneously perform pixel-level category determination, spatial location regression, and instance mask separation on each layer of the high-dimensional feature map set, outputting the type label, bounding box, and contour mask for each tree. Specifically, the six layers of high-dimensional feature maps are input into a parallel detection head with shared parameters. Each detection head includes four sequential convolutional module layers and one terminal convolutional layer, and all detection heads share the same network weight parameters. Each detection head has three parallel output branches: a segmentation branch, a classification branch, and a regression branch. The segmentation branch performs binary classification of foreground and background for each pixel in the feature map, generating an instance mask probability map for each tree and determining the precise contour range of the tree in the image. The classification branch performs multi-class probability prediction of tree type for each pixel in the feature map, outputting a confidence score for each tree type and determining the tree category to which each pixel belongs. The regression branch performs bounding box parameter regression for each pixel in the feature map, outputting four parameter values: center point horizontal and vertical coordinate offset, bounding box width, and bounding box height, to determine the spatial bounding box of the tree. The prediction results from the three branches are combined at the pixel level to finally output the type label, rectangular positioning box, and contour mask for each tree.
[0023] Step S4: Based on the positioning box and contour mask, crop the point cloud of a single tree from the corresponding point cloud image. Perform height difference statistics on the canopy points and non-canopy points to obtain the tree height. Simultaneously, perform differential feature mapping on the canopy of the single tree in the orthophoto image for color, canopy width, and texture to obtain a multidimensional health representation vector. In specific implementation, based on the positioning box coordinate parameters output in step S3, crop the minimum bounding rectangle subset of the point cloud of a single tree from the corresponding point cloud image. Based on the spatial distribution information of the contour mask, divide the points in the point cloud subset into a canopy point set and a non-canopy point set. The point cloud corresponding to the area inside the mask is marked as a canopy point, and the point cloud corresponding to the area outside the mask is marked as a non-canopy point. Sort the elevation values of each point in the canopy point set in descending order, and calculate the arithmetic mean of the top 5% of elevation values as the canopy elevation parameter. The elevation values of all points in the non-canopy point set are sorted in ascending order, and the arithmetic mean of the remaining 5% of elevation values is calculated as the ground elevation parameter. The canopy elevation parameter is subtracted from the ground elevation parameter, and the difference is taken as the height value of an individual tree. Simultaneously, three-dimensional feature extraction operations are performed on the canopy region defined by the contour mask in the orthophoto image. Color feature extraction uses convolutional layers combined with fully connected layers and pooling modules to output canopy color statistics. Canopy width feature extraction uses fully connected layers combined with a summation module to output the canopy diameter value. Texture feature extraction uses four concatenated convolutional layers to output the canopy texture encoding vector, which is then combined to form a multi-dimensional health representation vector.
[0024] Step S5: Input the multidimensional health representation vector into a fully connected metric network with a compression-expansion structure, perform multi-level association mapping, and output a continuous quantitative index representing the health level of a single tree. Specifically, a fully connected metric network with three hidden layers is constructed, with the number of neurons in each hidden layer set as follows: five neurons in the first layer, three neurons in the second layer, and five neurons in the third layer. The multidimensional health representation vector obtained in step S4 is input into the first hidden layer. This layer performs linear matrix multiplication on the input vector through five neurons, adds a bias term, and then processes it with a nonlinear activation function to expand the feature dimension to five-dimensional space, capturing the initial association relationship between the four health representation parameters. The output vector of the first hidden layer is passed to the second hidden layer. This layer performs linear transformation and nonlinear activation through three neurons, compressing the five-dimensional features into three-dimensional space, removing redundant association information, and retaining the core feature combination. The output vector of the second hidden layer is passed to the third hidden layer. This layer performs linear transformation and nonlinear activation through five neurons, expanding the three-dimensional features back into five-dimensional space, and refining the compressed association relationship. The output vector of the third hidden layer is mapped to a closed numerical range through an activation function, and a continuous value within that range is output as a quantitative indicator of the health of a single tree.
[0025] Step S6 involves cascading the outputs of Step S3 and Step S5 to generate an integrated monitoring result including the spatial location, type, and health status of individual trees. Specifically, the tree type label, rectangular bounding box, and contour mask output by each detection head in Step S3 are cached as structured recognition result data, and a unique spatial index number is assigned to each tree. The health status quantification index output in Step S5 is associated with the spatial index number of the tree. The type label, the four coordinate parameters of the rectangular bounding box, the pixel-by-pixel binary matrix of the contour mask, and the health status quantification index of each tree are integrated into a complete monitoring record. All tree monitoring records are arranged according to their spatial coordinates to form a tree distribution and health monitoring data matrix covering the entire mining area. The data matrix is mapped to the geospatial coordinate system of the original orthophoto image to generate an integrated monitoring result map including the spatial location, type category, and health quantification value of each tree, achieving an integrated output from tree detection to health assessment.
[0026] Preferably, the deep residual bidirectional multilayer feature extraction network in S2 performs a weighted fusion operation on the feature maps, adjusting the fusion weights of each layer's feature maps using the following calculation formula: , ,in, Indicates the first Layer feature map for the first The fusion weight coefficients of the layer feature maps For learnable log-scalar parameters, This indicates a bilinear upsampling operation. This indicates a max-pooling downsampling operation. For the first The output feature map after bidirectional fusion of the layers. and These are the original feature maps of the upper and lower layers that participated in the fusion.
[0027] Specifically, the bidirectional cross-scale feature fusion mechanism defines the calculation method of fusion weights in the first formula. For each feature map layer, the fusion weight coefficient between it and other layers is obtained by dividing the exponent value of a learnable logarithmic scalar parameter by the sum of the exponent values of all learnable parameters. This design draws on the idea of soft maximization normalization to ensure that the sum of the fusion weight coefficients of all sources in the same layer is equal to 1, so that features of different scales can be adaptively allocated according to their respective importance during fusion, avoiding the subjectivity of manually setting fixed weights. The second formula describes the specific operation of weighted fusion. For the output feature map of the l-th layer, its value is equal to the sum of the feature maps of all participating layers after upsampling or downsampling operations and the corresponding fusion weight coefficients. The high-level feature maps are upsampled to the spatial size of the l-th layer through bilinear interpolation, and the low-level feature maps are downsampled to the spatial size of the l-th layer through max pooling. Each learnable parameter is iteratively updated during network training through the backpropagation algorithm. The initial value is usually set to 0, and it converges to the optimal value after 500 iterations of training. The bidirectional fusion mechanism ensures the full interaction between high-level semantic information and low-level detail information at six scale levels.
[0028] Preferably, the The parallel multi-task detection branch in the model uses bounding box loss constraints for bounding box regression, and its loss function is: , ,in, For prediction boxes With real frame The intersection and union ratio, and These are the center coordinates of the predicted bounding box and the ground truth bounding box, respectively. It is a Euclidean distance metric. and These are the width and height of the prediction box, respectively. and These are the width and height of the actual bounding box, respectively. This is the diagonal length of the smallest bounding rectangle between the predicted bounding box and the ground truth bounding box. and These are the width and height of the smallest bounding rectangle, respectively.
[0029] Specifically, the bounding box regression loss constraint decomposes the bounding box loss into four components. The intersection-union ratio (IU) loss term is calculated by subtracting the IU of the predicted and ground truth bounding boxes from 1. The IU is defined as the area of the intersection of the predicted and ground truth bounding boxes divided by the area of their union. The closer this value is to 1, the more accurate the localization, and therefore the closer the loss term is to 0. The distance loss term is calculated by dividing the square of the Euclidean distance between the center points of the predicted and ground truth bounding boxes by the square of the diagonal length of their smallest bounding rectangle. This ratio normalizes the distance loss to the range of 0 to 1, eliminating the influence of different image scales on the distance dimension. The width loss term is calculated by dividing the square of the difference between the width of the predicted and ground truth bounding boxes by the square of the width of the smallest bounding rectangle. The height loss term is calculated by dividing the square of the difference between the height of the predicted and ground truth bounding boxes by the square of the height of the smallest bounding rectangle. These two terms constrain the aspect ratio of the bounding boxes. The second formula is the definition of the IU calculation: the area of the intersection of the predicted and ground truth bounding boxes divided by the area of their union. The total bounding box loss value is obtained by adding the four loss terms. This loss value is minimized in each iteration using the gradient descent algorithm. The learning rate is set to 0.0001, and the adaptive moment estimation optimizer is selected as the optimizer.
[0030] Preferably, the parallel multi-task detection branch in S3 uses a focus loss function to constrain the category determination, and its calculation formula is as follows: ,in, For the first The true category label for each pixel. The model predicts the first The probability value of each pixel belonging to the target category. To adjust the focusing parameters for the weights of easy and difficult samples, The loss value for class discrimination.
[0031] Specifically, the pixel-level class determination loss function calculates and sums the loss values for each pixel involved in training. For each pixel, when the true class label is 1, the loss term is the power of gamma of the difference between 1 and the model's predicted probability, multiplied by the negative logarithm of the predicted probability. When the true class label is 0, the loss term is the power of gamma of the predicted probability, multiplied by 1, minus the negative logarithm of the predicted probability. The gamma parameter is set to 0.9. This design introduces a modulating factor. For easily classified positive samples with predicted probabilities close to 1, the modulating factor approaches 0, significantly reducing the contribution of these samples to the total loss and allowing the network to focus its training on difficult-to-classify samples with lower predicted probabilities. A similar symmetrical processing mechanism is used for negative samples; when the predicted probability approaches 0, the modulating factor approaches 0, reducing the weight of easily classified negative samples. This loss function is used in conjunction with the bounding box regression loss throughout the network training process. The first four convolutional modules of the multi-task detection head are responsible for extracting features, while the last convolutional layer is responsible for outputting the class prediction probability. The total class loss is obtained by summing the loss values of all pixels and updating the network weight parameters through backpropagation.
[0032] Preferably, in step S4, when performing height difference statistics on canopy points and non-canopy points, the average of the top 5% of height values in the canopy point set is extracted as the canopy elevation parameter, and the average of the bottom 5% of height values in the non-canopy point set is extracted as the ground surface elevation parameter. The difference between the two is used as a measure of the height of a single tree, and its calculation method is as follows: ,in, The height of a single tree. This is the set of indices of the top 5% of tree crown points, sorted by height. The total number of points in this set. The first point in the set of tree canopy points Elevation values of each point This is the set of indexes for 5% of the non-crown point sets after height sorting. The total number of points in this set. The first point in the set of non-canopy points Elevation values of each point.
[0033] Specifically, in step S4, the height difference statistics between the canopy point and non-canopy points involve arithmetically averaging the elevation values of the top 5% of points in the canopy point set to obtain the canopy elevation parameter. Selecting the top 5%, rather than all canopy points, eliminates outliers caused by low branches or leaves obscuring the canopy, ensuring the canopy elevation parameter represents the true height of the top of the canopy. Simultaneously, the elevation values of the bottom 5% of points in the non-canopy point set are arithmetically averaged to obtain the ground elevation parameter. Selecting the bottom 5%, rather than all non-canopy points, eliminates outliers caused by weeds, low shrubs, or rocks, ensuring the ground elevation parameter represents the true ground elevation. Subtracting the ground elevation parameter from the canopy elevation parameter yields the height of a single tree. The first part of the formula represents the average elevation of the top 5% of points in the tree canopy point set by height sorting, and the second part represents the average elevation of the bottom 5% of points in the non-tree canopy point set by height sorting. Subtracting the two parts cancels out the systematic errors caused by ground undulations and terrain changes, making the extracted tree height highly adaptable to terrain and robust to measurement.
[0034] Preferably, the fully connected metric network in S5 uses the cross-entropy loss function for parameter optimization, the expression of which is: ,in, The total number of samples used in loss calculation. For the first The true health status label of each sample The first output of the fully connected metric network The predicted probability value of the health status of each sample This represents the loss value for the health diagnosis branch.
[0035] Specifically, the loss function used in the fully connected metric network calculates the loss value for each sample participating in the health diagnosis training, sums them, and then divides the sum by the total number of samples to obtain the average loss value. For each sample, when the true health status label is 1, the loss term is the negative logarithm of the predicted health status probability value output by the model; when the true health status label is 0, the loss term is 1 minus the negative logarithm of the predicted probability value. The predicted probability value is obtained by mapping the output of the third hidden layer of the fully connected metric network to the interval between 0 and 1 through an activation function. This activation function compresses any real number input into a continuous value between 0 and 1, which precisely matches the expression range of the health status quantification index. The summation symbol in the formula represents the accumulation of the loss terms for all samples, divided by the total number of samples to obtain the average loss value, which is used to characterize the overall error of the model under the current parameters. This loss function works in conjunction with the optimizer to calculate the gradient of the loss value with respect to all network weight parameters in each training iteration. The parameters are then updated along the negative gradient direction using the gradient descent algorithm with a learning rate of 0.0001 and the number of iterations is set to 500, so that the loss value gradually decreases until convergence, thus completing the optimization of the parameters of the fully connected metric network.
[0036] Preferably, step S2 includes the following sub-steps: S21, compressing the spatial dimension of the input orthophoto image through initial convolution and pooling operations to extract primary edge and texture response maps; S22, progressively downsampling the primary response maps through five residual modules to output five deep feature maps with different spatial resolutions; S23, performing channel unification mapping on each deep feature map to transform the number of channels of each feature map to the same dimension; S24, performing bidirectional cross-scale weighted aggregation operation on each layer of feature maps after channel unification, wherein the top-down path upsamples high-level features and superimposes them with low-level features, and the bottom-up path downsamples low-level features and superimposes them with high-level features, finally outputting a set of six high-dimensional feature maps fused with multi-scale context.
[0037] Specifically, in step S21, the input orthophoto image with a size of 1024 pixels by 1024 pixels by 3 channels is processed sequentially through an initial convolutional layer and a max pooling layer. The kernel size of the initial convolutional layer is 7 pixels by 7 pixels, the stride is set to 2, and the number of output channels is 64. The window size of the max pooling layer is 3 pixels by 3 pixels, and the stride is set to 2. This stage compresses the spatial size of the input image to 256 pixels by 256 pixels, while reducing the amount of information in the background area and initially extracting low-level image features of edges and textures. In step S22, the feature map output in step S21 is progressively downsampled through five residual modules. The five residual modules output feature maps with spatial sizes of 128 pixels by 128 pixels, 64 pixels by 64 pixels, 32 pixels by 32 pixels, 16 pixels by 16 pixels, and 8 pixels by 8 pixels, respectively, with corresponding number of channels of 256, 512, 1024, 2048, and 4096. Step S23: Perform channel unification mapping on the five feature maps output in step S22 using a convolution operation with a kernel size of 1 pixel by 1 pixel, transforming the number of channels in all five feature maps to 256, facilitating subsequent cross-layer fusion operations. Step S24: Input the five channel-unified feature maps into a bidirectional cross-scale aggregation structure. In this structure, circles represent convolution operations with a kernel size of 3 pixels by 3 pixels, downward and lower left dashed arrows represent downsampling the feature maps to match their size with the previous layer's feature maps and performing weighted superposition, and upward dashed arrows represent bilinear upsampling the feature maps to match their size with the next layer's feature maps and performing weighted superposition. Simultaneously, the smallest feature map is upsampled to generate an additional layer of feature maps, finally outputting a set of fused feature maps across six scales.
[0038] Preferably, step S3 includes the following sub-steps: S31, inputting the six-layer high-dimensional feature maps into a parallel detection head with shared parameters, each detection head including four serial convolutional modules and an end convolutional layer; S32, the segmentation branch in each detection head performs foreground and background binary classification on the feature map pixel by pixel, generating an instance mask probability map of a single tree; S33, the classification branch in each detection head performs multi-class probability prediction of tree type on the feature map pixel by pixel, outputting confidence scores for each category; S34, the regression branch in each detection head performs offset regression of the predicted bounding box center coordinates and width and height parameters on the feature map pixel by pixel, outputting the spatial positioning box of a single tree.
[0039] Specifically, in step S31, the six-layer high-dimensional feature map output from step S24 is input into six parallel detection heads with shared parameters for processing. Each detection head includes four serially connected convolutional module layers and one terminal convolutional layer. The kernel size of the four convolutional module layers is 3 pixels by 3 pixels, the stride is 1, and the padding is 1. The modules undergo non-linear transformation through activation functions. The network weight parameters are identical across all detection heads. This shared parameter mechanism allows feature maps of different scales to be processed through the same mapping function, enhancing the model's consistency in responding to targets of different scales. In step S32, the segmentation branch within each detection head performs foreground and background binary classification on the input feature map pixel by pixel, outputting an instance mask probability map for each tree. Pixels with a probability value greater than 0.5 are classified as foreground tree regions, and pixels with a probability value less than or equal to 0.5 are classified as background regions. In step S33, the classification branch within each detection head performs multi-class probability prediction of tree type on the input feature map pixel by pixel, outputting a confidence score for each tree type category, and selecting the category with the highest confidence score as the tree type determination result for that pixel. Step S34: The regression branch inside each detection head performs prediction box parameter regression on the input feature map pixel by pixel, and outputs four parameter values: the horizontal and vertical coordinate offset of the center point, the width of the prediction box, and the height of the prediction box. Based on the four parameters, the spatial location box of each tree is determined.
[0040] Preferably, step S4 includes the following sub-steps: S41, cropping the minimum bounding rectangle subset of the point cloud of a single tree from the point cloud image based on the positioning box; S42, separating the canopy point set and the non-canopy point set from the point cloud subset based on the spatial distribution of the contour mask; S43, performing height sorting on the canopy point set and the non-canopy point set respectively, and taking the difference between the mean values of the two ends to obtain the height parameter of a single tree; S44, performing color mean statistics, canopy diameter calculation, and texture encoding extraction through multiple concatenated convolutional layers on the canopy region defined by the contour mask in the orthophoto image, and combining them to form a multidimensional health representation vector.
[0041] Specifically, in step S41, based on the four coordinate parameters of the positioning box output in step S34, a minimum bounding rectangle subset of the point cloud for a single tree is cropped from the point cloud image at the corresponding spatial location. The boundary of this rectangular region completely coincides with the positioning box, ensuring that only the point cloud data around the tree is included. In step S42, based on the binary matrix of the contour mask output in step S32, each point in the point cloud subset cropped in step S41 is divided into a canopy point set and a non-canopy point set. The point cloud corresponding to the pixel position with a value of 1 in the mask matrix is marked as a canopy point, and the point cloud corresponding to the pixel position with a value of 0 in the mask matrix is marked as a non-canopy point. Step S43: Sort the elevation values of all points in the canopy point set in descending order, and calculate the arithmetic mean of the top 5% of elevation values as the canopy elevation parameter. Sort the elevation values of all points in the non-canopy point set in ascending order, and calculate the arithmetic mean of the bottom 5% of elevation values as the ground elevation parameter. Subtract the ground elevation parameter from the canopy elevation parameter to obtain the height value of a single tree. Step S44: Perform differential feature extraction on the canopy region defined by the contour mask in the orthophoto. The canopy width feature is output as the canopy coverage diameter value through a fully connected layer and a summing module. The color feature is output as the canopy three-channel color mean value through a convolutional layer, a fully connected layer, and a pooling module. The texture feature is output as the texture encoding vector through four cascaded convolutional layers with a kernel size of 3 pixels by 3 pixels.
[0042] Preferably, step S5 includes the following sub-steps: S51, inputting the multidimensional health representation vector into the first hidden layer, which includes five neurons, and performing linear transformation and nonlinear activation on the input vector to expand the feature dimension; S52, inputting the output of the first hidden layer into the second hidden layer, which includes three neurons, and performing dimensionality reduction compression on the features to remove redundant correlation information; S53, inputting the output of the second hidden layer into the third hidden layer, which includes five neurons, and performing expansion mapping on the compressed features to restore the comprehensiveness of the correlation expression; S54, mapping the output of the third hidden layer to a preset numerical range through an activation function, and outputting a continuous quantitative index representing the health status of a single tree.
[0043] Specifically, in step S51, the multidimensional health representation vector obtained in step S44 is input into the first hidden layer. This layer includes five neurons, and each neuron establishes fully connected weights with each dimension of the input vector. After performing linear matrix multiplication on the input vector, a bias term is added, and then processed by a nonlinear activation function to expand the input features to a five-dimensional space, which is used to capture the initial correlation and interaction effects between the four health representation parameters. In step S52, the output vector from step S51 is passed into the second hidden layer, which includes three neurons. Each neuron establishes fully connected weights with the five neurons in the first hidden layer. Linear transformation and nonlinear activation are performed to compress the five-dimensional features into a three-dimensional space. This dimensionality reduction operation removes redundant correlation information and noise interference, retaining the core feature combination that contributes most to health diagnosis. Step S53: The output vector from step S52 is fed into the third hidden layer, which consists of five neurons. Each neuron establishes fully connected weights with the three neurons in the second hidden layer. Linear transformation and nonlinear activation are performed to expand the three-dimensional features back to five-dimensional space, refining the compressed core feature combination and restoring the comprehensiveness of the association prediction. Step S54: The output of the third hidden layer is mapped to a closed numerical range of 0 to 1 using an activation function. The continuous values within this range are output as a quantitative indicator of the health of a single tree.
[0044] like Figure 2The diagram shows the framework of the single-tree health diagnosis model for mining areas according to the present invention. The overall architecture is an end-to-end cross-modal integrated diagnostic framework, encompassing six functional modules from top to bottom: a data input layer, a feature extraction layer, a multi-task recognition layer, a cross-modal representation layer, a health measurement layer, and a result integration layer. This fully covers the entire processing flow from raw images to diagnostic results. The input side simultaneously receives orthophotos and point cloud images collected by UAVs in the mining area. These images are segmented at the same scale to establish a one-to-one spatial pairing relationship, providing a unified spatial coordinate benchmark for subsequent cross-modal feature matching. The feature extraction layer uses a deep residual bidirectional multi-layer feature extraction network as its core carrier, performing multi-scale convolution operations and cross-layer weighted fusion operations on the orthophotos. It outputs a high-dimensional feature map set that fuses multi-level contextual responses, providing multi-dimensional feature support for tree identification at different scales. The multi-task recognition layer employs a parallel multi-task detection branch with shared parameters, simultaneously completing three tasks: pixel-level category determination, spatial location regression, and instance mask separation. It outputs three types of structured recognition results for single-tree identification: type label, bounding box, and contour mask. The cross-modal representation layer uses the bounding boxes and contour masks output from the recognition stage as spatial constraints. It accurately crops the point clouds of individual trees from the point cloud image and extracts tree height parameters through statistical analysis of the height difference between canopy points and non-canopy points. Simultaneously, it extracts color, canopy width, and texture features from the canopy region of the orthophoto image, combining them to form a multi-dimensional health representation vector. This achieves feature alignment and deep fusion between the orthophoto visual modality and the point cloud geometric modality. The health measurement layer relies on a fully connected measurement network with a compression-expansion structure to perform multi-level association mapping on the representation vector, outputting continuous quantitative indicators representing the degree of health. The results integration layer concatenates and binds the recognition results and health diagnosis results through spatial indexing, generating integrated monitoring results that include the spatial location, type, and health degree of the trees, achieving integrated processing throughout the entire process.
[0045] like Figure 3As shown in the diagram, the bidirectional multi-layer feature extraction network of this invention illustrates the internal hierarchical structure and feature flow path of the deep residual bidirectional multi-layer feature extraction network. The network is divided into four progressive stages: primary feature extraction, deep residual encoding, channel unification mapping, and bidirectional cross-scale aggregation. The primary feature extraction stage uses an initial convolutional layer and a max-pooling layer as front-end processing units to perform spatial dimension compression on the input orthophoto, suppress background redundancy, and extract primary edge and texture responses, outputting a low-level basic feature map. The deep residual encoding stage consists of five serially arranged residual modules. Each module sequentially performs a progressive downsampling operation on the feature map, outputting five sets of deep feature maps with progressively decreasing spatial resolution and progressively increasing channel dimensions, completing the hierarchical encoding process from shallow detail features to high-level semantic features. The channel unification mapping stage uses a 1×1 convolution operation to perform channel dimension transformation on the five sets of deep feature maps, unifying the number of channels in each layer to the same dimension. This eliminates the adaptation obstacles caused by channel dimension differences in subsequent cross-layer fusion, ensuring the computational feasibility of feature fusion. The bidirectional cross-scale aggregation stage constructs bidirectional feature transfer paths from top to bottom and bottom to top. The top-down path performs bilinear upsampling on high-level semantic features and then weights them together with low-level features. The bottom-up path performs max pooling downsampling on low-level detail features and then weights them together with high-level features. The fusion weights of each layer are adaptively and dynamically adjusted through learnable logarithmic scalar parameters. This stage adds a new feature layer generated by upsampling on top of the original five feature layers, and finally outputs a set of six high-dimensional feature maps that fuse multi-scale contextual information. This effectively makes up for the deficiency of single-scale features in representing the details of small-sized trees in mining areas and improves the feature completeness and recognition recall of trees at different scales.
[0046] like Figure 4The diagram illustrates the multi-task detection architecture of this invention, showcasing the structure and functional division of a parallel multi-task detection branch with shared parameters. This branch takes a six-layer fused high-dimensional feature map as input and achieves consistent mapping of multi-scale features through a parameter-sharing mechanism, balancing detection accuracy and computational efficiency. The main body of the detection branch consists of four sequential convolutional modules and an output layer. The four convolutional modules use 3×3 convolutional kernels as core computational units, combined with nonlinear activation functions to complete deep feature encoding. All parallel detection heads share completely consistent network weight parameters, significantly reducing the number of model parameters and the risk of overfitting while enhancing the consistency of detection responses for targets at different scales. The output layer is divided into three parallel functional branches, each undertaking an independent task: instance segmentation, category classification, and bounding box regression. The segmentation branch performs pixel-by-pixel foreground and background binary classification on the input feature map, outputting an instance mask probability map of a single tree to accurately define the contour boundary of the tree canopy. The classification branch performs multi-class probability prediction of tree type on the feature map pixel by pixel, outputting a confidence score for each tree type to complete the pixel-level determination of tree type. The regression branch performs spatial offset regression on the feature map pixel by pixel, outputting the coordinates of the predicted bounding box center point and width and height parameters to determine the spatial location box of a single tree. The three branches operate synchronously in parallel, optimizing parameters with focus loss function and bounding box loss function as constraints respectively. Finally, the pixel-level results are combined to output the type label, location box, and contour mask of a single tree, realizing multi-task synchronous and accurate tree identification in complex mining scenes, providing accurate spatial localization and contour constraints for subsequent cross-modal health diagnosis.
[0047] A real-time cross-modal matching method for single-tree health diagnosis in mining areas is proposed. This method utilizes a deep residual bidirectional multi-layer feature extraction network to perform multi-scale convolution and cross-layer weighted fusion on orthophotos, effectively mitigating the problem of lost detail information due to large scale differences in trees within deep networks. This improves the edge segmentation accuracy and recognition completeness of small-sized trees. Simultaneously, a parallel multi-task detection head with shared parameters is employed to synchronously complete pixel-level classification, position regression, and mask separation, reducing the number of model parameters and enhancing generalization ability. A cross-modal matching mechanism is introduced, accurately cropping single-tree point clouds from point cloud images based on the localization boxes and contour masks output during the recognition stage, and separating canopy points from non-canopy points. Tree height is obtained through height difference statistics. Furthermore, color, canopy width, and texture features are differentially extracted from the canopy region in the orthophotos. Data from different modalities are integrated into a unified multi-dimensional health representation vector through a differential feature extraction module. A fully connected metric network with a compression-expansion structure is constructed to perform multi-level association mapping on multi-dimensional health representations. The outputs of the identification model and the metric model are concatenated to achieve integrated output of tree spatial location, type and health status.
[0048] This invention addresses the problem of high false negative rates for small targets due to insufficient feature extraction capabilities of existing deep learning networks for targets at different scales. It employs a bidirectional cross-scale weighted aggregation strategy to fully integrate high-level semantic information with low-level detail information, ensuring that the features of small trees are preserved in the deep network and participate in the final discrimination, thereby effectively reducing false negatives and false negatives. To address the difficulty of existing methods in handling cross-modal matching and fusion between visual features and morphological parameters, this invention utilizes the precise localization boxes generated during the recognition stage as spatial constraints. Height information from point cloud data and visual information from orthophotos are collaboratively processed through a differentiated feature extraction channel. The extracted multimodal features are then uniformly input into a metric network, avoiding the fragmented nature of the independent recognition and diagnosis stages in traditional methods. Finally, a cascaded structure enables end-to-end real-time processing from tree detection to health assessment, improving the integration and efficiency of tree monitoring in mining areas.
[0049] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various equivalent changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for real-time cross-modal matching for the health diagnosis of individual trees in mining areas, characterized in that, Includes the following steps: S1. Acquire UAV images covering the mining area and generate orthophotos and point cloud images. Perform image segmentation on both at the same scale to construct a tree sample dataset that corresponds one-to-one with the orthophotos and point cloud images. S2, based on a deep residual bidirectional multi-layer feature extraction network, performs multi-scale convolution and cross-layer weighted fusion operations on the input orthophoto to obtain a set of high-dimensional feature maps that fuse multi-level contextual responses; S3, using a parallel multi-task detection branch with shared parameters, performs pixel-level category determination, spatial location regression, and instance mask separation on each layer of the high-dimensional feature map set simultaneously, and outputs the type label, localization box, and contour mask of a single tree; S4. Based on the positioning box and the contour mask, crop the point cloud of a single tree from the corresponding point cloud image, perform height difference statistics on the crown point and the non-crown point to obtain the tree height, and at the same time perform differential feature mapping of color, crown width and texture on the crown layer of a single tree in the orthophoto to obtain a multidimensional health representation vector. S5, input the multidimensional health representation vector into a fully connected metric network with a compression-expansion structure, perform multi-level association mapping, and output a continuous quantitative index representing the health status of a single tree. S6, cascade the outputs of S3 and S5 to generate integrated monitoring results including the spatial location, type and health status of individual trees.
2. The method for real-time cross-modal matching health diagnosis of a single tree in a mining area according to claim 1, characterized in that, The deep residual bidirectional multilayer feature extraction network in S2 performs a weighted fusion operation on the feature maps, adjusting the fusion weights of each layer's feature maps using the following calculation formula: , ,in, Indicates the first Layer feature map for the first The fusion weight coefficients of the layer feature maps For learnable log-scalar parameters, This indicates a bilinear upsampling operation. This indicates a max-pooling downsampling operation. For the first The output feature map after bidirectional fusion of the layers. and These are the original feature maps of the upper and lower layers that participated in the fusion.
3. The method for real-time cross-modal matching health diagnosis of individual trees in mining areas according to claim 1, characterized in that, The The parallel multi-task detection branch in the model uses bounding box loss constraints for bounding box regression, and its loss function is: , ,in, For prediction boxes With real frame The intersection and union ratio, and These are the center coordinates of the predicted bounding box and the ground truth bounding box, respectively. It is a Euclidean distance metric. and These are the width and height of the prediction box, respectively. and These are the width and height of the actual bounding box, respectively. This is the diagonal length of the smallest bounding rectangle between the predicted bounding box and the ground truth bounding box. and These are the width and height of the smallest bounding rectangle, respectively.
4. The method for real-time cross-modal matching health diagnosis of individual trees in mining areas according to claim 1, characterized in that, The parallel multi-task detection branch in S3 uses a focus loss function to constrain the category determination, and its calculation formula is as follows: ,in, For the first The true category label for each pixel. The model predicts the first The probability value of each pixel belonging to the target category. To adjust the focusing parameters for the weights of easy and difficult samples, The loss value for class discrimination.
5. The method for real-time cross-modal matching health diagnosis of individual trees in mining areas according to claim 1, characterized in that, In step S4, when performing height difference statistics on canopy points and non-canopy points, the average of the top 5% of height values in the canopy point set is extracted as the canopy elevation parameter, and the average of the bottom 5% of height values in the non-canopy point set is extracted as the ground surface elevation parameter. The difference between the two is used as a measure of the height of a single tree, and its calculation method is as follows: ,in, The height of a single tree. This is the set of indices of the top 5% of tree crown points, sorted by height. The total number of points in this set. The first point in the set of tree canopy points Elevation values of each point This is the set of indexes for 5% of the non-crown point sets after height sorting. The total number of points in this set. The first point in the set of non-canopy points Elevation values of each point.
6. The method for real-time cross-modal matching health diagnosis of single trees in mining areas according to claim 1, characterized in that, The fully connected metric network in S5 uses the cross-entropy loss function for parameter optimization, and its expression is: ,in, The total number of samples used in loss calculation. For the first The true health status label of each sample The first output of the fully connected metric network The predicted probability value of the health status of each sample This represents the loss value for the health diagnosis branch.
7. The method for real-time cross-modal matching health diagnosis of single trees in mining areas according to claim 1, characterized in that, S2 includes: The input orthophoto image is compressed in space by initial convolution and pooling operations to extract primary edge and texture response maps; the primary response maps are then progressively downsampled through five residual modules to output five deep feature maps with different spatial resolutions. Channel unification mapping is performed on each deep feature map to transform the number of channels in each feature map to the same dimension. Bidirectional cross-scale weighted aggregation operation is performed on each layer of feature maps after channel unification. The top-down path upsamples the high-level features and superimposes them with the low-level features, while the bottom-up path downsamples the low-level features and superimposes them with the high-level features. Finally, a set of high-dimensional feature maps with six layers of multi-scale context is output.
8. The method for real-time cross-modal matching health diagnosis of single trees in mining areas according to claim 1, characterized in that, S3 includes: The six high-dimensional feature maps are respectively input into a parallel detection head with shared parameters. Each detection head includes four serial convolutional modules and an end convolutional layer. Each detection head's segmentation branch performs foreground and background binary classification on the feature map pixel by pixel, generating an instance mask probability map of a single tree; Each detection head's classification branch performs multi-class probability prediction of tree type on the feature map pixel by pixel, and outputs confidence scores for each category; The regression branch within each detection head performs offset regression of the center coordinates and width and height parameters of the predicted bounding box pixel by pixel on the feature map, and outputs the spatial positioning bounding box of a single tree.
9. The method for real-time cross-modal matching health diagnosis of a single tree in a mining area according to claim 1, characterized in that, S4 includes: The minimum bounding rectangle subset of the point cloud for a single tree is cropped from the point cloud image based on the positioning box; the canopy point set and the non-canopy point set are separated from the point cloud subset based on the spatial distribution of the contour mask; the canopy point set and the non-canopy point set are sorted by height and the mean of the two ends is taken as the difference to obtain the height parameter of a single tree. Color mean statistics, crown diameter calculation, and texture encoding extraction via multiple concatenated convolutional layers are performed on the canopy region defined by the contour mask in the orthophoto, and combined to form a multidimensional health representation vector.
10. The method for real-time cross-modal matching health diagnosis of a single tree in a mining area according to claim 1, characterized in that, S5 includes: The multidimensional health representation vector is input into the first hidden layer, which includes five neurons. The input vector is subjected to linear transformation and nonlinear activation to expand the feature dimension. The output of the first hidden layer is fed into the second hidden layer, which includes three neurons. The feature is subjected to dimensionality reduction compression to remove redundant correlation information. The output of the second hidden layer is fed into the third hidden layer, which consists of five neurons. The compressed features are expanded and mapped to restore the comprehensiveness of the associated expression. The output of the third hidden layer is then mapped to a preset numerical range through an activation function to output a continuous quantitative index representing the health of a single tree.