Forest region picture-based basic data generation method
By combining multimodal data acquisition and ViT deep learning model, the problems of inefficient efficiency and insufficient recognition accuracy in forest area resource management are solved, and efficient, accurate acquisition and dynamic monitoring of forest area data are achieved.
Patent Information
- Application Number
- CN202510448162.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing technology is inefficient and costly in forest area resource management, and traditional convolutional neural networks are difficult to capture subtle feature differences in complex scenarios, resulting in insufficient recognition accuracy of trees and vegetation, unable to fully and accurately characterize terrestrial properties, and lack in-depth analysis of data at different periods.
ViT-based deep learning model combined with multimodal data acquisition, and pre-processed through optical images, thermal infrared and hyperspectral images, segmentation and labeling, shape, texture and spectral features are extracted, and the autoregressive model is used to analyze the changing trend of the features over time to generate basic data in the forest area.
It significantly improves the recognition accuracy of targets such as trees and vegetation, improves the quality and efficiency of forest area data acquisition, can fully characterize the terrain attributes, and achieves large-scale dynamic monitoring through automated processing.
Smart Images

Figure CN120298900A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and specifically relates to a method for generating basic data based on forest area pictures. Background Art
[0002] In modern forest area resource management, the application of image processing and analysis technology is becoming more and more extensive, especially in monitoring vegetation cover, tree growth, and ecological changes, etc.;
[0003] However, although on-site manual surveys can provide detailed ground information, they are inefficient and costly, especially in large-scale or inaccessible forest areas. In addition, human factors may lead to inaccuracy and subjectivity of data. And existing deep learning models often have difficulty capturing subtle feature differences in images when dealing with complex scenarios. Especially in an environment as highly complex and diverse as a forest area, traditional convolutional neural networks (CNNs) have limitations in capturing long-range dependencies, resulting in insufficient recognition accuracy for targets such as trees and vegetation. This lack of accuracy directly affects subsequent feature extraction and analysis, and cannot comprehensively and accurately describe the attributes of ground objects. At the same time, when existing technologies process forest area pictures to generate basic data, they mostly focus on analyzing single or a few pictures, and pay less attention to the association of data at different times. Even when analyzing time series data, it often only simply compares the changes in the data of two consecutive periods, and cannot deeply explore the evolution laws of vegetation cover, temperature distribution, etc. over a long time span.
[0004] Therefore, those skilled in the art have proposed a method for generating basic data based on forest area pictures, aiming to comprehensively improve the quality and efficiency of obtaining forest area basic data by combining advanced deep learning technology and multi-modal data collection, and improve the efficiency and accuracy of forest area resource management. Summary of the Invention
[0005] In order to solve the above technical problems, the present invention provides a method for generating basic data based on forest area pictures to solve the problems raised in the background art.
[0006] A method for generating basic data based on forest area pictures includes the following steps:
[0007] S1. Collect multi-data source images from multiple angles above the forest area; the multi-data source images include optical images, thermal infrared, and hyperspectral image data;
[0008] S2. Preprocess the multi-data source images to obtain a preprocessed comprehensive image;
[0009] S3. Use a deep learning model based on ViT to predict and label the comprehensive image to obtain a segmented image;
[0010] S4. Extract shape-based features, texture-based features, and spectrum-based features from the segmented image to obtain feature extraction data;
[0011] S5. Arrange the feature extraction data of different periods in chronological order, analyze the change trend of forest area features over time, and obtain an analysis result;
[0012] S6. Combine the feature extraction data and the analysis result to generate basic forest area data.
[0013] Preferably, the ViT-based deep learning model predicts and labels the comprehensive image, including:
[0014] Segment the comprehensive image I into non-overlapping image patches x p , each image patch has a size of p×p pixels, and each image patch is converted into an embedded vector of a fixed dimension through linear projection
[0015]
[0016] where, W p is the linear projection matrix, and M is the pixel value matrix of the image patch x p ;
[0017] Add positional encoding PE to retain the position information of the image patch in the image to obtain the initial input z0:
[0018]
[0019] Input z0 into the Transformer encoder, and the final output of the Transformer encoder is z L , where L is the total number of layers of the Transformer encoder;
[0020] The final output z L is predicted through a linear classification head, and the prediction result is:
[0021] y = Softmax(W c z L + b)
[0022] where, W c is the weight matrix of the linear classification head, b is the bias, y is the probability distribution vector, and each element corresponds to the probability that each pixel in the image belongs to different classes;
[0023] By performing threshold processing on the probability distribution, each pixel is assigned to the class with the highest probability to obtain the segmented image:
[0024] yi,k = max(y i,1 ,..., y i,C )
[0025] where k is the category labeled for pixel i, and C is the total number of categories.
[0026] Preferably, the Transformer encoder consists of multiple multi - head attention layers and a multi - layer perceptron. The multi - head attention layer includes:
[0027] The input z l-1 is linearly projected into query, key, and value matrices:
[0028] Q = W q z l-1
[0029] K = W k z l-1
[0030] V = W v z l-1
[0031] where l represents the layer number, Q, K, and V are the query, key, and value matrices respectively, and W q , W k , W v are the corresponding weight matrices;
[0032] For each head h, the attention score A h is calculated:
[0033]
[0034] where d k is the dimension of the key matrix K;
[0035] The concatenation of the attention scores of all heads is transformed through a linear layer to obtain the output of the multi - head self - attention:
[0036] MHA(z l-1 ) = Concat(A1,..., A H )W o
[0037] where MHA represents multi - head attention, H is the number of heads, and W o is the output weight matrix;
[0038] The multi - layer perceptron includes:
[0039] The output MHA(z l-1 ) of the multi - head self - attention is input into the multi - layer perceptron after residual connection and layer normalization:
[0040] MLP(x) = GELU(W2LayerNorm(x)W1)
[0041] Among them, MLP represents a multi-layer perceptron, GELU is an activation function, LayerNorm is layer normalization, and W1 and W2 are weight matrices;
[0042] The output of the Transformer encoder is: z l = MLP(MHA(z l-1 )) + z l-1 .
[0043] Preferably, obtaining the feature extraction data includes: extracting shape-based features from the segmented image, where the shape-based features include area, perimeter, and circularity;
[0044] extracting texture-based features from the segmented image, where the texture-based features are gray-level co-occurrence matrix features, including contrast, energy, and entropy;
[0045] extracting spectrum-based features from the segmented image, where the spectrum-based features include band mean and spectral index;
[0046] Combining the shape-based features, texture-based features, and spectrum-based features to obtain the feature extraction data.
[0047] Preferably, the feature extraction data is arranged in chronological order and includes:
[0048] At each time point t, n features are extracted for forest area features, denoted as: F t = [f t1 , f t2 ,..., f tn , where t = 1, 2,..., T, and f ti represents the i-th feature value at time point t;
[0049] Arranging the feature extraction data in chronological order into a time series dataset {F1, F2,..., F T};
[0050] Obtaining the analysis result includes:
[0051] Predicting future values from historical data of the time series {F t} through an autoregressive model. The p-order autoregressive model AR(p) is expressed as: Among them is the autoregressive coefficient, and ε t is the white noise term;
[0052] By estimating the autoregressive coefficients Predict the eigenvalues of the current period based on the eigenvalues of the past p periods to obtain the change trend results of the forest area feature objects over time.
[0053] Preferably, the basic forest area data includes the statistics of the number and types of trees, the vegetation coverage area and its changes, and the temperature distribution and change trend in the forest area.
[0054] Compared with the prior art, the present invention has the following beneficial effects:
[0055] 1. The present invention performs pixel-level segmentation annotation on the comprehensive image through a deep learning model based on ViT, and uses the global attention mechanism of Transformer to capture the subtle feature differences in complex scenes, significantly improving the recognition accuracy of targets such as trees and vegetation. Extract three types of features, namely shape, texture, and spectrum, from the segmented image to construct a multi-dimensional feature vector, which can comprehensively describe the attributes of ground objects. And through the autoregressive model, the feature data of different periods are sorted by time to realize the quantitative analysis of dynamic trends such as vegetation coverage changes and temperature distribution evolution, providing a long-term monitoring basis for forest area resource management.
[0056] 2. The present invention significantly reduces manual intervention through full-process automated processing, and combines the fast inspection ability of drones to significantly improve the efficiency of forest area data update, which is applicable to large-scale dynamic monitoring scenarios.
[0057] 3. The present invention combines traditional optical image acquisition with thermal infrared imaging and hyperspectral imaging technologies for multi-modal data fusion acquisition, and uses a deep learning model based on ViT to capture long-range dependencies in the image, making the segmentation and annotation of different elements in complex forest area scenes more accurate. Description of the Drawings
[0058] Figure 1 It is a flowchart of the method for generating basic data based on forest area pictures of the present invention. Detailed Embodiments
[0059] The following further describes in detail the embodiments of the present invention in conjunction with the drawings and examples. The following examples are used to illustrate the present invention, but cannot be used to limit the scope of the present invention.
[0060] As shown in the Figure 1 accompanying drawings
[0061] Embodiment 1: The present invention provides a method for generating basic data based on forest area pictures, including the following steps:
[0062] S1. By mounting a camera, a thermal infrared camera, and a hyperspectral imager on a drone, collect forest area optical image, thermal infrared, and hyperspectral image data from multiple angles over the forest area according to a preset flight path.
[0063] By equipping drones with a variety of devices, multi-modal image data of forest areas can be comprehensively and efficiently obtained. Optical images provide intuitive visual information of forest areas; thermal infrared images can detect temperature differences, which is conducive to discovering pest-infected trees; hyperspectral imaging can obtain fine spectral information to assist in tree species identification and vegetation health monitoring.
[0064] S2. Preprocess the optical images, thermal infrared, and hyperspectral image data to obtain a preprocessed composite image; for the optical images, stitching and correction are performed to generate a panoramic image, expanding the observation range and correcting image distortion to improve the usability of the images; for the thermal infrared and hyperspectral images, radiometric correction and geometric correction are carried out and matched with the optical images, which can unify the spatial positions and radiometric characteristics of different modal data, enabling subsequent analysis to comprehensively utilize multi-source information and improving the accuracy and reliability of the analysis results.
[0065] S3. Use a deep learning model based on ViT to predict and label the composite image to obtain the segmented image; the deep learning model based on ViT predicts and labels the composite image, including:
[0066] Segment the composite image I into non-overlapping image patches x p , each image patch has a size of p×p pixels, and each image patch is converted into an embedded vector of a fixed dimension through linear projection
[0067]
[0068] where, W p is the linear projection matrix, and M is the pixel value matrix of the image patch x p ;
[0069] Add positional encoding PE to retain the position information of the image patch in the image to obtain the initial input z0:
[0070]
[0071] Input z0 into the Transformer encoder, and the final output of the Transformer encoder is z L , where L is the total number of layers of the Transformer encoder;
[0072] The final output z L is predicted through a linear classification head, and the prediction result is:
[0073] y = Softmax(W c z L + b)
[0074] where, W cW is the weight matrix of the linear classification head, b is the bias, and y is the probability distribution vector, where each element corresponds to the probability that each pixel in the image belongs to different classes;
[0075] By thresholding the probability distribution, each pixel is assigned to the class with the highest probability to obtain the segmented image:
[0076] y i,k = max(y i,1 ,..., y i,C )
[0077] where k is the class labeled for pixel i, and C is the total number of classes.
[0078] The Transformer encoder consists of multiple multi - head attention layers and a multi - layer perceptron. The multi - head attention layer includes:
[0079] The input z l-1 is linearly projected into query, key, and value matrices:
[0080] Q = W q z l-1
[0081] K = W k z l-1
[0082] V = W v z l-1
[0083] where l represents the layer number, Q, K, and V are the query, key, and value matrices respectively, and W q , W k , W v are the corresponding weight matrices;
[0084] For each head h, calculate the attention score A h :
[0085]
[0086] where d k is the dimension of the key matrix K;
[0087] The concatenation of the attention scores of all heads is transformed through a linear layer to obtain the output of the multi - head self - attention:
[0088] MHA(z l-1 ) = Concat(A1,..., A H )W o
[0089] where MHA represents multi - head attention, H is the number of heads, and W ois the output weight matrix;
[0090] The multi-layer perceptron includes:
[0091] After passing the output MHA(z l-1 ) of the multi-head self-attention through residual connection and layer normalization, it is input into the multi-layer perceptron:
[0092] MLP(x) = GELU(W2LayerNorm(x)W1)
[0093] where MLP represents the multi-layer perceptron, GELU is the activation function, LayerNorm is the layer normalization, and W1 and W2 are the weight matrices;
[0094] The output of the Transformer encoder is: z l = MLP(MHA(z l-1 )) + z l-1 .
[0095] Utilize the powerful feature learning ability of the ViT model to accurately segment and label the comprehensive image. Its function is to accurately distinguish various ground objects in the complex forest area scene, provide clear and accurate target areas for subsequent feature extraction, and greatly improve the pertinence and effectiveness of feature extraction.
[0096] S4. Extract shape-based features, texture-based features, and spectrum-based features from the segmented image to obtain feature extraction data; the shape-based features include area, perimeter, and circularity; the texture-based features are gray-level co-occurrence matrix features, including contrast, energy, and entropy; the spectrum-based features include band mean and spectral index; combine the shape-based features, texture-based features, and spectrum-based features to obtain feature extraction data.
[0097] Extracting shape, texture, and spectrum features from the segmented image can characterize the characteristics of forest area ground objects from different angles. Shape features reflect the geometric shapes of trees and vegetation, which are helpful for tree counting and growth status assessment; texture features reflect surface characteristics and are used for tree species classification, providing rich quantitative information for subsequent analysis, and are the key to in-depth understanding of the forest area ecosystem.
[0098] S5. Arrange the feature extraction data of different periods in chronological order, analyze the change trend of forest area ground object features over time, and obtain the analysis results; the feature extraction data arranged in chronological order includes:
[0099] At each time point t, extract n kinds of features for the forest area ground objects, denoted as: F t = [f t1 , f t2 ,..., f tn, where \(t = 1, 2, \cdots, T\), \(f\) ti represents the \(i\)-th eigenvalue at time point \(t\);
[0100] Arrange the feature extraction data in chronological order into a time series dataset \(\{F_1, F_2, \cdots, F\}\) T \(\}\);
[0101] The analysis results are obtained, including:
[0102] Predict the future value through the autoregressive model for the historical data of the time series \(\{F\}\) t \(\}\), and the \(p\)-order autoregressive model \(AR(p)\) is expressed as: where is the autoregressive coefficient, and \(\varepsilon\) t is the white noise term;
[0103] By estimating the autoregressive coefficient Predict the eigenvalue of the current period based on the eigenvalues of the past \(p\) periods, and obtain the change trend result of the forest area ground object features over time.
[0104] Sort and analyze the feature data of different periods according to time, which can reveal the change trend of the forest area ground object features over time, and help managers understand the dynamic changes of the forest area, such as the growth rate of trees, the increase and decrease of vegetation coverage, and the change trend of temperature.
[0105] S6. Combine the feature extraction data and the analysis results to generate the basic data of the forest area. The basic data of the forest area includes the statistics of the number and types of trees, the vegetation coverage area and its changes, and the temperature distribution and change trend in the forest area.
[0106] Based on the feature extraction and analysis results of the previous steps, generate basic data such as the statistics of the number and types of trees, the vegetation coverage area and its changes, and the temperature distribution and change trend in the forest area. The generated basic data of the forest area is used to evaluate the ecological status, formulate protection plans, monitor resource changes, etc., providing strong data support for the sustainable development of the forest area.
[0107] Example 2: This example is basically the same as the previous example, except that in the basic data of the forest area, the statistics of the number of trees are obtained by identifying trees through the connected region labeling algorithm in the segmented images of different periods, counting the number of trees in each period's image, and summarizing the number of trees in each period to form a sequence of the number of trees changing over time, so as to reflect the dynamic changes of the number of trees in the forest area; then use the shape features extracted in S4 to screen the connected regions, excluding the interference of small-area and irregular regions caused by image noise or mis-segmentation to the statistics of the number of trees, and improving the statistical accuracy.
[0108] The tree species statistics in the forest area basic data use the spectral and texture features extracted in S4 to construct feature vectors, and classify the trees in the images of each period through a classification model. The classification results of each period are summarized to count the quantity changes of different tree species in each period. Combining with the time series analysis results of the tree features in S5, by analyzing the changing trends of these features over time, it can further assist in the identification and confirmation of tree species and improve the accuracy of classification.
[0109] The vegetation coverage area in the forest area basic data is obtained by binarizing the images in different periods to get the binary image of the vegetation area, and the vegetation coverage area of each period is obtained. The vegetation coverage areas of each period are summarized into a time series, which can visually display the change of the vegetation coverage area over time. Using the normalized difference vegetation index (NDVI) extracted in S4, calculate the NDVI value of each pixel in the hyperspectral image data of each period, so as to obtain a more accurate binary image of the vegetation area, and then calculate the vegetation coverage area, realizing a more accurate reflection of the actual vegetation coverage, especially in areas where the distinction between vegetation and other ground objects is not obvious.
[0110] The change situation of the vegetation coverage area in the forest area basic data is based on the analysis of the vegetation coverage area time series in S5, calculate the change rate of the vegetation coverage area, and organize the change rates of each period into a report, detailing the increase and decrease of the vegetation coverage area in different time periods and the analysis of the reasons. Then use the time series model in S5 to model the vegetation coverage area time series, and use the established model to predict the vegetation coverage area in the next few periods, providing early warning for forest area management in advance.
[0111] The temperature distribution in the forest area in the forest area basic data is obtained by using the camera calibration parameters to establish the conversion relationship between the gray value and the actual temperature for the thermal infrared images in different periods, and converting each pixel of the thermal infrared image of each period to obtain the temperature distribution images of the forest area in each period. Arrange these temperature distribution images in chronological order to form a sequence of the temperature distribution in the forest area changing over time. By comparing the temperature distribution images of different periods, the dynamic changes of the temperature hot spots and the temperature anomaly change areas in the forest area can be visually observed.
[0112] The temperature change trend in the forest area in the forest area basic data is to generate a detailed temperature change trend report according to the analysis results of the temperature change in the forest area in space and time in S5; based on the temperature change trend analysis results, evaluate the impact of temperature change on the forest ecosystem. Through research, it is found that the increase in temperature may lead to changes in the suitable growth areas of some tree species, or increase the occurrence frequency of pests and diseases. Feed back these evaluation results to the forest area managers, so that the managers can adjust the afforestation plan according to the impact of temperature change on the tree species distribution, and select tree species that are more adaptable to future temperature changes for planting.
[0113] Importantly, it should be noted that the construction and arrangement of the present application shown in multiple different exemplary embodiments are merely illustrative. Although only a few embodiments are described in detail in this disclosure, those who refer to this disclosure should easily understand that many modifications are possible without substantially departing from the novel teachings and advantages of the subject matter described in this application. Other substitutions, modifications, changes, and omissions may be made in the design, operating conditions, and arrangement of the exemplary embodiments without departing from the scope of the present invention. Therefore, the present invention is not limited to specific embodiments, but extends to various modifications that still fall within the scope of the appended claims.
[0114] In addition, in order to provide a concise description of the exemplary embodiments, not all features of the actual embodiments may be described (i.e., those features that are not relevant to the currently considered best mode of implementing the present invention or those features that are not relevant to the implementation of the present invention).
[0115] It should be understood that in the development of any actual implementation, in any engineering or design project, a large number of specific implementation decisions may be made. Such development efforts may be complex and time-consuming, but for those of ordinary skill in the art who benefit from this disclosure, without excessive experimentation, such development efforts will be a routine task of design, manufacture, and production.
[0116] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A method for generating basic data based on forest area pictures, characterized in that, It includes the following steps: S1. Collect multi-source images from multiple angles over the forest area; the multi-source images include optical images, thermal infrared, and hyperspectral image data; S2. Preprocess the multi-source images to obtain a preprocessed comprehensive image; S3. Use a deep learning model based on ViT to predict and annotate the comprehensive image to obtain a segmented image; S4. Extract shape-based features, texture-based features, and spectrum-based features from the segmented image to obtain feature extraction data; S5. Arrange the feature extraction data of different periods in chronological order, analyze the change trend of forest area features over time, and obtain an analysis result; S6. Combine the feature extraction data and the analysis result to generate forest area basic data.
2. The method for generating basic data based on forest area pictures as described in claim 1, characterized in that The deep learning model based on ViT predicts and annotates the comprehensive image, including: Divide the comprehensive image I into non-overlapping image patches x p , where the size of each image patch is p×p pixels, and convert each image patch into an embedding vector of a fixed dimension through linear projection Among them, W p is a linear projection matrix, and M is the pixel value matrix of the image block x p . Adding positional encoding PE to retain the position information of the image patches in the image to obtain an initial input z0: Input z0 into the Transformer encoder, and the final output of the Transformer encoder is z L , where L is the total number of layers of the Transformer encoder; The final output z L is predicted by a linear classification head, and the prediction result is: y = Softmax(W c z L + b) Among them, W c is the weight matrix of the linear classification head, b is the bias, and y is the probability distribution vector, where each element corresponds to the probability that each pixel in the image belongs to a different class; By performing threshold processing on the probability distribution, assign each pixel to the category with the highest probability to obtain a segmented image: y i,k = max(y i,1 ,..., y i,C ) where k is the category labeled by pixel i, and C is the total number of categories.
3. The method for generating basic data based on forest area pictures according to claim 2, characterized in that: The Transformer encoder consists of multiple multi-head attention layers and a multi-layer perceptron. The multi-head attention layer includes: Input z l-1 is linearly projected into query, key, and value matrices: Q = W q z l-1 K = W k z l-1 V = W v z l-1 Among them, l represents the number of layers, and Q, K, and V are the query, key, and value matrices respectively, and W q , W k , W v are the corresponding weight matrices respectively; For each header h, compute the attention score A h : where d k is the dimension of the key matrix K; Concatenating the attention scores of all heads and performing transformation through a linear layer to obtain the output of multi-head self-attention: MHA(z l-1 ) = Concat(A1,..., A H )W o Among them, MHA represents multi-head attention, H is the number of heads, and W o is the output weight matrix; The multi-layer perceptron includes: The output MHA(z of the multi-head self-attention l-1 ) After residual connection and layer normalization, it is input into the multi-layer perceptron: MLP(x) = GELU(W2LayerNorm(x)W1) where MLP represents the multi-layer perceptron, GELU is the activation function, LayerNorm is the layer normalization, and W1 and W2 are weight matrices; The output of the Transformer encoder is: z l = MLP(MHA(z l-1 )) + z l-1 .
4. The method for generating basic data based on forest area pictures according to claim 1, wherein, The obtaining of the feature extraction data includes: Extract shape-based features from the segmented image. The shape-based features include area, perimeter, and circularity; Extract texture-based features from the segmented image. The texture-based features are gray-level co-occurrence matrix features, including contrast, energy, and entropy; Extract spectrum-based features from the segmented image. The spectrum-based features include band mean and spectral index; Combine the shape-based features, texture-based features, and spectrum-based features to obtain feature extraction data.
5. The method for generating basic data based on forest area pictures according to claim 1, characterized in that The arrangement of the feature extraction data in chronological order includes: At each time point t, n features are extracted for forest area features, denoted as: F t = [f t1 , f t2 ,..., f tn , where t = 1, 2,..., T, and f ti represents the i-th feature value at time point t; Arrange the feature extraction data in chronological order to form a time series data set {F1, F2,..., F T}; The obtaining of the analysis result includes: Predict future values of the time series {F t} using an autoregressive model. The p-order autoregressive model AR(p) is expressed as: where are autoregressive coefficients, and ε t is a white noise term; By estimating the autoregressive coefficients Predict the eigenvalue of the current period based on the eigenvalues of the past p periods to obtain the result of the change trend of the forest area feature over time.
6. The method for generating basic data based on forest area pictures according to claim 1, characterized in that: The forest area basic data includes the statistics of the number and types of trees, the vegetation coverage area and its changes, and the temperature distribution and change trend in the forest area.