A tree species recognition method based on bag-of-visual-words representation and state-space modeling

Through visual bag of word representation and state space modeling methods, combined with scale invariant feature transformation and accelerated robust features, the problems of limited feature extraction capabilities and imbalance in tree species recognition are solved, and high-precision tree species recognition is achieved.

CN120431576BActive Publication Date: 2025-08-26SHANDONG JIANZHU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510933698.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-08-26
Estimated Expiration
2045-07-08

AI Technical Summary

Technical Problem

The prior art has problems of limited feature extraction capabilities and category imbalance in tree species recognition, which is difficult to meet the needs of high-precision and large-scale automation.

Method used

Visual bag of word representation and state space modeling methods are adopted, combined with scale invariant feature transformation and accelerated robust feature extraction, visual bag of word model is constructed, and global features are mined through state space model, and combined with the synthesis of a few categories of oversampling technology to improve the discriminatory ability of support vector machine classifiers.

Benefits of technology

It enhances the global feature modeling ability of tree species recognition, strengthens the ability to extract and express local feature, effectively alleviates the problem of category imbalance, and improves the classification performance of the model and the recognition performance of a few categories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431576B_ABST
    Figure CN120431576B_ABST
Patent Text Reader

Abstract

The present invention discloses a tree species identification method based on bag-of-visual-words representation and state-space modeling, which relates to the technical field of tree species cross-section microscopic image recognition and is characterized in that it includes the following steps: S1: data set construction; S2: local feature extraction using a bag-of-visual-words model; S3: global feature extraction using a state-space model; S4: multi-level feature fusion; S5: data oversampling and classifier training; S6: tree species identification. The technical problem to be solved by the present invention is to provide a tree species identification method based on bag-of-visual-words representation and state-space modeling, which introduces a state-space model to model the feature sequence of tree species images, mines their long-term dependencies, and extracts global feature information. By fusing local and global features, a multi-level feature representation with strong discriminative ability is constructed. The synthetic minority class oversampling technology is used to balance the samples to improve the discriminative ability of the support vector machine classifier.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of tree species cross-section microscopic image recognition, and in particular to a tree species recognition method based on visual word bag representation and state space modeling. Background Art

[0002] In forestry resource management and scientific research, accurate tree species identification is a key technical link in achieving efficient resource allocation, protecting endangered species, monitoring ecological changes, and combating illegal logging. Traditional tree species classification methods rely primarily on manual experience, with forestry experts identifying wood species based on macroscopic characteristics (such as color, grain direction, and gloss) and microstructural features (such as the arrangement of parenchyma tissue, the distribution of vessels, and cell wall structure). However, these methods are complex and inefficient, relying heavily on the accumulated knowledge and long-term experience of experts. Judgment criteria vary among experts, and they are highly subjective, making them difficult to meet the current forestry application needs for large-scale, automated, and high-precision tree species identification.

[0003] With the rapid development of artificial intelligence (AI), deep learning has been widely applied to image recognition tasks. Models based on convolutional neural networks (CNNs) and visual transformer architectures, in particular, have achieved remarkable results in natural image classification and object recognition. However, applying these models to the microstructural classification of tree cross-section images still faces numerous challenges, hindering the practical application of high-precision tree species identification. These challenges are reflected in the following two aspects:

[0004] (1) Limited feature extraction capabilities: Most deep learning methods focus on optimizing the structures of CNN and Transformer models. However, CNN is inherently inadequate in capturing long-range dependencies in images. While Transformer has strong global modeling capabilities, it is limited by its quadratic computational complexity and cannot be directly applied to high-resolution images. Furthermore, tree cross-section images themselves contain rich fine-grained structures and multi-scale semantic information, making it difficult for existing models to fully extract key features, which in turn limits classification performance.

[0005] (2) The problem of class imbalance is prominent: In actual forestry image data, the number of samples of different tree species is highly unevenly distributed, especially for some rare tree species or tree species in specific ecological environments, whose image samples are very limited. This significant data imbalance will make the model more inclined to the dominant class during training, resulting in insufficient learning of the features of the minority class, resulting in low recognition accuracy and high misclassification rate. In addition, deep learning models rely on large-scale balanced data for effective training. When there are too few minority class samples in the training set, the model finds it difficult to capture its discriminative features, which ultimately affects the overall classification performance and generalization ability, limiting the reliability and adaptability of the model in practical applications. Summary of the Invention

[0006] The technical problem addressed by this invention is to provide a tree species recognition method based on bag-of-visual-words representation and state-space modeling. This method extracts local region-of-interest features from an image using the Scale-Invariant Feature Transform (SIFT) and Speeded-Up Robust Features (SURF). A Bag-of-Visual-Words (BoVW) model is then constructed to capture local information within the image. Furthermore, a State-Space Model (SSM) is introduced to model the feature sequences of tree species images, mining their long-term dependencies and extracting global feature information. By fusing local and global features, a multi-level feature representation with strong discriminative power is constructed. To address the class imbalance in the training data, this invention employs the Synthetic Minority Over-sampling Technique (SMOTE) to balance the samples and enhance the discriminative power of the Support Vector Machine (SVM) classifier.

[0007] The present invention adopts the following technical solutions to achieve the invention objectives:

[0008] A tree species identification method based on bag-of-visual-words representation and state-space modeling, characterized by comprising the following steps:

[0009] S1: Dataset construction;

[0010] S2: Extract local features using the bag-of-visual-words model;

[0011] S3: Extract global features using state-space model;

[0012] S4: multi-level feature fusion;

[0013] S5: Data oversampling and classifier training;

[0014] S6: tree species identification;

[0015] S21: Scale-invariant feature transformation key point extraction and description;

[0016] S211: Generate scale space using Gaussian blur;

[0017] S212: using the Gaussian difference image to locate the key point position;

[0018] S213: Generate feature descriptors;

[0019] S22: Accelerated robust feature key point extraction and description;

[0020] S221: Hessian matrix calculation;

[0021] S222: key point detection;

[0022] S223: Generate feature descriptors; calculate Haar wavelet response statistics in each sub-region:

[0023] S23: K-means clustering algorithm to build visual dictionary;

[0024] S231: Collect feature descriptors of all images and construct the full feature set of the training set:

[0025] S232: K-means clustering generates a visual dictionary. K-means clustering is used to cluster the entire set to obtain a visual dictionary:

[0026] S233: Construct a visual word bag for each image. , assign all its joint descriptors to the nearest cluster center, build a bag-of-words histogram, weight the bag-of-words histogram, and finally obtain the bag-of-words vector representation of the image, that is, the local feature vector;

[0027] S31: Image segmentation and linear embedding;

[0028] S32: construct state space feature extraction block;

[0029] S321: Construct multi-directional sequence input;

[0030] S322: Unidirectional state space modeling;

[0031] S323: Four-directional output fusion;

[0032] S33: Building a state-space recognition network;

[0033] S331: stack multiple state space feature extraction blocks, and the output of each feature extraction block serves as the input of the next layer;

[0034] S332: Pooling obtains the global features of the image and average pooling is performed on the final stacked output sequence;

[0035] S333: Supervised training of classifiers.

[0036] As a further limitation of this technical solution, the specific steps of S1 are:

[0037] S11: Image collection, using a high-resolution microscope system to collect cross-section images of tree species and convert them into PNG format;

[0038] S12: image preprocessing;

[0039] S121: Image denoising;

[0040] S122: size normalization;

[0041] S13: Dataset division.

[0042] As a further limitation of this technical solution, the specific steps of S2 are:

[0043] S211: Use Gaussian blur to generate scale space, the image is , its Gaussian blur is expressed as:

[0044] (1);

[0045] in: is the standard deviation of the Gaussian function; and Respectively represent the pixel coordinates in the horizontal and vertical directions in the original spatial coordinate system of the image;

[0046] The generated scale space is obtained by Gaussian blurring images at different scales:

[0047] (2);

[0048] in: Represents a two-dimensional convolution operation;

[0049] S212: Use the Gaussian difference image to locate the key point position, which is expressed as follows:

[0050] (3);

[0051] S213: Generate feature descriptors, key points in image coordinates , first calculate the gradient information of the key point:

[0052] (4);

[0053] The direction of the gradient and amplitude They are:

[0054] (5);

[0055] S221: Hessian matrix calculation, expressed as:

[0056] (6);

[0057] in: For the original image at position Gray value at ; is the value of the integral image, indicating that arrive The sum of all pixels within the rectangle, the Hessian matrix describes the local curvature of the image, and the specific form is:

[0058] (7);

[0059] in: Indicates that the image is Second derivative of direction;

[0060] Indicates that the image is Second derivative of direction;

[0061] represents mixed second-order derivatives;

[0062] It means that the result calculated by formula (6) is brought into the scale space obtained by formula (2);

[0063] Approximate the Hessian matrix using a box filter with an integral image , and then calculate the determinant of the approximate Hessian matrix, which is expressed as:

[0064] (8);

[0065] in: , and is an approximation of box filtering;

[0066] S222: Key point detection uses interpolation fitting to accurately locate the coordinates and scales of key points, and proposes edge response points based on the following judgment conditions:

[0067] (9);

[0068] in: is the trace of the matrix, i.e. the sum of the elements on the main diagonal; is the threshold;

[0069] In the neighborhood of the key segment, calculate the Haar wavelet Horizontal response at and vertical response :

[0070] (10);

[0071] in: Indicates the position of a point in the neighborhood relative to the center of the key point, and Represent the horizontal and vertical coordinates in the local coordinate system respectively;

[0072] The Gaussian weighting function is expressed as:

[0073] (11);

[0074] in: and Respectively represent the horizontal and vertical coordinates of the center position in the local neighborhood of the key point;

[0075] Calculate the sum of the response amplitudes in each direction sector:

[0076] (12);

[0077] in: is the starting angle of the 6 sectors; Indicates the response direction;

[0078] The main direction of the key point is expressed as:

[0079] (13);

[0080] Among them: argmax represents the value of the independent variable corresponding to the maximum value of a function;

[0081] S223: Generate feature descriptors; calculate Haar wavelet response statistics in each sub-region:

[0082] (14);

[0083] in: ; ;

[0084] S231: Collect feature descriptors of all images and construct the full feature set of the training set:

[0085] (15);

[0086] in: is the total number of images; Indicates the total number of key points extracted from the image; Indicates the Joint descriptor of feature points; express The first in the collection Each image Descriptors of feature points;

[0087] S232: K-means clustering generates a visual dictionary. K-means clustering is used to cluster the entire set to obtain a visual dictionary:

[0088] (16);

[0089] in: Represents the cluster center set; R represents the symbolic interpretation of the real number set; Representation Descriptor The corresponding cluster number; Indicates the number of sight words; is the L2 norm;

[0090] S233: Construct visual word bag for each image , assign all its joint descriptors to the nearest cluster center, build the bag-of-words histogram, image Middle The frequency of visual words is expressed as:

[0091] (17);

[0092] in: Representing an image The number of descriptors in ; is the characteristic function;

[0093] Weight the bag-of-words histogram, and the weighted bag-of-words histogram value Expressed as:

[0094] (18);

[0095] in: Indicates how many images the visual word appears in ;

[0096] Finally, the bag-of-words vector representation of the image is obtained, that is, the local feature vector, which is expressed as:

[0097] (19).

[0098] As a further limitation of this technical solution, the specific steps of S3 are:

[0099] S31: Image segmentation and linear embedding. The input tree species image is divided into multiple non-overlapping image blocks of fixed size. Each image block is flattened and mapped to a uniform dimension through a linear layer to form a preliminary representation of the image:

[0100] (20);

[0101] in: Indicates the The embedding vector of the image patch; represents the dimension of the embedding space; Indicates the number of image blocks, For each image block size, and Represent the length and width of the image respectively;

[0102] S32: construct state space feature extraction block;

[0103] S321: Construct multi-directional sequence input. To capture different structural direction information in the image, the image blocks are serialized in four directions to form input sequences in different directions:

[0104] (twenty one);

[0105] in: ;

[0106] S322: One-way state space modeling. The formula of the state space model is expressed as:

[0107] (twenty two);

[0108] in: Indicates the input vectors; Indicates direction Previous The hidden state at a moment; The feature representation of the output; represents the state transition matrix; Represents the transformation matrix of the input; represents the mapping matrix from hidden state to output, A matrix representing a direct mapping of input to output;

[0109] S323: Four-directional output fusion: weighted fusion of the outputs from the four directions to generate the final output sequence of the state space model, expressed as:

[0110] (twenty three);

[0111] in: Represents the weighting coefficient for each direction;

[0112] S33: Building a state-space recognition network;

[0113] S331: stack multiple state space feature extraction blocks, and the output of each feature extraction block is used as the input of the next layer. The output of a feature extraction block is expressed as:

[0114] (twenty four);

[0115] in: Indicates the number of feature extraction blocks, represents the linear embedding of the input layer image patch, Indicates the The operation of the state space feature extraction block, each block contains an independent state parameter matrix ;

[0116] S332: Pooling obtains the global features of the image, and average pooling is performed on the final stacked output sequence, which is expressed as:

[0117] (25);

[0118] Among them: AVGPool represents the average pooling operation;

[0119] Next, the input is sent to the fully connected layer for feature dimension compression and semantic embedding to obtain the global features of the image:

[0120] (26);

[0121] in: Represents the activation function operation; and Represent the weight and bias of the fully connected layer respectively;

[0122] S333: Supervised training of the classifier, inputting the classification layer consisting of a 1×1 convolutional layer to predict the category score:

[0123] (27);

[0124] in: and Represent the weight and bias of the classification layer respectively; represents the logit score of each category;

[0125] Next, use the normalized exponential function to get the class probability:

[0126] (28);

[0127] in: Indicates the number of categories;

[0128] Using cross entropy loss To train:

[0129] (29);

[0130] in: Represents the tree species category labels annotated in the dataset.

[0131] As a further limitation of this technical solution, the specific steps of S4 are:

[0132] S41: Local and global feature normalization. Formula (19) obtains the local feature vector, and formula (26) obtains the global feature vector. First, the two-level feature vectors are normalized. The normalized local feature vector and global feature vector are expressed as:

[0133] (30);

[0134] S42: Local and global feature fusion, using linear weighting for feature fusion, the fused feature is expressed as:

[0135] (31);

[0136] in: represents the weight parameter;

[0137] S43: Fusion feature standardization, use L2 norm to standardize the fused features to obtain multi-level fusion features:

[0138] (32).

[0139] As a further limitation of this technical solution, the specific steps of S5 are:

[0140] S51: Identify minority class samples in the training set and classify them according to the image , select all minority class samples to form a set of feature vectors belonging to all minority classes:

[0141] (33);

[0142] in: The label identified as the minority class; | represents the conditional restriction symbol;

[0143] S52: Generate new samples using minority class oversampling technology;

[0144] S521: Calculate the nearest neighbor for each , find its Nearest neighbor samples:

[0145] (34);

[0146] in: Represents the fusion feature vector of any two samples; express The vector Dimensional component; Represents a vector No. Dimensional component;

[0147] S522: Generate a new sample and randomly select a sample from the nearest neighbor , generate new samples by linear interpolation:

[0148] (35);

[0149] in: Represents the current minority class sample; Indicates its Any neighbor among the nearest neighbors; represents the generated new sample, Represents a random number uniformly distributed between 0 and 1;

[0150] S53: Multi-class linear support vector machine classifier training;

[0151] S531: Input data, the expanded training data set obtained through S522, expressed as:

[0152] (36);

[0153] in: Indicates the first The fusion feature vector of samples, represents the true category label of the tree species image, represents the total number of training samples;

[0154] S532: Classification strategy: For each category, build a binary classifier, take the samples belonging to this category as positive class, and the samples belonging to all other categories as negative class. samples in the category c The label value on is represented as:

[0155] (37);

[0156] S533: Define the discriminant function. For each binary classification, train a linear support vector machine classifier. The output of the corresponding linear discriminant function is expressed as:

[0157] (38);

[0158] in: and Represent the weight and bias of the classifier respectively;

[0159] S534: Minimize the soft margin loss function, The optimization goal corresponding to the class is:

[0160] (39);

[0161] in: represents the regularization parameter; Indicates the The slack variables of samples are constrained as follows:

[0162] (40);

[0163] S54: determining the optimal parameters of the classifier;

[0164] S541: Construct multiple candidate hyperparameter sets and set a set of candidate regularization hyperparameters to form the set:

[0165] (41);

[0166] in: It is regularization hyperparameters to be evaluated;

[0167] S542: Train multiple support vector machine classifiers for each hyperparameter ,train A binary classifier is used to obtain the classification parameters:

[0168] (42);

[0169] Each classifier is trained using formula (39) of S534 as the optimization target;

[0170] S543: Evaluate the performance on the validation set. Make the classifier for each set of parameters predict on the validation set and calculate the accuracy:

[0171] (43);

[0172] in: Indicates the number of images in the validation set; Indicates the The predicted category of samples under hyperparameters; represents the indicator function;

[0173] S544: Select the optimal hyperparameters to determine the classifier parameters and find the hyperparameters that optimize the performance of the validation set. :

[0174] (44);

[0175] Finally determine the corresponding optimal support vector machine classifier parameters:

[0176] (45);

[0177] in: Represents the index that makes the parameter reach the optimal value, Indicates that the index is Parameters when , that is, satisfy ; and Respectively represent the index Time The weights and biases corresponding to the classes; and They represent the weight vector and bias of the corresponding optimal support vector machine classifier respectively.

[0178] As a further limitation of this technical solution, the specific steps of S6 are:

[0179] S61: For the tree species image to be identified, perform image denoising and size normalization preprocessing using the S12 alignment;

[0180] S62: extracting local features from the preprocessed image using the bag-of-visual-words model of S2, and extracting global features using the state-space model of S3;

[0181] S63: Combining the local and global features obtained in S62, and using S4 to obtain fused multi-level features;

[0182] S64: Input the multi-level features of S63 into the support vector machine classifier trained in S54, and calculate the discriminant function of each class respectively:

[0183] (46);

[0184] S65: Select the category with the maximum discriminant score as the final prediction result:

[0185] (47);

[0186] The output predicted category is the recognition result of the tree species image in actual application.

[0187] Compared with the prior art, the advantages and positive effects of the present invention are:

[0188] 1. This invention enhances global feature modeling capabilities. Using SSM to model the feature sequences of tree species images, it captures the structural information in cross-section images, effectively extracts global features, and improves the ability to distinguish morphologically similar tree species.

[0189] 2. This invention enhances local feature extraction and expression capabilities. By combining SIFT and SURF algorithms to extract key points and construct a BoVW model, it fully exploits the texture details and structural features in the image and adapts to the local feature changes of tree species at different scales and deformation conditions.

[0190] 3. This invention achieves effective fusion of local and global features. Through feature normalization and weighted fusion strategies, local and global information are uniformly encoded into multi-level discriminative features, significantly improving the model's ability to express complex images and classification performance.

[0191] 4. This invention improves the recognition performance of minority tree species. The SMOTE technique is used to oversample minority samples in the training data, effectively alleviating the training bias caused by class imbalance and enhancing the SVM classifier's ability to recognize minority tree species. BRIEF DESCRIPTION OF THE DRAWINGS

[0192] Figure 1 This is a flow chart of tree species identification according to the present invention.

[0193] Figure 2 This is a classification diagram of 75 tree species divided according to the principles of biological taxonomy in this invention.

[0194] Figure 3 This is the visual word bag model diagram of the present invention.

[0195] Figure 4 This is the state space identification network structure diagram of the present invention. DETAILED DESCRIPTION

[0196] A specific embodiment of the present invention is described in detail below with reference to the accompanying drawings, but it should be understood that the protection scope of the present invention is not limited by the specific embodiment.

[0197] The present invention comprises the following steps:

[0198] S1: Dataset Construction: First, cross-section microscopic images of tree species were collected using a high-resolution microscope. These images were denoised, grayscaled, and size-normalized. Subsequently, a category label was assigned to each tree species image based on the angiosperm phylogenetic classification system, and the dataset was divided into training, validation, and test sets.

[0199] The specific steps of S1 are:

[0200] S11: Image collection: Use a high-resolution microscope system to collect cross-section images of tree species and convert them into PNG format; assign a category label to each image based on the angiosperm phylogenetic classification system. This invention includes 75 most common tree species categories, and their corresponding names are as follows: Figure 2 shown.

[0201] S12: image preprocessing;

[0202] S121: Image denoising: Use the non-local means filtering algorithm to denoise the image, setting the search window size to 15×15 pixels and the matching window size to 5×5 pixels to improve the denoising effect and computational efficiency while preserving image details.

[0203] S122: Size normalization: Use bilinear interpolation to normalize the denoised image and adjust the image resolution to 224×224 pixels to facilitate subsequent feature extraction and model input;

[0204] S13: Dataset Partitioning. The dataset was divided into 70% training, 20% validation, and 10% test sets. A stratified random sampling strategy was used to ensure a balanced distribution of tree species across the subsets, preventing any disruption to model training and performance evaluation caused by imbalanced tree species.

[0205] S2: Extract local features using the bag-of-visual-words model; perform keypoint detection on the preprocessed image using the scale-invariant feature transform and the accelerated robust feature algorithm, and generate feature descriptors for each keypoint. A visual vocabulary library is constructed using the K-means clustering algorithm, and the feature descriptors are mapped to the visual vocabulary to obtain a bag-of-words representation of the image and generate a local feature vector.

[0206] The specific steps of S2 are:

[0207] S21: Scale-invariant feature transformation key point extraction and description;

[0208] S211: Generate scale space using Gaussian blur. In order to detect feature points at different scales, SIFT generates scale space by performing Gaussian blur on the image at different scales. Assume that the image is , its Gaussian blur can be expressed as:

[0209] (1);

[0210] in: is the standard deviation of the Gaussian function; and Respectively represent the pixel coordinates in the horizontal and vertical directions in the original spatial coordinate system of the image;

[0211] The generated scale space is obtained by Gaussian blurring images at different scales:

[0212] (2);

[0213] in: Represents a two-dimensional convolution operation;

[0214] S212: Use the Gaussian difference image to locate the key point position. In order to find the local extreme points in the image, SIFT calculates the Gaussian difference image, which can be expressed as follows:

[0215] (3);

[0216] SIFT uses local extrema detection to find potential keypoints in a Difference of Gaussian image. This process compares the current pixel with its neighbors (including those in the upper and lower layers and the same layer) and selects local extrema as candidate keypoints. By accurately locating keypoint locations, SIFT removes unstable and low-contrast keypoints.

[0217] S213: Generate feature descriptors. For each key point, SIFT calculates the descriptor of its local area. Assuming that the key point is at the image coordinate , first calculate the gradient information of the point:

[0218] (4);

[0219] The direction of the gradient and amplitude They are:

[0220] (5);

[0221] Within the neighborhood of the key point, the image is divided into 16 4×4 local regions, and the histogram of the gradient direction is calculated for each region (each histogram has 8 directions, representing a 128-dimensional feature vector). The gradient direction histograms of each small region are merged to form the 128-dimensional feature descriptor of SIFT.

[0222] S22: Accelerated robust feature key point extraction and description;

[0223] S221: Hessian matrix calculation. In order to quickly calculate the box filter response at different scales, the input image is first converted into an integral image, which can be expressed as:

[0224] (6);

[0225] in: For the original image at position Gray value at ; is the value of the integral image, indicating that arrive The sum of all pixels within the rectangle, the Hessian matrix describes the local curvature of the image, and the specific form is:

[0226] (7);

[0227] in: Indicates that the image is Second derivative of direction;

[0228] Indicates that the image is Second derivative of direction;

[0229] represents mixed second-order derivatives;

[0230] It means that the result calculated by formula (6) is brought into the scale space obtained by formula (2);

[0231] Approximate the Hessian matrix using a box filter with an integral image , and then calculate the determinant of the approximate Hessian matrix, which can be expressed as:

[0232] (8);

[0233] in: , and is an approximation of box filtering;

[0234] S223: Key point detection. In scale space, use formula (8) to calculate the approximate determinant of each point, perform 3D non-maximum suppression, and select local maximum points as candidate key points. Use interpolation fitting to accurately locate the coordinates and scales of the key points, and propose edge response points based on the following judgment conditions:

[0235] (9);

[0236] in: is the trace of the matrix, i.e. the sum of the elements on the main diagonal; As the threshold, the present invention sets ;

[0237] In the neighborhood of the key segment (radius ), calculate the Haar Wavelet in Horizontal response at and vertical response :

[0238] (10);

[0239] in: Indicates the position of a point in the neighborhood relative to the center of the key point, and Represent the horizontal and vertical coordinates in the local coordinate system respectively;

[0240] The Gaussian weighting function can be expressed as:

[0241] (11);

[0242] in: and Respectively represent the horizontal and vertical coordinates of the center position in the local neighborhood of the key point;

[0243] Calculate the sum of the response amplitudes in each direction sector:

[0244] (12);

[0245] in: is the starting angle of the 6 sectors; Indicates the response direction;

[0246] The main direction of the key point can be expressed as:

[0247] (13);

[0248] Among them: argmax represents the value of the independent variable corresponding to the maximum value of a function;

[0249] S224: Generate feature descriptors; rotate and align the keypoint neighborhood (size 20×20 pixels) around the main direction. Divide the neighborhood into 4×4 subregions, each 5×5 pixels. Calculate the Haar wavelet response statistics in each subregion:

[0250] (14);

[0251] in: ; ; Concatenate the 4-dimensional response statistics of the 16 sub-regions to form a 64-dimensional feature descriptor of SURF.

[0252] S23: K-means clustering algorithm to build visual dictionary;

[0253] S231: Collect feature descriptors of all images, extract SIFT and SURF descriptors of key points of each image, and concatenate them to form a joint descriptor. Collect joint descriptors from all training images to construct the full feature set of the training set:

[0254] (15);

[0255] in: is the total number of images; Indicates the total number of key points extracted from the image; Indicates the The joint descriptor of feature points has a dimension of 192; express The first in the collection The first Descriptors of feature points;

[0256] S232: K-means clustering generates a visual dictionary. K-means clustering is used to cluster the entire set to obtain a visual dictionary:

[0257] (16);

[0258] in: Represents the cluster center set; R represents the symbolic interpretation of the real number set; Representation Descriptor The corresponding cluster number; Indicates the number of sight words; is the L2 norm, which measures the distance between the descriptor and the cluster center;

[0259] S233: Construct visual word bag for each image , assign all its joint descriptors to the nearest cluster center, build the bag-of-words histogram, image Middle The frequency of visual words can be expressed as:

[0260] (17);

[0261] in: Representing an image The number of descriptors in ; is an indicator function, which is 1 if the condition is true, otherwise 0;

[0262] The bag-of-words histogram is weighted, and the weighted bag-of-words histogram value It can be expressed as:

[0263] (18);

[0264] in: Indicates how many images the visual word appears in ;

[0265] The final word bag vector representation of the image, that is, the local feature vector, can be expressed as:

[0266] (19).

[0267] The present invention sets .

[0268] S3: Extract global features using the state-space model; construct basic feature extraction blocks using the state-space model, and build a state-space recognition network by cascading several feature extraction blocks. Use the training set to optimize the network parameters to obtain the global feature vector of the tree species image.

[0269] The specific steps of S3 are:

[0270] S31: Image segmentation and linear embedding. The input tree species image is divided into multiple non-overlapping image blocks of fixed size. Each image block is flattened and mapped to a uniform dimension through a linear layer to form a preliminary representation of the image:

[0271] (20);

[0272] in: Indicates the The embedding vector of the image patch; represents the dimension of the embedding space; Indicates the number of image blocks, For each image block size, and Represent the length and width of the image respectively;

[0273] S32: construct state space feature extraction block;

[0274] S321: Construct multi-directional sequence input, in order to capture the different structural direction information in the image, according to the four directions (horizontal ,vertical , main diagonal , secondary diagonal ) Serialize the image blocks to form input sequences in different directions:

[0275] (twenty one);

[0276] in: ;

[0277] S322: Single-direction state space modeling: For the sequence data in each direction, a one-dimensional state space model is used to model the effect of each input block on the current state. The formula of the state space model can be expressed as:

[0278] (twenty two);

[0279] in: Indicates the input vectors; Indicates direction Previous The hidden state at a moment; The feature representation of the output; represents the state transition matrix; Represents the transformation matrix of the input; represents the mapping matrix from hidden state to output, A matrix representing a direct mapping of input to output;

[0280] S323: Four-directional output fusion: weighted fusion of the outputs in the four directions (horizontal, vertical, main diagonal, and sub-diagonal) to generate the final output sequence of the state space model, which can be expressed as:

[0281] (twenty three);

[0282] in: Represents the weighted coefficient in each direction, satisfying ;

[0283] S33: Building a state-space recognition network;

[0284] S331: Stack multiple state space feature extraction blocks. In order to improve the model's expressive power and learning level, a recognition network containing state space blocks is built. The output of each feature extraction block is used as the input of the next layer. The output of a feature extraction block is expressed as:

[0285] (twenty four);

[0286] in: Indicates the number of feature extraction blocks, represents the linear embedding of the input layer image patch, Indicates the The operation of the state space feature extraction block, each block contains an independent state parameter matrix ;

[0287] S332: Pooling obtains the global features of the image, and average pooling is performed on the final stacked output sequence, which can be expressed as:

[0288] (25);

[0289] Among them: AVGPool represents the average pooling operation;

[0290] Next, the input is sent to the fully connected layer for feature dimension compression and semantic embedding to obtain the global features of the image:

[0291] (26);

[0292] in: Represents the activation function operation; and Represent the weight and bias of the fully connected layer respectively;

[0293] S333: Supervised training of the classifier, inputting the classification layer consisting of a 1×1 convolutional layer to predict the category score:

[0294] (27);

[0295] in: and Represent the weight and bias of the classification layer respectively; represents the logit score of each category;

[0296] Next, use the normalized exponential function (Softmax function) to obtain the category probability:

[0297] (28);

[0298] in: Indicates the number of categories, i.e., 75 categories of the tree species of the present invention;

[0299] Using cross entropy loss To train:

[0300] (29);

[0301] in: Represents the tree species category labels annotated in the dataset. Through the back-propagation algorithm, the parameters of each state-space feature extraction block, fully connected layer, and classification layer are jointly optimized.

[0302] S4: Multi-level feature fusion; local feature vectors and global feature vectors are standardized separately, and then linear weighted feature fusion is used to form a multi-level feature vector.

[0303] The specific steps of S4 are:

[0304] S41: Local and global feature normalization. Formula (19) obtains the local feature vector, and formula (26) obtains the global feature vector. The present invention sets the dimensions of the local and global feature vectors to 256. First, the two-level feature vectors are normalized. The normalized local feature vector and global feature vector are expressed as:

[0305] (30);

[0306] S42: Local and global feature fusion, using linear weighting for feature fusion, the fused feature is expressed as:

[0307] (31);

[0308] in: Represents the weight parameter, which is used to balance local and global features;

[0309] S43: Fusion feature standardization, use L2 norm to standardize the fused features to obtain multi-level fusion features:

[0310] (32).

[0311] S5: Data oversampling and classifier training; To address the class imbalance problem in the training set, a synthetic minority oversampling technique algorithm is used to oversample minority class samples. A support vector machine is selected as the classifier, and the expanded training set is used to extract the fusion features of the tree species images. This feature is used as the input of the classifier to train the decision boundary of the classifier.

[0312] The specific steps of S5 are:

[0313] S51: Identify minority class samples in the training set and classify them according to the image , select all minority class samples to form a set of feature vectors belonging to all minority classes:

[0314] (33);

[0315] in: The label identified as the minority class; | represents the conditional restriction symbol;

[0316] S52: Generate new samples using minority class oversampling technology;

[0317] S521: Calculate the nearest neighbor for each , find its nearest neighbor samples (the present invention sets ):

[0318] (34);

[0319] in: Represents the fusion feature vector of any two samples; express The vector Dimensional component; Represents a vector No. Dimensional component;

[0320] S522: Generate a new sample and randomly select a sample from the nearest neighbor , generate new samples by linear interpolation:

[0321] (35);

[0322] in: Represents the current minority class sample; Indicates its Any neighbor among the nearest neighbors; represents the generated new sample, Represents a random number uniformly distributed between 0 and 1; the new samples synthesized by SMOTE are merged with the original training data to form an expanded training dataset.

[0323] S53: Multi-class linear support vector machine classifier training;

[0324] A multi-class linear SVM classifier is trained to learn the discriminant function between the fusion feature vector and the category. Since the tree species classification involved in this invention is a multi-classification task, a One-vs-Rest strategy will be used for training.

[0325] S531: Input data. The expanded training data set obtained through S522 can be expressed as:

[0326] Input data, the expanded training data set obtained through S522, is expressed as:

[0327] (36);

[0328] in: Indicates the first The fusion feature vector of samples, represents the true category label of the tree species image, represents the total number of training samples;

[0329] S532: Classification strategy: For each category, build a binary classifier, take the samples belonging to this category as positive class, and the samples belonging to all other categories as negative class. samples in the category The label value on is represented as:

[0330] (37);

[0331] S533: Define the discriminant function. For each binary classification, train a linear support vector machine classifier. The output of the corresponding linear discriminant function is expressed as:

[0332] (38);

[0333] in: and Represent the weight and bias of the classifier respectively;

[0334] S534: Minimize the soft margin loss function, The optimization goal corresponding to the class is:

[0335] (39);

[0336] in: represents the regularization parameter; Indicates the The slack variables of samples are constrained as follows:

[0337] (40);

[0338] S54: determining the optimal parameters of the classifier;

[0339] After training is completed, the classifier accuracy needs to be evaluated on the validation set entropy to determine the optimal regularization hyperparameters , and obtain the corresponding SVM classifier parameters.

[0340] S541: Construct multiple candidate hyperparameter sets and set a set of candidate regularization hyperparameters to form the set:

[0341] (41);

[0342] in: It is regularization hyperparameters to be evaluated, Alternatives

[0343] S542: Train multiple support vector machine classifiers for each hyperparameter ,train A binary classifier is used to obtain the classification parameters:

[0344] (42);

[0345] Each classifier is trained using formula (39) of S534 as the optimization target;

[0346] S543: Evaluate the performance on the validation set. Make the classifier for each set of parameters predict on the validation set and calculate the accuracy:

[0347] (43);

[0348] in: Indicates the number of images in the validation set; Indicates the The predicted category of samples under hyperparameters; Represents the indicator function, which is 1 when it holds, otherwise 0;

[0349] S544: Select the optimal hyperparameters to determine the classifier parameters and find the hyperparameters that optimize the performance of the validation set. :

[0350] (44);

[0351] Finally determine the corresponding optimal support vector machine classifier parameters:

[0352] (45);

[0353] in: Represents the index that makes the parameter reach the optimal value, Indicates that the index is Parameters when , that is, satisfy ; and Respectively represent the index Time The weights and biases corresponding to the classes; and They represent the weight vector and bias of the corresponding optimal support vector machine classifier respectively.

[0354] S6: Tree species identification: The tree species image to be identified is extracted using the bag-of-visual-words model and the state-space model, and the features are fused into multi-level features. The features are input into the trained support vector machine classifier to output the identification results of the tree species image.

[0355] The specific steps of S6 are:

[0356] S61: For the tree species image to be identified, perform image denoising and size normalization preprocessing using the S12 alignment;

[0357] S62: extracting local features from the preprocessed image using the bag-of-visual-words model of S2, and extracting global features using the state-space model of S3;

[0358] S63: Combining the local and global features obtained in S62, and using S4 to obtain fused multi-level features;

[0359] S64: Input the multi-level features of S63 into the support vector machine classifier trained in S54, and calculate the discriminant function of each class respectively:

[0360] (46);

[0361] S65: Select the category with the maximum discriminant score as the final prediction result:

[0362] (47);

[0363] The output predicted category is the recognition result of the tree species image in actual application.

[0364] The above disclosure is only a specific embodiment of the present invention, but the present invention is not limited thereto. Any changes that can be conceived by those skilled in the art should fall within the scope of protection of the present invention.

Claims

1. A tree species identification method based on bag-of-visual-words representation and state-space modeling, characterized in that: The following steps are involved: S1: Dataset construction; S2: Extract local features using the bag-of-visual-words model; S3: Extract global features using state-space model; S4: multi-level feature fusion; S5: Data oversampling and classifier training; S6: tree species identification; S21: Scale-invariant feature transformation key point extraction and description; S211: Generate scale space using Gaussian blur; S212: using the Gaussian difference image to locate the key point position; S213: Generate feature descriptors; S22: Accelerated robust feature key point extraction and description; S221: Hessian matrix calculation; S222: key point detection; S223: Generate feature descriptors; calculate Haar wavelet response statistics in each sub-region: S23: K-means clustering algorithm to build visual dictionary; S231: Collect feature descriptors of all images and construct the full feature set of the training set: S232: K-means clustering generates a visual dictionary. K-means clustering is used to cluster the entire set to obtain a visual dictionary: S233: Construct a visual word bag for each image. , assign all its joint descriptors to the nearest cluster center, build a bag-of-words histogram, weight the bag-of-words histogram, and finally obtain the bag-of-words vector representation of the image, that is, the local feature vector; S31: Image segmentation and linear embedding; S32: construct state space feature extraction block; S321: Construct multi-directional sequence input; S322: Unidirectional state space modeling; S323: Four-directional output fusion; S33: Building a state-space recognition network; S331: stack multiple state space feature extraction blocks, and the output of each feature extraction block serves as the input of the next layer; S332: Pooling obtains the global features of the image and average pooling is performed on the final stacked output sequence; S333: Supervised training of classifiers.

2. The tree species identification method based on bag-of-visual-words representation and state-space modeling according to claim 1 is characterized by: The specific steps of S1 are: S11: Image collection, using a high-resolution microscope system to collect cross-section images of tree species and convert them into PNG format; S12: image preprocessing; S121: Image denoising; S122: size normalization; S13: Dataset division.

3. The tree species identification method based on bag-of-visual-words representation and state-space modeling according to claim 1 is characterized in that: The specific steps of S2 are: S211: Use Gaussian blur to generate scale space, the image is , its Gaussian blur is expressed as: (1); in: is the standard deviation of the Gaussian function; and Respectively represent the pixel coordinates in the horizontal and vertical directions in the original spatial coordinate system of the image; The generated scale space is obtained by Gaussian blurring images at different scales: (2); in: Represents a two-dimensional convolution operation; S212: Use the Gaussian difference image to locate the key point position, which is expressed as follows: (3); S213: Generate feature descriptors, key points in image coordinates , first calculate the gradient information of the key point: (4); The direction of the gradient and amplitude They are: (5); S221: Hessian matrix calculation, expressed as: (6); in: For the original image at position Gray value at ; is the value of the integral image, indicating that arrive The sum of all pixels within the rectangle, the Hessian matrix describes the local curvature of the image, and the specific form is: (7); in: Indicates that the image is Second derivative of direction; Indicates that the image is Second derivative of direction; represents mixed second-order derivatives; It means that the result calculated by formula (6) is brought into the scale space obtained by formula (2); Approximate the Hessian matrix using a box filter with an integral image , and then calculate the determinant of the approximate Hessian matrix, which is expressed as: (8); in: , and is an approximation of box filtering; S222: Key point detection uses interpolation fitting to accurately locate the coordinates and scales of key points, and proposes edge response points based on the following judgment conditions: (9); in: is the trace of the matrix, i.e. the sum of the elements on the main diagonal; is the threshold; In the neighborhood of the key segment, calculate the Haar wavelet Horizontal response at and vertical response : (10); in: Indicates the position of a point in the neighborhood relative to the center of the key point, and Represent the horizontal and vertical coordinates in the local coordinate system respectively; The Gaussian weighting function is expressed as: (11); in: and Respectively represent the horizontal and vertical coordinates of the center position in the local neighborhood of the key point; Calculate the sum of the response amplitudes in each direction sector: (12); in: is the starting angle of the 6 sectors; Indicates the response direction; The main direction of the key point is expressed as: (13); Among them: argmax represents the value of the independent variable corresponding to the maximum value of a function; S223: Generate feature descriptors; calculate Haar wavelet response statistics in each sub-region: (14); in: ; ; S231: Collect feature descriptors of all images and construct the full feature set of the training set: (15); in: is the total number of images; Indicates the total number of key points extracted from the image; Indicates the Joint descriptor of feature points; express The first in the collection Each image Descriptors of feature points; S232: K-means clustering generates a visual dictionary. K-means clustering is used to cluster the entire set to obtain a visual dictionary: (16); in: Represents the cluster center set; R represents the symbolic interpretation of the real number set; Representation Descriptor The corresponding cluster number; Indicates the number of sight words; is the L2 norm; S233: Construct visual word bag for each image , assign all its joint descriptors to the nearest cluster center, build the bag-of-words histogram, image Middle The frequency of visual words is expressed as: (17); in: Representing an image The number of descriptors in ; is the indicator function; Weight the bag-of-words histogram, and the weighted bag-of-words histogram value Expressed as: (18); in: Indicates how many images the visual word appears in ; Finally, the bag-of-words vector representation of the image is obtained, that is, the local feature vector, which is expressed as: (19)。 4. The tree species identification method based on bag-of-visual-words representation and state-space modeling according to claim 3 is characterized by: The specific steps of S3 are: S31: Divide the input tree species image into multiple non-overlapping image blocks of fixed size. Each image block is flattened and mapped to a uniform dimension through a linear layer to form a preliminary representation of the image: (20); in: Indicates the The embedding vector of the image patch; represents the dimension of the embedding space; Indicates the number of image blocks, For each image block size, and Represent the length and width of the image respectively; S321: In order to capture different structural direction information in the image, the image blocks are serialized according to four directions to form input sequences in different directions: (21); in: ; S322: The formula of the state space model is expressed as: (22); in: Indicates the input vectors; Indicates direction Previous The hidden state at a moment; The feature representation of the output; represents the state transition matrix; Represents the transformation matrix of the input; represents the mapping matrix from hidden state to output, A matrix representing a direct mapping of input to output; S323: Perform weighted fusion on the outputs from the four directions to generate the final output sequence of the state space model, which is expressed as: (23); in: Represents the weighting coefficient for each direction; S331: The output of a feature extraction block is expressed as: (24); in: Indicates the number of feature extraction blocks, represents the linear embedding of the input layer image patch, Indicates the The operation of the state space feature extraction block, each block contains an independent state parameter matrix ; S332: Expressed as: (25); Among them: AVGPool represents the average pooling operation; (26); in: Represents the activation function operation; and Represent the weight and bias of the fully connected layer respectively; S333: Input the classification layer composed of 1×1 convolutional layers and predict the category score: (27); in: and Represent the weight and bias of the classification layer respectively; represents the logit score of each category; Next, use the normalized exponential function to get the class probability: (28); in: Indicates the number of categories; Using cross entropy loss To train: (29); in: Represents the tree species category labels annotated in the dataset.

5. The tree species identification method based on bag-of-visual-words representation and state-space modeling according to claim 4 is characterized in that: The specific steps of S4 are: S41: Local and global feature normalization. Formula (19) obtains the local feature vector, and formula (26) obtains the global feature vector. First, the two-level feature vectors are normalized. The normalized local feature vector and global feature vector are expressed as: (30); S42: Local and global feature fusion, using linear weighting for feature fusion, the fused feature is expressed as: (31); in: represents the weight parameter; S43: Fusion feature standardization, use L2 norm to standardize the fused features to obtain multi-level fusion features: (32)。 6. The tree species identification method based on bag-of-visual-words representation and state-space modeling according to claim 5, characterized in that: The specific steps of S5 are: S51: Identify minority class samples in the training set and classify them according to the image , select all minority class samples to form a set of feature vectors belonging to all minority classes: (33); in: The label identified as the minority class; | represents the conditional restriction symbol; S52: Generate new samples using minority class oversampling technology; S521: Calculate the nearest neighbor for each , find its Nearest neighbor samples: (34); in: Represents the fusion feature vector of any two samples; express The vector Dimensional component; Represents a vector No. Dimensional component; S522: Generate a new sample and randomly select a sample from the nearest neighbor , generate new samples by linear interpolation: (35); in: Represents the current minority class sample; Indicates its Any neighbor among the nearest neighbors; represents the generated new sample, Represents a random number uniformly distributed between 0 and 1; S53: Multi-class linear support vector machine classifier training; S531: Input data, the expanded training data set obtained through S522, expressed as: (36); in: Indicates the first The fusion feature vector of samples, represents the true category label of the tree species image, represents the total number of training samples; S532: Classification strategy: For each category, build a binary classifier, take the samples belonging to this category as positive class, and the samples belonging to all other categories as negative class. samples in the category The label value on is represented as: (37); S533: Define the discriminant function. For each binary classification, train a linear support vector machine classifier. The output of the corresponding linear discriminant function is expressed as: (38); in: and Represent the weight and bias of the classifier respectively; S534: Minimize the soft margin loss function, The optimization goal corresponding to the class is: (39); in: represents the regularization parameter; Indicates the The slack variables of samples are constrained as follows: (40); S54: determining the optimal parameters of the classifier; S541: Construct multiple candidate hyperparameter sets and set a set of candidate regularization hyperparameters to form the set: (41); in: It is regularization hyperparameters to be evaluated; S542: Train multiple support vector machine classifiers for each hyperparameter ,train A binary classifier is used to obtain the classification parameters: (42); Each classifier is trained using formula (39) of S534 as the optimization target; S543: Evaluate the performance on the validation set. Make the classifier for each set of parameters predict on the validation set and calculate the accuracy: (43); in: Indicates the number of images in the validation set; Indicates the The predicted category of samples under hyperparameters; represents the indicator function; S544: Select the optimal hyperparameters to determine the classifier parameters and find the hyperparameters that optimize the performance of the validation set. : (44); Finally determine the corresponding optimal support vector machine classifier parameters: (45); in: Represents the index that makes the parameter reach the optimal value, Indicates that the index is Parameters when , that is, satisfy ; and Respectively represent the index Time The weights and biases corresponding to the classes; and They represent the weight vector and bias of the corresponding optimal support vector machine classifier respectively.

7. The tree species identification method based on bag-of-visual-words representation and state-space modeling according to claim 6, characterized in that: The specific steps of S6 are: S61: For the tree species image to be identified, perform image denoising and size normalization preprocessing using the S12 alignment; S62: extracting local features from the preprocessed image using the bag-of-visual-words model of S2, and extracting global features using the state-space model of S3; S63: Combining the local and global features obtained in S62, and using S4 to obtain fused multi-level features; S64: Input the multi-level features of S63 into the support vector machine classifier trained in S54, and calculate the discriminant function of each class respectively: (46); S65: Select the category with the maximum discriminant score as the final prediction result: (47); The output predicted category is the recognition result of the tree species image in actual application.

Citation Information

Patent Citations

  • Remote sensing image classification and retrieval method

    CN108537238A

  • OCTA fundus image generation system, method, medium and device based on Mama and diffusion model

    CN120198318A