Method for generating aviation lifelike tree sample and identifying tree species by using image diffusion model

By integrating an image diffusion model that incorporates linguistic semantics, realistic tree samples are generated, solving the problems of illumination variation and environmental variation in existing tree species identification methods. This enables high-fidelity forest image generation and tree species identification, improving the accuracy and efficiency of forest monitoring.

CN121482601APending Publication Date: 2026-02-06NANJING FORESTRY UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511646486.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-11
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Existing technologies for tree species identification in forest monitoring rely on the discriminative features of RGB or hyperspectral images, which face challenges such as changes in light intensity, phenological changes, and environmental variations. Furthermore, they fail to fully utilize linguistic information for auxiliary identification, resulting in insufficient complexity and accuracy in identification.

Method used

By employing an image diffusion model that integrates language and semantics, and constructing a Language-Semantic Fusion Image Diffusion Network (LFINet), we combine language and semantics with the diffusion model to generate realistic tree samples that capture species-specific characteristics and seasonal phenological changes, thereby enhancing tree species identification in aerial imagery.

Benefits of technology

The generation of high-fidelity forest images improves the accuracy and efficiency of tree species identification, enables effective canopy detection and precise tree species identification, and enhances the ability to detect forest species.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121482601A_ABST
    Figure CN121482601A_ABST
Patent Text Reader

Abstract

The invention discloses a method for generating an aviation lifelike tree sample and identifying a tree species by using an image diffusion model. The method comprises the following steps: constructing a multi-modal data set; training an image-text comparison model and a Unet denoising network model; outputting predicted noise in the trained Unet denoising network model, and finally generating final image reconstruction which is in semantic alignment with the input text description; the obtained single tree crown images are synthesized; and constructing a YOLOv11 detection model, and training the YOLOv11 detection model by using the synthesized forest image to realize tree species identification. According to the method, language semantics and a diffusion model are combined to generate a vivid tree sample for capturing specific features of species and seasonal phenological changes, and the vivid tree sample is synthesized into a high-fidelity forest image, so that tree species identification in an aerial image is enhanced, and effective crown detection and accurate tree species identification are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of tree species identification technology, specifically a method for generating realistic tree samples and identifying tree species in aerial images using an image diffusion model that integrates language semantics. Background Technology

[0002] Forest ecosystems are vital to global biodiversity, carbon sequestration, and climate regulation; however, they face increasing threats from deforestation, climate change, invasive species, and habitat fragmentation. Accurate monitoring and species identification are crucial for advancing ecological conservation and sustainable forest management. The complexity and diversity of forest ecosystems present significant challenges to the application of deep learning in forestry, particularly due to the scarcity of training samples. Currently, multimodal data fusion—integrating different data sources such as images, textual semantics, and ecological metadata to provide a more comprehensive characterization of specific forest issues—is becoming a key technological trend and strategy for improving the accuracy and robustness of forest monitoring systems.

[0003] Despite the increasing exploration of multimodal remote sensing methods in forest monitoring, most current tree species identification methods still rely on discriminative features from RGB or hyperspectral images. Techniques such as YOLO, EfficientDet-D7, Mask R-CNN, and Segment Anything have been tailored to address ecological complexity. However, their effectiveness is limited by their heavy reliance on labor-intensive annotation and performance degradation under challenging conditions such as light variation, phenological change, and environmental variability. While some studies have attempted to enhance robustness through multimodal fusion (e.g., combining hyperspectral with aerial imagery or integrating multispectral and aerial data), spectral-based methods still face inherent limitations. These include “heterocoronamorphism” and “homophoreism” phenomena caused by interspecific similarity and intraspecific variation, which complicate accurate identification. A more fundamental limitation is the failure to fully utilize linguistic information for assisted identification, a problem that remains largely unresolved.

[0004] However, existing methods have not fully utilized large-scale semantic databases or tapped into the rich descriptive power inherent in human language. Large Language Models (LLMs), as core models in natural language processing, are essentially capable of understanding and generating natural language text by learning from massive text corpora. By modeling the probability distribution of word sequences, LLMs capture language patterns, structural, and semantic information, performing exceptionally well on various language tasks. Their strength lies in analyzing unstructured text data, a capability particularly valuable in contexts requiring the extraction of actionable insights from diverse linguistic sources. Pre-trained Transformer architectures, such as BERT and its derivatives (e.g., SciBERT for scientific literature, mBERT for multilingual tasks), represent fundamental advances in natural language processing. These models, initially developed for general language understanding, demonstrate remarkable ability to process text data through self-supervised learning on massive corpora. While LLMs are widely used across various industries, their application in forestry is still in its early stages. For example, recent work has begun exploring the application of LLMs in forestry. ForestryBERT, a customized BERT variant, addresses the dynamic nature of forestry knowledge by pre-training on 204,636 Chinese texts using a dynamic knowledge absorption mechanism. Despite their transformative potential, LLMs face key limitations in professional applications. First, domain-specific technical gaps persist: general-purpose models like BERT perform poorly in capturing specialized forestry knowledge due to their training on broad, non-specialized corpora. Second, LLMs typically exhibit poor interpretability, making it challenging to trace the reasoning behind their outputs, which is crucial for reliable decision-making in forestry. Third, deep reasoning within forestry-specific contexts is difficult: reliance on general-purpose corpora dilutes specialized, long-tail knowledge, and a lack of experience modeling domain-specific formats leads to the distortion or omission of critical details when addressing complex forestry problems.

[0005] On the other hand, the scarcity of labeled training data in deep learning applications has driven the adoption of generative models, which have revolutionized data synthesis across various fields and offer transformative potential for forestry and agriculture. These models can be broadly categorized into three main types: Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), and Diffusion Models, each bringing different mechanisms and applications to these fields. GANs, which contain a generator and a discriminator trained in an adversarial setting, remain crucial for dataset augmentation. In agriculture, GANs such as DCGAN have generated synthetic images of crop diseases, enabling robust training of diagnostic models on limited real-world data. In forestry, GANs have been used to generate synthetic forest images to assist tasks such as habitat mapping and species identification. However, GANs are prone to mode collapse and training instability—problems that become particularly pronounced in ecologically complex environments such as forests. High structural heterogeneity, occlusion, and varying lighting conditions pose challenges to both generators and discriminators: generators may fail to produce sufficiently diverse and realistic tree features, while discriminators may be overly sensitive to subtle texture and lighting variations, disrupting the adversarial balance. This often results in repetitive or unrealistic synthetic outputs that fail to adequately represent the true variability of the forest environment. Variational autoencoders (VAEs) encode data into a probabilistic latent space, enabling applications in agriculture and forestry through data-driven synthesis and cross-modal generation. In forestry, they have been used to generate synthetic canopy images conditioned on labeled species, thereby augmenting datasets for training deep learning classifiers. While advanced variants like β-VAEs and conditional VAEs (CVAEs) offer the potential to decouple latent factors such as seasonal variation or lighting conditions, their direct application to simulating forest growth dynamics or time-series canopy maps remains exploratory and relies on integration with ecological models or multimodal sensing data. A key limitation of Visual Image Processing (VAEs) is their tendency to trade off sharpness and fidelity, often producing smoothed or blurred outputs due to variational approximations in the latent space. In forest environments, this smoothing effect impairs the model's ability to preserve fine-grained texture and morphological features crucial for species identification. For example, when distinguishing visually similar species, such as Scots pine and Norway spruce—which differ subtly in crown shape, branching patterns, and bark texture—images generated by VAEs may lack the high-frequency details needed to capture these discriminative features, thus reducing the reliability of downstream classification tasks. Diffusion models, particularly denoising diffusion probabilistic models (DDPMs), have recently demonstrated excellent performance in generating high-fidelity forest scenes. Unlike GANs, which suffer from pattern collapse and training instability, DDPMs iteratively refine noise into structured plant details, producing images with accurate leaf vein patterns and bark textures.In forestry, DDPMs simulate rare ecological events by generating them conditioned on climatic variables such as temperature anomalies and precipitation variations. Emerging approaches integrating language and generative models, such as Rsdiff, and multimodal frameworks like Versatile Diffusion, demonstrate great potential for efficient and semantically rich urban building detection and forest scene generation, although their application in this field faces further challenges and room for development.

[0006] In summary, a key challenge in applying artificial intelligence to forest monitoring is tree species identification, which is hampered by complex visual variations and limited canopy data. Furthermore, how to utilize language-driven synthesis to generate high-fidelity forest images and how to accurately identify tree species to improve forest species detection capabilities remain unresolved issues. Summary of the Invention

[0007] The technical problem to be solved by this invention is to provide a method for generating realistic tree samples and identifying tree species in aerial imagery by integrating a language semantic image diffusion model to address the shortcomings of the prior art. This method proposes a language semantic image diffusion network that combines language semantics with a diffusion model to generate realistic tree samples that capture species-specific characteristics and seasonal phenological changes. The realistic tree samples are then synthesized into high-fidelity forest images, thereby enhancing tree species identification in aerial imagery and achieving effective canopy detection and accurate tree species identification.

[0008] To achieve the above-mentioned technical objectives, the technical solution adopted by the present invention is as follows:

[0009] An image diffusion model for generating realistic tree samples and identifying tree species in aerial photography includes the following steps:

[0010] Step 1: Collect forestry-related text semantic datasets and aerial image datasets containing tree canopies. Pair each tree canopy image in the aerial image dataset with a text semantic information annotation in the text semantic dataset to form a multimodal dataset.

[0011] Step 2: Train an image-text comparison model using a multimodal dataset, which includes an image encoder and a text encoder.

[0012] Step 3: Construct a diffusion model, which includes a noise-adding module, a time-step embedding module, and a Unet denoising network model. Use the text embedding tensor output by the trained text encoder, the high-dimensional vector output by the time-step embedding module, the noise image output by the noise-adding module, and the real noise to train the Unet denoising network model, and obtain the trained Unet denoising network model.

[0013] Step 4: Encode a batch of text into a text embedding tensor containing contextual semantics through the trained text encoder. This text embedding tensor, the Gaussian noise image, and the high-dimensional vector identifying the current denoising stage output by the time-step embedding module are injected into the trained diffusion model's Unet denoising network model, outputting the predicted noise, and finally producing the final image reconstruction aligned with the semantic description of the input text.

[0014] Step 5: Combine the individual tree canopy images obtained in Step 4 to form a forest image;

[0015] Step 6: Construct a YOLOv11 detection model by training the YOLOv11 detection model using the synthesized forest image and the tree species label information corresponding to each tree in the forest image.

[0016] Step 7: Input a forest image of the tree species to be identified into the trained YOLOv11 detection model, and output the tree species corresponding to each tree in the forest image.

[0017] As a further improvement to the present invention, step 2 is as follows:

[0018] Step 2.1: Construct an image-text comparison model, which includes a text encoder and an image encoder; the image encoder uses Vision Transformer; the text encoder includes the Word2Vec model and a Transformer encoder.

[0019] Step 2.2, the image encoder's processing procedure is as follows: the input image The image is divided into multiple image blocks, which are then combined with positional encodings and fed into the Transformer architecture to generate image embeddings. ;

[0020] The text encoder's processing procedure is as follows: Input text semantic information annotation Byte-to-byte encoding (BPE) is used for word segmentation, followed by text cleaning. The cleaned tokens are then mapped to a high-dimensional space using a pre-trained Word2Vec model. Each token is converted into a 512-dimensional vector, generating a text embedding matrix. This matrix is ​​then processed by a Transformer encoder with eight attention heads, outputting a text embedding tensor that includes contextual semantics. Each sentence in the batch is represented by a single 512-dimensional vector;

[0021] Step 2.3, Image Embedding With sentence embedding tensor The similarity matrix is ​​formed by combining the elements. Each row in the similarity matrix corresponds to an image embedding, and each column corresponds to a text embedding. The diagonal elements represent true pairings, and the off-diagonal elements correspond to incorrect pairings.

[0022] Step 2.4: Train the image-text comparison model using the multimodal dataset according to steps 2.2-2.3, use symmetric cross-entropy loss to optimize the similarity score, and update the image encoder and text encoder.

[0023] As a further improved technical solution of the present invention, the Unet denoising network model in step 3 includes three encoder layers, one bridging layer, three decoder layers, and a convolutional layer. The encoder layer includes two parallel groups and one downsampling layer. Each parallel group includes a residual block ResBlock for processing temporal embeddings and image features, and a spatial Transformer module for integrating text embeddings through cross-attention. The bridging layer includes three residual blocks ResBlock, one spatial Transformer module, and three residual blocks ResBlock arranged in sequence. The decoder layer includes an upsampling layer and two parallel groups, each containing a residual block ResBlock and a spatial Transformer module. The encoder layer and the decoder layer are skip-connected.

[0024] As a further improvement to the present invention, the processing procedure of the noise-adding module in step 3 is as follows:

[0025] The forward process modifies an original image by progressively adding Gaussian noise. Converted to standard normally distributed noise; its conditional probability is expressed as:

[0026] (1);

[0027] in, Indicates from arrive Image sequences with added noise Indicates that the i-th image has passed through... Images with Gaussian noise added at each time step; For hyperparameters, Describe an identity matrix. Indicates the image dimension;

[0028] The recursive property allows it to be further expressed as:

[0029] (2);

[0030] Indicates from image arrive Added noise;

[0031] conditional probability Represented as a Gaussian distribution:

[0032] (3);

[0033] conditional probability Represented as a Gaussian distribution:

[0034] (4).

[0035] As a further improvement to the present invention, the processing procedure of the time step embedding module in step 3 is as follows:

[0036] Time step The time-step embedding module processes the data and applies a sinusoidal embedding function to generate a position encoding matrix. This matrix is ​​then processed by a linear layer to expand its dimensions, resulting in a high-dimensional vector. This high-dimensional vector is then fed into each ResBlock in the Unet denoising network model.

[0037] As a further improvement to the present invention, the loss function of the Unet denoising network model in step 3 is:

[0038] (5);

[0039] in, Expressing expectations, For model parameters From noisy images and text embedding Predicted noise, This is real noise.

[0040] As a further improved technical solution of the present invention, the YOLOv11 detection model in step 6 includes an EfficientNet backbone network, a bidirectional feature pyramid network BiFPN, a category prediction network, a detection box prediction network, and a Dynamic Anchor module.

[0041] The beneficial effects of this invention are as follows:

[0042] To address the interrelated challenges of species diversity and phenology, this invention proposes a Linguistic Semantic Fusion Image Diffusion Network (LFINet). This overall framework synergistically combines linguistic semantics with a diffusion model to generate realistic tree samples that capture species-specific characteristics and seasonal phenological changes, thereby enhancing tree species identification in aerial imagery. LFINet presents an enhanced forest tree species diffusion model. This model integrates an improved U-Net-like architecture and incorporates guidance from a textual semantic network. This invention also generates high-fidelity forest images from these synthetic images. These high-fidelity forest images serve as training samples for effective canopy detection and tree species identification, driven by a custom YOLOv11 architecture optimized for large-scale forest monitoring. Attached Figure Description

[0043] Figure 1 This represents a subset of the aerial tree image dataset.

[0044] Figure 2 This diagram illustrates the theoretical process of forward noise injection and backward denoising applied to an image using a diffusion model.

[0045] Figure 3 This represents the network architecture diagram of the text semantic guided image generation component based on diffusion modeling in LFINet.

[0046] Figure 3 (a) in the diagram represents the contrastive language-image pre-training framework.

[0047] Figure 3 (b) in the diagram represents the training process of LFINet.

[0048] Figure 3 (c) in the diagram represents the generation process of LFINet.

[0049] Figure 4 This diagram shows the detailed structure of the residual block and spatial transformation module in the improved U-Net.

[0050] Figure 5 This diagram illustrates the improved YOLOv11 architecture (highlighting the Bidirectional Feature Pyramid Network (Bi-FPN) and the dynamic anchor mechanism).

[0051] Figure 6 A visual demonstration of the generative model.

[0052] Figure 6 In the example (a), the text prompts guide the process of gradually denoising from random noise (t = 500) to a clear image (t = 0).

[0053] Figure 6 (b) in the figure represents the tree species detection results of aerial photographs of Nanjing Forestry University and Shanghai Botanical Garden, as well as aerial images generated from text descriptions of the same species.

[0054] Figure 6 (c) in the figure represents canopy samples generated under different conditions and comparisons with GPT-4o, DALL-E 3 and Leonardo AI.

[0055] Figure 6 In the example, (d) represents multiple sample images generated by eight different semantic cues under different weight configurations.

[0056] Figure 7 (a) in the figure represents the trend of the loss function of the generated model under different parameter configurations.

[0057] Figure 7 (b) in the figure represents the ablation experiment results of the text encoder.

[0058] Figure 8 This is a flowchart illustrating the detection process of the improved YOLOv11.

[0059] Figure 9 This section compares the performance metrics of different experimental samples using the enhanced detection model (i.e., our model). Detailed Implementation

[0060] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings:

[0061] 1. A key challenge in applying artificial intelligence to forest monitoring is tree species identification, which is hampered by complex visual variations and limited canopy data. Furthermore, an AI-assisted identification framework that effectively integrates large language models with forestry human knowledge bases has not yet been fully established. Against this backdrop, this study introduces the Image Diffusion Model Integrating Language Semantics (LFINet), a framework that combines text-guided image generation with tree species detection, aiming to enhance forest applications through technological synergy. This framework comprises three key methodological components: (1) A module designed to leverage detailed textual descriptions of the forest environment, integrating elements such as species characteristics, spatial layout, and phenological features. It generates text embeddings aligned with forest knowledge semantics. (2) An innovative diffusion-based framework integrating an improved U-Net architecture, Markov chain theory, and textual semantic embeddings, fusing noisy images with forest-specific language semantics to generate highly realistic aerial tree images. (3) An optimized pipeline for detection using generated images, based on an enhanced YOLOv11 architecture with a context-aware feature extractor and adaptive anchor scaling, enabling efficient, large-scale canopy detection and tree species identification. Experimental evaluations validated the effectiveness of the proposed framework. The image generation module achieved a structural similarity index (SSIM) of 0.94 and a Fraser onset distance (FID) score of 6.42, demonstrating the excellent fidelity of the synthesized output. Furthermore, the detection pipeline achieved a mean accuracy (mAP50) of 0.868 in the tree species identification task, consistently outperforming all baseline models across all evaluation metrics. This integrated system enhances forest species detection capabilities by leveraging language-driven synthesis to generate high-fidelity forest images and perform tree species identification.

[0062] 2. Materials and Methods:

[0063] 2.1 Data Acquisition:

[0064] In this study, we employed a variety of image datasets to ensure broad geographical, ecological, and phenological coverage for multimodal forest analysis. The selected datasets included both international and domestic sources, supplemented by aerial imagery collected from publicly available internet repositories. International datasets formed the foundation of our research. The Prue Forest dataset provides images of temperate urban forests with representative canopy structures, while the Quebec Trees dataset offers high-resolution, multi-seasonal images of northern mixed forests, capturing significant phenological variations under spring, summer, and autumn conditions. Furthermore, specialized single-species datasets—Kolovai-Trees (coconut trees) and Larch_Casebearer (larch)—enhanced the characterization of specific tree morphologies. The Zhongshan Scenic Area Canopy Inventory provides high-resolution UAV imagery for managing the memorial landscape, while the Nanjing Forestry University Botanical Garden dataset offers multi-temporal observations of multiple woody species from 2018 to 2022. Additionally, our training corpus was augmented with aerial imagery of multiple tree species obtained from publicly available internet databases, expanding the diversity of canopy structures and environmental conditions. Table 1 summarizes the key features of the four main datasets, including species diversity, sample size, class distribution, spatial resolution, phenological diversity, and major tree species.

[0065] The language dataset comprises 150 academic articles and forestry reports (totaling approximately 1.2 million words), sourced from multiple databases: Google Scholar (80 articles), Web of Science (50 peer-reviewed studies on ecology and forestry), and CNKI (20 reports focusing on subtropical forestry data in China). Text data was retrieved through a combination of manual searching and automated tools, including web crawling using Python libraries such as BeautifulSoup and database API interaction using requests. Keywords included terms such as “forest ecology,” “tree species classification,” and “urban forestry.” Results were filtered to include publications from 2015 to 2022 to ensure relevance and timeliness. The raw text data underwent rigorous preprocessing to transform it into a structured format suitable for analysis. This includes text cleaning using regular expressions to remove HTML tags, special characters, and non-text elements; text segmentation into words and sentences using the NLTK library; lowercaseization and lemmatization using NLTK's WordNetLemmatizer to standardize terms; and domain-specific processing using a predefined vocabulary to extract forest-related terms such as species names and phenological terms. The final text dataset is then embedded using a pre-trained language model, which, when combined with image data, improves the interpretability of the diffusion model in a forest-specific context.

[0066] Table 1. Summary of the dataset used in this study:

[0067] Dataset Name Number of species Total sample size Average number of samples per species Spatial resolution (GSD) Seasonal / phenological diversity Main tree species PrueForest 18 69110 ~3893 average 20 cm Low (single-season winter imagery from January 2023) Deciduous oak, evergreen oak, beech, chestnut, black locust, coastal pine, fir, spruce QuebecTrees 14 22260 1590 average 1.81–2.02 cm High (Seven collections from late spring to early autumn in 2021, spanning multiple seasons, capturing phenological changes such as leaf senescence in temperate forests) Balsam fir, striped maple, red maple, paper-bark birch, American beech, spruce Zhongshan Scenic Area ~50 ~5,000 ~100 average 5 cm Medium (drone capture in the memorial landscape, partly due to phenological changes from seasonal flights) Cedar, Japanese cypress, Chinese juniper, ginkgo NJFU 47 12,347 ~262 average 2-5 cm High (multiple time phases from 2018 to 2022, capturing phenological changes such as leaf expansion and senescence) dawn redwood, juniper, sycamore, ginkgo

[0068] To construct a robust multimodal dataset, individual canopy images were first extracted from all source datasets and uniformly resized to 32×32×3 pixels. This image patch resolution was chosen to strike a balance between computational efficiency and preservation of key details (such as leaf texture), while maintaining consistency with established benchmarks such as DDPM on CIFAR-10. Each image then underwent a series of preprocessing operations to improve invariance and generalization ability. These operations included affine transformations (rotation, horizontal flipping) and radiometric adjustments (brightness and contrast varied within ±30%). Each processed image was paired with a descriptive text annotation detailing the collection angle, species classification, lighting conditions, and phenological stage. The final multimodal dataset encompasses a variety of tree species and is divided into training, validation, and test sets in a 7:2:1 ratio to facilitate reproducible evaluation of model performance. Figure 1 This dataset showcases approximately 30 representative species examples and their textual descriptions. It supports joint training of a CLIP-based visual-language encoder and a diffusion-based generator, achieving efficient cross-modal alignment between visual features and ecological semantics to enhance species detection in complex forest environments.

[0069] Figure 1 This is a subset of our aerial tree image dataset, with each image also accompanied by corresponding semantic information annotations (not shown in the figure), including the acquisition angle, species description, lighting conditions, and phenological characteristics, which are used as training samples for the network.

[0070] 2.2 Image generation guided by textual semantics using a diffusion model in LFINet:

[0071] This section describes the LFINet text-guided image generation framework, which integrates three core components: a denoising diffusion process for learning image representations, a contrastive image-language pre-training module for cross-modal alignment inspired by the Contrastive Language-Image Pre-training (CLIP) framework, and a noise prediction network based on an improved U-Net conditioned on text input. Starting with random noise, the model iteratively refines the image using textual cues to synthesize realistic forest scenes with detailed botanical features, ensuring ecological plausibility.

[0072] 2.2.1 The Markov chain theory basis of our diffusion model:

[0073] This section outlines the theoretical foundation of our proposed framework, such as... Figure 2 As shown, the forward process involves progressively injecting Gaussian noise into the input image via a Markov chain mechanism. (Where i represents the image index in the batch, i.e., the i-th image in a certain batch. At a predefined number of time steps...) This transforms the original data distribution into an approximate standard normal distribution. Conversely, the reverse process utilizes a learned neural network. The network iteratively refines the intermediate noise states. Starting with the terminal noise distribution, the network gradually reconstructs a high-fidelity image by predicting and removing noise. The trained model parameters... The current time step is used to generate a clearer image at each previous step.

[0074] Figure 2 This describes the theoretical process of forward noise injection and reverse denoising applied to a diffusion model of images.

[0075] a) Forward diffusion process ( Figure 2 (From left to right)

[0076] =Forward process Through Gaussian noise is gradually added at each time step to gradually transform an original image. Convert to standard normally distributed noise. This transformation is based on conditional probabilities following a Gaussian distribution. Control, corresponding to the predefined noise scheduling mechanism in the diffusion model, can be represented by the following equation:

[0077] ;

[0078] in Indicates from arrive Image sequences with added noise Indicates time step The image at that point has had noise added for t steps. We can flatten it. and Then, through reparameterization, we obtain... , The noise follows a Gaussian distribution and has a magnitude of It is a hyperparameter that controls the noise level at each diffusion step. It is linearly scheduled using a specific equation, starting from the initial value. Increase to the final value The formula for this linear scheduling is: . This represents an identity matrix whose dimensions correspond to the flattened dimensions. This ensures that the noise covariance matrix is ​​diagonal, and each diagonal element is... This indicates that the variances of each dimension are the same and the noise between dimensions is not correlated. This indicates the image dimensions, specifically 32×32×3. This represents the transpose of a matrix. It is a probability distribution, representing a given under conditions The distribution of the noise-adding process, i.e., the noise-adding process, follows a Gaussian distribution. , middle It is a random variable. The mean parameter, This is the variance parameter.

[0079] The recursive property allows it to be further expressed as:

[0080] (2);

[0081] noise matrix , , , and All are standard normally distributed noise. For is arrive Added noise, From arrive The added noises differ in specific values, but all follow a Gaussian distribution. Represents two independent standard normal variables and A linear combination of , with coefficients respectively and Since these noise terms are independent, their combination is also normally distributed with a mean of 0 and a variance of 0. , Right now Based on this recursive relationship, we can now derive the conditional probability distribution, and express the conditional probability... It can be represented as a Gaussian distribution, as shown in formula (3). Similarly, we can express the conditional probability as... It is represented by a Gaussian distribution, as shown in formula (4).

[0082] (3);

[0083] (4).

[0084] b) Reverse diffusion process ( Figure 2 (The process from right to left in the middle)

[0085] The reverse process involves recovering a sharper image from a noisy image by incorporating textual semantic guidance; it is the inverse of the forward process. The goal is to recover the original image from the noisy image while aligning it with linguistic semantics. In this process, the model learns text embeddings... To conditionally remove noise stepwise, a clean image sample matching the semantic meaning of the text is ultimately recovered. This process can be represented by the following equation:

[0086] (5).

[0087] Indicates text embedding Given the joint probability distribution, it is explained that the noise image is obtained from the standard normal distribution. Initially, through model parameters Gradually denoise to generate image sequences , ,…, Specifically, in this process, we hope to learn a parameterized conditional distribution. These are the initial conditions for the entire reverse process, defining the probability distribution of the noisy data. Conditional probability distribution. This indicates that the model is in a given current noise state. and Under the condition of the previous state The prediction. Mean Indicates in Under guidance from arrive The predicted transfer, while the variance is fixed at . , This represents the data distribution of the image after adding T steps of noise. Represents variance. Represents the identity matrix.

[0088] c) Loss function in the diffusion model:

[0089] The ultimate goal of the diffusion model is to make the distribution learned during the inverse denoising process... Restore the original data as faithfully as possible. This is equivalent to maximizing the log-likelihood of the training data. However, direct calculation It is difficult to handle. Instead, we optimize its variational lower bound and Jensen's inequality, as shown in equation (6).

[0090] (6).

[0091] yes The upper bound of the variational lower bound is the upper bound of the diffusion model, and the training objective is to minimize this upper bound. It refers to expectations. It is The two interactive positions within the parentheses are represented by Bayes' theorem, indicating... , Distribution under certain conditions.

[0092] The detailed derivation of the last step of formula (6) is as follows:

[0093] (a1);

[0094] Based on Bayes' theorem and the properties of Markov chains The second term on the right side of formula a1 can be expressed as:

[0095] (a2);

[0096] Substituting formula a2 into formula a1 and combining like terms, we get the simplified expression:

[0097] (a3);

[0098] Combining formula (a4) with the definition of KL divergence, formula (6) in the main text can be directly derived. .

[0099] By applying the Kullback-Leibler divergence (DKL)—which measures the difference between two probability distributions—to the first two terms on the right side of the last row of Equation 6, we can further elaborate the equation as follows:

[0100] (7).

[0101] This is the divergence formula, used to measure the distance between two distributions in parentheses. It is not a direct product; there are other corresponding formulas.

[0102] During the training process, the project and It can be treated as a constant and subsequently ignored because There is a lack of learnable parameters. Use D~KL~ directly Posterior probability of the forward process Compare them. According to the Markov property, when given... hour, Conditions independent of Using Bayes' theorem Incorporating these posterior probabilities as conditional information makes them easier to process:

[0103] (8).

[0104] mean and variance It can be calculated based on the probability density function. A detailed derivation is outlined below:

[0105] To clarify the specific derivation process of the parameters in formula (8), we ignore the common proportionality constant in the probability density function on both sides of the equation, and express the left side of formula (8) as:

[0106] (a4);

[0107] The right side of formula (8) can be expressed as:

[0108] (a5);

[0109] make and We can obtain:

[0110] (a6);

[0111] Based on formula (2) The derivation is Substitute the result into formula (a6) and simplify it to obtain the derivation process of formula (9).

[0112] (9).

[0113] Given two Gaussian distributions and and order We can substitute them into the KL divergence formula, that is... Will It is expressed as follows:

[0114] (10).

[0115] Our ultimate goal is to minimize the loss function, which involves making according to formula (9) Form and Formal matching , in It is determined by model parameters From noisy images and text embedding The predicted noise. Using Equation 10, it can be simplified to the following equation:

[0116] (11).

[0117] Formula (11) simplifies the loss function to directly using the prediction noise. and real noise This problem is solved by the mean square error (MSE) between the two sides, where C is a constant.

[0118] d) Generating images using a diffusion model:

[0119] Using a trained noise predictor We can also introduce And apply linear transformation to generate samples Here, we take the mean. and variance And substitute them into formula (5). Then we can... Perform reparameterization as follows:

[0120] (12).

[0121] Equation 12 represents a one-step denoising process. Each subsequent step systematically reduces the noise level while enhancing the complex details of the image, gradually aligning it with the learned image distribution. This process continues until all time steps are completed (from...). After reaching 0), the final latent variable It appears as a generated image.

[0122] 2.2.2 Contrastive Language-Image Pretraining Framework (CLIP):

[0123] In this study, we employ the CLIP framework (i.e., a text-image contrast model), which utilizes contrastive learning to train a model capable of aligning text and visual representations within an aligned embedding space. For example... Figure 3 As shown in (a), this image-text contrast model includes aligned image-text pairs. The training was performed on a large preprocessed dataset, where each image was associated with a corresponding text description. These pairings were of size [size missing]. The batch organization is achieved through a dual encoder—a text encoder and an image encoder.

[0124] Figure 3 This is a network architecture diagram of the text semantic guided image generation component based on diffusion modeling in LFINet. The model consists of three core modules: (a) a contrastive language-image pre-training framework, (b) the training process of LFINet, and (c) the generation process of LFINet.

[0125] a) ViT-based image encoder:

[0126] Image encoder, in Figure 3The blue module in (a) highlights the Vision Transformer (ViT), chosen for its low computational requirements and ease of integration. (Input image) It is divided into 4 image blocks, each with a size of (This refers to a small image with sizes of c, p, p) Figure 3 In (a), b, z, c, and p refer to the size of the entire batch, b is the batch size, z indicates that an image is divided into 4 parts (z=4), and c represents the number of channels in the image. The size of p is set to 16, for RGB images with 3 channels, and p indicates that each small image has a width and height of 16. These image blocks, combined with positional encoding, are input into the Transformer architecture to generate image embeddings. , The dimension is b is the batch size, and e represents the word vector size. Set to 512 to align with the text embedding generated in section b).

[0127] b) Text encoder using Word2Vec and Transformer:

[0128] like Figure 3 As described in the text encoder module (yellow area) of (a) in the diagram, for each input sentence First, word segmentation is performed using Byte-Pair Encoding (BPE). After text cleaning (removal of punctuation, numbers, and non-informative terms), these tokens are mapped to a high-dimensional space using a Word2Vec model pre-trained on a domain-specific forestry corpus (described in Section 2.1). Each token is transformed into a 512-dimensional vector. , generates a size of The text embedding matrix, where: Indicates batch size (number of sentences). This represents the maximum number of lexical units in each sentence, and 'e' represents the size of the word vector, which is 512. The Word2Vec model is trained using a Skip-gram method with a window size of 5 to maximize the probability of predicting context lexical units (the first two and the last two). The training objective is to maximize the probability of predicting the center lexical unit (using negative sampling loss). Cosine similarity between the lexical unit and its context:

[0129] (13).

[0130] here, This indicates the total number of tokens in the current text. This refers to contextual words. It is the central word element. yes The context words, if j=1, yes The word after j=-1 yes The word before that, Indicates word elements The transpose of the context word vectors. This represents the total number of unique lexical units in the entire training vocabulary. It is the first in the vocabulary list (that is, the first) Transpose of (n) word association vectors.

[0131] Finally, the size is The text embedding matrix is ​​processed by a Transformer encoder with 8 attention heads. The output is a contextualized sentence embedding tensor. Size is Each sentence in the batch is represented by a single 512-dimensional vector.

[0132] c) Joint image-text embeddings using contrastive learning:

[0133] like Figure 3 As shown in (a) of the similarity matrix, each row corresponds to an image embedding, each column corresponds to a text embedding, and the diagonal elements represent true pairings (high similarity), while the off-diagonal elements correspond to incorrect pairings (low similarity). We use symmetric cross-entropy loss to optimize these similarity scores, defined as follows:

[0134] (14);

[0135] in Represents image embedding and text embedding The cosine similarity between them. In our experimental setup, the CLIP model uses fixed temperature parameters. = 0.7 to balance the model's sharpness and generalization ability during training. For contrastive learning, we adopted data from... In-batch hard negative sampling of the similarity matrix (excluding diagonal elements) enables efficient contrastive optimization without the need for an external memory bank—a design choice crucial for handling our dataset size of 1.2 million text terms and over 30,000 images. The training scheme consists of 200 epochs (approximately 50,000 iterations, batch size 25) using the AdamW optimizer with a learning rate (lr) of 1e-4 and weight decay of 0.01. Convergence is considered achieved when the validation loss stabilizes (change < 0.05 over 10 consecutive epochs).

[0136] 2.2.3 Improve the training and generation process of U-Net:

[0137] The training process follows Figure 3 The improved U-Net architecture shown in (b) accepts three inputs: 1) a batch of noisy images By adding noise to the original image 2) The corresponding text embedding generated by the text encoder (yellow block) for the input sentence. 3) Time step .

[0138] Time step The time-step embedding module (dark green ellipse) processes the data, applying the sinusoidal embedding function defined in formulas (15) and (16) to generate the position encoding matrix. .here This represents the embedding dimension of the model. This dimension is determined by the size of the feature maps propagated through the improved U-Net architecture. This matrix... Then it is processed by a linear layer (light green ellipse), which expands its dimensions. It is then fed into each ResBlock to provide time step conditions throughout the network.

[0139] (15);

[0140] (16);

[0141] In the formula, for i=0,1,2…, This only indicates that i is a number from 0 to... The sequence index, in The value in the t-th row and 2i-th column of the matrix is ​​represented by .

[0142] Figure 3The improved U-Net architecture in (b) consists of three encoder layers, one bridging layer, and three decoder layers. Each encoder layer contains two parallel groups. Each group includes a ResBlock for processing temporal embeddings and image features, and a Spatial Transformer module for integrating text embeddings through cross-attention (see detailed architecture). Figure 4 , Figure 4 This diagram shows the detailed structure of the residual blocks and spatial transform module in the improved U-Net. Each encoder layer is followed by a downsampling operation via strided convolution. The bridging layer consists of three ResBlocks, one spatial Transformer module, and three more ResBlocks arranged sequentially. Each decoder layer begins with an upsampling module, followed by two parallel groups of ResBlocks and spatial Transformer modules. Skip connections (gray arrows) connect the corresponding encoder and decoder layers to preserve contextual information.

[0143] The final convolutional layer output predicts noise. The input image batches are used. These predictions are compared with the real noise. The mean squared error (MSE) loss, based on the simplified form of Equation 11, is used for comparison to optimize network parameters. The loss function is as follows:

[0144] (17);

[0145] for Figure 3 The generation process described in (c) is the reasoning process from... Proceed backwards. First, a batch of text prompts is encoded by a pre-trained text encoder into contextual embeddings, Gaussian noise samples (media / image149.wmf), and temporal embeddings (for each time step). The noise estimate is injected into the trained U-Net. Through iterative denoising steps controlled by Equation (12), the model progressively refines the noise estimate to produce intermediate latent representations. This recursive process continues until... This produces a final image reconstruction that is semantically aligned with the input text description.

[0146] To train the improved U-Net, we used the AdamW optimizer with an initial learning rate of 1×10⁻ 4A linear warm-up plan was employed for the first 1,000 steps to stabilize early training, and an exponential moving average (EMA) with a decay rate of 0.999 was used to improve generalization by averaging model weights between iterations. Training was performed for 70,000 steps on an NVIDIA A100 GPU with a batch size of 25, and convergence was ensured based on the validation FID score. During training, the text encoder was frozen to utilize domain-specific knowledge without overfitting, while the U-Net diffusion backbone was fully trained. Shared components included the variance scheduler. And noise samplers, which are fixed in both the forward and reverse processes to ensure consistency.

[0147] 2.3 Forest canopy detection in LFINet using diffusion-generated output:

[0148] To utilize synthetically generated canopy images for practical forest monitoring, 32×32 pixel generated image patches are first combined into a 640×640 pixel synthetic image (since the generated images are all single trees, they need to be stitched together) to simulate real forest scenes. This approach allows for the creation of larger scenes on a scalable basis and can be extended to real-world forest cover through iterative tiling or sliding window strategies, thereby reducing the attention complexity inherent in processing high-resolution aerial data. Based on this, we developed an enhanced YOLOv11-based detection pipeline (the training dataset is partly derived from real data and partly from augmented generated images. Detection is the input plot image and the output is the detection result, i.e., anchor boxes and tree species), which includes two key architectural improvements: (1) We implemented a Feature Pyramid Network (FPN) with bidirectional (BiFPN) cross-scale connections, enabling richer multi-scale feature fusion. This improvement is crucial for detecting canopies of different sizes in a single synthetic image, ensuring robust performance in different forest scenes. (2) We integrated a globally pooled attention-enhanced fully connected layer into the classification branch. These layers utilize semantic context from the entire image to modulate classification scores, improving species differentiation by combining local features with broader contextual cues. See [link to network structure] for details. Figure 5 .

[0149] The improved YOLOv11 architecture uses the EfficientNet backbone network to extract multi-scale features from a 640×640 input image, generating feature pyramids P1 to P7. Features from layers P3 to P7 are then fed into a Bidirectional Feature Pyramid Network (BiFPN) module, progressively optimizing feature representations through a multi-layered cross-scale fusion path from top to bottom and bottom to top. The enhanced features are then processed and selected by a dynamic anchor module, and finally, the parallel prediction heads—a category prediction network responsible for object classification and a detection box prediction network responsible for bounding box regression—process and output the detection results.

[0150] Specifically, the EfficientNet backbone network efficiently and hierarchically extracts increasingly abstract features from the input image. EfficientNet achieves an excellent balance between computational efficiency and accuracy through composite model scaling (coordinating scaling depth, width, and resolution), making it well-suited for tasks requiring high-resolution input, such as aerial imagery. The input is a three-channel (RGB) aerial image with a size of 640×640 pixels. The image sequentially passes through multiple stages of EfficientNet, each stage downsampling the feature map (reducing its size and increasing the number of channels). Finally, the backbone network outputs multi-scale feature maps, typically named P1 to P7, with a size of [missing information - likely a value]. Figure 5As shown, the BiFPN Layer (Bidirectional Feature Pyramid Network) receives multi-scale features generated by the backbone network and performs cross-scale information fusion through bidirectional paths from top to bottom and bottom to top. BiFPN uses learnable weights to balance the importance of different input features, ensuring that each layer of the fused feature map simultaneously contains high-resolution detail information and high-level semantic information. This is crucial for tree detection in aerial imagery with significant scale variations, significantly improving the detection capability for canopies of different sizes. Input: Feature maps at multiple levels output by the backbone. Output: A set of enhanced multi-scale feature pyramids or dynamic anchor modules. The Dynamic AnchorModule operates on each layer of enhanced feature maps output by BiFPN, primarily implementing two functions: Anchor box optimization: Dynamically fine-tunes the size and aspect ratio of preset anchor boxes based on the contextual information contained in the current feature map, making them better adaptable to the true shape of the target in the actual data (e.g., for typical canopy shapes of different tree species). Anchor box selection: Filters out a small number of high-quality candidate boxes most likely containing the target from a large number of possible anchor boxes, reducing the burden on subsequent prediction heads and improving the quality of positive and negative samples in the model. Input: Enhanced multi-scale feature maps output by BiFPN. Output: A series of optimized and filtered candidate regions (Proposals) generated for each layer of feature maps, along with their preliminary location information. Class Prediction Net outputs the probability distribution of each candidate box belonging to each tree species (and the background). Dynamic Anchor module outputs candidate regions and their corresponding features. Class confidence score for each anchor point. The Box Prediction Net predicts fine-tuned offsets relative to dynamic anchor points (such as center point coordinate shift, width and height scaling, etc.), similar to the Class Prediction Net, using candidate regions and features from the Dynamic Anchor module. The bounding box coordinates are for each anchor point. The complete end-to-end data flow is as follows: Input: 640×640 image → EfficientNet Backbone. Backbone outputs multi-scale features P3-P7 → BiFPN Layer. BiFPN outputs enhanced features → Dynamic Anchor Module. Dynamic Anchor outputs optimized candidate regions → Parallel prediction heads (ClassPrediction Net & Box Prediction Net). Final outputs: Class probabilities from the Class Prediction Net, and precise bounding box coordinates from the Box Prediction Net.

[0151] An improved version of YOLOv11 was pre-trained on our dataset and fine-tuned using our multimodal forest dataset. Training followed a curriculum-based learning approach, starting with detecting isolated canopies and progressively introducing synthetic images with increasingly overlapping canopies. Data augmentation techniques were applied, including photometric distortion, synthetic occlusion patterns simulating canopy gaps, and spatial transformations that preserve ecological validity to enhance robustness. The loss function combined a generalized intersection-union (GIoU) ​​term for bounding box regression with a hierarchical softmax classification loss, weighted by species frequency to address class imbalance.

[0152] 3. Results and Analysis:

[0153] 3.1 Visualization of generated image quality and denoising process:

[0154] To better understand the denoising process in our diffusion model, we Figure 6 The top panel visualizes the iterative refinement process from random noise to a coherent image. Each row represents a specific textual cue describing a different type of tree or an aerial view. The horizontal progression illustrates how the model progressively refines the noisy input into a high-quality image. For example, the right side of the first row shows the generation process of a ginkgo tree in autumn, starting with random noise and progressively refining the image to capture the golden canopy and detailed leaf structure. The right panels in rows 5-6 show that Chinese fir and oak species exhibit robust detail preservation under cloudy or low-light conditions, while the right panels in rows 7-8 highlight the ability of banyan and coconut trees to maintain clear texture generation even under intense sunlight.

[0155] Figure 6 A visual demonstration of the generated model. Figure 6 In the example (a), the text prompts guide the process of gradually denoising from random noise (t = 500) to a clear image (t = 0). Figure 6 In (b), the left side represents the tree species detection results from aerial photographs of Nanjing Forestry University and Shanghai Botanical Garden, and the right side represents the aerial image generated from the text description of the same species. Figure 6 (c) in the figure represents tree canopy samples generated under different conditions, and comparisons with GPT-4o, DALL-E 3 and Leonardo AI.

[0156] We first determined the optimal text cue weight to be 1.0, which effectively balances conditional fidelity and generation diversity. To qualitatively evaluate this trade-off, Figure 6Figure (d) shows 16 sample images generated by 8 different semantic cues under different weight configurations. The results show that unconditional generation (weight=0) produces coherent but semantically inconsistent samples; semantic fidelity is limited with a weight of 0.5 (FID=6.94, SSIM=0.90), while the optimal weight of 1.0 achieves a better balance (FID=6.42, SSIM=0.94). Although a weight of 1.5 slightly improves the metrics (FID=6.37, SSIM=0.96), it introduces subtle artifacts such as oversaturated foliage. (Note: FID stands for Fréchet Inception Distance, which measures the distance between the generated image and the real image distribution; a lower value indicates higher generation quality. SSIM is the structural similarity index, used to evaluate the structural similarity between the generated image and the reference image, ranging from 0 to 1; the closer to 1, the higher the consistency.)

[0157] To evaluate semantic controllability and cue robustness, we employed a classifier-free guidance (CFG) technique. This method achieves this by jointly training a language guidance conditional model and an unconditional diffusion model within the same network. We tested various text cue weights (0, 0.25, 0.5, 1, 1.5), consistent with current practices in the semantic guidance diffusion model field. (Note: CFG (Classifier-Free Guidance) is a technique that improves the alignment between generated results and text cue by jointly training a conditional model and an unconditional model. Text cue weights control the strength of the influence of text conditions during generation; higher weights result in stronger semantic consistency between the generated results and text cue, but may sacrifice diversity.)

[0158] The denoising process highlights the model's ability to integrate textual guidance, resulting in images that are both visually coherent and semantically highly consistent with the input prompts. For example... Figure 6 As shown in the right-middle panel, the model utilizes this text-guided synthesis capability to generate bird's-eye views that highly match the real data in the datasets from Nanjing Forestry University and Shanghai Botanical Garden. These synthesized images exhibit rich textures, accurate color distributions, and realistic structural details, and are seamlessly integrated into the corresponding tree species categories in the original datasets, naturally blending into the real ecological environment. Experimental results demonstrate that while consistently following semantic instructions, the model can robustly capture complex natural scene features—such as coordinated leaf distribution patterns, multi-layered canopy structures, and ecologically rational spatial layouts.

[0159] 3.2 Quantitative evaluation of generated images:

[0160] In the quantitative evaluation section, we employed a set of widely accepted comprehensive metrics to rigorously assess the quality of the generated images. These metrics collectively reflect the realism and diversity of the images from different perspectives. To objectively evaluate model performance, our model was compared with a series of benchmark models, including GAN, CycleGAN, BigGAN, StyleGAN, VAE, β-VAE, CVAE, DDPM, and the Denoising Diffusion Implicit Model (DDIM). The results are summarized in Table 2. Our model achieved an Initial Score (IS) of 9.89 on the forest image dataset, significantly outperforming the listed benchmark models and highlighting its ability to generate forest images with both clarity and diversity. Similarly, a Fraser Initial Distance (FID) score of 6.42 indicates that the images generated by the model are not only of excellent quality but also have a distribution that more closely resembles real forest scenes. Furthermore, a Structural Similarity Index (SSIM) of 0.94 demonstrates that the model not only reproduces the visual appearance of real forests but also preserves subtle structural features, achieving or surpassing the performance of contemporary mainstream methods. At the pixel level, a mean squared error (MSE) of 0.001 confirms the model's accuracy in reconstructing realistic forest images, while a peak signal-to-noise ratio (PSNR) of 30 dB highlights its robustness against distortion. These results collectively demonstrate that our model is a leader in generative modeling, exhibiting superior performance in pixel accuracy, structural consistency, and overall image quality.

[0161] We also compared our results with generative models such as GPT-4o, DALL-E 3, and Leonardo AI. Specific generation results are shown below. Figure 5 As shown in (c), GPT-4o generated two images in response to the prompt "an image of a green cedar canopy taken from a satellite perspective"—one with an incorrect perspective and the other with an incorrect number of trees. DALL-E 3 produced unrealistic and distorted results. While Leonardo's image was more realistic and had a higher resolution, it had inconsistencies in the number of trees and the shooting perspective. LFINet performed better in terms of fidelity and stability; however, due to our image generation size of 32×32 pixels, the results showed slight blurring.

[0162] Table 2 shows the performance evaluation results of different methods on this dataset under various evaluation metrics (the specific scores are shown below).

[0163] Model IS↑ FID↓ SSIM↑ MSE↓ PSNR↑ GT(s)↓ GFLOPs↑ Memory (MB) ↓ GAN 4.60 65.93 0.89 0.003 32 0.8 0.5 83 CycleGAN 7.02 15.03 0.72 0.008 26 1.2 1.2 156 BigGAN 9.22 14.73 0.86 0.004 28 1.5 3.5 257 StyleGAN 9.74 4.21 0.93 0.005 37 1.8 5.0 284 VAE 5.29 49.46 0.75 0.012 19 2.5 0.8 62 β-VAE 6.42 14.73 0.77 0.007 27 2.0 1.0 79 CVAE 8.22 21.7 0.83 0.005 27 2.2 1.5 101 DDPM 8.46 7.89 0.91 0.003 34 4.5 4.2 358 DDIM 8.87 7.54 0.92 0.002 38 4.0 3.8 343 Our 9.89 6.42 0.94 0.001 30 5.2 4.5 324

[0164] Note: GT: Generation time of a single image.

[0165] 3.3 Ablation Experiment:

[0166] To systematically evaluate the contributions of key components in the proposed diffusion model, we conducted ablation experiments on batch size, diffusion time steps, and the attention mechanism, focusing on their impact on generation quality and semantic consistency. Regarding batch size, we compared three settings—25, 50, and 100—to balance training stability and forest image fidelity; smaller batches can lead to training instability due to gradient noise, while larger batches may limit generalization ability due to excessive gradient smoothing. For diffusion time steps, we tested three configurations—500, 750, and 1000—to balance computational efficiency and ecological detail representation: shorter time steps reduce computational cost but may lose complex forest features, while longer time steps improve realism at a higher computational cost. For the attention mechanism, we evaluated the optimization effect of introducing a self-attention mechanism in the spatial transformation module on spatial coherence and image-text semantic alignment, a crucial mechanism for capturing forest-specific attributes such as species distribution and aerial layout patterns. These experiments quantify the independent contributions of each component, providing guidance for hyperparameter tuning to maximize visual fidelity and task-specific semantic consistency in aerial forest synthesis.

[0167] Ablation experiments confirm that the contributions of each component are significant: noise scheduling ensures sufficient granularity for high-quality denoising, positional encoding provides key temporal information, textual semantics aligns the generated image with semantic cues, self-attention mechanism enhances spatial and semantic coherence, and optimal batch size ensures efficient learning and generalization capabilities. Figure 7 Figure (a) shows how the loss changes with training epochs under different parameter settings. To further evaluate the effectiveness of different text encoder variants in LFINet, we conducted comparative experiments (results are shown in [link to results]). Figure 7 (b) covers four configurations: no text guidance, weak text guidance, original category labels (using only basic category labels as text guidance), and full text semantics (utilizing comprehensive text embeddings). Evaluation metrics include IS, FID, SSIM, MSE, and PSNR. The results show that as the text guidance mechanism is gradually improved, image quality and semantic consistency exhibit progressive enhancements, highlighting the advantages of the full text semantics method in enhancing the performance of LFINet forest applications. Figure 7 In the diagram, (a) represents the changing trend of the loss function of the generated model under different parameter configurations. Figure 7 (b) in the figure represents the ablation experiment results of the text encoder.

[0168] 3.4 Aerial Image Tree Species Detection Results:

[0169] Based on a proprietary dataset, we validated the detection performance on several state-of-the-art object detection models. Table 3 shows that our model achieves a highest mAP50 score of 0.868 while maintaining high efficiency with a computational cost of 21.4 GFLOPs, outperforming benchmark models such as YOLOv10s (0.864 mAP50, 24.7 GFLOPs) and YOLOv9s (0.858 mAP50, 26.8 GFLOPs).

[0170] We further validated this through targeted testing experiments across multiple scenarios. Figure 6 Aerial images of Nanjing Forestry University and Shanghai Botanical Garden (tree species and quantities are labeled in the left rectangle), and YOLOv11s were used to detect and analyze six typical areas (results are shown in...). Figure 8 These areas include pure stands of camellia, flame trees, poplar, and pine, as well as aerial photographs from Nanjing Forestry University and the Camp Grande urban area. Figure 8 The self-attention heatmap in the model highlights the region of interest, visually illustrating its detection mechanism. Results show that this method achieves accurate and interpretable canopy detection in various environments, from homogeneous plantations to complex urban forests. Figure 9 The test results show the robust performance of the model under different vegetation types and complexities: in pure forest detection, the detection accuracy of the four tree species (camellia oleifera, flame tree, poplar, and pine) all exceeded 0.87; in complex scene processing, the accuracy of the aerial image of Nanjing Forestry University containing mixed tree species was still maintained at 0.78; in urban forestry applications, the accuracy of 0.81 was achieved in the detection of Camp Grande urban area, demonstrating good adaptability to urban environment.

[0171] Table 3 shows the performance comparison of different detection models on this dataset:

[0172]

[0173] Note: FR-CNN: Faster R-CNN, MR-CNN: Mask R-CNN, EDet-D7: EfficientDet-D7, SAnything: Segment Anything, mAP50 represents the average precision when the intersection-union ratio threshold is 50% (higher is better); GFLOPs: billion floating-point operations per second; model size is in megabytes.

[0174] Figure 8 The improved YOLOv11 detection process is demonstrated, revealing the distribution of the model's region of interest during the detection process through the annotation of tree canopy detection bounding boxes of input / output image pairs and self-attention heatmaps.

[0175] 4. Discussion:

[0176] This research is based on a core assumption: structured ecological knowledge encoded through linguistic descriptions can guide probabilistic image generation, overcoming challenges such as inter-species similarity, environmental disturbances, and phenological changes. To achieve this goal, we propose LFINet, an innovative multimodal framework that leverages domain-specific text from a forestry corpus to guide a diffusion-based synthesis process. The framework's core employs a stable diffusion architecture based on Markov chain iterative denoising, continuously optimizing the synthesized image under text-derived semantic constraints. This method integrates a large-scale language model to parse complex and ambiguous botanical descriptions into dense ecological embeddings, effectively bridging the gap between linguistic knowledge and visual generation. Experimental results show that LFINet-generated tree images possess excellent realism (SSIM: 0.94; FID: 6.42), with the downstream detection module achieving an mAP50 of 0.868, outperforming baseline models. These findings highlight the potential of this method, contrasting sharply with the limitations of current mainstream data augmentation paradigms.

[0177] Previous research in this field has typically employed fusion strategies, such as combining GANs with CNNs or VAEs with Transformer architectures. For example, methods based on StyleGAN or BigGAN are combined with detection frameworks like Faster R-CNN to expand training datasets and improve species classification accuracy; similarly, VAEs are combined with CNNs to support detection tasks while modeling the complexity of forest structure. However, these fusion methods have significant drawbacks: GANs are prone to mode collapse and training instability, leading to blurred textures and semantic inconsistencies in synthesized images, which in turn affect downstream detection performance; VAEs, due to their inherent fidelity trade-offs, often produce outputs with insufficient sharpness, resulting in semantic drift and mismatch between generated images and ecological backgrounds. Furthermore, these fusion frameworks lack collaborative optimization strategies, limiting their effectiveness in handling dense canopies and tree occlusion, where detection performance typically degrades significantly.

[0178] While LFINet outperforms baseline GAN and VAE models in image fidelity (e.g., IS value of 9.89 vs. StyleGAN's 9.74), it still has several limitations that require improvement. The inherent iterative denoising process of the diffusion model leads to significant computational overhead: as shown in Table 2, on an A100 GPU, LFINet requires 5.2 seconds, 4.5 GFLOPs, and 324 MB of memory to process a single image, while StyleGAN requires only 1.8 seconds, 5.0 GFLOPs, and 280 MB, representing a 2.9-fold increase in latency. This computational characteristic currently hinders the deployment of its optimized version on embedded UAV chips, making real-time monitoring at the plot or regional scale difficult. Secondly, although CLIP improves ecological relevance, subtle semantic gaps remain when converting complex ecological descriptions into visual output, particularly noticeable when dealing with globally heterogeneous forest scenes. Thirdly, the framework lacks iterative language interaction capabilities, limiting the possibility of adjusting ecological parameters during scene generation. To address these limitations, our future research will focus on the following directions: developing lightweight variants using model compression techniques such as knowledge distillation; integrating taxonomic-ecological knowledge graphs to enhance semantic grounding; and designing multi-turn interaction frameworks to eliminate semantic gaps in the translation from ecological descriptions to visual output. Furthermore, we will introduce self-supervised learning strategies such as mask autoencoding and contrastive image-text alignment to improve ecological coherence while reducing reliance on labeled data.

[0179] 5. Conclusion:

[0180] This study proposes an innovative framework that integrates a forest-specific language model, a text-guided diffusion model, and a detection network to achieve semantically condition-controlled forest image synthesis and accurate tree species identification. The forest language model encodes ecological knowledge to guide the diffusion process, generating realistic forest scenes that are then used to train an improved YOLOv11 detector. Through image-text embedding alignment under the CLIP paradigm, this method ensures a high degree of semantic consistency between the generated visual content and the input description. The improved U-Net structure progressively optimizes the output image, achieving a PSNR of 30 dB, while the detection model achieves an mAP of 0.868. Experimental results demonstrate that the framework exhibits robust performance in complex canopy structures and species-rich environments, with its robustness and generalization ability benefiting from a multi-source training corpus. Future research will focus on expanding the model's adaptability to different geographical and ecological scenarios and exploring the potential of integrating large-scale language models to generate tree samples at a global scale.

[0181] The scope of protection of this invention includes, but is not limited to, the above embodiments. The scope of protection of this invention is defined by the claims. Any substitutions, modifications, or improvements to this technology that are easily conceived by those skilled in the art fall within the scope of protection of this invention.

Claims

1. An image diffusion model for generating realistic tree samples and identifying tree species in aerial photography, characterized in that, Includes the following steps: Step 1: Collect forestry-related text semantic datasets and aerial image datasets containing tree canopies. Pair each tree canopy image in the aerial image dataset with a text semantic information annotation in the text semantic dataset to form a multimodal dataset. Step 2: Train an image-text comparison model using a multimodal dataset, which includes an image encoder and a text encoder. Step 3: Construct a diffusion model, which includes a noise-adding module, a time-step embedding module, and a Unet denoising network model. Use the text embedding tensor output by the trained text encoder, the high-dimensional vector output by the time-step embedding module, the noise image output by the noise-adding module, and the real noise to train the Unet denoising network model, and obtain the trained Unet denoising network model. Step 4: Encode a batch of text into a text embedding tensor containing contextual semantics through the trained text encoder. This text embedding tensor, the Gaussian noise image, and the high-dimensional vector identifying the current denoising stage output by the time-step embedding module are injected into the trained diffusion model's Unet denoising network model, outputting the predicted noise, and finally producing the final image reconstruction aligned with the semantic description of the input text. Step 5: Combine the individual tree canopy images obtained in Step 4 to form a forest image; Step 6: Construct a YOLOv11 detection model by training the YOLOv11 detection model using the synthesized forest image and the tree species label information corresponding to each tree in the forest image. Step 7: Input a forest image of the tree species to be identified into the trained YOLOv11 detection model, and output the tree species corresponding to each tree in the forest image.

2. The image diffusion model according to claim 1 for generating realistic tree samples and identifying tree species in aerial photography, characterized in that, Step 2 is as follows: Step 2.1: Construct an image-text comparison model, which includes a text encoder and an image encoder; the image encoder uses Vision Transformer; the text encoder includes the Word2Vec model and a Transformer encoder. Step 2.2, the image encoder's processing procedure is as follows: the input image The image is divided into multiple image blocks, which are then combined with positional encodings and fed into the Transformer architecture to generate image embeddings. ; The text encoder's processing procedure is as follows: Input text semantic information annotation Byte-to-byte encoding (BPE) is used for word segmentation, followed by text cleaning. The cleaned tokens are then mapped to a high-dimensional space using a pre-trained Word2Vec model. Each token is converted into a 512-dimensional vector, generating a text embedding matrix. This matrix is ​​then processed by a Transformer encoder with eight attention heads, outputting a text embedding tensor that includes contextual semantics. Each sentence in the batch is represented by a single 512-dimensional vector; Step 2.3, Image Embedding With sentence embedding tensor The similarity matrix is ​​formed by combining the elements. Each row in the similarity matrix corresponds to an image embedding, and each column corresponds to a text embedding. The diagonal elements represent true pairings, and the off-diagonal elements correspond to incorrect pairings. Step 2.4: Train the image-text comparison model using the multimodal dataset according to steps 2.2-2.3, use symmetric cross-entropy loss to optimize the similarity score, and update the image encoder and text encoder.

3. The image diffusion model according to claim 1 for generating realistic tree samples and identifying tree species in aerial photography, characterized in that, The Unet denoising network model in step 3 includes three encoder layers, one bridging layer, three decoder layers, and a convolutional layer. The encoder layer includes two parallel groups and one downsampling layer. Each parallel group includes a ResBlock for processing temporal embeddings and image features, and a spatial Transformer module for integrating text embeddings through cross-attention. The bridging layer includes three ResBlocks, one spatial Transformer module, and three ResBlocks arranged in sequence. The decoder layer includes an upsampling layer and two parallel groups, each containing a ResBlock and a spatial Transformer module. The encoder layer and decoder layer are skip connections.

4. The image diffusion model according to claim 1 for generating realistic tree samples and identifying tree species in aerial photography, characterized in that, The processing procedure of the noise-adding module in step 3 is as follows: The forward process modifies an original image by progressively adding Gaussian noise. Convert to standard normally distributed noise; Its conditional probability is expressed as: (1); in, Indicates from arrive Image sequences with added noise Indicates that the i-th image has passed through... Images with Gaussian noise added at each time step; For hyperparameters, Describe an identity matrix. Indicates the image dimension; The recursive property allows it to be further expressed as: (2); Indicates from image arrive Added noise; conditional probability Represented as a Gaussian distribution: (3); conditional probability Represented as a Gaussian distribution: (4)。 5. The image diffusion model according to claim 1 for generating realistic tree samples and identifying tree species in aerial photography, characterized in that, The processing procedure of the time step embedding module in step 3 is as follows: Time step The time-step embedding module processes the data and applies a sinusoidal embedding function to generate a position encoding matrix. This matrix is ​​then processed by a linear layer to expand its dimensions, resulting in a high-dimensional vector. This high-dimensional vector is then fed into each ResBlock in the Unet denoising network model.

6. The image diffusion model according to claim 3 for generating realistic tree samples and identifying tree species in aerial photography, characterized in that, The loss function of the Unet denoising network model in step 3 is: (5); in, Expressing expectations, For model parameters From noisy images and text embedding Predicted noise, This is real noise.

7. The image diffusion model according to claim 1 for generating realistic tree samples and identifying tree species in aerial photography, characterized in that, The YOLOv11 detection model in step 6 includes an EfficientNet backbone network, a Bidirectional Feature Pyramid Network (BiFPN), a category prediction network, a bounding box prediction network, and a Dynamic Anchor module.

Citation Information

Cited By

  • Interpretable text semantic driving time sequence generation method based on Diffusion Transform model

    CN121981127A

  • A safety production abnormal image generation method and system based on text and images

    CN122289851A