Coastal zone mangrove biomass prediction method and device, storage medium and equipment

By integrating visual features, textual semantics, and knowledge graph information from high-resolution remote sensing images, this study addresses the issues of low accuracy and insufficient robustness of traditional remote sensing methods in mangrove biomass prediction, achieving high-precision biomass prediction in complex environments.

CN121095740BActive Publication Date: 2026-04-28GUANGDONG LABORATORY OF SOUTHERN OCEAN SCIENCE AND ENGINEERING (GUANGZHOU)
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGDONG LABORATORY OF SOUTHERN OCEAN SCIENCE AND ENGINEERING (GUANGZHOU)
Filing Date
2025-11-10
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Traditional remote sensing methods are subject to low accuracy and insufficient robustness when monitoring and predicting mangrove biomass in coastal zones due to tidal inundation, canopy saturation, and interference from complex ecosystems, and they are also highly dependent on a single remote sensing mode.

Method used

By combining the visual features, textual semantics, and knowledge graph information of high-resolution remote sensing images, a multimodal feature vector of image data is obtained through a pre-trained image feature extraction model, a visual language model, and a knowledge graph triple generation model. Biomass prediction is then performed using a multilayer perceptron.

Benefits of technology

It significantly improves the accuracy and stability of mangrove biomass prediction, accurately reflects biomass differences in complex environments, adapts to different regions and imaging conditions, and enhances the reliability of prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121095740B_ABST
    Figure CN121095740B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of coast zone mangrove biomass prediction method, device, storage medium and equipment, by obtaining the image data of target area and the text description information of the image data, based on pre-trained image feature extraction model, the visual feature vector of image data is extracted, based on normalization vegetation coefficient and image data, obtain spectral feature vector, based on text description information and knowledge graph information fusion obtains text feature vector, visual feature vector, spectral feature vector and text feature vector are spliced, obtain the fusion feature vector that multiple modal of vision, spectrum, text, knowledge graph is fused, then based on fusion feature vector, using the biomass prediction model trained obtains predicted biomass, improve the accuracy of coast zone mangrove biomass prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biomass prediction, and in particular to a method, apparatus, storage medium and device for predicting biomass in coastal mangroves. Background Technology

[0002] With the advent of global climate change and the goal of "carbon neutrality," coastal mangroves, as the core of the "blue carbon" ecosystem, play a significant role in carbon sequestration and storage. Therefore, accurate monitoring and dynamic prediction of their biomass are particularly important.

[0003] Traditional methods rely on remote sensing indices to monitor mangrove biomass in coastal zones. However, mangroves grow in the intertidal zone, and their spectral information is easily affected by periodic tidal inundation and turbid water background, resulting in low prediction accuracy. Summary of the Invention

[0004] This application provides a method, apparatus, storage medium, and device for predicting coastal mangrove biomass, which can improve the accuracy of coastal mangrove biomass prediction.

[0005] In a first aspect, embodiments of this application provide a method for predicting the biomass of coastal mangroves, including:

[0006] Acquire image data of the target area, textual description information of the image data, and knowledge graph triples of the textual description information;

[0007] Based on a pre-trained image feature extraction model, visual feature vectors are extracted from the image data;

[0008] Based on the normalized vegetation coefficient and the image data, spectral feature vectors are obtained;

[0009] A first text feature vector is obtained based on the text description information, a second text feature vector is obtained based on the knowledge graph triples, and a text feature vector is obtained based on the first text feature vector and the second text feature vector.

[0010] The visual feature vector, the spectral feature vector, and the text feature vector are concatenated to obtain a fused feature vector.

[0011] Based on the fused feature vector, the predicted biomass is obtained using a pre-trained biomass prediction model.

[0012] Secondly, embodiments of this application provide a coastal mangrove biomass prediction device, the device comprising:

[0013] The data acquisition module is used to acquire image data of the target area, text description information of the image data, and knowledge graph triples of the text description information;

[0014] The visual feature extraction module is used to extract visual feature vectors from the image data based on a pre-trained image feature extraction model.

[0015] The spectral feature acquisition module is used to acquire spectral feature vectors based on the normalized vegetation coefficient and the image data;

[0016] The text feature acquisition module is used to acquire a first text feature vector based on the text description information, acquire a second text feature vector based on the knowledge graph triples, and acquire a text feature vector based on the first text feature vector and the second text feature vector.

[0017] The fusion feature acquisition module is used to concatenate the visual feature vector, the spectral feature vector, and the text feature vector to obtain a fusion feature vector;

[0018] The biomass prediction module is used to obtain predicted biomass based on the fused feature vector using a pre-trained biomass prediction model.

[0019] Thirdly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the coastal mangrove biomass prediction method as described in any of the preceding claims.

[0020] Fourthly, embodiments of this application provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable by the processor;

[0021] When the processor executes the computer program, it implements the steps of the coastal mangrove biomass prediction method as described in any of the above.

[0022] In this embodiment, by acquiring image data of the target area and textual description information of the image data, a visual feature vector is extracted from the image data based on a pre-trained image feature extraction model. A spectral feature vector is obtained based on the normalized vegetation coefficient and the image data. A first textual feature vector is obtained based on the textual description information. A second textual feature vector is obtained based on the knowledge graph triples. A textual feature vector is obtained based on the first and second textual feature vectors. The visual feature vector, spectral feature vector, and textual feature vector are concatenated to obtain a fused feature vector that integrates multiple modalities including visual, spectral, textual, and knowledge graph features. Based on the fused feature vector, a pre-trained biomass prediction model is used to obtain predicted biomass, thereby improving the accuracy of coastal mangrove biomass prediction.

[0023] To better understand and implement this invention, the following detailed description is provided in conjunction with the accompanying drawings. Attached Figure Description

[0024] Figure 1 This is a flowchart of a method for predicting coastal mangrove biomass in one embodiment of the present invention;

[0025] Figure 2 This is a flowchart of step S120 in one embodiment of the present invention;

[0026] Figure 3 This is a schematic diagram of a coastal mangrove biomass prediction device according to one embodiment of the present invention;

[0027] Figure 4 This is a schematic diagram of the structure of a computer device according to one embodiment of the present invention. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0029] It should be understood that the described embodiments are merely some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of the embodiments of this application.

[0030] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the embodiments of this application. The singular forms “a,” “the,” and “the” used in the embodiments of this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0031] In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims. In the description of this application, it should be understood that the terms "first," "second," "third," etc., are used only to distinguish similar objects and are not necessarily used to describe a specific order or sequence, nor should they be construed as indicating or implying relative importance. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.

[0032] Furthermore, in the description of this application, unless otherwise stated, "several" refers to two or more. "And / or" describes the correspondence between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0033] Traditional methods for predicting and monitoring mangrove biomass in coastal zones typically employ remote sensing vegetation indices (such as NDVI and EVI). However, when these methods are applied to mangrove ecosystems with complex structures and unique environments, such as coastal zones, the following significant shortcomings exist:

[0034] 1. Remote sensing images are severely affected by tides and water bodies: Mangroves grow in the intertidal zone, and their spectral information is easily affected by periodic tidal inundation and turbid water background. This causes vegetation signals in remote sensing images to be confused with signals from water bodies and wet mudflats, resulting in serious deviations in the calculation results of traditional vegetation indices.

[0035] 2. Canopy closure and saturation issues: Mangrove vegetation is typically extremely dense, with high canopy closure. When monitoring high biomass areas, the traditional NDVI vegetation index easily reaches saturation, meaning the index value no longer changes sensitively with increasing biomass, thus failing to accurately reflect biomass differences within dense mangrove forests.

[0036] 3. Complex ecosystem and species structure: Mangrove ecosystems often exhibit characteristics of mixed tree species, diverse stand structures, and patchy distribution. Single spectral feature vectors are insufficient to capture subtle differences in biomass caused by factors such as tree species, canopy structure, and density, thus limiting monitoring accuracy.

[0037] 4. Traditional methods have severely inadequate predictive performance: Experiments and studies have shown that directly applying traditional methods such as NDVI to predict coastal mangrove biomass often results in low accuracy, and the predictive performance is insufficient for practical applications. For example, in some studies, the correlation coefficient ( The presence of outliers indicates a severe deficiency in its predictive ability.

[0038] When using deep learning models for remote sensing monitoring of mangroves, the performance of these models is highly dependent on a large quantity of high-quality, accurately labeled training samples. However, data collection in mangrove areas is difficult, leading to insufficient samples. Furthermore, semantic noise in remote sensing imagery (such as clouds, fog, and tidal changes) significantly impacts the model's generalization ability and stability.

[0039] Therefore, to address the aforementioned issues, this application provides a method for predicting mangrove biomass in coastal zones. This method can combine visual features, textual semantics, prior information from remote sensing indices, and knowledge graph information from high-resolution remote sensing images, thereby significantly improving the accuracy and stability of mangrove biomass prediction in complex coastal environments. It solves the limitations of existing technologies in predicting mangrove biomass in coastal zones, such as low accuracy, insufficient robustness, and over-reliance on a single remote sensing mode, due to interference from complex environmental factors like tidal inundation and canopy saturation.

[0040] Specifically, please refer to Figure 1 The coastal mangrove biomass prediction method of this application includes:

[0041] S110: Obtain image data of the target area, textual description information of the image data, and knowledge graph triples of the textual description information;

[0042] Image data can include multispectral imagery, high-resolution remote sensing imagery, and other imagery data. High-resolution remote sensing imagery refers to remote sensing imagery with a spatial resolution of 1 meter or higher.

[0043] In this embodiment, the image data may include Landsat images and high-resolution RGB image data. The high-resolution RGB image may be an RGB image with a spatial resolution of 1 meter or higher.

[0044] The text description information is a detailed text description generated based on the image data. The text description information is used to convert the visual elements in the image data into natural language. The text description information can include key objects, attributes and their spatial relationships in the scene.

[0045] S120: Extract the visual feature vector of the image data based on a pre-trained image feature extraction model;

[0046] Image feature extraction models are used to extract visual feature vectors from image data.

[0047] Image feature extraction models can be built from image recognition algorithms that are based on convolutional neural networks (CNN), ViT (Vision Transformer) models, etc., which can be used to extract visual feature vectors from image data.

[0048] S130: Based on the normalized vegetation coefficient and the image data, obtain the spectral feature vector;

[0049] The normalized vegetation coefficient is highly correlated with aboveground biomass (AGB), which can enhance the model's utilization of vegetation information and make up for the spectral details that may be missing in visual and text feature vectors. Therefore, in this embodiment, the spectral feature vector of image data is obtained by calculating the normalized vegetation coefficient.

[0050] In this embodiment of the application, the image data may include Landsat multispectral imagery, and the normalized vegetation coefficient is calculated using the reflectance of the near-infrared (NIR) and red (RED) bands.

[0051]

[0052] in, For spectral eigenvectors, This represents the normalized vegetation coefficient for pixel (x, y) in the image data. Let be the reflectance of the red (RED) band at pixel (x,y). Let be the reflectance of the pixel (x,y) in the near-infrared (NIR) band.

[0053] S140: Obtain a first text feature vector based on the text description information, obtain a second text feature vector based on the knowledge graph triplet, and obtain a text feature vector based on the first text feature vector and the second text feature vector;

[0054] Compared to using only visual and spectral features, text can describe advanced abstract concepts such as "canopy closure," "tidal inundation," and "species mixing," which are difficult to express directly with pixel-level visual features and spectral indices. This can provide models with a deeper understanding beyond pixels, thereby improving the accuracy of biomass prediction.

[0055] S150: Concatenate the visual feature vector, the spectral feature vector, and the text feature vector to obtain a fused feature vector;

[0056] Visual feature vectors are susceptible to the effects of lighting, shadows, and haze, while spectral feature vectors are susceptible to moisture. Textual descriptions can penetrate these low-level visual noises, grasp the "core semantics," and fuse features by combining the fine texture of high-resolution images, text, and physical indicators of spectral indices. This effectively avoids the information insufficiency problem of relying solely on a single data source such as spectroscopy or vision, and significantly improves the accuracy of biomass prediction.

[0057] S160: Based on the fused feature vector, obtain the predicted biomass using a pre-trained biomass prediction model.

[0058] In this embodiment, by acquiring image data of the target area and textual description information of the image data, a visual feature vector is extracted from the image data based on a pre-trained image feature extraction model. A spectral feature vector is obtained based on the normalized vegetation coefficient and the image data. A textual feature vector is extracted based on the textual description information. The visual feature vector, spectral feature vector, and textual feature vector are concatenated to obtain a fused feature vector that integrates multiple modalities of vision, spectroscopy, and text. Then, based on the fused feature vector, a pre-trained biomass prediction model is used to obtain the predicted biomass, thereby improving the accuracy of coastal mangrove biomass prediction.

[0059] In step S110, the text description information can be the text obtained by remote sensing experts interpreting the image data.

[0060] Alternatively, in another embodiment, obtaining the text description information of the image data and the knowledge graph triple of the text description information includes:

[0061] The image data is input into a visual language model to obtain the text description information generated by the visual language model.

[0062] The text description information is input into the knowledge graph triple generation model to obtain the knowledge graph triple output by the knowledge graph triple generation model.

[0063] The visual language model can be GPT-4o or other large-scale visual language models with image interpretation capabilities.

[0064] By inputting image data into a visual language model, the model simulates the interpretation process of remote sensing experts, transforming visual elements in the image into natural language to obtain textual descriptions of key objects, attributes, and spatial relationships in the scene, thereby improving the efficiency of obtaining textual description information.

[0065] For example, the text description of a mangrove image could be: "A dense mangrove forest grows along the edge of a turbid mudflat, with a dark green and highly closed canopy. Exposed supporting roots are visible in some areas, the water is turbid and yellowish-brown, and the sky is clear and cloudless."

[0066] The text description information includes at least the key objects, attributes, and spatial relationships of the remote sensing image, and the text feature vector can include features such as the key objects, attributes, and spatial relationships of the remote sensing image.

[0067] Knowledge graph triples can be used to describe various entities and the relationships between them. By forming triples of entities, relationships, and attributes, various types of knowledge can be clearly expressed. Knowledge graph triples can compensate for the information loss that may be caused by encoding long texts and capture the relationships between entities more accurately.

[0068] In this embodiment of the application, the knowledge graph triple may include key objects, attributes and their spatial relationships in the scene.

[0069] For example, when the text description is "The canopy is dark green and highly closed. Exposed supporting roots are visible in some areas", the corresponding knowledge graph triple can be "(canopy, attribute, dark green), (roots, state, exposed)".

[0070] Specifically, textual description information can be converted into structured knowledge graph triple information by using the FactualSceneGraph model or other knowledge graph triple generation models with the above-mentioned text transformation capabilities, thereby facilitating the subsequent application of textual description information.

[0071] FactualSceneGraph (FSG) is an open-source scene-graph parser that can automatically parse natural language text into structured triples in the form of "entity-attribute-relationship" to obtain a factual scene graph that can be directly used for downstream tasks. This allows text information to better serve downstream applications such as image retrieval, image description evaluation, and visual question answering.

[0072] In one embodiment, obtaining a first text feature vector based on the text description information includes:

[0073] Using a pre-trained first Transformer encoder, a first text feature vector is obtained based on the text description information;

[0074] The structure of the first Transformer encoder can be set according to the actual application scenario.

[0075] In this embodiment, the first Transformer encoder can be a 12-layer Transformer encoder, capturing information such as n-grams, syntax, semantics, and discourse of text description information sequentially from the lowest to the highest layer. The first Transformer encoder is based on CLIP pre-trained weights to improve the accuracy of the generated first text feature vector.

[0076] The second text feature vector is obtained based on the knowledge graph triples, including:

[0077] The knowledge graph triples are converted into knowledge graph vectors, and the knowledge graph vectors are input into the second Transformer encoder to obtain the second text feature vector.

[0078] Specifically, based on the triplet embedding method, knowledge graph triples can be converted into vector form to obtain knowledge graph vectors. The triplet embedding method uses an embedding function to convert knowledge graph triples into vector form, thus obtaining knowledge graph vectors.

[0079] The knowledge graph triple information is converted into a knowledge graph vector in the following manner:

[0080]

[0081] in, Let be the knowledge graph vector of the i-th triple. Let i be the head node, relation, and tail node of the i triples. For embedded functions;

[0082] The second Transformer encoder is a separate Transformer encoder, different from the first Transformer encoder.

[0083] The second Transformer encoder is used to perform contextual encoding on the knowledge graph vector and mean pooling along the time steps to obtain the second text feature vector.

[0084] Obtaining text feature vectors based on the first text feature vector and the second text feature vector includes:

[0085] The first text feature vector and the second text feature vector are added together to obtain the text feature vector.

[0086] Specifically, the text feature vector is obtained in the following manner:

[0087]

[0088] in, For text feature vectors, This is the first text feature vector. This is the second text feature vector.

[0089] While a single text vector contains rich descriptive information, it may fail to highlight key, structured facts during encoding and is prone to losing precise relationships between entities. Knowledge graph vectors, on the other hand, structure the scene and clarify objects and their relationships, but may lack the fluent context and detailed descriptions of text.

[0090] In this embodiment, text feature vectors are generated by fusing text vectors and knowledge graph vectors. This preserves the rich context of the text and strengthens the core semantic relationships through the knowledge graph, generating a more comprehensive and robust semantic representation than a single vector, which can effectively combat semantic noise.

[0091] like Figure 2 As shown, in one embodiment, the image data includes RGB high-resolution image data, which can be RGB images with a spatial resolution of 1 meter or higher.

[0092] The image feature extraction model can employ the ViT (Vision Transformer) model based on Contrastive Language-Image Pretraining (CLIP). Compared to traditional image feature extraction models, this model features zero-shot capability, strong robustness, cross-modal compatibility, and easy transferability.

[0093] Extracting the visual feature vectors from the image data includes:

[0094] S121: Divide the RGB high-resolution image data into several non-overlapping image blocks, and convert each image block into a one-dimensional vector sequence;

[0095] Specifically, the RGB high-resolution image data is divided into n non-overlapping image patches, and each image patch is flattened into a one-dimensional vector sequence. , where i = 1, 2, ..., n. i is the index of the image patch.

[0096] S122: Perform a linear projection on each of the one-dimensional vector sequences, and generate an input sequence based on the one-dimensional vector sequences of the linear projection;

[0097] The input sequence includes learnable classification labels and learnable location embeddings;

[0098]

[0099] in, Given the input sequence, For classification labels, For location embedding, Let E be the one-dimensional vector sequence of the 1st, 2nd...nth linear projections, and let E be the linear projection matrix.

[0100] Spatial location information is introduced by adding learnable classification labels and learnable location embeddings to the sequence.

[0101] S123: Input the input sequence into the ViT model, obtain the feature vector output by the ViT model corresponding to the classification label, and use the feature vector corresponding to the classification label as the visual feature vector.

[0102] In this embodiment, the ViT model includes L encoder modules, each encoder module including a multi-head self-attention (MSA) layer and a multilayer perceptron (MLP). The number of L can be set according to the actual application.

[0103]

[0104]

[0105] in, The feature vector output by the l-th encoder module, The feature vector output by the (l-1)th encoder module, Let l be the intermediate vector, l=1, ..., L, LN be the layer normalization.

[0106] After the input sequence passes through all L encoder modules of the ViT model, the feature vector of the last layer output is extracted. In and classification tags The corresponding feature vector is used as the global visual feature vector of the image.

[0107] In this embodiment, RGB high-resolution image data is divided into several non-overlapping image blocks, each image block is converted into a one-dimensional vector sequence, each one-dimensional vector sequence is linearly projected, an input sequence is generated based on the linearly projected one-dimensional vector sequence, a learnable classification label and a learnable position embedding are added to the sequence to introduce spatial position information, the input sequence is input into the ViT model, and by obtaining the feature vector output by the ViT model corresponding to the classification label, a global visual feature vector that can represent the entire image is obtained.

[0108] In one embodiment, the visual feature vector, the spectral feature vector, and the text feature vector are concatenated to obtain a fused feature vector, including:

[0109] The visual feature vector, the spectral feature vector, and the text feature vector are dimensionally aligned, and the dimensionally aligned visual feature vector, the spectral feature vector, and the text feature vector are concatenated to obtain a fused feature vector.

[0110] Optionally, the visual feature vector, spectral feature vector, and text feature vector can be dimensionally aligned by aligning their lengths and / or channels. For example, the above features can be mapped to features of the same length and / or the same number of channels, thereby facilitating feature fusion of multimodal features.

[0111] Specifically, the visual feature vector, the spectral feature vector, and the text feature vector are concatenated in the following manner:

[0112] .

[0113] in, To fuse feature vectors, For visual feature vectors, For spectral eigenvectors, Here, represents the text feature vector, and Concat is the concatenation function.

[0114] By concatenating visual feature vectors, spectral feature vectors, and text feature vectors, efficient integration of multi-source information from visual, spectral, and text sources is achieved. Furthermore, the concatenation preserves the original information from multiple modalities, thereby improving the accuracy of prediction.

[0115] In one embodiment, the biomass prediction model can be a model trained using a preset mangrove biomass dataset, wherein the mangrove biomass dataset may include remote sensing images and corresponding ground-measured biomass data.

[0116] In this embodiment, the biomass prediction model is a multilayer perceptron, which includes two linear transformation layers; based on the fused feature vector, the predicted biomass is obtained using a pre-trained biomass prediction model, including:

[0117] The fused feature vector is input into the multilayer perceptron, and the fused feature vector is mapped into a one-dimensional feature vector through two linear transformation layers of the multilayer perceptron to obtain the predicted biomass.

[0118] In this embodiment of the application, the fused features can be high-dimensional features of 1025 dimensions. The mapping from high dimension to one dimension is achieved through two linear transformation layers of a multilayer perceptron, and the output one-dimensional feature vector is the predicted biomass.

[0119] The activation function of a multilayer perceptron can be an existing activation function such as GELU or ReLU.

[0120] Preferably, the activation function of the multilayer perceptron in this embodiment is the Gaussian error linear unit activation function (GELU). The GELU is a commonly used activation function in deep learning, which activates the input data by multiplying it by the cumulative distribution function of a Gaussian distribution. The gradient of the GELU near the origin is non-zero, which helps reduce the gradient vanishing problem during training. Furthermore, the derivative of the GELU is relatively smooth and uninterrupted, making it easier to compute during backpropagation. Introducing the GELU can enhance the nonlinear expressive power of the model.

[0121] The multilayer perceptron further includes a normalization layer; the fused feature vector is mapped to a one-dimensional feature vector through two linear transformation layers of the multilayer perceptron, including:

[0122] The normalization layer is used to normalize the feature vector output by the linear transformation layer.

[0123] The normalization layer can normalize the features output by the linear transformation layer, so that the features output by each linear transformation layer are on the same scale, thereby accelerating model convergence.

[0124] This application's solution effectively integrates the fine texture of high-resolution images, the high-level semantics of text and knowledge graphs, and the physical indication information of spectral indices. Through multimodal complementarity, it solves the problem of insufficient information or bias in complex scenes from single remote sensing data sources (such as relying solely on spectral or visual data), thereby significantly improving the overall accuracy of biomass prediction. Simultaneously, it innovatively utilizes large-scale visual language models to generate text descriptions and combines them with knowledge graphs for semantic enhancement. Compared to using only visual and spectral feature vectors, text can describe pixel-level visual features such as "canopy closure," "tidal inundation," and "species mixing," as well as high-level abstract concepts that are difficult to directly express with spectral indices, providing the model with a deeper understanding beyond pixels. In high-biomass, high-density mangrove areas, spectral indices such as NDVI are easily saturated. Semantic information generated from images, such as "rich canopy layers" and "no gaps in the understory," can provide the model with crucial discriminative capabilities, enabling it to accurately distinguish subtle differences within high-biomass regions. Visual features are easily affected by light, shadow, and fog, while spectral feature vectors are easily affected by moisture. Textual descriptions generated by advanced large-scale visual language models (VLMs) can penetrate low-level visual noise and capture core semantics, making the model's predictions more stable under varying imaging conditions. By adaptively fusing visual details, semantic relationships, and spectral priors, the model can effectively combat interference from single information sources (such as spectral saturation and illumination variations), exhibiting stronger adaptability and generalization capabilities to mangrove ecosystems in different regions, at different growth stages, and under different tidal conditions, making it more reliable in practical applications.

[0125] like Figure 3 As shown in the illustration, this application also provides a coastal mangrove biomass prediction device, comprising:

[0126] Data acquisition module 110 is used to acquire image data of the target area, text description information of the image data, and knowledge graph triples of the text description information;

[0127] The visual feature extraction module 120 is used to extract the visual feature vector of the image data based on a pre-trained image feature extraction model.

[0128] The spectral feature acquisition module 130 is used to acquire spectral feature vectors based on the normalized vegetation coefficient and the image data;

[0129] The text feature acquisition module 140 is used to acquire a first text feature vector based on the text description information, acquire a second text feature vector based on the knowledge graph triplet, and acquire a text feature vector based on the first text feature vector and the second text feature vector.

[0130] The fusion feature acquisition module 150 is used to concatenate the visual feature vector, the spectral feature vector, and the text feature vector to obtain a fusion feature vector;

[0131] The biomass prediction module 160 is used to obtain predicted biomass based on the fused feature vector using a pre-trained biomass prediction model.

[0132] In one embodiment, the text feature acquisition module 140 is used to acquire a first text feature vector based on the text description information using a pre-trained first Transformer encoder; convert the knowledge graph triples into knowledge graph vectors; input the knowledge graph vectors into a second Transformer encoder to acquire a second text feature vector; and add the first text feature vector and the second text feature vector to acquire a text feature vector.

[0133] In one embodiment, the data acquisition module 110 includes:

[0134] The text description information acquisition unit is used to input the image data into the visual language model and acquire the text description information generated by the visual language model.

[0135] The triplet acquisition unit is used to input the text description information into the knowledge graph triplet generation model and obtain the knowledge graph triplet output by the knowledge graph triplet generation model.

[0136] In one embodiment, the image data includes RGB high-resolution image data, the image feature extraction model is a CLIP-based ViT model, and the visual feature extraction module 120 includes:

[0137] The vector sequence acquisition unit is used to divide the RGB high-resolution image data into several non-overlapping image blocks and convert each image block into a one-dimensional vector sequence.

[0138] An input sequence generation unit is used to perform a linear projection on each of the one-dimensional vector sequences and generate an input sequence based on the linearly projected one-dimensional vector sequences; wherein, the input sequence includes a classification label and a position embedding;

[0139]

[0140] in, Given the input sequence, For classification labels, For location embedding, Let E be the one-dimensional vector sequence of the 1st, 2nd...nth linear projections, and let E be the linear projection matrix;

[0141] The visual feature acquisition unit is used to input the input sequence into the ViT model, obtain the feature vector output by the ViT model corresponding to the classification label, and use the feature vector corresponding to the classification label as the visual feature vector.

[0142] In one embodiment, the biomass prediction model is a multilayer perceptron, which includes two linear transformation layers; the biomass prediction module 160 includes:

[0143] The predicted biomass acquisition unit is used to input the fused feature vector into the multilayer perceptron, and map the fused feature vector into a one-dimensional feature vector through two linear transformation layers of the multilayer perceptron to obtain the predicted biomass.

[0144] In one embodiment, the activation function of the multilayer perceptron is a Gaussian error linear unit activation function, and the multilayer perceptron further includes a normalization layer; the biomass prediction module 160 further includes:

[0145] The normalization processing unit is used to normalize the feature vector output by the linear transformation layer using the normalization layer.

[0146] In one embodiment, the fusion feature acquisition module 150 includes:

[0147] The splicing unit is used to perform dimensional alignment on the visual feature vector, the spectral feature vector, and the text feature vector, and to splice the dimensionally aligned visual feature vector, spectral feature vector, and text feature vector to obtain a fused feature vector.

[0148] It should be noted that the coastal mangrove biomass prediction device provided in the above embodiments is only illustrated by the division of the above functional modules when executing the coastal mangrove biomass prediction method. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the coastal mangrove biomass prediction device and the coastal mangrove biomass prediction method provided in the above embodiments belong to the same concept, and the implementation process is detailed in the method embodiments, which will not be repeated here.

[0149] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the coastal mangrove biomass prediction method as described in any of the above embodiments.

[0150] The embodiments of this application may take the form of a computer program product implemented on one or more storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing program code. Computer-readable storage media include permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information may be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to: phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0151] like Figure 4 As shown, this application embodiment also provides a computer device 200, including a memory 210, a processor 220, and a computer program stored in the memory 210 and executable by the processor 220;

[0152] When the processor 220 executes the computer program, it implements the steps of the coastal mangrove biomass prediction method as described in any of the above.

[0153] The memory 210 includes read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, disk storage, magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.

[0154] The processor 220 is the control unit of the computer device 200. It connects to various components of the computer device 200 via various interfaces and lines. By running or executing programs or modules stored in the memory 210, and by calling data stored in the memory 210, it performs various functions of the computer device 200 and processes data. For example, when the processor 220 executes the computer program stored in the memory 210, it implements all or part of the steps of the coastal mangrove biomass prediction method described in this application embodiment; or it implements all or part of the functions of the coastal mangrove biomass prediction device. The processor 220 can be composed of integrated circuits, such as a single packaged integrated circuit, or multiple integrated circuits with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips.

[0155] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0156] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for predicting mangrove biomass in coastal zones, characterized in that, include: The system acquires image data of the target area, textual description information of the image data, and knowledge graph triples of the textual description information; wherein, the textual description information is used to describe key objects, attributes, and spatial relationships in the image data scene, and the knowledge graph triples include key objects, attributes, and their spatial relationships; Based on a pre-trained image feature extraction model, visual feature vectors are extracted from the image data; Based on the normalized vegetation coefficient and the image data, spectral feature vectors are obtained; A first text feature vector is obtained based on the text description information, a second text feature vector is obtained based on the knowledge graph triples, and a text feature vector is obtained based on the first text feature vector and the second text feature vector. The visual feature vector, the spectral feature vector, and the text feature vector are concatenated to obtain a fused feature vector. Based on the fused feature vector, the predicted biomass is obtained using a pre-trained biomass prediction model. Obtaining text feature vectors based on the first text feature vector and the second text feature vector includes: The first text feature vector and the second text feature vector are added together to obtain the text feature vector.

2. The method for predicting coastal mangrove biomass according to claim 1, characterized in that, Obtaining the text description information of the image data and the knowledge graph triples of the text description information includes: The image data is input into a visual language model to obtain the text description information generated by the visual language model. The text description information is input into the knowledge graph triple generation model to obtain the knowledge graph triple output by the knowledge graph triple generation model.

3. The method for predicting coastal mangrove biomass according to claim 1, characterized in that, Obtaining the first text feature vector based on the text description information includes: Using a pre-trained first Transformer encoder, a first text feature vector is obtained based on the text description information; The second text feature vector is obtained based on the knowledge graph triples, including: The knowledge graph triples are converted into knowledge graph vectors, and the knowledge graph vectors are input into the second Transformer encoder to obtain the second text feature vector.

4. The method for predicting coastal mangrove biomass according to claim 1, characterized in that, The image data includes RGB high-resolution image data, and the image feature extraction model is a CLIP-based ViT model, which extracts visual feature vectors from the image data, including: The RGB high-resolution image data is divided into several non-overlapping image blocks, and each image block is converted into a one-dimensional vector sequence; A linear projection is performed on each of the one-dimensional vector sequences, and an input sequence is generated based on the one-dimensional vector sequences of the linear projection; wherein, the input sequence includes a classification label and a location embedding; in, Given the input sequence, For classification labels, For location embedding, Let E be the one-dimensional vector sequence of the 1st, 2nd...nth linear projections, and let E be the linear projection matrix; The input sequence is input into the ViT model to obtain the feature vector output by the ViT model corresponding to the classification label, and the feature vector corresponding to the classification label is used as the visual feature vector.

5. The method for predicting coastal mangrove biomass according to claim 1, characterized in that, The biomass prediction model is a multilayer perceptron, which includes two linear transformation layers; based on the fused feature vector, the predicted biomass is obtained using the pre-trained biomass prediction model, including: The fused feature vector is input into the multilayer perceptron, and the fused feature vector is mapped into a one-dimensional feature vector through two linear transformation layers of the multilayer perceptron to obtain the predicted biomass.

6. The method for predicting coastal mangrove biomass according to claim 5, characterized in that, The activation function of the multilayer perceptron is the Gaussian error linear unit activation function, and the multilayer perceptron also includes a normalization layer; the fused feature vector is mapped to a one-dimensional feature vector through two linear transformation layers of the multilayer perceptron, and the mechanism further includes: The normalization layer is used to normalize the feature vector output by the linear transformation layer.

7. The method for predicting coastal mangrove biomass according to claim 1, characterized in that, The visual feature vector, the spectral feature vector, and the text feature vector are concatenated to obtain a fused feature vector, including: The visual feature vector, the spectral feature vector, and the text feature vector are dimensionally aligned, and the dimensionally aligned visual feature vector, the spectral feature vector, and the text feature vector are concatenated to obtain a fused feature vector.

8. A device for predicting the biomass of coastal mangroves, characterized in that, The device includes: The data acquisition module is used to acquire image data of the target area, text description information of the image data, and knowledge graph triples of the text description information; wherein, the text description information is used to describe key objects, attributes, and spatial relationships in the image data scene, and the knowledge graph triples include key objects, attributes, and their spatial relationships; The visual feature extraction module is used to extract visual feature vectors from the image data based on a pre-trained image feature extraction model. The spectral feature acquisition module is used to acquire spectral feature vectors based on the normalized vegetation coefficient and the image data; The text feature acquisition module is used to acquire a first text feature vector based on the text description information, acquire a second text feature vector based on the knowledge graph triples, and acquire a text feature vector based on the first text feature vector and the second text feature vector; the text feature acquisition module is used to add the first text feature vector and the second text feature vector to acquire the text feature vector. The fusion feature acquisition module is used to concatenate the visual feature vector, the spectral feature vector, and the text feature vector to obtain a fusion feature vector; The biomass prediction module is used to obtain predicted biomass based on the fused feature vector using a pre-trained biomass prediction model.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When executed by a processor, the computer program implements the steps of the coastal mangrove biomass prediction method as described in any one of claims 1-7.

10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable by the processor; When the processor executes the computer program, it implements the steps of the coastal mangrove biomass prediction method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Prediction method and system for invasive plant biomass and storage medium

    CN118799758A

  • Endoscopic image description report generation method and device and medium

    CN119724464A

  • Remote sensing monitoring method and system for grassland biomass

    CN119964037A