Morphology and taxonomy-based few-sample coral intelligent identification method

By combining morphology and taxonomy adapters to optimize the CLIP model, the problem of insufficient generalization ability of coral recognition methods in complex marine environments is solved, efficient coral species recognition is achieved, and computing resource requirements are reduced.

CN120408325APending Publication Date: 2025-08-01SOUTH CHINA SEA PLANNING & ENVIRONMENT RES INST SOA +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510896457.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The existing coral recognition methods have insufficient generalization capabilities, making it difficult to accurately identify unknown coral species in complex marine environments, and the existing pre-trained models lack appropriate fine-tuning strategies for coral image classification, and the computing resource requirements are high.

Method used

The intelligent recognition method of corals based on morphology and taxonomy is adopted. The graphic and text data of corals are processed through CLIP text encoder and visual encoder. Combined with morphology and taxonomy adapters, the JS loss and similarity loss optimization model is used to extract the visual features of corals, and the texture projection layer is used for optimization to achieve small sample recognition.

Benefits of technology

The coral species recognition performance is improved, and a series of underwater coral species recognition can be achieved by only a small number of parameters and pictures, which improves the generalization ability of the model in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408325A_ABST
    Figure CN120408325A_ABST
Patent Text Reader

Abstract

The invention discloses a few-sample coral intelligent identification method based on morphology and taxonomy, and relates to the field of coral identification, and the method comprises the steps: obtaining image-text data of coral samples; training a few-sample coral intelligent identification model according to the image-text data of the coral samples to obtain a trained few-sample coral intelligent identification model; and according to the image-text data of the coral to be identified, coral species are identified based on the trained few-sample coral intelligent identification model. According to the method, knowledge enhancement is carried out on the vision-language large model by utilizing morphology and taxonomy knowledge of the coral, so that the recognition performance of coral species is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of coral recognition, and particularly to a few-shot coral intelligent recognition method based on morphology and taxonomy. Background Art

[0002] Coral reefs are one of the most biodiverse ecosystems on Earth, providing habitat for over a quarter of all marine species, including fish, mollusks, and crustaceans. Unfortunately, coral reefs are under significant threat from human activities and natural challenges such as global warming and ocean acidification. Studies have shown that over 70% of the world's coral reefs are in poor health. If current trends continue, all of the world's coral reefs could be bleached by the end of the century. To provide early warning, assess harmful factors, and ensure the resilience of coral reef ecosystems, it is necessary to continuously monitor their benthic communities. Accurate coral recognition is crucial in this monitoring process, so scientists use annotated coral data to statistically analyze coral habitat coverage, spatial relationships, and health status.

[0003] With the advancement of deep learning, computer vision, and autonomous underwater vehicle technologies, automated coral recognition has been increasingly applied to underwater coral monitoring. Although state-of-the-art methods have shown commendable accuracy, they are typically designed to predict a specific, pre-set group of coral species. Therefore, generalization ability has become a major obstacle to the application of artificial intelligence in coral reef conservation. The complex marine environment (e.g., factors such as temperature, light, and salinity) results in significant morphological differences even among the same coral species. These significant morphological differences among coral species mainly rely on expert knowledge, making it difficult for existing classification methods to generalize to unknown coral species in uncharted waters.

[0004] In the prior art, vision-language models (VLMs) pre-trained on hundreds of millions of pairs of image-text data, such as Contrastive Language-Image Pretraining (CLIP) and ALIGN Visual-Language Model (ALIGN), have demonstrated excellent generalization ability in cross-domain alignment of image and text modalities. They enable zero-shot or few-shot learning in a wide range of image classification tasks. However, despite their excellent performance in the general online image domain, due to the unique characteristics of corals, they cannot be directly used for coral image classification, mainly lacking a suitable fine-tuning strategy. Traditional methods of adapting pre-trained models to downstream tasks usually require fine-tuning the entire model, which has high computational resource requirements. The visual features generated by the emerging Contrastive Language-Image Pretraining (CLIP)-adapter model, a powerful tool for enhancing vision-language models based on parameter-efficient fine-tuning (PEFT), usually cannot handle the complex morphological features of coral reefs and their rich classification hierarchies.

[0005] Therefore, to solve the above problems, there is an urgent need for a few-shot coral intelligent recognition method based on morphology and taxonomy, which can improve the performance of coral species recognition. Summary of the Invention

[0006] The purpose of this application is to provide a few-shot coral intelligent recognition method based on morphology and taxonomy, which can improve the performance of coral species recognition.

[0007] To achieve the above purpose, this application provides the following solutions: This application provides a few-shot coral intelligent recognition method based on morphology and taxonomy, including: Obtain the graphic and text data of coral samples; the graphic and text data includes: coral images, morphological text prompts of corals, and taxonomic text prompts of corals; Train a few-shot coral intelligent recognition model according to the graphic and text data of coral samples to obtain a trained few-shot coral intelligent recognition model; the few-shot coral intelligent recognition model includes: a CLIP text encoder, a CLIP visual encoder, a morphological adapter, and a taxonomic adapter; Identify coral species based on the trained few-shot coral intelligent recognition model according to the graphic and text data of the coral to be identified; The training process of the few-shot coral intelligent recognition model is: Process the morphological text prompt and taxonomic text prompt of corals based on the CLIP text encoder to obtain the morphological text encoding representation and taxonomic text encoding representation of corals respectively; Process the coral image based on the CLIP visual encoder to obtain the visual encoding representation of the coral; Process the visual encoding representation of the coral based on the taxonomic adapter and morphological adapter to obtain the taxonomic visual features and morphological visual features of the coral; and optimize the few-shot coral intelligent recognition model by constraining the JS loss and similarity loss of the taxonomic adapter and morphological adapter; According to the morphological visual features of the coral output by the morphological adapter, adopt a texture projection layer and optimize it through CE loss to obtain a trained few-shot coral intelligent recognition model.

[0008] Optionally, the taxonomic adapter includes: a bottleneck layer and a residual layer.

[0009] Optionally, the morphological visual features of the coral include: the texture of the coral, the structure of the coral, and the color of the coral.

[0010] Optionally, the morphological adapter specifically includes: Use the formula Calculate the outer product of the feature vectors between two different layers in the morphological adapter of the few-shot coral intelligent recognition model ; Use the formula Max-pool the outer product of the feature vectors ; Highlight the texture differences between different coral species according to the max-pooled outer product of the feature vectors; Use the formula to obtain the global structure information of the coral based on the coral image ; Distinguish coral species according to the global structure information of the coral and the correlation between structures; wherein, is the transpose of the visual feature matrix, is a learnable projection matrix, is the outer product of the feature vectors between two different layers in the morphological adapter of the few-shot coral intelligent recognition model, is the max-pooling function; is the feature extraction function, is the coral structure building block, is the weight, is the projection block for obtaining channel dependencies, capturing the subtle changes in the color of the coral, is the number of coral images grouped based on structure location, j It is j Set of coral images, is the last step of the CLIP visual encoder i Visual features of the layer.

[0011] Alternatively, using the formula 、 and Calculating predicted probabilities for morphological adapters ; Using the formula 、 and Calculating predicted probabilities for taxonomic adapters ; in, is a given image, It's a category, are learnable text encoding parameters, is a fine-tunable trade-off parameter, is the morphological text encoding representation of coral, It is i Morphological text encoding representation of coral categories, is a taxonomic textual representation of corals, It is i The taxonomic text encoding representation of the coral categories, It is i The transpose of the morphological text encoding matrix of the coral categories, It is i The transpose of the taxonomic text encoding matrix of the coral categories, MorpAdapter is the morphological adapter, TaxoAdapter is the taxonomy adapter, and These are the taxonomic visual features of corals output by the taxonomic adapter and the morphological visual features of corals output by the morphological adapter, respectively.

[0012] Optionally, the optimization of the few-sample coral intelligent recognition model by constraining the JS loss and similarity loss of the taxonomic adapter and the morphological adapter specifically includes: Using the formula Calculating JS loss .

[0013] Optionally, the loss function of the trained few-sample coral intelligent recognition model is: ; in, is the loss function of the trained few-shot coral intelligent recognition model, is the similarity loss between the predicted probability of the taxonomic adapter and the predicted probability of the morphological adapter, is the JS loss, is the CE loss between the label and the predicted probability, and are the tunable parameters.

[0014] According to the specific embodiments provided by the present application, the present application has the following technical effects: The present application provides a few-shot coral intelligent recognition method based on morphology and taxonomy, which obtains coral images, morphological text prompts of corals, and taxonomic text prompts of corals; processes the coral images through a CLIP visual encoder to obtain visual encoding representations of corals; processes the visual encoding representations of corals through a taxonomic adapter and a morphological adapter to obtain taxonomic visual features of corals and morphological visual features of corals; utilizes the knowledge of coral morphology and biological taxonomy to enhance the knowledge of the vision-language large model, further improving the recognition performance of coral species; optimizes the few-shot coral intelligent recognition model by constraining the JS loss and similarity loss of the taxonomic adapter and the morphological adapter; through the morphological visual features of corals output by the morphological adapter, adopts a texture projection layer, and optimizes it through the CE loss to obtain a trained few-shot coral intelligent recognition model; inputs the graphic and text data of the coral to be recognized into the few-shot coral intelligent recognition model to identify the coral species. The present application combines the visual texture features and morphological structure features of coral species with the text knowledge of coral biology taxonomy and coral morphology, and can achieve a series of underwater coral species recognition tasks only by updating a small number of parameters and a small number of coral pictures. By using the taxonomic adapter and the morphological adapter, the present application can improve the coral species recognition performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0016] Figure 1 is a flowchart of a few-shot coral intelligent recognition method based on morphology and taxonomy in an embodiment of the present application; Figure 2 is a flowchart of training a few-shot coral intelligent recognition model in a few-shot coral intelligent recognition method based on morphology and taxonomy in an embodiment of the present application; Figure 3This is a model architecture diagram of a few-shot coral intelligent recognition method based on morphology and taxonomy provided by an embodiment of the present application. Detailed implementation manners

[0017] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0018] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0019] In an exemplary embodiment, as Figure 1 and Figure 3 shown, a few-shot coral intelligent recognition method based on morphology and taxonomy is provided. This method is executed by a computer device, specifically, it can be executed alone by a computer device such as a terminal or a server, or jointly executed by a terminal and a server. In the embodiments of the present application, it includes the following S101 to S103. Among them: S101: Obtain the graphic and text data of the coral sample; the graphic and text data includes: coral images, morphological text prompts of the coral, and taxonomic text prompts of the coral.

[0020] S102: Train a few-shot coral intelligent recognition model according to the graphic and text data of the coral sample to obtain a trained few-shot coral intelligent recognition model; the few-shot coral intelligent recognition model includes: a CLIP text encoder, a CLIP visual encoder, a morphological adapter, and a taxonomic adapter.

[0021] S103: Identify the coral species based on the trained few-shot coral intelligent recognition model according to the graphic and text data of the coral to be identified.

[0022] The few-shot coral intelligent recognition method based on morphology and taxonomy provided by the present application is proposed based on the CLIP method. Among them, CLIP is short for Contrastive Language-Image Pretraining, which is a machine learning framework that combines images and texts proposed by OpenAI. CLIP is trained on a large-scale image-text pair data and can achieve zero-shot recognition and few-shot recognition. CLIP consists of a CLIP text encoder and a CLIP visual encoder. During the computer training process, the visual representation of the CLIP visual encoder and the text representation of the CLIP text encoder are embedded into a new representation space, and then the cosine distance is used for alignment in the new representation space. Specifically, given an image and a manually crafted text template: a photo of [CLASS]. For images belonging to the category The predicted probability can be expressed as: where, is the visual feature of the last layer of the CLIP visual encoder; is the encoded representation of the text; is the transpose of the matrix of the encoded representation of the text; is the input image; is the category of the image, is the given image. However, the traditional CLIP method lacks detailed text captions for corals. The general text caption template in CLIP, "a photo of [CLASS]", cannot introduce the necessary domain knowledge to highlight the detailed differences between coral species. Therefore, this application further obtains a few-shot coral intelligent recognition method based on coral morphology and taxonomy knowledge by combining a morphological adapter, a taxonomic adapter, and a loss function on the basis of the CLIP method.

[0023] As

[0024] shown, the training process of the few-shot coral intelligent recognition model is replaced by the following S201~S204: Figure 2 S201: Process the morphological text caption of the coral and the taxonomic text caption of the coral based on the CLIP text encoder to obtain the morphological text encoded representation of the coral and the taxonomic text encoded representation of the coral, respectively. S202: Process the coral image based on the CLIP visual encoder to obtain the visual encoded representation of the coral.

[0025] S203: Process the visual encoded representation of the coral based on the taxonomic adapter and the morphological adapter to obtain the taxonomic visual feature of the coral and the morphological visual feature of the coral; and optimize the few-shot coral intelligent recognition model by the JS loss and the similarity loss that constrain the taxonomic adapter and the morphological adapter.

[0026] In an exemplary embodiment, the JS loss is calculated using the formula

[0027] The loss function of the trained few-shot coral intelligent recognition model is calculated using the formula according to the JS loss, and then the few-shot coral intelligent recognition model is optimized. Among them, is the loss function of the trained few-shot coral intelligent recognition model, is the loss function of the trained few-shot coral intelligent recognition model, is the similarity loss between the predicted probability of the taxonomic adapter and the predicted probability of the morphological adapter, is the JS loss, is the CE loss between the label and the predicted probability, and are tunable parameters, and are the taxonomic visual features of the coral and the morphological visual features of the coral output by the taxonomic adapter and the morphological adapter respectively.

[0028] :Among them, the reversed divergence loss function, that is, the JS loss can amplify the difference between the taxonomic adapter and the morphological adapter, thereby guiding the model to learn the biological taxonomy and morphological knowledge of the coral.

[0029] S204: According to the morphological visual features of the coral output by the morphological adapter, a texture projection layer is adopted and optimized through the CE loss to obtain a trained few-shot coral intelligent recognition model.

[0030] In an exemplary embodiment, the present application selects the ResNet 50 model as the CLIP visual encoder and selects Transformer as the CLIP text encoder to train the few-shot coral intelligent recognition model, including 100 training epochs, a batch size of 64, optimized using the Adam optimizer, and a learning rate of 0.001. The input of the few-shot coral intelligent recognition model is the coral image, the taxonomic text prompt of the coral, and the morphological text prompt of the coral obtained by manual collection. The taxonomic text prompt of the coral and the morphological text prompt of the coral are input into the CLIP text encoder to obtain the taxonomic text encoding representation and the morphological text encoding representation of the coral; the image of the coral is input into the CLIP visual encoder to obtain the visual encoding representation of the coral, and then the visual encoding representation of the coral is output to the taxonomic adapter and the morphological adapter, and then the few-shot coral intelligent recognition model is optimized by constraining the JS loss and the similarity loss of the taxonomic adapter and the morphological adapter; the morphological visual features of the coral output by the morphological adapter are input into the texture projection layer and then optimized through the CE loss. The input of the optimized model is the image of the coral, the taxonomic text prompt of the coral, and the morphological text prompt of the coral, and the output is the recognition result of the coral species.

[0031] To integrate the coral biological taxonomy knowledge, the present application designs a taxonomic adapter, which is composed of a bottleneck layer and a residual layer connected. The taxonomic adapter is expressed by the formula as: , where and are the weights of the bottleneck layer, is the transpose of the visual feature matrix. ReLU is the activation function. Add the newly learned taxonomic knowledge and the prior knowledge of CLIP through the residual network, and the classification adapter can be expressed as: , where is the adjustment factor for adjusting the newly learned coral biotaxonomic knowledge and the prior knowledge of CLIP, is the transpose of the taxonomic adapter matrix.

[0032] This application designs a morphological adapter to extract the morphological visual features of corals, including the texture, structure, and color of corals. Let and be the visual features of the last two layers of the CLIP visual encoder. By using the formula calculate the outer product of the feature vectors between two different layers in the morphological adapter of the coral species recognition model , the morphological adapter can capture the second-order relationship between different texture features in the local area, enabling it to learn the complex patterns and subtle differences in coral textures. Using the formula max-pool the outer product of the feature vectors ; the outer product of the max-pooled feature vectors can further retain the key texture information of the coral and highlight the texture differences between different coral species.

[0033] The global structural information of the coral image and the correlation between structures are considered the key to distinguishing coral species. This application proposes a learnable structural block integrating a non-local network. Using the formula , obtain the global structural information of the coral according to the coral image ; distinguish coral species according to the global structural information of the coral and the correlation between structures; where is the transpose of the visual feature matrix, is the learnable projection matrix, is the outer product of the feature vectors between two different layers in the morphological adapter of the few-shot coral intelligent recognition model, is the max-pooling function; is is the coral structure modeling module, is the weight, is the projection block for obtaining channel dependence, capturing the subtle changes in the coral color, is the number of groups for grouping the coral images based on the structural position, j is the j th group of coral images, is the visual feature of the last i th layer of the CLIP visual encoder, and They are the visual features of the last first layer and the last second layer of the CLIP visual encoder respectively.

[0034] In an exemplary embodiment, the predicted probability of the morphological adapter is: , and ; the predicted probability of the taxonomic adapter is: , and ; where is the given image, is the category, is the learnable text encoding parameter, is the fine-tunable trade-off parameter, is the morphological text encoding representation of the coral, is the i th morphological text encoding representation of the coral of the th category, is the taxonomic text encoding representation of the coral, i is the th taxonomic text encoding representation of the coral of the i th category, is the transpose of the morphological text encoding representation matrix of the coral of the i th category, MorpAdapter is the morphological adapter, TaxoAdapter is the taxonomic adapter, and are the taxonomic visual features of the coral and the morphological visual features of the coral output by the taxonomic adapter and the morphological adapter respectively.

[0035] This application proposes an adaptive fine-tuning method for CLIP, including two models: a morphological adapter and a taxonomic adapter. Specifically, by connecting the morphological cues with the captured visual features, the morphological adapter can distinguish the fine-grained visual features of different species. The morphological features of these two corals are similar, but they belong to different genera. In addition to the species labels, corals are also interrelated in a comprehensive taxonomic system. Even if the morphological features are similar, distinguishable representations can be learned through the taxonomic hierarchy (from genus to family) for classification. To this end, this application designs a taxonomic adapter to further align the visual representations with the hierarchical classification structure.

[0036] This application is the first to attempt to apply VLMs to the coral recognition task and proposes an adapter (CORAL-adapter) based on the model fine-tuning method, which is a novel framework that combines two complementary types of coral-specific knowledge (taxonomy and morphology) with the general knowledge learned by CLIP. In addition, CORAL-adapter can be applied with only a small number of parameter updates and can be used as a plug-in module for various CLIP-based methods.

[0037] There have been few-shot learning methods in the prior art, but few-shot learning has not been introduced into the field of coral intelligent recognition. For example, CLIP-adapter uses an additional bottleneck layer to learn new features integrated with the zero-shot prior knowledge of CLIP. Task Residual (TaskRes) retains the original text-based classifier and learns a new classifier for the target task by fine-tuning a set of prior-independent parameters as the residuals of the original parameters. LP++ generalizes the standard linear probing classifier by combining text knowledge from learnable text embeddings with class multipliers that mix image and text features. However, due to the unique coral features (e.g., texture, color, and structure) of corals and the influence of the complex marine environment, the general visual features extracted by traditional few-shot recognition methods often cannot handle the morphological features of coral species and the rich hierarchy in their taxonomy.

[0038] This application randomly divides the HSCR 16K coral dataset into a training set, a validation set, and a test set in a ratio of 6:2:2. This application uses 1, 2, 4, 8, and 16 samples of each coral species to train the adapter respectively, and then tests the model on the complete test split. In terms of coral species recognition, samples of some coral species may be few or difficult to obtain. Few-shot learning enables effective recognition with only a small number of labeled samples (e.g., 1 to 16 images per coral species), reducing the need for a large amount of labeled data. The detailed performance is shown in Table 1 (where, 0 to 16 represents the number of coral samples provided by each species). The CORAL-adapter based on the model fine-tuning method shows excellent performance superior to all other methods in all few-shot recognition settings. Specifically, the Zero-shot CLIP is not suitable for this task (only 13.5% accuracy), while the CORAL-adapter can achieve a 27.9% improvement even under the extreme condition of only providing 1 coral recognition sample for each species. Compared with other adaptation methods, the CORAL-adapter has achieved a huge performance improvement. For example, in the setting of only providing 1 coral recognition sample for each species, the accuracy has increased by 10% to 20%, and in the setting of only providing 8 coral recognition samples for each species, the recognition performance has increased by 10% to 40%. In summary, due to the learned classification relationships and morphological representations, the CORAL-adapter exhibits strong coral recognition capabilities.

[0039] This application further conducts ablation studies on each part to analyze the reasons for the accuracy improvement of the CORAL-adapter. Under the setting of providing 16 samples for each coral species, this application explores the effectiveness of different components. In the setting without any adapter, the coral recognition accuracy of the Zero-shot CLIP is 13.46%. When the taxonomic adapter is added, the recognition accuracy of the model reaches 28.88%. When the morphological adapter is added, the coral recognition accuracy is 46.50%. When both the taxonomic adapter and the morphological adapter are added to the model, the coral recognition accuracy reaches 77.27%. Compared with adding only the taxonomic adapter alone, the morphological adapter contributes more significantly to improving the accuracy.

[0040] Table 1 Performance comparison of CORAL-adapter and SOTA methods in coral species recognition

[0041] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0042] Specific examples are used in this article to elaborate on the principles and implementation methods of this application. The descriptions of the above embodiments are only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, based on the idea of this application, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. A few-shot coral intelligent recognition method based on morphology and taxonomy, characterized in that The few-shot coral intelligent recognition method based on morphology and taxonomy includes: Obtain the graphic and text data of coral samples; the graphic and text data includes: coral images, morphological text prompts of corals, and taxonomic text prompts of corals; Train a few-shot coral intelligent recognition model based on the graphic and text data of coral samples to obtain a trained few-shot coral intelligent recognition model; the few-shot coral intelligent recognition model includes: a CLIP text encoder, a CLIP visual encoder, a morphological adapter, and a taxonomic adapter; Based on the graphic and text data of the coral to be recognized, identify the coral species based on the trained few-shot coral intelligent recognition model; The training process of the few-shot coral intelligent recognition model is: Process the morphological text prompts and taxonomic text prompts of corals based on the CLIP text encoder to obtain the morphological text encoding representation and taxonomic text encoding representation of corals respectively; Process the coral image based on the CLIP visual encoder to obtain the visual encoding representation of the coral; Process the visual encoding representation of the coral based on the taxonomic adapter and the morphological adapter to obtain the taxonomic visual features and morphological visual features of the coral; and optimize the few-shot coral intelligent recognition model by constraining the JS loss and similarity loss of the taxonomic adapter and the morphological adapter; According to the morphological visual features of the coral output by the morphological adapter, use a texture projection layer and optimize it through CE loss to obtain a trained few-shot coral intelligent recognition model.

2. The few-shot coral intelligent recognition method based on morphology and taxonomy according to claim 1, characterized in that, The taxonomic adapter includes: a bottleneck layer and a residual layer.

3. The few-shot coral intelligent recognition method based on morphology and taxonomy according to claim 1, characterized in that, The morphological visual features of the coral include: the texture of the coral, the structure of the coral, and the color of the coral.

4. The few-shot coral intelligent recognition method based on morphology and taxonomy according to claim 3, wherein The morphological adapter specifically includes: Using the formula calculate the outer product of the feature vectors between two different layers in the morphological adapter of the few-shot coral intelligent recognition model ; Using the formula Outer product of the maximum pooling feature vectors ; Highlight the texture differences between different coral species according to the outer product of the maximum pooling feature vectors; Using the formula , obtain the global structure information of the coral based on the coral image ; Distinguish coral species according to the global structure information of the coral and the correlation between structures; Among them, is the transpose of the visual feature matrix, is a learnable projection matrix, is the outer product of the feature vectors between two different layers in the morphological adapter of the few-shot coral intelligent recognition model, is the max pooling function; is the feature extraction function, is the coral structure building block, is the weight, is the projection block for obtaining channel dependencies, capturing the subtle changes in the coral color, is the number of groups for grouping the coral images based on the structural positions, j is the j th group of coral images, is the visual feature of the last i th layer of the CLIP visual encoder.

5. The few-shot coral intelligent recognition method based on morphology and taxonomy according to claim 4, wherein Use the formula , and to calculate the predicted probability of the morphological adapter ; Using the formula , and calculate the predicted probability of the taxonomic adapter ; Among them, is the given image, is the category, is the learnable text encoding parameter, is the fine-tunable trade-off parameter, is the morphological text encoding representation of the coral, is the i morphological text encoding representation of the coral of the th category, is the taxonomic text encoding representation of the coral, i is the taxonomic text encoding representation of the coral of the th category, i is the transpose of the matrix of the morphological text encoding representation of the coral of the th category, i is the transpose of the matrix of the taxonomic text encoding representation of the coral of the MorpAdapter is the morphological adapter, TaxoAdapter is the taxonomic adapter, and are respectively the taxonomic visual feature of the coral and the morphological visual feature of the coral output by the taxonomic adapter and the morphological adapter.

6. The few-shot coral intelligent recognition method based on morphology and taxonomy according to claim 5, wherein The few-shot coral intelligent recognition model is optimized by constraining the JS loss and similarity loss of the taxonomic adapter and the morphological adapter, specifically including: Use the formula to calculate the JS loss .

7. The few-shot coral intelligent recognition method based on morphology and taxonomy according to claim 6, characterized in that, The loss function of the trained few-shot coral intelligent recognition model is: ; Among them, is the loss function of the trained few-shot coral intelligent recognition model, is the similarity loss between the prediction probability of the taxonomic adapter and the prediction probability of the morphological adapter, is the JS loss, is the CE loss between the label and the prediction probability, and are the tunable parameters.

Citation Information

Patent Citations

  • Text-guided multi-modal relationship extraction method and apparatus

    WO2025130069A1