Marine ship classification method based on multi-modal large model

By combining the pre-training ability of visual and language models and the prompt learning strategy of domain perception, the marine ship classification method with the "winding-unwinding" learning mechanism is adopted, and the problem of insufficient classification performance in the prior art in specific marine application scenarios and small samples is solved, achieving efficient and accurate ship classification.

CN119989101APending Publication Date: 2025-05-13SHANGHAI UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510196833.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing ship classification technology lacks adaptability when facing specific marine application scenarios, especially in the case of small samples, and its classification performance is insufficient.

Method used

The marine ship classification method based on multimodal large models is adopted, combined with the pre-training ability of visual and language models, and the domain-perceived prompt learning strategy and the "winding-unwinding" learning mechanism is used to achieve efficient classification of marine ships with few samples.

Benefits of technology

It significantly improves the accuracy and robustness of ship classification, solves the problem of insufficient generalization ability of traditional methods in the small sample scenario, and improves the model's adaptability to the ship field.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989101A_ABST
    Figure CN119989101A_ABST
Patent Text Reader

Abstract

The invention discloses an ocean ship classification method based on a multi-modal large model, and aims to solve the problem that an existing ship classification model is poor in performance in few-sample and cross-domain tasks. According to the method, a large-scale visual language model and a prompt learning technology are combined, so that effective utilization of a small number of labeled samples and efficient adaptability of cross-domain ship classification are realized. Specifically, through an entanglement-unentanglement strategy, image features of different domains are fused (entangled), and specific cues of the domains are used for unentanglement, so that classification prediction of the corresponding domains is generated. According to the method, domain specific prompt learning is adopted, and the generalization ability of the model can be enhanced by fully utilizing the knowledge difference between the source domain and the target domain. Through the method, the ship classification accuracy and robustness in a cross-domain scene are remarkably improved, the method is suitable for the fields of ocean monitoring, ship identification, intelligent maritime affair management and the like, and powerful technical support is provided for ocean traffic safety and environmental protection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of ship classification, and specifically to a method for classifying marine ships based on a multimodal large model. Specifically, the present invention proposes a few-sample domain adaptation method based on prompt learning by combining visual and language large models to achieve accurate classification of different types of marine ships. This method can make full use of specific knowledge in the source domain and the target domain, improve classification accuracy in data-scarce environments, and is applied to the fields of marine monitoring, ship identification, and marine safety. Background Art

[0002] With the continuous development of the marine transportation industry, the types and number of ships at sea are increasing. How to efficiently and accurately identify and classify ships has become an important issue in marine safety and management. Traditional ship classification methods mainly rely on ship identification systems, radars and image processing technologies, and classification is achieved by analyzing the appearance characteristics of the ship (such as hull length, width, etc.). However, these methods are highly dependent on data, and there is a problem of reduced recognition rate when dealing with different sea conditions, weather conditions and new types of ships.

[0003] In recent years, with the development of deep learning and large-scale pre-trained models, ship classification technology based on image recognition has made significant progress, especially the introduction of visual-language models (such as CLIP), which makes it possible to classify ships through multimodal information. Such models can effectively identify and classify ships in cross-modal situations by learning the features of joint images and texts. However, most of the existing ship classification technologies rely on pre-trained general models, which often lack adaptability to specific domains when facing specific marine application scenarios, especially in the case of few samples, and their classification performance still needs to be improved.

[0004] Existing studies have shown that the performance of pre-trained models in specific tasks can be improved through prompt learning. Prompt learning enables the model to better understand and adapt to the needs of specific tasks by adding additional contextual information to the model input. However, existing prompt learning methods mainly focus on learning general prompts, while ignoring the knowledge modeling of specific domains (such as the marine ship domain). This means that these methods may not be able to fully utilize the domain-specific knowledge and information when dealing with specific tasks such as marine ship classification, thereby limiting the performance of the model in practical applications.

[0005] Therefore, how to design and optimize the prompt learning method by combining the professional knowledge in the field of marine ships is an urgent problem to be solved in the current ship classification technology. Summary of the invention

[0006] The purpose of the present invention is to provide a marine ship classification method based on a multimodal large model to address the problems of weak cross-domain adaptability and insufficient classification accuracy in the current marine ship classification technology. By combining the pre-training capabilities of visual and language models and domain-specific prompt learning strategies, efficient classification of marine ships with few samples is achieved, and the generalization ability of the model in different marine fields is effectively improved. Specifically, by introducing a domain-aware prompt learning strategy, visual features and domain-specific text prompts are combined to enhance the adaptability of the model to the ship field. A learning mechanism of "entanglement first-disentanglement later" is adopted, that is, the features of different fields are first "entangled" together to form a feature set, and then these "entangled" features are disentangled through domain-specific prompt information, thereby achieving accurate classification of ship categories. It not only deepens the model's understanding of ship features, solves the problem of insufficient generalization ability of traditional methods in few-sample scenarios, but also significantly improves the accuracy and robustness of ship classification.

[0007] To achieve the above object, the technical solution of the present invention is as follows:

[0008] A method for classifying marine vessels based on a multimodal large model comprises the following steps:

[0009] Step A: Data preprocessing: preprocess the collected marine ship images and text data, including data cleaning, denoising and normalization, and extract the visual features and text description of the ship;

[0010] Step B: Optimization of learnable prompt words: Using the pre-trained visual-language multimodal large model, the preset text prompt words are converted into learnable continuous embedding vectors, and the embedding vectors are optimized by minimizing the cross-entropy loss function to achieve cross-modal feature alignment;

[0011] Step C: Construction of domain-aware cue words: The cue words are divided into domain cue words, context cue words and category cue words, wherein the domain cue words contain semantic tags related to the ship domain, and a classifier is trained independently for each domain by maximizing the prediction probability;

[0012] Step D: training domain prompt words through the “entanglement-disentanglement” strategy: mixing visual features from at least two different fields to generate entangled features, using the domain prompt words to perform domain decoupling prediction on the entangled features, and completing model training by jointly optimizing the classification loss function and the disentanglement loss function;

[0013] Step E: Deploy the trained model to the marine vessel perception system to achieve real-time classification of marine vessel types.

[0014] Furthermore, the data preprocessing in step A specifically includes:

[0015] First, the collected ship data or text materials are preprocessed and feature extracted. Preprocessing includes steps such as data cleaning, denoising and normalization to ensure the quality and availability of the data. Feature extraction focuses on extracting effective information that can reflect the general characteristics of the object from the raw data.

[0016] Furthermore, the step B may learn to optimize the prompt words, specifically including:

[0017] Use pre-trained multimodal large models to design prompt words to achieve cross-modal feature alignment and optimization of classification tasks. For example, the CLIP (Contrastive Language–Image Pre-training) model maps visual features and text descriptions into a unified feature space through contrastive learning, so that efficient knowledge transfer can be performed with limited data. For a given image and category text description, the CLIP model trains the image encoder and text encoder through contrastive learning. Specifically, given an image xi and a matching category text description ti (for example, "a picture belonging to [category]"), the model ensures that the visual and text features can be aligned in a unified space by maximizing the cosine similarity between the image and the corresponding text description, while minimizing the similarity between the image and irrelevant text. This goal can be expressed by the following formula:

[0018]

[0019] Where exp represents logarithmic calculation, g(·) represents text encoder, f(·) represents image encoder, <·,·> represents cosine similarity calculation, K is the number of categories, t i is the text description of the i-th category.

[0020] In the cue word design, we create a hand-crafted text cue word for each category, such as "a picture of [category]". These hand-crafted cues are converted into fixed vector representations through the text encoder g(ti). However, hand-crafted cue words are not always optimal, so the embedding representation of cue words needs to be further optimized.

[0021] In order to optimize the prompt word, it is represented as a learnable continuous embedding vector, that is, the contextual vocabulary in the prompt word is converted into a set of optimizable vector parameters θ. In this way, the prompt word is converted from a static manual design to a dynamic learnable representation, which better meets the needs of the target task. The optimization goal is to learn the best prompt word embedding by minimizing the cross entropy loss:

[0022]

[0023] Among them, t(θ) represents the learnable prompt word embedding, D represents the marine ship images in all training data sets, logP(y|x,t(θ)) represents the calculation of the prediction probability for each sample (x,y), and the specific calculation formula is as follows:

[0024]

[0025] Among them, exp represents logarithmic calculation, g(·) represents text encoder, f(·) represents image encoder, <·,·> represents cosine similarity calculation, K is the number of categories, and t(θ) k represents the learnable text embedding corresponding to k categories, t(θ) y Represents the text embedding of the y-th category.

[0026] Furthermore, the step C of constructing domain-aware prompt words specifically includes:

[0027] Based on the cue word learning process, domain cue words are designed for the few-shot domain adaptation task. Specifically, the cue words are divided into three parts: domain cue words, context cue words, and category cue words. Among them, domain cue words are used to model domain knowledge, context cue words are used to learn task-specific embeddings, and category cue words are responsible for providing category information. Each domain has a unique classifier. The definition of domain-specific cue words is as follows:

[0028]

[0029] Among them, ζ represents the domain cue word, θ represents the context cue word, ω represents the learnable parameters in the cue word, k represents the kth category, and l represents the lth domain; ζ l The prompt word representing domain l is composed of n context tokens (i.e., [d]1[d]2,...,[d]n[d]1[d]2,...,[d]n), where n is the number of context tokens used in the prompt word.

[0030] Based on the domain cue words, the cue is learned by maximizing the prediction probability. Since the labels and domain indexes are known during the cue word learning process, the only difference is that the domain-aware cue words make the classifier no longer shared in all domains, but a separate classifier is designed for each domain, which is consistent with the known target domain information in the few-shot domain adaptation task. The calculation formula is as follows:

[0031]

[0032] Where l represents the domain to which the image x belongs. The introduction of domain-specific cues means that each domain has a unique classifier.

[0033] Step D: The “entanglement-disentanglement” training process of the prompt words in a specific field includes:

[0034] By entangled visual features from different domains, these entangled features are disentangled using domain-specific knowledge, so as to effectively extract and learn domain-specific cue words. This step is divided into two main processes: the entanglement process and the disentanglement process.

[0035] In the intertwining process, we select image features from images of different domains and mix these features together to obtain an intertwined representation. First, we extract visual features from two images of different domains. For an image in domain l, The visual features are extracted through the image encoder f(·). Similarly, for the image in domain q The visual features are extracted through the image encoder f(·). Then, the image features from the two fields are mixed in a certain proportion to obtain an intertwined feature f ij , expressed as:

[0036]

[0037] in, is the i-th image in domain l, is the j-th image of domain q, f(·) represents the image encoder, λ is the mixing weight coefficient, and in this embodiment, λ is fixed to 0.5, which means that the features of the two domains are mixed with equal weights.

[0038] In the disentanglement process, domain-specific prompt words are used to disentangle the entangled features respectively, so as to learn the information of specific domains. For the entangled features, the prompt words of domain l and domain q are used for prediction respectively. In the prediction process, if the prompt word of domain l is used for prediction, the label of the entangled feature should be the image in domain l. Label y i If the cue word of domain q is used for prediction, the label of the entangled feature should be the image in domain q Label y j Finally, in order to learn domain-specific prompt words, a disentanglement loss function is defined. The goal of this function is to maximize the correct prediction of the model for the entangled features when using different domain prompt words, and then minimize the disentanglement loss function. The formula is as follows:

[0039]

[0040] in, is the set of training data, Indicates the prompt word t(ω) in the usage domain l lWhen , the model predicts the probability of label yi, which is calculated by the following formula:

[0041]

[0042] in, is the cue word t(ω) of domain l l Corresponding category y i The feature representation of , <·,> represents the similarity between two vectors (usually cosine similarity), and K is the total number of categories.

[0043] The final optimization goal is to minimize the common classification loss L CE (θ) and the disentanglement loss function L Dis (ω), its total loss function is:

[0044] min ω L final =L CE (θ)+γL Dis (ω)

[0045] Among them, γ is a hyperparameter used to balance the loss of ordinary classification tasks and disentanglement tasks.

[0046] Finally, the trained marine ship classification model is deployed to the perception system of the unmanned boat. By integrating with the shipboard color camera and computing module, the system can detect various types of ships in the ocean scene (such as cargo ships, fishing boats, and various military ships) and identify the category of the target ship in real time. Based on the identified target information, the unmanned boat can perform a variety of tasks, such as automatically recording the category of the target in cruise monitoring, or identifying suspicious ships in real time during maritime law enforcement. The system has good open vocabulary generalization capabilities, and can accurately identify even ship materials that did not appear in training, thereby effectively improving the execution capability and intelligence level of the unmanned boat in multi-task scenarios.

[0047] Compared with the prior art, the technical effects of the present invention are as follows:

[0048] 1) Through the "entanglement-disentanglement" strategy, a large-scale pre-trained visual language model is combined with specific domain prompt words and successfully applied to the marine ship classification task. Through the cross-domain ship visual feature entanglement and disentanglement process, the domain-specific information of the ship category is effectively extracted, which solves the problems of scarce samples, difficulty in fully utilizing domain information, and insufficient generalization of prompt words in ship classification in marine environments, and greatly improves the accuracy of ship classification under few sample conditions.

[0049] 2) The construction of domain prompt words combined with keywords in the field of ships (such as functional attributes and operating scenarios) enhances the model's modeling ability of domain knowledge and ensures that the classifier can adapt to complex scenarios of different sea conditions, weather and ship types.

[0050] 3) By dynamically optimizing the continuous embedding vectors of prompt words, efficient alignment of visual and textual features is achieved, overcoming the limitations of manually designed prompt words and improving the fusion efficiency of multimodal data. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0052] Figure 1 It is a flow chart of an embodiment of the method for classifying marine vessels based on a multi-modal large model according to the present invention. DETAILED DESCRIPTION

[0053] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0054] Example 1

[0055] like Figure 1 As shown, this embodiment provides a method for classifying marine vessels based on a multimodal large model, and the method specifically includes the following steps:

[0056] Step 1: Data cleaning and data preprocessing.

[0057] First, the collected ship data or text materials are preprocessed and feature extracted. Preprocessing includes steps such as data cleaning, denoising and normalization to ensure the quality and availability of the data. Feature extraction focuses on extracting effective information that can reflect the general characteristics of the object from the raw data.

[0058] Step 2: Design a traditional cue word learning process.

[0059] Use pre-trained multimodal large models (such as CLIP) to design prompt words to achieve cross-modal feature alignment and optimization of classification tasks. The CLIP model maps visual features and text descriptions into a unified feature space through contrastive learning. This means that the CLIP model can understand and associate images and related text descriptions, so that when given an image, the model can generate a matching text description, and vice versa.

[0060] Specifically, the CLIP model contains an image encoder f(·) and a text encoder g(·). Given an image x i and the matching category text description t i (e.g. "a picture belonging to [category]"), the model trains the two encoders by maximizing the cosine similarity between the image and the corresponding text description, while minimizing the similarity between the image and unrelated text, ensuring that the visual features and text features can be aligned in a unified space. This goal can be expressed as follows:

[0061]

[0062] Among them, exp represents logarithmic calculation, <·,·> represents cosine similarity calculation, K is the number of categories, t i is the text description of the i-th category.

[0063] Design of prompt words:

[0064] Prompts are the bridge between images and text descriptions. Through them, the model can understand and generate text descriptions related to images. Hand-crafted text prompts are created for each category, such as "a picture of [category]". These hand-crafted prompts are passed through the text encoder g(t i ) is converted into a fixed vector representation. Although simple and intuitive, it is not always optimal, because different tasks, datasets, and categories may require different prompt words to better express their characteristics, so it is necessary to optimize the embedding representation of the prompt word.

[0065] The prompt word is represented as a learnable continuous embedding vector, that is, the contextual vocabulary in the prompt word is converted into a set of optimizable vector parameters θ. In this way, the prompt word is transformed from a static manual design to a dynamic learnable representation, which can better adapt to the needs of the target task. The optimization goal is to learn the best prompt word embedding by minimizing the cross entropy loss. Specifically, for a given training data set D (for example, marine ship images), the predicted probability P(y|x,t(θ)) of each sample (x,y) is calculated, and the cross entropy loss function between the actual category and the predicted category is minimized. The formula is as follows:

[0066]

[0067] Where t(θ) y represents the text embedding for the current category y, t(θ) k Represents the text embedding of the k-th category.

[0068]

[0069] Step 3: Design domain prompt words specific to the ship category.

[0070] The prompt words are divided into three parts: domain prompt words, context prompt words and category prompt words. Domain prompt words are used to model domain knowledge. For ship categories, they need to reflect the unique attributes and characteristics of different ship types (such as cargo ships, tankers, warships, etc.). Context prompt words are used to learn task-specific embeddings, that is, to help the model understand the context information of the input data, so as to classify or predict more accurately. Category prompt words are used to provide information about the target category, which helps the model make more accurate judgments in classification tasks.

[0071] Field prompt word l , i.e., the prompt word of domain l, consists of n context tokens, i.e., [d]1[d]2,...,[d]n, where n is the number of context tokens used in the prompt word. These tokens can be keywords, phrases, or symbols related to the ship category. For example, for the cargo ship category, the domain prompt words may include words related to cargo transportation, such as "freight" and "container".

[0072] After the introduction of domain-specific prompts, each domain (i.e. each ship category) will have a unique classifier. The definition of domain prompts is as follows:

[0073]

[0074] Among them, ω represents the domain perception cue word, which is composed of the domain cue word ζ l and context cue word θ; l represents the lth field, k represents the kth category, and Category represents the category cue word.

[0075] In the process of learning the prompt words, the labels and domain indexes are known. The model learns these prompt words by maximizing the prediction probability, thereby continuously optimizing its performance in different fields and ensuring that the model can fully utilize the information provided by the domain-specific prompt words to improve the accuracy of classification or prediction. The specific calculation formula is as follows:

[0076]

[0077] Among them, l represents the field to which the image x belongs, K represents the total number of categories, and y represents that the current sample belongs to the yth category.

[0078] Step 4: Use domain cue words to perform data enhancement to complete feature winding.

[0079] During the entanglement process, our goal is to extract image features from images of different domains and mix these image features together to create a mixed entangled feature representation. This mixture is designed to simulate cross-domain feature interactions, thereby enhancing the model's understanding of the differences between domains. The specific steps are as follows:

[0080] Feature extraction: From images in domain l Extract visual features from domain q, usually through a pre-trained image encoder f(·). Similarly, Extract visual features.

[0081] Feature mixing: The visual features from the two domains are mixed in a certain ratio. This ratio is determined by the mixing weight coefficient λ. In this embodiment, λ is fixed to 0.5, that is, the features of the two domains are mixed with equal weights. This mixing produces a tangled feature representation f ij , expressed as:

[0082]

[0083] Data augmentation: This entangled feature representation is further enriched by using domain-cue words, which helps the model better understand and utilize these mixed features in subsequent steps.

[0084] Step 5: Training domain prompt words to complete feature disentanglement

[0085] In the disentanglement process, our goal is to use domain-specific cues to separate the parts of the entangled features that belong to different domains, thereby learning information in specific domains. The specific steps are as follows:

[0086] Prediction and label matching: For each entangled feature, we use the cue words of domain l and domain q for prediction. If the cue words of domain l are used for prediction, the expected prediction result should match the label of the image in domain l, that is, the label of the entangled feature should be the image in domain l. Label y i Similarly, if the cue word of domain q is used for prediction, the expected prediction result should match the label of the image in domain q, that is, the label of the entangled feature should be the image in domain q. Label y j .

[0087] Disentanglement loss function: Define the disentanglement loss function to learn domain-specific cues. The goal of this disentanglement loss function is to maximize the correct prediction probability of the model for the entangled features when using different domain cues. The disentanglement loss function works by calculating the difference between the model's predicted probability and the true label, and uses cosine similarity to measure the similarity between the feature representation and the category representation. The formula is as follows:

[0088]

[0089] in, is the set of training data, Indicates the prompt word t(ω) in the usage domain l l When the model is label y j The predicted probability is calculated by the following formula:

[0090]

[0091] in, is the cue word t(ω) of domain l l Corresponding category y i The feature representation of , <·,·> represents the similarity between two vectors (usually cosine similarity), and K is the total number of categories.

[0092] Optimization goal: Minimize the common classification loss function L at the same time CE (θ) and the disentanglement loss function L Dis (ω), its total loss function L final :

[0093] min ω L final =L CE (θ)+γL Dis (ω)

[0094] Among them, γ is a hyperparameter used to balance the loss of general classification tasks and disentanglement tasks to ensure that the model can simultaneously learn effective classification capabilities and domain-specific feature representations.

[0095] Step 6: Use the trained domain prompt words to classify marine vessels

[0096] Before the classification task begins, these trained domain cue words are loaded into the multimodal large model so that the model can utilize the domain knowledge contained in these cue words when processing input data.

[0097] The input ship data (such as images, text descriptions and other multi-modal data) is combined with the above-mentioned trained domain prompt words, and the multi-modal large model is used to extract features related to the ship category from the input ship data, and the input ship data is classified and predicted based on the extracted features. Based on the classification results, the model outputs the specific category to which the ship belongs. These categories may be divided according to the purpose of the ship (such as cargo ships, tankers, fishing boats, etc.), size (such as small ships, large ships, etc.) or other characteristics, thereby achieving high-precision marine ship classification.

[0098] The present invention combines a large multimodal model with trained domain prompt words, breaking through the limitations of traditional ship classification methods in processing multimodal data fusion and feature extraction. It can not only automatically adapt to the distribution characteristics of ship data in complex and changeable marine environments, but also improve the accuracy and efficiency of marine ship classification. It is of great significance to promote technological progress in marine traffic management, ship monitoring and related fields.

[0099] Finally, it should be noted that the above contents are only preferred embodiments of the present invention and are not intended to limit the present invention. Those skilled in the art may adjust and improve the above embodiments in accordance with the spirit and principles of the present invention, and all such modifications and equivalent substitutions are within the protection scope of the present invention.

Claims

1. A method for classifying marine vessels based on a multimodal large model, characterized in that: The following steps are involved: Step A: Data preprocessing: preprocess the collected marine ship images and text data, including data cleaning, denoising and normalization, and extract the visual features and text description of the ship; Step B: Optimization of learnable prompt words: Using the pre-trained visual-language multimodal large model, the preset text prompt words are converted into learnable continuous embedding vectors, and the embedding vectors are optimized by minimizing the cross-entropy loss function to achieve cross-modal feature alignment; Step C: Construction of domain-aware cue words: The cue words are divided into domain cue words, context cue words and category cue words, wherein the domain cue words contain semantic tags related to the ship domain, and a classifier is trained independently for each domain by maximizing the prediction probability; Step D: Training domain prompt words through the "entanglement-disentanglement" strategy: Mixing visual features from at least two different fields to generate entanglement features, using the domain prompt words to perform domain decoupling prediction on the entanglement features, and completing model training by jointly optimizing the classification loss function and the disentanglement loss function; Step E: Deploy the trained model to the marine vessel perception system to achieve real-time classification of marine vessel types.

2. The method for classifying marine vessels based on a multimodal large model according to claim 1 is characterized in that: The visual-language multimodal model in step B is a CLIP model, which maps visual features and text descriptions to a unified feature space through contrastive learning, and the learnable prompt word embedding is achieved by minimizing the cross entropy loss function L CE (θ) optimization, the formula is as follows: Where P(y|x,t(θ)) represents the predicted probability of each sample (x,y), which is calculated by the cosine similarity between visual features and text embeddings, t(θ) represents the learnable cue word embedding, and D is the training dataset.

3. The method for classifying marine vessels based on a multimodal large model according to claim 1 is characterized in that: In step C, the domain prompt words are composed of context tags related to the ship category, and an independent classifier is generated for each domain. The formula is: In the formula, ζ l represents the domain l prompt word, θ is the context prompt word, Category is the category prompt word; l represents the lth domain, and k represents the kth category.

4. The method for classifying marine vessels based on a multimodal large model according to claim 1 is characterized in that: The winding feature f in step D ij The generation formula is as follows: in, is an image in domain l, is the image of domain q, f(·) represents the image encoder, and λ is the mixing weight coefficient.

5. The method for classifying marine vessels based on a multimodal large model according to claim 1 or 2, characterized in that: In step D, the disentanglement loss function L is minimized. Dis (ω), the formula is as follows: In the formula, is the hint word t(ω) for the usage domain l l When the label y i The predicted probability of is the training data set.

6. The method for classifying marine vessels based on a multimodal large model according to claim 1 or 2, characterized in that: In step D, the total loss function L is minimized final , the formula is as follows: minutes ω L final =L CE (θ)+γL Dis (oh) Where γ is a hyperparameter.

7. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the classification method according to any one of claims 1 to 6 is implemented.

8. A marine vessel sensing system, characterized in that: It comprises a ship-borne camera, a computing module and the storage medium as claimed in claim 7, and is used for collecting ship images in real time and performing classification tasks.

Citation Information

Cited By

  • Remote sensing image domain adaptive classification method based on high-order prototype guidance

    CN120356014A

  • A domain-adaptive classification method for remote sensing images based on high-order prototype guidance

    CN120356014B