Underwater in-situ fish label-free sample detection method and equipment

By combining an image-text structured knowledge base with a visual-language model, the problems of category expansion and zero-shot detection in fish detection in underwater environments are solved. This enables matching detection and cross-domain reasoning for unseen categories, reducing costs and improving recognition capabilities.

CN120976535AActive Publication Date: 2025-11-18ZHEJIANG UNIV

Patent Information

Application Number
CN202511503809.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-21
Publication Date
2025-11-18
Estimated Expiration
2045-10-21

AI Technical Summary

Technical Problem

Traditional fish detection methods suffer from poor category scalability, high costs, lack of zero-sample detection capability, and inflexible human-computer interaction in underwater environments. Furthermore, existing visual-language models have not yet effectively solved the problem of open vocabulary detection in underwater in-situ environments.

Method used

A method for detecting unlabeled fish samples in situ underwater is adopted. By constructing an image-text structured knowledge base, using a visual-language model to align visual-text semantic space features, and combining a targeting-assisted model and a regularized loss function, the unlabeled detection of fish targets is achieved.

Benefits of technology

It achieves matching detection for categories not seen during the training phase, reducing detection costs and model maintenance difficulty, improving the ability to identify unknown categories, and possessing powerful open set detection and cross-domain reasoning capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976535A_ABST
    Figure CN120976535A_ABST
Patent Text Reader

Abstract

The invention discloses an underwater in-situ fish label-free sample detection method and equipment. The method comprises the following steps: acquiring an in-situ image shot by a monocular camera of an underwater observation platform, and constructing an image-text structured knowledge base; extracting a training and verification data set of the detection model from the knowledge base, and dividing the data set into a closed set category and an open world category; receiving a natural language text instruction which is input by a user and is used for describing the to-be-detected data set, and generating a text feature vector of a target category; a targeting auxiliary model is obtained through vision-text semantic space training, dynamic interactive fusion is carried out on the visual features of the in-situ image and the text features of the category, and end-to-end joint optimization is carried out on the target model; and performing label-free detection on fishes appearing in the underwater in-situ image by using the optimized target model. According to the method, new species which are not seen during training can be effectively identified, cross-domain reasoning can be implemented on a label-free observation sample, and the retraining and maintenance cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of general fish target detection, and particularly relates to an underwater in-situ fish label-free sample detection method and device. BACKGROUND

[0002] As the largest ecosystem on Earth, the health and stability of marine biodiversity are crucial for global climate regulation and human sustainable development. Fish, as a key component of marine ecosystems, their population changes are important indicators for assessing the distribution of marine fish resources and environmental health. Therefore, automatic and accurate detection of underwater fish is a key technology to realize smart fisheries, marine ecological monitoring and biodiversity protection, etc.

[0003] Traditional fish survey methods, such as catch survey method, are complex in operation, difficult to be scientific in sampling, and limited by survey conditions. This method is time-consuming and has large errors. In recent years, with the popularity of image and video collection platforms such as seafloor observation networks, remotely operated vehicles (ROVs) and autonomous underwater vehicles (AUVs), automated detection technology based on computer vision has been widely used in underwater target detection tasks. In deep water aquaculture net cages and open natural waters, underwater in-situ observation relies on fixed underwater platforms for real-time shooting and monitoring, which is less affected by underwater optical attenuation and suspended matter interference than mobile observation. And by laying a three-dimensional observation network composed of underwater camera equipment and sensor nodes, multi-dimensional spatiotemporal data of fish activity can be continuously obtained, and its non-intrusive nature effectively reduces the interference to biological behavior, which has important application value for species statistics and downstream ecological monitoring.

[0004] Target detection algorithms based on deep learning have achieved remarkable success in specific scenarios, but these algorithms usually rely on the setting of closed-set vocabulary. This means that the model can only recognize and detect a limited number of fish classes predefined in the training phase.

[0005] This closed-set paradigm faces serious challenges when applied to complex and variable underwater in-situ environments: 1. Poor class scalability and high cost: There are tens of thousands of fish species in the world. Whenever a new fish species needs to be detected, traditional supervised learning methods require collecting a large number of images of the new fish species and performing fine-grained manual labeling, and then retraining or incrementally training the entire model. This process not only has high cost and long cycle, but also as the number of classes increases, the difficulty of training and maintaining the model increases exponentially.

[0006] 2. Zero-shot detection capability is missing: Traditional models cannot recognize and locate fish species that have never appeared in the training dataset or extremely rare fish species, i.e., they do not have "zero-shot" detection capability. This greatly limits their application in biodiversity census, endangered species search and other exploratory tasks that need to handle open and unknown categories.

[0007] 3. Inflexible human-computer interaction: Traditional detection systems usually output preset category labels (such as "category_id_5"). Users cannot query the target of interest through natural and flexible language description, for example, they cannot directly issue instructions such as "detect all yellow and striped fish" or "find a guppy" to the system.

[0008] In recent years, large-scale vision-language pre-training models (VLMs), such as CLIP, have successfully aligned image visual features and text semantic features to a unified feature space through contrastive learning on massive image-text pairs. This cross-modal alignment capability provides a new approach to solving the above open vocabulary detection problem, but how to effectively apply it to the underwater in-situ scene with strong domain characteristics and extend from image-level classification to target-level precise positioning is a key problem faced by current technology development. SUMMARY

[0009] The purpose of the present application is to overcome the shortcomings of the prior art and provide a method and device for detecting unlabeled samples of fish in situ underwater.

[0010] To achieve the above-mentioned purpose, in a first aspect, the present application provides a method for detecting unlabeled samples of fish in situ underwater, comprising: Step 1: Obtain in-situ images taken by a monocular camera of an underwater observation platform, and based on prior knowledge in the fishing industry, decouple the fish species appearing in the images into multi-level classification labels and multiple combinations of visual attribute labels, and construct an image-text structured knowledge base; Step 2: Extract training and validation datasets for the detection model from the knowledge base, and divide the datasets into closed-class categories and open-world categories; Step 3: Receive a natural language text instruction input by a user to describe the dataset to be detected, and use a text encoder of a vision-language model to encode the natural language text instruction to generate a text feature vector of the target category; Step 4: Align the visual-text semantic space features of the image-text pair dataset using a pre-trained vision-language model to obtain a targeted auxiliary model; Step 5, construct a collaborative architecture containing a target model to be optimized and a targeted auxiliary model containing a regularization reference; dynamically and interactively fuse the visual features and the text features of the in-situ image; jointly optimize the target model through a composite loss composed of a multi-task learning module based on structured knowledge guidance and a feature space regularization loss guided by the targeted auxiliary model; Step 6, use the optimized target model to detect the fish in the underwater in-situ image without labels, locate the positions of the closed-set fish targets and the open-world fish targets, and classify them.

[0011] Further, the multi-level classification identifier includes a main classification identifier and a set of hierarchical classification identifiers, the main classification identifier is used to uniquely identify the category number and standard name of the fish class, and the hierarchical classification identifier is used to represent the field of biological hierarchical relationship, including the number and name of the order and family, and the visual attribute identifier is used to describe the attribute set of the observable morphological characteristics.

[0012] Further, the knowledge base includes an image sub-base and a text sub-base, the image sub-base is used to index and register the metadata of the in-situ image, and the text sub-base is used to provide a structured entry for each class in the fish dataset.

[0013] Further, the step 3 includes: embedding one or more pre-defined static text prompt templates in sequence to obtain a set of text sequences embedding one or more pre-defined static text prompt templates in sequence to obtain a set of text sequences inputting each text sequence into the text encoder of the visual-linguistic model to obtain a corresponding set of text feature vectors .

[0014] Further, the step 4 specifically includes: Step 4.1, initialize a pre-trained visual-linguistic model composed of an image encoder and a text encoder as a base model; Step 4.2, use the dataset of the closed-set classes and the text feature vectors of the target classes to form a graph-text pair for model training; Step 4.3, in the visual-text semantic space training, optimize the weights of the base model by conducting contrastive learning on the graph-text pair, maximize the similarity between the matched image visual features and text features, and minimize the similarity between the unmatched image visual features and text features, to learn the corresponding relationship between the visual concepts and text descriptions in the underwater fish field, and use the trained model as the targeted auxiliary model.

[0015] Further, the step 5 specifically includes: Step 5.1, the architecture of the target model adopts a visual-language model and a two-stage target detection framework to generate candidate regions by using a region proposal network in the target detection framework; Step 5.2, dynamically and interactively fusing visual features and text features of categories of candidate regions of the in-situ image by using a dynamic language perception module; Step 5.3, calculating a field task loss based on a structured knowledge guided multi-task learning module; Step 5.4, calculating a feature space regularization loss by calculating a distance between visual features extracted by the target model and the targeting auxiliary model for the same input image; Step 5.5, calculating the composite loss based on the field task loss and the feature space regularization loss as follows: ; wherein, is the calculated composite loss, is the field task loss, is the feature space regularization loss, and λ is a weighting coefficient.

[0016] Further, the step 5.2 specifically comprises: for each category c in the structured knowledge base, a dynamic prompt Prompt_c embedded with a learnable continuous vector V is constructed, and all dynamic prompts corresponding to all categories are input into a text encoder of the target model to generate a set of dynamic text features as text prototypes for the i-th candidate region feature a candidate region level cross-attention module is constructed: ; wherein, is a Query vector of the visual feature, 、 、 is a learned weight matrix, 、 are Key and Value vectors of the text feature, respectively; the attention weight of the candidate region level cross-attention module is represented as: ; wherein, is a normalized exponential function, is a transpose symbol of a matrix, is a key vector dimension; the fusion feature output by the candidate region level cross-attention module is represented as: ​​ .

[0017] Further, the step 5.3 specifically includes: Taking the candidate region features of the target model or linear projection thereof, the text prototype of the kth category is , and the temperature is , the cosine similarity between the candidate region features of the target model and the text prototype of the kth category is is calculated as: ; wherein, is a similarity calculation function; The hardest negative sample is used to select each The most similar top-K negative class set , and the main class contrast loss is calculated as: ; wherein, is the scaled cosine similarity between the ith candidate region and the text prototype of its positive class , and is the total number of candidate regions; The weighted sum of multiple cross-entropy loss terms is used to supervise the classification results of the order, family, and species levels, respectively, to obtain the predicted distribution and the real label of each level , and the cross-entropy is used to calculate the level loss as: ; wherein, is a level set, is a weight for different levels, is a cross-entropy loss function representing a standard; The multi-label binary cross-entropy loss is used to simultaneously predict multiple visual attributes of the target, to obtain the predicted label and the real label of the mth attribute of the ith candidate region, and the attribute loss is calculated as: ; wherein, is the total number of attributes, is a sigmoid activation function; The field task loss is calculated as: ;​ wherein, is a detection loss, , , , are weight coefficients of the main class contrast loss , the hierarchical loss , the attribute loss and the detection loss , respectively.

[0018] Further, the feature space regularization loss is calculated as follows: ; wherein, represents a feature vector extracted by the target model for the i-th candidate region, represents a feature vector extracted by the targeted auxiliary model for the i-th candidate region, is an L2 norm symbol.

[0019] In a second aspect, the present application provides an underwater in-situ fish unlabeled sample detection device, comprising a storage medium and a processor, the storage medium stores a computer program, and the computer program is executed by the processor to implement the above-mentioned method.

[0020] Beneficial effects: 1. The present application aligns the regional level visual features with the decomposable text semantics through the synergistic effect of the structured knowledge guidance and the dynamic language perception module; in the reasoning stage, the newly added class text prototype can be loaded on demand on the unlabeled in-situ observation image / video frame and participate in the regional level similarity calculation, without the need for retraining or changing the detection network structure, thereby realizing the matching detection of the class names or new attribute combinations not seen in the training stage; 2. The present application effectively solves the catastrophic forgetting problem existing in the field fine-tuning of the prior art. While improving the discriminability of visible classes, the general semantic structure from large-scale pre-training is preserved, which helps to maintain the recognition ability of unknown classes (classes not appearing in the training stage). The final obtained model has strong open set detection capability and can effectively identify new species not seen in training; at the same time, the model can successfully transfer knowledge to new application scenarios and implement cross-domain reasoning on unlabeled observation samples, thereby reducing the retraining and maintenance cost. BRIEF DESCRIPTION OF DRAWINGS

[0021] Figure 1 is a flowchart of the underwater in-situ fish unlabeled sample detection method of the embodiment of the present application; Figure 2 is a calculation process diagram of the composite loss of the embodiment of the present application. DETAILED DESCRIPTION

[0022] The present application will be further clarified by the following examples, which are intended to be purely exemplary of the present application. These examples should not be considered to limit the scope of the application.

[0023] As shown in Figure 1 and Figure 2 The embodiment of the present application provides an underwater in-situ fish label-free sample detection method, which comprises the following steps. Step 1, structured knowledge modeling: obtaining in-situ images captured by a monocular camera of an underwater observation platform, and based on prior knowledge in the field of fisheries, decoupling the fish species appearing in the images into multi-level classification labels and multiple combinations of visual attribute labels, and constructing an image-text structured knowledge base. Specifically, obtaining an in-situ observation video sequence captured by a monocular camera of an underwater observation platform, extracting representative images from the video sequence, and based on prior knowledge in the field of fisheries, decoupling the fish species appearing in the images into a main category label, a set of hierarchical classification labels, and a set of shared and combinable visual attribute labels, and constructing an image-text structured knowledge base. The structured knowledge base K is a pre-constructed database stored in JSON format, including an image sub-base and a text sub-base.

[0024] The image sub-base is used for indexing and metadata registration of in-situ images captured by a monocular camera of an underwater observation platform. The above metadata at least includes image identification, collection time, collection device information, etc. The text sub-base is used to provide a structured entry for each class c in the fish dataset. The main category label class_name is used to uniquely identify the class number and standard name (preferably at the species level) of the object class, such as the scientific name; the hierarchical classification label is used to represent the field of biological hierarchical relationship, at least including the number and name of order and family (optionally including genus / subfamily and other extended levels); the visual attribute label is used to describe the attribute set of observable morphological characteristics, at least including overall shape, main / auxiliary color, texture / pattern, key parts (such as dorsal fin / tail type), etc.

[0025] Step 2, creating an open vocabulary target detection dataset: extracting training and validation datasets for the detection model from the knowledge base, and dividing the dataset into closed set categories and open world categories. Specifically, representative images are extracted from the structured knowledge base image sub-library; 8000 frames of in-situ underwater images are manually extracted; the images are manually labeled using a labeling tool, and the training set and validation set are divided in a ratio of 8:2; the dataset is divided again according to the open vocabulary setting to obtain Base class (closed set category) and Novel class (open world category), wherein the Novel class does not participate in the training of the detection model, and is only used for zero-shot and / or open set evaluation in the validation / test stage.

[0026] Step 3, concept text instruction encoding: receiving user input of natural language text instructions describing the detection dataset, and using the text encoder of the visual-linguistic model to encode the natural language text instructions to generate text feature vectors of the target categories. Specifically, the class names in the target field are sequentially embedded in one or more predefined static text prompt templates, such as the English template "A photo of a [class_name]" (a photo of [class name]), to obtain a set of text sequences , input each text sequence into the text encoder of the visual-linguistic model to obtain a set of corresponding text feature vectors . In addition, when multiple templates are used for the same category, the category text prototype can be taken as : ; Step 4, construction of a targeted auxiliary model: using a pre-trained visual-linguistic model to perform visual-text semantic space feature alignment training on the image-text pair dataset to obtain a targeted auxiliary model. Specifically, the following steps are included: Step 4.1, initializing a pre-trained visual-linguistic model composed of an image encoder and a text encoder as a base model. The image encoder uses a ResNet-50 network structure, and the text encoder uses a Transformer structure.

[0027] Step 4.2, using the closed set category dataset divided in step 2 and the text feature vectors of the target categories generated in step 3 to form image-text pairs (Image-Text Pairs) for model training.

[0028] Step 4.3, in the visual-text semantic space training, the weights of the base model are optimized by contrastive learning on the image-text pairs, maximizing the similarity between matched image visual features and text features, and minimizing the similarity between unmatched image visual features and text features, to learn the correspondence between visual concepts and text descriptions in the underwater fish domain, and the trained model is used as a targeted auxiliary model for subsequent dynamic visual-language interactive training. The above scheme can effectively solve the Base-Novel Trade-off (BNT) problem. The stochastic gradient descent (SGD) optimizer is used for parameter update during model training, which is a preliminary and light fine-tuning. In the zero-shot detection setting, the matching probability of the visual features and text features of each class can be calculated as follows: ; where, is the calculated matching probability, is the natural exponential function, is the similarity calculation function, is the visual feature of the input image or candidate region.

[0029] Step 5, a collaborative architecture containing the target model to be optimized and the regularization reference of the targeted auxiliary model is constructed; the visual features of the in-situ image and the text features of the class are dynamically interactively fused; the structured knowledge guided multi-task learning module and the feature space regularization loss function guided by the targeted auxiliary model jointly constitute a composite loss to end-to-end jointly optimize the target model. Specifically as follows: Step 5.1, the architecture of the target model is constructed using a visual-language model (CLIP) and a two-stage target detection framework (Faster R-CNN). The targeted auxiliary model is the visual-language model obtained in step 4 and the parameters are frozen in this step. The target model to be optimized is a visual-language target detection model with updateable parameters. Specifically, this architecture can use the region proposal network (RPN) in the detection framework to generate candidate regions (ROI), and modify the recognition head (Recognition Head) into a visual-language alignment module that can perform region-level visual feature ( ) and text feature ( ) contrastive learning.

[0030] Step 5.2, design a dynamic language perception module, and utilize the dynamic language perception module to dynamically and interactively fuse the visual features and the text features of the candidate regions of the in-situ image. Step 5.2 specifically comprises: For each class c in the structured knowledge base, a dynamic prompt Prompt_c embedded with learnable continuous vectors V is constructed, for example: Prompt_c = { [V_context], "a photo of", [V_cls], V_cls, c.class_name, "a fish from the", [V_fam], V_fam, c.family, "family."}. Wherein [V_context], [V_cls], [V_fam] are continuous vectors optimized in the word embedding space together with other parameters of the model. And input all the dynamic prompts corresponding to the classes into the text encoder of the target model to generate a set of dynamic text features as text prototypes , which contains species prototypes, attribute prototypes and superordinate prototypes.

[0031] To realize the deep alignment of region-level visual features and text semantics, for the i-th candidate region feature , a candidate region-level cross-attention module is constructed: ; Wherein, is the Query vector of the visual feature, 、 、 is the learned weight matrix, 、 are the Key and Value vectors of the text feature respectively.

[0032] The attention weight of the candidate region-level cross-attention module is represented as: ; Wherein, is the normalization exponential function, is the transpose symbol of the matrix, is the key vector dimension.

[0033] The fusion feature output by the candidate region-level cross-attention module is represented as: .

[0034] ​​In preferred embodiments, a hierarchical prior mask can be imposed on the Key set according to the biological hierarchy information (masking non-ancestral / irrelevant prototypes) and a Top-K gate can be enabled to filter the text prototypes participating in the attention calculation, thereby improving training and inference efficiency.

[0035] Step 5.3, design a structured knowledge guided multi-task learning module, and calculate the field task loss based on the structured knowledge guided multi-task learning module. Step 5.3 specifically includes: a) main class contrast loss : a cross-modal InfoNCE loss function is adopted, which aims to pull the distance between the ROI feature and its positive class text prototype in the feature space, while pushing away the distance between the ROI feature and its negative class text prototype. In a preferred embodiment, to improve the model's ability to recognize fine-grained classes, negative samples can be preferentially sampled from neighboring classes of the same family or same purpose as the positive class (i.e., difficult negative sample mining). Take the candidate region feature of the target model or its linear projection, the text prototype of the kth class is , the temperature is , calculate the cosine similarity between the candidate region feature of the target model and the text prototype of the kth class : ; wherein is a similarity calculation function.

[0036] Hard-Negatives (difficult negative samples) are used to select the top-K negative class set for each , and calculate the main class contrast loss : ; wherein is the scaled cosine similarity between the ith candidate region and the text prototype of its positive class , and is the total number of candidate regions.

[0037] b) hierarchical loss : a weighted sum of multiple cross-entropy loss terms is used to supervise the classification results of the order, family, and species levels, respectively, to obtain the predicted distribution and the true label of each level , and the cross-entropy is used to calculate the hierarchical loss : ; wherein ​For the hierarchical set, , For the weights of different hierarchies, to embody the importance of different hierarchies (for example ), is the cross-entropy loss function representing the criterion. In a further preferred embodiment, to impose a higher penalty on the out-of-class error on the classification hierarchy, a path consistency constraint can be added, for example, by KL divergence regularization, to constrain the consistency between the predicted hierarchical path distribution and the true hierarchical path distribution.

[0038] c) Attribute loss : A multi-label binary cross-entropy loss (BCE-with-logits Loss) is used to simultaneously predict multiple visual attributes of the target, obtaining the predicted label of the mth attribute of the ith candidate region and the true label , and calculating the attribute loss is: ; wherein, is the total number of attributes, is the sigmoid activation function.

[0039] In a preferred embodiment, to alleviate the long-tail attribute distribution problem in the data set, a class-balancing strategy or focal loss (Focal Loss) can be combined.

[0040] d) Detection loss : Contains the bounding box regression loss.

[0041] The field task loss is calculated as: ; wherein, , , , are the weight coefficients of the main class contrast loss , the hierarchical loss , the attribute loss and the detection loss , respectively.

[0042] To cope with the dynamic changes of species and the deployment requirements across sea areas faced by underwater monitoring, the present application adopts a targeted regularization mechanism to suppress feature space drift and maintain cross-domain generalization ability during fine-tuning in underwater in-situ scenes.

[0043] Specifically, the feature space regularization loss is constructed by calculating the distance between the visual features extracted by the target model and the targeted auxiliary model for the same input image. This loss imposes constraints on the learning direction of the target model, preventing it from deviating too much from a general visual knowledge space with strong generalization ability when learning domain-specific knowledge.

[0044] Step 5.4, calculate the feature space regularization loss by calculating the distance between the visual features extracted by the target model and the targeted auxiliary model for the same input image. The feature space regularization loss is calculated as follows: ; where represents the visual feature vector extracted by the target model for the ith candidate region, represents the visual feature vector extracted by the targeted auxiliary model for the ith candidate region, is the L2 norm symbol.

[0045] Step 5.5, calculate the composite loss based on the domain task loss and the feature space regularization loss as follows: ; where, is the calculated composite loss, and λ is the weighting coefficient.

[0046] Step 6, unlabeled sample detection and inference: use the optimized target model to detect fish in underwater in-situ images, locate the positions of closed-set fish targets and open-world fish targets and classify them. Specifically as follows: a) Open set detection mode based on preset categories: In this mode, first, an underwater in-situ image to be detected is obtained as input. Second, the target model calculates the semantic similarity of its visual features with a preset library containing all Base and Novel text prototypes. Finally, this step can output the position and category information of the closed-set fish targets (i.e. labeled training samples) and open-world fish targets (i.e. unlabeled samples) in the image.

[0047] b) Open vocabulary detection mode based on user instructions: In this mode, first, a natural language text instruction describing one or more arbitrary target categories is received. Second, the text instruction is encoded into a set of temporary, task-driven text feature vectors in real time using a text encoder of the target model. Then, an image to be detected is obtained, and the target model detects the arbitrary target in the user instruction by calculating the similarity between the visual features of the image and the set of temporary text feature vectors. Finally, this step can output the location and category information of the target in the image that matches the user instruction.

[0048] Based on the above embodiments, those skilled in the art can easily understand that the present application also provides an underwater in-situ fish label-free sample detection device, comprising a storage medium and a processor, the storage medium stores a computer program, and the computer program is executed by the processor to realize the above-mentioned method.

[0049] Specific cases: (1) Experimental environment and data set a) Experimental environment Experimental platform. The experiment was carried out on a server equipped with Intel(R) Xeon(R) Gold 6342 CPU @ 2.80GHz and 4 NVIDIA A100-PCIE-40GB GPU, running Ubuntu 20.04 operating system. The deep learning framework uses PyTorch 1.12.

[0050] Evaluation index. In target detection, IoU (Intersection over Union, IoU) is used to determine whether the detection box belongs to the target to be detected. IoU represents the overlap between the detected location information and the real target box . At the same time, IoU is used to predict the accuracy of the location information of the result, and the calculation expression is as follows: ; Average precision (AP) is a core indicator to measure the detection performance of a single class. mAP represents the average of AP of all classes, and the expression is as follows: ; K represents the number of target categories to be detected, mAP@0.5 (or AP50) represents the mAP value calculated when the IoU threshold is fixed at 0.5. mAP@0.75 (or AP75) represents the mAP value calculated when the IoU threshold is fixed at 0.75. mAP (or AP) is obtained by calculating the AP of each category at 10 different IoU thresholds (from 0.5 to 0.95, step size 0.05), and then averaging all results.

[0051] b) dataset construction The experimental data is obtained from the underwater observation platform placed in Weizhou Island sea area (east longitude 109°7'32'', north latitude 21°4'57'') and Yutai pasture in Weihai, Shandong (east longitude 121°55'50'', north latitude 37°28'23''). The in-situ observation video sequence is obtained by the monocular camera of the underwater observation platform, and representative images are extracted from the video sequence. According to the method of the present application, an image-text structured knowledge base is constructed for 28 fish species in step one, and an open vocabulary object detection dataset FishDet is constructed in step two, obtaining 23 Base classes and 5 Novel classes.

[0052] (2) Experimental results The FishDet dataset is evaluated strictly in closed set and open set. Table 1 is the detection result of the target model on the FishDet dataset in the zero-shot setting. In terms of overall performance, the model achieves an AP value of up to 95.9% on 23 closed set classes and 54.9% on 5 open world classes. It is worth noting that the performance of conventional methods on the zero-shot task usually collapses to near zero, which proves that the method of the present application has made substantial progress in overcoming catastrophic forgetting and solving the BNT problem.

[0053] Table 1 is the zero-shot result of the target model on the FishDet dataset ; From the experimental results in Table 2, it can be seen that the method of the present application performs extremely excellent and robust performance on the closed set task. Among them, in 8 classes such as black snake, sea cucumber, chest spot snapper, and sky boat fish, the AP value reaches 100%, and the AP value of individual class Chauliodus maculatus is only 50.50%, which may be due to the high similarity between the species and the coral reef background, resulting in the visual features not obvious, easy to miss or confused with the background.

[0054] Table 2 is the zero-shot result of the target model on the FishDet dataset Base class ; From the experimental results in Table 3, it can be seen that the method of the present application exhibits excellent generalization ability in the open set zero sample detection task. The model exhibits excellent zero-shot reasoning ability on multiple Novel classes, achieving AP values of up to 75.2%, 58.6% and 50.5% on the D. punctatus, A. taeniata and S. chrysops respectively. This indicates that the model learns the unique knowledge of the closed set classes while retaining and utilizing the transferable and general visual semantic concepts (e.g. "red", "stripes", "spindle-shaped body", etc.) learned from visual-textual semantic space training, thereby identifying completely new species. Secondly, the analysis of the lower performance classes further reveals the challenge of the task and the potential of the method of the present application. For example, the AP value of the A. longipinnis is only 34.76%. This may be due to the "semantic confusion" of the visual features of these new classes with one or more closed set classes in the training set, making it difficult for the model to perform accurate zero-shot reasoning.

[0055] Table 3: Zero-shot results of the target model on the Novel classes of the FishDet dataset ; In summary, the experimental results on the open set task indicate that the method of the present application can effectively maintain the generalization ability of the model, enabling the model to move from "closed set" to "open set", providing a feasible technical path for realizing open world perception in underwater environments.

[0056] The above description is only the preferred embodiments of the present application, and it should be noted that other non-specifically described parts belong to the prior art or common knowledge for ordinary skilled in the art. Without departing from the principles of the present application, several improvements and refinements can be made, and these improvements and refinements should be considered as the protection scope of the present application.

Claims

1. A method for detecting unlabeled fish samples in situ underwater, characterized in that, include: Step 1: Acquire in-situ images captured by a monocular camera on an underwater observation platform, and based on prior knowledge in the fisheries field, decouple the fish categories appearing in the images into multi-level classification labels and multi-combination visual attribute labels to construct an image-text structured knowledge base. Step 2: Extract the training and validation datasets for the detection model from the knowledge base, and divide the datasets into closed-set categories and open-world categories; Step 3: Receive the natural language text instruction describing the dataset to be detected input by the user, and encode the natural language text instruction using the text encoder of the vision-language model to generate the text feature vector of the target category. Step 4: Use a pre-trained visual-language model to train the visual-text semantic space feature alignment of the image-text dataset to obtain a targeted auxiliary model; Step 5: Construct a collaborative architecture that includes the target model to be optimized and a targeted auxiliary model with regularization reference; dynamically and interactively fuse the visual features of the in-situ image and the text features of the category; the multi-task learning module guided by structured knowledge and the feature space regularization loss function guided by the targeted auxiliary model together form a composite loss to jointly optimize the target model end-to-end. Step 6: Use the optimized target model to perform unlabeled detection of fish appearing in underwater in-situ images, locate the positions of closed-set fish targets and open-world fish targets, and classify them.

2. The method for detecting unlabeled fish samples in situ underwater according to claim 1, characterized in that, The multi-level classification identifier includes a main category identifier and a set of hierarchical classification identifiers. The main category identifier is used to uniquely identify the category number and standard name of the fish category. The hierarchical classification identifier is used to represent fields of biological hierarchical relationship, including the order and family numbers and names. The visual attribute identifier is used to describe the set of attributes that describe observable morphological characteristics.

3. The method for detecting unlabeled fish samples in situ underwater according to claim 1, characterized in that, The knowledge base includes an image sub-base and a text sub-base. The image sub-base is used to index and register metadata for in-situ images, and the text sub-base is used to provide a structured entry for each category in the fish dataset.

4. The method for detecting unlabeled fish samples in situ underwater according to claim 1, characterized in that, Step 3 includes: Category names within the target domain By sequentially embedding one or more predefined static text prompt templates, a set of text sequences is obtained. Each text sequence is input into the text encoder of the vision-language model to obtain a set of corresponding text feature vectors. .

5. The method for detecting unlabeled fish samples in situ underwater according to claim 1, characterized in that, Step 4 specifically includes: Step 4.1: Initialize a pre-trained visual-language model consisting of an image encoder and a text encoder as the base model; Step 4.2: Construct image-text pairs for model training using the dataset of the closed set category and the text feature vector of the target category; Step 4.3: In the visual-text semantic space training, the weights of the base model are optimized by performing comparative learning on the image-text pairs to maximize the similarity between matching image visual features and text features, and minimize the similarity between mismatched image visual features and text features, so as to learn the correspondence between visual concepts and text descriptions in the underwater fish domain, and use the trained model as a targeted auxiliary model.

6. The method for detecting unlabeled fish samples in situ underwater according to claim 1, characterized in that, Step 5 specifically includes: Step 5.1: The architecture of the target model is constructed using a vision-language model and a two-stage object detection framework, which utilizes the region proposal network in the object detection framework to generate candidate regions. Step 5.2: Use the dynamic language perception module to dynamically and interactively fuse the visual features of candidate regions and the text features of categories in the in-situ image; Step 5.3: Calculate the domain task loss based on the multi-task learning module guided by structured knowledge; Step 5.4: Calculate the feature space regularization loss by calculating the distance between the visual features extracted by the target model and the targeting auxiliary model for the same input image; Step 5.5: Calculate the composite loss based on the domain task loss and the feature space regularization loss as follows: ; in, For the calculated composite loss, Loss due to domain mission λ is the feature space regularization loss, and λ is the tradeoff coefficient.

7. The method for detecting unlabeled fish samples in situ underwater according to claim 6, characterized in that, Step 5.2 specifically includes: For each category c in the structured knowledge base, a dynamic cue Prompt_c is constructed, which embeds a learnable continuous vector V. The set of dynamic cues corresponding to all categories is then input into the text encoder of the target model to generate a set of dynamic text features as text prototypes. For the i-th candidate region features Construct a candidate region-level cross-attention module: ; in, The query vector represents the visual features. , , The learned weight matrix, , These are the key and value vectors for the text features, respectively. The attention weights of the candidate region-level cross-attention module Represented as: ; in, For normalized exponential functions, The transpose symbol for a matrix. The dimension of the key vector; The fusion features output by the candidate region-level cross-attention module Represented as: 。 8. The method for detecting unlabeled fish samples in situ underwater according to claim 7, characterized in that, Step 5.3 specifically includes: Extract candidate region features from the target model Or its linear projection, the text prototype of the k-th category is Temperature is Calculate the candidate region features of the target model The text prototype of the k-th category is cosine similarity for: ; in, This is a similarity calculation function; Using hard-to-bear samples, for each Select the most similar top-K negative class set Calculate the main class contrast loss for: ; in, For the i-th candidate region and its positive class Scaled cosine similarity between text prototypes The total number of candidate regions; The classification results at the order, family, and species levels are supervised by a weighted sum of multiple cross-entropy loss terms to obtain the results at each level. Predicted distribution and real labels And use cross-entropy to calculate hierarchical loss. for: ; in, For hierarchical sets, For different levels of weight, This represents the standard cross-entropy loss function; A multi-label binary cross-entropy loss method is used to simultaneously predict multiple visual attributes of the target, obtaining the predicted label of the m-th attribute of the i-th candidate region. and real labels And calculate attribute loss. for: ; in, The total number of attributes. It is the sigmoid activation function; Computational domain task loss for: ; in, To detect the loss, , , , The main class contrast loss is respectively Hierarchical loss Attribute loss and detection loss The weighting coefficients.

9. The method for detecting unlabeled fish samples in situ underwater according to claim 6, characterized in that, The feature space regularization loss The calculation method is as follows: ; in, Let represent the visual feature vector extracted by the target model for the i-th candidate region. This represents the visual feature vector extracted by the targeting-assisted model for the i-th candidate region. This is the L2 norm symbol.

10. An underwater in-situ tagless fish sample detection device, comprising a storage medium and a processor, wherein the storage medium stores a computer program, characterized in that, When the computer program is executed by a processor, it is used to implement the method described in any one of claims 1-9.

Citation Information

Patent Citations

  • Multi-label microblog text classification method based on semi-supervised learning

    CN113254599A

  • Underwater fish target identification method and device, equipment and storage medium

    CN117197542A

  • Few-sample open set recognition method based on bimodal semantic feature task-driven learning

    CN118194017A

  • Rapid and label-free procedure for microbial community screening and profiling

    US20170138845A1

  • Visual question answering using model trained on unlabeled videos

    US20220067546A1

Cited By

  • Target detection method based on hierarchical prototype learning and related equipment

    CN121725230A

  • Target detection method based on hierarchical prototype learning and related device

    CN121725230B