Method and apparatus for fish target recognition

CN122391699BActive Publication Date: 2026-09-15UNIV OF SCI & TECH BEIJING +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610423592.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-04-01
Publication Date
2026-09-15
Estimated Expiration
2046-04-01

AI Technical Summary

Technical Problem

[0006]为了解决现有技术中在复杂水体场景下鱼类目标识别稳定性低的问题,本发明实施例提供了一种用于鱼类目标识别的方法与装置

Benefits of technology

[0016]本发明实施例提供的技术方案带来的有益效果至少包括:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122391699B_ABST
    Figure CN122391699B_ABST
Patent Text Reader

Abstract

The application provides a method and device for fish target recognition, and relates to the technical field of computer vision, and the method comprises the following steps: acquiring multi-view image information to be recognized and the geometric relationship of the multi-view image information; for each view image information, determining a detection frame and a fish target mask through a target positioning network; generating a cross-view fish target set according to the geometric relationship, the detection frame and the fish target mask; performing self-supervised feature learning based on the cross-view fish target set to generate fish target level feature representation, and using the fish target level feature representation as semantic prior information to perform differentiable rendering optimization of a three-dimensional Gaussian splash model to generate a three-dimensional Gaussian representation; determining the species category and three-dimensional specification level information of the fish target based on the three-dimensional Gaussian representation through a multi-task classification head, and combining the two-dimensional specification level information to determine the specification level information of the fish target. Therefore, fish target recognition can be realized in a complex water body scene, and the stability of the recognition method is relatively high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and in particular to a method and apparatus for fish target recognition. Background Technology

[0002] As aquaculture develops towards large-scale, intensive, and intelligent operations, fish species identification and size grading have become crucial technical aspects of aquaculture management, product grading and pricing, and quality control. Accurately obtaining key information such as fish species, body length, and size grade is essential for improving aquaculture management efficiency and ensuring product quality throughout the production and distribution processes.

[0003] Currently, fish identification and size determination still mainly rely on manual observation, supplemented by automated identification systems based on two-dimensional images. Manual methods depend on operator experience for species differentiation and size estimation, resulting in low efficiency, high subjectivity, and inconsistent judgment standards, making them unsuitable for the high-efficiency processing requirements of large-scale aquaculture. While existing automated methods based on two-dimensional images have improved processing efficiency to some extent, their reliance on single-view two-dimensional image information limits their adaptability to occlusion and non-rigid posture changes in complex aquatic environments.

[0004] In complex aquatic scenarios, fish frequently bend and turn during swimming, accompanied by mutual occlusion between individuals. Under a single viewpoint, local information of the target is easily lost or the morphological expression is distorted, resulting in insufficient stability of target recognition and size determination results based on two-dimensional images.

[0005] Therefore, there is an urgent need for a fish target recognition method that can maintain stable performance in complex aquatic scenarios. Summary of the Invention

[0006] To address the problem of low stability in fish target recognition under complex aquatic conditions in existing technologies, this invention provides a method and apparatus for fish target recognition. The technical solution is as follows: On the one hand, a method for fish target identification is provided, the method comprising: The process involves acquiring multi-view image information of the fish target to be identified and the geometric relationship of the multi-view image information; then, using a target localization network, detecting and locating the fish target in each view of the multi-view image information to determine the detection box and fish target mask in each view of the image information. Based on the geometric relationship, the detection boxes and fish target masks in the image information of each viewpoint are matched and associated to generate a cross-viewpoint fish target set. Based on the cross-viewpoint fish target set, self-supervised feature learning is performed to generate fish target-level feature representations. The fish target-level feature representation is used as semantic prior information for differentiable rendering optimization of the three-dimensional Gaussian splash model to generate a three-dimensional Gaussian representation of the fish target. Based on the three-dimensional Gaussian representation, the species category and three-dimensional size level information of the fish target are determined by a multi-task classification head. The multi-task classification head is generated by multi-modal training based on the multi-view image information of the fish target in the training dataset, the fish target-level feature representation determined by the multi-view image information of the fish target in the training dataset, and the three-dimensional Gaussian representation determined by the multi-view image information of the fish target in the training dataset. Based on the multi-view image information of the fish target to be identified and the geometric relationship, the two-dimensional size level information of the fish target is determined, and the size level information of the fish target is determined by combining the three-dimensional size level information and the two-dimensional size level information.

[0007] Optionally, the multi-task classification head is generated through multimodal training in the following manner: For multi-view image information of fish targets in the training dataset, the target localization network is used to detect and locate fish targets based on the image information of each view, and to determine the detection box and fish target mask in the image information of each view. Based on the geometric relationship of multi-view image information, the detection boxes and fish target masks in each view image information are matched and associated to generate a cross-view fish target set. Based on the cross-view fish target set, self-supervised feature learning is performed to generate fish target-level feature representations. The fish target-level feature representation is used as semantic prior information for differentiable rendering optimization of the three-dimensional Gaussian splash model to generate a three-dimensional Gaussian representation of the fish target. Multimodal training is performed based on the three-dimensional Gaussian representation of fish targets, multi-view image information of fish targets, and fish target-level feature representation to generate semantic embedding representations corresponding to fish targets; Based on aquaculture application scenarios, species category descriptions and size level descriptions corresponding to fish targets are constructed. These species category descriptions and size level descriptions are used as semantic reference information, and semantic consistency constraints are constructed between the fish target-level feature representations in the semantic embedding representation set.

[0008] Optionally, the loss for constructing semantic consistency constraints is expressed as: ,in, This represents the target-level features of the fish species. This represents the semantic embedding representation.

[0009] Optionally, the step of using the fish target-level feature representation as semantic prior information to perform differentiable rendering optimization of the 3D Gaussian splash model to generate a 3D Gaussian representation of the fish target includes: The fish target-level feature representation is used as semantic prior information to construct a three-dimensional Gaussian set of fish targets. The three-dimensional Gaussian set is composed of multiple three-dimensional Gaussian units, and each three-dimensional Gaussian unit contains the spatial center position, spatial distribution parameters and appearance attributes. Based on the geometric relationship, the three-dimensional Gaussian set is projected to each viewpoint through a differentiable rendering mechanism to generate a rendered image at the corresponding viewpoint. The pixel domain reprojection error between the multi-view image information and the rendered image is constructed. The spatial parameters and appearance parameters of the three-dimensional Gaussian unit are iteratively optimized by minimizing the pixel domain reprojection error to generate a three-dimensional Gaussian representation of the fish target.

[0010] Optionally, the method provided in this embodiment of the invention further includes: An encoder obtained through self-supervised feature learning via a masked autoencoder is used as a feature extraction network to extract features from the multi-view image information, thereby obtaining a first fish target-level feature representation. The encoder is used as a feature extraction network to extract features from the rendered image and obtain a second fish target-level feature representation. In the feature space, a consistency constraint is applied to the first fish target-level feature representation and the second fish target-level feature representation to construct the feature domain semantic consistency error; The pixel domain reprojection error and the feature domain semantic consistency error are jointly optimized to construct a joint loss function: , in, This refers to pixel-domain reprojection error. For the semantic consistency error of the feature domain, These are the weighting coefficients; Minimize the joint loss function so that the 3D Gaussian representation of the fish target generated by the 3D Gaussian splash model maintains the semantic expression consistent with the fish target-level feature representation.

[0011] Optionally, the step of generating fish target-level feature representations based on the cross-view fish target set through self-supervised feature learning includes: Based on the cross-view fish target set, cropped images of the same fish target at each view are obtained, and the fish target mask is applied to the cropped images to determine the fish target images in the cropped images; A random masking operation is performed on the fish target image to obtain a masked input image. The masked input image is then input into a mask autoencoder model, which performs feature encoding on the masked image. The decoder then reconstructs the occluded areas in the fish target image to generate a reconstructed image. By minimizing the reconstruction loss between the original image of the fish target and the reconstructed image, self-supervised feature learning based on context information is performed. Represented as: ,in, This represents the original image of the fish target. This refers to the reconstructed image.

[0012] Optionally, a cross-view consistency mechanism is introduced in the self-supervised feature learning process to aggregate the view features extracted by the mask autoencoder from the same fish target under different viewpoints, and to constrain the consistency of the feature representations of the same fish target under different viewpoints through cross-view consistency loss. The cross-perspective consistency loss Represented as: ,in, Indicates the fish target in the field of view The following fish target-level feature representation, Indicates the fish target in the field of view The fish target-level features below, the view index And satisfy ; The overall optimization objective of the masked autoencoder in the self-supervised feature learning stage is: ,in, The weighting coefficients for cross-perspective consistency constraints.

[0013] On the other hand, embodiments of the present invention also provide an apparatus for fish target identification, wherein the apparatus for fish target identification is used to implement the method for fish target identification provided in embodiments of the present invention, the apparatus comprising: The acquisition module is used to acquire multi-view image information of the fish target to be identified and the geometric relationship of the multi-view image information. For each view image information in the multi-view image information, the fish target is detected and located by the target localization network, and the detection box and fish target mask in each view image information are determined. The first generation module is used to match and associate the detection boxes and fish target masks in the image information from each viewpoint according to the geometric relationship, generate a cross-viewpoint fish target set, and perform self-supervised feature learning based on the cross-viewpoint fish target set to generate fish target-level feature representations. The second generation module is used to use the fish target-level feature representation as semantic prior information to perform differentiable rendering optimization of the three-dimensional Gaussian splash model, and generate a three-dimensional Gaussian representation of the fish target. The first determining module is used to determine the species category and three-dimensional size level information of the fish target through a multi-task classification head based on the three-dimensional Gaussian representation. The multi-task classification head is generated by multi-modal training based on the multi-view image information of the fish target in the training dataset, the fish target-level feature representation determined based on the multi-view image information of the fish target in the training dataset, and the three-dimensional Gaussian representation determined based on the multi-view image information of the fish target in the training dataset. The second determining module is used to determine the two-dimensional size grade information of the fish target based on the multi-view image information of the fish target to be identified and the geometric relationship, and to determine the size grade information of the fish target by combining the three-dimensional size grade information and the two-dimensional size grade information.

[0014] On the other hand, embodiments of the present invention also provide a device for fish target identification, the device for fish target identification comprising: processor; A memory storing computer-readable instructions, which, when executed by the processor, implement the method provided in the embodiments of the present invention.

[0015] On the other hand, embodiments of the present invention also provide a computer-readable storage medium storing program code, which can be invoked by a processor to execute the method provided in the embodiments of the present invention.

[0016] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: This invention achieves robust detection and consistency matching of fish targets in complex occlusion and pose change scenarios through multi-view target localization and cross-view association mechanisms. Building upon this, it introduces cross-view self-supervised feature learning to generate fish target-level feature representations with viewpoint invariance and semantic consistency without manual annotation. These features are used as semantic priors to guide the differentiable rendering optimization of the 3D Gaussian splash model, achieving collaborative modeling of geometric structure and semantic features, overcoming the shortcomings of traditional methods where 3D reconstruction and semantic representation are separated. A multi-task classification head is constructed based on the 3D Gaussian representation, simultaneously outputting species category and 3D size level. Combined with a 3D-2D dual-path size reading fusion mechanism, the reliability and fault tolerance of size determination are effectively improved. Finally, an integrated processing flow of "3D reconstruction—species identification—size assessment" is formed, significantly reducing reliance on manual annotation, enhancing cross-environment generalization ability, and providing high-precision and high-stability key technical support for intelligent management of aquaculture. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart of a method for fish target recognition provided by an embodiment of the present invention; Figure 2 This is an overall flowchart of a method for fish target recognition provided in an embodiment of the present invention; Figure 3 This is a three-dimensional Gaussian splash reconstruction and semantic co-optimization graph provided in an embodiment of the present invention; Figure 4 This is a self-supervised pre-training graph of a mask autoencoder provided in an embodiment of the present invention; Figure 5 This is a three-dimensional fish target-level joint identification and size discrimination map provided by an embodiment of the present invention; Figure 6 This is a schematic diagram illustrating the principle of a multi-view synchronous imaging device for observing fish targets, as provided in an embodiment of the present invention. Figure 7 This is a data collection diagram provided in an embodiment of the present invention; Figure 8 This is a data preprocessing and perspective consistency correction diagram provided in an embodiment of the present invention; Figure 9 This is a fish target localization and cross-view correlation map provided in an embodiment of the present invention; Figure 10 This is a semantic alignment and parameter fine-tuning diagram of a large-scale multimodal model provided in an embodiment of the present invention; Figure 11 This is a schematic diagram of the structure of a device for fish target identification provided in an embodiment of the present invention; Figure 12 This is a schematic diagram of the structure of a device for fish target identification provided in an embodiment of the present invention. Detailed Implementation

[0019] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0020] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0021] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.

[0022] In this embodiment of the invention, sometimes a subscript such as W1 may be mistakenly written as a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.

[0023] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0024] To address the problem of low stability in fish target recognition under complex aquatic conditions in existing technologies, this invention provides a method and apparatus for fish target recognition. The technical solution is as follows: like Figure 1 As shown, this embodiment of the invention provides a method for fish target identification, the method comprising: S1. Obtain multi-view image information of the fish target to be identified and the geometric relationship of the multi-view image information. For each view image information in the multi-view image information, detect and locate the fish target through a target localization network, and determine the detection box and fish target mask in each view image information.

[0025] S2. Based on the geometric relationship, the detection boxes and fish target masks in the image information of each viewpoint are matched and associated to generate a cross-viewpoint fish target set. Based on the cross-viewpoint fish target set, self-supervised feature learning is performed to generate fish target-level feature representations.

[0026] S3. The fish target-level feature representation is used as semantic prior information to perform differentiable rendering optimization of the three-dimensional Gaussian splash model, generating a three-dimensional Gaussian representation of the fish target.

[0027] S4. Based on the three-dimensional Gaussian representation, the species category and three-dimensional size level information of the fish target are determined by a multi-task classification head. The multi-task classification head is generated by multi-modal training based on the multi-view image information of the fish target in the training dataset, the fish target-level feature representation determined by the multi-view image information of the fish target in the training dataset, and the three-dimensional Gaussian representation determined by the multi-view image information of the fish target in the training dataset.

[0028] S5. Based on the multi-view image information of the fish target to be identified and the geometric relationship, determine the two-dimensional size level information of the fish target, and combine the three-dimensional size level information and the two-dimensional size level information to determine the size level information of the fish target.

[0029] Optionally, the multi-task classification head is generated through multimodal training in the following manner: For multi-view image information of fish targets in the training dataset, the target localization network is used to detect and locate fish targets based on the image information of each view, and to determine the detection box and fish target mask in the image information of each view. Based on the geometric relationship of multi-view image information, the detection boxes and fish target masks in each view image information are matched and associated to generate a cross-view fish target set. Based on the cross-view fish target set, self-supervised feature learning is performed to generate fish target-level feature representations. The fish target-level feature representation is used as semantic prior information for differentiable rendering optimization of the three-dimensional Gaussian splash model to generate a three-dimensional Gaussian representation of the fish target. Multimodal training is performed based on the three-dimensional Gaussian representation of fish targets, multi-view image information of fish targets, and fish target-level feature representation to generate semantic embedding representations corresponding to fish targets; Based on aquaculture application scenarios, species category descriptions and size level descriptions corresponding to fish targets are constructed. These species category descriptions and size level descriptions are used as semantic reference information, and semantic consistency constraints are constructed between the fish target-level feature representations in the semantic embedding representation set.

[0030] Optionally, the loss for constructing semantic consistency constraints is expressed as: ,in, This represents the target-level features of the fish species. This represents the semantic embedding representation.

[0031] Optionally, the step of using the fish target-level feature representation as semantic prior information to perform differentiable rendering optimization of the 3D Gaussian splash model to generate a 3D Gaussian representation of the fish target includes: The fish target-level feature representation is used as semantic prior information to construct a three-dimensional Gaussian set of fish targets. The three-dimensional Gaussian set is composed of multiple three-dimensional Gaussian units, and each three-dimensional Gaussian unit contains the spatial center position, spatial distribution parameters and appearance attributes. Based on the geometric relationship, the three-dimensional Gaussian set is projected to each viewpoint through a differentiable rendering mechanism to generate a rendered image at the corresponding viewpoint. The pixel domain reprojection error between the multi-view image information and the rendered image is constructed. The spatial parameters and appearance parameters of the three-dimensional Gaussian unit are iteratively optimized by minimizing the pixel domain reprojection error to generate a three-dimensional Gaussian representation of the fish target.

[0032] Optionally, the method provided in this embodiment of the invention further includes: An encoder obtained through self-supervised feature learning via a masked autoencoder is used as a feature extraction network to extract features from the multi-view image information, thereby obtaining a first fish target-level feature representation. The encoder is used as a feature extraction network to extract features from the rendered image and obtain a second fish target-level feature representation. In the feature space, a consistency constraint is applied to the first fish target-level feature representation and the second fish target-level feature representation to construct the feature domain semantic consistency error; The pixel domain reprojection error and the feature domain semantic consistency error are jointly optimized to construct a joint loss function: ,in, This refers to pixel-domain reprojection error. For the semantic consistency error of the feature domain, These are the weighting coefficients; Minimize the joint loss function so that the 3D Gaussian representation of the fish target generated by the 3D Gaussian splash model maintains the semantic expression consistent with the fish target-level feature representation.

[0033] Optionally, the step of generating fish target-level feature representations based on the cross-view fish target set through self-supervised feature learning includes: Based on the cross-view fish target set, cropped images of the same fish target at each view are obtained, and the fish target mask is applied to the cropped images to determine the fish target images in the cropped images; A random masking operation is performed on the fish target image to obtain a masked input image. The masked input image is then input into a mask autoencoder model, which performs feature encoding on the masked image. The decoder then reconstructs the occluded areas in the fish target image to generate a reconstructed image. By minimizing the reconstruction loss between the original image of the fish target and the reconstructed image, self-supervised feature learning based on context information is performed. Represented as: ,in, This represents the original image of the fish target. This refers to the reconstructed image.

[0034] Optionally, a cross-view consistency mechanism is introduced in the self-supervised feature learning process to aggregate the view features extracted by the mask autoencoder from the same fish target under different viewpoints, and to constrain the consistency of the feature representations of the same fish target under different viewpoints through cross-view consistency loss. The cross-perspective consistency loss Represented as: ,in, Indicates the fish target in the field of view The following fish target-level feature representation, Indicates the fish target in the field of view The fish target-level features below, the view index And satisfy ; The overall optimization objective of the masked autoencoder in the self-supervised feature learning stage is: ,in, The weighting coefficients for cross-perspective consistency constraints.

[0035] In some implementations, embodiments of the present invention provide an exemplary implementation.

[0036] In an exemplary implementation, such as Figures 2 to 10 As shown, this embodiment of the invention uses a multi-view synchronous imaging device to acquire synchronized image data of a fish target from at least two different viewpoints at the same time, and obtains the intrinsic and extrinsic parameters of the corresponding viewpoint cameras and the initial three-dimensional information output by the imaging device. The number of viewpoints in the multi-view synchronous imaging device is not limited, and its number of viewpoints is... ,and That is, a multi-view synchronous imaging device can be a binocular, trinocular, quadrinocular, pentanocular or more-view imaging structure.

[0037] In practical applications, distortion correction, brightness and color consistency processing, and resolution and scale unification can be sequentially performed on the multi-view image information acquired by the multi-view synchronous imaging device to obtain standardized multi-view image information.

[0038] In the image information from various perspectives, a target localization network oriented towards fish targets is used to detect and locate fish targets, obtaining corresponding detection boxes and fish target masks. Based on the known geometric relationships of the multi-view cameras, the detection results of the same fish target under different perspectives are matched and associated to form a cross-view fish target set. For the cross-view fish target set, cropped images of fish targets under each perspective are extracted, and the target region pixels are preserved by combining the fish target mask. After performing random masking on the cropped images, the images are input into a mask autoencoder for self-supervised feature learning. Through image reconstruction constraints and cross-view feature consistency constraints, a cross-view fish target-level feature representation for characterizing the semantic consistency of the same fish target under different perspectives is obtained.

[0039] The cross-view fish target-level feature representation is introduced as semantic prior information into the differentiable rendering optimization process of the 3D Gaussian splash model. Based on the known geometric relationships of the multi-view cameras, the 3D Gaussian set is rendered differentlyiably under each viewpoint. The 3D Gaussian parameters are optimized through joint constraints of pixel domain reprojection error and feature domain semantic consistency error, thereby constructing a 3D Gaussian representation of the fish target. After completing the 3D representation construction, a large-scale multimodal pre-trained model is introduced to semantically align the multi-view real 2D images of the fish target, the rendered images obtained from the 3D Gaussian set, and the fish target-level feature representation. Relevant parameters are fine-tuned to match the obtained semantic features with the species category semantics and size grade semantics of the fish target. Based on the obtained 3D fish target representation, a 3D fish target-level multi-task classification head is constructed, simultaneously outputting the species category and size grade of the fish target. Two size readings are generated based on the 3D reconstruction results and multi-view 2D information, respectively. The two size readings are fused to obtain the final size determination result of the fish target, thus achieving integrated processing of 3D reconstruction, species identification, and size assessment of the fish target. In some implementations, the embodiments of the present invention can be achieved through the following steps: Step 1: Data collection.

[0040] Data is acquired using a multi-view synchronous imaging device. A unified triggering mechanism is used to simultaneously obtain image information of the target fish from multiple perspectives. The multi-view images acquired at the same time are then synchronized and matched according to their timestamps to form a set of multi-view observation data. This data set includes images acquired synchronously from multiple perspectives. and the intrinsic parameter matrix of the corresponding camera. Camera extrinsic rotation matrix Translation vector The device acquires initial three-dimensional data output by the device. The multi-view synchronous imaging device includes... Each imaging perspective, and N is a positive integer. Indicates the first Two-dimensional images of fish targets acquired simultaneously from multiple perspectives. Indicates the first Each viewpoint corresponds to a camera's intrinsic parameter matrix. and They represent the first The rotation matrix and translation vector of each camera relative to the world coordinate system. The above parameters are all obtained from the camera calibration process and are used to establish the geometric correspondence between multi-view images.

[0041] The imaging perspectives of the multi-view synchronous imaging device have fixed and pre-calibrated geometric relationships. Its camera intrinsic and extrinsic parameters are obtained through a calibration process before the system runs and remain unchanged during the system operation. Therefore, there is no need to perform additional online view registration or feature alignment processing on the multi-view images or the features extracted from them.

[0042] Step 2: Data preprocessing and perspective consistency correction.

[0043] For image data acquired by a multi-view synchronous imaging device, based on the camera calibration parameters of the corresponding viewpoints, the first... Fish target images acquired from multiple perspectives Perform distortion correction processing, where ,and Subsequently, white balance correction, color cast correction, and dehazing enhancement are performed on the images from each viewpoint in sequence to reduce the impact of differences in imaging conditions from different viewpoints on subsequent processing. Based on this, cross-viewpoint brightness normalization and histogram matching are performed on the multi-view images to achieve consistency in brightness and color distribution across different viewpoints. After completing the brightness and color consistency processing, resolution and scale are unified for each viewpoint image, and the preprocessed multi-view image information is output.

[0044] Step 3: Fish target localization and cross-view correlation.

[0045] Based on the preprocessed multi-view image information, a target localization network is used in each view image to detect and localize fish targets, and the corresponding detection boxes for fish targets in each view are obtained. Fish target mask Subsequently, based on the known intrinsic and extrinsic parameters of the multi-view synchronous imaging device, a geometric correspondence between images from different viewpoints is established. Then, according to epipolar geometric constraints, fish targets satisfying the spatial correspondence in images from different viewpoints are matched. The matched fish targets are considered as the same fish target, and their detection boxes and fish target masks from multiple viewpoints are associated to form a corresponding cross-view fish target set. . .

[0046] Step 4: Self-supervised feature learning by masked autoencoder.

[0047] For each cross-view fish target set obtained in the above steps, the same fish target is obtained in the [missing information - likely a specific step or step]. Cropped images corresponding to each viewpoint are used, and the fish target mask is applied to the cropped images to preserve pixel information within the target area. ,and Random masking is performed on fish target images obtained from various viewpoints to obtain masked input images. These masked input images are then input into a mask autoencoder model, where the encoder encodes the features of the masked image, and the decoder reconstructs the occluded areas. This process achieves self-supervised learning of the visual representation of fish targets without manual annotation.

[0048] During training, context-based self-supervised feature learning is achieved by minimizing the reconstruction error between the original image of the fish target and the image reconstructed by the masked autoencoder. Its reconstruction loss... It can be represented as: .in This represents the original image of the fish target. This represents the image reconstructed by a mask autoencoder.

[0049] To further utilize the cross-view information constraints brought about by multi-view synchronous acquisition, a cross-view consistency mechanism is introduced in the self-supervised feature learning process of the mask autoencoder. Specifically, the view features extracted by the mask autoencoder from the same fish target under different viewpoints are aggregated, and a consistency constraint is applied to the feature representations of the same fish target under different viewpoints. The cross-view consistency loss... This can be summarized as follows: .in and These represent the same fish target from different perspectives. Perspective The following is a fish target-level feature representation, the view index And satisfy .

[0050] Combining the above two types of constraints, the overall optimization objective of the self-supervised feature learning stage of the mask autoencoder can be expressed as: .in These are the weight coefficients for cross-view consistency constraints. Through the self-supervised feature learning process of the mask autoencoder described above, a cross-view fish target-level feature representation is obtained to characterize the semantic consistency of the same fish target under different viewpoints, providing feature input for subsequent 3D Gaussian splash model and category and size discrimination.

[0051] Step 5: Reconstruction and Semantic Co-optimization Method of 3D Gaussian Splash Model

[0052] For the cross-view fish target-level feature representation obtained through self-supervised feature learning by a masked autoencoder, a 3D Gaussian splash model is introduced in the 3D reconstruction stage to model the 3D representation of the fish target. Specifically, the cross-view fish target-level feature representation is used as semantic prior information in the 3D Gaussian modeling process to guide the construction of the 3D representation of the fish target. For each fish target, a Gaussian set consisting of several 3D Gaussian units is constructed in 3D space, where each 3D Gaussian unit includes its spatial center position, spatial distribution parameters, and appearance attributes, which are used to collectively describe the geometric structure and appearance features of the fish target in 3D space.

[0053] In the 3D reconstruction process, based on the known and fixed geometric relationships between the various viewpoints of the multi-view synchronous imaging device, the 3D Gaussian set is projected onto the first... Each viewpoint generates a rendered image from that viewpoint, which is then compared with a real image simultaneously acquired from the same viewpoint. ,and By constructing the pixel-domain reprojection error between the real image and the rendered image, the spatial and appearance parameters of the 3D Gaussian unit are iteratively updated to obtain a 3D Gaussian set for representing fish targets. The above comparison and optimization process is based on known camera geometry and does not require additional online viewpoint registration operations on multi-view images or features.

[0054] Building upon this foundation, to achieve consistent and collaborative optimization between the 3D reconstruction process and semantic representation, a semantic constraint mechanism based on feature consistency is introduced during the optimization of the 3D Gaussian splash model. In practical applications, an encoder obtained through self-supervised feature learning via a masked autoencoder is used as the feature extraction network to extract features from the multi-view image information, obtaining a first fish target-level feature representation. The encoder is then used as the feature extraction network to extract features from the rendered image, obtaining a second fish target-level feature representation. Consistency constraints are applied to the first and second fish target-level feature representations in the feature space, ensuring that the 3D Gaussian model maintains a consistent semantic expression with the cross-view fish target-level feature representations while optimizing its geometric structure.

[0055] In the implementation, the optimization objectives in the 3D reconstruction stage include pixel-domain reprojection error and feature-domain semantic consistency error based on the cross-view fish target-level feature representation, and their joint loss function... It can be represented as: .in, Represents the real image and the image formed by the three-dimensional Gaussian set at the th... Pixel domain reprojection error between differentiable rendered images obtained from differentiable rendering at different viewpoints This represents the semantic consistency error based on the encoded features of a mask autoencoder. These are the weighting coefficients.

[0056] In practical applications, the pixel domain reprojection error can be expressed as: .in, Indicating the first step in the aforementioned steps Two-dimensional images of fish targets acquired simultaneously from multiple perspectives. This represents the result obtained by differentiating a three-dimensional Gaussian set, and is related to the first... The rendered image corresponding to each viewpoint ,and .

[0057] The semantic consistency error of the feature domain can be expressed as: .in, It is an encoder obtained through self-supervised feature learning using a masked autoencoder, used to extract fish target-level feature representations. By constraining the consistency between the real image and the rendered image in the feature space, the 3D Gaussian model maintains a semantic expression consistent with the cross-view fish target-level feature representation while optimizing the geometric structure.

[0058] Through the above-mentioned three-dimensional reconstruction and semantic co-optimization process, the obtained three-dimensional Gaussian representation is used to uniformly characterize the three-dimensional geometric structure and corresponding semantic feature information of fish targets, and serves as the input for subsequent species identification and size determination steps.

[0059] Step Six: Semantic alignment and parameter fine-tuning of large-scale multimodal models.

[0060] After completing the self-supervised feature learning based on mask autoencoder and the 3D semantic co-modeling of the 3D Gaussian splash model, a large-scale multimodal pre-trained model is introduced to perform semantic alignment and parameter fine-tuning on the 3D Gaussian representation and its corresponding cross-view fish target-level feature representation obtained in the aforementioned steps.

[0061] In practical applications, multi-view image information of fish targets, rendered images obtained by differentiable rendering of 3D Gaussian sets, and fish target-level feature representations are used as visual modal inputs to the large-scale multimodal pre-trained model to obtain semantic embedding representations corresponding to fish targets. Simultaneously, species category descriptions and size classifications corresponding to fish targets are constructed based on aquaculture application scenarios and used as semantic reference information in the semantic alignment process.

[0062] In the semantic alignment process, a semantic consistency constraint is constructed between the fish target-level feature representation and the semantic embedding representation to establish a correspondence between the 3D Gaussian representation and the species category and size level semantics of the fish target in the semantic space. Semantic alignment is achieved by minimizing the distance between the fish target-level feature representation obtained through self-supervised feature learning based on a mask autoencoder and semantic co-optimization using a 3D Gaussian splashing model, and the semantic embedding representation output by the large-scale multimodal pre-trained model. The alignment loss is... It can be represented as: .in, This is the fish target-level feature representation obtained through semantic co-optimization of the 3D Gaussian splash model, corresponding to the 3D Gaussian representation. This represents the semantic embedding representation of the fish target output by the large-scale multimodal pre-trained model.

[0063] During semantic alignment, some parameters of the large-scale multimodal pre-trained model are fine-tuned to adapt to the species category and size semantic system based on the 3D representation of fish targets. After semantic alignment and parameter fine-tuning, semantic features corresponding to the 3D representation of fish targets are obtained and used for subsequent species classification and size determination.

[0064] Step 7: 3D fish target-level multi-task classification.

[0065] After completing the reconstruction and semantic alignment of the 3D Gaussian splash model, a 3D fish target-level multi-task classification head is constructed based on the 3D Gaussian representation of the fish target. This head is used to simultaneously perform species category discrimination and size attribute discrimination on a single fish target.

[0066] In practical applications, for each fish target, the spatial location, feature representation, and weight information of each 3D Gaussian unit in its corresponding 3D Gaussian set are read. Then, a fish target-level feature aggregation operation is performed on the 3D Gaussian set to obtain a unified 3D feature representation of the fish target. During feature aggregation, based on the distribution relationship and visibility information of each 3D Gaussian unit in 3D space, the features of each Gaussian unit are weighted and aggregated to form a target-level representation vector that characterizes the overall geometric structure and semantic information of the fish target.

[0067] The three-dimensional feature representation of the fish target is input into the three-dimensional fish target multi-task classification head. Based on the shared three-dimensional feature representation, discrimination processing is performed through species classification branch and size discrimination branch respectively. The species classification branch outputs the species category result corresponding to the fish target, and the size discrimination branch outputs the size grade result corresponding to the fish target.

[0068] Step 8: Generate dual-path specification read values.

[0069] After obtaining the multi-task classification output results of the three-dimensional fish target, a two-way independent specification reading generation mechanism is constructed based on the specification information of the fish target to perform dual-path estimation of the fish target's specification.

[0070] In the first path, specification readings are generated based on the 3D reconstruction results. In practical applications, a 3D representation of the fish target is constructed based on a 3D Gaussian splash model. Geometric information reflecting the spatial scale characteristics of the fish target is extracted, and the dimensions of the geometric information are calculated to obtain the corresponding 3D size parameters. Combining camera calibration parameters and scale correction relationships, the 3D size parameters are converted into 3D specification level information of the fish target.

[0071] In the second path, specification readings are generated based on multi-view two-dimensional information. Specifically, based on the projection information of the fish target in the multi-view image information under each viewpoint, combined with the geometric correspondence between different viewpoints, the two-dimensional size of the fish target is estimated, and the two-dimensional size information is mapped to the two-dimensional specification level information of the fish target.

[0072] Step Nine: Specification Reading Value Fusion and Final Judgment.

[0073] After obtaining the three-dimensional and two-dimensional specification level information, the two specification level information are fused to obtain the final specification determination result of the fish target.

[0074] In one implementation, the three-dimensional size rating information and the two-dimensional size rating information are arithmetically averaged to obtain the size rating information of the fish target. In another implementation, the three-dimensional size rating information and the two-dimensional size rating information are weighted and fused based on the confidence and stability information corresponding to the three-dimensional size rating information and the two-dimensional size rating information to obtain the size rating information of the fish target.

[0075] During the fusion process, a consistency check is performed on the 3D and 2D specification level information based on the differences between them. When the difference between the 3D and 2D specification level information exceeds a preset threshold, abnormal specification level information in both the 3D and 2D information is corrected to reduce the impact of single-path anomalies on the final determined specification level information. After fusion, the system outputs the specification level information corresponding to the fish target, and simultaneously outputs the species category result and corresponding specification confidence information of the fish target. In some implementations, the output results can be combined with the 3D representation results of the fish target for visualization or statistical analysis.

[0076] This invention constructs an integrated processing framework for fish target "3D reconstruction—species identification—size assessment" from a holistic perspective. This invention acquires collaborative image information of fish targets from multiple viewpoints and corresponding camera intrinsic and extrinsic parameters using a multi-view synchronous imaging device. Preprocessing operations such as distortion correction and brightness and color consistency processing are performed on the multi-view image information to reduce cross-view imaging differences and provide a consistent data foundation for subsequent feature learning and 3D modeling. Furthermore, a self-supervised feature learning mechanism combining a mask autoencoder with cross-view consistency constraints is introduced to learn robust visual feature representations of fish targets to occlusion and non-rigid pose changes without manual annotation. Based on this, the visual features of the fish targets are introduced as prior information into the differentiable rendering optimization process of the 3D Gaussian splash model. Through joint constraints of pixel-domain reprojection error and feature-domain semantic consistency error, collaborative modeling and optimization of the 3D geometric structure and semantic features of the fish targets are achieved. Finally, by constructing a fish target-level 3D multi-task classification model to synchronously output the species category and size level of fish targets, and combining a dual-path size reading generation and fusion mechanism based on 3D reconstruction results and multi-view 2D information, the size determination results are fused and corrected, thereby effectively improving the stability of 3D modeling of fish targets, the accuracy of species identification, and the reliability of size determination results under occlusion and non-rigid posture change conditions.

[0077] This invention acquires multiple view images of fish targets at the same time using a multi-view synchronous imaging device, and introduces a self-supervised feature learning mechanism with cross-view consistency constraints. This enables the model to comprehensively utilize complementary information between different viewpoints, effectively alleviating the information loss problem of single-view two-dimensional images in scenarios with target occlusion and non-rigid pose changes, thereby improving the consistency and stability of fish target recognition and three-dimensional reconstruction results.

[0078] In this embodiment of the invention, the semantic features of fish targets obtained through self-supervised learning are introduced as prior information into the differentiable rendering optimization process of the three-dimensional Gaussian splash model. Through the joint constraint of pixel domain reprojection error and feature domain semantic consistency error, the constructed three-dimensional Gaussian representation can accurately reflect the spatial geometric structure of fish targets while maintaining a stable and consistent semantic expression, thus providing a reliable three-dimensional semantic foundation for subsequent species identification and size determination.

[0079] This invention introduces a large-scale multimodal pre-trained model for semantic alignment and combines it with a parameter fine-tuning mechanism to effectively transfer general semantic knowledge to the fine-grained identification and size determination scenarios of fish, reducing the dependence on a large number of manually labeled samples and enabling the model to better adapt to the application needs of different aquaculture environments and different fish species.

[0080] This invention constructs a fish target-level feature aggregation representation based on a three-dimensional Gaussian set, and designs a multi-task classification structure on this basis to simultaneously output the species category and size level of the fish target. It comprehensively utilizes multi-view information and three-dimensional structural features to reduce the interference of single-view or local area features on the recognition results and improve the overall consistency and stability of the joint recognition results.

[0081] This invention constructs a geometric size estimation path based on 3D reconstruction results and a size estimation path based on multi-view 2D information to independently calculate and fuse the size information of fish targets. In the event of errors or instability in a single path, abnormal results can be corrected through fusion and consistency correction mechanisms, thereby significantly improving the reliability and stability of the size determination results.

[0082] On the other hand, such as Figure 11 As shown, this embodiment of the invention also provides an apparatus for fish target identification. The apparatus is used to implement the fish target identification method provided in this embodiment of the invention. The apparatus includes: The acquisition module 1101 is used to acquire multi-view image information of the fish target to be identified and the geometric relationship of the multi-view image information. For each view image information in the multi-view image information, the fish target is detected and located by the target localization network, and the detection box and fish target mask in each view image information are determined. The first generation module 1102 is used to match and associate the detection boxes and fish target masks in the image information of each view according to the geometric relationship, generate a cross-view fish target set, and generate a fish target-level feature representation based on the cross-view fish target set through self-supervised feature learning. The second generation module 1103 is used to use the fish target-level feature representation as semantic prior information to perform differentiable rendering optimization of the three-dimensional Gaussian splash model, and generate a three-dimensional Gaussian representation of the fish target. The first determining module 1104 is used to determine the species category and three-dimensional size level information of the fish target through a multi-task classification head based on the three-dimensional Gaussian representation. The multi-task classification head is generated by multi-modal training based on the multi-view image information of the fish target in the training dataset, the fish target-level feature representation determined based on the multi-view image information of the fish target in the training dataset, and the three-dimensional Gaussian representation determined based on the multi-view image information of the fish target in the training dataset. The second determining module 1105 is used to determine the two-dimensional size grade information of the fish target based on the multi-view image information of the fish target to be identified and the geometric relationship, and to determine the size grade information of the fish target by combining the three-dimensional size grade information and the two-dimensional size grade information.

[0083] On the other hand, embodiments of the present invention also provide a device for fish target identification, the device for fish target identification comprising: processor; A memory storing computer-readable instructions, which, when executed by the processor, implement the method provided in the embodiments of the present invention.

[0084] On the other hand, embodiments of the present invention also provide a computer-readable storage medium storing program code, which can be invoked by a processor to execute the method provided in the embodiments of the present invention.

[0085] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: This invention achieves robust detection and consistency matching of fish targets in complex occlusion and pose change scenarios through multi-view target localization and cross-view association mechanisms. Building upon this, it introduces cross-view self-supervised feature learning to generate fish target-level feature representations with viewpoint invariance and semantic consistency without manual annotation. These features are used as semantic priors to guide the differentiable rendering optimization of the 3D Gaussian splash model, achieving collaborative modeling of geometric structure and semantic features, overcoming the shortcomings of traditional methods where 3D reconstruction and semantic representation are separated. A multi-task classification head is constructed based on the 3D Gaussian representation, simultaneously outputting species category and 3D size level. Combined with a 3D-2D dual-path size reading fusion mechanism, the reliability and fault tolerance of size determination are effectively improved. Finally, an integrated processing flow of "3D reconstruction—species identification—size assessment" is formed, significantly reducing reliance on manual annotation, enhancing cross-environment generalization ability, and providing high-precision and high-stability key technical support for intelligent management of aquaculture.

[0086] Figure 12 This is a schematic diagram of the structure of a device for fish target identification provided in an embodiment of the present invention, as shown below. Figure 12 As shown, optionally, the device 1210 for fish target identification may include a first processor 2001.

[0087] Optionally, the device 1210 for fish target identification may also include a memory 2002 and a transceiver 2003.

[0088] The first processor 2001, memory 2002, and transceiver 2003 can be connected via a communication bus.

[0089] The following is combined with Figure 12 The various components of the device 1210 for fish target identification are described in detail below: The first processor 2001 is the control center of the device 1210 for fish target identification. It can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 can be one or more central processing units (CPUs), application-specific integrated circuits (ASICs), or one or more integrated circuits configured to implement embodiments of the present invention, such as one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs).

[0090] Optionally, the first processor 2001 can perform various functions of the device 1210 for fish target identification by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.

[0091] In a specific implementation, as one example, the first processor 2001 may include one or more CPUs, for example... Figure 12 CPU0 and CPU1 are shown in the diagram.

[0092] In a specific implementation, as one example, the device 1210 for fish target identification may also include multiple processors, for example... Figure 12The first processor 2001 and the second processor 20012 are shown in the diagram. Each of these processors may be a single-core processor or a multi-core processor. Here, a processor may refer to one or more devices, circuits, and / or processing cores used to process data (such as computer program instructions).

[0093] The memory 2002 is used to store the software program that executes the present invention, and is controlled by the first processor 2001 to execute it. The specific implementation method can be referred to the above method embodiment, and will not be repeated here.

[0094] Optionally, the memory 2002 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. The memory 2002 may be integrated with the first processor 2001 or may exist independently, and may be connected via the interface circuit of the device 1210 for fish target identification. Figure 12 (Not shown in the image) is coupled to the first processor 2001, and this embodiment of the invention does not specifically limit this.

[0095] The transceiver 2003 is used to communicate with network devices or with terminal devices.

[0096] Alternatively, transceiver 2003 may include a receiver and a transmitter. Figure 12 (Not shown separately). The receiver is used to implement the receiving function, and the transmitter is used to implement the transmitting function.

[0097] Alternatively, the transceiver 2003 can be integrated with the first processor 2001 or exist independently, and can be connected via the interface circuit of the device 1210 for fish target identification. Figure 12 (Not shown in the image) is coupled to the first processor 2001, and this embodiment of the invention does not specifically limit this.

[0098] It should be noted that, Figure 12 The structure of the device 1210 for fish target identification shown in the figure does not constitute a limitation on the device for fish target identification. Actual devices for fish target identification may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0099] Furthermore, the technical effects of the device 1210 for fish target recognition can be referred to the technical effects of the method for federated learning described in the above method embodiments, and will not be repeated here.

[0100] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0101] It should also be understood that the memory in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0102] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, motor drive, or data center to another website, computer, motor drive, or data center via infrared, microwave, or other means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a motor drive or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.

[0103] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.

[0104] In this invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single item or multiple items.

[0105] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0106] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0107] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0108] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0109] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0110] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0111] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a motor driver, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0112] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for fish target recognition, characterized in that, The method includes: The process involves acquiring multi-view image information of the fish target to be identified and the geometric relationship of the multi-view image information; then, using a target localization network, detecting and locating the fish target in each view of the multi-view image information, and determining the detection box and fish target mask in each view of the image information. Based on the geometric relationship, the detection boxes and fish target masks in the image information of each viewpoint are matched and associated to generate a cross-viewpoint fish target set. Based on the cross-viewpoint fish target set, self-supervised feature learning is performed to generate fish target-level feature representations. The fish target-level feature representation is used as semantic prior information for differentiable rendering optimization of the three-dimensional Gaussian splash model to generate a three-dimensional Gaussian representation of the fish target. Based on the three-dimensional Gaussian representation, the species category and three-dimensional size level information of the fish target are determined by a multi-task classification head. The multi-task classification head is generated by multi-modal training based on the multi-view image information of the fish target in the training dataset, the fish target-level feature representation determined by the multi-view image information of the fish target in the training dataset, and the three-dimensional Gaussian representation determined by the multi-view image information of the fish target in the training dataset. Based on the multi-view image information of the fish target to be identified and the geometric relationship, the two-dimensional size level information of the fish target is determined, and the size level information of the fish target is determined by combining the three-dimensional size level information and the two-dimensional size level information. The step of using the fish target-level feature representation as semantic prior information to perform differentiable rendering optimization of the 3D Gaussian splash model to generate a 3D Gaussian representation of the fish target includes: The fish target-level feature representation is used as semantic prior information to construct a three-dimensional Gaussian set of fish targets. The three-dimensional Gaussian set is composed of multiple three-dimensional Gaussian units, and each three-dimensional Gaussian unit contains the spatial center position, spatial distribution parameters and appearance attributes. Based on the geometric relationship, the three-dimensional Gaussian set is projected to each viewpoint through a differentiable rendering mechanism to generate a rendered image at the corresponding viewpoint. The pixel domain reprojection error between the multi-view image information and the rendered image is constructed. The spatial parameters and appearance parameters of the three-dimensional Gaussian unit are iteratively optimized by minimizing the pixel domain reprojection error to generate a three-dimensional Gaussian representation of the fish target. The method further includes: An encoder obtained through self-supervised feature learning via a masked autoencoder is used as a feature extraction network to extract features from the multi-view image information, thereby obtaining a first fish target-level feature representation. The encoder is used as a feature extraction network to extract features from the rendered image and obtain a second fish target-level feature representation. In the feature space, a consistency constraint is applied to the first fish target-level feature representation and the second fish target-level feature representation to construct the feature domain semantic consistency error; The pixel domain reprojection error and the feature domain semantic consistency error are jointly optimized to construct a joint loss function: ,in, This refers to pixel-domain reprojection error. For the semantic consistency error of the feature domain, These are the weighting coefficients; Minimize the joint loss function so that the 3D Gaussian representation of the fish target generated by the 3D Gaussian splash model maintains the semantic expression consistent with the fish target-level feature representation.

2. The method according to claim 1, characterized in that, The multi-task classification head is generated through multimodal training in the following manner: For multi-view image information of fish targets in the training dataset, the target localization network is used to detect and locate fish targets based on the image information of each view, and to determine the detection box and fish target mask in the image information of each view. Based on the geometric relationship of multi-view image information, the detection boxes and fish target masks in each view image information are matched and associated to generate a cross-view fish target set. Based on the cross-view fish target set, self-supervised feature learning is performed to generate fish target-level feature representations. The fish target-level feature representation is used as semantic prior information for differentiable rendering optimization of the three-dimensional Gaussian splash model to generate a three-dimensional Gaussian representation of the fish target. Multimodal training is performed based on the three-dimensional Gaussian representation of fish targets, multi-view image information of fish targets, and fish target-level feature representation to generate semantic embedding representations corresponding to fish targets; Based on aquaculture application scenarios, species category descriptions and size level descriptions corresponding to fish targets are constructed. These species category descriptions and size level descriptions are used as semantic reference information, and semantic consistency constraints are constructed between the fish target-level feature representations in the semantic embedding representation set.

3. The method according to claim 2, characterized in that, The loss for constructing semantic consistency constraints is expressed as: ,in, This represents the target-level features of the fish species. This represents the semantic embedding representation.

4. The method according to claim 1, characterized in that, The process of generating fish target-level feature representations through self-supervised feature learning based on the cross-view fish target set includes: Based on the cross-view fish target set, cropped images of the same fish target at each view are obtained, and the fish target mask is applied to the cropped images to determine the fish target images in the cropped images; A random masking operation is performed on the fish target image to obtain a masked input image. The masked input image is then input into a mask autoencoder model, which performs feature encoding on the masked image. The decoder then reconstructs the occluded areas in the fish target image to generate a reconstructed image. By minimizing the reconstruction loss between the original image of the fish target and the reconstructed image, self-supervised feature learning based on context information is performed. Represented as: , in, Represents the original image of the fish target. This refers to the reconstructed image.

5. The method according to claim 4, characterized in that, In the self-supervised feature learning process, a cross-view consistency mechanism is introduced to aggregate the view features extracted by the mask autoencoder from different viewpoints of the same fish target, and to constrain the consistency of the feature representations of the same fish target under different viewpoints through cross-view consistency loss. The cross-perspective consistency loss Represented as: , in, Indicates the fish target in the field of view The following fish target-level feature representation, Indicates the fish target in the field of view The fish target-level features below, the view index And satisfy ; The overall optimization objective of the masked autoencoder in the self-supervised feature learning stage is: ,in, The weighting coefficients for cross-perspective consistency constraints.

6. A device for fish target identification, characterized in that, The apparatus for fish target identification is used to implement the method for fish target identification as described in any one of claims 1-5, the apparatus comprising: The acquisition module is used to acquire multi-view image information of the fish target to be identified and the geometric relationship of the multi-view image information. For each view image information in the multi-view image information, the fish target is detected and located by the target localization network, and the detection box and fish target mask in each view image information are determined. The first generation module is used to match and associate the detection boxes and fish target masks in the image information from each viewpoint according to the geometric relationship, generate a cross-viewpoint fish target set, and perform self-supervised feature learning based on the cross-viewpoint fish target set to generate fish target-level feature representations. The second generation module is used to use the fish target-level feature representation as semantic prior information to perform differentiable rendering optimization of the three-dimensional Gaussian splash model, and generate a three-dimensional Gaussian representation of the fish target. The first determining module is used to determine the species category and three-dimensional size level information of the fish target through a multi-task classification head based on the three-dimensional Gaussian representation. The multi-task classification head is generated by multi-modal training based on the multi-view image information of the fish target in the training dataset, the fish target-level feature representation determined based on the multi-view image information of the fish target in the training dataset, and the three-dimensional Gaussian representation determined based on the multi-view image information of the fish target in the training dataset. The second determining module is used to determine the two-dimensional size grade information of the fish target based on the multi-view image information of the fish target to be identified and the geometric relationship, and to determine the size grade information of the fish target by combining the three-dimensional size grade information and the two-dimensional size grade information.

7. A device for fish target identification, characterized in that, The device for fish target identification includes: processor; A memory storing computer-readable instructions that, when executed by the processor, implement the method as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium contains program code that can be invoked by a processor to execute the method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Object attitude estimation method based on scene-level semantic three-dimensional Gaussian splash

    CN121330053A

  • Three-dimensional scene segmentation method and device, equipment and storage medium

    CN121564702A