Multi-modal fish phenotype plasticity classification method fusing two-dimensional key point detection and three-dimensional structured light

By integrating two-dimensional keypoint detection and three-dimensional structured light into a multimodal fish phenotypic plasticity classification method, the problem of traditional methods being unable to distinguish the phenotypic plasticity differences between sea-caught and farmed large yellow croaker is solved, achieving high-precision fish origin identification and supporting intelligent management of aquatic resources.

CN122023951AActive Publication Date: 2026-05-12EAST CHINA SEA FISHERIES RES INST CHINESE ACAD OF FISHERY SCI +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202610479025.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-13
Publication Date
2026-05-12
Estimated Expiration
2046-04-13

AI Technical Summary

Technical Problem

Traditional measurement methods are insufficient to accurately distinguish the phenotypic plasticity differences between sea-caught and farmed large yellow croaker, and existing two-dimensional vision methods are insufficient to capture the three-dimensional geometric details of the fish, resulting in technical bottlenecks for market supervision and resource protection.

Method used

A multimodal fish phenotypic plasticity classification method integrating 2D keypoint detection and 3D structured light is proposed. By constructing a multimodal classification model and combining 2D images and 3D structured light data, statistically significant key discriminant features are selected, and a logistic regression discriminant formula is constructed to achieve rapid identification between wild and farmed fish.

Benefits of technology

The model has achieved accurate identification of the origin of large yellow croaker, with an accuracy rate of 97.14%, providing technical support for the intelligent management of aquatic resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122023951A_ABST
    Figure CN122023951A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal fish phenotype plasticity classification method fusing two-dimensional key point detection and three-dimensional structured light, which constructs a large yellow croaker two-dimensional and three-dimensional point cloud data set, designs a feature engineering process, selects a key point detection and geometric feature extraction algorithm, lays a foundation for subsequent modeling, and improves the classification efficiency. The performance of various machine learning models is compared through systematic experiments, key discrimination characteristics with statistical significance are screened out, a simplified formula suitable for on-site rapid identification is derived, the biological connotation of the fish body curvature serving as a phenotypic plasticity sensitive index is deeply discussed according to the experimental result, and the method is suitable for rapid identification of the fish body. The improvement effect of multi-modal feature fusion on precision is analyzed, and the application value of the result in resource protection and industrial optimization is explained. According to the method, the communication from the microscopic curvature characteristic to the macroscopic industrial application is realized, the fish surface plasticity difference of the sea-caught and cultured pseudosciaena crocea is researched for the first time, and the key form distinguishing characteristic with high distinguishing power is extracted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0002] This invention relates to the application of computer vision technology in fisheries, specifically to a multimodal fish phenotypic plasticity classification method that integrates two-dimensional key point detection and three-dimensional structured light. Background Technology

[0004] The striking similarity in appearance between wild-caught large yellow croaker and those from mainstream aquaculture (workboats, deep-sea cages) poses challenges to market regulation and presents technical bottlenecks for wildlife conservation. There is an urgent need to establish precise identification criteria based on phenotypic plasticity. Fish phenotypic plasticity, the ability of an individual with the same genotype to exhibit reversible phenotypic characteristics under different environmental conditions, is a key mechanism for short-term adaptation to environmental changes and survival. Within the framework of ecology and evolutionary biology, this mechanism is considered a bridge connecting environmental pressure, individual response, and long-term adaptive evolution. Traditional measurement methods struggle to fully capture these subtle geometric differences, leading to technical bottlenecks in the accurate identification of wild-caught and aquaculture populations. Therefore, there is an urgent need to develop precise quantitative methods based on three-dimensional reconstruction to systematically analyze the differences in their surface geometric features, providing support for overcoming the technical challenges of market regulation and resource conservation.

[0005] Computer vision technology has become the mainstream method for automating and non-invasively acquiring fish phenotypic parameters. Existing passive 3D methods are still insufficient for acquiring high-resolution surface details and struggle to capture the key local curvature features needed to distinguish different populations. Therefore, there is an urgent need to introduce more precise and robust active 3D phenotypic acquisition technologies to deeply analyze the impact of different living environments on the geometric morphology of fish bodies, thus providing a new technical path for the accurate identification of wild-caught and farmed large yellow croaker. Structured light and other 3D vision technologies effectively compensate for the shortcomings of traditional 2D methods in terms of feature dimensions by acquiring high-precision surface geometric information. Although deep learning-based 2D image methods are widely used in fish identification and size measurement, they rely on projected images and struggle to capture the subtle geometric features of the fish's surface, such as three-dimensional undulations and curvature—key indicators reflecting the fish's plumpness, surface texture, and other physiological states and plasticity phenotypes. Therefore, using structured light and other 3D vision technologies to acquire high-precision geometric data has become an advanced and necessary means to accurately quantify the plasticity of the body surface morphology of economically important fish such as large yellow croaker and to systematically compare subtle phenotypic differences between wild-caught and farmed individuals.

[0006] As an important economic fish species in my country, the large yellow croaker faces an increasingly prominent contradiction between the booming aquaculture industry and the continuous decline of wild resources. Fish phenotypic plasticity, as a key mechanism for environmental adaptation, contains rich ecological and aquaculture-related information in its external manifestations (such as body shape and surface texture). However, the field currently faces the dual challenges of difficult market supervision and technological bottlenecks in wild resource protection: on the one hand, it is necessary to accurately identify vendors claiming to be fishing large yellow croaker to prevent false labeling and maintain market fairness; on the other hand, it is necessary to prevent illegally caught individuals from entering aquaculture channels to evade supervision and ensure resource security. Traditional two-dimensional vision methods are insufficient to accurately capture the three-dimensional geometric details of fish, limiting technological breakthroughs in addressing these issues. Summary of the Invention

[0008] This invention provides a multimodal fish phenotypic plasticity classification method that integrates two-dimensional keypoint detection and three-dimensional structured light. By constructing a multimodal classification model that integrates two-dimensional images and three-dimensional structured light data, it achieves accurate identification of the origin of large yellow croaker based on phenotypic features, with a model accuracy of 97.14%, providing technical support for the intelligent identification and management of aquatic resources.

[0009] This invention is achieved through the following technical solution:

[0010] A multimodal fish phenotypic plasticity classification method integrating 2D keypoint detection and 3D structured light includes the following steps: (1) Select the data acquisition equipment and collect data on large yellow croaker; (2) Use acquisition equipment to acquire two-dimensional and three-dimensional geometric features of the fish body, and define and label key points; (3) By comparing the performance of various machine learning models through systematic experiments, key discriminant features with statistical significance were selected; (4) A simplified formula suitable for rapid on-site identification was derived. Based on the logistic regression modeling framework of 5-fold cross-validation, a formula for distinguishing wild fish from farmed fish suitable for rapid on-site application was constructed. This formula is based on five easily measurable and statistically significant morphological features. While ensuring high discrimination accuracy, it also takes into account the ease of calculation. The five feature variables used were all screened through three correlation analyses (p value < 0.01), covering the segmented measurement values ​​of body height and body length and their standardized ratios, specifically including: body_high, body_length_1, body_length_2, body_length_3, and body_length_total. Among them, all five features are significantly negatively correlated with the farming status, indicating that their values ​​tend to point to farmed fish. On this basis, by performing Z-score standardization on the training set in each fold cross-validation and fitting a logistic regression model, the following discrimination function was finally obtained by summing the mean values ​​of the 5-fold coefficients: , The discrimination rule is set based on the decision boundary of the logistic regression output probability: when At that time, it was determined to be a wild fish; when When a fish is identified as farmed, this rule is equivalent to predicting the probability. This avoids exponential calculations and significantly improves the feasibility of manual calculation; (5) The model performance was verified by 5-fold cross-validation, which showed that the formula has excellent generalization ability and stability.

[0011] As a preferred embodiment, in step (1), the acquisition device uses the Zivid TwoM70 high-precision structured light 3D camera manufactured by the Norwegian company Zivid to acquire a total of 1408 large yellow croaker data, including 416 farmed large yellow croaker data and 992 sea-fished large yellow croaker data, including large yellow croaker point cloud data and color RGB data.

[0012] As a preferred embodiment, the points collected in step (1) are stored in cloud PLY format. Each file contains approximately 2,332,800 vertices, and each vertex accurately records floating-point (x, y, z) spatial coordinates and unsigned character RGB color components.

[0013] In a preferred embodiment, in step (2), the midline of the fish body is discretized into multiple line segments, and the body shape of different types of fish is described by segmented arc length. A labeling scheme for key points of the fish body is constructed, and the height of the fish body and the pectoral fins are also labeled.

[0014] As a preferred embodiment, in step (3), the three-dimensional fish geometric features are extracted, and the mapping relationship between the image and the point cloud data is established. The dataset is acquired using a structured light three-dimensional scanning system and a synchronously triggered two-dimensional RGB camera for multimodal data acquisition. The structured light camera generates point cloud data containing spatial coordinates (X, Y, Z) and RGB color information by projecting an coded grating and capturing the deformation pattern. The data is stored in PLY format, while the two-dimensional camera records a high-resolution PNG projection image of the corresponding viewpoint.

[0015] In a preferred embodiment, step (3) involves establishing a precise geometric mapping relationship to achieve multimodal data fusion. The three-dimensional point cloud is back-projected back onto the two-dimensional image plane based on the imaging model, ensuring pixel-level alignment between the point cloud data and the original image. The specific technical implementation is as follows: The origin of the structured light camera coordinate system coincides with the center of the imaging plane, and the optical axis along… The positive half-axis direction, therefore we referenced the principle of structured light cameras based on any point in the given point cloud. Its back projection onto the target plane The process is achieved through ray parameterization, where, To preset the target depth, 50% of the maximum Z-value of the point cloud is taken to balance perspective distortion. The calculation formula is as follows: , Back projection point The coordinates are: , Then The XY range is linearly normalized to the original image resolution (width W, height H): , , Indicates rounding down. Using pixel coordinates, the point cloud color values ​​are ultimately written directly to the corresponding pixel positions to generate the reconstructed image.

[0016] As a preferred embodiment, in step (3), the performance of six mainstream machine learning algorithms on classification tasks is evaluated. Accuracy, Precision, Recall, F1 score and AUC value are used as core evaluation indicators. During the training process, 5-fold true values ​​are used for training. After training, the test set is used for testing. All models are trained and validated using features calculated from the true value labels.

[0017] As a preferred embodiment, in step (3), the traditional machine learning model evaluation metrics commonly used in machine learning and classification tasks include accuracy, precision, recall, F1 score, and AUC (Area Under the ROC Curve). Their mathematical definitions are as follows, applicable to the "Evaluation Metrics" section of the paper: 1. Accuracy represents the proportion of correct predictions out of all predictions: , in: True Positives True Negatives False positives False Negatives 2. Precision represents the proportion of samples predicted as positive that are actually positive: , 3. Recall represents the proportion of samples that are actually positive but were correctly predicted as positive: , 4. The F1 score is the harmonic mean of precision and recall, used to balance the two: , 5. AUC is the area under the ROC curve, measuring the model's ability to distinguish between positive and negative samples at different classification thresholds. Its mathematical expression is usually defined using integral or ranking statistics: , in: , ,

[0018] In practice, AUC can also be interpreted as the probability that the model scores the positive sample higher than the negative sample when a positive sample and a negative sample are randomly selected.

[0019] In a preferred embodiment, in step (3), the tree-based ensemble model LightGBM stands out with its superior overall performance: its accuracy reaches 0.9502, precision 0.9441, F1 score 0.9649, and AUC value 0.9875, ranking first in all four key indicators; its recall (0.9881) is only slightly lower than Random Forest (0.9907) by a mere 0.0026, indicating that the model effectively balances false positive control while maintaining high positive example recognition capability. XGBoost (accuracy 0.9417) and Random Forest (accuracy 0.9406) rank second and third respectively, with both achieving high recall (XGBoost: 0.9850, Random Forest: 0.9907) and AUC (XGBoost: 0.9778, Random Forest: 0.9907). The robust performance of the Gradient Boosting Tree (GBR) model (0.9804) demonstrates its strong fitting ability for complex nonlinear relationships. Notably, while the GBR achieves a high recall (0.9859), its precision (0.8183) and accuracy (0.8189) are significantly low, revealing the risk of false positives associated with high recall strategies. SVM performs the worst across all metrics (accuracy 0.8589, AUC 0.9017), indicating its insufficient adaptability to the feature distribution of this dataset. LightGBM achieves the best balance between overall performance, robustness to classification thresholds, and computational efficiency, making it the preferred model for this task. When business scenarios are extremely sensitive to the cost of missed detections, Random Forest can serve as an alternative for high recall.

[0020] In a preferred embodiment, step (5) achieves an overall accuracy of 90.28%, an AUC value of 95.39%, and category-specific analysis shows a recall rate of 92.74% for wild fish. The cross-validation operation includes the following steps: (a) measuring nine morphological indicators using a digital caliper with an accuracy of not less than 0.1 mm; (b) calculating the discriminant function value by substituting the values ​​into the above formula. (c) Based on The classification decision is made based on the symbol, for boundary samples. It is recommended to retest key ratio features or combine auxiliary methods for verification.

[0021] Technical Principles: In the Materials and Methods section, two-dimensional and three-dimensional point cloud datasets of large yellow croaker were constructed. A feature engineering process was designed, and key point detection and geometric feature extraction algorithms were selected to lay the foundation for subsequent modeling. Subsequently, through systematic experiments, the performance of various machine learning models was compared, key discriminative features with statistical significance were selected, and a simplified formula suitable for rapid on-site identification was derived. Furthermore, based on the experimental results, the biological implications of fish body curvature as a sensitive indicator of phenotypic plasticity were explored, the role of multimodal feature fusion in improving accuracy was analyzed, and the application value of the results in resource conservation and industrial optimization was explained. Finally, the innovative contributions of this research to the intelligent quantification of three-dimensional phenotypic features are summarized, and future directions such as multidimensional database construction, model transfer, and lightweight technology are discussed.

[0022] Beneficial effects: This invention not only deepens the theoretical understanding of the environment-phenotype coupling mechanism, but also, using multimodal three-dimensional intelligent analysis technology as a bridge, achieves the connection from microscopic curvature characteristics to macroscopic industrial applications. Employing high-precision structured light three-dimensional reconstruction technology, it systematically elucidates the phenotypic differences between sea-caught and farmed large yellow croaker, specifically in the following aspects:

[0023] 1. This study is the first to investigate the differences in surface plasticity between wild-caught and farmed large yellow croaker, and extracts key morphological distinguishing features with high discriminative power.

[0024] 2. We innovatively introduced the local curvature of the fish body surface as a key parameter for quantifying phenotypic plasticity, and verified through feature importance assessment that its identification power is significantly better than that of traditional morphological parameters.

[0025] 3. A multimodal classification model integrating two-dimensional images and three-dimensional structured light data was constructed, which achieved accurate identification of the origin of large yellow croaker based on phenotypic features. The model accuracy reached 97.14%, providing technical support for the intelligent identification and management of aquatic resources. Attached Figure Description

[0026] Figure 1 This is a schematic diagram of a large yellow croaker data acquisition device in one embodiment of the present invention.

[0027] Figure 2 This is a schematic diagram of a large yellow croaker data acquisition device in one embodiment of the present invention.

[0028] Figure 3 This is a schematic diagram of the fish body data and labeling definition of large yellow croaker in one embodiment of the present invention.

[0029] Figure 4 This is a schematic diagram of the back projection mapping from structured light point cloud to a two-dimensional image in one embodiment of the present invention.

[0030] Figure 5 This is a schematic diagram of the back projection mapping from structured light point cloud to a two-dimensional image in one embodiment of the present invention.

[0031] Figure 6 This is a schematic diagram of the back projection mapping result from structured light point cloud to two-dimensional image in one embodiment of the present invention.

[0032] Figure 7 This is a schematic diagram of the back projection mapping result from structured light point cloud to two-dimensional image in one embodiment of the present invention. Detailed Implementation

[0034] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings: These embodiments are implemented based on the technical solution of the present invention, and provide detailed implementation methods and specific operation processes, but the protection scope of the present invention is not limited to the following embodiments.

[0035] A multimodal fish phenotypic plasticity classification method integrating 2D keypoint detection and 3D structured light includes the following steps: (1) Data acquisition equipment was selected to collect large yellow croaker data. The acquisition equipment used was the ZividTwo M70 high-precision structured light 3D camera manufactured by the Norwegian company Zivid. A total of 1408 large yellow croaker data points were collected, including 416 farmed large yellow croaker data points and 992 sea-fished large yellow croaker data points. These data points included point cloud data and color RGB data. The acquired points were stored in cloud PLY format. Each file contained approximately 2,332,800 vertices. Each vertex accurately recorded floating-point (x, y, z) spatial coordinates and unsigned character RGB color components, such as... Figure 1 , 2 As shown, Figure 1 The data acquisition equipment and methods were demonstrated. Figure 2 The on-site data acquisition process was demonstrated. The acquired point cloud was stored in PLY format, with each file containing approximately 2,332,800 vertices. Each vertex accurately recorded floating-point (x, y, z) spatial coordinates and unsigned character RGB color components, fully preserving the geometric shape and true color information of the large yellow croaker.

[0036] (2) Two-dimensional and three-dimensional geometric features of the fish body were collected using acquisition equipment, and key points were defined and labeled. The midline of the fish body was discretized into multiple line segments, and the body shape of different types of fish was described using segmented arc lengths. A labeling scheme for key points of the fish body was constructed, and the height of the fish body and the pectoral fins were also labeled. For example Figure 3 As shown, the key points marked on the image are used to characterize the fish's outline and plasticity. Specifically, the fish body is assumed to have 9 anatomical key points. Each of the key points These nine key points represent the pixel coordinates on the image plane. Through their spatial distribution and relative relationships, the shape of the fish can be quantitatively described and analyzed. The locations of these key points are carefully selected to effectively capture the main morphological features and structural changes of the fish, thus reflecting its overall outline and shape characteristics. The definitions of the key points are shown in Table 1. Table 1 Key Point Symbols and Definitions

[0037] Geometric measurement of key points:

[0038] Based on the spatial locations of the nine key anatomical points in Table 1, to achieve digital representation of fish morphology, discrete coordinates need to be mapped into geometric indicators with clear biological connotations. The spatial configuration between key points encodes the core morphological information of the fish outline, while Euclidean distance, as an intuitive geometric metric, can effectively characterize the linear scale between key anatomical structures. By constructing a distance matrix for specific key point pairs, the original coordinate data is transformed into a set of interpretable morphological feature parameters, thereby abstracting the two-dimensional fish outline into a low-dimensional and comparable numerical feature vector. This distance feature set not only preserves the spatial constraints of key anatomical structures but also lays a solid quantitative foundation for subsequent morphological variation analysis, growth dynamic monitoring, and behavioral phenotypic recognition. Specifically, any two points... and The Euclidean distance between them is defined as shown in Table 2:

[0039] Table 2 Definition of Euclidean distance between any two points

[0040] (3) By comparing the performance of various machine learning models through systematic experiments, the key discriminant features with statistical significance were selected, the geometric features of the three-dimensional fish body were extracted, and the mapping relationship between the image and the point cloud data was established. The dataset was acquired using a structured light three-dimensional scanning system and a synchronously triggered two-dimensional RGB camera for multimodal data acquisition. The structured light camera generated point cloud data (stored in PLY format) containing spatial coordinates (X, Y, Z) and RGB color information by projecting an encoded grating and capturing deformation patterns. At the same time, the two-dimensional camera recorded high-resolution PNG projection images of the corresponding viewpoints.

[0041] To achieve multimodal data fusion, a precise geometric mapping relationship needs to be established. The 3D point cloud is back-projected back onto the 2D image plane according to the imaging model, ensuring that the point cloud data is aligned with the original image at the pixel level. The specific technical implementation is as follows: The origin of the structured light camera's coordinate system coincides with the center of the imaging plane, and the optical axis along... The positive half-axis direction. Therefore, we referenced the principle of structured light cameras and used any point in the point cloud as a reference. Its back projection onto the target plane The process is achieved through ray parameterization

[29] . To preset the target depth, 50% of the maximum Z-value of the point cloud is taken to balance perspective distortion. The calculation formula is as follows: , Back projection point The coordinates are: , Then The XY range is linearly normalized to the original image resolution (width W, height H): , , Indicates rounding down. Using pixel coordinates, the point cloud color values ​​are ultimately written directly to the corresponding pixel locations to generate the reconstructed image, such as... Figure 4 , 5 As shown, Figure 4 For 3D point cloud visualization, Figure 5 This is a schematic diagram of the back projection process, with the camera origin and target plane marked. and the ray parameter t. The final two-dimensional reconstruction result is as follows: Figure 6 , 7 As shown, where Figure 6 This is the original image; Figure 7During the process of geometrically aligning the 3D point cloud image (orthogonally or from perspective along the Z-axis) to the 2D image plane after reconstructing the point cloud data, a distinct concentric circle pattern appeared on the surface of the large yellow croaker. This phenomenon is a typical moiré pattern, originating from the fact that the high-frequency spatial structure implicit in the point cloud data is similar to but not completely synchronized with the discrete sampling frequency of the pixel grid of the 2D image sensor, thus producing low-frequency interference fringes during projection. These concentric circles with the center of the field of view as the origin are essentially a "beat frequency" effect between two periodic structures, belonging to the sampling aliasing phenomenon in signal processing. Under the premise of fixed projection rules and consistent data preprocessing, although the 2D image generated by point cloud projection has visual distortions or artifacts compared to the real image due to aliasing effects such as moiré patterns, these distortions are essentially products of deterministic geometric transformations and sampling processes, possessing high repeatability and consistency. For traditional machine learning models, their classification performance depends on the relative distribution of input features in the feature space rather than absolute visual fidelity. As long as the training and test sets undergo the same projection method and a unified preprocessing procedure, the model can effectively learn discriminative patterns in the transformed feature space. Therefore, such systematic and consistent deformation will not destroy the discriminative structure of the features, nor will it have a substantial impact on the final classification performance.

[0042] To evaluate the performance of six mainstream machine learning algorithms on classification tasks, accuracy, precision, recall, F1 score, and AUC were used as core evaluation metrics. During training, 5-fold true values ​​were used for training, and the test set was used for testing after training. The classification results of different machine learning models for all keypoint detection model recognition results are shown in Table 3. All models were trained and validated using features calculated from the true value labels.

[0043] Table 3 Machine Learning Model Training Results

[0044] ml_model accuracy precision recall f1 auc Gradient_boosting 0.8189 0.8183 0.9859 0.8885 0.9427 Lightgbm 0.9502 0.9441 0.9881 0.9649 0.9875 Logistic_regression 0.9131 0.9246 0.9586 0.9398 0.9586 Random_forest 0.9406 0.9320 0.9907 0.9596 0.9804 Svm 0.8589 0.8801 0.9383 0.9050 0.9017 Xgboost 0.9417 0.9373 0.9850 0.9598 0.9778

[0045] Table 3 shows that the tree-based ensemble models significantly outperform traditional models overall. LightGBM stands out with the best overall performance: its accuracy is 0.9502, precision is 0.9441, F1 score is 0.9649, and AUC is 0.9875, ranking first in all four key metrics. Its recall (0.9881) is only slightly lower than Random Forest (0.9907) by a mere 0.0026, indicating that this model effectively balances false positive control while maintaining high positive example recognition capability. XGBoost (accuracy 0.9417) and Random Forest (accuracy 0.9406) rank second and third respectively. Both models demonstrate robust performance in recall (XGBoost: 0.9850, Random Forest: 0.9907) and AUC (XGBoost: 0.9778, Random Forest: 0.9804), confirming the strong fitting ability of ensemble tree models for complex nonlinear relationships. It is worth noting that while Gradient Boosting Tree achieved a high recall rate (0.9859), its precision (0.8183) and accuracy (0.8189) were significantly low, revealing the risk of false positives associated with high recall strategies. SVM performed the worst across all metrics (accuracy 0.8589, AUC 0.9017), indicating its insufficient adaptability to the feature distribution of this dataset. LightGBM achieves the best balance between overall performance, robustness to classification thresholds, and computational efficiency, making it the preferred model for this task. When business scenarios are extremely sensitive to the cost of missed detections, Random Forest can serve as an alternative with high recall.

[0046] Traditional machine learning model evaluation metrics, commonly used in machine learning and classification tasks, include accuracy, precision, recall, F1 score, and AUC (Area Under the ROC Curve). Their mathematical definitions are as follows, applicable to the "Evaluation Metrics" section of the paper: 1. Accuracy represents the proportion of correct predictions out of all predictions: , in: True Positives True Negatives False positives False Negatives 2. Precision represents the proportion of samples predicted as positive that are actually positive: , 3. Recall represents the proportion of samples that are actually positive but were correctly predicted as positive: , 4. The F1 score is the harmonic mean of precision and recall, used to balance the two: , 5. AUC is the area under the ROC curve, measuring the model's ability to distinguish between positive and negative samples at different classification thresholds. Its mathematical expression is usually defined using integral or ranking statistics: , in: , , In practice, AUC can also be interpreted as the probability that the model scores the positive sample higher than the negative sample when a positive sample and a negative sample are randomly selected.

[0047] The tree-based ensemble model LightGBM stands out with its superior overall performance: its accuracy reaches 0.9502, precision 0.9441, F1 score 0.9649, and AUC 0.9875, ranking first in all four key metrics. Its recall (0.9881) is only slightly lower than Random Forest (0.9907), a difference of only 0.0026, indicating that the model effectively balances false positive control while maintaining high positive example recognition capability. XGBoost (accuracy 0.9417) and Random Forest (accuracy 0.9406) rank second and third respectively, with both achieving high recall (XGBoost: 0.9850, Random Forest: 0.9907) and AUC (XGBoost: 0.9778, Random Forest: 0.9907). The robust performance of the Gradient Boosting Tree (GBR) model (0.9804) demonstrates its strong fitting ability for complex nonlinear relationships. Notably, while the GBR achieves a high recall (0.9859), its precision (0.8183) and accuracy (0.8189) are significantly low, revealing the risk of false positives associated with high recall strategies. SVM performs the worst across all metrics (accuracy 0.8589, AUC 0.9017), indicating its insufficient adaptability to the feature distribution of this dataset. LightGBM achieves the best balance between overall performance, robustness to classification thresholds, and computational efficiency, making it the preferred model for this task. When business scenarios are extremely sensitive to the cost of missed detections, Random Forest can serve as an alternative for high recall.

[0048] (4) A simplified formula suitable for rapid on-site identification was derived. Based on the logistic regression modeling framework of 5-fold cross-validation, a formula for distinguishing wild fish from farmed fish suitable for rapid on-site application was constructed. This formula is based on five easily measurable and statistically significant morphological features. While ensuring high discrimination accuracy, it also takes into account the ease of calculation. The five feature variables used were all screened through three correlation analyses (p value < 0.01), covering the segmented measurement values ​​of body height and body length and their standardized ratios, specifically including: body_high, body_length_1, body_length_2, body_length_3, and body_length_total. Among them, all five features are significantly negatively correlated with the farming status, indicating that their values ​​tend to point to farmed fish. On this basis, by performing Z-score standardization on the training set in each fold cross-validation and fitting a logistic regression model, the following discrimination function was finally obtained by summing the mean values ​​of the 5-fold coefficients: , The discrimination rule is set based on the decision boundary of the logistic regression output probability: when At that time, it was determined to be a wild fish; when At that time, it was determined to be farmed fish. This rule is equivalent to predicting probability. This avoids exponential calculations and significantly improves the feasibility of manual calculation.

[0049] (5) The model performance was verified by 5-fold cross-validation, which showed that the formula has excellent generalization ability and stability, with an overall accuracy of 90.28% and an AUC value as high as 95.39%. The category specificity analysis showed that the recall rate for wild fish was 92.74%. The operation includes the following steps: (a) Nine morphological indicators were measured using a digital caliper with an accuracy of not less than 0.1 mm; (b) Substitute into the above formula to calculate the discriminant function value. ; (c) Based on The classification decision is made based on the symbol, for boundary samples. It is recommended to retest key ratio features or combine auxiliary methods for verification.

[0050] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.

Claims

1. A multimodal fish phenotypic plasticity classification method integrating two-dimensional keypoint detection and three-dimensional structured light, characterized in that, The specific steps include the following: (1) Select the data acquisition equipment and collect data on large yellow croaker; (2) Use acquisition equipment to acquire two-dimensional and three-dimensional geometric features of the fish body, and define and label key points; (3) By comparing the performance of various machine learning models through systematic experiments, key discriminant features with statistical significance were selected; (4) A simplified formula suitable for rapid on-site identification was derived. Based on the logistic regression modeling framework of 5-fold cross-validation, a formula for distinguishing wild fish from farmed fish suitable for rapid on-site application was constructed. This formula is based on five easily measurable and statistically significant morphological features. While ensuring high discrimination accuracy, it also takes into account the ease of calculation. The five feature variables used were all screened by three correlation analyses with p-values ​​< 0.

01. They cover the segmented measurement values ​​of body height and body length and their standardized ratios, specifically including: body_high, body_length_1, body_length_2, body_length_3, and body_length_total. Among them, all five features are significantly negatively correlated with the farming status, indicating that their values ​​tend to point to farmed fish. On this basis, the training set was standardized by Z-score in each fold cross-validation and a logistic regression model was fitted. Finally, the mean values ​​of the 5-fold coefficients were summarized to obtain the following discrimination function: , The discrimination rule is set based on the decision boundary of the logistic regression output probability: when At that time, it was determined to be a wild fish; when When the fish is classified as farmed, this rule is equivalent to predicting a probability Pwild > 0.5, which avoids exponential calculations and significantly improves the feasibility of manual calculation. (5) The model performance was verified by 5-fold cross-validation, which showed that the formula has excellent generalization ability and stability.

2. The multimodal fish phenotypic plasticity classification method integrating two-dimensional keypoint detection and three-dimensional structured light according to claim 1, characterized in that, In step (1), the data acquisition device used was the Zivid Two M70 high-precision structured light 3D camera manufactured by the Norwegian company Zivid. A total of 1,408 large yellow croaker data were collected, including 416 farmed large yellow croaker data and 992 sea-fished large yellow croaker data, which included large yellow croaker point cloud data and color RGB data.

3. The multimodal fish phenotypic plasticity classification method integrating two-dimensional keypoint detection and three-dimensional structured light according to claim 1, characterized in that, The points collected in step (1) are stored in cloud PLY format. Each file contains approximately 2,332,800 vertices, and each vertex accurately records the floating-point X, Y, Z spatial coordinates and unsigned character RGB color components.

4. The multimodal fish phenotypic plasticity classification method integrating two-dimensional keypoint detection and three-dimensional structured light according to claim 1, characterized in that, In step (2), the midline of the fish body is discretized into multiple line segments. The segmented arc length is used to describe the body shape of different types of fish, and a labeling scheme for key points of the fish body is constructed. At the same time, the height of the fish body and the pectoral fins are labeled.

5. The multimodal fish phenotypic plasticity classification method integrating two-dimensional keypoint detection and three-dimensional structured light according to claim 1, characterized in that, In step (3), the three-dimensional fish geometric features are extracted, and the mapping relationship between the image and the point cloud data is established. The dataset uses a structured light three-dimensional scanning system and a synchronously triggered two-dimensional RGB camera to acquire multimodal data. The structured light camera generates point cloud data containing spatial coordinates X, Y, Z and RGB color information by projecting an encoded grating and capturing deformation patterns. The data is stored in PLY format. At the same time, the two-dimensional camera records high-resolution PNG projection images of the corresponding viewpoint.

6. The multimodal fish phenotypic plasticity classification method integrating two-dimensional keypoint detection and three-dimensional structured light according to claim 1, characterized in that, In step (3), to achieve multimodal data fusion, a precise geometric mapping relationship is established. The 3D point cloud is back-projected back onto the 2D image plane according to the imaging model, ensuring that the point cloud data and the original image are aligned at the pixel level. The specific technical implementation is as follows: The origin of the structured light camera coordinate system coincides with the center of the imaging plane, and the optical axis is along... The positive half-axis direction, therefore we referenced the principle of structured light cameras based on any point in the given point cloud. Its back projection onto the target plane The process is achieved through ray parameterization, where, To preset the target depth, 50% of the maximum Z-value of the point cloud is taken to balance perspective distortion. The calculation formula is as follows: , Back projection point The coordinates are: , Then The XY range is linearly normalized to the original image resolution: width W, height H. , , This indicates rounding down. Using pixel coordinates, the point cloud color values ​​are ultimately written directly to the corresponding pixel positions to generate the reconstructed image.

7. The multimodal fish phenotypic plasticity classification method integrating two-dimensional keypoint detection and three-dimensional structured light according to claim 1, characterized in that, In step (3), the performance of six mainstream machine learning algorithms on classification tasks is evaluated. Accuracy, Precision, Recall, F1 score and AUC value are used as core evaluation indicators. During training, 5-fold true values ​​are used for training. After training, the test set is used for testing. All models are trained and validated using features calculated from the true value labels.

8. The multimodal fish phenotypic plasticity classification method integrating two-dimensional keypoint detection and three-dimensional structured light according to claim 1, characterized in that, In step (3), the traditional machine learning model evaluation metrics commonly used in machine learning and classification tasks include accuracy, precision, recall, F1-score, and AUC. Their mathematical definitions are as follows, applicable to the evaluation metrics section of the paper:

1. Accuracy represents the proportion of correct predictions out of all predictions: , in: TP A real example, TN : True negative example, FP False positive example FN : False negative example; 2. Precision represents the proportion of samples predicted as positive that are actually positive: , 3. Recall represents the proportion of samples that are actually positive but were correctly predicted as positive: , 4. The F1 score is the harmonic mean of precision and recall, used to balance the two: , 5. AUC is the area under the ROC curve, measuring the model's ability to distinguish between positive and negative samples at different classification thresholds. Its mathematical expression is defined through integral or ranking statistics: , in: , , In practice, AUC can also be interpreted as the probability that the model scores the positive sample higher than the negative sample when a positive sample and a negative sample are randomly selected.

9. The multimodal fish phenotypic plasticity classification method integrating two-dimensional keypoint detection and three-dimensional structured light according to claim 1, characterized in that, In step (3), the tree-based ensemble model LightGBM stands out with its best overall performance. LightGBM achieves the best balance between overall performance, classification threshold robustness and computational efficiency, making it the preferred model for this task. When the business scenario is extremely sensitive to the cost of missed detections, random forest can be used as an alternative with high recall.

10. The multimodal fish phenotypic plasticity classification method integrating two-dimensional keypoint detection and three-dimensional structured light according to claim 1, characterized in that, In step (5), the cross-validation operation includes the following steps: (a) Nine morphological indicators were measured using a digital caliper with an accuracy of not less than 0.1 mm; (b) Substitute the values ​​into the above formula to calculate the discriminant function value; (c) Make classification decisions based on the symbols used. For boundary samples, it is recommended to retest key ratio features or combine auxiliary methods for verification.