Live fish injury detection method based on AI and multi-vision underwater imaging technology

Through the improved ResNet-50 convolutional neural network, combined with multi-eye visual underwater imaging technology, the problems of interference factors such as light change, noise and blur in underwater live fish damage detection are solved, and high-precision damage area identification and feature extraction are achieved, enhancing the generalization ability and robustness of the model.

CN120259865AActive Publication Date: 2025-07-04SHANDONG AGRI TECH EXTENSION GENERAL STATION
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510381873.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-04
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

The prior art is difficult to effectively deal with interfering factors such as light changes, noise and blur in the detection of underwater live fish, resulting in insufficient identification accuracy and detection accuracy, and easy to overfit, making it impossible to effectively extract distinguished features.

Method used

Using AI and multi-eye visual underwater imaging technology, an improved ResNet-50 convolutional neural network is constructed. Feature extraction and classification are optimized through expanded convolution, Inception module, global context pooling, combination regularization, dynamic learning rate and step size adjustment, category weight, multi-stage optimization and adaptive sample reweighting strategies.

Benefits of technology

It improves the accuracy and robustness of underwater live fish damage detection, reduces the overfitting problem, enhances the adaptability and generalization ability to multi-scale damage areas, and solves the problem of category imbalance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259865A_ABST
    Figure CN120259865A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence and visual detection, in particular to a live fish injury detection method based on AI and a multi-view visual underwater imaging technology, and the method specifically comprises the following steps: setting a multi-view visual imaging device, collecting a clear and complete fish body image, carrying out the image denoising, graying, contrast enhancement and normalization operation of the collected image, and carrying out the detection of the damage of a live fish. Obtaining a pre-processed fish body image; a convolutional neural network based on ResNet-50 serving as a main body frame is constructed to serve as a live fish injury grading detection model, the convolutional neural network based on ResNet-50 serving as the main body frame is improved, and then a preprocessed fish body image is input into the live fish injury grading detection model for training. According to the method, the overfitting problem can be effectively reduced, the feature extraction effect is optimized, the classification precision is improved, the problem of class imbalance in the underwater image is solved, and the detection accuracy is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence and visual detection, and particularly relates to a method for detecting live fish injuries based on AI and multi-camera underwater imaging technology. Background Art

[0002] The global aquaculture scale is growing increasingly large. According to statistics in 2024, the injury rate during the transportation and temporary cultivation of live fish is as high as 15%-30%, causing huge economic losses. Therefore, it is necessary to detect the injuries of underwater live fish to avoid greater economic losses. The shooting of underwater images is affected by factors such as water environment, lighting conditions, and water surface fluctuations, and usually presents characteristics such as blurriness, low contrast, and uneven lighting. These factors bring great difficulties to image processing and injury detection. Especially in the task of detecting live fish injuries, the dynamic changes of the fish body and the subtle differences in the injury areas increase the complexity of detection.

[0003] Existing technologies generally include traditional optical detection technologies relying on monocular vision and near-infrared imaging, acoustic and mechanical sensor technologies, and artificial intelligence technologies based on convolutional neural networks; traditional convolutional neural networks have poor adaptability to interference factors such as lighting changes, noise, and blurriness in underwater imaging, resulting in the inability to effectively capture the characteristics of injury areas and affecting the recognition accuracy; traditional methods are prone to bias towards the majority class when facing minority class samples, resulting in poor recognition ability for injury areas and affecting the detection accuracy; traditional training methods do not have effective regularization strategies, which are prone to overfitting phenomena. Especially in the case of insufficient data volume, it is difficult to improve the generalization ability; existing technologies usually generate a large number of redundant features in underwater imaging, leading to the network being prone to overfitting and unable to effectively extract discriminative features; Therefore, the present invention proposes a method for detecting live fish injuries based on AI and multi-camera underwater imaging technology to solve the above problems. Summary of the Invention

[0004] In view of the deficiencies of the existing technology, the present invention develops a method for detecting live fish injuries based on AI and multi-camera underwater imaging technology. Through artificial intelligence and multi-camera underwater imaging technology, the present invention can accurately identify and extract meaningful local and global features, while dealing with the problems of redundant features in high-dimensional data, inconsistent feature scales, and avoiding overfitting.

[0005] The technical solution for the present invention to solve the technical problem is a method for detecting live fish injuries based on AI and multi-camera underwater imaging technology, including the following steps: S1. Image acquisition: Set up a multi-camera imaging device, collect clear and complete fish body images, and adjust the shooting angle and focal length of the multi-camera imaging device in real time according to the size and movement state of the fish body; S2. Image preprocessing: Perform image denoising, grayscale conversion, contrast enhancement, and normalization operations on the collected images to obtain the preprocessed fish body images; S3. Construct a live fish injury grading detection model for live fish injury detection: Construct a convolutional neural network based on ResNet-50 as the main framework as the live fish injury grading detection model. Improve the convolutional neural network based on ResNet-50 as the main framework, adjust the shallow convolution of ResNet-50, replace the first three layers of convolution with dilated convolution, and reset the spacing between the convolutional kernel elements. Insert the lightweight deep learning convolutional neural network architecture Inception after the third stage Stage3 of ResNet-50, replace the fully connected layer of the output layer of ResNet-50 with global context pooling, and then input the preprocessed fish body images into the live fish injury grading detection model for training.

[0006] In the specific implementation manner, the improvement of the convolutional neural network based on ResNet-50 as the main framework is as follows: Adopt a heuristic method to initialize the convolutional kernel according to the feature distribution of the fish body image, and then calculate the convolutional kernel weight W; Calculate the combined regularization term , and embed the combined regularization term into the feature extraction process of each layer of the convolutional neural network; Dynamically adjust the learning rate and step size according to the error change situation of the fish body image data; Dynamically adjust the importance of feature extraction by the convolutional layer according to the multi-view visual data of the fish body image; Adopt class weights in the loss function; Pay attention to feature extraction and classification optimization of different classes through multi-stage optimization; Reduce redundancy through a weight sharing mechanism; Adopt an adaptive sample reweighting strategy and a multi-stage verification mechanism.

[0007] In the specific implementation manner, during the dynamic update process of the convolutional kernel weight, dynamically adjust the learning rate and step size according to the error change situation of the fish body image data. The update formula of the convolutional kernel weight is as follows: , where, represents the convolutional kernel weight of the th iteration, represents the convolutional kernel weight of the th iteration, represents the gradient of the loss function with respect to the convolutional kernel weight of the th iteration, represents the learning rate of the th iteration, Indicates the step size factor for the th iteration.

[0008] In the specific implementation, the learning rate and step size in the process of updating the weights of the convolutional kernel are automatically adjusted according to the error change; For the learning rate , it is calculated through the initial learning rate, error dynamic adjustment coefficient, and error metric; For the step size , it is calculated through the initial step size factor, adjustment coefficient of the step size factor, and error metric; Among them, the calculation of the error dynamic adjustment coefficient adopts the cosine annealing strategy; the calculation of the error metric adopts a composite error function, including cross-entropy loss and training error; the adjustment coefficient of the step size factor is calculated by dynamically adjusting the step size amplitude in combination with the current detection accuracy.

[0009] In the specific implementation, during the adaptive fusion process, the importance of feature extraction by the convolutional layer is dynamically adjusted according to the multi-view visual data of the fish body image, and the features extracted by the convolutional layer of each layer of the convolutional neural network are weighted and fused, and the weighted coefficients of each layer of features are dynamically calculated by combining local and global information, where the local information is further represented by the correlation between the features and the input samples, and the correlation is calculated by the Pearson correlation coefficient.

[0010] In the specific implementation, class weights are used in the loss function to balance the minority classes and the majority classes, and the calculation formula is as follows: , Among them, represents the weighted loss, represents the total number of input samples, represents the dynamic weighted coefficient of the class corresponding to the th sample, represents the weight of the th iteration of the th sample, represents the standard loss function, represents 's predicted label, represents the th input sample, represents 's true label.

[0011] In the specific implementation, a dynamic weighted factor based on the data distribution is used to adjust the weighted coefficients of each category. Specifically, the data distribution is determined according to the number of samples with the label and the total number of input samples, and the corresponding dynamic weighted coefficient is calculated by combining the density of the label .

[0012] In the specific implementation manner, in different training stages, multi-stage optimization is adopted to focus on feature extraction and classification optimization of different categories, and the calculation formula is as follows: , where represents the total loss after integration, represents the total number of training stages, represents the weighted coefficient of the th training stage, represents the th training stage, and

[0013] represents the loss value of the In the specific implementation manner, the weighted coefficients of each stage are corrected according to the strength of the association between the stage loss and the features.

[0014] Specifically, the Pearson correlation coefficient is calculated to determine the correlation between the features output by each stage of the convolutional neural network and the input samples, and the weighted values of the input samples for the features output by the convolutional neural network are calculated to determine the association between the loss of each stage and the features. , where represents the weight of the th sample in the th iteration, represents the exponential function, represents the smoothing coefficient, represents the function for measuring the overlap degree between the prediction of the convolutional neural network for the th sample and the true damage area, represents the entropy calculation function, represents the calculation of the th sample feature map chaos degree to evaluate the sample clarity, represents the maximum value of the entropy value among all samples; The training process includes three stages, namely the feature convergence period, the regularization strengthening period, and the fine-tuning period; Feature convergence period: First, optimize the backbone network. The backbone network is ResNet-50 and the Inception module, and a high initial learning rate is used to quickly capture the basic features of the damage area; Regularization Enhancement Phase: Suppress overfitting by randomly discarding 20% of the low-weight feature channels, and at the same time enhance robustness by generating adversarial samples; Fine Tuning Phase: Freeze the parameters of the shallow network, only fine-tune the high-level modules and classification head after Stage 3 of the third stage, adopt a low initial learning rate, force the model to map the multi-view image features from different perspectives to a unified semantic space, and constrain the feature consistency of the same-class multi-view samples.

[0015] The effects provided in the invention content are only the effects of the embodiments, rather than all the effects of the invention. The above technical solutions have the following advantages or beneficial effects: The present invention adjusts the convolutional neural network structure, uses dilated convolution to expand the receptive field, can effectively capture the large-scale damage features in underwater blurred images, and by adding the Inception module to ResNet-50, enables the network to process damage areas of different scales, enhances the joint perception of low-frequency (contour) and high-frequency (texture) features, and improves the adaptability to multi-scale damage areas. In addition, by replacing the traditional fully connected layer with global context pooling, redundant parameters can be reduced, thereby optimizing the network structure and improving the generalization ability of the model; in terms of the heuristic optimization of convolutional kernel initialization, the convolutional kernel is optimized and initialized according to the characteristics of underwater images (such as illumination changes, blur, and noise), which can enhance the response ability of the convolutional kernel to underwater images and overcome the limitations of traditional methods.

[0016] The present invention combines a regularization strategy, combines L1 regularization and L2 regularization, embeds a combined regularization term in the feature extraction process of each layer, suppresses redundant features and strengthens discriminative features, thereby effectively reducing the overfitting problem; it also dynamically adjusts the regularization coefficient, dynamically adjusts the regularization coefficient of each layer during the training process, enabling the model to adapt to different noise levels and feature importance, and can further optimize the feature extraction effect.

[0017] The present invention also dynamically adjusts the learning rate and step size, dynamically adjusts the learning rate and step size according to the error, can ensure that the network converges quickly in the initial stage of training, finely adjusts the weights in the later stage, avoids overfitting, and improves the classification accuracy; for the handling of the class imbalance problem, by using class weighting in the loss function, the attention to minority class samples (such as damage areas) is strengthened, thereby solving the class imbalance problem in underwater images; the present invention also relates to a multi-stage optimization strategy, which is divided into a feature convergence phase, a regularization enhancement phase, and a fine tuning phase. By gradually optimizing the model in each stage, the detection ability for damage areas can be improved and the robustness can be enhanced.

[0018] In summary, the method proposed in the present invention is different from conventional machine vision tasks. When using multi-view vision underwater imaging technology for live fish injury detection, the challenges mainly lie in dealing with interference factors such as blurring, noise, and multiple illuminations in underwater images, as well as the dynamic changes of the fish body, which require the detection model to accurately identify and extract meaningful local and global features in a multi-dimensional feature space, while coping with the problems of redundant features in high-dimensional data, inconsistent feature scales, and avoiding overfitting. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention.

[0020] Figure 1 It is a schematic flowchart of the method of the present invention.

[0021] Figure 2 It is an image of a live fish collected by a multi-view vision imaging device.

[0022] Figure 3 It is an effect diagram of feature extraction after adding the Inception module.

[0023] Figure 4 It is a comparison example of feature maps with and without the constraint of the combined regularization term.

[0024] Figure 5 It is a feature map at different training stages.

[0025] Figure 6 It is a comparison chart of the recognition accuracy of different network structures.

[0026] Figure 7 It is a comparison chart of the influence of different hyperparameter combinations on the model performance. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] In order to clearly illustrate the technical features of the present solution, the present invention will be described in detail below through specific embodiments and in conjunction with their accompanying drawings. The following disclosure provides many different embodiments or examples for implementing different structures of the present invention. In order to simplify the disclosure of the present invention, the components and settings of specific examples are described below.

[0028] Embodiment 1 As Figures 1 - 5 shown, a method for live fish injury detection based on AI and multi-view vision underwater imaging technology includes the following steps: S1. Image acquisition: Set up a multi-view vision imaging device to collect clear and complete fish body images, and adjust the shooting angle and focal length of the multi-view vision imaging device in real time according to the size and movement state of the fish body; Among them, the live fish image captured by the multi-view vision imaging device is as Figure 2 shown; S2. Image preprocessing: Perform image denoising, grayscale conversion, contrast enhancement, and normalization operations on the collected images to obtain the preprocessed fish body images; S3. Construct a live fish injury grading detection model for live fish injury detection: Construct a convolutional neural network based on ResNet-50 as the main framework as the live fish injury grading detection model, improve the convolutional neural network based on ResNet-50 as the main framework, adjust the shallow convolution of ResNet-50, replace the first three layers of convolution with dilated convolution, and reset the spacing between the convolutional kernel elements. Insert the lightweight deep learning convolutional neural network architecture Inception after the third stage Stage3 of ResNet-50, replace the fully connected layer of the output layer of ResNet-50 with global context pooling, and then input the preprocessed fish body images into the live fish injury grading detection model for training; Among them, the feature extraction effect diagram after inserting the lightweight deep learning convolutional neural network architecture Inception is as Figure 3 shown. The addition of the Inception module can significantly enhance the significant difference between the injury area and other areas.

[0029] In the specific implementation manner, the improvement of the convolutional neural network based on ResNet-50 as the main framework is as follows: Adopt a heuristic method to initialize the convolutional kernel according to the feature distribution of the fish body image; Embed the combined regularization term into the feature extraction process of each layer of the convolutional neural network; Dynamically adjust the learning rate and step size according to the error change situation of the fish body image data; Dynamically adjust the importance of feature extraction by the convolutional layer according to the multi-view vision data of the fish body image; Adopt class weights in the loss function; Pay attention to the feature extraction and classification optimization of different classes through multi-stage optimization; Reduce redundancy through the weight sharing mechanism; Adopt an adaptive sample reweighting strategy and a multi-stage verification mechanism.

[0030] In the specific implementation manner, initialize the convolutional neural network structure. In the multi-view vision underwater imaging task, the image data distribution under each camera view varies greatly. Therefore, the initialization of the convolutional kernel of the convolutional neural network needs to be optimized for these differences. Adopt a heuristic method to initialize the convolutional kernel of the convolutional neural network according to the feature distribution of the fish body image, so that the convolutional kernel can better adapt to the feature differences brought by factors such as light changes, fluctuations, and object movements in the image. The calculation formula is as follows: , Among them, represents the convolution kernel weight, which characterizes the learning parameters of each convolution filter. N represents the dimension of the input feature map, which characterizes the number of channels of each image in the convolution operation. represents a normal distribution with a mean of 0 and a covariance of . represents the covariance matrix of the input feature map, which is used to initialize according to the feature distributions of different underwater images; The heuristic method refers to the initialization method for the convolution kernel weight to adapt to the feature distribution of underwater images. Traditional convolutional neural networks usually use Gaussian distribution or uniform distribution for initialization. However, due to the particularity of underwater imaging (such as light changes and water surface fluctuations), using the heuristic method can better adapt to the feature distribution of images. The heuristic method adjusts the weight initialization according to the characteristics of underwater images (such as blurring and noise) (affecting the convolution kernel weight through the covariance matrix of image features), and optimizes the response ability of the convolution kernel in underwater images.

[0031] In the specific implementation manner, underwater imaging technology may cause a large amount of redundant information in the high-dimensional feature space of images. Traditional convolutional neural networks are difficult to effectively handle these redundant features and are prone to overfitting. A combined regularization term is embedded into the feature extraction process of each layer of the convolutional neural network to solve the multicollinearity and high-dimensional redundancy problems in underwater images, suppress redundant features, and strengthen discriminative features. The combined regularization term is obtained by combining L1 regularization and L2 regularization; The calculation formula of the convolution layer in the convolutional neural network is as follows: , Among them, represents the image features output by the convolution layer, represents the image features output by the previous layer, and b represents the bias term; The combined regularization term is embedded into the process of extracting image features by each convolution layer, and the regularization strength in the feature extraction process of each layer is dynamically adjusted. The calculation formula of the combined regularization term is as follows: , Among them, represents the combined regularization term, which characterizes the constraint of the weights of each convolution layer, represents the number of layers of the convolutional neural network, represents the -th layer convolution L1 regularization coefficient, which is used to control the sparsity of features, represents the -th layer convolution L2 regularization coefficient, which is used to control the smoothness of features, It represents L1 regularization. By penalizing the absolute values of the weights, it promotes some weights to become zero, thus achieving feature selection. It represents L2 regularization. By penalizing the sum of the squares of the weights, it promotes the weights to become smaller, increases the smoothness of the model, and avoids overfitting. It represents the convolution kernel weights of the layer, which characterize the learning ability of each layer of convolutional features. By adopting a calculation method of adaptive regularization coefficients, the convolutional neural network can dynamically select the most discriminative features according to the synergistic effects of different image features during the training process, suppress the interference of redundant information and noise on the identification of damaged areas, and dynamically calculate the regularization coefficients according to the features and noise levels of the images. The calculation formula is as follows: , , Among them, represents the covariance matrix of the feature image of the layer, represents the noise level, represents the feature importance, which is measured by calculating the information gain of the feature on the model output.

[0032] It can be seen from Figure 4 that the L1 regularization and L2 regularization constraints included in the combined regularization term can effectively eliminate noise features, and at the same time can effectively highlight the damaged area, which is beneficial to the accurate identification of live fish damage. It can be seen from the binarized feature map without the combined regularization term constraint in the traditional method that not only the noise is not effectively reduced, but also new noise areas are introduced.

[0033] In the specific implementation, during the dynamic update process of the convolution kernel weights, the learning rate and step size are dynamically adjusted according to the error change situation of the fish body image data, so that the convolutional neural network converges quickly in the initial stage of training, and at the same time finely adjusts the weight parameters in the later stage, thus avoiding overfitting and improving the classification accuracy in the detection task of the live fish damaged area. The update formula of the convolution kernel weights is as follows: , Among them, represents the convolution kernel weights of the represents the convolution kernel weights of the represents the gradient of the loss function with respect to the convolution kernel weights at the represents the learning rate at the represents the step size factor at the Automatically adjust the learning rate and step size according to the error change, and the calculation formula is as follows: , , wherein, represents the initial learning rate, represents the initial step size factor, represents the error dynamic adjustment coefficient of the th iteration, represents the error metric of the th iteration, represents the adjustment coefficient of the step size factor of the th iteration; The error dynamic adjustment coefficient adopts a cosine annealing strategy to prevent late oscillation, and the calculation formula is as follows: , wherein, represents the total number of iterations; The error metric adopts a composite error function, and the calculation formula is as follows: , wherein, represents the cross-entropy loss, represents the training error of the th iteration; The adjustment coefficient of the step size factor dynamically adjusts the step size amplitude in combination with the current detection accuracy, and the calculation formula is as follows: , wherein, represents the training accuracy of the th iteration.

[0034] In the specific implementation manner, during the adaptive fusion process, the importance of feature extraction by the convolutional layer is dynamically adjusted according to the multi-view visual data of the fish body image. The damaged area of the fish body may contain subtle feature changes. Therefore, it is necessary to accurately weight the feature extraction effects of each layer to maximize the recognition ability of the damaged area. The calculation formula of the weighted fusion process is as follows: , wherein, is the fused feature map, representing the comprehensive representation of the features extracted by the multi-layer convolutional network, represents the weighted coefficient of the features of the th layer of the convolutional neural network, representing the contribution of the features of this layer to damage detection, represents the features extracted by the th convolutional layer of the convolutional neural network, representing the encoding of local or global features of the image by each layer; To refine the weighted process of each layer's features, the weighted coefficient of each layer is dynamically calculated by combining local and global information, and the calculation formula is as follows: , where, represents the number of input live fish images, that is, the number of samples, represents the th sample, represents the L2 norm of the features extracted by the th sample in the th convolutional layer of the convolutional neural network, represents the correlation between the features of the th layer of the convolutional neural network and the th sample; The correlation is calculated by computing the Pearson correlation coefficient, and the calculation formula is as follows: , where, represents the mean value of the feature , represents the mean value of the sample .

[0035] In the specific implementation manner, class weights are adopted in the loss function to balance the minority classes and the majority classes, and the calculation formula is as follows: , where, represents the weighted loss, represents the dynamic weighted coefficient of the class corresponding to the th sample, represents the weight of the th iteration of the th sample, represents the standard loss function, represents 's predicted label, represents 's true label; To precisely adjust the weights of each class, a dynamic weighting factor based on the data distribution is adopted to adjust the weighted coefficients of each class, which can not only improve the accuracy and robustness of live fish damage detection, but also effectively solve the interference problems unique to underwater imaging such as noise and blurring in the image, and the calculation formula is as follows: , where, represents the set of various class labels, represents the number of samples with the label , represents 's density.

[0036] In the specific implementation manner, in different training stages, a multi-stage optimization is adopted to focus on feature extraction and classification optimization of different categories in order to maximize the performance of the model. The present invention adopts a multi-stage ensemble learning method. By separately focusing on different subsets of features or the performance of specific categories in multiple training stages, the results learned in each stage are fused, so as to obtain higher accuracy and stability under complex data distributions. The calculation formula is as follows: , Among them, represents the total loss after integration, represents the total number of training stages, represents the weighting coefficient of the th training stage, and represents the loss value of the th training stage; To take into account the influence of each stage on the features of different samples, the weighting coefficient is corrected according to the strength of the association between the stage loss and the features. The calculation formula is as follows: , Among them, represents the weighting of the output features of the convolutional neural network by the th sample in the stage, is the correlation between the features output by the convolutional neural network in the stage and the th sample.

[0037] In the specific implementation manner, different damaged areas of the fish body in the image are expressed differently in the feature space using multiple convolutional layers. For this, a weight sharing mechanism is introduced to reduce redundancy. A shared matrix is established among multiple convolutional kernels to reduce duplicate parameters and strengthen the attention to common features. The calculation formula is as follows: , Among them, represents the shared convolutional kernel weight, represents the number of convolutional kernels, represents the weight of the th convolutional kernel layer, is the number of layers of the convolutional neural network.

[0038] In the specific implementation manner, during the iterative process of the live fish damage grading detection model, an adaptive sample reweighting strategy and a multi-stage verification mechanism are adopted to dynamically screen high-value training samples and adjust the training direction, and the weights are dynamically adjusted according to the prediction confidence and feature consistency of the samples, reducing the contribution of noise samples and fuzzy samples and strengthening the training weights of samples with clear damaged areas. The calculation formula is as follows: , Among them, represents the weight of the th iteration of the th sample, represents the exponential function, represents the smoothing coefficient, represents a function to measure the overlap degree between the prediction of the convolutional neural network for the th sample and the true damage area, represents the entropy value calculation function, represents calculating the confusion degree of the feature map of the th sample to evaluate the clarity of the sample, represents the maximum value of the entropy value among all samples; The training process includes three stages, namely the feature convergence period, the regularization enhancement period, and the fine-tuning period; Feature convergence period: Optimize the backbone network first. The backbone network is ResNet-50 and the Inception module. Use a high initial learning rate to quickly capture the basic features of the damage area; Regularization enhancement period: Suppress overfitting by randomly discarding 20% of the low-weight feature channels. At the same time, use adversarial sample generation to enhance robustness; Fine-tuning period: Freeze the parameters of the shallow network. Only fine-tune the high-level modules and the classification head after Stage 3 of the third stage. Use a low initial learning rate to force the model to map the multi-view image features from different perspectives to a unified semantic space and constrain the feature consistency of the same-class multi-view samples; During the training process, if the performance improvement in consecutive iterations is lower than the preset threshold, stop the iteration, which means the training of the convolutional neural network is completed.

[0039] As Figure 5 shown, visualize the features extracted in the three stages. It can be seen from the figure that as the number of training iterations increases, the extracted features become clearer and the damaged parts of the live fish become more prominent.

[0040] Example 2 Compare the recognition accuracy of the benchmark residual network ResNet-50 and the improved network ResNet-Inception with an added multi-branch structure with the method in the present invention. The comparison results are as Figure 6As shown, ResNet-50 has a weak ability to identify damage types with subtle textures, ResNet-Inception has improved the ability to capture multi-scale features, while the method in the present invention achieves the optimal detection effect for all damage types, verifying the synergistic effect of the multi-branch structure and dynamic feature weighted fusion in the present technology. By paralleling convolutional operations of different scales, the model's joint perception ability of damage contours and textures is enhanced. At the same time, the contributions of each layer are dynamically adjusted in combination with feature correlation, enabling the model to capture large-scale structural damages such as fin cracks and identify local subtle lesions such as body surface ulcers, ultimately achieving precise positioning of multi-form damages.

[0041] Example 3 As Figure 7 shown, the parameter sensitivity of the traditional method and the parameter robustness of the method in the present invention are analyzed. The performance surface of the traditional method shows steep island characteristics and only achieves better performance in a specific parameter interval, indicating its high dependence on fine parameter tuning. The performance surface of the technology in the present invention shows a broad plateau region and maintains stable performance within a large variation range of the regularization coefficient and learning rate, indicating the technical advantages of the adaptive regularization mechanism and dynamic learning rate adjustment. By establishing a self-adjusting relationship between parameters, the model can automatically adapt to the data distribution characteristics of different underwater imaging conditions, reducing the dependence on manual parameter tuning and enhancing the engineering feasibility in practical applications.

[0042] Although the specific implementation mode of the invention is described above in conjunction with the drawings, it is not a limitation on the protection scope of the present invention. Based on the technical solution of the present invention, various modifications or deformations that can be made by those skilled in the art without creative labor are still within the protection scope of the present invention.

Claims

1. A method for detecting live fish injuries based on AI and multi-camera underwater imaging technology, characterized in that, It includes the following steps: S1. Image acquisition: Set up a multi - vision imaging device to acquire clear and complete fish body images, and adjust the shooting angle and focal length of the multi - vision imaging device in real time according to the size and movement state of the fish body; S2. Image pre - processing: Perform image denoising, grayscale conversion, contrast enhancement and normalization operations on the acquired images to obtain pre - processed fish body images; S3. Construct a live fish injury grading detection model for live fish injury detection: Construct a convolutional neural network based on ResNet - 50 as the main framework as the live fish injury grading detection model. Improve the convolutional neural network based on ResNet - 50 as the main framework, adjust the shallow - layer convolution of ResNet - 50, replace the first three layers of convolution with dilated convolution, and reset the spacing between the elements of the convolution kernel. Insert the lightweight deep - learning convolutional neural network architecture Inception after the third stage Stage3 of ResNet - 50, replace the fully - connected layer of the output layer of ResNet - 50 with global context pooling, and then input the pre - processed fish body images into the live fish injury grading detection model for training.

2. The method for detecting live fish damage based on AI and multi-camera underwater imaging technology according to claim 1, characterized in that, The improvement of the convolutional neural network based on ResNet - 50 as the main framework is as follows: Adopt a heuristic method to initialize the convolution kernel according to the feature distribution of the fish body image, and then calculate the convolution kernel weight W; Calculate the combined regularization term , and embed the combined regularization term into the feature extraction process of each layer of the convolutional neural network; Dynamically adjust the learning rate and step size according to the error change of the fish body image data; Dynamically adjust the importance of feature extraction by the convolutional layer according to the multi - vision data of the fish body image; Adopt class weights in the loss function; Pay attention to feature extraction and classification optimization of different classes through multi - stage optimization; Reduce redundancy through a weight sharing mechanism; Adopt an adaptive sample re - weighting strategy and a multi - stage verification mechanism.

3. A live fish injury detection method based on AI and multi - vision underwater imaging technology according to claim 2, characterized in that: In the process of dynamically updating the convolution kernel weight, the learning rate and step size are dynamically adjusted according to the error change of the fish body image data. The update formula of the convolution kernel weight is as follows: , Among them, represents the convolution kernel weight of the th iteration, represents the convolution kernel weight of the th iteration, represents the gradient of the loss function with respect to the convolution kernel weight at the th iteration, represents the learning rate at the th iteration, represents the step size factor at the th iteration.

4. A live fish injury detection method based on AI and multi - vision underwater imaging technology according to claim 3, characterized in that: Automatically adjust the learning rate and step size in the process of updating the convolution kernel weight according to the error change; For the learning rate , it is calculated through the initial learning rate, the error dynamic adjustment coefficient, and the error metric; For the step size , it is calculated through the initial step size factor, the adjustment coefficient of the step size factor, and the error metric; Among them, the calculation of the error dynamic adjustment coefficient adopts a cosine annealing strategy; the calculation of the error metric adopts a composite error function, including cross - entropy loss and training error; the adjustment coefficient of the step size factor is calculated by dynamically adjusting the step size amplitude in combination with the current detection accuracy.

5. A live fish injury detection method based on AI and multi - vision underwater imaging technology according to claim 4, characterized in that: In the adaptive fusion process, the importance of feature extraction by the convolutional layer is dynamically adjusted according to the multi - vision data of the fish body image, and the features extracted by the convolutional layer of each layer of the weighted fusion convolutional neural network are weighted and fused, and the weighted coefficients of each layer of features are dynamically calculated by combining local and global information. Among them, the local information is further represented by the correlation between the feature and the input sample, and the correlation is calculated by the Pearson correlation coefficient.

6. A method for detecting live fish damage based on AI and multi-view vision underwater imaging technology according to claim 5, characterized in that: Class weights are used in the loss function to balance minority and majority classes, and the calculation formula is as follows: , Among them, represents the weighted loss, represents the total number of input samples, represents the dynamic weighted coefficient of the category corresponding to the th sample, represents the weight of the th sample in the th iteration, represents the predicted label of represents the th input sample, represents the true label of 7. A method for detecting live fish damage based on AI and multi-view vision underwater imaging technology according to claim 6, characterized in that: Adjust the weighting coefficients of each category by using a dynamic weighting factor based on data distribution. Specifically, determine the data distribution according to the number of samples with label and the total number of input samples. Then, calculate the corresponding dynamic weighting coefficient in combination with the density of label . .

8. A method for detecting live fish damage based on AI and multi-view vision underwater imaging technology according to claim 7, characterized in that: In different training stages, multi-stage optimization is adopted to focus on feature extraction and classification optimization of different classes, and the calculation formula is as follows: , Among them, represents the total integrated loss, represents the total number of training phases, represents the weighting coefficient of the th training phase, and represents the loss value of the 9. A method for detecting live fish damage based on AI and multi-view vision underwater imaging technology according to claim 8, characterized in that: Modify the weighted coefficients of each stage according to the strength of the association between the stage loss and the features , determine the correlation between the features output by each stage of the convolutional neural network and the input samples by calculating the Pearson correlation coefficient, and calculate the weighting of the input samples of each stage on the output features of the convolutional neural network, thereby determining the association between the loss of each stage and the features.

10. A method for detecting live fish damage based on AI and multi-view vision underwater imaging technology according to claim 9, characterized in that: In the iterative process of the live fish damage grading detection model, an adaptive sample reweighting strategy and a multi-stage verification mechanism are adopted to dynamically screen high-value training samples and adjust the training direction, and the weights are dynamically adjusted according to the prediction confidence and feature consistency of the samples to reduce the contribution of noise samples and fuzzy samples and strengthen the training weights of clear samples in the damage area. The calculation formula is as follows: , Among them, represents the weight of the th iteration of the th sample, represents the exponential function, represents the smoothing coefficient, represents a function to measure the overlap degree between the prediction of the convolutional neural network for the th sample and the true damage area, represents the entropy value calculation function, represents calculating the confusion degree of the feature map of the th sample to evaluate the clarity of the sample, represents the maximum value of the entropy value among all samples; The training process includes three stages, namely the feature convergence period, the regularization strengthening period, and the fine-tuning period; Feature convergence period: Prioritize optimizing the backbone network, which is the ResNet-50 and Inception modules, and use a high initial learning rate initial value to quickly capture the basic features of the damage area; Regularization strengthening period: Suppress overfitting by randomly discarding 20% of the low-weight feature channels, and at the same time use adversarial sample generation to enhance robustness; Fine-tuning period: Freeze the parameters of the shallow network, only fine-tune the high-level modules and classification heads after the third stage Stage3, use a low learning rate initial value, and force the model to map the multi-view image features from different perspectives to a unified semantic space, and constrain the feature consistency of multi-view samples of the same class.

Citation Information

Patent Citations

  • Underwater fish target detection method and device based on convolutional neural network, and storage medium

    CN113837104A

  • Video subtitle generation method based on hierarchical semantic representation and aggregation network

    CN118590598A

  • Green coffee bean rating and classifying method based on hybrid convolutional neural network structure

    CN118608827A

  • Fish identification method based on convolutional neural network

    CN119206462A

  • Microchip appearance defect detection method based on convolutional neural network

    CN119693363A