Ocean catch abnormal feature detection device and method
Multimodal detection using vision and multiple sensors combined with an improved YOLOv8 model solves the low efficiency problem of traditional methods, achieves fast and accurate detection of deep-sea catches, and meets the automation needs of deep-sea fisheries.
Patent Information
- Application Number
- CN202510970685.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-09-26
AI Technical Summary
Traditional methods for identifying abnormal features of fish catches are inefficient and difficult to cope with the actual situation where fish catches are diverse in type, large in quantity, and have complex features. They also rely on manually designed non-robust features, have limited generalization capabilities, and poor real-time performance, and cannot meet the fast, accurate, and automated processing needs of offshore fisheries.
Multimodal detection is achieved using vision and multiple sensors. Combined with the improved YOLOv8 model, data is collected through electronic tongue sensors, cameras and gas sensors, and multimodal fusion processing is performed to screen out abnormal catches.
It achieves rapid and accurate detection of distant-water catches, improves detection efficiency, and can automatically extract features to meet the needs of modern fishery management.
Smart Images

Figure CN120702536A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of aquatic product detection technology, and in particular to a device and method for detecting abnormal characteristics of deep-sea fish catches. Background Art
[0002] Distant-water fishing plays a crucial role in the development and utilization of marine resources. It not only provides abundant seafood resources for human society but also plays a crucial role in the sustainable development of the fishery industry and the protection of the marine ecosystem. However, as distant-water fishing activities continue to expand and deepen, their management and monitoring face increasingly complex and challenging challenges. One of these challenges is the rapid and accurate identification and management of abnormal characteristics in fish catches.
[0003] Traditional methods for identifying abnormal features in fish catches rely primarily on manual inspection and empirical judgment. This approach is not only inefficient but also struggles to fully cover all fishing operations. Furthermore, it struggles to cope with the diverse, large, and complex nature of fish catches. Therefore, leveraging modern technologies, such as image processing, machine learning, and big data analysis, to achieve automated and intelligent monitoring and anomaly detection of distant-water fish catches has become a research hotspot and a pressing need in distant-water fishery management. While traditional object detection methods played a crucial role in the early days of computer vision, they suffer from significant limitations, including high computational cost, reliance on handcrafted, non-robust features, limited generalization, poor real-time performance, complex multi-stage processing, sensitivity to image quality, poor scalability, complex post-processing, strong dataset dependence, and a lack of ability to capture depth and semantic information. These shortcomings limit their effectiveness in modern applications, particularly those requiring fast, accurate, and automated processing. Consequently, with the development of YOLO and multimodal detection techniques, these traditional methods are no longer able to meet the needs of distant-water fisheries and are being gradually replaced by more efficient, accurate, and automated object detection technologies. Summary of the Invention
[0004] In response to the deficiencies in the prior art, the present invention aims to provide a device and method for detecting abnormal features of deep-sea fish catches. The present invention uses vision and multiple sensors to achieve multimodal detection, quickly and accurately detecting abnormal features in the catches, greatly improving efficiency.
[0005] In order to solve the above problems, the technical solution of the present invention is:
[0006] A device for detecting abnormal characteristics of deep-sea fish catches includes a conveying channel, an electronic tongue sensor, a first camera, a transparent box, a second camera, a classification channel, and a gas sensor device. The electronic tongue sensor is embedded in the conveying channel at the connection with the transparent box. The first camera and the second camera constitute an image acquisition module. An improved YOLOv8 model is constructed in the image acquisition module. The gas sensor device is embedded in the inner wall of the transparent box and is used to detect gas data of the catch. The catch image data and sensor data are multimodally fused and input into the constructed improved YOLOv8 model to perform multimodal abnormal characteristic detection. Catch with abnormal characteristics is then screened out through the classification channel.
[0007] Preferably, the delivery channel is composed of three sections of pipes, the first part is a flat channel, the second part is a slightly sloping channel, and the third part is a flat channel. The purpose of the slightly sloping second section of the channel is to enable the catch to utilize the high drop to more easily retain the body fluids required for detection by the electronic tongue sensor.
[0008] Preferably, the electronic tongue sensor consists of two slightly protruding soft structures and a recessed body fluid storage detector. The two slightly protruding soft structures are used to collect body fluids on the surface of the catch without causing damage to the surface of the catch. The body fluid storage detector is used to collect body fluids from the catch and test the collected body fluids in a timely manner.
[0009] Preferably, the gas sensor device includes a data transmission module and a gas sensor module. The data transmission module is used to transmit the data detected by the gas sensor for uploading for multimodal detection; the gas sensor module is composed of four gas sensors, which are used to detect the unique gas emitted by abnormal characteristics of the catch. The gas sensor is detachable and can be replaced according to the unique gas of different catches.
[0010] Preferably, the first camera is located below the transparent box and is equipped with a high-definition lens. The second camera is located above the transparent box and is equipped with a scanner. The scanner consists of three rows of near-infrared cameras in parallel, which are used to detect the freshness of the catches in the three conveying channels. The first and second cameras are used to simultaneously shoot both sides of the catches in the transparent box to further improve the accuracy of the detection results.
[0011] Preferably, the improved YOLOv8 model includes a backbone feature extraction network Backbone, a neck network Neck, and an output end Head. The first two Conv modules of the backbone feature extraction network Backbone of the improved YOLOv8 model are replaced with a new C2f-EIEM module; the detail enhancement convolution in DEA-Net is used to improve the C2f module to a New-C2f module, and the second Conv module is improved to a Deconv module; the SPPF module in the backbone network is improved to an FPSC module.
[0012] Preferably, the new C2f-EIEM module includes: a Conv module, which is then divided into two paths, one path is a SobelConv module, and the other path is a MaxPool module. The two paths are then spliced into the Concat module, and then the input image is convolved through two Conv modules; the Deconv module contains five parallel deployed convolution layers, including Vanilla convolution, center difference convolution, angle difference convolution, horizontal difference convolution and vertical difference convolution; the FPSC module sequentially includes the initial convolution layer cv1, the shared convolution layer ShareConv, the feature splicing concat final convolution layer cv2
[0013] Furthermore, the present invention also provides a method for detecting abnormal characteristics of deep-sea fish catches, comprising the following steps:
[0014] The deep-sea fish catch is conveyed through a conveying channel into a device for detecting abnormal characteristics of the catch;
[0015] The caught fish contacts the electronic tongue sensor in the conveying channel and the detection data is collected;
[0016] The catch enters a transparent box, undergoes near-infrared scanning, and collects images and gas data;
[0017] Build an improved YOLOv8 model;
[0018] All sensor data are spliced together to form multimodal data, and trained using the improved YOLOv8 model;
[0019] The trained improved YOLOv8 model is used to detect abnormal features of deep-sea fish catches and screen out fish with abnormal features.
[0020] Preferably, the step of constructing an improved YOLOv8 model specifically includes: the improved YOLOv8 model includes a backbone feature extraction network Backbone, a neck network Neck, and an output end Head, and the first two Conv modules of the backbone feature extraction network Backbone are replaced with a new C2f-EIEM module; the second Conv module in the C2f module is improved using the detail enhancement convolution in DEA-Net, and the second Conv module is improved to a Deconv module; the SPPF module in the backbone network is improved to an FPSC module.
[0021] Preferably, the step of splicing all sensor data to form multimodal data and training with the improved YOLOv8 model specifically includes: processing the image features and the electronic tongue sensor, gas sensor, and scanner data of the same time period into a tensor format, linearly mapping the extracted image features to a high-dimensional space, and mapping the obtained sensor data to the same high-dimensional space at the same time, then connecting these high-dimensional features and passing them through another linear layer to obtain the final required high-dimensional fused multimodal data.
[0022] Compared to existing technologies, this invention uses detection equipment to collect image data of fish caught and data collected by sensor modules, then combines these data to train an improved YOLOv8 model. After training, the image and sensor data are input into the trained YOLOv8 model to determine the categories and locations of abnormal features in the fish caught, and to screen out pelagic fish with these features. This invention utilizes vision and multiple sensors to achieve multimodal detection, enabling rapid and accurate detection of abnormal features in fish caught, significantly improving efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments with reference to the following drawings:
[0024] Figure 1 This is a schematic structural diagram of a device for detecting abnormal characteristics of deep-sea fish catches according to the present invention;
[0025] Figure 2 This is a flowchart of the method for detecting abnormal characteristics of deep-sea fish catches according to the present invention;
[0026] Figure 3 This is a structural diagram of the C2f-EIEM module of the present invention;
[0027] Figure 4 This is a structural diagram of the New-C2f module of the present invention;
[0028] Figure 5is a structural diagram of the FPSC module of the present invention;
[0029] Figure 6 This is a structural diagram of the multimodal data fusion of the present invention. DETAILED DESCRIPTION
[0030] The present invention will be described in detail below with reference to specific embodiments. The following examples will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those skilled in the art, several changes and improvements can be made without departing from the scope of the present invention. These all fall within the scope of protection of the present invention.
[0031] Specifically, the present invention provides a device for detecting abnormal characteristics of deep-sea fish catches, such as Figure 1 As shown, the device includes: a conveying channel 1, an electronic tongue sensor 2, a rotating shaft 3, a first camera 4, a high-definition lens 5, a transparent box 6, a scanner 7, a second camera 8, a classification channel 9, a base 10, and a gas sensor device. The gas sensor device includes a data transmission module 11 and a gas sensor module 12. The data transmission module 11 is used to transmit data detected by the gas sensor for uploading for multimodal detection. The gas sensor module 12 is composed of four gas sensors for detecting unique gases emitted by abnormal characteristics of the catch. The gas sensors are detachable and can be flexibly replaced to meet the unique gases of different catches.
[0032] Among them, the conveying channel 1 is composed of three sections of pipes spliced together, and the deep-sea catch enters the equipment environment for detecting abnormal characteristics of the catch through three catch conveying channels 1. Each conveying pipeline is first composed of a flat channel, followed by a slightly sloping channel, and finally a flat channel. The purpose of the slightly sloping channel is to enable the catch to take advantage of the high drop to more easily retain the body fluids required for electronic tongue detection, so as to meet the needs of electronic tongue detection of abnormal characteristics of the catch.
[0033] The electronic tongue sensor 2 is embedded in the connection between the conveyor channel 1 and the transparent housing 6. The second camera 8 is equipped with a scanner 7, and the first camera 4 is equipped with a high-definition lens 5. The first and second cameras 4 and 8 are supported by a base 10 and connected to the rotating shaft 3. The gas sensor device is embedded in the upper inner wall of the transparent housing 6 in an inverted position to detect gas data from the fish. All detected data is subjected to multimodal fusion and processing, and pelagic fish exhibiting abnormal characteristics are screened out of the pipeline through a sorting channel 9.
[0034] The obtained catch image data and the remaining sensor data are spliced and input into the trained YOLOv8 improved model to perform multimodal abnormal feature detection, and the catch with abnormal features is screened out through the classification channel 9 of the conveyor belt.
[0035] The first camera 4 and the second camera 8 constitute an image acquisition module, which is equipped with an improved YOLOv8 model. The improved YOLOv8 model includes a backbone feature extraction network (Backbone), a neck network (Neck), and an output terminal (Head). The present invention replaces the first two Conv modules of the backbone feature extraction network of the improved YOLOv8 model with a new C2f-EIEM module; uses the detail enhancement convolution in DEA-Net to improve the C2f module to a New-C2f module; improves the second Conv module to a Deconv module; and improves the SPPF module in the backbone network to an FPSC module.
[0036] The new C2f-EIEM module includes: a Conv module, which is then divided into two paths, one is a SobelConv module and the other is a MaxPool module. The two paths are then spliced into a Concat module, and then the input image is convolved through two Conv modules.
[0037] The New-C2f module improves the second Conv module into a Deconv module. The Deconv module contains five parallel deployed convolutional layers, including: Vanilla convolution, center difference convolution, angle difference convolution, horizontal difference convolution and vertical difference convolution.
[0038] The FPSC (feature pyramid shared convolution) module sequentially includes an initial convolution layer cv1, a shared convolution layer ShareConv, a feature concatenation concat final convolution layer cv2.
[0039] The image data of the fish to be inspected is fed into the trained improved YOLOv8 model. Combined with the sensor data analysis results, multimodal anomaly feature detection is performed. This involves fusing the image features with data from the electronic tongue sensor, gas sensor, and scanner collected over the same time period to generate data. The extracted image features are then linearly mapped into a high-dimensional space. The remaining three sensor data are then simultaneously mapped into the same high-dimensional space. These high-dimensional features are then concatenated and passed through another linear layer to obtain the final high-dimensional fused multimodal data. Finally, this fused multimodal feature is fed into the improved YOLOv8 model.
[0040] Further, if Figure 2As shown, the present invention also provides a method for detecting abnormal characteristics of deep-sea fish catches, comprising the following steps:
[0041] S1: The ocean-going fish are transported through a conveying channel into a device for detecting abnormal characteristics of the fish;
[0042] Specifically, each transport channel is composed of three parts. The first part is a flat channel, the second part is a slightly sloping channel, and the third part is a flat channel. The purpose of the second part of the channel with a slight slope is to enable the catch to utilize the high drop to more easily retain the body fluids required for electronic tongue detection, so as to meet the needs of the electronic tongue to detect abnormal characteristics of the catch.
[0043] S2: The fish catch contacts the electronic tongue sensor in the conveying channel and the detection data is collected;
[0044] Specifically, the electronic tongue sensor is located in the second section of the delivery channel, with one sensor installed in each of the three delivery channels. The sensor consists of two slightly raised flexible structures and a recessed fluid storage detector. The two slightly raised flexible structures are used to collect fluids from the surface of the fish without damaging the surface, while the fluid storage detector is used to collect and promptly test the collected fluids. The electronic tongue sensor of this invention can quickly, effectively, and non-destructively detect surface fluid data from fish. It can also measure the content and activity of enzymes, proteins, and their derivatives in the fish fluids, thereby determining the freshness of the fish.
[0045] S3: The fish enters the transparent box, undergoes near-infrared scanning, and collects images and gas data;
[0046] Specifically, the transparent box is located at the third section of the conveyor channel. The inverted gas sensor device on the upper inner wall of the box consists of a data transmission module and a gas sensor module. A near-infrared scanner, consisting of three parallel rows of near-infrared cameras, is mounted on the camera above the transparent box. This scanner is used to rapidly monitor the freshness of the catch from the three conveyor channels. The first and second cameras, located above and below the transparent box, capture images of both sides of the catch.
[0047] S4: Build an improved YOLOv8 model;
[0048] Specifically, the improved YOLOv8 model includes a backbone feature extraction network (Backbone), a neck network (Neck), and an output terminal (Head). The first two Conv modules in the backbone feature extraction network of the model are replaced with the new C2f-EIEM module. The second Conv module in the C2f module is improved using the detail enhancement convolution in DEA-Net, and the Conv module is improved to a Deconv module. The SPPF module in the backbone network is improved to an FPSC module.
[0049] like Figure 3 As shown in the figure, the new C2f-EIEM module includes: a Conv module, which is then divided into two paths, one is a SobelConv module, and the other is a MaxPool module. The two paths are then spliced into the Concat module, and then the input image is convolved through two Conv modules.
[0050] The following combination Figure 3 The above improvements are described in detail:
[0051] The C2f-EIEM module is an efficient image recognition front-end. By combining edge detection and spatial feature extraction, it learns richer and more comprehensive image features. Through this innovative fusion strategy, the module can more accurately understand and identify image content.
[0052] Edge Feature Capture: Although traditional convolutional neural networks excel at processing spatial information, they can be insensitive when it comes to identifying edge features in images. The C2f-EIEM module addresses this issue through the SobelConv branch. SobelConv utilizes the Sobel filter, a widely used edge detection technique that accurately identifies locations with brightness changes in an image.
[0053] Spatial feature preservation: The spatial features of an image are also crucial for understanding its overall structure. The C2f-EIEM module extracts these features through an independent convolutional branch. Unlike the SobelConv branch that focuses on edges, the convolutional branch focuses on extracting spatial details from the original image.
[0054] Feature Fusion: The C2f-EIEM module integrates the features extracted by the SobelConv branch and the convolution branch. This integration process, typically accomplished through a concatenation operation, ensures that the final feature representation incorporates both edge and spatial information, providing a richer and more comprehensive perspective for image recognition.
[0055] like Figure 4As shown in Figure 1, the detail enhancement convolution in DEA-Net is used to improve the C2f module to New-C2f, and the second Conv module is improved to the Deconv module.
[0056] The Deconv module contains five convolutional layers deployed in parallel, including: Vanilla convolution, center difference convolution, angle difference convolution, horizontal difference convolution and vertical difference convolution.
[0057] The following combination Figure 4 The above improvements are described in detail:
[0058] The Deconv module deploys these convolutional layers in parallel for feature extraction and integrates traditional local descriptors into the convolutional layers through center differential convolution, angular differential convolution, horizontal differential convolution, and vertical differential convolution. By parallelizing ordinary and differential convolutions, DEConv enhances the feature representation capability without adding additional parameters and computational cost.
[0059] like Figure 5 As shown in the figure, the SPPF module in the backbone feature network is improved into the FPSC (feature pyramid shared convolution) module, which sequentially includes the initial convolution layer cv1, the shared convolution layer ShareConv (using the same convolution kernel to perform multiple convolutions with different dilation parameters on the initial feature map), the feature concatenation concat final convolution layer cv2.
[0060] The following combination Figure 5 The above improvements are described in detail:
[0061] By using convolutional layers with different dilation rates, the FPSC module is able to extract features of different scales, which is very beneficial for capturing information of different sizes and contexts in the image. Low dilation rates capture local details, while high dilation rates capture global context.
[0062] The use of shared convolutional layers greatly reduces the number of parameters that need to be trained. Compared with using independent convolutional layers for each expansion rate, shared convolutional layers can reduce redundancy, improve model efficiency, reduce the storage and computational overhead of the model, and improve computational efficiency.
[0063] Through the 1x1 convolutional layers cv1 and cv2, the module can efficiently adjust the number of channels and perform feature fusion. The 1x1 convolutional layer can retain important feature information while reducing the number of parameters.
[0064] FPSC uses convolution operations for feature extraction, which can capture more fine-grained features. In contrast, the pooling operation of SPPF may lose some detailed information.
[0065] S5: All sensor data are spliced together to form multimodal data, and trained using the improved YOLOv8 model;
[0066] Specifically, if Figure 6 As shown, the multimodal abnormal feature detection involves first processing image features and data from the electronic tongue sensor, gas sensor, and near-infrared scanner over the same time period into a tensor format. The extracted image features are then linearly mapped into a high-dimensional space. The remaining three sensor data are then simultaneously mapped into the same high-dimensional space. These high-dimensional features are then concatenated and passed through another linear layer to obtain the final high-dimensional fused multimodal data.
[0067] S6: Use the trained improved YOLOv8 model to detect abnormal features of deep-sea fish catches and filter out fish with abnormal features.
[0068] The above describes specific embodiments of the present invention. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art may make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. The embodiments of this application and the features in the embodiments may be combined with each other in any manner unless there is a conflict.
Claims
1. A device for detecting abnormal characteristics of deep-sea fish catches, characterized in that: The device includes a conveying channel, an electronic tongue sensor, a first camera, a transparent box, a second camera, a classification channel, and a gas sensor device. The electronic tongue sensor is embedded in the connection between the conveying channel and the transparent box. The first camera and the second camera constitute an image acquisition module. An improved YOLOv8 model is constructed in the image acquisition module. The gas sensor device is embedded in the inner wall of the transparent box and is used to detect gas data of the catch. The catch image data and sensor data are multimodally fused and input into the constructed improved YOLOv8 model to perform multimodal abnormal feature detection, and the catch with abnormal features is screened out through the classification channel.
2. The device for detecting abnormal characteristics of deep-sea fish catches according to claim 1, characterized in that: The delivery channel is composed of three sections of pipes. The first section is a flat channel, the second section is a slightly sloping channel, and the third section is a flat channel. The purpose of the slightly sloping second section of the channel is to enable the catch to utilize the high drop to more easily retain the body fluids required for detection by the electronic tongue sensor.
3. The device for detecting abnormal characteristics of deep-sea fish catches according to claim 1, characterized in that: The electronic tongue sensor consists of two slightly protruding soft structures and a recessed body fluid storage detector. The two slightly protruding soft structures are used to collect body fluids on the surface of the catch without causing damage to the surface of the catch. The body fluid storage detector is used to collect body fluids from the catch and test the collected body fluids in a timely manner.
4. The device for detecting abnormal characteristics of deep-sea fish catches according to claim 1, characterized in that: The gas sensor device includes a data transmission module and a gas sensor module. The data transmission module is used to transmit the data detected by the gas sensor for uploading for multimodal detection; the gas sensor module is composed of four gas sensors, which are used to detect the unique gas emitted by abnormal characteristics of the catch. The gas sensor is detachable and can be replaced according to the unique gas of different catches.
5. The device for detecting abnormal characteristics of deep-sea fish catches according to claim 1, characterized in that: The first camera is located below the transparent box and is equipped with a high-definition lens. The second camera is located above the transparent box and is equipped with a scanner. The scanner consists of three rows of near-infrared cameras in parallel, which are used to detect the freshness of the catch in the three conveying channels. The first and second cameras simultaneously shoot both sides of the catch in the transparent box to further improve the accuracy of the detection results.
6. The device for detecting abnormal characteristics of deep-sea fish catches according to claim 1, characterized in that: The improved YOLOv8 model includes a backbone feature extraction network Backbone, a neck network Neck, and an output end Head. The first two Conv modules of the backbone feature extraction network Backbone of the improved YOLOv8 model are replaced with a new C2f-EIEM module; the C2f module is improved to a New-C2f module using the detail enhancement convolution in DEA-Net, and the second Conv module is improved to a Deconv module; the SPPF module in the backbone network is improved to an FPSC module.
7. The device for detecting abnormal characteristics of deep-sea fish catches according to claim 6, characterized in that: The new C2f-EIEM module includes: a Conv module, which is then divided into two paths, one for the SobelConv module and the other for the MaxPool module. The two paths are then spliced into the Concat module, and then convolve the input image through two Conv modules; the Deconv module contains five parallel deployed convolutional layers, including Vanilla convolution, center difference convolution, angle difference convolution, horizontal difference convolution and vertical difference convolution; the FPSC module sequentially includes the initial convolution layer cv1, the shared convolution layer ShareConv, the feature splicing concat final convolution layer cv2.
8. A method for detecting abnormal characteristics of deep-sea fish catches, characterized in that: The method comprises the following steps: The deep-sea fish catch is conveyed through a conveying channel into a device for detecting abnormal characteristics of the catch; The caught fish contacts the electronic tongue sensor in the conveying channel and the detection data is collected; The catch enters a transparent box, undergoes near-infrared scanning, and collects images and gas data; Build an improved YOLOv8 model; All sensor data are spliced together to form multimodal data, and trained using the improved YOLOv8 model; The trained improved YOLOv8 model is used to detect abnormal features of deep-sea fish catches and screen out fish with abnormal features.
9. The method for detecting abnormal characteristics of deep-sea fish catches according to claim 8, characterized in that: The step of constructing the improved YOLOv8 model specifically includes: the improved YOLOv8 model includes a backbone feature extraction network Backbone, a neck network Neck, and an output end Head, replacing the first two Conv modules of the backbone feature extraction network Backbone with a new C2f-EIEM module; using the detail enhancement convolution in DEA-Net to improve the second Conv module in the C2f module, and improving the second Conv module to a Deconv module; and improving the SPPF module in the backbone network to an FPSC module.
10. The method for detecting abnormal characteristics of deep-sea fish catches according to claim 8, characterized in that: The steps of splicing all sensor data to form multimodal data and training using the improved YOLOv8 model specifically include: processing image features and electronic tongue sensor, gas sensor, and scanner data in the same time period into a tensor format, linearly mapping the extracted image features to a high-dimensional space, and mapping the obtained sensor data to the same high-dimensional space. These high-dimensional features are then connected and passed through another linear layer to obtain the final required high-dimensional fused multimodal data.
Citation Information
Cited By
Fish multi-part freshness detection method and device based on improved YOLO model
CN121616584A