Improved fishery illegal fishing behavior distinguishing system, method and equipment based on YOLOv10 algorithm and medium
By introducing the HWD-ADown module and the AM attention module into the image detection model of the YOLOv10 algorithm, the problem of small target missed detection in the water background is solved, and higher detection accuracy and recognition effect are achieved.
Patent Information
- Application Number
- CN202510141041.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-02-08
AI Technical Summary
When applying the YOLOv10 algorithm in water background, it is prone to missed detection problems of small targets, which are poor in recognition and difficult to effectively identify illegal fishing behaviors in waters.
The image detection model of YOLOv10 algorithm is improved, and the downsampling module in the original Backbone network module is replaced by the introduction of the HWD-ADown module, and the AM attention module is introduced on the basis of the C2f module to enhance the feature extraction and attention mechanism to improve the detection accuracy of the model.
Through the improved model, ships and personnel of different scales can be more effectively identified, small target missed inspections can be reduced, and the detection accuracy of illegal fishing behaviors in waters can be improved.
Smart Images

Figure CN120217036A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a system, method, device and medium for discriminating illegal fishing behaviors of fishery administration based on the improved YOLOv10 algorithm. Background Art
[0002] In recent years, the fishery resources in many important waters such as the main stream of the Yangtze River, Poyang Lake, and Dongting Lake in the middle and lower reaches of the Yangtze River have shown an obvious decline. After multiple generations of reproduction, artificially cultured fish will inevitably experience genetic diversity degradation. Therefore, protecting wild fish resources is crucial for the future of our aquaculture industry. Currently, there are illegal fishing behaviors in various basins of the Yangtze River. Cracking down on such illegal fishing behaviors is the core of fish resource protection. Currently, the inspection work on various basins of the Yangtze River mainly relies on the manual inspection method of law enforcement officers. Law enforcement officers are regularly arranged to inspect important waters, and the inspection results are summarized and recorded in a book. Considering from aspects such as manpower and material resources, this method has various drawbacks: First, manual inspection cannot cover 24 hours; Second, illegal personnel will take chances and continue to commit crimes. Therefore, by installing Internet of Things devices such as cameras and speakers on the banks of key waters and analyzing the targets and behaviors in the 24-hour surveillance videos to determine whether it is illegal fishing behavior, it is a relatively promising fishery administration law enforcement method.
[0003] With the development of image processing technology, deep learning has made good progress in the field of object detection. Currently, classic object detection algorithms are mainly divided into two categories: single-stage and two-stage. Single-stage ones include YOLO, SDD, Retina-Net, etc., and two-stage ones include R-CNN, Fast R-CNN, Mask R-CNN, etc.; among them, the YOLO algorithm has the fastest running speed and relatively high accuracy. Especially for the latest version of the YOLOv10 model, it has made significant breakthroughs in speed and accuracy. However, YOLOv10 is prone to missing small targets when applied in water area backgrounds, and the recognition effect is not very good, and further research is still needed. Summary of the Invention
[0004] The purpose of the present invention is to provide a system, method, device and medium for discriminating illegal fishing behaviors of fishery administration based on the improved YOLOv10 algorithm for the above problems.
[0005] To achieve the above purpose, the technical solution provided by the present invention is:
[0006] The first aspect of the present application provides a system for discriminating illegal fishing behaviors of fishery administration based on the improved YOLOv10 algorithm, including:
[0007] A water area real-time image acquisition module, configured to acquire real-time water area images through a pan-tilt monitoring camera;
[0008] An image detection model improved based on the YOLOv10 algorithm, including a Backbone network module for encoding the input image in the form of features, a Neck network module for processing the encoded features in the form of features, and a Head network module for generating the predicted output of the model;
[0009] The Backbone network module includes an HWD-ADown module for decomposing the feature map and a C2f module for feature extraction and feature fusion. The output end of the HWD-ADown module is connected to the C2f module;
[0010] The HWD-ADown module includes an HWD module and an ADown module; the HWD module is used to obtain the feature map decomposition processing result by performing Haar wavelet transform on the input image; the ADown module is used to perform downsampling operation on the feature map decomposition processing result of the HWD module;
[0011] Based on the C2f module connected to the output end of the HWD-ADown module, an AM attention module is introduced. The AM attention module includes a context attention module CxAM and a content attention module CnAM. The context attention module CxAM is used to focus on the specified local area or block on the feature map, and the content attention module CnAM is used to focus on the features at relatively shallower levels on the feature map; the output features of the context attention module CxAM and the content attention module CnAM are fused to obtain the final feature expression of the feature map;
[0012] A detection and discrimination module is used to input the real-time image of the water area into the image detection model improved based on the YOLOv10 algorithm for recognition to obtain the detection result. If there is an illegal fishing behavior in the detection result discrimination information of the discrimination model, the alarm function is triggered.
[0013] To optimize the above technical solution, the specific measures taken also include:
[0014] The HWD module includes a lossless feature encoding block and a feature representation learning block. The lossless feature encoding block includes a Haar wavelet transform layer, and the Haar wavelet transform layer is used to perform Haar wavelet transform on the input image to obtain the image decomposition information; the feature representation learning block is used to extract discriminative features from the image decomposition information obtained by the Haar wavelet transform layer and output the feature map decomposition processing result.
[0015] In the ADown module, the input feature map is divided into two branches after passing through the average pooling layer. One branch only passes through the CBS module, and the other branch passes through the max pooling layer and the CBS module in sequence. The results obtained from the two branches are fused and output.
[0016] The process of the Haar wavelet transform layer performing the Haar wavelet transform on the input image to obtain image decomposition information is specifically as follows:
[0017] After the image with a resolution of h*w is decomposed by the Haar wavelet transform, where h represents the image height and w represents the image width, four components are generated, including: a low-frequency component A, a high-frequency component H in the horizontal direction, a high-frequency component V in the vertical direction, and a high-frequency component D in the diagonal direction. The size of each component is h / 2×w / 2, the resolution is reduced to 1 / 4 of the original image resolution, and the number of channels of the feature map becomes four times larger;
[0018] The feature representation learning block includes a 1×1 convolutional layer, a normalization layer, and a ReLU activation function; the process of the feature representation learning block extracting discriminative features from the image decomposition information obtained by the Haar wavelet transform and outputting the feature map decomposition processing result is specifically as follows:
[0019] The feature representation learning block uses a 1×1 convolutional layer to adjust the number of channels of the feature map to adapt to the subsequent layer and filter redundant information, and then combines the normalization layer and the ReLU activation function to extract discriminative features to obtain the feature map decomposition processing result.
[0020] The process of the context attention module CxAM focusing on a specified local region or block on the feature map is specifically as follows:
[0021] The context attention module CxAM transforms the feature map feature mapping F∈R C×h×w respectively using the convolutional layer W q and the convolutional layer W k into the potential space, where C is the number of channels, h is the image height, and w is the image width; the convolutional layer W q is used to guide the attention to a specified local region or block in the input feature mapping, and the generated feature vectors form the first query space Q; the convolutional layer W k is used to calculate the correlation or similarity of the feature mapping of the specified local region or block in the first query space Q, and the generated feature vectors form the first key space K; the transformation formula is as follows:
[0022]
[0023] where, {Q,K}∈R C×h×w , and then the first query space Q and the first key space K are transformed into vectors. In order to obtain the relationship between local regions or blocks, the following relationship matrix is obtained:
[0024] R = Q T K
[0025] where, R∈R N×N, N = h × w, and then normalize R through the sigmoid activation function and average pooling to obtain the attention matrix R1 ∈ R C×h×w ; then use the convolutional layer W V to transform the feature map F into the feature V, where V ∈ R C ×h×w , and then perform element-wise multiplication of R1 and V to obtain the final attention representation E. The attention representation E i of the i-th feature map with the number of channels C is:
[0026] E i = R1 ⊙ V i
[0027] where V i represents the feature transformed by the convolution of the i-th feature map with the number of channels C;
[0028] The process by which the content attention module CnAM focuses on the relatively shallow-level features on the feature map is as follows:
[0029] The content attention module CnAM transforms the relatively shallow-level feature map C4 ∈ S C×h×w on the feature map into the potential space using the convolutional layer W p and the convolutional layer W z respectively. The convolutional layer W p is used to guide the attention to the relatively shallow-level features in the input feature map, and the generated feature vectors form the second query space P; the convolutional layer W k is used to calculate the correlation or similarity of the relatively shallow-level feature map in the second query space P, and the generated feature vectors form the second key space Z; the transformation formula is as follows:
[0030]
[0031] where {P, Z} ∈ S C×h×w , and then transform the second query space P and the second key space Z into vectors; in order to obtain the relationship between local regions or blocks, the following relationship matrix is obtained:
[0032] S = P T Z
[0033] where S ∈ S N×N , N = h × w, and then normalize S through the sigmoid activation function and average pooling to obtain the attention matrix S1 ∈ S C×h×w ; finally, multiply the feature V transformed from the feature map F using the convolutional layer W V element-wise with S1 and finally use the convolutional layer W V, the final attention representation D, the attention representation D of the i-th feature map with the number of channels C is obtained i is:
[0034] D i = S1 ⊙ V i .
[0035] Among them, V i represents the feature transformed by the feature mapping convolution of the i-th feature map with the number of channels C.
[0036] The second aspect of this application provides a method for discriminating illegal fishing behaviors based on the improved YOLOv10 algorithm, including the following steps:
[0037] Collect real-time images of the water area through a pan-tilt monitoring camera;
[0038] Construct an image detection model based on the improved YOLOv10 algorithm, and perform model training and evaluation;
[0039] Input the real-time water area image into the image detection model based on the improved YOLOv10 algorithm for recognition, obtain the detection result, and if there is illegal fishing behavior discrimination information in the detection result, trigger the alarm function;
[0040] The above-mentioned model training and evaluation include: obtaining an illegal fishing image dataset and dividing it into a training set and a test set; performing iterative training on the training set, retaining the best network weights; using the test set to test the model;
[0041] The above-mentioned iterative training on the training set and retaining the best network weights are specifically: input the training set into the Backbone network module through the input layer, and then obtain the feature map of the training set through the Neck network module; perform prediction in the Head network module based on the feature map of the training set, calculate the loss, and update the model parameters; retrain until the loss function converges or reaches the maximum number of iterations, and save the best network weights.
[0042] The above-mentioned illegal fishing images include images of operating a ship for illegal fishing in key waters and images of fishing with live bait on the shore; use annotation software to annotate the illegal fishing images to obtain a preliminary selected dataset, and then perform augmentation processing on each original picture in the preliminary selected dataset. The above-mentioned augmentation processing includes simulating weather environment data augmentation.
[0043] The third aspect of this application provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the method for discriminating illegal fishing behaviors based on the improved YOLOv10 algorithm as described above.
[0044] The fourth aspect of the present application provides a computer-readable storage medium storing a computer program, which causes a computer to execute the method for discriminating illegal fishing behaviors improved based on the YOLOv10 algorithm as described above.
[0045] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0046] The present invention constructs an image detection model improved based on the YOLOv10 algorithm, improves the original network model of YOLOv10, replaces the downsampling module in the original Backbone network module with the HWD-ADown module, and introduces the AM attention module on the basis of the C2f module in the Backbone network module, trains the model and evaluates the model.
[0047] Since there is a possibility of information loss in the original downsampling module, the present invention can retain image information as much as possible by using the HWD-ADown module, so that subsequent layers can extract more discriminative features, thereby improving the model performance.
[0048] The HWD-ADown module includes an HWD module and an ADown module. The HWD module includes a lossless feature encoding block and a feature representation learning block. The lossless feature encoding block decomposes the image into four components: a low-frequency component A, a high-frequency component H in the horizontal direction, a high-frequency component V in the vertical direction, and a high-frequency component D in the diagonal direction by using Haar wavelet transform, which can effectively reduce the spatial resolution of the feature map while retaining all information.
[0049] The present invention introduces the AM attention module on the basis of the C2f module to generate an attention matrix. The AM attention module can adaptively focus on and weight the specified local region or block on the feature map and the relatively shallower-level features containing more detailed and position information on the feature map, so that the network can more accurately capture the position of each object to improve the performance of the network.
[0050] The AM attention module includes two sub-modules: a context attention module CxAM and a content attention module CnAM. CxAM mainly focuses on the specified local region or block on the feature map, and CnAM focuses on the relatively shallower-level features on the feature map, thereby reducing the impact of deformable convolution on the target position. After the features output by the two sub-modules are fused, a more comprehensive feature expression can be obtained.
[0051] By adopting the illegal fishing behavior discrimination model that improves the Backbone network module of YOLOv10, the present invention can identify ships and personnel of different scales, improve the problems of redundant calculation and missed detection of small targets in the original algorithm, and improve the detection accuracy of illegal fishing behaviors in key waters. Description of the Drawings
[0052] Figure 1 : Schematic flow chart of the method for discriminating illegal fishing behaviors based on the improved YOLOv10 algorithm of the present invention.
[0053] Figure 2 : Schematic flow chart of the construction of the image detection model based on the improved YOLOv10 algorithm of the present invention.
[0054] Figure 3 : Schematic structural diagram of the improved Backbone network module of the present invention.
[0055] Figure 4 : Schematic structural diagram of the HWD module in the HWD-ADown module.
[0056] Figure 5 : Schematic structural diagram of the ADown module in the HWD-ADown module.
[0057] Figure 6 : Schematic structural diagram of the CxAM module in the AM attention module.
[0058] Figure 7 : Schematic structural diagram of the CnAM module in the AM attention module.
[0059] Figure 8 : Schematic diagram of the detection results of illegal fishing in the embodiment of the present invention. Detailed implementation manners
[0060] The above content of the present invention will be further described in detail below in the form of embodiments, but it should not be understood that the scope of the above subject matter of the present invention is limited to the following embodiments. All technologies implemented based on the above content of the present invention belong to the scope of the present invention.
[0061] Some noun explanations in the present invention are as follows:
[0062] The features on the relatively shallow levels of the feature maps involved in this article refer to the features that have gone through fewer processing levels in the neural network. They are relative concepts compared to the features that have gone through more processing levels, and their division range can be set according to specific needs during application. Among them, the features that have gone through more processing levels, due to going through more layers of processing, have a larger receptive field and can capture the global context information in the image; while the features that have gone through fewer levels of processing can retain more detailed information.
[0063] In one embodiment of the present invention, an illegal fishing behavior discrimination system based on the improved YOLOv10 algorithm is proposed, including:
[0064] A real-time water area image acquisition module for collecting real-time water area images through a pan-tilt monitoring camera;
[0065] An image detection model improved based on the YOLOv10 algorithm, including a Backbone network module for encoding the input image in the form of features, a Neck network module for processing the feature-form encoding, and a Head network module for generating the model's prediction output;
[0066] As Figure 3 shown, the Backbone network module includes an HWD-ADown module for decomposing the feature map and a C2f module for feature extraction and feature fusion. The output end of the HWD-ADown module is connected to the C2f module;
[0067] The HWD-ADown module includes an HWD module and an ADown module; the HWD module is used to obtain the feature map decomposition processing result by performing Haar wavelet transform on the input image; the ADown module is used to perform downsampling operation on the feature map decomposition processing result of the HWD module;
[0068] Based on the C2f module connected to the output end of the HWD-ADown module, an AM attention module is introduced. The AM attention module includes a context attention module CxAM and a content attention module CnAM. The context attention module CxAM is used to focus on the specified local area or block on the feature map, and the content attention module CnAM is used to focus on the features at relatively shallower levels on the feature map; the output features of the context attention module CxAM and the content attention module CnAM are fused to obtain the final feature expression of the feature map;
[0069] A detection and discrimination module for inputting the real-time water area image into the image detection model improved based on the YOLOv10 algorithm for recognition to obtain the detection result. If there is an illegal fishing behavior in the detection result discrimination information of the discrimination model, the alarm function is triggered.
[0070] The HWD module is as Figure 4 shown, including a lossless feature encoding block and a feature representation learning block. The lossless feature encoding block includes a Haar wavelet transform layer, and the Haar wavelet transform layer is used to perform Haar wavelet transform on the input image to obtain the image decomposition information; the feature representation learning block is used to extract discriminative features from the image decomposition information obtained by the Haar wavelet transform layer and output the feature map decomposition processing result.
[0071] The ADown module is as Figure 5As shown in the figure, it includes an average pooling layer, and two branches are connected to the backend of the average pooling layer; in the ADown module, the input feature map is divided into two branches after passing through the average pooling layer. One branch only passes through the CBS module, and the other branch passes through the max pooling layer and the CBS module in sequence. The results obtained from the two branches are output after feature fusion; the convolutional layers of the two branches are both composed of a Conv module, a BN module, and a SiLU module.
[0072] The process of the Haar wavelet transform layer performing the Haar wavelet transform on the input image to obtain image decomposition information is specifically as follows:
[0073] After the image with a resolution of h*w is decomposed by the Haar wavelet transform, where h represents the image height and w represents the image width, four components are generated, including: a low-frequency component A, a high-frequency component H in the horizontal direction, a high-frequency component V in the vertical direction, and a high-frequency component D in the diagonal direction. The size of each component is h / 2×w / 2, the resolution is reduced to 1 / 4 of the original image resolution, and the number of channels of the feature map becomes four times larger. While retaining all information, it can effectively reduce the spatial resolution of the feature map.
[0074] The feature representation learning block includes a 1×1 convolutional layer, a normalization layer, and a ReLU activation function; the process of the feature representation learning block extracting discriminative features from the image decomposition information obtained by the Haar wavelet transform layer and outputting the feature map decomposition processing result is specifically as follows:
[0075] The feature representation learning block uses a 1×1 convolutional layer to adjust the number of channels of the feature map to adapt to the subsequent layer and filter redundant information, and then combines the normalization layer and the ReLU activation function to extract discriminative features to obtain the feature map decomposition processing result.
[0076] As Figure 6 shown in the figure, the process of the context attention module CxAM focusing on a specified local area or block on the feature map is specifically as follows:
[0077] The context attention module CxAM transforms the feature map F∈R C×h×w into the potential space using convolutional layers W q and convolutional layer W k respectively. C is the number of channels, h is the image height, and w is the image width; convolutional layer W q is used to guide the attention to a specified local area or block in the input feature map, and the generated feature vectors form the first query space Q; convolutional layer W k is used to calculate the correlation or similarity of the feature map of the specified local area or block in the first query space Q, and the generated feature vectors form the first key space K; the transformation formula is as follows:
[0078]
[0079] where {Q, K} ∈ ℝ C×h×w , then the first query space Q and the first key space K are transformed into vectors. To obtain the relationships between local regions or blocks, the following relationship matrix is obtained:
[0080] R = Q T K
[0081] where R ∈ ℝ N×N , N = h × w, and then R is normalized through the sigmoid activation function and average pooling to obtain the attention matrix R1 ∈ ℝ C×h×w ; then the convolutional layer W V is used to transform the feature map F into the feature V, where V ∈ ℝ C ×h×w , and then R1 and V are element-wise multiplied to obtain the final attention representation E. The attention representation E i of the i-th feature map with the number of channels C is:
[0082] E i = R1 ⊙ V i
[0083] where V i represents the feature transformed by the convolution of the i-th feature map with the number of channels C.
[0084] The process by which the content attention module CnAM focuses on the relatively shallow-level features on the feature map is specifically as follows:
[0085] The content attention module CnAM transforms the relatively shallow-level feature map C4 ∈ 𝒮 C×h×w on the feature map into the potential space using the convolutional layer W p and the convolutional layer W z respectively. The convolutional layer W p is used to guide the focus on the relatively shallow-level features in the input feature map, and the generated feature vectors form the second query space P; the convolutional layer W k is used to calculate the correlation or similarity of the relatively shallow-level feature map in the second query space P, and the generated feature vectors form the second key space Z; the transformation formula is as follows:
[0086]
[0087] where {P, Z} ∈ 𝒮 C×h×w , and then the second query space P and the second key space Z are transformed into vectors; to obtain the relationships between local regions or blocks, the following relationship matrix is obtained:
[0088] S = P T Z
[0089] Among them, S ∈ S N×N , N = h × w, and then S is normalized through the sigmoid activation function and average pooling to obtain the attention matrix S1 ∈ S C×h×w ; Finally, the feature map F is transformed into the feature V by the convolutional layer W V , and the element multiplication of V and S1 is performed, and finally the convolutional layer W V is used to obtain the final attention representation D. The attention representation D of the i-th feature map with the number of channels C is i :
[0090] D i = S1 ⊙ V i .
[0091] Among them, V i represents the feature transformed by the feature map convolution of the i-th feature map with the number of channels C
[0092] The present invention improves the original YOLOv10 network model. As Figure 2 shown, it includes: replacing the downsampling module (SCDown module) in the original YOLOv10 network model structure with the HWD-ADown module, and introducing the AM attention module on the basis of the C2f module after the HWD-ADown module to construct an improved image detection model based on the YOLOv10 algorithm
[0093] Due to the possible loss of information in the original downsampling module, the present invention can retain image information as much as possible by using the HWD-ADown module, so that subsequent layers can extract more discriminative features, thereby improving the model performance
[0094] The present invention introduces the AM attention module on the basis of the C2f module to adaptively focus on and weight the specified local region or block on the feature map and the relatively shallow-level features containing more detailed and positional information on the feature map, so as to improve the performance of the network
[0095] In some preferred embodiments, the present invention uses the HWD-ADown module to replace the last downsampling module in the original YOLOv10 network model structure, and introduces the AM attention module in the C2f module connected to the backend of the HWD-ADown module
[0096] In another embodiment, the present invention proposes a method for discriminating illegal fishing behaviors based on the improved YOLOv10 algorithm. As Figure 1 shown, it includes the following steps
[0097] Collect real-time images of the water area through the pan-tilt monitoring camera
[0098] Construct an improved image detection model based on the YOLOv10 algorithm, and perform model training and evaluation;
[0099] Input the real-time water area image into the improved image detection model based on the YOLOv10 algorithm for recognition to obtain the detection result. If there is illegal fishing behavior discrimination information in the detection result, trigger the alarm function;
[0100] Among them, the improved image detection model based on the YOLOv10 algorithm is the improved image detection model based on the YOLOv10 algorithm in the above-mentioned embodiment.
[0101] Performing model training and evaluation includes: obtaining an illegal fishing image dataset, dividing it into a training set and a test set, performing iterative training on the training set, retaining the best network weights, and using the test set to test the model.
[0102] Performing iterative training on the training set and retaining the best network weights specifically means: inputting the training set into the Backbone network module through the input layer, and then obtaining the feature map of the training set through the Neck network module; performing prediction in the Head network module based on the feature map of the training set, calculating the loss, and updating the model parameters; retraining until the loss function converges or reaches the maximum number of iterations, and saving the best network weights.
[0103] Illegal fishing images include images of operating illegal fishing in key water areas by driving a ship and images of fishing with live bait on the shore; use the LabelImg software to annotate the illegal fishing images to obtain a preliminary selected dataset, and then perform augmentation processing on each original image in the preliminary selected dataset. The augmentation processing includes simulating weather environment data augmentation.
[0104] In another embodiment of the present invention, an electronic device is proposed, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the above-mentioned method for discriminating illegal fishing behavior based on the improved YOLOv10 algorithm.
[0105] In another embodiment of the present invention, a computer-readable storage medium is proposed, storing a computer program, and the computer program causes the computer to execute the above-mentioned method for discriminating illegal fishing behavior based on the improved YOLOv10 algorithm.
[0106] In the embodiments disclosed in the present application, a computer storage medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The computer storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the computer storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0107] The technical solutions of the present invention will be further described in detail below in conjunction with specific embodiments:
[0108] Monitoring and discriminating illegal fishing behavior in a certain water area includes constructing an improved image detection model based on the YOLOv10 algorithm, and performing model training and evaluation. The specific implementation process includes the following steps:
[0109] Step 1: Use illegal fishing images with a size of 960 pixels × 960 pixels shared by the network platform and provided by government departments. The illegal fishing images include images of operating a ship for illegal fishing in key waters and images of fishing with live bait on the shore. According to the standards of the PASCAL VOC 2007 dataset, use the LabelImg software to annotate the illegal fishing images to obtain a preliminary dataset, and then perform two amplification processes on each original image in the preliminary dataset; the first data amplification process is to amplify the original image by simulating the weather environment, and the second data amplification process is to use one to three traditional data amplification methods to amplify the data again; the dataset is divided into a training set and a test set according to a ratio of 3:2.
[0110] Step 2: Construct an HWD-ADown module and replace the last SCDown module in the original Backbone network module with it.
[0111] The HWD-ADown module is obtained by replacing the Conv module in the ADown module with the HWD module (Haar Wavelet Downsampling) structure. The HWD structure refers to Figure 4 the schematic diagram of the HWD module structure shown, including a lossless feature encoding block and a feature representation learning block; the ADown module refers to Figure 5The schematic diagram of the A Down module structure shown is divided into two branches after passing through an average pooling layer once. One branch passes through a convolutional layer, and the other branch passes through a max pooling layer and a convolutional layer once. The convolutional layer consists of a Conv module, a BN module, and a SiLU module.
[0112] The lossless feature encoding block utilizes a Haar wavelet transform layer, as Figure 4 shown. After an image with a resolution of h * w is decomposed by the Haar wavelet transform, where h represents the image height and w represents the image width, four components are generated, including: a low-frequency component A, a high-frequency component H in the horizontal direction, a high-frequency component V in the vertical direction, and a high-frequency component D in the diagonal direction. The size of each component is h / 2 × w / 2, which can effectively reduce the spatial resolution of the feature map while retaining all information.
[0113] The feature representation learning block includes a standard convolutional layer, a normalization layer, and a ReLU activation function.
[0114] Step 3: Introduce an AM attention module into the C2f module at the output end of the HWD - A Down module.
[0115] The AM attention module includes a context attention module CxAM and a content attention module CnAM. CxAM mainly focuses on specified local regions or blocks on the feature map, and CnAM focuses on features at relatively shallower levels of the feature map, thereby reducing the impact of deformable convolution on the target position. After the features output by the two sub-modules are fused, a more comprehensive feature representation can be obtained.
[0116] Step 4: Input the training set into the Backbone network through the input layer, and then obtain the feature map of the training set through the Neck network; make predictions in the Head network module based on the feature map of the training set, calculate the loss, and update the model parameters; retrain until the loss function converges or reaches the maximum number of iterations, and save the best network weights.
[0117] In this embodiment, the model is built using the Tensorflow 2.2 framework, and the graphics card is NVIDIA GeForce RTX 3060. The environment is configured with CUDA 11.8, cuDNN 7.6.5, Python 3.7, Pytorch 1.10.1, and opencv 4.4. The initial learning rate is 0.01, the momentum factor is 0.9, and its decay method is a fixed-step learning rate decay strategy to ensure that the learning rate can be appropriately adjusted after each training cycle. The specific adjustment multiple is 0.3.
[0118] Step 5: Evaluate the model, and evaluate the detection accuracy of the model according to the improved illegal fishing discrimination model based on YOLOv10 after training;
[0119] Select precision P, recall R, F1 score, mean average precision mAP, and processing time as the verification metrics for the improved YOLOv10 model. The calculation formulas are as follows:
[0120]
[0121] Where: TP is the number of defects correctly detected; FP is the number of misdetected defects; FN is the number of undetected defects; n is the number of classes; AP is the enclosed area formed by the P value and the R value under different confidence levels.
[0122] Then, collect real-time images of the water area through the pan-tilt monitoring camera, and use the image detection model improved based on the YOLOv10 algorithm constructed in this embodiment to identify the real-time images of the water area to obtain the detection results. If there is an illegal fishing behavior in the detection result discrimination information of the discrimination model, the alarm function is triggered. The specific implementation process is as follows:
[0123] Use the pan-tilt monitoring camera to collect real-time images of a certain water area, obtain images of illegal fishing operations by driving a ship in key waters and images of fishing with live bait on the shore, and perform data augmentation. The enhanced dataset has a total of 1259 images, which are divided into a training set and a test set according to a ratio of 3:2, and are respectively input into the image detection model improved based on the YOLOv10 algorithm and the original YOLOv10 model in this embodiment for training and testing. Select average accuracy, average recall, F1 score, mean average precision mAP, and processing time as the verification metrics for the model, and compare the performance of the improved algorithm in this paper with the model before improvement. The evaluation comparison results are shown in Table 1. The results show that the performance of the improved YOLOv10 detection method of the present invention is better than other detection methods, and all indicators are higher than those of the method before improvement. In practical applications, it is measured that mAP@0.5:0.95 has increased by 3.72%.
[0124] Table 1 Model Performance Evaluation Comparison
[0125]
[0126] Input the illegal fishing image to be detected into the trained illegal fishing discrimination model improved based on YOLOv10 for identification to obtain the detection results, referring to Figure 8 the schematic diagram of the illegal fishing detection results shown
[0127] If it is determined that there is an illegal fishing behavior, trigger the real-time voice alarm function of the smart speaker through the intelligent fishery supervision platform, and send a text message to notify relevant personnel to go to the law enforcement for disposal. Among them, the intelligent fishery supervision platform is an integrated system platform developed by China Telecom, which accesses China Telecom's NB-IoT devices through the LWM2M protocol and supports 24-hour unified management and control of smart speakers and text message sending.
[0128] In the embodiments disclosed in the present application, the computer storage medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The computer storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any suitable combination of the foregoing. More specific examples of the computer storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0129] The above are only the preferred embodiments of the present invention and do not impose any formal limitations on the present invention. Any person skilled in the art, without departing from the scope of the technical solution of the present invention and based on the technical essence of the present invention, any simple modifications, equivalent replacements, and improvements made to the above embodiments shall still fall within the scope of protection of the technical solution of the present invention.
Claims
1. A fishery illegal fishing behavior identification system based on the improved YOLOv10 algorithm, characterized in that: include: A water area real-time image acquisition module is used to acquire real-time images of the water area through a PTZ monitoring camera; An image detection model improved based on the YOLOv10 algorithm, including a Backbone network module for encoding the input image in feature form, a Neck network module for processing feature form encoding, and a Head network module for generating model prediction output; The Backbone network module includes a HWD-ADown module for decomposing feature graphs and a C2f module for feature extraction and feature fusion, and the output end of the HWD-ADown module is connected to the C2f module; The HWD-ADown module includes an HWD module and an ADown module; the HWD module is used to obtain a feature map decomposition processing result by subjecting the input image to Haar wavelet transform; the ADown module is used to perform a downsampling operation on the feature map decomposition processing result of the HWD module; An AM attention module is introduced based on the C2f module connected to the output end of the HWD-ADown module. The AM attention module includes a context attention module CxAM and a content attention module CnAM. The context attention module CxAM is used to focus on a specified local area or block on the feature map, and the content attention module CnAM is used to focus on relatively shallow features on the feature map. The output features of the context attention module CxAM and the content attention module CnAM are fused to obtain the final feature expression of the feature map. The detection and discrimination module is used to input the real-time image of the water area into the image detection model improved based on the YOLOv10 algorithm for identification to obtain the detection result. If the detection result discrimination information of the discrimination model contains illegal fishing behavior, the alarm function is triggered.
2. The fishery illegal fishing behavior identification system based on the improved YOLOv10 algorithm according to claim 1 is characterized in that: The HWD module includes a lossless feature coding block and a feature representation learning block. The lossless feature coding block includes a Haar wavelet transform layer. The Haar wavelet transform layer is used to perform Haar wavelet transform on the input image to obtain image decomposition information; the feature representation learning block is used to extract discriminative features from the image decomposition information obtained by the Haar wavelet transform layer and output the feature map decomposition processing result.
3. The fishery illegal fishing behavior identification system based on the improved YOLOv10 algorithm according to claim 1 is characterized in that: In the ADown module, the input feature map is divided into two branches after passing through the average pooling layer, one of which only passes through the CBS module, and the other passes through the maximum pooling layer and the CBS module in sequence. The results obtained by the two branches are output after feature fusion.
4. The fishery illegal fishing behavior identification system based on the improved YOLOv10 algorithm according to claim 2 is characterized in that: The process of the Haar wavelet transform layer performing Haar wavelet transform on the input image to obtain image decomposition information is specifically as follows: After the image with a resolution of h*w is decomposed by Haar wavelet transform, h represents the image height and w represents the image width, and four components are generated, including: low-frequency component A, high-frequency component H in the horizontal direction, high-frequency component V in the vertical direction and high-frequency component D in the diagonal direction. The size of each component is h / 2×w / 2, and the resolution is reduced to 1 / 4 of the original image resolution, and the number of channels of the feature map is quadrupled; The feature representation learning block includes a 1×1 convolution layer, a normalization layer, and a ReLU activation function. The feature representation learning block extracts discriminative features from the image decomposition information obtained by the Haar wavelet transform layer, and the process of outputting the feature map decomposition processing result is specifically as follows: The feature representation learning block uses a 1×1 convolutional layer to adjust the number of channels of the feature map, and then combines a normalization layer and a ReLU activation function to extract discriminative features to obtain a feature map decomposition processing result.
5. The fishery illegal fishing behavior identification system based on the improved YOLOv10 algorithm according to claim 1 is characterized in that: The process of the contextual attention module CxAM focusing on a specified local area or block on the feature map is specifically as follows: The contextual attention module CxAM converts the feature map F∈R C×h×w Use convolutional layer W respectively q And the convolutional layer W k Transformed into potential space, C is the number of channels, h is the image height, and w is the image width; the convolution layer W q It is used to guide attention to the specified local area or block in the input feature map, and the generated feature vector constitutes the first query space Q; the convolution layer W k It is used to calculate the relevance or similarity of the feature map of the specified local area or block in the first query space Q. The generated feature vector constitutes the first key space K. The conversion formula is as follows: Where {Q,K}∈R C×h×w , and then convert the first query space Q and the first key space K into vectors. In order to obtain the relationship between local areas or blocks, the following relationship matrix is obtained: R=Q T K Where R∈R N×N , N = h × w, and then normalize R through the sigmoid activation function and average pooling to obtain the attention matrix R1∈R C×h×w ; Then use the convolutional layer W V Transform the feature map F into feature V, where V∈R C×h×w , then perform element-wise multiplication of R1 and V to obtain the final attention representation E, the i-th feature map attention representation E with the number of channels C i for: AND i =R1⊙V i Among them, V i Represents the features of the convolution transformation of the feature map of the i-th feature map with C channels; The process of the content attention module CnAM focusing on relatively shallow features on the feature map is as follows: The content attention module CnAM takes the relatively shallow feature map C4∈S on the feature map. C×h×w Use convolutional layer W respectively p And the convolutional layer W z Transformed into potential space, the convolutional layer W p It is used to guide the attention to the relatively shallow features in the input feature map, and the generated feature vector constitutes the second query space P; the convolution layer W k It is used to calculate the relevance or similarity of relatively shallow feature maps in the second query space P. The generated feature vectors constitute the second key space Z. The conversion formula is as follows: Among them, {P,Z}∈S C×h×w , and then transform the second query space P and the second key space Z into vectors; in order to obtain the relationship between each local area or block, the following relationship matrix is obtained: S=P T Z Among them, S∈S N×N , N = h × w, and then normalize S through the sigmoid activation function and average pooling to obtain the attention matrix S1∈S C×h×w ; Finally, the feature map F uses the convolution layer W V The converted feature V is element-wise multiplied with S1 and finally the convolution layer W is used V , get the final attention representation D, the i-th feature map with C channels has the attention representation D i for: D i =S1⊙V i Among them, V i Represents the features of the convolution transformation of the feature map of the i-th feature map with C channels.
6. A method for identifying illegal fishing behavior based on the improved YOLOv10 algorithm, characterized in that: The following steps are involved: Collect real-time images of the water area through PTZ surveillance cameras; Build an image detection model based on the improved YOLOv10 algorithm, and perform model training and evaluation; The real-time image of the water area is input into the image detection model improved based on the YOLOv10 algorithm for recognition to obtain the detection result. If there is any information on illegal fishing behavior in the detection result, the alarm function is triggered; Among them, the image detection model improved based on the YOLOv10 algorithm is the image detection model improved based on the YOLOv10 algorithm as described in any one of claims 1-5.
7. The method for distinguishing illegal fishing behavior based on the improved YOLOv10 algorithm according to claim 6 is characterized in that: The model training and evaluation includes: obtaining an illegal fishing image dataset and dividing it into a training set and a test set; performing iterative training on the training set and retaining the best network weights; and testing the model using the test set; The iterative training is performed on the training set to retain the best network weights, specifically: the training set is input into the Backbone network module through the input layer, and then the feature map of the training set is obtained through the Neck network module; prediction is performed in the Head network module based on the feature map of the training set, the loss is calculated, and the model parameters are updated; retraining is performed until the loss function converges or the maximum number of iterations is reached, and the best network weights are saved.
8. The method for distinguishing illegal fishing behavior based on the improved YOLOv10 algorithm according to claim 7 is characterized in that: The illegal fishing images include images of ships operating illegal fishing operations in key waters and images of fishing using live bait on the shore; The illegal fishing images are annotated using annotation software to obtain a preliminary data set, and then each original image in the preliminary data set is amplified, wherein the amplification includes simulating weather environment data amplification.
9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for distinguishing illegal fishing behavior in fisheries based on the improved YOLOv10 algorithm as described in any one of claims 6 to 8 is implemented.
10. A computer-readable storage medium storing a computer program, wherein the computer program enables a computer to execute the method for distinguishing illegal fishing behavior in fisheries administration based on the improved YOLOv10 algorithm as described in any one of claims 6 to 8.
Citation Information
Patent Citations
YOLOv7-MO-based acete chinensis catching target identification and counting system and operation method thereof
CN116311096A
Improved Yolov5 pedestrian detection method and system and storage medium
CN117542013A
Construction site safety helmet detection method, computer equipment and storage medium
CN118644761A
Method for detecting number of people in high-place operation hanging basket based on improved YOLOv9 model
CN118942033A
Foggy day illegal sand dredger identification method and system based on deep learning
CN119169500A