Diatom small target detection method and device for drowning diagnosis

Through multi-layer convolution and feature enhancement modules to improve the detection accuracy of small objects in complex contexts, the problem of low detection accuracy of small objects in the existing technology is solved, and effective extraction and detection of fine-grained features of small objects is achieved.

CN120032166AActive Publication Date: 2025-05-23GUANGDONG UNIV OF TECH

Patent Information

Application Number
CN202510089779.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-23
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

The prior art small- and medium-sized object detection has low accuracy and weak feature expression in complex contexts, making it difficult to effectively capture and utilize subtle features of small objects.

Method used

Multi-layer convolution is used to extract image features, generate multi-scale feature maps, and local enhancement and cross-scale information integration of feature maps by constructing feature enhancement and selection modules SFE and multi-scale forward and reverse convolution module MPNC.

Benefits of technology

The detection accuracy of small target diatoms in complex backgrounds is improved, the ability to extract the fine-grained features of small targets is enhanced, and the problem of low detection accuracy of small targets in the existing technology is effectively solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032166A_ABST
    Figure CN120032166A_ABST
Patent Text Reader

Abstract

The invention relates to a diatom small target detection method and device for drowning diagnosis, and the method comprises the steps: S1, collecting a diatom image data set, and carrying out the processing of an image and a format; s2, extracting image features by using multilayer convolution, and generating a multi-scale feature map; s3, constructing a feature enhancement and selection module SFE to enhance the response of the degradation area and extracting edge information by using discrete wavelet transform; s4, constructing a multi-scale forward and reverse convolution module MPNC to perform local enhancement on the feature map; s5, performing cross-scale information integration on the feature maps with different resolutions by using weighted fusion; and S6, outputting an accurate target bounding box and a corresponding category prediction result. According to the invention, the problems of weak feature expression in detection of small target diatom and poor detection effect under a complex background can be effectively solved. The small target diatom detection method and device based on multi-scale feature perception can be widely applied to the fields of microscopic image processing, classification and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target detection, and in particular to a method and device for detecting small diatom targets for drowning diagnosis. Background Art

[0002] Diatom detection is an important diagnostic method in drowning diagnosis. It can be used to determine whether the body salvaged from the water is drowned or dumped after death, and even the point where the body fell into the water can be inferred based on its type. The type and quantity of diatoms will be affected by the water quality environment, that is, different water quality environments will have different types and quantities of diatoms. Therefore, diatom detection provides a reliable basis for diagnosing drowning cases. In the study, it was found that due to the low image resolution, unclear features, and easy occlusion of diatom small targets, it is easy to be missed. It is a challenge to design a model that is both accurate and efficient. Therefore, it is crucial to improve the accuracy of small target detection under complex backgrounds.

[0003] There are four main methods for small target detection. The first is the multi-scale feature fusion method, which relies on high-resolution feature maps for small target detection, but the fusion process may cause the fine-grained features of shallow features to be diluted or weakened, reducing the accuracy of small target detection. The second is the context-aware enhancement method, which can provide more background information, but is easily disturbed by complex backgrounds, resulting in the submersion of small target features. The third is the data enhancement method, which can only supplement the distribution of training data and cannot fundamentally solve the problem of insufficient small target feature extraction capabilities from the model structure. The fourth is the special convolution design, which reduces the amount of calculation, but its expression ability is limited, especially for the capture of fine-grained features of small targets.

[0004] In summary, the existing technologies for small target detection have problems such as multi-scale targets and weak feature representation. Therefore, how to effectively capture and utilize the subtle features of small targets to improve the detection accuracy of the model has become a key issue that needs to be solved in current research. Summary of the invention

[0005] In order to solve the above technical problems, the present invention provides a small target diatom detection method and device for drowning diagnosis, which can effectively solve the problems of weak feature expression of small target diatoms during detection and poor detection effect under complex background.

[0006] Specifically, the method comprises the following steps:

[0007] S1: Collect diatom detection image datasets and process images and formats;

[0008] S2: Use multi-layer convolution to extract image features and generate multi-scale feature maps;

[0009] S3: Construct the feature enhancement and selection module SFE (Selective Feature Enhancer) to enhance the response of the degraded area and extract edge information using discrete wavelet transform;

[0010] S4: Construct a multi-scale positive and negative convolution module MPNC (Multi-scale Positive-Negative Convolution) to locally enhance the feature map;

[0011] S5: Use weighted fusion to integrate cross-scale information of feature maps with different resolutions;

[0012] S6: Output accurate target bounding box and corresponding category prediction results.

[0013] Preferably, S1 comprises the following steps:

[0014] S1.1: Collect multiple types of diatom detection image datasets and use annotation tools to accurately annotate the targets;

[0015] S1.2: resize all images uniformly and normalize the image pixel values;

[0016] S1.3: Convert the annotation file into .txt format, including the target's center coordinates (x, y), width and height (w, h), and rotation angle (theta).

[0017] Preferably, S2 comprises the following steps:

[0018] S2.1: Extract and fuse image features;

[0019] Preferably, S2.1 comprises the following steps:

[0020] S2.1.1: Input image f input Extract preliminary features f through multiple YOLO11 CBS modules (Convolution BatchNorm-SiLU) 1 , the expression is as follows:

[0021] f 1 =CBS(f input )

[0022] =SiLU(BN(Conv2D(f input ))),

[0023] In the above formula, SiLU(.) represents SiLU activation function, BN(.) represents batch normalization, and Conv2D(.) represents two-dimensional convolution operation;

[0024] S2.1.2: Preliminary features are obtained by fusing deep and shallow features through the multi-scale feature extraction module C3K2 (Cross-Stage PartialKernel2) of YOLO11 to obtain feature f 2 , the expression is as follows:

[0025] f 11 ,f 12 =Spilt(f 1 ),

[0026] f 2 =C3K2(f 1 )

[0027] =Conv 3×3 (Concat(f 11 ,C3Stack(f 12 ))),

[0028] In the above formula, Spilt(.) represents the segmentation operation of the input feature, f 11 and f 12 Yes 1 The features after segmentation and the sum of the number of channels is equal to f 1 The number of channels, Conv n×n represents a convolution operation with a convolution kernel of n, Concat(.) represents a concatenation operation, and C3Stack(.) represents stacking n C3 modules. In the C3 module, a main branch has a direct skip connection and a sub-branch includes a convolution and a Bottleneck structure;

[0029] S2.1.3: High-level features are obtained through the long-distance modeling module C2PSA (CSP+PSABlock block) of YOLO11, using the global attention mechanism to obtain the feature f 3 , the expression is as follows:

[0030] f 21 ,f 22 =Spilt(f 2 ),

[0031] f 3 =C2PSA(f 2 )

[0032] =Conv 3×3 (Concat(f 21 ,PSAStack(f 22 ))),

[0033] In the above formula, f 21 and f 22 Yes 2 The features after segmentation and the sum of the number of channels is equal to f2 The number of channels, PSAStack(.) is the result of stacking n PSABlock blocks, each of which includes Attention, FNN and convolution modules;

[0034] S2.1.4: The global feature is aggregated through the spatial pyramid pooling module SPPF (Spatial Pyramid Pooling-Fast) to obtain the feature f 4 , the expression is as follows:

[0035] f 4 =SPPF(f 3 )

[0036] =Conv2D(Concat(MaxPool 5×5 ,MaxPool 3×3 ,f 3 )),

[0037] In the above formula, MaxPool n×n Represents a maximum pooling operation with a pooling window size of n×n.

[0038] S2.2: From the multi-layer feature output set {f 2 ,f 3 ,f 4}Extract multi-scale features.

[0039] Preferably, S3 comprises the following steps:

[0040] S3.1: Construct feature enhancement and selection module SFE (Selective Feature Enhancer);

[0041] S3.2: Use pooling and convolution operations to extract the global features of the degraded area from the feature map output by the backbone network

[0042] Preferably, S3.2 comprises the following steps:

[0043] S3.2.1: Use pooling to extract the global features of the degraded area. The expression is as follows:

[0044]

[0045] In the above formula, f i (i=2,3,4) represents the feature maps of multiple scales output from the backbone network. Represents feature maps of multiple scales after the pooling operation.

[0046] S3.2.2: Use 1×1 convolution to extract channel-level features and model the degraded region features channel by channel. The expression is as follows:

[0047]

[0048] In the above formula, DepthConv(.) represents the depth convolution operation, Indicates the characteristics of degraded areas.

[0049] S3.2.2: Multiply the degraded region features with the original feature map point by point to generate enhanced features The expression is as follows:

[0050]

[0051] In the above formula, ⊙ represents the point-by-point multiplication operation.

[0052] S3.3: Using discrete wavelet transform (DWT) to enhance edge information of multi-scale feature maps And weighted fusion to generate the final feature

[0053] Preferably, S3.3 comprises the following steps:

[0054] S3.3.1: Use discrete wavelet transform (DWT) to decompose the input features into low-frequency components and three high-frequency components (horizontal, vertical, and diagonal). The expression is as follows:

[0055]

[0056] In the above formula, DWT(.) represents the discrete wavelet transform operation, F iL Represents the low-frequency components of each scale feature, They represent the horizontal, vertical and diagonal components of the high-frequency components of each scale feature respectively.

[0057] S3.3.2: Fusion and enhancement of high-frequency components, slight suppression of low-frequency components, re-fusion of high-frequency and low-frequency features to generate enhanced features The expression is as follows:

[0058]

[0059] In the above formula, Represents the features after reintegration.

[0060] S3.3.3: The fused features Upsample to the size of the original feature map, and the output is weighted fused to generate the final enhanced features. The expression is as follows:

[0061]

[0062]

[0063] In the above formula, Upasample(.) represents the upsampling operation, and α and β are learnable weight parameters used to balance the weights of the degraded area and edge information.

[0064] Preferably, S4 comprises the following steps:

[0065] S4.1: Construct a multi-scale positive and negative convolution module MPNC (Multi-scale Positive-Negative Convolution);

[0066] S4.2: Use a multi-branch convolutional structure to extract multi-scale and multi-directional features, and each branch outputs features

[0067] Preferably, S4.2 comprises the following steps:

[0068] S4.2.1: Branch 1 extracts the global features between channels through 1×1 convolution, which is expressed as follows:

[0069]

[0070] S4.2.2: Branches 2 and 3 gradually pass through different paths (1×3 and 3×1) convolution to capture horizontal, vertical and local context information, and combine deconvolution to improve resolution. The expression is as follows:

[0071]

[0072]

[0073] S4.2.3: Branch 4 directly extracts high-resolution features through 1×1 convolution + deconvolution + 3×3 convolution, and the expression is as follows:

[0074]

[0075] In the above formula, Deconv2D(.) represents the deconvolution operation.

[0076] S4.3: The enhanced features of all branches are fused to generate the final output features, which are expressed as follows;

[0077]

[0078] In the above formula, is the final output feature.

[0079] Preferably, S5 comprises the following steps:

[0080] S5.1: Integrate information from feature maps of different resolutions;

[0081] Preferably, S5.1 comprises the following steps:

[0082] S5.1.1: Align the feature map sizes by upsampling high-level feature maps and downsampling low-level feature maps.

[0083] S5.1.2: Integrate information at different scales by fusing features across scales;

[0084] S5.2: Extract multi-scale information from the fused features and output multi-scale features.

[0085] Preferably, S5.2 comprises the following steps:

[0086] Standard convolution and dilated convolution are used to further extract multi-scale features, and all scale features (high resolution, medium resolution, and low resolution) are integrated and sent to the detection head for task processing. The expression is as follows:

[0087]

[0088] In the above formula, They are respectively represented as the integrated high, medium and low resolution feature maps, Represented as feature maps of different scales fed into the Head network.

[0089] Preferably, S6 comprises the following steps:

[0090] S6.1: Perform classification prediction and bounding box parameter prediction for each input feature map;

[0091] Preferably, S6.1 comprises the following steps:

[0092] S6.1.1: Using CBS module and multi-scale convolution MSCD (Multi-Scale Convolutional Decomposition), the expression is as follows:

[0093]

[0094] In the above formula, Represents the output of the CBS module, Represents the output of the MSDC module.

[0095] S6.1.2: Use independent convolutional layers to predict the target category for the features at each position, as expressed as follows:

[0096] Pclass =Conv2D(F MSDC ,N class ),

[0097] In the above formula, N class represents the number of categories, P class Represents classification prediction parameters.

[0098] S6.1.3: Use another independent convolutional layer to predict the regression parameters of the bounding box, expressed as follows:

[0099] P bbox =Conv2D(F MSDC ),

[0100] In the above formula, P bbox Represents the predicted bounding box parameters.

[0101] S6.2: Integrate the classification and bounding box prediction results to generate the final detection result.

[0102] Preferably, S6.2 comprises the following steps:

[0103] Decode the bounding box, convert the offset into the actual coordinates in the image, use NMS to remove redundant bounding boxes, and retain the target box with the highest confidence. The expression is as follows:

[0104] Bbox decoded =Decode(P bbox ),

[0105] Bbox final =NMS(Bbox decoded ,P class ),

[0106] In the above formula, NMS(.) represents non-maximum suppression.

[0107] The second technical solution adopted by the present invention is: a diatom small target detection device for drowning diagnosis, comprising:

[0108] SFE module: enhances the degraded area and extracts edge information to optimize the recognition ability of shallow features;

[0109] MPNC module: combines multi-scale convolution and positive and negative feature decomposition mechanism to highlight the characteristics of the target area and suppress background interference;

[0110] Feature fusion module: Integrates cross-scale contextual information through upsampling, downsampling, and weighted fusion of multi-resolution features;

[0111] Prediction module: Completes target classification and bounding box regression with independent branches, and combines decoding with non-maximum suppression (NMS) to finally output accurate target bounding boxes and corresponding category prediction results.

[0112] The present invention first extracts image features by using multi-layer convolution to generate a multi-scale feature map, strengthens the response of the degraded area by constructing a feature enhancement and selection module SFE and extracts edge information by using discrete wavelet transform, and locally enhances the feature map by constructing a multi-scale forward and reverse convolution module MPNC. Finally, weighted fusion is used to integrate cross-scale information of feature maps of different resolutions, and outputs an accurate target bounding box and corresponding category prediction results. At the same time, the present invention can effectively cope with the challenges of existing neural network models in small target detection. On the one hand, the feature enhancement and selection module is used to highlight the degraded area and effectively strengthen the edge information; on the other hand, the multi-branch convolution structure effectively increases local information and suppresses the background; at the same time, the CBS module and multi-scale decomposition convolution MSCD are used to enhance the fine-grained feature extraction capability of small targets, and finally improve the detection accuracy of small target diatom images. BRIEF DESCRIPTION OF THE DRAWINGS

[0113] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the prior art and the drawings required for use in the embodiments. The following drawings are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0114] Figure 1 It is a flow chart of a method for detecting small diatom targets for drowning diagnosis according to the present invention;

[0115] Figure 2 It is a schematic diagram of a specific embodiment of a diatom small target detection method for drowning diagnosis of the present invention;

[0116] Figure 3 It is a schematic diagram of the structure of an SFE module of a diatom small target detection device for drowning diagnosis of the present invention;

[0117] Figure 4 This is a schematic diagram of the structure of an MPCN module of a diatom small target detection device for drowning diagnosis according to the present invention; Specific implementation plan

[0118] In order to make the purpose, features and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. It should be noted that the following detailed descriptions are illustrative and are intended to provide further explanation of the present application. Unless otherwise specified, all other embodiments obtained by ordinary technicians in this field without creative work based on the embodiments of the present invention are within the scope of protection of the present invention.

[0119] The embodiments of the present invention provide a method and device for detecting small diatom targets for drowning diagnosis, which are used to improve the fine-grained feature extraction capability of small targets at the algorithm level, thereby improving the detection accuracy of small target diatom images.

[0120] A typical embodiment of the present invention takes the diatom image under a microscope magnified 1800 times as an example, referring to Figure 1 , Figure 2 , the method comprises the following steps:

[0121] S1: Collect diatom detection image datasets and process images and formats;

[0122] S2: Use multi-layer convolution to extract image features and generate multi-scale feature maps;

[0123] S3: Construct the feature enhancement and selection module SFE (Selective Feature Enhancer) to enhance the response of the degraded area and extract edge information using discrete wavelet transform;

[0124] S4: Construct a multi-scale positive and negative convolution module MPNC (Multi-scale Positive-Negative Convolution) to locally enhance the feature map;

[0125] S5: Use weighted fusion to integrate cross-scale information of feature maps with different resolutions;

[0126] S6: Output accurate target bounding box and corresponding category prediction results.

[0127] Further, S1 comprises the following steps:

[0128] S1.1: Collect multiple types of diatom detection image datasets and use annotation tools to accurately annotate the targets;

[0129] S1.2: resize all images uniformly and normalize the image pixel values;

[0130] S1.3: Convert the annotation file into .txt format, including the center coordinates (x, y), width and height (w, h), and rotation angle (theta) of the target.

[0131] Further, S2 comprises the following steps:

[0132] S2.1: Extract and fuse image features;

[0133] Further, S2.1 includes the following steps:

[0134] S2.1.1: Input image f input Extract preliminary features f through multiple YOLO11 CBS modules (Convolution BatchNorm-SiLU) 1 , the expression is as follows:

[0135] f 1 =CBS(f input )

[0136] =SiLU(BN(Conv2D(f input ))),

[0137] In the above formula, SiLU(.) represents SiLU activation function, BN(.) represents batch normalization, and Conv2D(.) represents two-dimensional convolution operation;

[0138] S2.1.2: Preliminary features are obtained by fusing deep and shallow features through the multi-scale feature extraction module C3K2 (Cross-Stage PartialKernel2) of YOLO11 to obtain feature f 2 , the expression is as follows:

[0139] f 11 ,f 12 =Spilt(f 1 ),

[0140] f 2 =C3K2(f 1 )

[0141] =Conv 3×3 (Concat(f 11 ,C3Stack(f 12 ))),

[0142] In the above formula, Spilt(.) represents the segmentation operation of the input feature, f 11 and f 12 Yes 1 The features after segmentation and the sum of the number of channels is equal to f 1 The number of channels, Conv n×nrepresents a convolution operation with a convolution kernel of n, Concat(.) represents a concatenation operation, and C3Stack(.) represents stacking n C3 modules. In the C3 module, a main branch has a direct skip connection and a sub-branch includes a convolution and a Bottleneck structure;

[0143] S2.1.3: High-level features are obtained through the long-distance modeling module C2PSA (CSP+PSABlock block) of YOLO11, using the global attention mechanism to obtain the feature f 3 , the expression is as follows:

[0144] f 21 ,f 22 =Spilt(f 2 ),

[0145] f 3 =C2PSA(f 2 )

[0146] =Conv 3×3 (Concat(f 21 ,PSAStack(f 22 ))),

[0147] In the above formula, f 21 and f 22 Yes 2 The features after segmentation and the sum of the number of channels is equal to f 2 The number of channels, PSAStack(.) is the result of stacking n PSABlock blocks, each of which includes Attention, FNN and convolution modules;

[0148] S2.1.4: The global feature is aggregated through the spatial pyramid pooling module SPPF (Spatial Pyramid Pooling-Fast) to obtain the feature f 4 , the expression is as follows:

[0149] f 4 =SPPF(f 3 )

[0150] =Conv2D(Concat(MaxPool 5×5 ,MaxPool 3×3 ,f 3 )),

[0151] In the above formula, MaxPool n×n Represents a maximum pooling operation with a pooling window size of n×n.

[0152] S2.2: From the multi-layer feature output set {f2 ,f 3 ,f 4}Extract multi-scale features.

[0153] Further, refer to Figure 3 , S3 includes the following steps:

[0154] S3.1: Construct feature enhancement and selection module SFE (Selective Feature Enhancer);

[0155] S3.2: Use pooling and convolution operations to extract the global features of the degraded area from the feature map output by the backbone network

[0156] Further, S3.2 includes the following steps:

[0157] S3.2.1: Use pooling to extract the global features of the degraded area. The expression is as follows:

[0158]

[0159] In the above formula, f i (i=2,3,4) represents the feature maps of multiple scales output from the backbone network. Represents feature maps of multiple scales after the pooling operation.

[0160] S3.2.2: Use 1×1 convolution to extract channel-level features and model the degraded region features channel by channel. The expression is as follows:

[0161]

[0162] In the above formula, DepthConv(.) represents the depth convolution operation, Indicates the characteristics of degraded areas.

[0163] S3.2.2: Multiply the degraded region features with the original feature map point by point to generate enhanced features The expression is as follows:

[0164]

[0165] In the above formula, ⊙ represents the point-by-point multiplication operation.

[0166] S3.3: Using discrete wavelet transform (DWT) to enhance edge information of multi-scale feature maps And weighted fusion to generate the final feature

[0167] Further, S3.3 includes the following steps:

[0168] S3.3.1: Use discrete wavelet transform (DWT) to decompose the input features into low-frequency components and three high-frequency components (horizontal, vertical, and diagonal). The expression is as follows:

[0169]

[0170] In the above formula, DWT(.) represents the discrete wavelet transform operation, F iL Represents the low-frequency components of each scale feature, They represent the horizontal, vertical and diagonal components of the high-frequency components of each scale feature respectively.

[0171] S3.3.2: Fusion and enhancement of high-frequency components, slight suppression of low-frequency components, re-fusion of high-frequency and low-frequency features to generate enhanced features The expression is as follows:

[0172]

[0173] F′ iH =Conv2D 3×3 (F iH ),

[0174] F′ iL =Conv2D 3×3 (F iL ),

[0175]

[0176] In the above formula, Represents the features after reintegration.

[0177] S3.3.3: The fused features Upsample to the size of the original feature map, and the output is weighted fused to generate the final enhanced features. The expression is as follows:

[0178]

[0179] In the above formula, Upsample(.) represents the upsampling operation, and α and β are learnable weight parameters used to balance the weights of the degraded area and edge information.

[0180] Further, refer to Figure 4 , S4 comprises the following steps:

[0181] S4.1: Construct a multi-scale positive and negative convolution module MPNC (Multi-scale Positive-Negative Convolution);

[0182] S4.2: Use a multi-branch convolutional structure to extract multi-scale and multi-directional features, and each branch outputs features

[0183] Further, S4.2 includes the following steps:

[0184] S4.2.1: Branch 1 extracts the global features between channels through 1×1 convolution, which is expressed as follows:

[0185]

[0186] S4.2.2: Branches 2 and 3 gradually pass through different paths (1×3 and 3×1) convolution to capture horizontal, vertical and local context information, and combine deconvolution to improve resolution. The expression is as follows:

[0187]

[0188] S4.2.3: Branch 4 directly extracts high-resolution features through 1×1 convolution + deconvolution + 3×3 convolution, and the expression is as follows:

[0189]

[0190] In the above formula, Deconv2D(.) represents the deconvolution operation.

[0191] S4.3: The enhanced features of all branches are fused to generate the final output features, which are expressed as follows;

[0192]

[0193] In the above formula, is the final output feature.

[0194] Further, S5 comprises the following steps:

[0195] S5.1: Integrate information from feature maps of different resolutions;

[0196] Further, S5.1 includes the following steps:

[0197] S5.1.1: Align the feature map sizes by upsampling high-level feature maps and downsampling low-level feature maps.

[0198] S5.1.2: Integrate information at different scales by fusing features across scales;

[0199] S5.2: Extract multi-scale information from the fused features and output multi-scale features.

[0200] Further, S5.2 includes the following steps:

[0201] Standard convolution and dilated convolution are used to further extract multi-scale features, and all scale features (high resolution, medium resolution, and low resolution) are integrated and sent to the detection head for task processing. The expression is as follows:

[0202]

[0203] In the above formula, They are respectively represented as the integrated high, medium and low resolution feature maps, Represented as feature maps of different scales fed into the Head network.

[0204] Further, S6 comprises the following steps:

[0205] S6.1: Perform classification prediction and bounding box parameter prediction for each input feature map;

[0206] Further, S6.1 includes the following steps:

[0207] S6.1.1: Using CBS module and multi-scale convolution MSCD (Multi-Scale Convolutional Decomposition), the expression is as follows:

[0208]

[0209]

[0210]

[0211] In the above formula, Represents the output of the CBS module, Represents the output of the MSDC module.

[0212] S6.1.2: Use independent convolutional layers to predict the target category for the features at each position, as expressed as follows:

[0213] P class =Conv2D(F MSDC ,N class ),

[0214] In the above formula, N class represents the number of categories, P class Represents classification prediction parameters.

[0215] S6.1.3: Use another independent convolutional layer to predict the regression parameters of the bounding box, expressed as follows:

[0216] P bbox =Conv2D(F MSDC ),

[0217] In the above formula, P bbox Represents the predicted bounding box parameters.

[0218] S6.2: Integrate the classification and bounding box prediction results to generate the final detection result.

[0219] Further, S6.2 includes the following steps:

[0220] Decode the bounding box, convert the offset into the actual coordinates in the image, use NMS to remove redundant bounding boxes, and retain the target box with the highest confidence. The expression is as follows:

[0221] Bbox decoded =Decode(P bbox ),

[0222] Bbox final =NMS(Bbox decoded ,P class ),

[0223] In the above formula, NMS(.) represents non-maximum suppression.

[0224] The above is a specific description of the preferred implementation of the present invention, but the invention is not limited to the embodiments. Those skilled in the art may make other equivalent modifications or substitutions without violating the spirit of the invention, and these equivalent modifications or substitutions are included in the scope defined by the application claims.

Claims

1. A method for detecting small diatom targets for drowning diagnosis, characterized in that: The method comprises the following steps: S1: Collect diatom detection image datasets and process images and formats; S2: Use multi-layer convolution to extract image features and generate multi-scale feature maps; S3: Construct the feature enhancement and selection module SFE (Selective Feature Enhancer) to enhance the response of the degraded area and extract edge information using discrete wavelet transform; S4: Construct a multi-scale positive and negative convolution module MPNC (Multi-scale Positive-Negative Convolution) to locally enhance the feature map; S5: Use weighted fusion to integrate cross-scale information of feature maps with different resolutions; S6: Output accurate target bounding box and corresponding category prediction results.

2. A diatom small target detection method for drowning diagnosis according to claim 1, characterized in that: S1 includes the following steps: S1.1: Collect multiple types of diatom detection image datasets and use annotation tools to accurately annotate the targets; S1.2: resize all images uniformly and normalize the image pixel values; S1.3: Convert the annotation file into .txt format, including the target's center coordinates (x, y), width and height (w, h), and rotation angle (theta).

3. A diatom small target detection method for drowning diagnosis according to claim 1, characterized in that: S2 includes the following steps: S2.1: Extract and fuse image features; Among them, S2.1 includes the following steps: S2.1.1: Input image f input The preliminary feature f1 is extracted through multiple YOLO11 CBS modules (ConvolutionBatchNorm-SiLU), and the expression is as follows: f1=CBS(f input ) =SiLU(BN(Conv2D(f input ))), In the above formula, SiLU(.) represents SiLU activation function, BN(.) represents batch normalization, and Conv2D(.) represents two-dimensional convolution operation; S2.1.2: Preliminary features are obtained by fusing deep and shallow features through the multi-scale feature extraction module C3K2 (Cross-Stage PartialKernel2) of YOLO11, and the expression is as follows: f 11 ,f 12 =Play(f1), f2=C3K2(f1) =Conv 3×3 (Concat(f 11 ,C3Stack(f 12 ))), In the above formula, Spilt(.) represents the segmentation operation of the input feature, f 11 and f 12 It is the feature after f1 segmentation and the sum of the number of channels is equal to the number of channels of f1. Conv n×n represents a convolution operation with a convolution kernel of n, Concat(.) represents a concatenation operation, and C3Stack(.) represents stacking n C3 modules. In the C3 module, a main branch has a direct skip connection and a sub-branch includes a convolution and a Bottleneck structure; S2.1.3: The high-level features are passed through the long-distance modeling module C2PSA (CSP+PSABlock block) of YOLO11, and the feature f3 is obtained using the global attention mechanism. The expression is as follows: f 21 ,f 22 =Play(f2), f3=C2PSA(f2) =Conv 3×3 (Concat(f 21 ,PSAStack(f 22 ))), In the above formula, f 21 and f 22 is the feature after f2 segmentation and the sum of the number of channels is equal to the number of channels of f2. PSAStack(.) is the result of stacking n PSABlock blocks. Each PSABlock block includes Attention, FNN and convolution modules. S2.1.4: The global feature is aggregated through the spatial pyramid pooling module SPPF (Spatial Pyramid Pooling-Fast) to obtain the feature f4, which is expressed as follows: f4=SPPF(f3) =Conv2D(Concat(MaxPool 5×5 ,MaxPool 3×3 ,f3)), In the above formula, MaxPool n×n Indicates the maximum pooling operation with a pooling window size of n×n; S2.2: Extract multi-scale features from the multi-layer feature output set {f2,f3,f4} of the backbone network.

4. A diatom small target detection method for drowning diagnosis according to claim 1, characterized in that: S3 includes the following steps: S3.1: Construct feature enhancement and selection module SFE (Selective Feature Enhancer); S3.2: Use pooling and convolution operations to extract the global features of the degraded area from the feature map output by the backbone network Among them, S3.2 includes the following steps: S3.2.1: Use pooling to extract the global features of the degraded area. The expression is as follows: In the above formula, f i (i=2,3,4) represents the feature maps of multiple scales output from the backbone network. Represents feature maps of multiple scales after the pooling operation. S3.2.2: Use 1×1 convolution to extract channel-level features and model the degraded region features channel by channel. The expression is as follows: In the above formula, DepthConv(.) represents the depth convolution operation, Indicates the characteristics of degraded areas. S3.2.2: Multiply the degraded region features with the original feature map point by point to generate enhanced features The expression is as follows: In the above formula, ⊙ represents the point-by-point multiplication operation. S3.3: Using discrete wavelet transform (DWT) to enhance edge information of multi-scale feature maps And weighted fusion to generate the final feature Among them, S3.3 includes the following steps: S3.3.1: Use discrete wavelet transform (DWT) to decompose the input features into low-frequency components and three high-frequency components (horizontal, vertical, and diagonal). The expression is as follows: In the above formula, DWT(.) represents the discrete wavelet transform operation, F iL Represents the low-frequency components of each scale feature, Respectively represent the horizontal, vertical and diagonal components of the high-frequency components of each scale feature; S3.3.2: Fusion and enhancement of high-frequency components, slight suppression of low-frequency components, re-fusion of high-frequency and low-frequency features to generate enhanced features The expression is as follows: F’ iH =Conv2D 3×3 (F iH ), F’ iL =Conv2D 3×3 (F iL ), In the above formula, Represents the features after reintegration. S3.3.3: The fused features Upsample to the size of the original feature map, and the output is weighted fused to generate the final enhanced features. The expression is as follows: In the above formula, Upsample(.) represents the upsampling operation, and α and β are learnable weight parameters used to balance the weights of the degraded area and edge information.

5. A diatom small target detection method for drowning diagnosis according to claim 1, characterized in that: S4 includes the following steps: S4.1: Construct a multi-scale positive and negative convolution module MPNC (Multi-scale Positive-Negative Convolution); S4.2: Use a multi-branch convolutional structure to extract multi-scale and multi-directional features, and each branch outputs features Among them, S4.2 includes the following steps: S4.2.1: Branch 1 extracts the global features between channels through 1×1 convolution, which is expressed as follows: S4.2.2: Branches 2 and 3 gradually pass through different paths (1×3 and 3×1) convolution to capture horizontal, vertical and local context information, and combine deconvolution to improve resolution. The expression is as follows: S4.2.3: Branch 4 directly extracts high-resolution features through 1×1 convolution + deconvolution + 3×3 convolution, and the expression is as follows: In the above formula, Deconv2D(.) represents the deconvolution operation; S4.3: The enhanced features of all branches are fused to generate the final output features, which are expressed as follows; In the above formula, is the final output feature.

6. The method for detecting small diatom targets for drowning diagnosis according to claim 1, characterized in that: S5 includes the following steps: S5.1: Integrate information from feature maps of different resolutions; Among them, S5.1 includes the following steps: S5.1.1: Align the feature map sizes by upsampling high-level feature maps and downsampling low-level feature maps. S5.1.2: Integrate information at different scales by fusing features across scales; S5.2: Extract multi-scale information from the fused features and output multi-scale features; Among them, S5.2 includes the following steps: Standard convolution and dilated convolution are used to further extract multi-scale features, and all scale features (high resolution, medium resolution, and low resolution) are integrated and sent to the detection head for task processing. The expression is as follows: In the above formula, They are respectively represented as the integrated high, medium and low resolution feature maps, Represented as feature maps of different scales fed into the Head network.

7. The method for detecting small diatom targets for drowning diagnosis according to claim 1, characterized in that: S6 includes the following steps: S6.1: Perform classification prediction and bounding box parameter prediction for each input feature map; Wherein, S6.1 comprises the following steps: S6.1.1: Using CBS module and multi-scale convolution MSCD (Multi-Scale Convolutional Decomposition), the expression is as follows: In the above formula, Represents the output of the CBS module, Represents the output of the MSDC module; S6.1.2: Use independent convolutional layers to predict the target category for the features at each position, as expressed as follows: P class =Conv2D(F MSDC ,N class ), In the above formula, N class represents the number of categories, P class represents the classification prediction parameter; S6.1.3: Use another independent convolutional layer to predict the regression parameters of the bounding box, expressed as follows: P bbox =Conv2D(F MSDC ), In the above formula, P bbox represents the predicted bounding box parameters; S6.2: Integrate the classification and bounding box prediction results to generate the final detection result; Among them, S6.2 includes the following steps: Decode the bounding box, convert the offset into the actual coordinates in the image, use NMS to remove redundant bounding boxes, and retain the target box with the highest confidence. The expression is as follows: Bbox decoded =Decode(P bbox ), Bbox final =NMS(Bbox decoded ,P class ), In the above formula, NMS(.) represents non-maximum suppression.

8. A diatom small target detection device for drowning diagnosis, characterized in that: include: Preprocessing module: realizes normalization processing of multi-resolution and multi-scale input; SFE module: enhances the degraded area and extracts edge information to optimize the recognition ability of shallow features; MPNC module: combines multi-scale convolution and positive and negative feature decomposition mechanism to highlight the characteristics of the target area and suppress background interference; Feature fusion module: Integrates cross-scale contextual information through upsampling, downsampling, and weighted fusion of multi-resolution features; Prediction module: Completes target classification and bounding box regression with independent branches, and combines decoding with non-maximum suppression (NMS) to finally output accurate target bounding boxes and corresponding category prediction results.

Citation Information

Patent Citations

  • Target detection enhancement model, target detection method, target detection device and electronic device

    CN113298080A

  • Drowning diatom detection method based on multi-scale feature fusion

    CN116229435A

Cited By

  • Helicopter rotor crack fault identification method and device based on video semantic segmentation

    CN121616921A

  • Helicopter rotor crack fault recognition method and device for video semantic segmentation

    CN121616921B