Drowning diatom detection method based on efficient feature extraction and fusion

By using efficient feature extraction and fusion technology in the diatom detection method, the diatom detection model is optimized using SERAM and SLO Smooth L1 Loss modules, the problem of detection interference and low detection accuracy of elongated diatom objects in complex backgrounds is solved, and the effect of efficient identification of diatom images is achieved.

CN120032367AActive Publication Date: 2025-05-23GUANGDONG UNIV OF TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510089535.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-23
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

The existing diatom detection methods detect interference problems in complex backgrounds and the detection accuracy of elongated diatom objects is low, making it difficult to effectively identify diatom images.

Method used

Using a detection method based on efficient feature extraction and fusion, the diatom detection model is optimized and the detection accuracy is improved by constructing a compressed excitation residual attention module (SERAM) and a smooth L1 loss function (SLO Smooth L1 Loss) based on the aspect ratio feedback of the elongated target.

Benefits of technology

Effectively identify diatom images in complex backgrounds, improve the detection accuracy of elongated diatom objects, and overcome the problems of detection interference and low accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032367A_ABST
    Figure CN120032367A_ABST
Patent Text Reader

Abstract

The invention relates to a drowning diatom detection method based on efficient feature extraction and fusion, and the method comprises the steps: S1, obtaining a diatom image, and carrying out the preprocessing of the obtained diatom image, so as to generate input data; s2, optimizing a network structure of the diatom detection model based on the preprocessed input data to obtain prediction distribution; s3, optimizing a loss function of the diatom detection model based on the predicted distribution and the true value; and S4, outputting the type of the diatom image based on the optimized diatom detection model. According to the invention, diatom images collected under the scanning electron microscope can be efficiently detected and classified. According to the invention, the problems of detection interference in a diatom image under a complex background and low detection precision of a slender diatom target can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the field of image processing, and specifically designs a drowned diatom detection method based on efficient feature extraction and fusion. Background Art

[0002] Diatoms are widely distributed in the water bodies of the earth and are highly sensitive to the water environment in which they are located. Diatoms not only react sensitively to changes in the concentration of nutrients such as nitrogen and phosphorus, but also respond significantly to a variety of environmental factors. Therefore, they are often used as important indicators in environmental monitoring and forensic research. In forensic practice, forensic scientists often face great challenges when determining the cause of death of bodies salvaged from the water, drowned people, or bodies thrown into the water after death. In this case, diatom testing is regarded as the gold standard for diagnosing drowning. This not only provides an important basis for determining the cause of death of drowning victims, but also can infer the location of drowning through the results of diatom testing, thereby providing key clues for solving the case.

[0003] Common diatom detection methods are mainly divided into two categories: molecular biology methods and morphological methods. Molecular biology methods usually use DNA barcodes or DNA arrays for detection, but the technology is not yet mature. And as the body remains for a long time, the internal DNA of diatoms will degrade, which often leads to detection deviations. Morphological methods classify diatoms by observing the morphological characteristics of diatoms. Morphological methods can be further divided into methods based on manual observation and computer-based automatic detection and identification methods. Methods based on manual observation rely on experienced forensic experts to observe and classify. However, a major disadvantage of this method is that it is manpower and time-consuming. Computer-based methods are divided into two technical paths: traditional machine learning and deep learning. Traditional machine learning methods usually extract morphological differences between diatom species as input to the classifier. However, this method is more effective for diatom images with simple backgrounds, but the recognition effect is limited when facing diatoms with high category similarity. With the development of deep learning technology, researchers have gradually tried to apply it to diatom detection. Deep learning can automatically extract feature information through a large amount of training data, significantly improving detection performance. At present, many studies have improved and optimized diatom image detection in different aspects based on deep neural networks, but these methods have not yet solved the problems of detection interference in complex backgrounds and the low detection accuracy of slender diatom targets. Summary of the invention

[0004] In order to solve the above technical problems, the present invention provides a drowned diatom detection method based on efficient feature extraction and fusion, which can accurately identify diatom images in large quantities and overcome the detection interference problem under the complex background under the scanning electron microscope and the problem of low detection accuracy of slender diatom targets.

[0005] Specifically, the method comprises the following steps:

[0006] S1: acquiring diatom images, and preprocessing the acquired diatom images to generate input data;

[0007] S2: Based on the preprocessed input data, the network structure of the diatom detection model is optimized to obtain the predicted distribution;

[0008] S3: Based on the predicted distribution and the true value, the loss function of the diatom detection model is optimized to obtain the optimized diatom detection model;

[0009] S4: Based on the optimized diatom detection model, the type of diatom image is output.

[0010] Preferably, S1 comprises the following steps:

[0011] S1.1: Build image acquisition module;

[0012] S1.2: annotate the diatom image with a rotating frame according to the image acquisition module;

[0013] S1.3: Divide the diatom image dataset according to the image acquisition module.

[0014] Preferably, S2 comprises the following steps:

[0015] S2.1: Construct the Squeeze-and-Excitation Residual Attention Module SERAM (Squeeze-and-Excitation Residual Attention Module);

[0016] S2.2: Extract local features from the input feature x according to the SERAM module to obtain the local residual feature f residual ;

[0017] S2.3: According to the SERAM module, the local residual feature f residual Perform maximum pooling and upsampling operations to capture multi-scale feature relationships and obtain the global feature f global ;

[0018] S2.4: According to the SERAM module, the global feature f global Adopting the channel attention mechanism, the attention weight w between channels is generated by adaptive global average pooling attention ;

[0019] S2.5: For the global feature f global and the attention weight w attention Combined, we get the feature f that weights and strengthens the input features. enhanced .

[0020] Preferably, S2.1 comprises the following steps:

[0021] f residual =RELU(BN(Conv3×3(RELU(BN(Conv3×3(x))))))+x,

[0022] f global =Upsample(MaxPool(f residual )),

[0023] w attention =SE(f global ) = f global ⊙σ(W 2 RELU(W 1 ·GlobalAvgPool(f global ))),

[0024] f enhanced =(1+w attention )⊙f global .

[0025] In the above formula, x represents the input feature, RELU represents the activation function, BN represents the normalization process, Conv n×n(·) represents the convolution kernel of size n, Upsample represents the upsampling operation, MaxPool represents the maximum pooling operation, SE represents the channel attention mechanism, ⊙ represents the XOR operation, σ is the Sigmoid activation function, and There are two fully connected layers, C is the number of channels, r is the channel compression factor, and GlobalAvgPool represents the global average pooling operation.

[0026] Preferably, S3 includes:

[0027] S3.1: Construct SLO Smooth L1 Loss (Slender and Long Object Feedback Smooth L1 Loss) based on the aspect ratio feedback of slender objects;

[0028] S3.2: Obtain the aspect ratio of the target according to the SLO Smooth L1 Loss module;

[0029] S3.3: Dynamically determine the target loss function weight w based on the slender target judgment standard Mask determined by the SLO Smooth L1 Loss module and the aspect ratio Ratio of the acquired target i ;

[0030] S3.4: Loss function weight w obtained from the SLO Smooth L1 Loss module i Adjust the loss function.

[0031] Preferably, S3.1 comprises the following steps:

[0032]

[0033] L SLOSmoothL1 =∑(1+w i )·SmoothL1(pred i -target i ),

[0034]

[0035] In the above formula, x represents the input error value, δ is the threshold for controlling the error segmentation, i represents the i-th sample, and w i represents the weight of the i-th sample, pred i and target i They represent the predicted value and the target value respectively, Ratio represents the aspect ratio or height-to-width ratio of the target box, Mask is the criterion for judging slender targets, and exp(·) represents the exponential function.

[0036] Preferably, S4 includes:

[0037] S4.1: Construct data processing module;

[0038] S4.2: according to the data processing module, input the diatom image into the optimized diatom detection model;

[0039] S4.3: output the rotation bounding box, type and confidence of the diatom image according to the optimized diatom detection model;

[0040] S4.4: Count the output detection results and output the type of diatom images.

[0041] The second technical solution adopted by the present invention is: a drowned diatom detection method based on multi-scale feature fusion, comprising:

[0042] An image acquisition module, used to acquire diatom images and perform annotation and classification;

[0043] SERAM module, used to capture the correlation between channels and spatial features and optimize the diatom detection model;

[0044] SLO Smooth L1 Loss, used to improve the model's detection ability and accuracy for slender targets;

[0045] The data processing module is used to input the diatom image into the optimized diatom detection model, and the diatom detection model outputs the type of the diatom image.

[0046] The beneficial effects of the method of the present invention are as follows: the present invention first extracts residual features by performing local feature extraction on the input features, then performs maximum pooling and upsampling operations on the input features to capture multi-scale feature relationships to obtain global features, and obtains attention weights through a channel attention mechanism, and finally overcomes the problem of detection difficulties under complex backgrounds by feature fusion of residual features and attention weights. At the same time, the present invention can effectively cope with the challenge of low detection accuracy of slender diatom targets, and dynamically determine the weight of the target loss function by obtaining the aspect ratio Ratio of the target and the slender target judgment standard Mask, so as to overcome the problem of low detection accuracy of slender diatom targets without affecting the detection accuracy of other types of diatoms. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical methods in the embodiments of the present invention, the drawings required for the prior art and the embodiments are incorporated into the specification and constitute a part of the specification. The following drawings are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0048] Figure 1 is a flow chart of the present invention;

[0049] Figure 2 This is a structural diagram of a diatom detection algorithm based on efficient feature extraction and fusion of the present invention;

[0050] Figure 3 This is a comparison chart of the effects of the diatom detection algorithm based on efficient feature extraction and fusion of the present invention and the original Oriented R-CNN algorithm;

[0051] Figure 4 Diatom detection algorithm diagram based on efficient feature extraction and fusion of the present invention;

[0052] Figure 5 It is an example diagram of each algae in the diatom dataset of the present invention;

[0053] Figure 6 Schematic diagram of the attention residual module of the present invention;

[0054] Figure 7 It is a schematic diagram of slender target determination of the SLO SmoothL1 loss function of the present invention; Specific implementation plan

[0055] In order to make the purpose, features and advantages of the present invention more obvious and easy to understand, the technical methods in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. It should be noted that the following detailed descriptions are illustrative and are intended to provide further explanation of the present application. Unless otherwise specified, all other embodiments obtained by ordinary technicians in this field without creative work based on the embodiments of the present invention are within the scope of protection of the present invention.

[0056] The embodiment of the present invention provides a drowned diatom detection method based on efficient feature extraction and fusion, which is used to reduce detection interference under complex backgrounds and improve the detection accuracy of slender diatom targets at the algorithm level, thereby improving the detection effect of diatom images under a scanning electron microscope.

[0057] In a typical embodiment of the present invention, a diatom image dataset collected at a magnification of 1500 times under a scanning electron microscope is used as an example. Figure 1 , Figure 2 , the method comprises the following steps:

[0058] S1: acquiring diatom images, and preprocessing the acquired diatom images to generate input data;

[0059] S2: Based on the preprocessed input data, the network structure of the diatom detection model is optimized to obtain the predicted distribution;

[0060] S3: Based on the predicted distribution and the true value, the loss function of the diatom detection model is optimized to obtain the optimized diatom detection model;

[0061] S4: Based on the optimized diatom detection model, the type of diatom image is output.

[0062] Further, as a preferred embodiment of the present method, S1 comprises the following steps:

[0063] S1.1: Build image acquisition module;

[0064] S1.2: annotate the diatom image with a rotating frame according to the image acquisition module;

[0065] S1.3: Divide the diatom image dataset according to the image acquisition module.

[0066] According to the standard of DOTA dataset, the LabelImage tool was used to annotate the seven types of diatom images collected at 1500 times magnification under a scanning electron microscope with rotating boxes to obtain the diatom dataset, which was then divided into training set, validation set and test set in a ratio of 6:2:2. The specific division is shown in Table 1.

[0067] Table 1 Division of training set and test set in diatom dataset

[0068]

[0069] References to various algae images in the dataset Figure 5 .

[0070] The diatom dataset is created according to the DOTA data format. The diatom data is mainly saved in two folders: Labels and Images. Labels is used to store the annotation information of each image in the storage format of txt files, and Images stores all images.

[0071] After the DOTA format dataset is created, the dataset is divided into training set: validation set: test set = 6:2:2 ratio and then stored in the corresponding folder.

[0072] Further, refer to Figure 6 , S2 includes the following steps:

[0073] S2.1: Construct the Squeeze-and-Excitation Residual Attention Module SERAM (Squeeze-and-Excitation Residual Attention Module);

[0074] S2.2: Extract local features from the input feature x according to the SERAM module to obtain the local residual feature f residual ;

[0075] S2.3: According to the SERAM module, the local residual feature f residual Perform maximum pooling and upsampling operations to capture multi-scale feature relationships and obtain the global feature f global ;

[0076] S2.4: According to the SERAM module, the global feature f global Adopting the channel attention mechanism, the attention weight w between channels is generated by adaptive global average pooling attention ;

[0077] S2.5: For the global feature f global and the attention weight w attention Combined, we get the feature f that weights and strengthens the input features. enhanced .

[0078] Further, S2.1 includes the following steps:

[0079] The residual unit and channel attention mechanism are used to extract the results and perform weighted enhancement to obtain the SERAM structure:

[0080] fresidual =RELU(BN(Conv3×3(RELU(BN(Conv3×3(x))))))+x,

[0081] f global =Upsample(MaxPool(f residual )),

[0082] w attention =SE(f global )=f global ⊙σ(W 2 RELU(W 1 ·GlobalAvgPool(f global ))),

[0083] f enhanced =(1+w attention )⊙f global .

[0084] In the above formula, x represents the input feature, RELU represents the activation function, BN represents the normalization process, Conv n×n(·) represents the convolution kernel of size n, Upsample represents the upsampling operation, MaxPool represents the maximum pooling operation, SE represents the channel attention mechanism, ⊙ represents the XOR operation, σ is the Sigmoid activation function, and There are two fully connected layers, C is the number of channels, r is the channel compression factor, and GlobalAvgPool represents the global average pooling operation.

[0085] Further, refer to Figure 7 , S3 includes the following steps:

[0086] S3.1: Construct SLO Smooth L1 Loss (Slender and Long Object Feedback Smooth L1 Loss) based on the aspect ratio feedback of slender objects;

[0087] S3.2: Obtain the aspect ratio of the target according to the SLO Smooth L1 Loss module;

[0088] S3.3: Dynamically determine the target loss function weight w based on the slender target judgment standard Mask determined by the SLO Smooth L1 Loss module and the aspect ratio Ratio of the acquired target i ;

[0089] S3.4: Loss function weight w obtained from the SLO Smooth L1 Loss module iAdjust the loss function.

[0090] Further, S3.1 includes the following steps:

[0091]

[0092] L SLOSmoothL1 =∑(1+w i )·SmoothL1(pred i -target i ),

[0093]

[0094] In the above formula, x represents the input error value, δ is the threshold for controlling the error segmentation, i represents the i-th sample, and w i represents the weight of the i-th sample, pred i and target i They represent the predicted value and the target value respectively, Ratio represents the aspect ratio or height-to-width ratio of the target box, Mask is the criterion for determining slender targets, and exp(·) represents the exponential function.

[0095] Further, S4 comprises the following steps:

[0096] S4.1: Construct data processing module;

[0097] S4.2: according to the data processing module, input the diatom image into the optimized diatom detection model;

[0098] S4.3: output the rotation bounding box, type and confidence of the diatom image according to the optimized diatom detection model;

[0099] S4.4: Count the output detection results and output the type of diatom images.

[0100] The above is a specific description of the preferred implementation of the present invention, but the invention is not limited to the embodiments. Those skilled in the art may make other equivalent modifications or substitutions without violating the spirit of the invention, and these equivalent modifications or substitutions are included in the scope defined by the application claims.

Claims

1. A drowned diatom detection method based on efficient feature extraction and fusion, characterized in that: The method comprises the following steps: S1: acquiring diatom images, and preprocessing the acquired diatom images to generate input data; S2: Based on the preprocessed input data, the network structure of the diatom detection model is optimized to obtain the predicted distribution; S3: Based on the predicted distribution and the true value, the loss function of the diatom detection model is optimized to obtain the optimized diatom detection model; S4: Based on the optimized diatom detection model, the type of diatom image is output.

2. The drowned diatom detection method based on multi-scale feature fusion according to claim 1 is characterized in that: S1 includes the following steps: S1.1: Build image acquisition module; S1.2: annotate the diatom image with a rotating frame according to the image acquisition module; S1.3: Divide the diatom image dataset according to the image acquisition module.

3. The drowned diatom detection method based on efficient feature extraction and fusion according to claim 1 is characterized in that: S2 includes the following steps: S2.1: Construct the Squeeze-and-Excitation Residual Attention Module SERAM (Squeeze-and-Excitation Residual Attention Module); S2.2: Extract local features from the input feature x according to the SERAM module to obtain the local residual feature f residual ; S2.3: According to the SERAM module, the local residual feature f residual Perform maximum pooling and upsampling operations to capture multi-scale feature relationships and obtain the global feature f global ; S2.4: According to the SERAM module, the global feature f global Adopting the channel attention mechanism, the attention weight w between channels is generated by adaptive global average pooling attention ; S2.5: For the global feature f global and the attention weight w attention Combined, we get the feature f that weights and strengthens the input features. enhanced ; Among them, S2.1 includes the following steps: f residual =RELU(BN(Conv3×3(RELU(BN(Conv3×3(x))))))+x, f global =Upsample(MaxPool(f residual )), w attention =SE(f global )=f global ⊙σ(W2·RELU(W1·GlobalAvgPool(f global ))), f enhanced =(1+w attention )⊙f global . In the above formula, x represents the input feature, RELU represents the activation function, BN represents the normalization process, Conv n×n(·) represents the convolution kernel of size n, Upsample represents the upsampling operation, MaxPool represents the maximum pooling operation, SE represents the channel attention mechanism, ⊙ represents the XOR operation, σ is the Sigmoid activation function, and There are two fully connected layers, C is the number of channels, r is the channel compression factor, and GlobalAvgPool represents the global average pooling operation.

4. The drowned diatom detection method based on efficient feature extraction and fusion according to claim 1, characterized in that S3 include: S3.1: Construct a smooth L1 loss function SLO Smooth L1 Loss (Slenderand Long Object Feedback Smooth L1 Loss) based on the aspect ratio feedback of the slender target; S3.2: Obtain the aspect ratio of the target according to the SLO Smooth L1 Loss module; S3.3: Dynamically determine the target loss function weight w based on the slender target judgment standard Mask determined by the SLO Smooth L1 Loss module and the aspect ratio Ratio of the acquired target i ; S3.4: Loss function weight w obtained from the SLO Smooth L1 Loss module i Adjust the loss function; Among them, S3.1 includes the following steps: L SLOSmoothL1 =∑(1+w i )·SmoothL1(pred i -target i ), In the above formula, x represents the input error value, δ is the threshold for controlling the error segmentation, i represents the i-th sample, and w i represents the weight of the i-th sample, pred i and target i They represent the predicted value and the target value respectively, Ratio represents the aspect ratio or height-to-width ratio of the target box, Mask is the criterion for judging slender targets, and exp(·) represents the exponential function.

5. The drowned diatom detection method based on efficient feature extraction and fusion according to claim 1, characterized in that S4 include: S4.1: Construct data processing module; S4.2: according to the data processing module, input the diatom image into the optimized diatom detection model; S4.3: output the rotation bounding box, type and confidence of the diatom image according to the optimized diatom detection model; S4.4: Count the output detection results and output the type of diatom images.

6. A drowned diatom detection method based on multi-scale feature fusion, characterized in that: include: An image acquisition module, used to acquire diatom images and perform annotation and classification; SERAM module, used to capture the correlation between channels and spatial features and optimize the diatom detection model; SLO Smooth L1 Loss, used to improve the model's detection ability and accuracy for slender targets; The data processing module is used to input the diatom image into the optimized diatom detection model, and the diatom detection model outputs the type of the diatom image.

Citation Information

Patent Citations

  • Algae detection method and device and terminal equipment

    CN115187982A

  • Domain adaptation method using residual attention module

    CN115578593A

  • Visual detection method and system for micro surface defects based on multi-scale feature fusion

    CN115775236A

  • Drowning diatom detection method based on multi-scale feature fusion

    CN116229435A

  • Remote sensing image target detection method based on attention mechanism weighted feature fusion

    CN117611994A