Flotation foam anomaly classification method and system based on multi-scale feature fusion

By adopting a multi-scale feature fusion method in flotation foam anomaly classification, combining deep convolutional neural network and spatial pyramid pooling module, the problem of low accuracy in complex environments in the existing technology is solved, and anomaly classification with high accuracy and robustness is achieved.

CN119625440BActive Publication Date: 2025-06-06CHANGSHA RES INST OF MINING & METALLURGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510154339.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-06-06
Estimated Expiration
2045-02-12

AI Technical Summary

Technical Problem

The existing flotation foam anomaly classification methods have low accuracy in complex environments and are difficult to adapt to multi-scale features and complex backgrounds.

Method used

The flotation bubble anomaly classification method based on multi-scale feature fusion is adopted. By constructing an exception classification model including multi-scale feature extraction layer, fully connected layer and output layer, combined with deep convolutional neural network and spatial pyramid pooling module, the effective fusion and processing of multi-scale features are achieved.

Benefits of technology

It significantly improves the accuracy and robustness of flotation foam anomaly classification, can maintain high accuracy in complex environments, reduce errors, and adapt to foam characteristics of different sizes and morphologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119625440B_ABST
    Figure CN119625440B_ABST
Patent Text Reader

Abstract

The present invention discloses a flotation foam anomaly classification method and system based on multi-scale feature fusion, the method comprising: constructing an anomaly classification model for flotation foam, the anomaly classification model comprising an input layer, a multi-scale feature extraction layer, a fully connected layer and an output layer connected in sequence; the multi-scale feature extraction layer comprises a plurality of feature extraction sublayers and an SPPF module connected in sequence; each feature extraction sublayer comprises a downsampling module and an IEFBlock module, the output end of each downsampling module is connected to the input end of the IEFBlock module of the same feature extraction sublayer, the input end of the downsampling module at the head end of the multi-scale feature extraction layer is connected to the input layer, and the output end of the IEFBlock module at the end is connected to the input end of the SPPF module; the trained anomaly classification model sorts the target foam image. The present invention improves the accuracy of flotation foam detection through multi-level feature extraction and fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent monitoring of flotation foam, and in particular to a flotation foam anomaly classification method and system based on multi-scale feature fusion. Background Art

[0002] The flotation foam process has important application value in mineral separation, wastewater treatment and other related fields. In the flotation process, the state, size and distribution of the foam directly affect the recovery rate and flotation efficiency of the mineral. Therefore, accurate monitoring and analysis of the abnormal conditions and size changes of the flotation foam are of great significance for improving the automation of the flotation process, optimizing process parameters and improving production efficiency.

[0003] Traditional flotation foam monitoring methods mostly rely on manual visual observation or traditional algorithms based on image processing, which have many limitations. First, manual inspection is not only inefficient, but also easily affected by the subjective factors of the observer, and cannot monitor the dynamic changes of foam in real time. Secondly, traditional image processing methods mostly rely on threshold-based, edge detection and other technologies. These methods have poor detection effects when the flotation foam has complex morphology, large differences in foam size, or large changes in the background environment. They are easily disturbed by noise and are difficult to meet the requirements of high precision and high efficiency. Therefore, how to achieve high-precision and automated foam anomaly classification and size measurement in a complex flotation foam environment has become a technical problem that needs to be solved urgently.

[0004] With the development of deep learning technology, especially the widespread application of convolutional neural networks (CNN) in the field of image segmentation and target detection, automated foam detection methods based on deep learning have gradually become a research hotspot. As a classic image segmentation network, U-Net has achieved remarkable results in medical image segmentation by virtue of its encoder-decoder structure and skip connection advantages. However, the application of U-Net in flotation foam images still faces challenges, especially when dealing with multi-scale features and complex backgrounds, its segmentation accuracy is often limited, and its performance is unstable when facing irregular shapes and large size differences in foam images.

[0005] YOLOv5 and YOLOv8 are deep learning models that are currently widely used in the field of target detection. Although they perform well in real-time target detection tasks, their application in flotation foam detection also has certain limitations. First, flotation foam usually presents a smaller target size, and the overlap between foams is strong. The YOLO series models have difficulties in detecting small targets and separating dense targets. Secondly, the background of flotation foam is usually more complex and may contain multiple elements such as liquids and mineral particles, resulting in more background noise. YOLOv5 and YOLOv8 have poor robustness in complex backgrounds, and may miss or misdetect. In addition, the scale adaptability of the YOLO series models is weak, and it is difficult to handle the multi-scale problem in foam images, resulting in fluctuations in detection accuracy at different sizes.

[0006] Therefore, existing deep learning models still have certain limitations in the anomaly classification and size measurement of flotation foam, and a new approach is urgently needed to overcome these challenges. Summary of the invention

[0007] The present invention provides a flotation foam anomaly classification method and system based on multi-scale feature fusion, which is used to solve the technical problem of low accuracy caused by the inability of existing flotation foam anomaly classification methods to adapt to complex flotation foam environments.

[0008] In order to solve the above technical problems, the technical solution proposed by the present invention is:

[0009] A flotation foam anomaly classification method based on multi-scale feature fusion comprises the following steps:

[0010] Construct an abnormal classification model for flotation foam, the abnormal classification model includes an input layer, a multi-scale feature extraction layer, a fully connected layer and an output layer connected in sequence; the multi-scale feature extraction layer includes a plurality of feature extraction sublayers and an SPPF module connected in sequence; each feature extraction sublayer is composed of a downsampling module and an IEFBlock module, the output end of each downsampling module is connected to the input end of the IEFBlock module of the same feature extraction sublayer, the input end of the downsampling module at the head end of the multi-scale feature extraction layer is connected to the input layer, and the output end of the IEFBlock module at the end of the multi-scale feature extraction layer is connected to the input end of the SPPF module;

[0011] The IEFBlock module includes: a spilt segmentation module, a standard convolution module, a depth-separable convolution module, a splicing module A, an enhancement module and a splicing module B;

[0012] The output end of the spilt segmentation module is connected to the input end of the standard convolution module, the depthwise separable convolution module and the splicing module B respectively, the output end of the standard convolution module and the depthwise separable convolution module is connected to the input end of the splicing module A, the output end of the splicing module A is connected to the input end of the enhancement module, and the output end of the enhancement module is connected to the input end of the splicing module B;

[0013] The spilt segmentation module is used to separate the input feature map by channel to obtain a first separation feature, a second separation feature and a third separation feature; the standard convolution module is used to extract the local core feature of the first separation feature; the depth-separable convolution module is used to extract the high-resolution feature of the second separation feature; the splicing module A is used to splice the local core feature with the high-resolution feature to obtain a first splicing feature; the enhancement module is used to enhance the first splicing feature; the splicing module B is used to splice the enhanced first splicing feature with the third separation feature.

[0014] Collect foam images of different states under different flotation conditions to construct a training set; the states of the foam images include at least normal and abnormal; use the training set to train an abnormal classification model for the flotation foam, and use the trained abnormal classification model to sort the target foam image.

[0015] Preferably, the output layer includes an FCN module and a CLS module, the input end of the FCN module is connected to the output end of the fully connected layer, the output end of the FCN module is connected to the input end of the CLS module, and the CLS module is used for anomaly classification.

[0016] Preferably, a segmentation and positioning module is also included, which includes an IFPN layer and multiple SEGBlock modules. The input end of the IFPN layer is respectively connected to the output ends of the multiple IEFBlock modules and the SPPF module of the multi-scale feature extraction layer, and the output end of the IFPN layer is respectively connected to the input ends of the multiple SEGBlock modules.

[0017] Preferably, the feature extraction sublayer includes a first feature extraction sublayer, a second feature extraction sublayer, a third feature extraction sublayer and a fourth feature extraction sublayer connected in sequence, and the output ends of the IEFBlock module and the SPPF module of the second feature extraction sublayer and the third feature extraction sublayer are connected to the input end of the IFPN layer.

[0018] Preferably, the IFPN layer includes: a first splicing module, a second splicing module, a first fusion module, and a second fusion module;

[0019] The input end of the first splicing module is respectively connected to the output end of the first feature extraction sublayer and the output end of the second splicing module, and the output end of the first splicing module is connected to the input end of the first fusion module; the first splicing module is used to splice the first feature map output by the first feature extraction sublayer with the second splicing map output by the second splicing module to obtain a first splicing map;

[0020] The input end of the first fusion module is also jump-connected to the output end of the first feature extraction sublayer, the input end of the first fusion module is also connected to the output end of the second splicing module, and the output end of the first fusion module is also connected to the input end of the second fusion module; the first fusion module is used to fuse the first splicing image, the second splicing image output by the second splicing module, and the jump-connected feature image of the first feature image, and output a first fusion image;

[0021] The input end of the second splicing module is connected to the output ends of the second feature extraction sublayer and the third feature extraction sublayer, and is used to splice the second feature map output by the second feature extraction sublayer and the third feature map output by the third feature extraction sublayer to obtain a second splicing map;

[0022] The input end of the second fusion module is also jump-connected to the output end of the second feature extraction sublayer, and the input end of the second fusion module is also connected to the output end of the third feature extraction sublayer; the second fusion module is used to fuse the first fusion image, the second splicing image output by the second splicing module, and the jump-connected feature image of the first feature image, and output a second fusion image.

[0023] Preferably, the first splicing module includes:

[0024] A first upsampling unit, a first splicing unit and a first enhancement unit, wherein an input end of the first upsampling unit is connected to an output end of the second splicing module, an output end of the first upsampling unit is connected to an input end of the first splicing unit, and an input end of the first splicing unit is also connected to an output end of the first feature extraction sublayer; an output end of the first splicing unit is connected to an input end of the first enhancement unit, and an output end of the first enhancement unit is connected to an input end of the first fusion module;

[0025] and / or

[0026] The second splicing module includes:

[0027] A second upsampling unit, a second splicing unit and a second enhancing unit, wherein the input end of the second upsampling unit is connected to the output end of the third feature extraction sublayer, the output end of the second upsampling unit is connected to the input end of the second splicing unit, and the input end of the second splicing unit is also connected to the output end of the second feature extraction sublayer; the output end of the second splicing unit is connected to the input end of the second enhancing unit, and the output end of the second enhancing unit is connected to the input ends of the first splicing module and the first fusion module.

[0028] Preferably, the first fusion module includes:

[0029] A first fusion unit, a third channel adjustment unit and a third enhancement unit, wherein the input end of the first fusion unit is jump-connected to the output end of the first feature extraction sublayer, and the input end of the first fusion unit is also connected to the output ends of the first splicing module and the second splicing module; the output end of the first fusion unit is connected to the input end of the third channel adjustment unit, the output end of the third channel adjustment unit is connected to the input end of the third enhancement unit, and the output end of the third enhancement unit is connected to the input end of the second fusion module;

[0030] and / or

[0031] The second fusion module includes:

[0032] a second fusion unit, a fourth channel adjustment unit, and a fourth enhancement unit, wherein the input end of the second fusion unit is jump-connected to the output end of the second feature extraction sublayer, and the input end of the second fusion unit is also connected to the output ends of the first fusion module and the third feature sublayer; the output end of the second fusion unit is connected to the input end of the fourth channel adjustment unit, and the output end of the fourth channel adjustment unit is connected to the input end of the fourth enhancement unit;

[0033] and / or

[0034] The IFPN layer also includes a first channel adjustment module, a second channel adjustment module and a third channel adjustment module;

[0035] The input end of the first channel adjustment module is connected to the output end of the first enhancement unit, and is used to perform channel adjustment on the first spliced ​​image to obtain a high-level feature map F1;

[0036] The input end of the second channel adjustment module is connected to the output end of the third enhancement unit, and is used to perform channel adjustment on the first fusion image to obtain a middle-level feature map F2;

[0037] The input end of the third channel adjustment module is connected to the output end of the fourth enhancement unit, and is used to perform channel adjustment on the first fusion image to obtain the bottom feature image F3.

[0038] Preferably, the SEGBlock module comprises:

[0039] A fifth channel adjustment unit, a fifth enhancing unit, a third fusion unit, a sixth enhancing unit, a third splicing unit and a sixth channel adjustment unit; the output end of the fifth channel adjustment unit is connected to the input end of the fifth enhancing unit, and the output end of the fifth enhancing unit is connected to the input end of the third fusion unit; the output end of the third fusion unit is connected to the input end of the sixth enhancing unit; the output end of the sixth enhancing unit is connected to the input end of the third splicing unit, and the output end of the third splicing unit is connected to the input end of the sixth channel adjustment unit; the output end of the fifth channel adjustment unit is also connected to the input end of the third fusion unit, and the output end of the fifth enhancing unit is also connected to the input end of the third splicing unit.

[0040] Preferably, the loss function of the abnormal classification model satisfies:

[0041] ;

[0042] ;

[0043] ;

[0044] ;

[0045] Among them, CLS Loss represents classification loss, rect Loss represents region positioning loss, and Mask Loss represents mask loss. , , are hyperparameters that control the importance of classification, localization, and segmentation losses in the total loss respectively; is the actual category label, is the probability of the predicted category, is the number of samples; are the predicted bounding box coordinates, are the actual bounding box coordinates, is the predicted pixel probability, is the actual pixel label.

[0046] A computer system comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the steps of the above method are implemented when the processor executes the computer program.

[0047] The present invention has the following beneficial effects:

[0048] 1. The present invention can simultaneously consider the characteristics of foams of different sizes, shapes and complex backgrounds, greatly improving the accuracy of flotation foam anomaly classification. In complex flotation environments, such as bubble overlap, background interference and other factors, the system can still maintain a high level of robustness, significantly reducing the errors caused by illumination changes, background interference and other factors in traditional methods. The present invention not only improves the accuracy of flotation foam detection, but also improves the robustness and adaptability of the model to a certain extent, thereby providing reliable technical support for the automated monitoring and optimization of flotation foam.

[0049] 2. In the preferred embodiment, the present invention combines advanced deep learning technology and multi-scale feature fusion methods, which can quickly and in real time classify and measure the flotation foam abnormalities, and meet the needs of online monitoring. While ensuring high accuracy, it significantly improves the efficiency of flotation foam monitoring, and can detect problems in time and take corresponding measures to optimize the flotation process.

[0050] In addition to the above-described purposes, features and advantages, the present invention has other purposes, features and advantages. The present invention will be further described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] The drawings constituting a part of this application are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0052] Figure 1 A structural block diagram of an abnormal classification model provided by an embodiment of the present invention;

[0053] Figure 2 IEFBlock structure diagram provided by an embodiment of the present invention;

[0054] Figure 3 is a structural diagram of an IFPN provided by an embodiment of the present invention;

[0055] Figure 4 is a structural diagram of SEGBlock provided by an embodiment of the present invention; DETAILED DESCRIPTION

[0056] The embodiments of the present invention are described in detail below with reference to the accompanying drawings, but the present invention can be implemented in many different ways as defined and covered by the claims.

[0057] In this embodiment, a flotation foam anomaly classification method based on multi-scale feature fusion is provided, including:

[0058] 1. Construct an abnormal classification model for flotation foam (also known as FlotationFoamNet network);

[0059] like Figure 1 As shown, the abnormal classification model includes an input layer, a multi-scale feature extraction layer, a fully connected layer and an output layer connected in sequence; the multi-scale feature extraction layer includes a plurality of feature extraction sublayers and an SPPF module connected in sequence; each feature extraction sublayer is composed of a downsampling module and an IEFBlock module, the output end of each downsampling module is connected to the input end of the IEFBlock module of the same feature extraction sublayer, the input end of the downsampling module at the head end of the multi-scale feature extraction layer is connected to the input layer, and the output end of the IEFBlock module at the end is connected to the input end of the SPPF module;

[0060] Further, such as Figure 2 As shown, the IEFBlock module includes: a spilt segmentation module, a standard convolution module, a depth-separable convolution module, a splicing module A, an enhancement module and a splicing module B;

[0061] In the preferred embodiment, both the splicing module A and the splicing module B perform the Concatenate operation. Figure 2 In the expression, it is concat;

[0062] The output end of the spilt segmentation module is connected to the input end of the standard convolution module, the depthwise separable convolution module and the splicing module B respectively, the output end of the standard convolution module and the depthwise separable convolution module is connected to the input end of the splicing module A, the output end of the splicing module A is connected to the input end of the enhancement module, and the output end of the enhancement module is connected to the input end of the splicing module B;

[0063] The spilt segmentation module is used to separate the input feature map by channel to obtain a first separation feature, a second separation feature and a third separation feature; the standard convolution module is used to extract the local core feature of the first separation feature; the depth-separable convolution module is used to extract the high-resolution feature of the second separation feature; the splicing module A is used to splice the local core feature with the high-resolution feature to obtain a first splicing feature; the enhancement module is used to enhance the first splicing feature; the splicing module B is used to splice the enhanced first splicing feature with the third separation feature.

[0064] In a preferred embodiment, the standard convolution module is connected to the splicing module A via a standard convolution enhancement module, and the depthwise separable convolution module is connected to the splicing module A via a depthwise convolution enhancement module;

[0065] In a preferred embodiment, the standard convolution module is a 3x3 convolution layer; the depthwise separable convolution module is a 5x5 depthwise separable convolution layer; and the standard convolution enhancement module and the depthwise convolution enhancement module are both 1x1 convolution layers.

[0066] The standard convolution module is used for standard convolution operations to extract the core features of the input data. The standard convolution module captures local features through 3x3 convolution and extracts detail information of the flotation foam image through sliding window operations. Subsequently, the standard convolution enhancement module performs dimensionality reduction and feature compression through 1x1 convolution, effectively reducing the amount of calculation while maintaining the integrity and recognizability of the features.

[0067] The depth-separable convolution module uses the depth-separable convolution technology to decompose the standard convolution into depth convolution and point convolution, and performs independent spatial convolution operations on each channel of the input data. Then the depth-separable convolution module combines the channels through 1x1 convolution. This not only retains the high-resolution features of the input data, but also significantly reduces the number of parameters and the amount of calculation, thereby improving the efficiency of feature extraction, which is particularly suitable for complex morphological analysis of flotation foam.

[0068] In the preferred embodiment, the concatenation module A concatenates the local core features with the high-resolution features. By comprehensively utilizing multi-path features, a richer and more diverse feature representation can be obtained, thereby improving the recognition and classification capabilities of flotation foam size and morphology. The concatenation module A performs a Concatenate operation.

[0069] In the preferred solution, the enhancement module preferably uses a 1x1 convolution layer. After the feature fusion is completed, the fused feature map is convolved through a 1x1 convolution layer to further enhance the feature expression capability. This step can effectively reduce the number of channels of the feature map while retaining important feature information.

[0070] In this embodiment, the third separation feature retains the original data characteristics to ensure that the splicing module B retains and utilizes the original information during the feature fusion process, thereby improving the diversity and expression ability of the final output features; the splicing module B finally splices the feature map of the post-fusion convolution enhancement block with the direct transfer path data to generate the final output of the downstream task. In this way, it is possible to provide rich and high-quality feature representation while maintaining efficient calculation, meeting the needs of flotation foam size measurement and feature extraction.

[0071] This comprehensive feature extraction and fusion block not only plays a key role in the entire system structure, but also provides a strong guarantee for the overall performance improvement through its efficient feature extraction and fusion technology.

[0072] In a preferred embodiment, the output layer includes an FCN module and a CLS module, the input end of the FCN module is connected to the output end of the fully connected layer, the output end of the FCN module is connected to the input end of the CLS module, and the CLS module is used for anomaly classification.

[0073] In this embodiment, if Figure 1 As shown in the figure, based on the requirements of flotation foam sorting, the feature extraction sublayer of the anomaly classification model includes four layers. The workflow of the anomaly classification model is as follows:

[0074] First, the input image is input into the input layer, which is a 3×3 convolution layer with a stride of 1, and the convolution layer generates a feature map F-1. This operation preliminarily extracts feature information and enhances the feature expression capability of the network while keeping the spatial resolution of the input image unchanged.

[0075] Next, feature map F-1 is downsampled through a 3×3 convolutional layer with a stride of 2 to generate feature map F-2. The downsampling operation significantly reduces the spatial size of the feature map and increases the receptive field, laying the foundation for subsequent feature extraction.

[0076] The feature map F-2 is input into the first IEFBlock module, which achieves efficient multi-scale feature extraction through feature enhancement and reorganization operations. Specifically, IEFBlock effectively integrates local details with global context information through parallel convolution branches and fusion strategies to generate an enhanced feature map F-3.

[0077] Feature map F-3 then enters a 3×3 convolution layer with a stride of 2 for further downsampling, and outputs feature map F-4. Feature map F-4 is input into the second IEFBlock module, and feature map F-5 is generated through feature extraction and fusion operations. This stage further expands the receptive field and extracts deep semantic information.

[0078] In order to capture higher-level feature information, feature map F-5 is downsampled through a 3×3 convolution layer with a stride of 2 to generate feature map F-6. Subsequently, feature map F-6 enters the third IEFBlock module, which extracts richer multi-scale features through feature enhancement and fusion strategies, and outputs feature map F-7.

[0079] Feature map F-7 is downsampled again through a 3×3 convolutional layer with a stride of 2 to generate feature map F-8. Feature map F-8 is input into the fourth IEFBlock module, which further fuses local and global feature information and outputs the final feature map F-9.

[0080] Feature map F-9 is input into the SPPF module, i.e., the spatial pyramid pooling module, which captures spatial information of different scales through multi-scale pooling operations and generates a high-level multi-scale fused feature map F-10. The feature map F-10 output by the SPPF module contains global context information, which provides support for subsequent processing of the network.

[0081] Based on the multi-scale feature fusion, the feature map F-10 output by SPPF (spatial pyramid pooling module) is further processed in depth:

[0082] First, feature map F-10 is input into a 1×1 convolutional layer, which performs channel compression and dimension adjustment on the feature map to generate a highly expressive feature map F-11. This step aims to reduce redundant information, extract more representative features, and retain global context information, providing a more concise and effective feature representation for subsequent classification tasks.

[0083] Subsequently, the feature map F-11 enters a fully connected layer (FCN) to further enhance the feature expression capability through global feature integration. The fully connected layer maps the feature map F-11 to a high-dimensional space through weight learning, capturing the global semantic information of the foam anomaly features as a whole, ensuring a more accurate classification process.

[0084] Finally, the feature vector processed by the fully connected layer is passed to the CLS classification head, which uses the final global feature vector to output the classification result of foam anomaly (CLS). Through this process, the network can accurately classify the abnormal type of foam.

[0085] In a preferred embodiment, in order to measure the foam size, a segmentation and positioning module is further included, the segmentation and positioning module includes an IFPN layer and multiple SEGBlock modules, the input end of the IFPN layer is respectively connected to the output ends of the multiple IEFBlock modules and the SPPF module of the multi-scale feature extraction layer, and the output end of the IFPN layer is respectively connected to the input ends of the multiple SEGBlock modules. The output ends of the IEFBlock modules and the SPPF modules of the second feature extraction sublayer and the third feature extraction sublayer are connected to the input end of the IFPN layer.

[0086] After feature extraction is completed, the network realizes cross-level fusion of multi-scale features through IFPN (Integrated Feature Pyramid Network, IFPN), ensuring that feature maps of different depths are effectively combined in terms of spatial resolution and semantic information, thereby improving the segmentation and positioning accuracy of the target area. IFPN adopts a step-by-step backhaul and fusion mechanism to integrate shallow detail features with deep semantic features to form an efficient multi-scale feature expression. This fusion method can effectively deal with targets of different sizes and complex shapes, and improve the overall performance of target detection and region segmentation.

[0087] Specifically, Figure 3 As shown, the IFPN layer includes: a first splicing module, a second splicing module, a first fusion module, and a second fusion module;

[0088] The input end of the first splicing module is respectively connected to the output end of the first feature extraction sublayer and the output end of the second splicing module, and the output end of the first splicing module is connected to the input end of the first fusion module; the first splicing module is used to splice the first feature map output by the first feature extraction sublayer with the second splicing map output by the second splicing module to obtain a first splicing map;

[0089] The input end of the first fusion module is also jump-connected to the output end of the first feature extraction sublayer, the input end of the first fusion module is also connected to the output end of the second splicing module, and the output end of the first fusion module is also connected to the input end of the second fusion module; the first fusion module is used to fuse the first splicing image, the second splicing image output by the second splicing module, and the jump-connected feature image of the first feature image, and output a first fusion image;

[0090] The input end of the second splicing module is connected to the output ends of the second feature extraction sublayer and the third feature extraction sublayer, and is used to splice the second feature map output by the second feature extraction sublayer and the third feature map output by the third feature extraction sublayer to obtain a second splicing map;

[0091] The input end of the second fusion module is also jump-connected to the output end of the second feature extraction sublayer, and the input end of the second fusion module is also connected to the output end of the third feature extraction sublayer; the second fusion module is used to fuse the first fusion image, the second splicing image output by the second splicing module, and the jump-connected feature image of the first feature image, and output a second fusion image.

[0092] Specifically, the first splicing module includes:

[0093] A first upsampling unit, a first splicing unit and a first enhancement unit, wherein an input end of the first upsampling unit is connected to an output end of the second splicing module, an output end of the first upsampling unit is connected to an input end of the first splicing unit, and an input end of the first splicing unit is also connected to an output end of the first feature extraction sublayer; an output end of the first splicing unit is connected to an input end of the first enhancement unit, and an output end of the first enhancement unit is connected to an input end of the first fusion module;

[0094] The second splicing module includes:

[0095] A second upsampling unit, a second splicing unit and a second enhancing unit, wherein the input end of the second upsampling unit is connected to the output end of the third feature extraction sublayer, the output end of the second upsampling unit is connected to the input end of the second splicing unit, and the input end of the second splicing unit is also connected to the output end of the second feature extraction sublayer; the output end of the second splicing unit is connected to the input end of the second enhancing unit, and the output end of the second enhancing unit is connected to the input ends of the first splicing module and the first fusion module.

[0096] The first fusion module includes:

[0097] A first fusion unit, a third channel adjustment unit and a third enhancement unit, wherein the input end of the first fusion unit is jump-connected to the output end of the first feature extraction sublayer, and the input end of the first fusion unit is also connected to the output ends of the first splicing module and the second splicing module; the output end of the first fusion unit is connected to the input end of the third channel adjustment unit, the output end of the third channel adjustment unit is connected to the input end of the third enhancement unit, and the output end of the third enhancement unit is connected to the input end of the second fusion module;

[0098] The second fusion module includes:

[0099] a second fusion unit, a fourth channel adjustment unit, and a fourth enhancement unit, wherein the input end of the second fusion unit is jump-connected to the output end of the second feature extraction sublayer, and the input end of the second fusion unit is also connected to the output ends of the first fusion module and the third feature sublayer; the output end of the second fusion unit is connected to the input end of the fourth channel adjustment unit, and the output end of the fourth channel adjustment unit is connected to the input end of the fourth enhancement unit;

[0100] The IFPN layer also includes a first channel adjustment module, a second channel adjustment module and a third channel adjustment module;

[0101] The input end of the first channel adjustment module is connected to the output end of the first enhancement unit, and is used to perform channel adjustment on the first spliced ​​image to obtain a high-level feature map F1;

[0102] The input end of the second channel adjustment module is connected to the output end of the third enhancement unit, and is used to perform channel adjustment on the first fusion image to obtain a middle-level feature map F2;

[0103] The input end of the third channel adjustment module is connected to the output end of the fourth enhancement unit, and is used to perform channel adjustment on the first fusion image to obtain the bottom feature image F3.

[0104] Specifically, Figure 4 As shown, the SEGBlock module includes:

[0105] A fifth channel adjustment unit, a fifth enhancing unit, a third fusion unit, a sixth enhancing unit, a third splicing unit and a sixth channel adjustment unit; the output end of the fifth channel adjustment unit is connected to the input end of the fifth enhancing unit, and the output end of the fifth enhancing unit is connected to the input end of the third fusion unit; the output end of the third fusion unit is connected to the input end of the sixth enhancing unit; the output end of the sixth enhancing unit is connected to the input end of the third splicing unit, and the output end of the third splicing unit is connected to the input end of the sixth channel adjustment unit; the output end of the fifth channel adjustment unit is also connected to the input end of the third fusion unit, and the output end of the fifth enhancing unit is also connected to the input end of the third splicing unit.

[0106] Among them, the first upsampling unit and the second upsampling unit are preferably transposed convolution layers ConvTranspose, the first splicing unit, the second splicing unit, and the third splicing unit are preferably Concatenate operation units, the first fusion unit, the second fusion unit, and the third fusion unit are preferably ADD operation units, the first enhancement unit, the second enhancement unit, the third enhancement unit, the fourth enhancement unit, the fifth enhancement unit, and the sixth enhancement unit are preferably IEFBlock modules, and the first channel adjustment unit, the second channel adjustment unit, the third channel adjustment unit, the fourth channel adjustment unit, the fifth channel adjustment unit, and the sixth channel adjustment unit are preferably 1×1 convolution layers;

[0107] Specifically, IFPN is an efficient multi-scale feature extraction and fusion network. The network achieves efficient integration of multi-scale information through step-by-step feature enhancement, connection and fusion strategies, ensuring that feature maps at different levels can interact with each other with high quality. Finally, the network outputs three feature maps F1, F2 and F3, providing rich multi-level feature expressions for subsequent tasks (such as classification, segmentation or detection). The specific workflow is as follows:

[0108] 1. Input data and initial feature extraction

[0109] The IFPN network starts with three input paths P1, P2 and P3, each of which corresponds to an initial feature map of a different scale. This multi-input design ensures that the network can cover spatial details and semantic information of different scales, laying the foundation for subsequent multi-scale feature fusion.

[0110] 2. IEFBlock module and multi-scale feature enhancement

[0111] IFPN (Iterative Feature Pyramid Network) is an efficient multi-scale feature extraction and fusion network. The network ensures high-quality information interaction between feature maps of different scales and realizes cross-level feature fusion through step-by-step feature enhancement, connection and fusion strategies. Finally, the network outputs three feature maps F3, F2 and F1, providing comprehensive multi-scale feature expression for subsequent tasks (such as classification, detection and segmentation).

[0112] P3 path: initial feature extraction and IEFBlock enhancement

[0113] The P3 path is the lowest layer of the network, which mainly processes feature maps with low spatial resolution and large receptive field. The input feature map is first downsampled through a 3×3 convolution layer with a stride of 2 to generate a feature map with the lowest resolution, thereby capturing deep semantic information. Subsequently, the feature map enters the IEFBlock module, which performs feature reorganization and context information fusion through parallel convolution branches (such as 3×3 convolution and 1×1 convolution), effectively extracting local and global information and enhancing feature expression capabilities. The feature map processed by IEFBlock is channel-adjusted through a 1×1 convolution layer, and finally outputs the underlying feature map F3. F3 has a large receptive field, contains rich deep semantic features, and is suitable for capturing global information.

[0114] P2 path: mid-level feature fusion and IEFBlock enhancement

[0115] The P2 path is built on the basis of the P3 path and mainly processes feature maps with medium resolution. First, the feature map F3 output by the P3 path is restored to the spatial resolution of the P2 path through upsampling operations (such as the transposed convolution layer ConvTranspose). The upsampled F3 is concatenated with the initial feature map of the P2 path to combine feature information at different levels. The concatenated feature map enters the IEFBlock module, which further performs feature enhancement and multi-scale information extraction, and effectively integrates context information and local details through multi-branch convolution and feature fusion mechanisms. Subsequently, the feature map processed by IEFBlock is added to the jump connection feature map from the P3 path through the ADD fusion operation to complete the efficient transmission and fusion of information. Finally, the channel is adjusted through the 1×1 convolution layer to output the middle-level feature map F2. F2 combines deep semantic information with medium-scale detail features, and is suitable for capturing feature expressions of medium receptive fields.

[0116] P1 path: high-level feature fusion and IEFBlock enhancement

[0117] The P1 path is the highest layer of the network, which mainly processes feature maps with higher resolution and retains more spatial detail information. The feature map F2 output by the P2 path is first restored to the spatial resolution of the P1 path through an upsampling operation (such as the transposed convolution layer ConvTranspose). The upsampled F2 is concatenated with the initial input feature map of the P1 path to effectively fuse the high-resolution spatial details with the mid-level semantic information. Subsequently, the concatenated feature map enters the IEFBlock module for feature enhancement and reorganization. The IEFBlock effectively integrates contextual information with high-level local details through multi-scale convolution and fusion strategies. The fused feature map is added to the jump connection feature map from the P2 path through the ADD operation to ensure efficient cross-layer information flow and feature sharing. Finally, after channel adjustment through a 1×1 convolution layer, the high-level feature map F1 is output. F1 has the highest resolution, retains rich detail information, and fuses deep semantic features from the P2 and P3 paths, making it suitable for fine feature extraction and positioning tasks.

[0118] Multi-level feature output summary

[0119] Through the step-by-step feature fusion and enhancement of the P3 → P2 → P1 path, the IFPN network finally outputs three feature maps of different scales: the bottom feature map F3 (hereinafter referred to as F3), the middle feature map F2 (hereinafter referred to as F2) and the high-level feature map F1 (hereinafter referred to as F1). Among them, F3 has the lowest resolution and the deepest semantic information, with a large receptive field, suitable for capturing global features; F2 combines deep semantics with medium-scale contextual information, with medium resolution and receptive field; F1 has the highest resolution, contains rich spatial detail information, and combines middle-level and deep semantic features, suitable for tasks that require high-precision detail expression.

[0120] The IFPN network achieves efficient multi-scale feature integration through a bottom-up step-by-step feature fusion strategy. As the core component of the network, the IEFBlock module effectively extracts and enhances local and global information at different levels through multi-scale convolution branches and fusion mechanisms. In addition, combined with the Concatenate operation, ADD fusion and cross-layer jump connection, IFPN ensures the efficient flow and sharing of feature information. Finally, the three feature maps F3, F2 and F1 output by the network provide powerful multi-scale feature support for tasks such as classification, detection, and segmentation, with both deep semantics and fine detail information.

[0121] Specifically, the SEGBlock module is the output module of the FlotationFoamNet network, which is mainly used to generate the positioning result (rect) and segmentation result (mask) of the target area. As the key output part of the entire network, SEGBlock accurately extracts features and outputs the final result through feature enhancement, cross-layer feature fusion and multi-scale information integration.

[0122] First, the input feature map is preliminarily processed by a 1×1 convolution layer. The function of this convolution layer is to adjust the number of channels of the feature map while retaining the spatial resolution of the input features. This step effectively removes redundant information and improves the network's ability to express target features. The preliminarily processed feature map is sent to the IEFBlock module, which is the first feature enhancement module in SEGBlock. IEFBlock effectively extracts local details and contextual semantic information through multi-branch convolution (such as 1×1 and 3×3 convolution) and fusion operations, further optimizing the input features.

[0123] Then, the output features of IEFBlock are fused with the initial input features through jump connections through the ADD operation. Jump connections can establish a direct connection between deep features and shallow features, retain shallow detail information, and introduce deep semantic information extracted by IEFBlock, thereby enhancing feature expression capabilities. This fusion mechanism enables the network to take into account both accurate edge information and semantic understanding when performing region positioning and segmentation.

[0124] Subsequently, the fused feature map enters the second IEFBlock module for further processing. The IEFBlock integrates and reorganizes multi-scale features through deep feature extraction and optimization operations, further enhancing the network's perception of the target area, especially the recognition of complex edges and details. After being processed by IEFBlock, the feature map is sent to the Concatenate operation. In this stage, the output features from the second IEFBlock are concatenated with the initial input features of the jump connection, ensuring the effective fusion of shallow feature details and deep semantic features.

[0125] Finally, the concatenated feature map is processed through a 1×1 convolutional layer to adjust the number of channels and output. This convolutional layer is responsible for generating the final network output, including the location result (rect) and segmentation result (mask) of the target area. Through this structure, the SEGBlock module achieves efficient fusion and expression of multi-level features, ensuring the accuracy of regional location and segmentation.

[0126] In summary, the SEGBlock module completes the multi-scale integration and optimization of features through a series of structured steps, including initial 1×1 convolution processing, IEFBlock module feature enhancement, ADD operation of jump connection, Concatenate operation of feature splicing, and the final 1×1 convolution output. While retaining shallow detail information, this module combines deep semantic features to generate accurate target area positioning and segmentation results, providing strong support for the FlotationFoamNet network in flotation foam anomaly classification and size measurement tasks.

[0127] In the preferred solution, the final output of the anomaly classification model includes the following: CLS: classification results of flotation foam (anomaly detection), used to identify whether there is an anomaly in the foam state. Three rect+masks: rectangular frame positioning and segmentation masks of the foam area at three different scales, used to accurately measure the foam size and spatial position.

[0128] FlotationFoamNet achieves the abnormal classification and size measurement tasks of flotation foam through a series of efficient modules. First, the network extracts multi-scale features through step-by-step downsampling and IEFBlock modules to ensure that the spatial resolution and semantic information of feature maps of different depths are effectively expressed. Subsequently, after processing by SPPF (spatial pyramid pooling module), multi-scale features are further aggregated to enhance the network's ability to understand the global context. Next, cross-level feature fusion is performed through IFPN (integrated feature pyramid network), which effectively combines shallow spatial details with deep semantic information, and improves the network's detection and segmentation accuracy of foam areas of different scales.

[0129] Finally, the CLS classification head outputs the classification results of flotation foam to achieve accurate identification of abnormal states; the SEGBBlock module outputs the rect+mask of three regions on feature maps of different scales to ensure accurate measurement and positioning of foam size. Overall, the FlotationFoamNet structure efficiently combines local detail features with global semantic information, while ensuring classification accuracy, and achieves high-precision detection and segmentation of foam regions at multiple scales, with strong robustness and practical value.

[0130] Second, collect foam images of different states under different flotation conditions to construct a training set;

[0131] Data collection is the basis for training high-performance models. It is necessary to ensure the diversity and representativeness of the data to cover all situations that may occur in actual application scenarios. In the flotation foam anomaly classification and size measurement tasks, a large number of high-quality flotation foam images need to be collected, including:

[0132] Multiple foam states: Collect foam state images under different flotation conditions, such as normal foam, abnormal foam (such as excessive rupture, foam accumulation, excessive foaming, etc.).

[0133] Different foam sizes: Covers foam samples of different sizes, including tiny foam, larger foam, and mixed-size foam, ensuring that the model can adapt to multi-scale target detection and measurement needs.

[0134] Diversified collection conditions: Data collection needs to take into account different lighting conditions, shooting angles, and background complexity, simulate lighting changes, noise interference, and other problems that may occur in industrial environments, and increase data diversity.

[0135] Abnormal category coverage: Collect bubble abnormal category data to ensure that each abnormal category (such as bubble breakage, non-uniform distribution, etc.) has enough samples in the dataset to avoid underfitting of the model for categories with few samples.

[0136] Data annotation:

[0137] Data labeling is an important part of training supervised learning models to ensure that the network can learn accurate classification, positioning and segmentation information. The specific steps include:

[0138] Abnormal classification label (CLS): An abnormal category label is annotated for each image or foam area. The classification label can be "normal foam" or various abnormal foams (such as excessive bursting, foam accumulation, etc.). This provides supervision signals for the network's classification task.

[0139] Region positioning box (rect): The boundary information of the foam target is marked with a rectangular box, and the position and size of the box are defined (coordinate format such as x, y, width, height). This annotation form helps the network learn the precise positioning of the target during training.

[0140] Segmentation mask: Generate a binary segmentation mask for each foam area. The foreground pixels in the mask represent the foam area, and the background pixels represent the non-foam area. The segmentation mask provides pixel-level supervision to help the network learn more detailed boundary and region information.

[0141] Labeling tools and standards: Use professional labeling tools (such as LabelImg, CVAT, or VGG Image Annotator) to ensure the consistency and accuracy of data labeling. At the same time, establish data labeling standards and standardize the labeling format and naming rules.

[0142] 3. Using the training set to train the flotation foam anomaly classification model

[0143] When designing the loss function in the FlotationFoamNet network, we need to consider multi-task learning (classification, localization, and segmentation), so we can jointly optimize the performance of the network by weighted combination of multiple loss functions.

[0144] 1. Classification Loss (CLS)

[0145] The classification loss is mainly used to guide the network to classify foam anomalies. In order to classify the foam type (such as whether there is abnormal foam), the cross-entropy loss function can be used to measure the gap between the predicted result and the true label.

[0146] Cross entropy loss formula:

[0147] ,in, is the actual category label, is the probability of the predicted category, is the number of samples.

[0148] 2. Region positioning loss (rect)

[0149] The region localization loss is used to ensure that the network can accurately locate the bounding box of the foam region. In this task, L1 loss or Smooth L1 loss are common choices. L1 loss is simple and direct, while Smooth L1 loss is more stable when the bounding box error is small.

[0150] L1 loss formula:

[0151] ,in, are the predicted bounding box coordinates, are the actual bounding box coordinates.

[0152] Smooth L1 loss formula:

[0153] in, .

[0154] 3. Segmentation loss (mask)

[0155] The segmentation loss is used to optimize the accuracy of the mask area to ensure that the network can accurately predict the pixels in the foam area. Commonly used segmentation loss functions are Dice loss and cross entropy loss.

[0156] (1) Dice loss

[0157] Dice loss is a commonly used indicator for binary segmentation, especially in unbalanced data sets. Its formula is:

[0158] in, is the predicted area, It’s a real area.

[0159] (2) Cross Entropy Loss

[0160] Cross entropy loss can also be applied in segmentation tasks, especially in pixel-level classification tasks:

[0161] ,in, is the predicted pixel probability, is the actual pixel label.

[0162] 4. Total loss function

[0163] Finally, the total loss function of the network can be optimized by weighted combination of the above loss functions in a multi-task learning manner. Assuming that the loss function of each task has a weight coefficient, the total loss function is defined as:

[0164] ;

[0165] in, , , are hyperparameters that control the importance of classification, localization, and segmentation losses in the total loss, respectively. These weights can be adjusted through cross-validation or hyperparameter search.

[0166] 5. Pre-trained model initialization

[0167] In order to improve the convergence efficiency of the network, especially in the feature extraction part, it is a common practice to use a pre-trained model (such as the feature extraction part from ImageNet) to initialize the basic network layer. This can help the model have better representation capabilities from the beginning, especially when processing image data, the pre-trained weights can capture useful low-level features (such as edges, textures, etc.).

[0168] In this way, FlotationFoamNet can get better initial values ​​at the beginning of training, thereby speeding up the training process, reducing oscillations during training, and improving convergence speed.

[0169] By weighted combination of classification, localization and segmentation loss functions, we are able to optimize FlotationFoamNet on all three tasks simultaneously. The initialization method of the pre-trained model further improves the performance and training efficiency of the network.

[0170] 4. Training Process

[0171] (1) Forward propagation

[0172] In the forward propagation process of FlotationFoamNet, the input image is first sent to the network. The image undergoes preliminary feature extraction through the downsampling layer, and then enters the IEFBlock module for feature enhancement and further extraction. Next, after the multi-scale feature fusion stage, the network can effectively fuse feature maps of different scales to enhance the processing capability of multi-scale information. Then, it enters the IFPN (Image Feature Pyramid Network) module to further strengthen the fusion of features at different levels and improve the semantic understanding and positioning accuracy of the image. Finally, through the SEGBlock module, the network outputs the classification result (CLS), the regional positioning result (rect), and the segmentation mask (mask).

[0173] (2) Loss function calculation

[0174] After obtaining the network output in the forward propagation, the loss functions of the classification, positioning and segmentation tasks are calculated respectively by comparing with the true annotation value.

[0175] (3) Back propagation

[0176] The back propagation algorithm is used to calculate the gradients of all parameters in the network. Through the chain rule, the partial derivatives of the loss function with respect to each network parameter are calculated layer by layer to obtain gradient information. This gradient information will be used to optimize the weights of the network, so that the network can gradually reduce the value of the loss function in each iteration, thereby improving the performance of the model.

[0177] (4) Optimizer update

[0178] After back propagation calculates the gradient, the optimizer Adam is used to update the network parameters based on the gradient information. The Adam optimizer combines the momentum method and adaptive learning rate to effectively accelerate the convergence process and avoid the problem of gradient explosion or disappearance. The SGD optimizer updates the network parameters through small batches of data and performs well when processing large-scale data. Through these optimizers, the network parameters are gradually optimized to minimize the total loss function and improve the model performance.

[0179] (5) Learning rate strategy

[0180] In order to ensure the stability and convergence of the training process, a dynamic learning rate adjustment strategy is adopted. Common strategies include cosine annealing and learning rate decay mechanisms. Cosine annealing simulates the changes of the cosine function to make the learning rate larger in the early stage of training to accelerate training, and gradually decreases it in the later stage of training, thereby improving the fine-tuning ability of the model. Learning rate decay regularly reduces the learning rate, which helps avoid oscillations during training and ensures that the network converges stably when approaching the optimal solution.

[0181] (6) Model verification and adjustment

[0182] After each training cycle (epoch), the model is evaluated using the validation set to monitor the model's performance in classification, localization, and segmentation tasks. Specifically, the evaluation indicators include:

[0183] Classification accuracy: Evaluate the performance of the model on the classification task by calculating the match between the model's predicted categories and the true labels.

[0184] Positioning accuracy: The accuracy of the model in the target positioning task is evaluated by calculating the difference between the predicted bounding box and the true bounding box. The commonly used evaluation metric is IoU (Intersection over Union), which is used to measure the overlap between the predicted box and the true box.

[0185] Segmentation IoU: Calculate the IoU value between the predicted segmentation mask and the true mask to evaluate the accuracy of the model in the image segmentation task.

[0186] Through these evaluation indicators, the training progress of the model can be monitored in real time, and the model can be tuned according to the results of the validation set to avoid overfitting and improve the generalization ability of the model.

[0187] 6. Testing and performance evaluation

[0188] After training is completed, the FlotationFoamNet model is finally evaluated using the test set, focusing on the following key performance indicators:

[0189] Classification accuracy: This metric is used to evaluate the performance of the model in the foam anomaly classification task. It measures the classification accuracy by calculating the match between the network predicted category and the true label. A high accuracy indicates that the model can effectively distinguish different types of foam anomalies.

[0190] Positioning accuracy: The positioning ability of the model is evaluated by calculating the difference between the predicted regional positioning box and the real annotation box. Positioning accuracy reflects the ability of the model to accurately predict the location and size of the foam anomaly area. Generally, IoU (Intersection over Union) is used as a measurement standard. The higher the IoU value, the greater the overlap between the predicted box and the real box, and the more accurate the positioning.

[0191] Segmentation performance (IoU and Dice coefficient): This metric evaluates the degree of overlap between the predicted mask area and the true mask. IoU (Intersection over Union) and Dice coefficient are commonly used evaluation metrics for image segmentation tasks, measuring the ratio of the intersection to the union between the predicted area and the true area, and the similarity between the predicted area and the true area, respectively. Higher IoU and Dice coefficients indicate that the model has higher accuracy in segmentation tasks.

[0192] Through these evaluation indicators, the actual application performance of the FlotationFoamNet model can be comprehensively measured. According to the evaluation results of the test set, if it is found that the model does not perform well in certain tasks (such as classification, positioning or segmentation), the model structure can be further adjusted or the training strategy can be optimized to ensure stable and high-precision results in practical applications.

[0193] 4. The trained anomaly classification model sorts the target foam image.

[0194] After completing model quantization and optimization, the FlotationFoamNet model is deployed to the target platform, such as industrial equipment or edge computing equipment. The deployment process includes:

[0195] Platform adaptation: Adapt according to the hardware configuration and software environment of the target platform to ensure that the model can run efficiently on the specified hardware (such as embedded devices, GPU, TPU, etc.).

[0196] Real-time detection and analysis: The deployed FlotationFoamNet model will enable real-time classification and size measurement of flotation foam anomalies. Through real-time video stream or image input, the model can quickly determine the type of foam anomaly, locate the abnormal area and calculate its size, providing real-time feedback for industrial process control.

[0197] The present invention has the following characteristics:

[0198] The present invention combines multi-scale feature extraction and fusion technology to perform multi-level and multi-angle feature analysis on the flotation foam image, thus overcoming the limitation of traditional methods that only rely on single-scale features. Foam features of different scales are effectively fused, and can handle foam samples with inconsistent sizes and diverse shapes. In the prior art, most methods only extract features at a single scale and fail to fully utilize multi-scale information, resulting in low detection accuracy in complex flotation foam environments. Therefore, the use of multi-scale feature fusion technology makes the performance of the present invention in foam anomaly classification and size measurement more accurate and robust.

[0199] The present invention innovatively combines deep convolutional neural networks (CNNs) with multi-scale feature fusion to achieve efficient feature extraction and classification of flotation foam. Traditional methods mostly rely on manual feature selection or shallow learning models, while the present invention can automatically extract and optimize the key features of foam through deep learning models to achieve high-precision abnormal classification. The hierarchical learning ability of deep CNN enables the present invention to cope with the complexity and variability of foam images in the flotation process to a greater extent.

[0200] The present invention combines the two tasks of foam anomaly classification and size measurement for the first time. Through multi-scale feature fusion, the system can not only identify whether there is anomaly in flotation foam, but also accurately measure the size of the foam, providing a more comprehensive foam monitoring solution. In the prior art, most methods only focus on a single task of foam anomaly detection or size measurement, and fail to combine these two important tasks, resulting in poor results in practical applications.

[0201] In the image processing of flotation foam, background noise and interference are often the key factors affecting the detection accuracy. In order to improve the detection accuracy, the present invention proposes an adaptive background interference suppression method, which effectively filters out the background noise by combining multi-scale features. Traditional methods are usually unable to adapt well to dynamic changes in the flotation process, such as illumination changes, bubble overlap and other interferences, resulting in low classification and measurement accuracy. The background suppression method of the present invention can effectively deal with these dynamic interferences and improve the accuracy and reliability of flotation foam detection.

[0202] The present invention combines advanced deep learning technology and multi-scale feature fusion to make real-time online detection and processing possible in complex flotation environments. Compared with traditional rule-based detection methods, the present invention can achieve real-time abnormal classification and size measurement of flotation foam, meeting the requirements of high real-time and high efficiency of flotation foam monitoring systems in industrial applications. This feature is not common in the prior art, and traditional methods often face problems of processing delay and large amount of calculation.

[0203] The present invention deeply explores the diversity and hierarchical characteristics of flotation foam through multi-scale feature fusion, which can not only accurately describe the surface characteristics of the foam such as shape and size, but also analyze the abnormal conditions that may occur in the flotation process from a deeper level, such as bubble gaps, density distribution, etc. These hierarchical features can significantly improve the accuracy of anomaly detection, especially under complex flotation conditions, where traditional methods cannot effectively distinguish these subtle changes.

[0204] In summary, the innovation of the present invention lies in the technical application of combining multi-scale feature fusion, deep convolutional neural network and foam feature extraction. It has high efficiency, accuracy and adaptability that traditional methods do not have, especially in the abnormal classification and size measurement of flotation foam, showing significant technical advantages.

[0205] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A flotation foam anomaly classification method based on multi-scale feature fusion, characterized in that: The following steps are involved: Construct an abnormal classification model for flotation foam, the abnormal classification model includes an input layer, a multi-scale feature extraction layer, a fully connected layer and an output layer connected in sequence; the multi-scale feature extraction layer includes a plurality of feature extraction sublayers and an SPPF module connected in sequence; each feature extraction sublayer is composed of a downsampling module and an IEFBlock module, the output end of each downsampling module is connected to the input end of the IEFBlock module of the same feature extraction sublayer, the input end of the downsampling module at the head end of the multi-scale feature extraction layer is connected to the input layer, and the output end of the IEFBlock module at the end of the multi-scale feature extraction layer is connected to the input end of the SPPF module; Collect foam images of different states under different flotation conditions to construct a training set; the states of the foam images include at least normal and abnormal; use the training set to train an abnormal classification model for the flotation foam, and use the trained abnormal classification model to sort target foam images; The IEFBlock module includes: a spilt segmentation module, a standard convolution module, a depth-separable convolution module, a splicing module A, an enhancement module and a splicing module B; The output end of the spilt segmentation module is connected to the input end of the standard convolution module, the depthwise separable convolution module and the splicing module B respectively, the output end of the standard convolution module and the depthwise separable convolution module is connected to the input end of the splicing module A, the output end of the splicing module A is connected to the input end of the enhancement module, and the output end of the enhancement module is connected to the input end of the splicing module B; The spilt segmentation module is used to separate the input feature map by channel to obtain a first separation feature, a second separation feature and a third separation feature; the standard convolution module is used to extract the local core feature of the first separation feature; the depth-separable convolution module is used to extract the high-resolution feature of the second separation feature; the splicing module A is used to splice the local core feature with the high-resolution feature to obtain a first splicing feature; the enhancement module is used to enhance the first splicing feature; the splicing module B is used to splice the enhanced first splicing feature with the third separation feature.

2. The flotation foam anomaly classification method based on multi-scale feature fusion according to claim 1 is characterized in that: The output layer includes an FCN module and a CLS module, the input end of the FCN module is connected to the output end of the fully connected layer, the output end of the FCN module is connected to the input end of the CLS module, and the CLS module is used for anomaly classification.

3. The flotation foam anomaly classification method based on multi-scale feature fusion according to claim 1 is characterized in that: It also includes a segmentation and positioning module, which includes an IFPN layer and multiple SEGBlock modules. The input end of the IFPN layer is respectively connected to the output ends of the multiple IEFBlock modules and the SPPF module of the multi-scale feature extraction layer, and the output end of the IFPN layer is respectively connected to the input ends of the multiple SEGBlock modules.

4. The flotation foam anomaly classification method based on multi-scale feature fusion according to claim 3 is characterized in that: The feature extraction sublayer includes a first feature extraction sublayer, a second feature extraction sublayer, a third feature extraction sublayer and a fourth feature extraction sublayer connected in sequence, and the output ends of the IEFBlock module and the SPPF module of the second feature extraction sublayer and the third feature extraction sublayer are connected to the input end of the IFPN layer.

5. The flotation foam anomaly classification method based on multi-scale feature fusion according to claim 4 is characterized in that: The IFPN layer includes: a first splicing module, a second splicing module, a first fusion module, and a second fusion module; The input end of the first splicing module is respectively connected to the output end of the first feature extraction sublayer and the output end of the second splicing module, and the output end of the first splicing module is connected to the input end of the first fusion module; the first splicing module is used to splice the first feature map output by the first feature extraction sublayer with the second splicing map output by the second splicing module to obtain a first splicing map; The input end of the first fusion module is also jump-connected to the output end of the first feature extraction sublayer, the input end of the first fusion module is also connected to the output end of the second splicing module, and the output end of the first fusion module is also connected to the input end of the second fusion module; the first fusion module is used to fuse the first splicing image, the second splicing image output by the second splicing module, and the jump-connected feature image of the first feature image, and output a first fusion image; The input end of the second splicing module is connected to the output ends of the second feature extraction sublayer and the third feature extraction sublayer, and is used to splice the second feature map output by the second feature extraction sublayer and the third feature map output by the third feature extraction sublayer to obtain a second splicing map; The input end of the second fusion module is also jump-connected to the output end of the second feature extraction sublayer, and the input end of the second fusion module is also connected to the output end of the third feature extraction sublayer; the second fusion module is used to fuse the first fusion image, the second splicing image output by the second splicing module, and the jump-connected feature image of the first feature image, and output a second fusion image.

6. The flotation foam anomaly classification method based on multi-scale feature fusion according to claim 5 is characterized in that: The first splicing module includes: A first upsampling unit, a first splicing unit and a first enhancement unit, wherein an input end of the first upsampling unit is connected to an output end of the second splicing module, an output end of the first upsampling unit is connected to an input end of the first splicing unit, and an input end of the first splicing unit is also connected to an output end of the first feature extraction sublayer; an output end of the first splicing unit is connected to an input end of the first enhancement unit, and an output end of the first enhancement unit is connected to an input end of the first fusion module; and / or The second splicing module includes: A second upsampling unit, a second splicing unit and a second enhancing unit, wherein the input end of the second upsampling unit is connected to the output end of the third feature extraction sublayer, the output end of the second upsampling unit is connected to the input end of the second splicing unit, and the input end of the second splicing unit is also connected to the output end of the second feature extraction sublayer; the output end of the second splicing unit is connected to the input end of the second enhancing unit, and the output end of the second enhancing unit is connected to the input ends of the first splicing module and the first fusion module.

7. The flotation foam anomaly classification method based on multi-scale feature fusion according to claim 6 is characterized in that: The first fusion module includes: A first fusion unit, a third channel adjustment unit and a third enhancement unit, wherein the input end of the first fusion unit is jump-connected to the output end of the first feature extraction sublayer, and the input end of the first fusion unit is also connected to the output ends of the first splicing module and the second splicing module; the output end of the first fusion unit is connected to the input end of the third channel adjustment unit, the output end of the third channel adjustment unit is connected to the input end of the third enhancement unit, and the output end of the third enhancement unit is connected to the input end of the second fusion module; and / or The second fusion module includes: a second fusion unit, a fourth channel adjustment unit, and a fourth enhancement unit, wherein the input end of the second fusion unit is jump-connected to the output end of the second feature extraction sublayer, and the input end of the second fusion unit is also connected to the output ends of the first fusion module and the third feature sublayer; the output end of the second fusion unit is connected to the input end of the fourth channel adjustment unit, and the output end of the fourth channel adjustment unit is connected to the input end of the fourth enhancement unit; and / or The IFPN layer also includes a first channel adjustment module, a second channel adjustment module and a third channel adjustment module; The input end of the first channel adjustment module is connected to the output end of the first enhancement unit, and is used to perform channel adjustment on the first spliced ​​image to obtain a high-level feature map F1; The input end of the second channel adjustment module is connected to the output end of the third enhancement unit, and is used to perform channel adjustment on the first fusion image to obtain a middle-level feature map F2; The input end of the third channel adjustment module is connected to the output end of the fourth enhancement unit, and is used to perform channel adjustment on the first fusion image to obtain the bottom feature image F3.

8. The flotation froth anomaly classification method based on multi-scale feature fusion according to claim 7 is characterized in that: The SEGBlock module includes: A fifth channel adjustment unit, a fifth enhancing unit, a third fusion unit, a sixth enhancing unit, a third splicing unit and a sixth channel adjustment unit; the output end of the fifth channel adjustment unit is connected to the input end of the fifth enhancing unit, and the output end of the fifth enhancing unit is connected to the input end of the third fusion unit; the output end of the third fusion unit is connected to the input end of the sixth enhancing unit; the output end of the sixth enhancing unit is connected to the input end of the third splicing unit, and the output end of the third splicing unit is connected to the input end of the sixth channel adjustment unit; the output end of the fifth channel adjustment unit is also connected to the input end of the third fusion unit, and the output end of the fifth enhancing unit is also connected to the input end of the third splicing unit.

9. The flotation froth anomaly classification method based on multi-scale feature fusion according to claim 8, characterized in that: The loss function of the anomaly classification model satisfies: ; ; ; ; Among them, CLS Loss represents classification loss, rect Loss represents region positioning loss, and Mask Loss represents mask loss. , , are hyperparameters that control the importance of classification, localization, and segmentation losses in the total loss respectively; is the actual category label, is the probability of the predicted category, is the number of samples; are the predicted bounding box coordinates, are the actual bounding box coordinates, is the predicted pixel probability, is the actual pixel label.

10. A computer system comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Cross-scale target real-time detection method and device for large-scene image and medium

    CN117496127A

  • Copper-molybdenum ore froth flotation segmentation method and system

    CN117593747A