Industrial pellet image segmentation method and device, medium and product

Through the lightweight dual-branch encoder and an image segmentation network with adaptive weight feature fusion strategy, the problem of insufficient image segmentation accuracy and speed of industrial pellets in the prior art is solved, and a fast and effective image segmentation effect is achieved.

CN120451539APending Publication Date: 2025-08-08YUNNAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510528748.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

When the existing industrial pellet image segmentation method faces complex mineral background interference, non-uniform lighting conditions and heterogeneous particle morphology diversity, the segmentation accuracy is insufficient and the processing speed is insufficient in a dynamic production environment, making it difficult to quickly and effectively segment industrial pellet images.

Method used

Image segmentation is performed using a lightweight dual-branch encoder. The first branch encoder uses a lightweight convolutional neural network for feature extraction. The second branch encoder extracts low-frequency information through wavelet transformation and combines the improved lightweight ResNet18 network for deep feature extraction. The decoder integrates deep and low-frequency features through an adaptive weight feature fusion strategy to build an image segmentation network.

Benefits of technology

While ensuring computing efficiency, the segmentation accuracy and robustness of the image segmentation network for dense pellet particles is improved, and industrial pellet images can be quickly and effectively segmented.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451539A_ABST
    Figure CN120451539A_ABST
Patent Text Reader

Abstract

The invention discloses an industrial pellet image segmentation method and device, a medium and a product, and relates to the field of industrial pellet image segmentation, and the method comprises the steps: constructing an image segmentation network; wherein the encoder adopts a lightweight double-branch encoder comprising a first branch encoder and a second branch encoder; the first branch encoder adopts a lightweight convolutional neural network to carry out feature extraction; after a second branch encoder extracts low-frequency information of the industrial pellet particle image by using wavelet transform, deep feature extraction is carried out by using an improved lightweight ResNet18 network; the decoder performs feature fusion on features extracted by the second branch encoder and the first branch encoder through an adaptive weight feature fusion strategy; training and optimizing the image segmentation network by using an open-source industrial pellet particle image data set; and inputting a to-be-segmented industrial pellet particle image into the optimized image segmentation network to obtain a segmented result image. According to the method, the industrial pellet image can be quickly and effectively segmented.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of industrial pellet image segmentation, and in particular to an industrial pellet image segmentation method, equipment, medium and product. Background Art

[0002] Particles are ubiquitous in nature and industry, playing a vital role. Natural particles come in a wide variety of forms, including sand, soil, ice, and snow. They are widely distributed across diverse ecosystems and have a profound impact on natural processes and geological activities. For example, sand is constantly eroded and transported by weathering, while soil particles are formed into diverse topographic structures by water and wind. Particles are also widely used in industry. Industrial particles typically include pellets, ores, pulverized coal, and fertilizers. These particles are not only components of raw materials but also play an important role in processing, transportation, and storage. For example, the particle component in ores is an important source for extracting non-ferrous metals, while fertilizer particles are used to improve soil and optimize crop growth. In short, as an important form of matter, the application of particles in various fields not only affects production efficiency and product quality, but also profoundly influences our lifestyle and environment.

[0003] Given that particles are one of the most ubiquitous substances in the world, measuring parameters related to their morphology is crucial for many industrial processes, such as fertilizer production from diammonium phosphate and pelletization from iron ore processing. These parameters include particle count measurement, shape analysis, particle size distribution, and color and texture evaluation. In the iron ore pelletization process, particle size distribution testing is a crucial step. Traditional particle size testing methods rely on the subjective judgment of on-site personnel, providing feedback on particle size trends to control personnel, who then make operational adjustments based on this feedback. However, this approach not only increases the workload of personnel, but also, due to the subjectivity and lag in information feedback, often leads to operational errors, resulting in a large number of substandard particles and potential losses for the company.

[0004] With the rapid development of artificial intelligence, the focus of existing segmentation techniques for industrial pellet images (i.e., images of industrial pellet particles acquired during iron ore processing) has shifted from manual methods to machine vision. Machine vision captures pellet images using high-speed cameras and image acquisition systems. Image processing techniques are then used to analyze and process the images, ultimately accurately segmenting the pellet particles contained within each image. Segmentation accuracy directly impacts the reliability of particle size detection. However, due to the varying pellet sizes, shapes, overlaps, and powder interference, the segmentation of industrial pellet images is challenging. Researchers have proposed various pellet segmentation methods, which can be broadly categorized into two categories: traditional morphological methods and deep learning methods. Traditional morphological methods rely on manually defined features, such as particle size, morphology, and edge recognition, and employ specific algorithms (e.g., threshold segmentation and edge detection) to perform pellet segmentation. While these methods can handle simple pellet segmentation tasks to a certain extent, their performance is limited by three major challenges: complex mineral background interference, non-uniform illumination conditions, and the diverse morphology of heterogeneous particles. Furthermore, they rely on empirically driven feature engineering. The current deep learning-based industrial pellet image segmentation method uses a convolutional neural network to build an end-to-end architecture, and achieves a coupled representation of the apparent characteristics of pellet particles and contextual information through the collaborative extraction of multi-level visual features (edges, textures) and high-order semantic features (shapes, spatial associations). Compared with traditional methods, this data-driven feature extraction mechanism greatly improves the robustness of segmentation under complex working conditions. However, the existing deep learning-based pellet image segmentation method mainly optimizes the model structure to extract effective features, but ignores the low-frequency information in the pellet image. The low-frequency information in the pellet particle image (i.e., pellet image) is very helpful for pellet segmentation, especially the segmentation of dense pellets. At the same time, due to the rapid changes in the distribution and state of pellets, the existing methods often face the problems of delay and insufficient processing speed, which limits their application in dynamic production environments. In summary, traditional morphological methods and existing deep learning-based pellet image segmentation methods are unable to quickly and effectively segment industrial pellet images. Summary of the Invention

[0005] The purpose of this application is to provide an industrial pellet image segmentation method, equipment, medium and product, which can quickly and effectively segment industrial pellet images.

[0006] To achieve the above objectives, this application provides the following solutions:

[0007] In a first aspect, the present application provides an industrial pellet image segmentation method, the industrial pellet image segmentation method comprising:

[0008] Obtain an open source industrial pellet particle image dataset; the open source industrial pellet particle image dataset includes multiple industrial pellet particle images and a pixel-level GT label corresponding to each industrial pellet particle image;

[0009] Construct an image segmentation network; the encoder in the image segmentation network adopts a lightweight dual-branch encoder; the lightweight dual-branch encoder includes a first branch encoder and a second branch encoder; the first branch encoder uses a lightweight convolutional neural network to extract features of the industrial pellet particle image input to the first branch encoder; the second branch encoder uses wavelet transform to perform multi-scale decomposition on the industrial pellet particle image input to the second branch encoder, extracts the low-frequency information of the industrial pellet particle image, and then adopts an improved lightweight ResNet18 network to extract deep features; the improved lightweight ResNet18 network includes a first convolutional layer, a first Bottleneck module, a first maximum pooling layer, a second Bottleneck module, a third Bottleneck module and a second convolutional layer connected in sequence; the decoder in the image segmentation network performs feature fusion on the deep features extracted by the second branch encoder and the features extracted by the first branch encoder through an adaptive weight feature fusion strategy;

[0010] Using the open source industrial pellet particle image dataset to train and optimize the image segmentation network to obtain an optimized image segmentation network;

[0011] The image of the industrial pellet particles to be segmented is input into the optimized image segmentation network, and the optimized image segmentation network is used to segment the boundaries of the industrial pellet particles to obtain a segmented result image.

[0012] Optionally, a dynamic adaptive soft threshold module is arranged between the first convolutional layer and the first Bottleneck module; the dynamic adaptive soft threshold module is connected to the first convolutional layer and the first Bottleneck module respectively; the dynamic adaptive soft threshold module is used to perform a noise reduction operation on the feature map extracted by the first convolutional layer using dynamic adaptive soft threshold processing; the feature map is obtained after the first convolutional layer performs initial feature extraction on the low-frequency information.

[0013] Optionally, the number of output channels of the first convolutional layer is 20, the number of output channels of the first Bottleneck module is 40, the number of output channels of the second Bottleneck module is 80, the number of output channels of the third Bottleneck module is 160, and the number of output channels of the second convolutional layer is 160;

[0014] The convolution kernel size of the first convolution layer is 7×7, and the convolution stride of the first convolution layer is 2; the first convolution layer uses BatchNorm2d normalization and ReLU activation; the convolution kernel size of the second convolution layer is 3×3, and the convolution stride of the second convolution layer is 2; the second convolution layer uses BatchNorm2d normalization and ReLU activation;

[0015] The pooling size of the first maximum pooling layer is 3×3, and the pooling stride of the first maximum pooling layer is 2;

[0016] The first Bottleneck module, the second Bottleneck module, and the third Bottleneck module all include a fusion operation of cubic convolution, BatchNorm2d normalization, and ReLU activation.

[0017] Optionally, the lightweight convolutional neural network includes a third convolutional layer, a second maximum pooling layer, a fourth convolutional layer, a third maximum pooling layer, a fifth convolutional layer and a sixth convolutional layer connected in sequence;

[0018] The number of output channels of the third convolutional layer is 20, the number of output channels of the fourth convolutional layer is 40, the number of output channels of the fifth convolutional layer is 80, and the number of output channels of the sixth convolutional layer is 80;

[0019] The convolution kernel sizes of the third convolution layer, the fourth convolution layer, and the fifth convolution layer are all 1×1, and the convolution strides of the third convolution layer, the fourth convolution layer, and the fifth convolution layer are all 3. The convolution kernel size of the sixth convolution layer is 1×1, and the convolution stride of the sixth convolution layer is 1;

[0020] The third convolutional layer, the fourth convolutional layer, the fifth convolutional layer and the sixth convolutional layer all adopt InstanceNorm2d normalization operation and LeakyReLU activation;

[0021] The pooling sizes of the second maximum pooling layer and the third maximum pooling layer are both 2×2, and the pooling strides of the second maximum pooling layer and the third maximum pooling layer are both 2.

[0022] Optionally, the decoder in the image segmentation network includes a first adaptive weight feature fusion module, a first upsampling layer, a first decoding layer, a second upsampling layer, a second decoding layer, a third upsampling layer, a third decoding layer, a second adaptive weight feature fusion module, a fourth upsampling layer and a convolution block;

[0023] The first adaptive weight feature fusion module is respectively connected to the third Bottleneck module and the first upsampling layer; the first upsampling layer, the first decoding layer, the second upsampling layer, the second decoding layer, the third upsampling layer, the third decoding layer, the second adaptive weight feature fusion module, the fourth upsampling layer and the convolution block are connected in sequence; the second convolution layer is connected to the first upsampling layer; the first Bottleneck module is connected to the third upsampling layer; the second Bottleneck module is connected to the second upsampling layer; the sixth convolution layer is connected to the second adaptive weight feature fusion module.

[0024] Optionally, the second branch encoder uses wavelet transform to perform multi-scale decomposition on the industrial pellet particle image input to the second branch encoder, extracting the low-frequency information of the industrial pellet particle image while also extracting the high-frequency information of the industrial pellet particle image; the high-frequency information is input into the first adaptive weight feature fusion module; the first adaptive weight feature fusion module is used to perform feature fusion on the high-frequency information and the output of the third Bottleneck module through an adaptive weight feature fusion strategy.

[0025] Optionally, the second adaptive weight feature fusion module is used to perform feature fusion on the output of the third decoding layer and the output of the sixth convolutional layer through an adaptive weight feature fusion strategy.

[0026] In a second aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement any one of the above-described industrial pellet image segmentation methods.

[0027] In a third aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described industrial pellet image segmentation methods.

[0028] In a fourth aspect, the present application provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the above-described industrial pellet image segmentation methods.

[0029] According to the specific embodiments provided in this application, this application has the following technical effects:

[0030] The present application provides an industrial pellet image segmentation method, device, medium and product, which adopts a lightweight dual-branch encoder as an encoder in an image segmentation network. The first branch encoder in the lightweight dual-branch encoder uses a lightweight convolutional neural network to extract features from the input image. The second branch encoder in the lightweight dual-branch encoder uses a wavelet transform to perform multi-scale decomposition on the input image. After extracting low-frequency information, an improved lightweight ResNet18 network including a first convolutional layer, a first Bottleneck module, a first maximum pooling layer, a second Bottleneck module, a third Bottleneck module and a second convolutional layer is used to perform deep feature extraction. By introducing The wavelet transform decomposes the input image at multiple scales, extracts low-frequency components that represent the global structure, and combines a dual-branch architecture to capture low-frequency structural features and deep semantic information respectively. While ensuring computational efficiency, it effectively preserves the topological structure and key features of the image, thereby improving the accuracy of the image segmentation network in segmenting dense pellet particles. The deep features extracted by the second-branch encoder and the features extracted by the first-branch encoder are fused and decoded using an adaptive weight feature fusion strategy to obtain the segmented result image. This effectively integrates the significant information contained in the particle image and preserves the detailed information of the image, thereby improving the overall performance and robustness of the image segmentation network. The constructed image segmentation network can quickly and effectively segment industrial pellet images. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0032] Figure 1 A schematic flow chart of an industrial pellet image segmentation method provided in one embodiment of the present application;

[0033] Figure 2 This is a flow chart of the lightweight pellet segmentation method based on image decomposition to extract low-frequency information in this application;

[0034] Figure 3 This is an example of the pellet particle image used in this application;

[0035] Figure 4 This is a diagram of the structure of the lightweight dual-branch encoder used in this application;

[0036] Figure 5 This is a visualization of the results of extracting low-frequency information from the image decomposition for this application;

[0037] Figure 6This is the adaptive weight feature fusion structure diagram used in this application;

[0038] Figure 7 This is a visualization of the test results of the optimal model for this application;

[0039] Figure 8 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0040] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0041] The purpose of this application is to provide an industrial pellet image segmentation method, equipment, medium and product, which can quickly and effectively segment industrial pellet images.

[0042] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0043] like Figure 1 As shown, the present application provides an industrial pellet image segmentation method, comprising:

[0044] Step 101: Obtain an open-source industrial pellet particle image dataset; the open-source industrial pellet particle image dataset includes multiple industrial pellet particle images and a pixel-level GT label corresponding to each industrial pellet particle image.

[0045] Step 102: Construct an image segmentation network; the encoder in the image segmentation network adopts a lightweight dual-branch encoder; the lightweight dual-branch encoder includes a first branch encoder and a second branch encoder; the first branch encoder uses a lightweight convolutional neural network to extract features of the industrial pellet particle image input to the first branch encoder; the second branch encoder uses wavelet transform to perform multi-scale decomposition on the industrial pellet particle image input to the second branch encoder, extracts the low-frequency information of the industrial pellet particle image, and then adopts an improved lightweight ResNet18 network to extract deep features; the improved lightweight ResNet18 network includes a first convolutional layer, a first Bottleneck module, a first maximum pooling layer, a second Bottleneck module, a third Bottleneck module and a second convolutional layer connected in sequence; the decoder in the image segmentation network performs feature fusion on the deep features extracted by the second branch encoder and the features extracted by the first branch encoder through an adaptive weight feature fusion strategy.

[0046] Among them, a dynamic adaptive soft threshold module is arranged between the first convolutional layer and the first Bottleneck module; the dynamic adaptive soft threshold module is connected to the first convolutional layer and the first Bottleneck module respectively; the dynamic adaptive soft threshold module is used to perform noise reduction operation on the feature map extracted by the first convolutional layer using dynamic adaptive soft threshold processing; the feature map is obtained after the first convolutional layer performs initial feature extraction on the low-frequency information.

[0047] The first convolutional layer has 20 output channels, the first bottleneck module has 40 output channels, the second bottleneck module has 80 output channels, the third bottleneck module has 160 output channels, and the second convolutional layer has 160 output channels. The convolution kernel size of the first convolutional layer is 7×7, and the convolution stride of the first convolutional layer is 2. The first convolutional layer uses BatchNorm2d normalization and ReLU activation. The convolution kernel size of the second convolutional layer is 3×3, and the convolution stride of the second convolutional layer is 2. The second convolutional layer uses BatchNorm2d normalization and ReLU activation. The pooling size of the first max pooling layer is 3×3, and the pooling stride of the first max pooling layer is 2. The first, second, and third bottleneck modules all contain a fusion of cubic convolution, BatchNorm2d normalization, and ReLU activation.

[0048] The lightweight convolutional neural network consists of a third convolutional layer, a second maximum pooling layer, a fourth convolutional layer, a third maximum pooling layer, a fifth convolutional layer, and a sixth convolutional layer, connected in sequence. The third convolutional layer has 20 output channels, the fourth convolutional layer has 40 output channels, the fifth convolutional layer has 80 output channels, and the sixth convolutional layer has 80 output channels. The convolution kernel size of the third, fourth, and fifth convolutional layers is 1×1, and the convolution stride of the third, fourth, and fifth convolutional layers is 3. The convolution kernel size of the sixth convolutional layer is 1×1, and the convolution stride of the sixth convolutional layer is 1. The third, fourth, fifth, and sixth convolutional layers all use InstanceNorm2d normalization and LeakyReLU activation. The pooling size of the second and third maximum pooling layers is 2×2, and the pooling stride of the second and third maximum pooling layers is 2.

[0049] The decoder in the image segmentation network includes a first adaptive weight feature fusion module, a first upsampling layer, a first decoding layer, a second upsampling layer, a second decoding layer, a third upsampling layer, a third decoding layer, a second adaptive weight feature fusion module, a fourth upsampling layer, and a convolution block. The first adaptive weight feature fusion module is connected to the third Bottleneck module and the first upsampling layer respectively; the first upsampling layer, the first decoding layer, the second upsampling layer, the second decoding layer, the third upsampling layer, the third decoding layer, the second adaptive weight feature fusion module, the fourth upsampling layer, and the convolution block are connected in sequence; the second convolution layer is connected to the first upsampling layer; the first Bottleneck module is connected to the third upsampling layer; the second Bottleneck module is connected to the second upsampling layer; and the sixth convolution layer is connected to the second adaptive weight feature fusion module.

[0050] The second branch encoder uses wavelet transform to perform multi-scale decomposition on the industrial pellet particle image input into the second branch encoder, extracting the low-frequency information of the industrial pellet particle image while also extracting the high-frequency information of the industrial pellet particle image; the high-frequency information is input into the first adaptive weight feature fusion module; the first adaptive weight feature fusion module is used to perform feature fusion on the high-frequency information and the output of the third Bottleneck module through an adaptive weight feature fusion strategy.

[0051] The second adaptive weight feature fusion module is used to perform feature fusion on the output of the third decoding layer and the output of the sixth convolutional layer through an adaptive weight feature fusion strategy.

[0052] Step 103: Use the open source industrial pellet particle image dataset to train and optimize the image segmentation network to obtain an optimized image segmentation network.

[0053] Step 104: Input the image of the industrial pellet particles to be segmented into the optimized image segmentation network, and use the optimized image segmentation network to segment the boundaries of the industrial pellet particles to obtain a segmented result image.

[0054] The technical solution of this application is described below with a specific embodiment:

[0055] The industrial pellet image segmentation method provided in this application is a lightweight pellet segmentation method based on image decomposition to extract low-frequency information (i.e., a pellet image segmentation method based on image decomposition to extract low-frequency information). The overall flow chart of the lightweight pellet segmentation method based on image decomposition to extract low-frequency information in this application is as follows: Figure 2 As shown, the method specifically includes:

[0056] (1) Data preparation (i.e. preparing the dataset):

[0057] Before detecting the particle size distribution, the particle image needs to be segmented. This application is an open source industrial pellet particle image dataset obtained from the literature (i.e., an open source industrial pellet particle image dataset). The dataset includes industrial pellet images and corresponding labels. The dataset is labeled, that is, GT labels made by annotation tools. The dataset takes into account factors such as uneven lighting, particle shadows, particle overlap, and material adhesion in the background. The pixels of shadows and dust particles are classified as background, and the pixels of transparent particles are classified as foreground. Accordingly, a binary classification problem is designed for this application. The pellet dataset mentioned (i.e., an open source industrial pellet particle image dataset) can be downloaded from the literature "DeoAJ, Behera SK, Das D P. Online Monitoring of Iron Ore Pellet Size Distribution Using Lightweight Convolutional Neural Network [J]. IEEE Transactions on Automation Science and Engineering, 2023, 21 (2): 1974-1985." The data example of this dataset is as follows Figure 3 As shown, Figure 3 Pellets and example diagrams used in this application. Figure 3 Four images from this dataset are selected for example. Figure 3 Parts (a), (b), (c), and (d) represent four different images selected from the open-source industrial pellet image dataset. The pellet image training set (i.e., the open-source industrial pellet image dataset) is used for model training and performance evaluation to obtain the optimal segmentation model.

[0058] (2) Data division:

[0059] The pellet dataset (i.e., pellet image dataset) obtained from the literature was cleaned and randomly divided into a training set and a test set in a 4:1 ratio. In order to prove the effectiveness of the method proposed in this application, the pellet dataset was randomly divided twice for training (i.e., the pellet image dataset obtained from the literature was randomly divided into two training sets and test sets in a 4:1 ratio), and the results obtained were used to evaluate the method of this application. Among them, the training set is used for model training and performance evaluation to obtain the optimal segmentation model, and the test set is used to test the segmentation effect of the model.

[0060] (3) Construct a lightweight dual-branch encoder (i.e., a dual-branch encoder based on the Pytorch deep learning framework) to perform semantic segmentation on the pellet image:

[0061] Based on the PyTorch deep learning framework, a pruned dual-branch encoder (i.e., a lightweight dual-branch encoder) that extracts low-frequency information based on image decomposition is built. The specific operations are as follows:

[0062] The first-branch encoder directly extracts features from the input image using a lightweight convolutional neural network and prunes the input and output channels (i.e., reducing the number of input channels for feature extraction). The first-branch encoder consists of an input layer, a third convolutional layer, a second maximum pooling layer, a fourth convolutional layer, a third maximum pooling layer, a fifth convolutional layer, and a sixth convolutional layer. The convolutional layer directly extracts features from the input image, while the pooling layer compresses the spatial dimensions of the feature map while retaining important information and enhancing the robustness of the model (i.e., the image segmentation network). For the second branch encoder, the input image is first decomposed using wavelet transform to extract the low-frequency information of the image, and then the feature map corresponding to the extracted low-frequency information is subjected to dynamic soft threshold processing (i.e., dynamic adaptive soft threshold processing) for noise reduction. Specifically, the extracted low-frequency information is subjected to initial feature extraction using a convolutional layer, and the extracted feature map is subjected to adaptive dynamic soft threshold processing (i.e., dynamic adaptive soft threshold processing) for noise reduction. Finally, an improved lightweight ResNet18 network is selected to extract the deep features (i.e., deep features) of the pellet image. The improved lightweight ResNet18 network includes a first convolutional layer, a first maximum pooling layer, three basic Bottleneck modules (i.e., a first Bottleneck module, a second Bottleneck module, and a third Bottleneck module) and a second convolutional layer. After obtaining the feature map extracted by the second branch encoder, the decoder is composed of a first adaptive feature fusion layer (i.e., a first adaptive weight feature fusion module), a first upsampling layer, a first decoding layer, a second upsampling layer, a second decoding layer, a third upsampling layer, a third decoding layer, a second adaptive feature fusion layer (i.e., a second adaptive weight feature fusion module), a fourth upsampling layer, and a convolution block. First, the feature map output from the third basic Bottleneck module (i.e., the third Bottleneck module) and the decomposed high-frequency information are reconstructed by feature fusion through the first adaptive feature fusion layer, and then upsampled with the feature map output by the second convolutional layer for fusion decoding. The decoded features are upsampled and then iteratively fused with the feature map extracted by the basic Bottleneck module (i.e., Bottleneck module) in the second branch encoder through a jump connection. Finally, the feature map extracted by the first branch encoder is fused through the second adaptive feature fusion layer, and then upsampled and a convolution block operation is performed to finally output the feature map of the specified number of categories. The structure of the lightweight dual-branch encoder (i.e., lightweight dual-branch encoder) is as follows: Figure 4 shown. Figure 4A flow chart of the lightweight dual-branch encoder module (ie, lightweight dual-branch encoder) used in the present application is shown.

[0063] The input image has 3 channels. The third convolutional layer of the first branch encoder has 20 output channels, the fourth convolutional layer has 40 output channels, the fifth convolutional layer has 80 output channels, and the sixth convolutional layer has 80 output channels. The convolution kernel size of the first three layers is 1×1, with a convolution stride of 3, and the convolution kernel size of the fourth layer is 1×1. All four convolutional layers use InstanceNorm2d normalization and LeakyReLU activation. In the first branch encoder, the maximum pooling layer has a pooling size of 2×2 and a pooling stride of 2. The first convolutional layer of the second branch encoder has 20 output channels, the first basic Bottleneck module (i.e., the first Bottleneck module) has 40 output channels, the second basic Bottleneck module (i.e., the second Bottleneck module) has 80 output channels, the third basic Bottleneck module has 160 output channels, and the second convolutional layer has 160 output channels. In the second-branch encoder, the first convolutional layer has a 7×7 kernel size and a stride of 2, with BatchNorm2d normalization and ReLU activation. In the second-branch encoder, the maximum pooling layer has a pooling size of 3×3 and a stride of 2. The basic Bottleneck module consists of a fusion of cubic convolution, BatchNorm2d normalization, and ReLU activation. The second convolutional layer has a 3×3 kernel size and a stride of 2. The decoder's first adaptive feature fusion layer has 40 output channels, the first decoding layer has 40 output channels, the second decoding layer has 40 output channels, the third decoding layer has 20 output channels, the second adaptive feature fusion layer has 20 output channels, and the convolution block has 1 output channel.

[0064] After obtaining the feature map extracted by the dual branches, the present application reconstructs the extracted feature map and high-dimensional information through an adaptive weight feature fusion operation, and fuses the features in combination with jump connections to obtain the feature map extracted by the second branch encoder. Finally, after being fused with the feature map extracted by the first branch encoder, a convolution operation is performed to finally output the feature map of the specified number of categories.

[0065] Among them, the low-frequency information and high-frequency information obtained by decomposing the image through wavelet transform are as follows: Figure 5 shown. Figure 5 This is a visualization of the results of extracting low-frequency information from the image decomposition used in this application. Figure 5 Part (a) represents the input image. Figure 5 Part (b) represents low-frequency information (i.e., low-frequency components). Figure 5Part (c) represents high-frequency information (i.e., high-frequency components).

[0066] Then the low-frequency information feature map is subjected to soft threshold processing, and the steps are as follows:

[0067] Step 1: First calculate the median, standard deviation std and variance var of each channel of the input low-frequency information feature map in the spatial dimension.

[0068] Step 2: Calculate the dynamic threshold of each channel feature map through the median, standard deviation and variance. The calculation formula is as follows:

[0069] threshold=median+k·std·e -u·var ;

[0070] Among them, k is the initial threshold factor, u represents the nonlinear adjustment parameter, and e represents the natural logarithm base.

[0071] Step 3: Adaptively adjust the initial threshold factor based on the calculated dynamic threshold, namely:

[0072]

[0073] in, Indicates the threshold factor after adaptive adjustment.

[0074] After obtaining the adaptively adjusted threshold factor, update the dynamic threshold, that is:

[0075]

[0076] in, represents the updated dynamic threshold, and β represents the adaptive threshold factor adjustment coefficient.

[0077] Step 4: After soft thresholding the input feature map, the noise below the threshold is removed to retain the stronger features. This is the feature map after denoising that is finally obtained in this application. That is:

[0078]

[0079] in, represents the feature map after denoising, and F represents the input feature map.

[0080] The extracted feature maps and high-dimensional information are then processed as follows Figure 6 The adaptive weight feature fusion operation shown is reconstructed, and the steps are as follows:

[0081] Step S1: First, L1 normalization is performed on the feature maps extracted from low-frequency information and high-frequency information (i.e., low-frequency feature map and high-frequency feature map) to obtain an activity level map, and the final activity level map is calculated based on a fast average operator.

[0082] Step S2: Generate an attention weight map through 1*1 convolution for the activity level map corresponding to the extracted high-dimensional information, and multiply the map with the input feature map element by element, so as to achieve adaptive weighting and aggregation of the feature map.

[0083] Step S3: Finally, a soft maximum operation is used to generate feature weights and then a splicing operation is performed to complete the adaptive weight feature fusion.

[0084] (4) Training the model and evaluating model performance:

[0085] After setting hyperparameters such as the number of output channels of the convolutional layer, the number of hidden layer neurons, the size of the convolution kernel, the activation function, the loss function, the optimizer, the learning rate, the number of training iterations, and the number of rounds of evaluation for the lightweight dual-branch encoder (i.e., the lightweight dual-branch encoder), the model is trained using a public pellet image dataset (i.e., an open source industrial pellet particle image dataset). The training process is divided into two stages: forward propagation and backpropagation. In the forward propagation stage, the input data undergoes a series of transformations, and the output result (i.e., the feature map before classification by the output layer) is obtained. (i) It can be expressed as:

[0086]

[0087] Among them, F(·) represents this series of linear and nonlinear operations. The linear in this series represents the convolution operation, the nonlinear represents the activation function, normalization and other operations, and C represents the total number of data categories. represents the input features, represents the score value of the i-th sample on the j-th output category, The last step of the forward process is a 1*1 convolution operation, which converts the number of channels of the neuron's output feature map into the number of categories of the segmented area. It can be expressed as:

[0088]

[0089] in, represents the original score of the i-th sample in the c-th category, where c represents different categories.

[0090] The backward process is the process of gradient descent and model weight parameter update. It is connected with the forward process through the improved adaptive interval binary cross entropy loss function. The specific operations are as follows:

[0091]

[0092] in, represents the improved adaptive margin binary cross entropy loss function, y i is the label corresponding to the input x (marked using the Labelme tool), represents the predicted value of the i-th sample after passing the Sigmoid activation function, N represents the total number of samples participating in the loss calculation, x represents the input image, m i is the adaptive interval of the i-th sample, which is used to dynamically adjust the interval parameter in the loss function according to the difficulty of predicting the sample, so that the model pays more attention to the difficult-to-predict samples. It is defined as follows:

[0093]

[0094] Wherein, a represents a constant term, which is set to 0.1 in this application. After sufficient iterations of the forward and backward propagation phases, the network learning is stopped, and the model is automatically saved during the learning period.

[0095] Load the trained model and input the validation set data into the model to evaluate the model segmentation performance.

[0096] Among them, the evaluation indicators use precision and IoU (intersection over union), and the calculation formula is as follows:

[0097]

[0098] Among them, TP, TN, FP and FN represent the number of correctly predicted positive samples, the number of correctly predicted negative samples, the number of incorrectly predicted positive samples, and the number of incorrectly predicted negative samples, respectively.

[0099] By evaluating the model, the optimal segmentation model (the segmentation model is the image segmentation network) can be obtained.

[0100] Specifically, the model built by step (3) is set with a learning rate of 0.001, a training batch size of 2, and a total of 150 iterations, in which the model saves the weights every 10 iterations. In addition, the improved adaptive interval binary cross entropy loss function is selected to dynamically adjust the interval parameter in the loss function according to the difficulty of the prediction sample, so that the model pays more attention to the difficult-to-predict samples, as shown in the formula As shown, the losses are summed to obtain the total loss. This application uses ADAM as the optimizer. The process of gradient optimization using ADAM can be expressed as follows:

[0101] m t =β1m t-1 +(1-β1)g t ;

[0102]

[0103] Among them, α represents the learning rate (0.001), and denote the deviation correction of the first-order moment and the second-order moment, respectively, and ε denotes a small constant term, m t represents the first-order moment estimate at the t-th time step, β1 represents the momentum decay rate corresponding to the first-order moment, and m t-1 represents the first-order moment estimate at the t-1th time step, g t represents the gradient at the t-th time step, v t represents the second-order moment estimate at the t-th time step, β2 represents the momentum decay rate of the second-order moment, and v t-1 represents the second-order moment estimate at the t-1th time step, represents the momentum decay rate of the first-order moment at the t-th time step, represents the momentum decay rate of the second-order moment at the t-th time step, θ t+1 represents the model parameter value after the t+1th time step is updated according to the gradient information, θ t Represents the model parameter value at the tth time step.

[0104] After setting the above hyperparameters, the model was trained for 150 iterations, with the trained weights saved every 10 iterations. The optimal weights saved every 10 iterations were selected and used to evaluate and compare the segmented pellet image test data one by one to obtain the weights that achieved the best segmentation effect. The evaluation metrics used were Accuracy and Intersection over Union (IoU). The evaluation results are shown in Table 1.

[0105] Table 1 Segmentation effect of the model in this application on the pellet image test set

[0106]

[0107] (5) Visualization results of the test model:

[0108] Load the optimal segmentation model weights, randomly extract images from the pellet test dataset divided by this application, and input them into the optimal segmentation model. After the forward process and the backward process, the optimal segmentation result output by the proposed model after the input image is obtained. The performance evaluation results of the optimal model of this application are visualized as follows Figure 7 shown.

[0109] Specifically, the optimal detection weights in the above evaluation process are loaded, and 3 images are selected from the pre-processed pellet image dataset to input the model to obtain the test results. The visualization diagram of the optimal model test results of this application is as follows: Figure 7 As shown in the figure, the segmentation effect of the proposed method is very good under conditions such as uneven lighting, dense particles, and stacked particles. After visualizing several images, the effectiveness of the proposed segmentation method is demonstrated. The next step is to fit the obtained particle mask image to obtain the particle size and finally draw a distribution trend chart of the pellet particle size. Figure 7 Parts (a), (b), and (c) represent three different input images selected from the preprocessed pellet image dataset. Figure 7 Part (d) is Figure 7 The labels corresponding to the images shown in part (a) are: Figure 7 Part (e) is Figure 7 The labels corresponding to the images shown in part (b) are Figure 7 Part (f) is Figure 7 The labels corresponding to the images shown in part (c) are Figure 7 Part (g) is Figure 7 The segmentation result of the image shown in part (a) is Figure 7 The middle (h) part is Figure 7 The segmentation result of the image shown in part (b) is Figure 7 Part (i) is Figure 7 The segmentation result of the image shown in part (c) is shown in Figure 1. The labels produced are the GT labels produced by the annotation tool.

[0110] (6) Use the segmented mask image for granularity analysis:

[0111] The most important step before particle size detection in this application is to segment the dense particles. Finally, the segmented image is first read through the connected region analysis method to read the mask image and binarize it, then the connected region is marked to extract the data of the region, and finally the particle size distribution trend can be obtained by calculating the equivalent diameter of the particles.

[0112] The lightweight pellet segmentation method based on image decomposition to extract low-frequency information provided in this application is a method for quickly and effectively segmenting dense large-scale pellet particles. It achieves high-precision particle segmentation through frequency domain information decoupling and dynamic noise reduction mechanism. Its core lies in: First, a dual-branch encoder architecture is constructed. The first branch uses a lightweight convolutional neural network for feature extraction, and the second branch performs multi-scale decomposition of the input image through wavelet transform. After obtaining the low-frequency components, an improved lightweight ResNet18 network is used for deep feature learning. Secondly, in order to solve the noise interference problem in the low-frequency feature map, a dynamic adaptive soft threshold mechanism is proposed for noise reduction. Finally, an adaptive weight feature fusion strategy is proposed to fuse and reconstruct the low-frequency features of the second branch with the corresponding high-frequency components, and the probability map after segmentation is obtained through a multi-scale decoder, which provides reliable technical support for particle size analysis.

[0113] This application addresses the shortcomings of traditional morphological methods and existing deep learning-based pellet image segmentation methods by providing a lightweight pellet segmentation method based on image decomposition to extract low-frequency information. The following introduces the pellet image segmentation method based on morphology and the pellet image segmentation algorithm based on deep learning:

[0114] (1) Segmentation method of pellet particle image based on morphology:

[0115] The pellet size detection method based on traditional image processing includes the following steps: first, the collected original pellet image is filtered and grayed, and the target area and background area are segmented by threshold value. Then, the target area is re-segmented using the watershed segmentation model to obtain the area where each pellet particle is located.

[0116] Advantages: The algorithm has a simple structure, is easy to implement, does not require prior knowledge, and has a high degree of adaptability.

[0117] Disadvantages: For complex industrial processes, there are often problems of over-segmentation and breaking of regional boundaries. Redundant parameter adjustments are required based on the local pixel value distribution to obtain satisfactory results.

[0118] (2) Pellet image segmentation algorithm based on deep learning:

[0119] The deep learning-based pellet image segmentation algorithm includes the following steps: first, the pellet image is normalized through image preprocessing. Then, a convolutional neural network is used to extract features and train the model. By optimizing the loss function, the model learns to accurately segment the pellet particle boundaries and ultimately generates a segmentation result.

[0120] Advantages: It can automatically learn complex features, provide high-precision and robust segmentation results, and reduce manual intervention.

[0121] Disadvantages: They primarily optimize the model structure to extract effective features, but neglect low-frequency information in pellet images. Furthermore, due to the rapid changes in pellet distribution and state, existing methods often suffer from latency and insufficient processing speed, limiting their application in dynamic production environments.

[0122] The shortcomings of the existing technology include:

[0123] Disadvantage 1: Existing deep learning methods often face the problem of large-scale particle overlap and blurred boundaries when processing dense pellet images. This makes it difficult for traditional convolutional neural networks to accurately distinguish adjacent particles. To address this challenge, researchers have used networks such as modified U-Net and modified Mask-RCNN to improve particle segmentation accuracy. However, the high complexity of dense particle images still requires further optimization of network models to improve processing efficiency and accuracy.

[0124] Disadvantage 2: Due to the rapid changes in the distribution and state of pellet particles in industrial processes, some existing methods often face delays or insufficient processing speed while achieving high performance, which limits their application in dynamic production environments.

[0125] Disadvantage 3: Low-frequency information often contains the main features of the image, but may also be accompanied by noise or irrelevant interference. Directly using low-frequency features may cause the model to over-rely on these noises, affecting the final segmentation effect.

[0126] Disadvantage 4: Due to information imbalance, quality differences and noise problems in industrial pellet production scenarios, important features cannot be efficiently integrated, which will affect the accuracy of subsequent particle size segmentation and make it impossible to accurately estimate the particle size distribution.

[0127] In response to shortcoming 1, this application proposes a lightweight dual-branch encoder based on multi-scale decomposition of wavelet transform for pellet image segmentation. By introducing wavelet transform to perform multi-scale decomposition on the input image, low-frequency components that represent the global structure are extracted, and the dual-branch architecture is combined to capture low-frequency structural features and deep semantic information respectively. While ensuring computational efficiency, the topological structure and key features of the image are effectively retained, thereby improving the accuracy of the model in segmenting dense pellet particles.

[0128] To address shortcoming 2, this application pruned the parameters of the proposed dual-branch encoder and adjusted the number of output channels, thereby achieving effective segmentation accuracy while ensuring the segmentation speed.

[0129] In response to shortcoming 3, this application proposes a dynamic adaptive soft threshold processing method to reduce the noise component contained in the low-frequency information extracted by wavelet transform, and solves the problems of shadows and uneven lighting in some pellet images.

[0130] To address shortcoming 4, this application proposes an adaptive weighted fusion strategy, which improves the overall performance and robustness of the model by dynamically adjusting the weights between the feature maps corresponding to the low-frequency components and the high-frequency components and then reconstructing them.

[0131] This application addresses the accuracy-efficiency imbalance in existing deep learning algorithms for the segmentation of densely overlapping pellets. By proposing a lightweight pellet image segmentation method based on image decomposition to extract low-frequency information, this method can generate a particle size distribution map based on particle images collected at industrial sites, thereby enabling real-time monitoring of production status. Starting from the perspective of extracting low-frequency information from particle images, this application constructs a lightweight dual-branch encoder to achieve the collaborative extraction of global semantic information and low-frequency structural dual-stream features. This method not only improves segmentation accuracy but also allows for real-time dynamic application in real-world scenarios.

[0132] Based on existing deep learning technology, this application proposes a lightweight dual-branch encoder method for extracting low-frequency information based on image decomposition. The core process is as follows: First, the acquired open-source industrial pellet particle dataset is cleaned and re-partitioned into training and test sets. Second, a lightweight dual-branch encoder is constructed by combining wavelet transform decomposition to extract low-frequency information. After feature fusion using an adaptive feature fusion strategy, the segmented result image is decoded. Finally, the constructed model is trained using the divided training set, evaluated on the test set, and then the optimized model is deployed in the industrial pellet image segmentation application scenario.

[0133] The core of this application is:

[0134] 1. This application proposes a lightweight end-to-end dual-branch encoder structure to solve the particle image segmentation problem, which is summarized as follows:

[0135] (1) Global perception branch: Through parameterized and redesigned lightweight convolutional neural networks, multi-level progressive extraction of granular semantic features is achieved.

[0136] (2) Frequency domain analysis branch: Innovatively combines the wavelet multi-scale decomposition mechanism with the deep residual network. By improving the ResNet18 network and performing channel compression optimization, the fusion and collaborative representation of the image's high-frequency detail features and low-frequency structural features are achieved.

[0137] 2. This application introduces wavelet transform to decompose images for the first time in the field of particle image segmentation, which provides great help for subsequent particle segmentation tasks, especially the segmentation of dense particles.

[0138] 3. The dynamic adaptive soft threshold method proposed in this application can perform denoising on the feature map corresponding to the extracted low-frequency information.

[0139] 4. This application proposes a new adaptive weight feature fusion strategy to effectively integrate the significant information and detail information contained in particle images.

[0140] 5. This application proposes a new loss function to solve the problem of excessive focus on easy-to-predict samples during training.

[0141] Compared with the prior art, this application has the following advantages:

[0142] 1. This application proposes a dual-branch encoder for particle granularity segmentation based on image decomposition to extract low-frequency information. This method extracts low-frequency information by performing wavelet transform decomposition on the input image, and then extracts features from the extracted low-frequency information separately, thereby improving the robustness and accuracy of the model and providing assistance for subsequent particle segmentation tasks.

[0143] 2. The lightweight dual-branch encoder proposed in this application prunes the network structure and parameters, thereby achieving effective segmentation accuracy while ensuring the segmentation speed, so that it can be applied to the production process in real time.

[0144] 3. The dynamic adaptive soft threshold method proposed in this application improves the quality of the feature map by reducing the noise component of the feature map corresponding to the low-frequency information extracted by wavelet transform, thereby optimizing the subsequent segmentation process.

[0145] 4. The adaptive weighted fusion strategy proposed in this application reconstructs the feature map of the extracted low-frequency information and the high-frequency information, effectively integrating the significant information contained in the particle image and preserving the detailed information of the image.

[0146] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 8As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store industrial pellet image segmentation data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, an industrial pellet image segmentation method is implemented.

[0147] Those skilled in the art will understand that Figure 8 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present application and does not constitute a limitation on the computer device to which the solution of the present application is applied. A specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement. In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps of the above-mentioned method embodiments when executing the computer program.

[0148] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0149] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0150] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0151] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0152] The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may include, but are not limited to, general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units based on quantum computing, and the like.

[0153] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0154] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A method for industrial pellet image segmentation, characterized in that: The industrial pellet image segmentation method comprises: Obtain an open source industrial pellet particle image dataset; the open source industrial pellet particle image dataset includes multiple industrial pellet particle images and a pixel-level GT label corresponding to each industrial pellet particle image; Construct an image segmentation network; the encoder in the image segmentation network adopts a lightweight dual-branch encoder; the lightweight dual-branch encoder includes a first branch encoder and a second branch encoder; the first branch encoder uses a lightweight convolutional neural network to extract features of the industrial pellet particle image input to the first branch encoder; the second branch encoder uses wavelet transform to perform multi-scale decomposition on the industrial pellet particle image input to the second branch encoder, extracts the low-frequency information of the industrial pellet particle image, and then adopts an improved lightweight ResNet18 network to extract deep features; the improved lightweight ResNet18 network includes a first convolutional layer, a first Bottleneck module, a first maximum pooling layer, a second Bottleneck module, a third Bottleneck module and a second convolutional layer connected in sequence; the decoder in the image segmentation network performs feature fusion on the deep features extracted by the second branch encoder and the features extracted by the first branch encoder through an adaptive weight feature fusion strategy; Using the open source industrial pellet particle image dataset to train and optimize the image segmentation network to obtain an optimized image segmentation network; The image of the industrial pellet particles to be segmented is input into the optimized image segmentation network, and the optimized image segmentation network is used to segment the boundaries of the industrial pellet particles to obtain a segmented result image.

2. The industrial pellet image segmentation method according to claim 1, characterized in that: A dynamic adaptive soft threshold module is arranged between the first convolutional layer and the first Bottleneck module; the dynamic adaptive soft threshold module is connected to the first convolutional layer and the first Bottleneck module respectively; the dynamic adaptive soft threshold module is used to perform a noise reduction operation on the feature map extracted by the first convolutional layer using dynamic adaptive soft threshold processing; the feature map is obtained after the first convolutional layer performs initial feature extraction on the low-frequency information.

3. The industrial pellet image segmentation method according to claim 1, characterized in that: The number of output channels of the first convolutional layer is 20, the number of output channels of the first Bottleneck module is 40, the number of output channels of the second Bottleneck module is 80, the number of output channels of the third Bottleneck module is 160, and the number of output channels of the second convolutional layer is 160; The convolution kernel size of the first convolution layer is 7×7, and the convolution stride of the first convolution layer is 2; the first convolution layer uses BatchNorm2d normalization and ReLU activation; the convolution kernel size of the second convolution layer is 3×3, and the convolution stride of the second convolution layer is 2; the second convolution layer uses BatchNorm2d normalization and ReLU activation; The pooling size of the first maximum pooling layer is 3×3, and the pooling stride of the first maximum pooling layer is 2; The first Bottleneck module, the second Bottleneck module, and the third Bottleneck module all include a fusion operation of cubic convolution, BatchNorm2d normalization, and ReLU activation.

4. The industrial pellet image segmentation method according to claim 1, characterized in that: The lightweight convolutional neural network includes a third convolutional layer, a second maximum pooling layer, a fourth convolutional layer, a third maximum pooling layer, a fifth convolutional layer and a sixth convolutional layer connected in sequence; The number of output channels of the third convolutional layer is 20, the number of output channels of the fourth convolutional layer is 40, the number of output channels of the fifth convolutional layer is 80, and the number of output channels of the sixth convolutional layer is 80; The convolution kernel sizes of the third convolution layer, the fourth convolution layer, and the fifth convolution layer are all 1×1, and the convolution strides of the third convolution layer, the fourth convolution layer, and the fifth convolution layer are all 3. The convolution kernel size of the sixth convolution layer is 1×1, and the convolution stride of the sixth convolution layer is 1; The third convolutional layer, the fourth convolutional layer, the fifth convolutional layer and the sixth convolutional layer all adopt InstanceNorm2d normalization operation and LeakyReLU activation; The pooling sizes of the second maximum pooling layer and the third maximum pooling layer are both 2×2, and the pooling strides of the second maximum pooling layer and the third maximum pooling layer are both 2.

5. The industrial pellet image segmentation method according to claim 4, characterized in that: The decoder in the image segmentation network includes a first adaptive weight feature fusion module, a first upsampling layer, a first decoding layer, a second upsampling layer, a second decoding layer, a third upsampling layer, a third decoding layer, a second adaptive weight feature fusion module, a fourth upsampling layer and a convolution block; The first adaptive weight feature fusion module is respectively connected to the third Bottleneck module and the first upsampling layer; the first upsampling layer, the first decoding layer, the second upsampling layer, the second decoding layer, the third upsampling layer, the third decoding layer, the second adaptive weight feature fusion module, the fourth upsampling layer and the convolution block are connected in sequence; the second convolution layer is connected to the first upsampling layer; the first Bottleneck module is connected to the third upsampling layer; the second Bottleneck module is connected to the second upsampling layer; the sixth convolution layer is connected to the second adaptive weight feature fusion module.

6. The industrial pellet image segmentation method according to claim 5, characterized in that: The second branch encoder uses wavelet transform to perform multi-scale decomposition on the industrial pellet particle image input to the second branch encoder, extracting the low-frequency information of the industrial pellet particle image while also extracting the high-frequency information of the industrial pellet particle image; the high-frequency information is input into the first adaptive weight feature fusion module; the first adaptive weight feature fusion module is used to perform feature fusion on the high-frequency information and the output of the third Bottleneck module through an adaptive weight feature fusion strategy.

7. The industrial pellet image segmentation method according to claim 5, characterized in that: The second adaptive weight feature fusion module is used to perform feature fusion on the output of the third decoding layer and the output of the sixth convolutional layer through an adaptive weight feature fusion strategy.

8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the industrial pellet image segmentation method according to any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the industrial pellet image segmentation method according to any one of claims 1 to 7 is implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the industrial pellet image segmentation method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Variable-granularity self-adaptive coding method for smoke and dust disaster environment

    CN120751131A