Lightweight industrial vertical domain image segmentation method, device and equipment
By introducing a cascaded multi-receptive field module and channel regulator in the U-Net model, the problem of insufficient feature information acquisition in complex scenarios of lightweight industrial vertical image segmentation network is solved, and more efficient feature representation and segmentation performance is achieved.
Patent Information
- Application Number
- CN202510215946.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-24
AI Technical Summary
The existing lightweight industrial vertical image segmentation network is difficult to fully learn rich and diverse feature information in complex industrial scenarios, which makes it difficult to reach an ideal level of segmentation accuracy.
Using the improved U-Net model, the feature encoding branch and decoding branch are both set up with a cascading multi-receptive field module. The channel regulator and feature extractor dynamically adjust the number of channels and feature extraction strategies to achieve mining and fusion of multiple feature scales.
Effectively mining richer and more representative feature information, reduces the computational complexity and parameter quantity, improves segmentation performance, and achieves efficient and reliable segmentation effect in industrial vertical image segmentation.
Smart Images

Figure CN120198660A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image segmentation, and particularly to a lightweight industrial vertical domain image segmentation method, device and equipment. Background Art
[0002] In today's era, industrial vertical domain image segmentation technology plays a crucial role in numerous fields. From the high-precision production inspection of electronic components to the large-scale quality control of automotive parts, from the defect identification of complex aerospace components to the fine screening of textile defects, the accuracy and efficiency of image segmentation are directly related to product quality, production efficiency and production cost.
[0003] Traditional image segmentation methods, such as threshold-based segmentation, region growing method and edge detection algorithms, etc., once dominated in early industrial applications. However, with the development of industrial production towards high precision, high complexity and intelligence, these traditional methods gradually expose insurmountable limitations.
[0004] Lightweight industrial vertical domain image segmentation methods have emerged as the times require, attempting to solve the image segmentation problem under limited resources. For example, the UNeXt model has made certain progress in reducing the number of parameters and floating-point operation amounts by introducing multi-layer perceptrons and depthwise separable convolutions. Another example is that the ConvUNeXt model further integrates lightweight attention mechanisms and large kernel convolutions, reducing the parameter burden while maintaining the segmentation performance. The U-Lite model adopts axial depth convolution, effectively reducing the computational pressure while expanding the receptive field of the model. The CMUNeXt model uses large kernel functions and inverted bottleneck designs to successfully integrate long-range spatial and positional information, improving the diagnostic speed and accuracy in real industrial scenarios. The double-branch structure of the FBSNet model can capture a wide range of receptive field information and establish local pixel dependencies, which is beneficial to realizing real-time semantic segmentation on edge devices. Nevertheless, there are still many problems with lightweight networks. Due to the reduction of parameters and computational complexity, their feature representation capabilities are limited. In complex and changeable industrial scenarios, the shapes, textures, scales and lighting conditions of industrial targets vary greatly, and the above-mentioned lightweight networks are difficult to fully learn these rich and diverse feature information, resulting in the segmentation accuracy being difficult to reach the ideal level. For example, in the detection of surface defects of tiny electronic components, fine scratches, holes and other defects require precise feature characterization to be accurately identified, while lightweight networks may miss these key details due to insufficient feature representation, affecting the accuracy of product quality inspection.
[0005] Modern feature extraction modules based on multiple receptive fields, such as techniques like feature pyramids and parallel multi-group convolutions, provide new ideas for improving segmentation performance. The feature pyramid constructs multi-scale feature maps, enabling effective capture of target features at different levels and enhancing the detection ability for targets of different scales. Parallel multi-group convolutions, on the other hand, use convolutional kernels of different sizes to process the image in parallel, obtaining multi-scale feature responses and enriching the feature representation. However, while these techniques improve performance, they inevitably increase the computational resource requirements and the number of parameters of the model, making the model more complex and heavy, which runs counter to the actual needs of industrial environments with limited resources. In an industrial environment, it is necessary to reduce the model complexity and resource consumption as much as possible while ensuring segmentation accuracy to achieve efficient and reliable image segmentation.
[0006] In view of this, there is an urgent need to provide a lightweight industrial vertical domain image segmentation method for multi-feature scale mining. Summary of the Invention
[0007] To overcome the problems in the related art, the present disclosure provides a lightweight industrial vertical domain image segmentation method, device, and equipment to solve the technical problems in the related art.
[0008] One or more embodiments of this specification provide a lightweight industrial vertical domain image segmentation method, which acquires an industrial vertical domain image segmentation dataset;
[0009] Construct an improved U-Net model. Among them, the encoder of the feature encoding branch of the improved U-Net model includes a first cascaded multi-receptive field module and a downsampling module connected thereto. Each cascaded decoder of the feature decoding branch includes a second cascaded multi-receptive field module, a channel connection block, and an upsampling module. Each decoder splices and integrates the feature map processed by the upsampling module and the feature map output from the corresponding layer encoder in the feature decoding branch in the channel direction, and the spliced feature map is input to the second cascaded multi-receptive field module. The first cascaded multi-receptive field module and the second cascaded multi-receptive field module are respectively provided with a first channel regulator, a cascaded strategy module composed of multiple cascaded feature extractors, and a second channel regulator;
[0010] The first channel regulator is used to keep the resolution size of the input feature map unchanged and adjust the number of output channels to There are N (N≥1) of them. The feature map processed by the first channel regulator is divided into channels on average to obtain a first feature map and a second feature map. The first feature map and the second feature map are fused by element-wise summation to obtain a third feature map. The first feature map / second feature map is input into the cascade strategy module. The cascade strategy module uses the depth convolution in the depthwise separable convolution to apply convolution kernels to each channel of the first feature map / second feature map respectively for feature extraction. And the feature maps extracted by each level of feature extractor are concatenated with the third feature map along the channel direction to obtain a fourth feature map, and the fourth feature map is output with the number of channels restored to the input channel number through the second channel regulator.
[0011] The industrial vertical domain image segmentation dataset is input into the improved U-Net model for training until the convergence condition is reached, so as to obtain a segmentation model for realizing industrial vertical domain image segmentation.
[0012] One or more embodiments of this specification provide a lightweight industrial vertical domain image segmentation device, including:
[0013] A training data acquisition module, which is used to acquire the industrial vertical domain image segmentation dataset;
[0014] A model construction module, which is used to construct an improved U-Net model. Among them, the encoder of the feature encoding branch of the improved U-Net model includes a first cascade multi-receptive field module and a downsampling module connected thereto. Each level of cascade decoder of the feature decoding branch includes a second cascade multi-receptive field module, a channel connection block and an upsampling module. Each level of decoder splices and integrates the feature map processed by the upsampling module with the feature map output from the corresponding layer encoder in the feature decoding branch along the channel direction. The spliced feature map is input into the second cascade multi-receptive field module. The first cascade multi-receptive field module and the second cascade multi-receptive field module are respectively provided with a first channel regulator, a cascade strategy module composed of multiple cascade feature extractors and a second channel regulator;
[0015] The first channel regulator is used to keep the resolution size of the input feature map unchanged and adjust the output channel number to There are N, N≥1, and the feature map processed by the first channel regulator is divided equally in channels to obtain a first feature map and a second feature map. The first feature map and the second feature map are fused by element-wise summation to obtain a third feature map. The first feature map / second feature map is input into the cascaded strategy module. The cascaded strategy module uses the depth convolution in depthwise separable convolution to apply convolution kernels to the first feature map / second feature map channel by channel for feature extraction. And the feature maps extracted by each level of feature extractor are concatenated with the third feature map along the channel direction to obtain a fourth feature map, and the channel of the fourth feature map is restored to the input channel number by the second channel regulator for output;
[0016] A model training module, configured to input an industrial vertical domain image segmentation data set into the improved U-Net model for training until the convergence condition is reached, so as to obtain a segmentation model for implementing industrial vertical domain image segmentation.
[0017] One or more embodiments of this specification provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the lightweight industrial vertical domain image segmentation method as described above.
[0018] The lightweight industrial vertical domain image segmentation method, device, equipment, and medium provided by the present disclosure have the advantage that by using a U-shaped neural network, each feature encoder and feature decoder are both provided with a cascaded multi-receptive field module that can realize the mining of deep multi-scale information of the image feature map. This module, through the set channel regulator and feature extractor, the first channel regulator at the input end adjusts the output channel number to N according to the preset N feature extractors, and the second channel regulator at the output end restores the channels of the feature map, realizing that the channel number for custom channel upsampling and downsampling can dynamically adjust the fusion weights of different scale features according to the image content, so as to effectively mine richer, more representative, and more detailed feature information, and explore a more compact convolutional structure to reduce the amount of computation; furthermore, the cascaded strategy composed of multiple feature extractors set realizes to configure C×C convolution kernels for each channel to perform feature extraction operations, so as to realize in-depth exploration of different scale information in the feature map, and cleverly maintain the lightweight feature while effectively enhancing the feature representation ability by fusing information from different receptive fields in a single network layer, thereby greatly improving the overall performance. In addition, the method of this embodiment realizes excellent segmentation performance under the conditions of using a small number of parameters and low computational complexity through an efficient U-shaped neural network. And the cascaded multi-receptive field module set in this embodiment has universality and flexibility, is easy to be extended to other networks, while reducing the parameters and computational complexity and improving the segmentation performance. Description of the Drawings
[0019] To more clearly illustrate the technical solutions in one or more embodiments of this specification or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0020] Figure 1 Flowchart of a lightweight industrial vertical domain image segmentation method provided for one or more embodiments of this specification;
[0021] Figure 2 Improved U-Net model architecture diagram provided for one or more embodiments of this specification;
[0022] Figure 3 Cascade multi-receptive field module structure diagram provided for one or more embodiments of this specification;
[0023] Figure 4 First and second channel regulator structure diagrams provided for one or more embodiments of this specification;
[0024] Figure 5 Feature extractor structure diagram provided for one or more embodiments of this specification;
[0025] Figure 6 Class regulator structure diagram provided for one or more embodiments of this specification;
[0026] Figure 7 Block diagram of a lightweight industrial vertical domain image segmentation device provided for one or more embodiments of this specification; and
[0027] Figure 8 Structural schematic diagram of a computer device provided for one or more embodiments of this specification. Detailed implementation manners
[0028] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the following will clearly and completely describe the technical solutions in one or more embodiments of this specification in conjunction with the drawings in one or more embodiments of this specification. Obviously, the described embodiments are only some embodiments of this specification, rather than all embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0029] The following will make a detailed description of the present invention in conjunction with the detailed implementation manners and the drawings of the specification.
[0030] Method Embodiment
[0031] According to an embodiment of the present invention, a lightweight industrial vertical domain image segmentation method is provided, as Figure 1-3 shown Figure 1 is the flowchart of the lightweight industrial vertical domain image segmentation method provided in this embodiment, Figure 2 is the improved U-Net model architecture diagram provided in this embodiment, Figure 3 is the structure diagram of the cascaded multi-receptive field module provided in this embodiment. The lightweight industrial vertical domain image segmentation method according to an embodiment of the present invention includes the steps:
[0032] Step S1, obtaining an industrial vertical domain image segmentation data set;
[0033] Step S2, constructing an improved U-Net model. Among them, the encoder of the feature encoding branch of the improved U-Net model includes a first cascaded multi-receptive field module and a downsampling module connected thereto. Each cascaded decoder of the feature decoding branch includes a second cascaded multi-receptive field module, a channel connection block, and an upsampling module. Each decoder splices and integrates the feature map processed by the upsampling module and the feature map from the output of the corresponding layer encoder in the feature decoding branch in the channel direction. The integrated feature map is input to the second cascaded multi-receptive field module;
[0034] The first cascaded multi-receptive field module and the second cascaded multi-receptive field module are respectively provided with a first channel regulator, a cascaded strategy module composed of multiple cascaded feature extractors, and a second channel regulator. The first channel regulator is used to maintain the resolution size of the input feature map unchanged and adjust the output channel number to according to N feature extractors preset in the cascaded strategy module, N≥1, and perform channel average division on the feature map processed by the first channel regulator to obtain a first feature map and a second feature map. The first feature map and the second feature map are fused by element-wise summation to obtain a third feature map, and the first feature map / second feature map is input to the cascaded strategy module. The cascaded strategy module uses depth convolution in depthwise separable convolution to perform feature extraction operations on each channel of the first feature map / second feature map by using a C×C convolution kernel respectively, so as to perform feature extraction on each channel of the first feature map / second feature map respectively, keeping the resolution size and channel dimension of the feature map unchanged. And the feature maps extracted by each level of feature extractor are connected to the third feature map in the channel direction to obtain a fourth feature map, and the channel of the fourth feature map is restored to the input channel number by the second channel regulator for output.
[0035] Step S3: Input the industrial vertical domain image segmentation dataset into the improved U-Net model for training until the convergence condition is reached, thereby obtaining a segmentation model for realizing industrial vertical domain image segmentation.
[0036] A lightweight industrial vertical domain image segmentation method provided in this embodiment uses a U-shaped neural network. Cascade multi-receptive field modules capable of mining deep multi-scale information of image feature maps are set in each feature encoder and feature decoder. Through the set channel regulator and feature extractor, the first channel regulator at the input end adjusts the output channel number to a certain number according to N preset feature extractors, and the second channel regulator at the output end restores the feature map channels, realizing that the channel numbers for custom channel upsampling and downsampling can dynamically adjust the fusion weights of different scale features according to the image content, thereby effectively mining richer, more representative, and more detailed feature information, and exploring a more compact convolutional structure to reduce the computational amount; then, the cascade strategy composed of multiple feature extractors is set to configure C×C convolutional kernels for each channel to perform feature extraction operations to deeply explore different scale information in the feature map, and cleverly maintain the lightweight characteristics while effectively enhancing the feature representation ability by fusing information from different receptive fields in a single network layer, thereby greatly improving the overall performance. In addition, the method in this embodiment realizes excellent segmentation performance under the conditions of using a small number of parameters and low computational complexity through an efficient U-shaped neural network. The cascade multi-receptive field module set in this embodiment has generality and flexibility, is easy to be extended to other networks, and at the same time reduces the parameters and computational complexity and improves the segmentation performance.
[0037] In this embodiment, for the acquisition of the industrial vertical domain image segmentation dataset, mainly select a representative public industrial vertical domain image segmentation dataset, which should cover image data under various industrial scenarios, such as images for detecting surface defects of machined parts (surface defects of tiny electronic components), images for identifying materials on industrial production lines, etc. Conduct comprehensive and detailed preprocessing on the dataset to ensure the quality and consistency of the data and improve the effect of model training.
[0038] In this embodiment, it also includes image preprocessing on the industrial vertical domain image segmentation dataset. The specific steps of image preprocessing are as follows:
[0039] Step 11: Image normalization; adjust the size of all images so that their resolutions all reach a unified pixel H×H, such as 256×256 pixels, then convert the sized images into the PyTorch tensor format, and for the image data with pixel values in the [0, H] interval, use the maximum-minimum normalization algorithm for processing, and finally obtain the preprocessed images as input feature maps.
[0040] Step 12, Image segmentation category encoding: Set the category ID of the background pixels of each preprocessed image to 0. For the different category pixels segmented from the image, different category pixel IDs from 1 to D are sequentially assigned, where D represents the total number of categories other than the background. Then convert each ID into the form of a PyTorch tensor, which is used as the true label and combined with the loss function to carry out the model training work.
[0041] In this embodiment, referring to Figure 4 As shown, it is the structural diagram of the first and second channel regulators provided by this embodiment. An innovative channel regulator is designed. By setting the 1x1 convolution operation for feature fusion to mine non-linear feature information, its main function is to flexibly set and adjust the number of channel dimensions according to requirements while maintaining the resolution size of the feature map unchanged, so as to effectively mine richer and more representative feature information. Referring to Figure 4 , the specific channel adjustment processing procedures of the first and second channel regulators are as follows:
[0042] Step 21, Adjust the number of channels in the dimension of the input feature map through a pointwise convolution layer according to the preset number of channel upsampling or downsampling, and then obtain the feature map after channel adjustment processing; for example, assume that it is necessary to adjust the output channel number of the input feature map to according to the number of feature extractors set in the cascaded multi-receptive field module; the pointwise convolution processing can refer to the convolution operation performed using a convolution kernel of 1×1×C, where C represents the number of channels of the convolution kernel. In this embodiment, the feature information extracted from each channel is integrated through 1×1 pointwise convolution, and the output channel number is reduced to further compress the number of parameters;
[0043] Step 22, Then perform batch normalization operation on the feature map obtained by adjusting the channel dimension in Step 21 to obtain the normalized feature map;
[0044] Step 23, Perform non-linear processing on the batch-normalized feature map using the Gaussian Error Linear Unit (GELU) activation function to obtain the feature map corresponding to the number of channels.
[0045] In this embodiment, the first cascaded multi-receptive field module and the second cascaded multi-receptive field module include a first channel regulator, a cascaded policy module composed of multiple cascaded feature extractors, and a second channel regulator. The structure of the feature extractor is specifically referred to Figure 5 As shown, the feature extraction process according to the structure of the feature extractor is as follows:
[0046] Step 31: Through the depth convolution in the set depthwise separable convolution, perform feature extraction operations on each channel of the feature map by applying a 3×3 convolution kernel respectively; during this process, the resolution size and channel dimension of the feature map both maintain a stable state and do not change. Use the depth convolution in the depthwise separable convolution to apply the convolution kernel to each channel of the feature map for feature extraction, keeping the resolution size and channel dimension of the feature map unchanged.
[0047] Step 32: Perform batch normalization operation on the feature extraction result in Step 31 to obtain the normalized feature map.
[0048] In this embodiment, the constructed feature extractor has a lightweight structure, and through multiple cascaded feature extractors, it is possible to extract feature maps with effective receptive field information at a lower computational cost, laying a foundation for subsequent multi-feature scale mining.
[0049] In this embodiment, referring to Figure 3 , the specific processing steps of the first cascaded multi-receptive field module and the second cascaded multi-receptive field module are as follows:
[0050] Step 41: Adjust the channels of the input feature map through the first channel regulator; where the input feature map is represented as:
[0051]
[0052] where B is the batch size, C represents the number of channels, and h and w are the height and width of the current feature map respectively. Since in this embodiment, the cascaded policy module sets N feature extractors, the first channel regulator realizes the dimensionality reduction of the output channels, and the number is adjusted to channels, so the feature map obtained after being processed by the first channel regulator The processing formula of the first channel regulator is expressed as follows:
[0053]
[0054] where C out is the final output channel number of the cascaded multi-receptive field module; represents the pointwise convolution in a depthwise separable convolution, which accepts the number of input channels C in of the feature map X in , and adjusts the feature map X′ to output channels; BN(·) represents the batch normalization operation; Act(·) represents the non-linear activation function. In this embodiment, to improve the convergence speed and performance during model training, the first channel regulator uses the Gaussian error linear unit as the activation function.
[0055] Step 42: Utilize the characteristic of information redundancy in multiple channels of the feature map to perform a channel division operation on the feature map X′, that is, perform an average channel division on the feature map X′ to obtain a first feature map and a second feature map;
[0056] Specifically, starting from the leftmost channel of the feature map X′, numbering from 1 to Divide the feature map X′ into a first feature map and a second feature map according to the parity of the channel numbers
[0057] Step 43: Fuse the first feature map and the second feature map by element-wise summation to obtain a third feature map; Specifically, fuse the first feature map X′ odd with the second feature map X′ even to enhance the richness of the feature information and obtain a third feature map The element-wise summation formula is expressed as follows:
[0058]
[0059] where represents the element-wise summation operation
[0060] Step 44: Input the first feature map X′ odd / the second feature map X′ even into the cascaded strategy module, and extract feature maps with information of multiple different receptive fields through N cascaded feature extractors. The implementation formula of the feature extractor is expressed as follows:
[0061]
[0062] where represents a depth convolution with a convolution kernel size of 3, where both the number of input and output channels are
[0063] Step 45: Perform a channel concatenation operation on the feature maps extracted by each level of the feature extractor and the third feature map along the channel direction to obtain a fourth feature map X″′, and output by restoring the channels of the fourth feature map to the number of input channels through a second channel regulator, that is where the calculation formula of the second channel regulator is expressed as follows:
[0064]
[0065] This embodiment carefully designs a cascaded multi-receptive field module for efficient multi-feature scale mining. The core principle of this module is to fully mine the redundant information contained in multiple channels in the feature map, and use a multi-scale feature extraction strategy to deeply explore the different scale information in the feature map. In terms of design, it cleverly maintains the lightweight characteristics, and can achieve low computational cost to extract feature maps with effective receptive field information, so as to integrate the information of multiple receptive fields in a network layer to improve feature representation, and finally achieve excellent segmentation performance and feature representation capabilities, thereby greatly improving the overall performance.
[0066] In this embodiment, a category regulator is also provided at the output end of the feature decoding branch, referring to Figure 6 As shown, this is the structural diagram of the category regulator provided in this embodiment. The category regulator uses the set point convolution layer to implement the point convolution operation in the depth-separable convolution, and adjusts the final output feature map channel of the feature decoding branch to D+1, that is, D+1 represents the sum of the number of pixel categories of the input feature map and the background, so as to adapt to the subsequent classification task requirements and lay the foundation for accurate image segmentation.
[0067] This embodiment refers to Figure 1 The figure shows the structure of the improved U-Net model, which is a U-shaped network consisting of four layers of corresponding encoders and decoders. Each encoder consists of a first cascade and a downsampling module, and each decoder consists of an upsampling module, a channel connection block and a second cascade multi-receptive field module.
[0068] In the encoder, the output feature map of the cascaded multi-receptive field module not only points to the downsampling module below, but also feeds to the channel connection block through skip connections. The downsampling module uses a max-pooling operation, where the pooling kernel and pooling stride are both set to 2 for double downsampling.
[0069] The encoder from top to bottom gradually expands the channel output dimension and gradually reduces the scale of the feature map resolution. The channel output dimensions are 64, 128, 256, and 512 respectively. The outputs of the first cascade multi-receptive field module in the encoder from top to bottom are (H, W), H∈N + and W∈N + are the height and width of the image respectively.
[0070] In the decoder, an upsampling module is used to perform a two-fold upsampling operation on the feature map using a bicubic interpolation algorithm. Subsequently, the feature map obtained by double upsampling and the feature map of the corresponding encoder number from the feature decoding branch are input into the channel connection block for splicing and integration in the channel direction, and finally input into the second cascade multi-receptive field module;
[0071] In the feature decoding branch, the decoder from bottom to top gradually reduces the channel output dimension and gradually enlarges the scale of the feature map resolution. The channel output dimensions are successively: 512, 256, 128, 64. The outputs of the second cascaded multi-receptive field module in the encoder from bottom to top are successively (H, W), where H ∈ N + and W ∈ N + which are the height and width of the image respectively.
[0072] In this embodiment, the feature encoding branch and the feature decoding branch are successively concatenated to form a U-shaped neural network structure. The skip connection method is adopted to connect and concatenate the encoder i in the feature encoding branch and the decoder i in the feature decoding branch by means of a channel connection block, thereby constructing an associated path between the two.
[0073] In this embodiment, in order to further optimize the training effect, a cosine annealing learning rate strategy is introduced during the model training process. The minimum learning rate is set to 1×10 -6 , so that the model can dynamically adjust the learning rate according to different stages during the training process, and thus converge to the optimal solution faster; and the weighted sum of cross entropy and Dice loss is used as the final loss function to guide the training of the model. This combination can fully consider different types of errors in the image segmentation task, enabling the model to more comprehensively optimize the segmentation results during the learning process.
[0074] The implementation process of the method in this embodiment is described below through a specific case.
[0075] This case is implemented in a specific hardware environment. Specifically, an NVIDIA GeForce RTX4090 GPU with a memory of 24GB is selected as the computing platform, and this method is constructed and run based on the advanced programming language PyTorch. During the model training process, the Adam optimizer is selected, and its learning rate is set to 1×10 -4 , and at the same time, the initial value of momentum decay is determined to be 0.9; in the industrial vertical domain image segmentation dataset experiment, it is divided into a training set, a validation set, and a test set in a ratio of 6:2:2.
[0076] For the input image, its resolution is uniformly set to 256×256 (i.e., H = W = 256) to ensure that the image data has a consistent format and scale when entering the model; the number of iterations for model training is set to 300 times. After 300 iterations, the training process stops to avoid overfitting.
[0077] To improve the generalization ability of the model, various data augmentation operations are applied during each training to increase the diversity of the training data. These operations include random rotation with an angle between ±25°, random horizontal and vertical shift operations with a shift ratio of 15%, and random flipping operations in the horizontal and vertical directions. These operations can effectively expand the training dataset, making the model more adaptable when processing industrial images with different shapes, shooting conditions, and poses. At the same time, to balance the data augmentation effect and training efficiency, after 240 epochs, the data augmentation operations were stopped, and then the model was continuously trained until the preset number of iterations was reached. In addition, according to the characteristics and requirements of industrial images, the initial image input channel in this case was set to 3 to ensure that the model can make full use of the image color information to achieve accurate feature extraction and segmentation.
[0078] The beneficial effects of the method of this embodiment are as follows:
[0079] 1) The method of this embodiment designs a lightweight industrial vertical domain image segmentation method with efficient multi-feature scale mining. This network is an efficient U-shaped neural network that achieves excellent segmentation performance under the conditions of using a small number of parameters and low computational complexity.
[0080] 2) The method of this embodiment is different from traditional lightweight industrial vertical domain image segmentation methods. This method designs an efficient multi-feature scale mining module based on a multi-scale feature extraction strategy to fuse the information of multiple receptive fields in a network layer to enhance feature representation, so as to achieve image segmentation with lighter weights and better performance, and finally achieve excellent segmentation performance.
[0081] 3) The proposed efficient multi-feature scale mining module of the method of this embodiment has generality and flexibility, is easy to be extended to other networks, and at the same time reduces the parameters and computational complexity and improves the segmentation performance.
[0082] Device embodiment
[0083] According to an embodiment of the present invention, a lightweight industrial vertical domain image segmentation device is provided, as Figure 7 shown, which is a block diagram of the lightweight industrial vertical domain image segmentation device provided in this embodiment. The lightweight industrial vertical domain image segmentation device according to an embodiment of the present invention includes:
[0084] A training data acquisition module 10, configured to acquire an industrial vertical domain image segmentation dataset;
[0085] The model construction module 20 is used to construct an improved U-Net model. The encoder of the feature encoding branch of the improved U-Net model includes a first cascaded multi-receptive field module and a downsampling module connected thereto. Each cascaded decoder of the feature decoding branch includes a second cascaded multi-receptive field module, a channel connection block, and an upsampling module. Each decoder splices and integrates the feature map processed by the upsampling module and the feature map output from the corresponding layer encoder in the feature decoding branch in the channel direction. The integrated feature map is input to the second cascaded multi-receptive field module.
[0086] The first cascaded multi-receptive field module and the second cascaded multi-receptive field module are respectively provided with a first channel regulator, a cascaded strategy module composed of multiple cascaded feature extractors, and a second channel regulator.
[0087] The first channel regulator is used to keep the resolution size of the input feature map unchanged and adjust the number of output channels to according to the preset N feature extractors in the cascaded strategy module, and perform channel average division on the feature map processed by the first channel regulator to obtain a first feature map and a second feature map. The first feature map and the second feature map are fused by element-wise summation to obtain a third feature map, and the first feature map / second feature map is input to the cascaded strategy module. The cascaded strategy module uses depth convolution in depthwise separable convolution to perform feature extraction operations on each channel of the first feature map / second feature map by using a C×C convolution kernel respectively, so as to perform feature extraction on each channel of the first feature map / second feature map respectively, keeping the resolution size and channel dimension of the feature map unchanged. The feature maps extracted by each feature extractor and the third feature map are subjected to channel connection operation along the channel direction to obtain a fourth feature map, and the second channel regulator restores the number of channels of the fourth feature map to the input channel number for output.
[0088] The model training module 30 is used to input an industrial vertical domain image segmentation data set into the improved U-Net model for model training until the convergence condition is reached, so as to obtain a segmentation model for realizing industrial vertical domain image segmentation.
[0089] A lightweight industrial vertical domain image segmentation device provided in this embodiment. The model construction module 20 uses a U-shaped neural network, and each feature encoder and feature decoder are provided with a cascaded multi-receptive field module that can realize the mining of deep multi-scale information of the image feature map. This module is provided with a channel regulator and a feature extractor. The first channel regulator at the input end adjusts the number of output channels to The number of channels is [number], and the second-channel regulator at the output end restores the feature map channels, enabling the number of channels for custom channel upsampling and downsampling to dynamically adjust the fusion weights of different-scale features according to the image content, thereby effectively mining richer, more representative, and more detailed feature information, and exploring a more compact convolutional structure to reduce the computational amount. Additionally, the cascading strategy formed by multiple feature extractors is set up to configure a C×C convolutional kernel for each channel to perform feature extraction operations, so as to deeply explore different-scale information within the feature map. Moreover, it cleverly maintains the lightweight property while effectively enhancing the feature representation ability by fusing information from different receptive fields in a single network layer, thus significantly improving the overall performance. Furthermore, the method of this embodiment achieves excellent segmentation performance under the conditions of using a small number of parameters and low computational complexity through an efficient U-shaped neural network. The cascaded multi-receptive field module set in this embodiment has generality and flexibility, is easy to extend to other networks, and at the same time reduces the parameters and computational complexity and improves the segmentation performance.
[0090] This embodiment further includes an image preprocessing module 40, which is configured to perform the following image preprocessing operations:
[0091] Step 11, image normalization: Adjust the size of all images so that their resolutions all reach a unified pixel size of H×H, such as 256×256 pixels. Then convert the sized images into the PyTorch tensor format. For the image data with pixel values in the range of [0, H], use the min-max normalization algorithm for processing, and finally obtain the preprocessed image as the input feature map.
[0092] Step 12, image segmentation category encoding: Set the category ID of the background pixels of each preprocessed image to 0, and sequentially assign different category pixel IDs from 1 to D to the different category pixels segmented in the image, where D represents the total number of pixel categories other than the background. Convert the result obtained from the pixel category encoding into the PyTorch tensor form, which is used as the ground truth label and combined with the loss function to carry out the model training work.
[0093] In this embodiment, the first and second channel regulators mine non-linear feature information through the set 1x1 convolution operation for feature fusion. Their main function is to flexibly set and adjust the number of channel dimensions according to requirements while maintaining the resolution size of the feature map unchanged, thereby effectively mining richer, more representative feature information. Refer to Figure 4 , the first and second channel regulators include a pointwise convolution layer, a batch normalization operation layer, and an activation function module; among them,
[0094] The pointwise convolution layer adjusts the number of channels in the input feature map according to the preset number of channel upsampling or channel downsampling, and then obtains the feature map after channel adjustment processing; the batch normalization operation layer performs batch normalization operation on the feature map obtained after adjusting the channel dimension to obtain the normalized feature map; finally, the Gaussian error linear unit is used in the activation function module to perform nonlinear processing on the batch-normalized feature map to obtain the feature map corresponding to the number of channels.
[0095] In the cascade strategy module composed of multiple cascaded feature extractors in this embodiment, each feature extractor is provided with a convolution layer and a batch normalization operation layer;
[0096] The convolution layer uses each channel of the 3×3 convolution kernel feature map to perform feature extraction operations; during this process, the resolution size and channel dimension of the feature map both maintain a stable state and do not change. The depth convolution in the depthwise separable convolution is used to apply the convolution kernel to each channel of the feature map separately for feature extraction, keeping both the resolution size and channel dimension of the feature map unchanged.
[0097] Then the batch normalization operation layer performs batch normalization operation on the features extracted by the convolution layer to obtain the normalized feature map.
[0098] In this embodiment, the feature map processed by the first channel regulator is divided by channel average to obtain the first feature map and the second feature map, specifically, according to the feature map obtained after being processed by the first channel regulator, the first feature map and the second feature map are obtained according to the parity of the channel numbers of the feature map.
[0099] In this embodiment, a class regulator is further provided at the output end of the feature decoding branch. Refer to Figure 6 As shown, it is the structure diagram of the class regulator provided in this embodiment. The class regulator realizes the point convolution operation in the depthwise separable convolution through the set point convolution layer, and adjusts the channels of the feature map finally output by the feature decoding branch to D + 1, that is, D + 1 represents the sum of the pixel categories of the input feature map and the background, so as to adapt to the subsequent classification task requirements and lay a foundation for accurate image segmentation.
[0100] The embodiment of the present invention is the device embodiment corresponding to the above method embodiment. The specific operations of each module processing step can be understood with reference to the description of the method embodiment, and will not be elaborated here.
[0101] As Figure 8 shown, the present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the lightweight industrial vertical domain image segmentation method in the above embodiment.
[0102] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the lightweight industrial vertical domain image segmentation method in the above embodiments, or when the computer program is executed by a processor, it implements the lightweight industrial vertical domain image segmentation method in the above embodiments.
[0103] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0104] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the device or system embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, refer to the description of the method embodiments. The device and system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.
[0105] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention, and the content not described in detail in the specification of the present invention belongs to the well-known technology of those skilled in the art.
Claims
1. A lightweight industrial vertical image segmentation method, characterized in that: The following steps are involved: Obtain industrial vertical image segmentation dataset; An improved U-Net model is constructed, wherein the encoder of the feature encoding branch of the improved U-Net model includes a first cascade multi-receptive field module and a downsampling module connected thereto, and the decoders of each cascade of the feature decoding branch include a second cascade multi-receptive field module, a channel connection block and an upsampling module; the decoders of each stage splice and integrate the feature map processed by the upsampling module with the feature map output from the encoder of the corresponding layer in the feature decoding branch according to the channel direction, and the spliced feature map is input into the second cascade multi-receptive field module; the first cascade multi-receptive field module and the second cascade multi-receptive field module are respectively provided with a first channel regulator, a cascade strategy module based on multiple cascade feature extractors, and a second channel regulator; The first channel regulator is used to maintain the resolution size of the input feature map unchanged and adjust the number of output channels to N according to the preset N feature extractors in the cascade strategy module. , N ≥ 1, and perform channel average division on the feature map processed by the first channel regulator to obtain a first feature map and a second feature map, fuse the first feature map and the second feature map by element-by-element summation to obtain a third feature map, input the first feature map / the second feature map to the cascade strategy module, the cascade strategy module uses the depth convolution in the depthwise separable convolution to apply the convolution kernel to the first feature map / the second feature map channel by channel for feature extraction, and the feature maps extracted by the feature extractors at all levels are connected to the third feature map along the channel direction to obtain a fourth feature map, and the channels of the fourth feature map are restored to the number of input channels by the second channel regulator for output; The industrial vertical image segmentation dataset is input into the improved U-Net model for training until the convergence condition is reached, thereby obtaining a segmentation model for realizing industrial vertical image segmentation.
2. The lightweight industrial vertical image segmentation method according to claim 1, characterized in that: Also includes: The image preprocessing is performed on the industrial vertical image segmentation dataset, which specifically includes the following preprocessing steps: Image standardization: resize all images to achieve a uniform resolution of H×H pixels, and then use the maximum and minimum normalization algorithm to process image data with pixel values in the interval [0, H], and finally obtain the preprocessed image as the input feature map; Image segmentation category encoding; A unique category code is set for the background pixels of each preprocessed image, and different unique codes are assigned to the pixels of different categories segmented from the image in sequence, and then the corresponding unique codes are converted into PyTorch tensor form as the real label.
3. The lightweight industrial vertical image segmentation method according to claim 1, characterized in that: The first channel regulator and the second channel regulator implement the channel regulation process as follows: The input feature map is adjusted through a point-by-point convolution layer to adjust the number of channel dimensions according to the preset number of channel dimension increases or channel dimension decreases, thereby obtaining a feature map after channel adjustment processing; Then, the feature map obtained after adjusting the channel dimension is batch normalized to obtain the normalized feature map; The batch normalized feature map is nonlinearly processed using the Gaussian error linear unit activation function to obtain a feature map with the corresponding number of channels.
4. The lightweight industrial vertical image segmentation method according to claim 1, characterized in that: The feature extraction process of the feature extractor is as follows: Through the depthwise convolution in the depthwise separable convolution, a 3×3 convolution kernel is used to perform feature extraction for each channel of the feature map. The extracted features are batch normalized to obtain the normalized feature map.
5. The lightweight industrial vertical image segmentation method according to claim 1, characterized in that: The specific implementation process of performing channel average division on the feature map processed by the first channel regulator to obtain the first feature map and the second feature map is as follows: The feature map obtained after the channel processing of the first channel regulator is divided according to the parity of the feature map channel number to obtain the first feature map and the second feature map.
6. The lightweight industrial vertical image segmentation method according to claim 1, characterized in that: A category regulator is also set at the output end of the feature decoding branch, and the set point convolution layer implements the point convolution operation in the depth-separable convolution, and adjusts the final output feature map channel of the feature decoding branch to the sum of the number of pixel categories of the input feature map and the background.
7. A lightweight industrial vertical image segmentation device, characterized in that: include: Training data acquisition module, used to obtain industrial vertical image segmentation dataset; A model construction module is used to construct an improved U-Net model, wherein the encoder of the feature encoding branch of the improved U-Net model includes a first cascade multi-receptive field module and a downsampling module connected thereto, and the decoders of each cascade of the feature decoding branch include a second cascade multi-receptive field module, a channel connection block and an upsampling module; the decoders of each stage splice and integrate the feature map processed by the upsampling module with the feature map output from the corresponding layer encoder in the feature decoding branch according to the channel direction, and the spliced feature map is input into the second cascade multi-receptive field module; the first cascade multi-receptive field module and the second cascade multi-receptive field module are respectively provided with a first channel regulator, a cascade strategy module based on multiple cascade feature extractors, and a second channel regulator; The first channel regulator is used to maintain the resolution size of the input feature map unchanged and adjust the number of output channels to N according to the preset N feature extractors in the cascade strategy module. , N ≥ 1, and perform channel average division on the feature map processed by the first channel regulator to obtain a first feature map and a second feature map, fuse the first feature map and the second feature map by element-by-element summation to obtain a third feature map, input the first feature map / the second feature map to the cascade strategy module, the cascade strategy module uses the depth convolution in the depthwise separable convolution to apply the convolution kernel to the first feature map / the second feature map channel by channel for feature extraction, and the feature maps extracted by the feature extractors at all levels are connected to the third feature map along the channel direction to obtain a fourth feature map, and the channels of the fourth feature map are restored to the number of input channels by the second channel regulator for output; The model training module is used to input the industrial vertical image segmentation dataset into the improved U-Net model for training until the convergence condition is reached, so as to obtain a segmentation model for realizing industrial vertical image segmentation.
8. The lightweight industrial vertical image segmentation device according to claim 7, characterized in that: The first cascade multi-receptive field module and the second cascade multi-receptive field module perform channel average division on the feature map processed by the first channel regulator to obtain the first feature map and the second feature map. The specific implementation process is as follows: The feature map obtained after the first channel regulator channel processing is divided according to the parity of the feature map channel number to obtain the first feature map and the second feature map.
9. The lightweight industrial vertical image segmentation device according to claim 7, characterized in that: A category regulator is also set at the output end of the feature decoding branch, which is used to implement the point convolution operation in the depth-separable convolution according to the set point convolution layer, and adjust the final output feature map channel of the feature decoding branch to the sum of the number of pixel categories of the input feature map and the background.
10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the lightweight industrial vertical image segmentation method as described in any one of claims 1 to 7 is implemented.