Metal section image segmentation method based on deep learning
Through overlapping blocking and the improved U-Net++ model combined with mixed loss function and edge detection operator, the problem of insufficient feature extraction in metal section image segmentation is solved, and high-precision metal section image segmentation is achieved, and continuous and consistent segmentation results are generated.
Patent Information
- Application Number
- CN202510338497.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-04
AI Technical Summary
Traditional image segmentation methods are difficult to effectively extract complex multi-scale features and fine structures in metal section images, especially the U-Net model is insufficient in metal section image segmentation.
The image is cropped using overlapping chunking, combined with the improved U-Net++ model for training, the model is optimized using a hybrid loss function, and the image edge is processed using edge detection operators and closed operations. Image stitching is realized through the pixel value averaging processing of the overlapping area, and metal section image segmentation is performed.
Continuous and consistent segmentation of metal section images is realized, segmentation accuracy and completeness are improved, multi-level features can be accurately captured, clear and continuous edge textures are generated, and the accuracy and continuity of segmentation results are improved.
Smart Images

Figure CN120259659A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image segmentation, and in particular to a method for segmenting metal cross-section images based on deep learning. Background Technique
[0002] In the field of materials science and engineering, the study of metal cross-section behavior is a key link in optimizing the performance of metal materials, and the study of cross-section behavior is based on cross-section image segmentation. Based on the results of metal cross-section image segmentation, characteristic parameters such as dimple size and dimple depth can be extracted therefrom, so as to quantitatively analyze the mechanical properties of metal materials such as fracture toughness, tensile strength, and yield strength.
[0003] However, traditional image segmentation methods usually require artificial design of algorithms to extract the features of different regions in the image. When facing the complex structure in metal cross-section images, it is difficult to extract detailed features. With the development of deep learning technology in recent years, it has opened up a brand-new path for the field of image segmentation. The U-Net model, as a landmark network in the field of image segmentation, has been widely used in biomedical image segmentation. This network integrates four levels of progressive downsampling and corresponding four levels of upsampling paths to form a U-shaped symmetric structure, which can effectively recover feature extraction and spatial information. However, in the face of the complex and variable multi-scale features and fine structures in metal cross-section images, the performance of the traditional U-Net model in feature capture and detail retention still needs to be improved. Summary of the Invention
[0004] Object of the Invention: To solve the problem that existing image segmentation methods are difficult to extract detailed features when facing the complex structure in metal cross-section images, the present invention proposes a method for segmenting metal cross-section images based on deep learning.
[0005] Technical Solution: A method for segmenting metal cross-section images based on deep learning includes the following steps:
[0006] Step 1: Collect metal cross-section images, perform segmentation annotation on each metal cross-section image to form a corresponding label image; randomly divide and crop the metal cross-section images and the corresponding label images to form a data set;
[0007] Step 2: Use the data set formed in Step 1 to train the U-Net++ model to determine the model weight parameters;
[0008] Step 3: Perform block processing on the metal cross-section image to be segmented to form an image of the same size as the training set image, and ensure that there is a certain overlapping area between adjacent blocks. Use the trained U-Net++ model to segment the block images, splice the segmentation results of all the block images obtained, and form a complete segmentation image. During the splicing process, perform an average processing on the pixel values in the overlapping area between adjacent block images;
[0009] Step 4: Threshold the complete segmentation image to convert it into a binary image, perform smoothing processing on the binary image using Gaussian filtering, and perform edge detection on the image after Gaussian filtering using an edge operator to obtain an edge image;
[0010] Step 5: Perform skeletonization processing and closing operation processing on the edge image to obtain the final metal cross-section image segmentation result.
[0011] Further, in Step 2, when training the U-Net++ model, the following hybrid loss function is used:
[0012]
[0013] In the formula, N represents the number of samples, y i represents the true label of the i-th sample, taking values of 0 or 1, x i represents the output of the model, σ(·) represents the sigmoid function, X represents the predicted value, and Y represents the true value.
[0014] Further, in Step 3, the average processing on the pixel values in the overlapping area between adjacent block images includes:
[0015] For each pixel in the overlapping area, superimpose the pixel values of the segmentation images of each block at this position, and record the number of superimpositions at the same time. Then, the pixel value of the finally formed complete segmentation image can be expressed as:
[0016]
[0017] In the formula, (x, y) represents the coordinates of a certain pixel, n(x, y) represents the number of block superimpositions at the coordinates (x, y). For pixels in non-overlapping areas, n(x, y) = 1, p j (x, y) represents the pixel value of the j-th block image at the coordinate position (x, y).
[0018] Beneficial effects: Compared with the prior art, the present invention has the following advantages:
[0019] (1) The present invention uses an overlapping block method to crop the input image, so that it can process metal cross-section images of any size, and uses overlapping splicing to overcome the obvious discontinuity at the splicing of each block image, and a continuous and consistent segmented image can be obtained;
[0020] (2) As an improved U-Net model, the U-Net++ model significantly enhances the network's ability to represent complex image features and segmentation accuracy by introducing multi-level skip connections and dense feature fusion mechanisms, breaking through the limitations of the traditional U-Net model. The present invention applies the U-Net++ model to metal cross-section image segmentation, which can capture multi-level feature information in metal cross-section images, thereby obtaining more accurate segmentation results;
[0021] (3) The present invention uses a hybrid loss function, which combines two loss terms: cross-entropy loss and Dice loss. By optimizing this hybrid loss function, it can effectively guide the model to pay more attention to the key features of metal cross-section images during training and improve the accuracy of segmentation results;
[0022] (4) The present invention uses edge detection operators and closing operations. The edge detection operator effectively refines and highlights the edge parts in the image by calculating the gradients of the image in the vertical and horizontal directions, improving the accuracy of edge detection. The closing operation can suppress the information in non-edge regions and generate clearer and continuous edge textures; combining the two can improve the integrity and continuity of the generated segmented image while ensuring the accuracy of image segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 is a processing flow chart of a method for segmenting metal cross-section images based on deep learning according to the present invention;
[0024] Figure 2 is a schematic diagram of the block and overlapping splicing processing of metal cross-section images according to the present invention;
[0025] Figure 3 is a schematic diagram of the network structure used in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] The technical solutions of this embodiment will be further described below in conjunction with the accompanying drawings and embodiments.
[0027] This embodiment proposes a method for segmenting metal cross-section images based on deep learning, as Figure 1 shown, mainly including the following steps:
[0028] Step 1: Collect metal cross-section images and use the image annotation tool LabelImg to perform edge segmentation annotation on the collected metal cross-section images to generate corresponding label images. Since the U-Net++ model requires input images of a fixed size, the original metal cross-section images and the corresponding label images need to be randomly intercepted according to a certain block size, and then data augmentation processing such as random flipping, rotation, hue, brightness, and contrast adjustment is performed to form a dataset for model training.
[0029] Step 2: Adopt U-Net++ as the image segmentation model. As Figure 3 shown, the U-Net++ model is an improvement on the ordinary U-Net structure, enabling feature extraction from images at multiple scales. Among them, the encoder gradually extracts the features of the input image through multiple convolutional operations and pooling layers, capturing semantic information at different levels. While in the decoder, upsampling is used to add skip connections to connect the feature maps, fusing deep-level information with shallow-level information to make up for the lost details in the image and restore the pixel information of the image. The U-Net++ model also introduces nested skip connections, which include not only the traditional direct connections in the U-Net model but also multiple intermediate connections, thus forming a "nested" structure. Each layer of the decoder is connected to multiple layers of the encoder through multiple paths. These connections help the model fuse at different feature levels, enhancing the ability to capture image details.
[0030] Use the dataset formed in Step 1 to train the U-Net++ model to determine the model weight parameters. During the model training process, a hybrid loss function BCEDiceLoss is introduced. This hybrid loss function combines binary cross-entropy loss and Dice loss. Binary cross-entropy loss evaluates the difference between the predicted probability distribution and the target binary label, and is used to measure the prediction accuracy of the model during training; while Dice loss calculates the degree of overlap between the prediction result and the target ground truth in space, thereby enhancing the sensitivity of the model to the target boundary. The hybrid loss function assigns different weight coefficients to the BCE loss and Dice loss to adjust the proportion of the two. To avoid numerical instability problems caused by the denominator being zero, a small smoothing term is introduced into the hybrid loss function, as shown in the following formula:
[0031]
[0032] In the formula, N represents the number of samples, y i represents the true label of the i-th sample, taking values of 0 or 1, x i represents the output of the model, σ(·) represents the sigmoid function, X represents the predicted value, and Y represents the true value.
[0033] Step 3: The metal cross-section image to be segmented is divided into blocks according to the same size as the training set images. To avoid obvious discontinuities at the joints of the segmented images of each block, there is a certain overlapping area between adjacent blocks. The trained U-Net++ model is used to segment the block images to obtain the segmentation results of the blocks, and then the segmentation results of each block are stitched together to form a complete segmented image. During the stitching process, the overlapping area is averaged, that is, in the overlapping area, the sum of pixel values is divided by the number of overlaps to obtain the average pixel value, realizing a smooth transition at the block joints, reducing the stitching traces, and forming a continuous and complete metal cross-section segmented image. The specific stitching process is as follows:
[0034] As Figure 2 shown, the upper left corner of the figure is the segmented images of four adjacent blocks (represented by four square frames in red, green, blue, and yellow respectively). For each pixel in the overlapping area, the pixel values of the segmented images of each block at that position are superimposed, and the number of superimpositions is recorded at the same time. Then the pixel value of the final complete segmented image formed by stitching can be expressed as:
[0035]
[0036] In the formula, (x, y) represents the coordinates of a certain pixel, n(x, y) represents the number of block superimpositions at the coordinates (x, y). For pixels in non-overlapping areas, n(x, y) = 1, p j (x, y) represents the pixel value of the jth block image at the coordinate position (x, y).
[0037] In the specific implementation of the program, first, two two-dimensional arrays of all zeros with the same dimension as the image to be segmented are initialized. One is used to store the result of pixel value superposition (denoted as A), and the other is used to record the number of superpositions (denoted as B). Then, the pixel values of each segmented block image are superimposed on the corresponding element positions of the A array one by one, and the corresponding number of superpositions in the B array is updated at the same time. After the superposition is completed, the elements at the same positions in the two arrays are divided to obtain the average value, which is the value at the corresponding pixel position of the complete segmented image.
[0038] Step 4: Threshold processing is performed on the complete image formed by stitching to convert it into a binary image, that is, pixels greater than the set threshold are set to white, while pixels less than or equal to the threshold are set to black. Then, Gaussian filtering is used to smooth the binary image to reduce the noise in the image. Gaussian filtering uses the following Gaussian kernel function to perform two-dimensional filtering on the image:
[0039]
[0040] Wherein, a and b are respectively the offsets of the current position relative to the center, σ represents the standard deviation of the Gaussian kernel function, and the shape and scope of action of the filter are controlled by adjusting σ. The weight G(a,b) of the Gaussian kernel function is positively correlated with the distance from the current position to the center. The normalization factor 1 / 2πσ 2 makes the sum of all weights of the Gaussian kernel function equal to 1, thereby ensuring that the brightness of the filtered image remains unchanged. Then, the Sobel operator is used to perform edge detection on the filtered image in the vertical and horizontal directions respectively to identify the vertical and horizontal edges in the image. The expression for edge detection using the Sobel operator is:
[0041]
[0042] Wherein, G x is obtained by convolving the horizontal Sobel operator with the image M and is used to extract the horizontal edge information of the image;
[0043]
[0044] Wherein, G y is obtained by convolving the vertical Sobel operator with the image M and is used to extract the vertical edge information of the image.
[0045] Combining the horizontal gradient G x and the vertical gradient G y , the approximate gradient of the image M is obtained:
[0046] G = |G x | + |G y |
[0047] G is used as the calculation result of the image pixels after being processed by the Sobel operator.
[0048] Step 5: Perform skeletonization processing on the image after Sobel edge detection, convert the image into its skeleton structure, and retain the center line and key contour features of the original image. Then, perform closing operation processing. The closing operation is a commonly used technique in morphological processing. It first performs dilation operation on the image and then performs erosion operation to fill the small gaps in the image, remove the sharp features in the image, and obtain a more continuous and smooth image. In the dilation and erosion operations of the closing operation, a 2×2 elliptical structural element is used, and the number of iterations is set to 1. After the closing operation processing, the final metal cross-section image segmentation result can be obtained.
[0049] Using the method disclosed in this embodiment can greatly improve the accuracy of metal cross-section image segmentation, help extract the characteristic parameters of metal materials from the segmentation result more accurately, and quantitatively analyze the mechanical properties of metal materials.
Claims
1. A method for segmenting metal cross-section images based on deep learning, characterized in that: It includes the following steps: Step 1: Collect metal cross-section images, perform segmentation annotation on each metal cross-section image to form a corresponding label image; randomly crop and process the metal cross-section images and the corresponding label images to form a dataset; Step 2: Use the dataset formed in Step 1 to train the U-Net++ model to determine the model weight parameters; Step 3: Process the metal cross-section image to be segmented into block images of the same size as the images in the training set, and there is a certain overlapping area between adjacent block images. Use the trained U-Net++ model to segment the block images, and splice the segmentation results of all the obtained block images to form a complete segmentation image. During the splicing process, average the pixel values in the overlapping area between adjacent block images; Step 4: Threshold the complete segmentation image to convert it into a binary image, perform smoothing processing on the binary image using Gaussian filtering, and perform edge detection on the Gaussian-filtered image using an edge operator to obtain an edge image; Step 5: Perform skeletonization processing and closing operation processing on the edge image to obtain the final metal cross-section image segmentation result.
2. The method for segmenting metal cross-section images based on deep learning according to claim 1, characterized in that: In Step 2, when training the U-Net++ model, the following hybrid loss function is adopted: where N represents the number of samples, y i represents the true label of the i-th sample, taking values of 0 or 1, x i represents the output of the model, σ(·) represents the sigmoid function, X represents the predicted value, and Y represents the true value.
3. A method for metal cross-section image segmentation based on deep learning according to claim 1, characterized in that: In Step 3, averaging the pixel values in the overlapping area between adjacent block images includes: For each pixel in the overlapping area, add up the pixel values of the segmentation results of each block image at this position, and record the number of times of addition at the same time. Then the pixel value of the finally formed complete segmentation image can be expressed as: In the formula, (x, y) represents the coordinates of a certain pixel, n(x, y) represents the number of times of block addition at the coordinates (x, y). For the pixels in the non-overlapping area, n(x, y) = 1, and pj(x, y) represents the pixel value of the jth block image at the coordinate position (x, y).