Metal fracture image semantic segmentation method and device based on transfer learning
Through the transfer learning-based method, pre-trained model and data augmentation technology, the problems of insufficient data and long training time in the semantic segmentation task of metal fracture image are solved, and efficient semantic segmentation is achieved, which significantly improves the accuracy and saves costs.
Patent Information
- Application Number
- CN202510205853.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-10
AI Technical Summary
The semantic segmentation task of metal fracture images is difficult to complete efficiently due to insufficient data and long training time, and metal fractures involve complex crystallographic relationships and texture features, resulting in heavy segmentation tasks.
Using a transfer learning-based method, the pre-trained ResNet-101 model and ImageNet1K dataset are used to combine semantic segmentation neural networks and data enhancement technology to perform semantic segmentation of metal fracture images. Specific steps include collection of image data sets, semantic segmentation mask annotation and data augmentation processing, migration of pre-trained image feature processing models to extract low-resolution features, setting training strategies, and training semantic segmentation neural network models.
Through the transfer learning method, the complexity and time consumption of model training are significantly reduced, and the semantic segmentation accuracy of metal fracture images is improved. The average accuracy is 94.71%, and the average cross-border ratio is 89.93%, which greatly saves labor costs.
Smart Images

Figure CN120125819A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of metal fracture semantic segmentation, and particularly to a metal fracture image semantic segmentation method and device based on transfer learning. Background Art
[0002] The analysis of metal fracture images is of extremely important significance in materials science and engineering practice, especially in failure analysis and material optimization. By analyzing the fracture images, the failure mechanism of materials can be understood, which can further help improve material properties and structural design. Its importance lies in providing valuable guidance for materials science, structural design, manufacturing processes, and failure prevention by revealing the microscopic mechanism of material fracture. It is an indispensable tool in engineering failure analysis and is crucial for ensuring the safety, reliability, and extending the service life of equipment.
[0003] Semantic segmentation is an important technology in the field of computer vision, and its goal is to identify and classify each pixel in an image. Different from object detection and classification tasks, semantic segmentation not only needs to detect objects in the image but also accurately determine the pixel range of each object. Therefore, semantic segmentation can generate a "pixel-level" annotation map with the same size as the input image, and each pixel in the map will be assigned to a specific category. Semantic segmentation tasks have been successfully applied in fields such as autonomous driving, medical image analysis, and scene understanding. Due to the complex crystallographic relationships and texture features involved in metal fractures, and the diverse fracture types, the entire process from collecting to analyzing metal fracture images is a very heavy and difficult task. Transfer learning, as a machine learning method, refers to applying the knowledge learned from one task to another related task. Its basic idea is to use a model that has been trained on a large-scale dataset (usually called a "pre-trained model") and apply it to another dataset with different but related tasks. This method reduces the need to train from scratch and can improve the performance of the model on the new task, solving the problems of insufficient data or excessive training time in the target task. Summary of the Invention
[0004] To solve the above technical problems or at least partially solve the above technical problems, the present invention provides a metal fracture image semantic segmentation method and device based on transfer learning.
[0005] In a first aspect, the present invention provides a metal fracture image semantic segmentation method based on transfer learning, including: Step 1: Collect metal fracture images of different fracture types as an image dataset;
[0006] Step 2: Perform semantic segmentation mask annotation and data augmentation processing on the metal fracture images in the image dataset;
[0007] Step 3: Migrate the pre-trained image feature processing model to extract low-resolution features from the metal fracture images in the image dataset. Specifically, preprocess the image datasets of various fractures and input them into the image feature processing model to obtain low-resolution features;
[0008] Step 4: Set the training strategy for the fracture image semantic segmentation task;
[0009] Step 5: Input the low-resolution features into the constructed semantic segmentation neural network for training so that the semantic segmentation neural network model can predict a segmentation result that matches the semantic segmentation mask annotation;
[0010] Step 6: Evaluate the semantic segmentation neural network model obtained from the transfer learning training.
[0011] Furthermore, the fracture types included in the metal fracture images for training the semantic segmentation task are: cleavage, quasi-cleavage, ductile, fatigue, and intergranular, a total of 5 fracture types. When collecting metal fracture images of different fracture types: select metal fracture images with mixed fracture types, that is, the metal fracture images contain 2 or more fracture types; the resolution of the selected metal fracture images is not less than 512×512 pixels; the texture features and context information contained in the metal fracture images support obtaining a clear fracture diagnosis conclusion; in the collected image dataset, the proportion difference between various fracture types does not exceed 10%.
[0012] Furthermore, the semantic segmentation mask annotation and data augmentation processing of the metal fracture images in the image dataset include:
[0013] Use the Real-ESRGAN model to perform super-resolution processing on each metal fracture image to obtain a super-resolution metal fracture image with a resolution greater than 1024×1024 pixels;
[0014] Use the LabelImg toolbox to perform semantic segmentation mask annotation on each super-resolution metal fracture image;
[0015] Use the Albumentations toolbox to perform image enhancement operations on each super-resolution metal fracture image. The image enhancement operations include flipping, random contrast and brightness, optical distortion, Gaussian noise, random blur, and translation and scaling.
[0016] Furthermore, in Step 3, transfer learning uses the pre-training results obtained by training based on the ImageNet1K dataset and the ResNet101 model as the weights of the image feature processing model to extract low-resolution features from the metal fracture images in the image dataset. The image feature processing model compresses the input data into a compact representation and extracts basic features while reducing the dimension.
[0017] Further, the preprocessing of the fracture image dataset divides the dataset into a training set and a validation set at a ratio of 9:1, and completes the preprocessing of grayscale conversion and normalization.
[0018] Further, the training strategy for setting the fracture image semantic segmentation task in step 4 includes the settings of the optimizer, learning rate, and number of training epochs. Among them, the optimizer type is Adam; the number of training epochs is 60,000 rounds; a dynamic learning rate is adopted, with a linear learning rate in the range of [0, 1000) rounds and a multi-step decay learning rate in the range of [1000, 60000) rounds; the loss function is the sum of the cross-entropy loss between the true mask annotation and the predicted mask and the cross-entropy loss between the predicted type and the type corresponding to the true mask.
[0019] Further, the semantic segmentation neural network described in step 5 includes: a pixel decoder that gradually upsamples low-resolution features. The pixel decoder is a feature pyramid network. The results of each level of the feature pyramid network are combined with position encoding to generate high-resolution embeddings of different levels. The high-resolution embeddings are flattened into a sequence. Among them, B, C, H, and W are the batch size, channel size, height, and width of the high-resolution embeddings respectively. The dimension of the flattened sequence is H×W, B, C; the calculation process of the high-resolution embeddings includes: calculating the position encoding of the results of each level of the feature pyramid network through cosine position encoding and flattening the position encoding; after projecting the results of each level of the feature pyramid network through a projection layer, obtaining the metal fracture image features of each level by adding the embedding level encoding weights, flattening the metal fracture image features, and combining them with the position encoding to obtain high-resolution embeddings; constructing learnable query features through an embedding layer, and using the self-attention and mask-guided cross-attention mechanisms in the attention decoder to incorporate the localized features of the high-resolution embeddings into the valid region of the mask predicted by the query features, extracting the localized features, focusing on the target local area, enabling the model to focus on the associated regions in the image, and considering the relationships between different objects and their features at the same time;
[0020] The mask prediction head obtains pixel weights based on the query features output by the mask attention decoder, and the result of the level with the maximum resolution of the pyramid network is combined with the pixel weights to obtain the predicted mask;
[0021] The type prediction head obtains the fracture type probability based on the query features output by the mask attention decoder, and takes the type with the maximum probability as the predicted fracture type.
[0022] Further, the mask prediction head uses a multi-layer perceptron, and the type prediction head uses a linear layer.
[0023] Further, in step 6, two metrics, namely the average accuracy rate and the average intersection over union, are used to evaluate the semantic segmentation neural network model obtained by transfer learning training.
[0024] In a second aspect, the present invention provides a semantic segmentation device for metal fracture images based on transfer learning, including: at least one processing unit, the processing unit is connected to a storage unit and a collection unit through a bus unit, the storage unit stores a computer program, and when the computer program is executed by the processing unit, the semantic segmentation method for metal fracture images based on transfer learning as described above is implemented.
[0025] The above technical solutions provided by the embodiments of the present invention have the following advantages compared with the prior art:
[0026] The present invention has the following advantages and effects compared with the prior art:
[0027] The collection of SEM images of metal fractures in the present invention includes 5 of the most common fracture types, and the images are all mixed fracture types, that is, each image contains at least two fracture types. After data augmentation processing on the basis of small-scale annotation, augmented images can be obtained to obtain a larger-scale data set. This greatly reduces the workload of collecting new data set images and their subsequent annotation, thus greatly saving labor costs.
[0028] The transfer learning of the present invention for the semantic segmentation task of fracture images is based on the pre-training results of the ResNet-101 deep neural network on the ImageNet1K data set, thus avoiding the complex process of training from scratch. This can enable the model to achieve rapid convergence under the condition of a short training cycle, significantly reducing the computing power overhead and time consumption.
[0029] The image feature processing model and the semantic segmentation neural network of the present invention, as the algorithm architecture of the transfer learning model, can obtain a relatively high model prediction accuracy rate, where the average accuracy rate is 94.71% and the average intersection over union is 89.93%. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present invention and used together with the specification to explain the principles of the present invention.
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0032] Figure 1Flowchart of a semantic segmentation method for metal fracture images based on transfer learning provided by an embodiment of the present invention;
[0033] Figure 2 Schematic diagram of the process of a semantic segmentation method for metal fracture images based on transfer learning provided by an embodiment of the present invention;
[0034] Figure 3 Schematic diagram of the semantic segmentation neural network provided by an embodiment of the present invention;
[0035] Figure 4 Schematic diagram showing the variation of the average accuracy mAcc and the average intersection over union mIoU with the number of iterations Epochs during the training process provided by an embodiment of the present invention;
[0036] Figure 5 Schematic diagram of a semantic segmentation device for metal fracture images based on transfer learning provided by an embodiment of the present invention. Detailed implementation manners
[0037] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0038] It should be noted that in this article, the terms "include", "comprise", or any other variation thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or device including a series of elements includes not only those elements but also other elements that are not explicitly listed, or elements that are inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device including the element.
[0039] Embodiment 1
[0040] As Figure 1 and Figure 2 shown, the present invention technically implements a semantic segmentation method for metal fracture images based on transfer learning, including:
[0041] Step 1: Collect different types of metal fracture images as the training dataset for the deep neural network;
[0042] Specifically, the collection of the metal fracture images is carried out by consulting public books, academic literature, conference reports, etc. in related fields such as mechanical tests of metal materials and failure analysis, and obtaining 5 types of SEM images of fracture surfaces for training semantic segmentation tasks. The fracture types included in the metal fracture images for training semantic segmentation tasks are: cleavage, quasi-cleavage, ductile, fatigue, and intergranular, a total of 5 fracture types. When collecting metal fracture images of different fracture types: select metal fracture images with mixed fracture types, that is, the metal fracture images contain 2 or more fracture types; the resolution of the selected metal fracture images is not less than 512×512 pixels; the texture features and context information contained in the metal fracture images support obtaining a clear fracture diagnosis conclusion; in the collected image dataset, the proportion difference between various fracture types does not exceed 10%.
[0043] Step 2: Perform semantic segmentation mask annotation and data augmentation processing on the metal fracture images in the image dataset.
[0044] Specifically, first use the Real-ESRGAN model to perform super-resolution processing on the fracture images to obtain super-resolution metal fracture images with a resolution greater than 1024×1024. Improving the resolution of the fracture images through the Real-ESRGAN model can effectively improve the quality of the original data and provide a high-quality dataset for subsequent segmentation tasks.
[0045] Use the LabelImg toolbox to perform semantic segmentation mask annotation on each super-resolution metal fracture image; the semantic segmentation mask annotation information is saved as an XML file in the PASCAL VOC format.
[0046] Use the Albumentations toolbox to perform image enhancement operations on each super-resolution metal fracture image. The image enhancement operations include flipping, random contrast and brightness, optical distortion, Gaussian noise, random blur, and translation and scaling. After data augmentation processing based on small-scale annotation, augmented images can be obtained to get a larger-scale dataset. This greatly reduces the workload of collecting new dataset images and their subsequent annotation, thus greatly saving labor costs.
[0047] Step 3: Transfer the pre-trained image feature processing model to extract low-resolution features from the metal fracture images in the image dataset. Among them, the image datasets of various fractures are preprocessed and then input into the image feature processing model to obtain low-resolution features.
[0048] Specifically, transfer learning uses the pre-training results obtained by training based on the ImageNet1K dataset and the ResNet101 model as the weights of the image feature processing model to extract low-resolution features from the metal fracture images in the image dataset. The image feature processing model compresses the input data into a compact representation and extracts basic features while reducing the dimension. Moreover, by adopting the method of transfer learning, the complex process of training from scratch is avoided, enabling the model to achieve fast convergence under the condition of a short training cycle, significantly reducing the computing power overhead and time consumption.
[0049] Before extracting the low-resolution features, the preprocessing of the fracture image dataset after super-resolution and data augmentation includes: dividing the dataset into a training set and a validation set in a ratio of 9:1, and completing the preprocessing of grayscale conversion and normalization.
[0050] Step 4: Set the training strategy for the semantic segmentation task of the fracture image.
[0051] Specifically, the setting of the training strategy mainly includes: the optimizer type is Adam; the number of training epochs is 60,000; a dynamic learning rate is adopted, with a linear learning rate in the range of [0, 1000) epochs and a multi-step decay learning rate in the range of [1000, 60000) epochs; the loss function is the sum of the cross-entropy loss between the true mask annotation and the predicted mask and the cross-entropy loss between the predicted type and the type corresponding to the true mask.
[0052] Step 5: Input the low-resolution features into the constructed semantic segmentation neural network for training so that the semantic segmentation neural network model can predict a segmentation result that matches the semantic segmentation mask annotation.
[0053] Such as Figure 3As shown in the figure, the semantic segmentation neural network includes: a pixel decoder that gradually upsamples low-resolution features. The pixel decoder is a feature pyramid network. The results of each layer of the feature pyramid network are combined with position encoding to generate high-resolution embeddings of different layers. The high-resolution embeddings are flattened into a sequence. Here, B, C, H, and W are the batch size, channel size, height, and width of the high-resolution embeddings respectively. The dimension of the flattened sequence is H×W, B, C. The calculation process of the high-resolution embeddings includes: calculating the position encoding of the results of each layer of the feature pyramid network through cosine position encoding and flattening the position encoding; after projecting the results of each layer of the feature pyramid network through a projection layer, obtaining the metal fracture image features of each layer by adding the embedding layer encoding weights, flattening the metal fracture image features, and combining with the position encoding to obtain high-resolution embeddings; constructing learnable query features through an embedding layer, and using the self-attention and mask-guided cross-attention mechanisms in the attention decoder to integrate the localized features of the high-resolution embeddings into the valid region of the mask predicted by the query features, extracting the localized features, focusing on the target local area, enabling the model to focus on the associated regions in the image, and at the same time considering the relationships between different objects and their features.
[0054] The mask prediction head obtains pixel weights based on the query features output by the mask attention decoder, and the result of the layer with the maximum resolution of the pyramid network is combined with the pixel weights to obtain the predicted mask.
[0055] The type prediction head obtains the fracture type probability based on the query features output by the mask attention decoder, and takes the type with the maximum probability as the predicted fracture type.
[0056] As Figure 3 shown, the mask prediction head uses a multi-layer perceptron, and the type prediction head uses a linear layer.
[0057] By minimizing the loss function, the mask predicted by the semantic segmentation neural network is close to the real mask, and the predicted fracture type is close to the real fracture type.
[0058] Step 6: Evaluate the semantic segmentation neural network model obtained by transfer learning training.
[0059] Specifically, within the range of the validation set, two metrics, average accuracy and mean intersection over union, are used to evaluate the semantic segmentation neural network model obtained by transfer learning training. The semantic segmentation neural network model that passes the evaluation can be used for semantic segmentation of metal fracture images. As Figure 4 shown, during the training process, the average accuracy mAcc and the mean intersection over union mIoU change with the number of iteration rounds Epochs.
[0060] Example 2
[0061] Refer toFigure 5 As shown in Figure 5 , an embodiment of the present invention provides a semantic segmentation device for metal fracture images based on transfer learning, including: at least one processing unit, which is connected to a storage unit through a bus unit. The storage unit, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the software programs, computer-executable programs, and modules corresponding to a semantic segmentation method for metal fracture images based on transfer learning in an embodiment of the present invention. The processing unit realizes the above-mentioned semantic segmentation method for metal fracture images based on transfer learning by running the software programs, computer-executable programs, and modules stored in the storage unit, including:
[0062] Step 1: Collect metal fracture images of different fracture types as an image data set;
[0063] Step 2: Perform semantic segmentation mask annotation and data augmentation processing on the metal fracture images in the image data set;
[0064] Step 3: Transfer a pre-trained image feature processing model to extract low-resolution features from the metal fracture images in the image data set. Among them, the image data sets of various fractures are pre-processed and then input into the image feature processing model to obtain low-resolution features;
[0065] Step 4: Set the training strategy for the fracture image semantic segmentation task;
[0066] Step 5: Input the low-resolution features into the constructed semantic segmentation neural network for training so that the semantic segmentation neural network model can predict a segmentation result that matches the semantic segmentation mask annotation;
[0067] Step 6: Evaluate the semantic segmentation neural network model obtained by transfer learning training.
[0068] Of course, for the storage unit in the semantic segmentation device for metal fracture images based on transfer learning provided by an embodiment of the present invention, the computer program stored therein is not limited to the method operations described above, and can also execute related operations in a semantic segmentation method for metal fracture images based on transfer learning provided by any embodiment of the present invention.
[0069] Embodiment 3
[0070] An embodiment of the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed, it realizes the semantic segmentation method for metal fracture images based on transfer learning, including:
[0071] Step 1: Collect metal fracture images of different fracture types as an image data set;
[0072] Step 2: Perform semantic segmentation mask annotation and data augmentation processing on the metal fracture images in the image dataset;
[0073] Step 3: Transfer the pre-trained image feature processing model to extract low-resolution features from the metal fracture images in the image dataset. Among them, the image datasets of various fractures are pre-processed and then input into the image feature processing model to obtain low-resolution features;
[0074] Step 4: Set the training strategy for the fracture image semantic segmentation task;
[0075] Step 5: Input the low-resolution features into the constructed semantic segmentation neural network for training so that the semantic segmentation neural network model can predict a segmentation result that matches the semantic segmentation mask annotation;
[0076] Step 6: Evaluate the semantic segmentation neural network model obtained by transfer learning training. A computer-readable storage medium provided by an embodiment of the present invention stores a computer program that is not limited to the method operations described above, and can also execute related operations in a method for semantic segmentation of metal fracture images based on transfer learning provided by any embodiment of the present invention.
[0077] In the embodiments provided by the present invention, it should be understood that the disclosed structures and methods can be implemented in other ways. For example, the structural embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of structures or units can be in an electrical, mechanical or other form.
[0078] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0079] In addition, each functional unit in various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0080] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A method for semantic segmentation of metal fracture images based on transfer learning, characterized in that: include: Step 1: Collect metal fracture images of different fracture types as image datasets; Step 2: Perform semantic segmentation, mask annotation and data enhancement processing on the metal fracture images in the image dataset; Step 3: Migrate the pre-trained image feature processing model to extract low-resolution features from the metal fracture images in the image data set, wherein the image data sets of various fractures are pre-processed and input into the image feature processing model to obtain low-resolution features; Step 4: Set the training strategy for the fracture image semantic segmentation task; Step 5: Input the low-resolution features into the constructed semantic segmentation neural network for training so that the semantic segmentation neural network model can predict the segmentation results that match the semantic segmentation mask annotations; Step 6: Evaluate the semantic segmentation neural network model trained by transfer learning.
2. The method for semantic segmentation of metal fracture images based on transfer learning according to claim 1, characterized in that: The fracture types contained in the metal fracture images used to train the semantic segmentation task include: cleavage, quasi-cleavage, toughness, fatigue and intergranular fracture types; when collecting metal fracture images of different fracture types: select metal fracture images of mixed fracture types, that is, the metal fracture images contain 2 or more fracture types; the resolution of the selected metal fracture images is not less than 512×512 pixels; the texture features and contextual information contained in the metal fracture images support clear fracture diagnosis conclusions; in the collected image data set, the difference in the proportion of various fracture types does not exceed 10%.
3. The method for semantic segmentation of metal fracture images based on transfer learning according to claim 1, characterized in that: The semantic segmentation, mask annotation and data enhancement processing of the metal fracture images in the image data set includes: The Real-ESRGAN model is used to perform super-resolution processing on each metal fracture image to obtain a super-resolution metal fracture image with a resolution greater than 1024×1024 pixels; Use the LabelImg toolbox to perform semantic segmentation and mask annotation on each super-resolution metal fracture image; The Albumentations toolbox is used to perform image enhancement operations on each super-resolution metal fracture image. The image enhancement operations include flipping, random contrast and brightness, optical distortion, Gaussian noise, random blur, and translation and scaling.
4. The method for semantic segmentation of metal fracture images based on transfer learning according to claim 1, characterized in that: In step 3, transfer learning uses the pre-training results obtained based on the ImageNet1K dataset and the ResNet101 model as the weights of the image feature processing model to extract low-resolution features from the metal fracture images in the image dataset. The image feature processing model compresses the input data into a compact representation and extracts basic features while reducing the dimension.
5. The method for semantic segmentation of metal fracture images based on transfer learning according to claim 1, characterized in that: The fracture image data set preprocessing is to divide the data set into a training set and a validation set in a ratio of 9:1, and complete grayscale and normalization preprocessing.
6. The method for semantic segmentation of metal fracture images based on transfer learning according to claim 1, characterized in that: The training strategy for setting the fracture image semantic segmentation task described in step 4 includes the settings of the optimizer, learning rate and training cycle, wherein the optimizer type is Adam; the training cycle is 60,000 rounds; a dynamic learning rate is adopted, [0,1000) rounds are linear learning rates, and [1000,60000) rounds are multi-step decay learning rates; the loss function is the sum of the cross entropy loss between the true mask annotation and the predicted mask and the cross entropy loss between the predicted type and the type corresponding to the true mask.
7. The method for semantic segmentation of metal fracture images based on transfer learning according to claim 1, characterized in that: The semantic segmentation neural network described in step 5 includes: a pixel decoder for gradually upsampling low-resolution features, the pixel decoder is a feature pyramid network, the results of each level of the feature pyramid network are combined with the position encoding to generate high-resolution embeddings of different levels, and the high-resolution embeddings are flattened into a sequence, wherein B, C, H, and W are the batch size, channel size, height, and width of the high-resolution embedding, respectively, and the dimensions of the unfolded sequence are H×W, B, and C; the calculation process of the high-resolution embedding includes: calculating the position encoding of the results of each level of the feature pyramid network through cosine position encoding, and flattening the position encoding; after projecting the results of each level of the feature pyramid network through the projection layer, the metal fracture image features of each level are obtained by adding the embedding level encoding weights, the metal fracture image features are flattened, and the high-resolution embedding is obtained in combination with the position encoding; a learnable query feature is constructed through the embedding layer, and the self-attention and mask-guided cross-attention mechanism in the attention decoder is used to integrate the high-resolution embedded localized features into the effective area of the mask predicted by the query feature, extract the localized features, and focus on the target local area, so that the model can focus on the associated area in the image, while considering the relationship between different objects and their features; The mask prediction head obtains pixel weights based on the query features output by the mask attention decoder, and the maximum resolution level result of the pyramid network is combined with the pixel weights to obtain the predicted mask; The type prediction head obtains the fracture type probability based on the query features output by the masked attention decoder, and takes the type with the maximum probability as the predicted fracture type.
8. The method for semantic segmentation of metal fracture images based on transfer learning according to claim 7, characterized in that: The mask prediction head adopts a multi-layer perceptron, and the type prediction head adopts a linear layer.
9. The method for semantic segmentation of metal fracture images based on transfer learning according to claim 1, characterized in that: Step 6 uses the average accuracy and average intersection-over-union ratio to evaluate the semantic segmentation neural network model trained by transfer learning.
10. A metal fracture image semantic segmentation device based on transfer learning, characterized in that: include: At least one processing unit, wherein the processing unit is connected to a storage unit and an acquisition unit via a bus unit, wherein the storage unit stores a computer program, and when the computer program is executed by the processing unit, the method for semantic segmentation of metal fracture images based on transfer learning as described in any one of claims 1 to 9 is implemented.
Citation Information
Cited By
High-precision holder control method and system
CN120669764A
Multispectral image semantic segmentation method based on deep learning
CN121883833A