Rail Surface Defect Segmentation Method Based on the SAM2-CLIP Model

Through the fusion characteristics of the cross attention mechanism of the SAM2-CLIP model, pixel-level segmentation of rail surface defects is realized, solving the problem of insufficient richness of image features in the prior art, and improving detection accuracy and efficiency.

CN120107604BActive Publication Date: 2025-07-04SOUTHWEST PETROLEUM UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510581602.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-07-04
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

The existing rail surface defect detection methods extracted insufficient image features, resulting in inaccurate defect segmentation, loss of edge details, and high labor costs and low efficiency.

Method used

Using the rail surface defect segmentation method based on the SAM2-CLIP model, features are extracted through the image and text encoder of CLIP, and the cross attention mechanism is used to fuse with the features of the SAM2 image encoder, and input it to the mask decoder for segmentation to realize pixel-level defect segmentation.

Benefits of technology

The precise segmentation of rail surface defects is achieved, the problems of inefficiency and high labor costs in traditional methods are solved, and more refined image features are provided for subsequent defect repair work.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107604B_ABST
    Figure CN120107604B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of rail defect detection, and specifically discloses a method for segmenting rail surface defects based on the SAM2-CLIP model, including: inputting the original track image and text prompt into the image encoder and text encoder of CLIP respectively to obtain CLIP image features and CLIP text features; performing weighted fusion on the CLIP image features and CLIP text features to obtain CLIP weighted features; inputting the original track image into the image encoder of SAM2, and outputting SAM2 image features; performing cross-attention on the CLIP weighted features and SAM2 image features; inputting the cross-attention fusion features into the mask decoder of SAM2 for segmentation to complete the rail surface defect segmentation result. The present invention solves the problem that the richness of image features extracted by the existing rail surface defect detection methods is insufficient, resulting in inaccurate defect segmentation and loss of edge details.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of rail defect detection, and particularly relates to a method for segmenting rail surface defects based on the SAM2-CLIP model. Background Art

[0002] Rail is an important part of the railway system, and its quality directly affects the safety and stability of train operation. With the gradual expansion of the scale of China's high-speed rail construction and the acceleration of the construction speed, the train operation speed is constantly increasing, and safety guarantee has become increasingly important. With the increase of operation time, the rail is affected by factors such as the repeated action of train loads, the erosion of the natural environment, and the aging of the material itself for a long time, and various defects are likely to appear on the track. For example, rail cracks, wear, pitting, etc. will pose a serious threat to the safe operation of trains. If these defects cannot be discovered and processed in time, they will gradually deteriorate, and may even lead to major safety accidents such as train derailment and overturning in severe cases, causing huge casualties and property losses. According to relevant statistical data, a considerable part of the past railway accidents were caused by track defects, fully highlighting the importance and urgency of railway track safety detection.

[0003] Existing detection methods for rail surface defects include: manual inspection, which means that railway inspection workers with certain experience conduct defect detection on the rails. This method has problems such as high labor costs, being greatly affected by personal subjectivity, low efficiency, and a high false detection rate; eddy current detection, when detecting the rails, by observing the changes in the signals of eddy current sensors to determine whether there are defects in the rails. However, eddy current detection is easily affected by the electromagnetic field, unable to determine the type of defects, and has a slow detection speed; traditional image processing methods, such as Canny edge detection, achieve precise edge extraction through multi-stage processing, suppress noise through Gaussian filtering, and at the same time enhance the image contrast by combining histogram equalization, calculate the pixel gradient amplitude and direction field using the Sobel operator, and construct the original edge response map. This method has essential defects such as being sensitive to noise, having fixed parameters, and poor dynamic adaptability, and it is difficult to meet the requirements of modern railway intelligent detection and has poor effects in complex environments. Currently, deep learning research methods have also been proposed, such as constructing defect image segmentation based on the U-Net model, using historical rail images to input into the U-net model to achieve real-time defect area positioning, and improving the rail surface defect detection method of the YOLO model. It improves the model effect by using the full-dimensional dynamic convolution ODConv to replace the traditional convolution of Yolov8 and embedding the double-layer context enhancement module CAM, etc. However, the existing deep learning research methods require a large amount of data to be labeled for training when performing rail defect detection, consuming high costs. When the samples are insufficient, problems such as serious overfitting and unstable gradients are likely to occur during the training process. Moreover, in actual rail defect segmentation, there are deficiencies such as insufficient richness of the extracted image features, resulting in inaccurate defect segmentation and loss of edge details. Summary of the Invention

[0004] The purpose of the present invention is to address the problem that the existing rail surface defect detection methods have insufficient richness of the extracted image features, resulting in inaccurate defect segmentation and loss of edge details, and a rail surface defect segmentation method based on the SAM2-CLIP model is proposed.

[0005] The technical solution of the present invention is as follows: A rail surface defect segmentation method based on the SAM2-CLIP model, comprising the following steps:

[0006] Obtain a track image dataset;

[0007] Input the original track images in the track image dataset into the image encoder of CLIP, and output the CLIP image features;

[0008] Perform weighted fusion on the CLIP image features and the CLIP text features output by the text encoder of CLIP to obtain the CLIP weighted features;

[0009] Input the original track image into the image encoder of SAM2 to obtain SAM2 image features;

[0010] Process the CLIP weighted features and SAM2 image features through the cross-attention mechanism to obtain cross-attention fusion features;

[0011] Input the cross-attention fusion features into the mask decoder of SAM2 to obtain the segmentation result of the original track image;

[0012] Segment all the original track images in the track image dataset to obtain the segmentation results of all the original track images, and complete the segmentation results of the rail surface defects.

[0013] Preferably, the inputting the original track image into the image encoder of CLIP to obtain CLIP image features is specifically as follows:

[0014] Input the original track image into the image encoder of CLIP, and split and flatten the original track image into a first sequence of 2D patches;

[0015] Map the dimension of the first sequence of 2D patches to D dimensions to obtain the first patch embedding vector;

[0016] Add position encoding to the first patch embedding vector, and splice the CLS Token to the first patch embedding vector with position encoding added to obtain the first input embedding vector;

[0017] Input the first input embedding vector into the Transformer encoder, and output a vector , and then map the dimension of the vector to obtain CLIP image features.

[0018] Preferably, the method for obtaining the CLIP text features is specifically as follows:

[0019] In the training stage of the SAM2-CLIP model, set a text prompt, input the text prompt into the text encoder of CLIP, and then use the tokenizer of CLIP to split the text prompt into a token sequence :

[0020]

[0021] wherein, represents the vocabulary index of the th token in the text prompt, and represents the number of tokens generated after splitting;

[0022] Input the token sequence Map it to a vector space to obtain a sequence :

[0023]

[0024] where represents the vocabulary index of the th token in the text prompt in the vector space, and there is , represents the word embedding matrix, represents the set of real numbers, represents the vocabulary size, represents the embedding dimension;

[0025] Add positional encoding to each token in the sequence to obtain a sequence :

[0026]

[0027] where represents the vocabulary index of the th token after adding positional encoding, and there is , represents the positional encoding of the th token;

[0028] Input the sequence into the Transformer layer in the text encoder of CLIP, and the output is a sequence :

[0029]

[0030] where represents the final feature representation of the th token after going through the complete Transformer encoding process;

[0031] Extract the CLIP text feature according to the sequence :

[0032] ;

[0033] Map the CLIP text feature to the same dimension as the CLIP image feature.

[0034] Preferably, the expression formula of the CLIP weighted feature is:

[0035]

[0036] Among them, represents the CLIP weighted feature, represents the CLIP image feature, represents the weight of the CLIP text feature, represents the CLIP text feature.

[0037] Preferably, the original track image is input into the image encoder of SAM2, and the SAM2 image feature is output, specifically:

[0038] Input the original track image into the image encoder of SAM2, and split and flatten the original track image into a second 2D patches sequence;

[0039] Map the dimension of the second 2D patches sequence to D dimensions to obtain a second Patch embedding vector;

[0040] Add positional encoding to the second Patch embedding vector, and splice the CLS Token to the second Patch embedding vector with positional encoding added to obtain a second input embedding vector;

[0041] Input the second input embedding vector into the Transformer encoder to output a multi-scale feature map;

[0042] Input the multi-scale feature map into the FPN structure for multi-scale feature fusion to obtain a two-dimensional high-resolution feature;

[0043] Map the two-dimensional high-resolution feature to the same dimension as the CLIP image feature, flatten it into a sequence, and then transform the dimension to obtain the SAM2 image feature.

[0044] Preferably, the CLIP weighted feature and the SAM2 image feature are processed through a cross-attention mechanism to obtain a cross-attention fusion feature, specifically:

[0045] Use the SAM2 image feature as the query, the CLIP weighted feature as the key and value, and then perform cross-attention calculation to adjust the features of each pixel point according to the CLIP weighted feature, enhance the regional features, and obtain the feature output after cross-attention;

[0046] Perform inverse dimension adjustment on the feature output after cross-attention, transform it back to a 2D feature map to match the dimension of the incoming mask decoder structure, and obtain the cross-attention fusion feature .

[0047] Preferably, the cross-attention fusion feature is input into the mask decoder of SAM2, and the rail surface defect segmentation result is output, specifically:

[0048] Obtain point prompt information from the mask label and input it into the mask encoder for encoding processing, and output the prompt information ; The mask label is obtained by data annotation of the original track image during the training stage of the SAM2-CLIP model;

[0049] Input the cross-attention fusion feature, prompt information and two-dimensional high-resolution feature into the mask decoder to obtain the segmentation result of the original track image.

[0050] Preferably, the input of the cross-attention fusion feature, prompt information and two-dimensional high-resolution feature into the mask decoder to obtain the segmentation result of the original track image is specifically as follows:

[0051] Construct an initial output token sequence , and splice the initial output token sequence with the prompt information to obtain a token sequence , and its calculation formula is:

[0052]

[0053] where represents splicing;

[0054] Based on the Transformer multi-head attention mechanism, use the token sequence as the query, and the result of adding the two-dimensional high-resolution feature and the position encoding as the key and value, perform attention calculation to obtain the updated image feature src , and output the IoU prediction token and the mask tokens of the candidate mask;

[0055] Transpose and reshape the updated image feature src to restore it to a two-dimensional feature map;

[0056] Upsample the two-dimensional feature map and fuse it with the cross-attention fusion feature to obtain the upsampled feature ;

[0057] Based on the mask tokens of the candidate mask, use a multi-layer perceptron to generate a weight vector , and perform a reshape operation on the upsampled feature to obtain ;

[0058] Based on the weight vector and , use matrix multiplication to generate a low-resolution mask, construct a set of all low-resolution masks, i.e., candidate masks, and normalize the candidate masks;

[0059] Input the IoU prediction token into the IoU prediction multi-layer perceptron to generate the quality scores of each candidate mask, and obtain the optimal mask;

[0060] Perform a multi-layer convolutional downsampling operation on the optimal mask to obtain a mask. After performing 1*1 convolution on the two-dimensional high-resolution feature , perform matrix addition with the memory feature and an object pointer representing the object motion trend to obtain the memory representation of the current image , and complete the segmentation of the original track image.

[0061] The beneficial effects of the present invention are as follows:

[0062] 1. The present invention uses an improved SAM2-CLIP deep learning model to segment the defects on the rail surface, achieving semantic segmentation of the defects. Different from traditional methods that only locate the defect positions, it performs pixel-level segmentation of the defects, facilitating subsequent defect repair work, and solving the problems of low efficiency and high labor cost of traditional detection techniques.

[0063] 2. The present invention fuses the features extracted by the text encoder and image encoder of CLIP with the features extracted by the SAM2 image encoder through a cross-attention mechanism to strengthen the image information, providing more refined image features for subsequent mask generation, thereby obtaining a more accurate segmentation result of the original track image. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 Shown is a flowchart of a method for segmenting rail surface defects based on the SAM2-CLIP model.

[0065] Figure 2 Shown is a flow block diagram of a method for segmenting rail surface defects based on the SAM2-CLIP model.

[0066] Figure 3 Shown is a schematic structural diagram of the image encoder of CLIP.

[0067] Figure 4 Shown is a schematic structural diagram of the Transformer encoder.

[0068] Figure 5 Shown is a schematic structural diagram of the memory encoder. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0069] Exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be understood that the embodiments shown and described in the drawings are merely exemplary, intended to illustrate the principles and spirit of the present invention, and not to limit the scope of the present invention.

[0070] Term Explanation:

[0071] SAM2 Model: SAM2 full name Segment Anything Model 2, a large vision model for image segmentation. Based on the original SAM (Segment Anything Model), it adds a memory mechanism (Memory module), improving the segmentation accuracy, generalization ability, and adaptability. It has a powerful Zero-shot ability and can be more efficiently applied to various complex vision tasks.

[0072] CLIP Model (Contrastive Language-Image Pre-Training), is a multi-modal pre-trained neural network that uses a large amount of paired image and text data for pre-training to learn the alignment relationship between images and texts.

[0073] Cross-attention mechanism: A special form of multi-head attention used for information interaction between different inputs. It can effectively align and focus on the context from different sources, helping the model better capture the correlation between two inputs.

[0074] SAM2 Image Encoder Module: An image feature extraction module, that is, the image encoder of SAM2, used to extract image features. Specifically, it uses the Hiera image encoder pre-trained by MAE.

[0075] CLIP Image Encoder Module: An image feature extraction module, that is, the image encoder of CLIP, used to extract image features. Specifically, it uses the VIT pre-trained model.

[0076] Memory Attention Module: Memory attention module.

[0077] Prompt Encoder Module: A prompt encoder that segments regions through point, box, and mask prompts.

[0078] Mask Decoder Module: A mask decoder that outputs an image mask.

[0079] Memory Encoder Module: Memory encoder.

[0080] Memory Bank Module: Memory storage module.

[0081] Example 1:

[0082] As Figure 1 shown and Figure 2 shown, a method for segmenting rail surface defects based on the SAM2-CLIP model includes the following steps:

[0083] S1. Obtain a track image dataset;

[0084] In the training stage of the SAM2-CLIP model, data annotation is performed on the original track image to generate corresponding text prompts; the text prompts describe what types of defects exist in the original track image;

[0085] Specifically, obtain track image data to construct a dataset, divide the dataset into a training set and a test set, collect railway track pictures through an industrial camera, and through data cleaning and data annotation, divide the processed dataset into a training set and a test set according to 8:2.

[0086] Data annotation uses the labelme data annotation tool to perform the following process: create category labels, define defect categories, perform pixel-level annotation on the image, return the annotated json file, and convert the json file into a binary mask image, where the pixel value of 0 represents the background and the pixel value of 1 represents the defect. The dataset consists of the original track image, the mask label, and the text label. The text label describes the defect type corresponding to the original picture. There are three types of defects in the collected track pictures: depression, rolling injury, and friction.

[0087] Perform non-local mean denoising on the track image data. Specifically, use the fastNlMeansDenoisingColored function in opencv, set the h value and hcolor value to 3, the templateWindowSize to 7, and the searchWindowSize value to 21. Use the resize function to convert the image size to 224*224 to ensure that the dimensions passed into CLIP and SAM2 are the same.

[0088] S2. Input the original track image in the track image dataset into the image encoder of CLIP. The image encoder of CLIP is specifically the VIT-B / 16 pre-trained model, and CLIP image features are output;

[0089] In this embodiment, the step of inputting the original track image into the image encoder of CLIP and outputting CLIP image features is specifically:

[0090] Input the first image of the original track image into the CLIP image encoder, and its structure is as Figure 3As shown, the original size of the orbital image is 3*224*224. For the input original orbital image, the image is represented as a tensor, , where represents the set of real numbers, represents the height of the original orbital image, represents the width of the original orbital image, = =224, represents the number of channels, =3, and B represents the batch size, that is, the number of pictures input into the model at one time.

[0091] Patch of the VIT-B / 16 pre-trained model: (16, 16), and the resulting number of image patches: , where represents the Patch size;

[0092] The original orbital image is sliced and flattened into a sequence of 2D patches ; specifically, each original orbital image is sliced and flattened into image blocks with the number of channels C = 3, Patch size P = 16, and the number of image patches N = 196. Each image block has P 2 C =16*16*3 = 768 pixels;

[0093] The sequence of 2D patches is fed into a linear projection layer (Linear Projection) , and the dimension of the sequence of 2D patches P 2 C is mapped to D dimensions while keeping the number of image patches N unchanged, obtaining Patch embedding vectors . After linear projection, the resulting Patch embedding vectors have the shape of (b, 196, 768);

[0094] Position encoding is added to each Patch embedding vector so that the model can capture the position information of each Patch. The position encoding is extracted using standard learnable / trainable 1-D position encoding embeddings, treating the 2-D image blocks as 1-D sequences; a learnable CLS Token is concatenated to the Patch embedding vectors with added position encoding for aggregating global image information, facilitating subsequent classification and feature alignment tasks. The size of the CLS Token is (1, 768). After concatenation, the overall sequence length is N + 1, and the resulting input embedding vectors have the shape of (b, 197, 768);

[0095] Input the input embedding vector into the Transformer encoder as shown in Figure 4 and continuously pass it forward through the Transformer encoder composed of a serial stack of Transformer Encoder Blocks, and the output obtained is of size (b, 197, 768). ;

[0096] The output vector projects 768 dimensions to 512 dimensions through a linear projection layer to obtain a 512-dimensional vector. That is, the CLIP image feature. is a trainable weight matrix of shape (768, 512), and finally an image feature representation of (b, 512) is obtained.

[0097] In this embodiment, the method for obtaining the CLIP text feature is specifically as follows:

[0098] During the training stage of the SAM2-CLIP model, input the artificially given text prompt into the text encoder of CLIP, and use the tokenizer of CLIP to split the text prompt into a token sequence. :

[0099]

[0100] Among them, represents the vocabulary index of the th token in the text prompt. At the same time, the sequence head is represented as CLS, and the sequence tail is represented as Seq. represents the number of patches of the original image of the track corresponding to the text prompt;

[0101] Map the token sequence to the vector space to obtain the sequence :

[0102]

[0103] Among them, represents the vocabulary index of the th token in the text prompt in the vector space. There is , represents the word embedding matrix. represents the set of real numbers. represents the vocabulary size. represents the embedding dimension, and the default 512 dimensions are used in this embodiment;

[0104] Add position encoding to each token in the sequence to represent the position information, and obtain the sequence :

[0105]

[0106] Among them, represents the vocabulary index of the th token with positional encoding added, and there is , represents the positional encoding of the th token;

[0107] Input the sequence into the L Transformer layers in the text encoder of CLIP. Each Transformer layer includes a multi-head self-attention mechanism, a feed-forward network, as well as residual connections and layer normalization, and output the sequence , and then extract the CLIP text features , and map the CLIP text features to the same dimension as the CLIP image features, that is, 512 dimensions, through a linear projection layer.

[0108] S3. Perform weighted fusion on the CLIP image features and the CLIP text features output by the text encoder of CLIP to obtain the CLIP weighted features;

[0109] In this embodiment, the expression formula of the CLIP weighted features is:

[0110]

[0111] Among them, represents the CLIP weighted features, represents the CLIP image features, represents the weight of the CLIP text features. In this embodiment, is taken, which represents the CLIP text features as auxiliary feature information, represents the CLIP text features.

[0112] S4. Input the original orbital image into the image encoder of SAM2. The image encoder of SAM2 is specifically the MAE-Hiera pre-trained model, and output the SAM2 image features;

[0113] In this embodiment, the original track image is input into the image encoder of SAM2, and the SAM2 image features are output. Specifically: the first image in the original track image is input into the image encoder of SAM2, that is, the MAE-Hiera pre-trained model. A 224*224 original track image is input, and the operation of its feature encoding is the same as the encoding process of the above-mentioned CLIP's VIT-B / 16 pre-trained model. However, the output of the MAE-Hiera pre-trained model is not a single sequence, but features are extracted at multiple scales, obtaining feature maps at four scales , where , , , , where, represents the batch size, that is, the number of pictures input into the model at one time; the shallow features retain high-resolution information for fine segmentation, and the specific steps are as follows:

[0114] Input the original track image into the image encoder of SAM2, and cut and flatten the original track image into a 2D patches sequence;

[0115] Map the dimension of the 2D patches sequence to D dimensions to obtain Patch embedding vectors;

[0116] Add position encoding to the Patch embedding vectors, and splice the classification token (CLS Token) to the Patch embedding vectors with position encoding added to obtain input embedding vectors;

[0117] Input the input embedding vectors into the Transformer encoder to output multi-scale feature maps;

[0118] Input the multi-scale feature maps into the FPN structure for multi-scale feature fusion to obtain two-dimensional high-resolution features; specifically, first use a 1×1 convolution at each scale to align the channels of the features, and its formula is: . The highest-level feature is directly used as the highest feature of the FPN , that is , and the information is propagated from top to bottom. Each layer of feature is fused with the horizontal connection feature of the current layer after upsampling the feature map of the previous layer, that is , , , , and finally two-dimensional high-resolution features ;

[0119] Convert the two-dimensional high-resolution features through a 1×1 convolution The dimension is mapped to 512 dimensions and flattened into a sequence, and then the dimension is transformed, from to become the SAM2 image features , , to match the input format of the Transformer.

[0120] S5. Process the CLIP weighted features and SAM2 image features through the cross-attention mechanism to obtain cross-attention fusion features;

[0121] In this embodiment, the process of processing the CLIP weighted features and SAM2 image features through the cross-attention mechanism to obtain cross-attention fusion features is specifically as follows:

[0122] Take the SAM2 image features as the query, and the CLIP weighted features as the key and value, and then perform cross-attention calculation, so that the features of each pixel point are adjusted according to the CLIP weighted features, enhancing the regional features, being able to avoid incorrect segmentation, improving the segmentation granularity, and making the segmentation more accurate;

[0123] The formula for the cross-attention calculation is:

[0124]

[0125] where, represents the cross-attention weight, represents the query, represents the key, represents the value, is the activation function, represents the transpose of the matrix, represents the key vector dimension.

[0126] The features output after cross-attention are still , perform reverse dimension adjustment on the features output after cross-attention, transform back to a 2D feature map to match the dimension of the incoming mask decoder structure, and obtain the cross-attention fusion features , , where , .

[0127] S6. Input the cross-attention fusion features into the mask decoder of SAM2, and output the segmentation result of the original image of the track;

[0128] In this embodiment, the process of inputting the cross-attention fusion features into the mask decoder of SAM2 and outputting the segmentation result of the rail surface defect is specifically as follows:

[0129] The point hint information is obtained from the mask label and input into the mask encoder for encoding processing, and the output is the 512-dimensional hint information ;

[0130] The cross-attention fusion feature , the hint information and the two-dimensional high-resolution feature are input into the mask decoder for processing, which specifically includes the following steps:

[0131] Construct the initial output token sequence , including the tokens for IoU prediction and the mask tokens for generating candidate masks. The initial output token sequence is concatenated with the hint information to obtain the token sequence , and its calculation formula is:

[0132]

[0133] where represents concatenation;

[0134] Based on the Transformer multi-head attention mechanism, the token sequence is used as the query, and the result of adding the two-dimensional high-resolution feature to the position encoding is used as the key and value for attention calculation to obtain the updated image feature src , and the IoU prediction token and the mask tokens of the candidate mask are output: and ;

[0135] The updated image feature src is transposed and reshaped to restore it to a two-dimensional feature map;

[0136] The two-dimensional feature map is upsampled and fused with the feature to obtain the upsampled feature , specifically:

[0137]

[0138] where represents the feature obtained by the first-stage upsampling, represents the feature obtained by the second-stage upsampling, and there is , and represent the activation function. In this embodiment, the GELU activation function is used, , represents the error function, LayerNorm2d that normalizes the channels, and represent the transposed convolutional upsampling operations ConvTranspose2d for the first stage and the second stage respectively, and represent the corresponding branch features extracted from the high-resolution features respectively.

[0139] Based on , use the multi-layer perceptron MLP to generate the weight vector , represents the th , and perform a reshape operation on the upsampled feature to obtain , represents the number of feature channels of the upsampled feature , ;

[0140] Based on the weight vector and , use matrix multiplication to generate the low-resolution mask , construct the set of all low-resolution masks, i.e., the candidate masks, and normalize all the candidate masks using the formula: , represents the normalized candidate mask, represents the Sigmoid activation function;

[0141] Input the IoU prediction token into the IoU prediction multi-layer perceptron to generate the quality scores of each candidate mask and obtain the optimal mask;

[0142] Input the optimal mask and the two-dimensional high-resolution feature into the memory storage module. The optimal mask undergoes a convolutional operation to obtain the memory feature. The two-dimensional high-resolution feature undergoes a 1*1 convolution and then performs a matrix addition with the memory feature and an object pointer representing the object motion trend to obtain the memory representation of the t th image, where t represents which image it is currently, and store it in the Memory Bank for subsequent images to use.

[0143] S7. Segment all the original track images in the track image dataset to obtain the segmentation results of all the original track images, and complete the segmentation results of the rail surface defects. Specifically, for the next input image, the previous running process is the same. After the original image enters the SAM2 Image Encoder, the obtained image features enter the Memory Attention module. Self-attention is performed on the original track image features to extract feature information, and then cross-attention is performed. The original image features serve as the Query, and the Key and Value are extracted from the first six times (if there are less than six times, take until the first image) from the Memory bank. The subsequent operations are the same as before. After obtaining the current image mask, it is stored in the MemoryBank through memory encoding by the Memory Encoder. Repeat the above operations until all image segmentations are completed to obtain the segmentation results of the rail surface defects. The structure of the Memory Encoder is as shown. (If there are less than six times, take until the first image ). The subsequent operations are the same as before. After obtaining the current image mask , it is stored in the MemoryBank through memory encoding by the Memory Encoder. Repeat the above operations until all image segmentations are completed to obtain the segmentation results of the rail surface defects. The structure of the Memory Encoder is as shown Figure 5 .

[0144] Example 2:

[0145] Based on Example 1, the embodiments of the present invention conduct comparative experiments and ablation experiments to illustrate the technical effects of the present invention. Four parameters, namely mIoU, Dice coefficient, number of parameters, and inference time, are compared. The experimental results are shown in Tables 1 and 2. It can be obtained that the SAM2-CLIP proposed by the present invention has better performance compared to the traditional DeepLabv3+, PSPNet, and UNet.

[0146] Table 1 Comparative experiment results

[0147]

[0148] Table 2 Ablation experiment results

[0149]

[0150] Those of ordinary skill in the art will realize that the embodiments described herein are for helping readers understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations without departing from the essence of the present invention based on the technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.

Claims

1. A method for segmenting rail surface defects based on the SAM2-CLIP model, characterized in that, It includes the following steps: Obtain an orbital image dataset; Input the original orbital images in the orbital image dataset into the image encoder of CLIP, and output CLIP image features; Perform weighted fusion on the CLIP image features and the CLIP text features output by the text encoder of CLIP to obtain CLIP weighted features; Input the original orbital images into the image encoder of SAM2, and output SAM2 image features, specifically: Input the original orbital images into the image encoder of SAM2, and split and flatten the original orbital images into a second 2D patches sequence; Map the dimension of the second 2D patches sequence to D dimensions to obtain a second Patch embedding vector; Add position encoding to the second Patch embedding vector, and splice the CLS Token to the second Patch embedding vector with position encoding added to obtain a second input embedding vector; Input the second input embedding vector into a Transformer encoder to output a multi-scale feature map; Input the multi-scale feature map into an FPN structure for multi-scale feature fusion to obtain a two-dimensional high-resolution feature; Map the two-dimensional high-resolution feature to the same dimension as the CLIP image features, flatten it into a sequence, and then transform the dimension to obtain SAM2 image features; Process the CLIP weighted features and the SAM2 image features through a cross-attention mechanism to obtain cross-attention fusion features; Input the cross-attention fusion features, the prompt information, and the two-dimensional high-resolution features into a mask decoder to obtain the segmentation result of the original orbital images; Segment all the original orbital images in the orbital image dataset to obtain the segmentation results of all the original orbital images, and complete the segmentation result of the rail surface defects.

2. The rail surface defect segmentation method based on the SAM2-CLIP model according to claim 1, wherein The specific process of inputting the original orbital images into the image encoder of CLIP and outputting CLIP image features is as follows: Input the original orbital images into the image encoder of CLIP, and split and flatten the original orbital images into a first 2D patches sequence; Map the dimension of the first 2D patches sequence to D dimensions to obtain a first Patch embedding vector; Add position encoding to the first Patch embedding vector, and splice the CLS Token to the first Patch embedding vector with position encoding added to obtain a first input embedding vector; Input the first input embedding vector into the Transformer encoder, and output a vector , and then perform dimensional mapping on the vector to obtain CLIP image features.

3. The method for segmenting rail surface defects based on the SAM2-CLIP model according to claim 1, wherein The specific method for obtaining the CLIP text features is as follows: During the training phase of the SAM2-CLIP model, set a text prompt, input the text prompt into the text encoder of CLIP, and then use the tokenizer of CLIP to split the text prompt into a sequence of tokens : Among them, represents the vocabulary index of the th token in the text prompt, represents the number of tokens generated after splitting; Map the token sequence to the vector space to obtain the sequence : wherein, represents the vocabulary index of the th token in the text prompt in the vector space, and there is , represents the word embedding matrix, represents the set of real numbers, represents the vocabulary size, represents the embedding dimension; Add positional encoding to each token in the sequence to obtain the sequence : Among them, represents the vocabulary index of the th token with positional encoding added, and there is , represents the positional encoding of the th token; Input the sequence into the Transformer layer in the text encoder of CLIP, and output the sequence : Among them, represents the final feature representation obtained by the n-th token after the complete Transformer encoding process; According to the sequence the CLIP text features are extracted : ; Map the CLIP text features to the same dimension as the CLIP image features.

4. The method for segmenting rail surface defects based on the SAM2-CLIP model according to claim 1, wherein, The expression formula of the CLIP weighted features is: Among them, represents the CLIP weighted feature, represents the CLIP image feature, represents the weight of the CLIP text feature, represents the CLIP text feature.

5. The rail surface defect segmentation method based on the SAM2-CLIP model according to claim 1, wherein The specific process of processing the CLIP weighted features and the SAM2 image features through a cross-attention mechanism to obtain cross-attention fusion features is as follows: Use the SAM2 image features as the query, the CLIP weighted features as the key and value, and then perform cross-attention calculation to adjust the features of each pixel point according to the CLIP weighted features, enhance the regional features, and obtain the features output after cross-attention; Reverse the dimensionality adjustment of the features output by the cross-attention to transform them back into a 2D feature map to match the dimensions of the incoming mask decoder structure, obtaining the cross-attention fusion features .

6. The rail surface defect segmentation method based on the SAM2-CLIP model according to claim 1, characterized in that, The specific process of inputting the cross-attention fusion features into the mask decoder of SAM2 and outputting the segmentation result of the rail surface defects is as follows: Obtain point hint information from the mask label and input it into the mask encoder for encoding processing to output the hint information ; The mask label is obtained by data annotation of the original track image during the training stage of the SAM2-CLIP model; Input the cross-attention fusion features and prompt information and two-dimensional high-resolution features into the mask decoder to obtain the segmentation result of the original orbital image.

7. The method for segmenting rail surface defects based on the SAM2-CLIP model according to claim 6, wherein, The cross-attention fusion feature and the prompt information and the two-dimensional high-resolution feature are input into the mask decoder to obtain the segmentation result of the original orbital image, specifically as follows: Construct the initial output token sequence , and splice the initial output token sequence with the prompt message to obtain the token sequence , and its calculation formula is: Among them, represents splicing; Based on the Transformer multi-head attention mechanism, the token sequence is used as the query, and the result of adding the two-dimensional high-resolution feature to the position encoding is used as the key and value for attention calculation to obtain the updated image feature src , and the IoU prediction token and the mask tokens of the candidate mask are output; The updated image features will be src Transposed and reshaped to restore a two-dimensional feature map; Upsample the two-dimensional feature map and fuse it with the cross-attention fusion feature to obtain the upsampled feature ; Mask tokens based on candidate masks , generate a weight vector using a multi-layer perceptron , and perform a reshape operation on the upsampled features to obtain ; Based on the weight vector and , use matrix multiplication to generate a low-resolution mask, construct a set of all low-resolution masks, i.e., candidate masks, and normalize the candidate masks; Input the IoU prediction token into the IoU prediction multi-layer perceptron to generate the quality scores of each candidate mask and obtain the optimal mask; Perform multi-layer convolutional downsampling processing on the optimal mask to obtain the mask, and perform 1*1 convolution on the two-dimensional high-resolution feature After convolution, perform matrix addition with the memory feature and an object pointer representing the object motion trend to obtain the memory representation of the current image , and complete the segmentation of the original orbital image.

Citation Information

Patent Citations

  • Image indication segmentation method based on pre-training model migration and prompt learning

    CN117808819A

  • Open word list segmentation method based on multi-base large model

    CN118799876A