Steel rail surface defect segmentation method based on SAM2-CLIP model

Through the rail surface defect segmentation method based on the SAM2-CLIP model, the feature fusion is performed using the cross attention mechanism, which solves the problem of inaccurate defect segmentation in the prior art, and realizes high-precision pixel-level segmentation of rail surface defects.

CN120107604AActive Publication Date: 2025-06-06SOUTHWEST PETROLEUM UNIV

Patent Information

Application Number
CN202510581602.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-06-06
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

The existing rail surface defect detection methods have problems such as insufficient image feature richness, resulting in inaccurate defect segmentation and loss of edge details.

Method used

The rail surface defect segmentation method based on the SAM2-CLIP model is adopted. By acquiring the track image data set, the CLIP and SAM2 image encoder extracts features, and the feature fusion is performed through the cross attention mechanism, and the final input mask decoder generates the segmentation result.

Benefits of technology

The pixel-level segmentation of rail surface defects is realized, the accuracy and efficiency of detection are improved, and the problems of inefficient and high labor costs are solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107604A_ABST
    Figure CN120107604A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of steel rail defect detection, and particularly discloses a steel rail surface defect segmentation method based on an SAM2-CLIP model, and the method comprises the steps: inputting a track original image and a text prompt into an image encoder and a text encoder of a CLIP respectively, and obtaining CLIP image features and CLIP text features; performing weighted fusion on the CLIP image features and the CLIP text features to obtain CLIP weighted features; inputting the original track image into an image encoder of the SAM2, and outputting to obtain an SAM2 image feature; performing cross attention on the CLIP weighted features and the SAM2 image features; and inputting the cross attention fusion features into a SAM2 mask decoder for segmentation, and completing a steel rail surface defect segmentation result. According to the method, the problems that defect segmentation is inaccurate and edge details are lost due to the fact that image features extracted by an existing steel rail surface defect detection method are not rich enough are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of rail defect detection, and in particular relates to a rail surface defect segmentation method based on a SAM2-CLIP model. Background Art

[0002] Rails are an important part of the railway system, and their quality directly affects the safety and stability of train operation. With the gradual expansion of the scale of high-speed rail construction in my country and the acceleration of construction speed, the speed of train operation continues to increase, and safety assurance has become increasingly important. With the increase in operating time, rails are subject to the repeated effects of train loads, erosion by the natural environment, and aging of the materials themselves. Various defects are prone to appear on the tracks, such as cracks, wear, pitting, etc., which pose a serious threat to the safe operation of trains. If these defects are not discovered and handled in time, they will gradually develop and deteriorate. In severe cases, they may lead to major safety accidents such as train derailment and subversion, causing huge casualties and property losses. According to relevant statistics, a considerable number of railway accidents that have occurred in the past were caused by track defects, which fully highlights the importance and urgency of railway track safety inspection.

[0003] Existing methods for detecting rail surface defects include: manual inspection, which refers to rail defect detection by experienced railway inspectors, which has the problems of high labor cost, great influence of personal subjectivity, low efficiency and high false detection rate; eddy current detection, when detecting rails, by observing the changes in eddy current sensor signals to determine whether the rails have defects, but eddy current detection is easily affected by electromagnetic fields, cannot determine the type of defects, and has a slow detection speed; traditional image processing methods, such as Canny edge detection, achieve accurate edge extraction through multi-stage processing, suppress noise through Gaussian filtering, and enhance image contrast by combining histogram equalization, and use Sobel operator to calculate pixel gradient amplitude and direction field to construct the original edge response map. This method has essential defects such as noise sensitivity, parameter fixation, and poor dynamic adaptability. It is difficult to meet the needs of modern railway intelligent detection and has poor effect in complex environments. At present, deep learning research methods have also been proposed, such as constructing defect image segmentation based on the U-Net model, using historical rail images to input the U-net model to achieve real-time defect area positioning, and improving the rail surface defect detection method of the YOLO model, which uses the full-dimensional dynamic convolution ODConv to replace the traditional convolution of Yolov8, and embeds a double-layer context enhancement module CAM to improve the model effect. However, the existing deep learning research methods need to annotate a large amount of data for training when performing track defect detection, which is costly. When there are insufficient samples, serious overfitting and gradient instability are prone to occur during the training process. In addition, in the actual track defect segmentation, the extracted image features are not rich enough, resulting in inaccurate defect segmentation and loss of edge details. Summary of the invention

[0004] The purpose of the present invention is to solve the problem that the image features extracted by the existing rail surface defect detection method are not rich enough, resulting in inaccurate defect segmentation and loss of edge details. A rail surface defect segmentation method based on the SAM2-CLIP model is proposed.

[0005] The technical solution of the present invention is: a rail surface defect segmentation method based on the SAM2-CLIP model, comprising the following steps: Get orbital image dataset; The original track image in the track image dataset is input into the CLIP image encoder, and the CLIP image features are output; The CLIP image features and the CLIP text features output by the CLIP text encoder are weightedly fused to obtain the CLIP weighted features; The original track image is input into the image encoder of SAM2, and the SAM2 image features are output; The CLIP weighted features and SAM2 image features are processed through the cross-attention mechanism to obtain the cross-attention fusion features; The cross-attention fusion features are input into the mask decoder of SAM2, and the segmentation result of the original track image is output; All the original rail images in the rail image data set are segmented to obtain the segmentation results of all the original rail images and complete the rail surface defect segmentation results.

[0006] Preferably, the track original image is input into the CLIP image encoder, and the CLIP image features are output, specifically: Input the original track image into the image encoder of CLIP, split and flatten the original track image into the first 2Dpatches sequence; Map the dimension of the first 2D patches sequence to D dimensions to obtain the first Patch embedding vector; Add a position code to the first Patch embedding vector, and concatenate the CLS Token to the first Patch embedding vector with the position code added, to obtain a first input embedding vector; The first input embedding vector is input to the Transformer encoder, and the output vector is obtained , and then the vector Perform dimension mapping to obtain CLIP image features.

[0007] Preferably, the method for acquiring the CLIP text feature is specifically as follows: During the SAM2-CLIP model training phase, set the text prompt, input the text prompt into the CLIP text encoder, and then use the CLIP tokenizer to split the text prompt into a token sequence. :

[0008] in, Indicates the first The vocabulary index of the token, Indicates the number of tokens generated after splitting; Sequence the token Mapped into the vector space, we get the sequence :

[0009] in, Represents the first The vocabulary index of tokens is , represents the word embedding matrix, represents the set of real numbers, represents the vocabulary size, represents the embedding dimension; For sequence Add position code to each token in to get sequence :

[0010] in, Indicates that the position code is added. The vocabulary index of tokens is , Indicates The position code of each token; will sequence Input to the Transformer layer in the text encoder of CLIP, and output the sequence :

[0011] in, Indicates The final feature representation of a token after the complete Transformer encoding process; According to the sequence Extract CLIP text features : ; CLIP text features Mapped to the same dimension as the CLIP image features.

[0012] Preferably, the expression formula of the CLIP weighted feature is:

[0013] in, represents the CLIP weighted feature, represents the CLIP image features, represents the weight of CLIP text features, Represents a CLIP text feature.

[0014] Preferably, the original track image is input to the image encoder of SAM2, and the SAM2 image features are output, specifically: Input the original track image into the image encoder of SAM2, split and flatten the original track image into a second 2Dpatches sequence; Map the dimension of the second 2D patches sequence to D dimensions to obtain the second Patch embedding vector; Add position code to the second Patch embedding vector, and concatenate CLS Token to the second Patch embedding vector with position code added to obtain a second input embedding vector; Input the second input embedding vector into the Transformer encoder and output a multi-scale feature map; The multi-scale feature map is input into the FPN structure for multi-scale feature fusion to obtain two-dimensional high-resolution features; The two-dimensional high-resolution features are mapped to the same dimension as the CLIP image features, flattened into a sequence, and then the dimension is converted to obtain the SAM2 image features.

[0015] Preferably, the CLIP weighted features and SAM2 image features are processed by a cross attention mechanism to obtain a cross attention fusion feature, specifically: The SAM2 image features are used as queries, and the CLIP weighted features are used as keys and values, and then cross-attention calculation is performed to adjust the features of each pixel according to the CLIP weighted features, enhance the regional features, and obtain the features output by cross-attention; The features output by the cross-attention are reversely dimensioned and transformed back to a 2D feature map to match the dimension of the incoming mask decoder structure to obtain the cross-attention fusion feature .

[0016] Preferably, the cross-attention fusion feature is input into the mask decoder of SAM2, and the rail surface defect segmentation result is output, specifically: Get the point prompt information from the mask label and input it into the mask encoder for encoding, and output the prompt information ; The mask label is obtained by annotating the original track image during the SAM2-CLIP model training phase; Cross-attention fusion features and prompt information and 2D high-resolution features Input to the mask decoder to obtain the segmentation result of the original track image.

[0017] Preferably, the cross-attention fusion feature and prompt information and 2D high-resolution features Input to the mask decoder to obtain the segmentation result of the original track image, specifically: Constructing the initial output token sequence , and the initial output token sequence With prompt information Splice to get the token sequence , and its calculation formula is:

[0018] in, Indicates splicing; Based on the Transformer multi-head attention mechanism, the token sequence As a query, a 2D high-resolution feature The result of adding the position code is used as the key and value to perform attention calculation to obtain the updated image features. src , and output the IoU predicted token and the mask tokens of the candidate mask; Image features will be updated src Perform transposition and reshape operations to restore it to a two-dimensional feature map; Upsample the 2D feature map and fuse features with cross attention Fusion to obtain upsampled features ; Mask tokens based on candidate masks , using a multilayer perceptron to generate the weight vector , and upsample the features Perform reshape operation to obtain ; Based on the weight vector and , generate low-resolution masks using matrix multiplication, construct a set of all low-resolution masks, namely candidate masks, and normalize the candidate masks; Input the IoU prediction token into the IoU prediction multi-layer perceptron to generate the quality score of each candidate mask and obtain the optimal mask; Perform multi-layer convolution downsampling on the optimal mask to obtain a mask for two-dimensional high-resolution features. After 1*1 convolution, matrix addition is performed with the memory feature and an object pointer representing the movement trend of the object to obtain the memory representation of the current image. , completing the segmentation of the original track image.

[0019] The beneficial effects of the present invention are: 1. The present invention uses an improved SAM2-CLIP deep learning model to segment rail surface defects, realizing semantic segmentation of defects. Unlike traditional methods that simply locate defect positions, the present invention performs pixel-level segmentation of defects, which facilitates subsequent defect repair work and solves the problems of low efficiency and high labor cost of traditional detection technology.

[0020] 2. The present invention uses the features extracted by the text encoder and image encoder of CLIP and the features extracted by the SAM2 image encoder through a cross-attention mechanism to perform feature fusion, strengthen the image information, and provide more refined image features for subsequent mask generation, thereby obtaining a more accurate track original image segmentation result. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 Shown is a flow chart of a rail surface defect segmentation method based on the SAM2-CLIP model.

[0022] Figure 2 Shown is a flowchart of a rail surface defect segmentation method based on the SAM2-CLIP model.

[0023] Figure 3 Shown is a schematic diagram of the image encoder structure of CLIP.

[0024] Figure 4 Shown is a schematic diagram of the Transformer encoder structure.

[0025] Figure 5 Shown is a schematic diagram of the structure of the memory encoder. DETAILED DESCRIPTION

[0026] Now, exemplary embodiments of the present invention will be described in detail with reference to the accompanying drawings. It should be understood that the embodiments shown and described in the accompanying drawings are only exemplary and are intended to explain the principles and spirit of the present invention, rather than to limit the scope of the present invention.

[0027] Terminology explanation: SAM2 model: SAM2 stands for Segment Anything Model 2, a large image segmentation vision model. Based on the original SAM (Segment Anything Model), it adds a memory mechanism (Memory module), improves the segmentation accuracy, generalization ability and adaptability, has a strong zero-shot capability, and can be more efficiently applied to various complex visual tasks.

[0028] The CLIP model (Contrastive Language-Image Pre-Training) is a multimodal pre-training neural network that uses a large amount of paired image and text data for pre-training to learn the alignment relationship between images and text.

[0029] Cross-attention mechanism: A special form of multi-head attention used for information interaction between different inputs. It can effectively align and focus on context from different sources, helping the model better capture the correlation between two inputs.

[0030] SAM2 Image Encoder module: Image feature extraction module, i.e., SAM2’s image encoder, is used to extract image features. Specifically, it uses the Hiera image encoder pre-trained with MAE.

[0031] CLIP Image Encoder module: The image feature extraction module, i.e., CLIP’s image encoder, is used to extract image features, specifically using the VIT pre-trained model.

[0032] Memory Attention module: memory attention module.

[0033] Prompt Encoder module: Prompt encoder, which prompts the segmentation area through points, boxes, and masks.

[0034] Mask Decoder module: mask decoder, output image mask.

[0035] Memory Encoder module: memory encoder.

[0036] Memory Bank module: memory storage module.

[0037] Embodiment 1: like Figure 1 Shown and Figure 2 As shown, a rail surface defect segmentation method based on the SAM2-CLIP model includes the following steps: S1. Obtain orbital image dataset; During the SAM2-CLIP model training phase, data annotation is performed on the original rail images to generate corresponding text prompts; the text prompts describe the types of defects present in the original rail images; Specifically, track image data is obtained to construct a data set, and the data set is divided into a training set and a test set. Railway track images are collected by an industrial camera, and the processed data set is divided into a training set and a test set at a ratio of 8:2 through data cleaning and data labeling.

[0038] Data annotation uses the labelme data annotation tool to perform the following process: create category labels, define defect categories, annotate images at the pixel level, return annotation json files, and convert json files into binary mask images, where pixel values ​​of 0 represent background and pixel values ​​of 1 represent defects. The dataset consists of original rail images, mask labels, and text labels. The text labels describe the defect type corresponding to the original images. The collected rail images have three types of defects: dents, bruises, and frictions.

[0039] Perform non-local mean denoising on the track image data. Specifically, use the fastNlMeansDenoisingColored function of opencv, set the h value and hcolor value to 3, the templateWindowSize to 7, and the searchWindowSize to 21. Use the resize function to convert the image size to 224*224 to ensure that the dimensions passed to CLIP and SAM2 are the same.

[0040] S2. Input the original track image in the track image dataset into the image encoder of CLIP, which is specifically the VIT-B / 16 pre-trained model, and output the CLIP image features; In this embodiment, the original track image is input into the CLIP image encoder, and the CLIP image features are output, specifically: The first image of the original track image is input into the CLIP image encoder, whose structure is as follows Figure 3 As shown, the size of the original track image is 3*224*224. For the input original track image, the image is represented as a tensor. ,in, represents the set of real numbers, Indicates the height of the original image of the track, Indicates the width of the original image of the track, = =224, Indicates the number of channels, =3, B represents the batch size, that is, the number of images input into the model at one time.

[0041] The patch of the VIT-B / 16 pre-trained model is (16, 16). The number of image patches generated is: ,in, Indicates the Patch size; Split and flatten the original track image into a sequence of 2D patches Specifically, each track original image is segmented and flattened into image blocks with the number of channels C = 3, the patch size P = 16, and the number of image patches N = 196. Each image block has P 2 C =16*16*3=768 pixels; Feed the 2D patches sequence into the Linear Projection layer , the dimension of the 2Dpatches sequence P 2 C Mapped to D dimension, while keeping the number of image patches N unchanged, we get the Patch embedding vector , after linear projection, the obtained Patch embedding vector The shape is (b, 196, 768); Add positional encoding to each patch embedding vector so that the model can capture the positional information of each patch. The positional encoding is extracted using standard learnable / trainable 1-D positional encoding embedding, treating the 2-D image block as a 1-D sequence. Concatenate a learnable CLS Token to each patch embedding vector with positional encoding to aggregate global image information for subsequent classification and feature alignment tasks. The size of the CLS Token is (1, 768). After concatenation, the overall sequence length is N+1, and the shape of the generated input embedding vector is (b, 197, 768). Input the input embedding vector into Figure 4 The Transformer encoder shown in the figure continuously passes through the Transformer encoder composed of a series of Transformer encoder blocks, and the output is an output of size (b, 197, 768) ; The output vector is projected from 768 dimensions to 512 dimensions through a linear projection layer to obtain a 512-dimensional vector , namely CLIP image features, It is a trainable weight matrix with a shape of (768, 512), and finally obtains the (b, 512) image feature representation.

[0042] In this embodiment, the method for acquiring the CLIP text feature is specifically as follows: During the SAM2-CLIP model training phase, the artificially given text prompts are input into the CLIP text encoder, and the CLIP tokenizer is used to split the text prompts into token sequences. :

[0043] in, Indicates the first The vocabulary index of a token, with the sequence head represented as CLS and the sequence tail represented as Seq. Indicates the patch number of the original image of the track corresponding to the text prompt; Sequence the token Mapped into the vector space, we get the sequence :

[0044] in, Represents the first The vocabulary index of tokens is , represents the word embedding matrix, represents the set of real numbers, represents the vocabulary size, Indicates the embedding dimension. In this embodiment, the default dimension is 512. For sequence Each token in the sequence is encoded to indicate the position information. :

[0045] in, Indicates that the position code is added. The vocabulary index of tokens is , Indicates The position code of each token; will sequence Input to the L Transformer layers in the text encoder of CLIP, each Transformer layer includes a multi-head self-attention mechanism, a feedforward network, a residual connection and layer normalization, and the output is a sequence , and then extract the CLIP text features , and the CLIP text features are mapped to the same dimension as the CLIP image features, i.e. 512 dimensions, through a linear projection layer.

[0046] S3. Perform weighted fusion on the CLIP image features and the CLIP text features output by the CLIP text encoder to obtain the CLIP weighted features; In this embodiment, the expression formula of the CLIP weighted feature is:

[0047] in, represents the CLIP weighted feature, represents the CLIP image features, Represents the weight of the CLIP text feature. In this embodiment, , indicating that CLIP text features are used as auxiliary feature information, Represents a CLIP text feature.

[0048] S4. Input the original track image to the image encoder of SAM2, where the image encoder of SAM2 is specifically a MAE-Hiera pre-trained model, and output the SAM2 image features; In this embodiment, the track original image is input to the image encoder of SAM2, and the SAM2 image features are output. Specifically, the first image in the track original image is input to the image encoder of SAM2, that is, the MAE-Hiera pre-trained model, and the 224*224 track original image is input. The feature encoding operation is the same as the encoding process of the VIT-B / 16 pre-trained model of CLIP above, but the output of the MAE-Hiera pre-trained model is not a single sequence, but features are extracted at multiple scales to obtain feature maps of four scales. ,in , , , ,in, Indicates the batch size, that is, the number of images input to the model at one time; shallow features High-resolution information is retained for fine segmentation, which includes the following steps: Input the original track image into the image encoder of SAM2, split and flatten the original track image into a 2Dpatches sequence; Map the dimension of the 2D patches sequence to D dimensions to obtain the patch embedding vector; Add position encoding to the Patch embedding vector, and concatenate the classification tag (CLS Token) to the Patch embedding vector with position encoding to obtain the input embedding vector; Input the input embedding vector to the Transformer encoder and output a multi-scale feature map; Multi-scale feature map Input to the FPN structure for multi-scale feature fusion to obtain two-dimensional high-resolution features; specifically, first use 1×1 convolution at each scale to align the features channels, and the formula is: . Top Features Directly used as the highest feature of FPN ,Right now , propagating information from top to bottom, each layer of features The feature map of the previous layer is upsampled and laterally connected to the current layer. To merge, , , , and finally obtain the two-dimensional high-resolution feature ; The two-dimensional high-resolution features are transformed into The dimension is mapped to 512 dimensions and flattened into a sequence, and then the dimension is converted from Transformed into SAM2 image features , , to match the input format of the Transformer.

[0049] S5. The CLIP weighted features and SAM2 image features are processed through the cross attention mechanism to obtain the cross attention fusion features; In this embodiment, the CLIP weighted features and SAM2 image features are processed by the cross attention mechanism to obtain the cross attention fusion features, specifically: SAM2 image features As a query, CLIP weighted features As the key and value, cross attention calculation is performed to adjust the features of each pixel according to the CLIP weighted features, enhance regional features, avoid wrong segmentation, improve segmentation granularity, and make segmentation more accurate; The formula for calculating the cross attention is:

[0050] in, represents the cross attention weight, Indicates a query, Indicates the key, Indicates the value, is the activation function, represents the transpose of a matrix, Indicates the key vector dimension.

[0051] The features output by cross attention are still , the features output by the cross attention are reversely dimensioned and transformed back to the 2D feature map to match the dimension of the incoming mask decoder structure to obtain the cross attention fusion feature , ,in , .

[0052] S6. Input the cross-attention fusion features into the mask decoder of SAM2, and output the segmentation result of the original track image; In this embodiment, the cross-attention fusion feature is input into the mask decoder of SAM2, and the rail surface defect segmentation result is output, specifically: Get the point prompt information from the mask label and input it into the mask encoder for encoding processing, and output the 512-dimensional prompt information ; Cross-attention fusion features , prompt information and 2D high-resolution features Input to the mask decoder for processing, which includes the following steps: Constructing the initial output token sequence , Including tokens for IoU prediction and mask tokens for generating candidate masks, the initial output token sequence With prompt information Splice to get the token sequence , and its calculation formula is:

[0053] in, Indicates splicing; Based on the Transformer multi-head attention mechanism, the token sequence As a query, a 2D high-resolution feature The result of adding the position code is used as the key and value to perform attention calculation to obtain the updated image features. src , and output the IoU predicted token and the mask tokens of the candidate mask: and ; Image features will be updated src Perform transposition and reshape operations to restore it to a two-dimensional feature map; Upsample the two-dimensional feature map and compare it with the feature Fusion to obtain upsampled features , specifically:

[0054] in, represents the features obtained by upsampling in the first stage, represents the features obtained by upsampling in the second stage, , and represents the activation function. In this embodiment, the GELU activation function is used. , represents the error function, Represents LayerNorm2d that normalizes the channels. and Respectively represent the upsampling transposed convolution operation ConvTranspose2d in the first and second stages, and Respectively represent the high-resolution features The corresponding branch features extracted from .

[0055] based on , using multi-layer perceptron MLP to generate weight vector , Indicates indivual , and upsample the features Perform reshape operation to obtain , Represents upsampled features The number of feature channels, ; Based on the weight vector and , using matrix multiplication to generate a low-resolution mask , construct all low-resolution mask sets, namely candidate masks, for all candidate masks Normalized, the formula is: , represents the normalized candidate mask, Represents the Sigmoid activation function; Input the IoU prediction token into the IoU prediction multi-layer perceptron to generate the quality score of each candidate mask and obtain the optimal mask; Combine the optimal mask and 2D high-resolution features Input to the memory storage module, the optimal mask is convolved to obtain memory features and two-dimensional high-resolution features After 1*1 convolution, matrix addition is performed with the memory feature and an object pointer representing the movement trend of the object to obtain the t Memory representation of an image ,in t Indicates the current image number, which is stored in the memory bank for subsequent images.

[0056] S7. Segment all the original rail images in the rail image dataset to obtain the segmentation results of all the original rail images and complete the rail surface defect segmentation results. Specifically, for the next input image, the previous operation process is the same. After the original image enters the SAM2 Image Encoder, the image features obtained enter the memory attention module, and self-attention is performed on the original rail image features to extract feature information, and then cross-attention is performed. The original image features are used as the query, and the Key and Value are the first six times extracted from the Memory bank. (If less than six times, the first picture is taken The subsequent operations are the same as before. After obtaining the current image mask Then, the memory encoder is used to encode the image and store it in the MemoryBank. Repeat the above operation until all the image segmentation is completed to obtain the rail surface defect segmentation result. The structure of the memory encoder is as follows: Figure 5 shown.

[0057] Embodiment 2: On the basis of Example 1, the embodiment of the present invention conducts comparative experiments and ablation experiments to illustrate the technical effects of the present invention. The four parameters of mIoU, Dice coefficient, parameter amount and inference time are compared. The experimental results are shown in Tables 1 and 2. It can be obtained that the SAM2-CLIP proposed in the present invention has better performance than the traditional DeepLabv3+, PSPNet and UNet.

[0058] Table 1 Comparative experimental results

[0059] Table 2 Ablation experiment results

[0060] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific variations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed by the present invention, and these variations and combinations are still within the protection scope of the present invention.

Claims

1. A rail surface defect segmentation method based on the SAM2-CLIP model, characterized in that: The following steps are involved: Get orbital image dataset; The original track image in the track image dataset is input into the CLIP image encoder, and the CLIP image features are output; The CLIP image features and the CLIP text features output by the CLIP text encoder are weightedly fused to obtain the CLIP weighted features; The original track image is input into the image encoder of SAM2, and the SAM2 image features are output; The CLIP weighted features and SAM2 image features are processed through the cross-attention mechanism to obtain the cross-attention fusion features; The cross-attention fusion features are input into the mask decoder of SAM2, and the segmentation result of the original track image is output; All the original rail images in the rail image data set are segmented to obtain the segmentation results of all the original rail images and complete the rail surface defect segmentation results.

2. The rail surface defect segmentation method based on the SAM2-CLIP model according to claim 1, characterized in that: The original track image is input into the CLIP image encoder, and the CLIP image features are output, specifically: Input the original track image into the image encoder of CLIP, split and flatten the original track image into the first 2Dpatches sequence; Map the dimension of the first 2D patches sequence to D dimensions to obtain the first Patch embedding vector; Add a position code to the first Patch embedding vector, and concatenate the CLS Token to the first Patch embedding vector with the position code added, to obtain a first input embedding vector; The first input embedding vector is input to the Transformer encoder, and the output vector is obtained , and then the vector Perform dimension mapping to obtain CLIP image features.

3. The rail surface defect segmentation method based on the SAM2-CLIP model according to claim 1, characterized in that: The method for obtaining the CLIP text feature is specifically as follows: During the SAM2-CLIP model training phase, set the text prompt, input the text prompt into the CLIP text encoder, and then use the CLIP tokenizer to split the text prompt into a token sequence. : in, Indicates the first The vocabulary index of the token, Indicates the number of tokens generated after splitting; Sequence the token Mapped into the vector space, we get the sequence : in, Represents the first The vocabulary index of tokens is , represents the word embedding matrix, represents the set of real numbers, represents the vocabulary size, represents the embedding dimension; For sequence Add position code to each token in to get sequence : in, Indicates that the position code is added. The vocabulary index of tokens is , Indicates The position code of each token; will sequence Input to the Transformer layer in the text encoder of CLIP, and output the sequence : in, Indicates The final feature representation of a token after the complete Transformer encoding process; According to the sequence Extract CLIP text features : ; CLIP text features Mapped to the same dimension as the CLIP image features.

4. The rail surface defect segmentation method based on the SAM2-CLIP model according to claim 1, characterized in that: The expression formula of the CLIP weighted feature is: in, represents the CLIP weighted feature, represents the CLIP image features, represents the weight of CLIP text features, Represents a CLIP text feature.

5. The rail surface defect segmentation method based on the SAM2-CLIP model according to claim 1, characterized in that: The original track image is input to the image encoder of SAM2, and the SAM2 image features are output, specifically: Input the original track image into the image encoder of SAM2, split and flatten the original track image into a second 2Dpatches sequence; Map the dimension of the second 2D patches sequence to D dimensions to obtain the second Patch embedding vector; Add position code to the second Patch embedding vector, and concatenate CLS Token to the second Patch embedding vector with position code added to obtain a second input embedding vector; Input the second input embedding vector to the Transformer encoder and output a multi-scale feature map; The multi-scale feature map is input into the FPN structure for multi-scale feature fusion to obtain two-dimensional high-resolution features; The two-dimensional high-resolution features are mapped to the same dimension as the CLIP image features, flattened into a sequence, and then the dimension is converted to obtain the SAM2 image features.

6. The rail surface defect segmentation method based on the SAM2-CLIP model according to claim 5, characterized in that: The CLIP weighted features and SAM2 image features are processed through the cross attention mechanism to obtain the cross attention fusion features, specifically: The SAM2 image features are used as queries, and the CLIP weighted features are used as keys and values, and then cross-attention calculation is performed to adjust the features of each pixel according to the CLIP weighted features, enhance the regional features, and obtain the features output by cross-attention; The features output by the cross-attention are reversely dimensioned and transformed back to a 2D feature map to match the dimension of the incoming mask decoder structure to obtain the cross-attention fusion feature .

7. The rail surface defect segmentation method based on the SAM2-CLIP model according to claim 6, characterized in that: The cross-attention fusion feature is input into the mask decoder of SAM2, and the rail surface defect segmentation result is output, which is specifically: Get the point prompt information from the mask label and input it into the mask encoder for encoding, and output the prompt information ; The mask label is obtained by annotating the original track image during the SAM2-CLIP model training phase; Cross-attention fusion features and prompt information and 2D high-resolution features Input to the mask decoder to obtain the segmentation result of the original track image.

8. The rail surface defect segmentation method based on the SAM2-CLIP model according to claim 7, characterized in that: The cross-attention fusion feature and prompt information and 2D high-resolution features Input to the mask decoder to obtain the segmentation result of the original track image, specifically: Constructing the initial output token sequence , and the initial output token sequence With prompt information Splice to get the token sequence , and its calculation formula is: in, Indicates splicing; Based on the Transformer multi-head attention mechanism, the token sequence As a query, a 2D high-resolution feature The result of adding the position code is used as the key and value to perform attention calculation to obtain the updated image features. src , and output the IoU predicted token and the mask tokens of the candidate mask; Image features will be updated src Perform transposition and reshape operations to restore it to a two-dimensional feature map; Upsample the 2D feature map and fuse features with cross attention Fusion to obtain upsampled features ; Mask tokens based on candidate masks , using a multilayer perceptron to generate the weight vector , and upsample the features Perform reshape operation to obtain ; Based on the weight vector and , generate low-resolution masks using matrix multiplication, construct a set of all low-resolution masks, namely candidate masks, and normalize the candidate masks; Input the IoU prediction token into the IoU prediction multi-layer perceptron to generate the quality score of each candidate mask and obtain the optimal mask; Perform multi-layer convolution downsampling on the optimal mask to obtain a mask for two-dimensional high-resolution features. After 1*1 convolution, matrix addition is performed with the memory feature and an object pointer representing the movement trend of the object to obtain the memory representation of the current image. , completing the segmentation of the original track image.

Citation Information

Patent Citations

  • Image indication segmentation method based on pre-training model migration and prompt learning

    CN117808819A

  • Open word list segmentation method based on multi-base large model

    CN118799876A

  • Image segmentation method, device, equipment and storage medium

    US20220207742A1

Cited By

  • Visual large model efficient fine tuning and semantic segmentation method for rail transit

    CN120510387A

  • A visual large model fine-tuning and semantic segmentation method for rail transit

    CN120510387B

  • Steel surface defect detection method and system based on CLIP model cross-domain learning

    CN120782754A

  • A steel surface defect detection method and system based on CLIP model cross-domain learning

    CN120782754B

  • Visual inspection method and system for intelligent blank carrying line

    CN121392214A