A visual large model fine-tuning and semantic segmentation method for rail transit

By constructing a semantic segmentation dataset for the orbital operating environment and combining memory attention and efficient adaptive fine-tuning methods, the problem of insufficient image segmentation accuracy in the orbital operating environment is solved, the robustness and adaptability of the model are improved, and more efficient semantic segmentation results are achieved.

CN120510387BActive Publication Date: 2026-01-06BEIHANG UNIV +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510998144.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2026-01-06
Estimated Expiration
2045-07-21

AI Technical Summary

Technical Problem

Existing methods suffer from insufficient image segmentation accuracy, robustness, and generalization ability in orbital operating environments. They perform poorly, especially under complex lighting conditions, high-speed motion, and extreme environments, making it difficult to effectively identify and segment key regions.

Method used

By constructing a semantic segmentation dataset for the orbital operating environment, and utilizing automated annotation and image enhancement techniques, combined with a memory attention module and an efficient adaptive fine-tuning method, the visual large model is adjusted to adapt to the specific needs of the orbital operating environment.

Benefits of technology

It improves the accuracy of orbital operating environment identification and semantic segmentation, enhances the robustness and generalization ability of the model under rare events and abnormal conditions, and reduces the computational resource requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120510387B_ABST
    Figure CN120510387B_ABST
Patent Text Reader

Abstract

The application discloses a visual large model fine-tuning and semantic segmentation method for rail transit, relates to the fields of image feature processing, visual large model fine-tuning and semantic segmentation in the field of artificial intelligence deep learning, and the method constructs a rail operation environment semantic segmentation dataset, and trains a visual large model by using the rail operation environment semantic segmentation dataset; high-dimensional multi-scale features of rail transit images are extracted by a mask autoencoder; the high-dimensional multi-scale features are input into a memory attention module to obtain cross-attention calculation results; according to prompt encoding, the encoded image features are decoded by using a visual large model decoder, a target pointer list is determined, and the weights of the mask decoder are initialized; the visual large model is adjusted, the rail operation environment image to be measured is recognized, and the semantic segmentation of the rail image is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of image feature extraction, large visual model fine-tuning, and semantic segmentation in the field of deep learning artificial intelligence, and in particular to an efficient method for large visual model fine-tuning and semantic segmentation for rail transit. Background Technology

[0002] In the rail transit operating environment, an efficient fine-tuning and semantic segmentation method for large-scale visual models is crucial for ensuring the safe operation, fault detection, and autonomous navigation of rail transit. However, current methods still have significant limitations in handling the complex lighting conditions unique to rail transit, blurring caused by high-speed motion, and image quality degradation in extreme environments. Furthermore, the robustness and generalization ability of existing models in the face of rare events or anomalies urgently need improvement. Therefore, developing an efficient fine-tuning and semantic segmentation method for large-scale visual models in rail transit can not only improve the success rate and efficiency of tasks but also enhance the reliability and adaptability of the system, which is of critical significance for the development of rail transit.

[0003] On the one hand, the Segment Anything Model (SAM), the foundational model for image segmentation, stands out for its powerful versatility and flexibility, capable of handling a variety of complex segmentation tasks, from simple objects in images to fine-grained instance segmentation. It introduces a cueing mechanism, allowing users to guide the model to generate accurate segmentation masks through simple clicks, bounding boxes, or text descriptions, greatly improving user experience and segmentation accuracy. Furthermore, SAM is capable of handling multimodal inputs, supports multiple cueing methods, and can output multiple reasonable segmentation results even in the face of ambiguity, enabling it to perform excellently in uncertain environments. However, it still has some shortcomings. First, for large-scale video data, SAM has limited performance in cross-frame consistency, especially when processing long-term video sequences, which may lead to unstable object tracking. Second, SAM has high computational resource requirements, and deployment on resource-constrained devices may face challenges. Finally, while SAM performs well in common vision tasks, its application in specific domains, particularly in orbital operating environments, requires further fine-tuning to adapt to the unique needs and complexities of these domains.

[0004] On the other hand, efficient adaptive fine-tuning is a highly efficient and resource-friendly method that can optimize for specific tasks by introducing lightweight adapter modules without altering the original large model architecture. This approach not only significantly reduces the number of parameters and computational costs but also preserves the general knowledge of the large model, enabling the fine-tuned model to perform well on new tasks while maintaining strong generalization ability. Furthermore, efficient adaptive fine-tuning supports multi-task learning, allowing the model to share some parameters across different tasks, further improving resource utilization efficiency. However, the design and selection of the adapter module have a significant impact on the final performance; improper design may lead to limited or even decreased performance improvement. Secondly, for very complex tasks or datasets, lightweight adapters may be insufficient to capture all necessary features, resulting in model performance inferior to full-parameter fine-tuning. Finally, efficient adaptive fine-tuning typically requires careful tuning of hyperparameters and training strategies to ensure the adapter effectively learns task-specific knowledge, which increases the complexity of model development and deployment.

[0005] In summary, semantic segmentation of orbital operating environments still faces the following key challenges: how to design a large-scale visual model fine-tuning method, and how to construct a comprehensive dataset of orbital operating environments for fine-tuning training of the large-scale visual model. These issues limit the performance of existing methods in orbital operating environments, resulting in insufficient segmentation accuracy of traditional SAM when performing semantic segmentation on images, and consequently, low recognition accuracy of orbital operating environments. Summary of the Invention

[0006] The purpose of this application is to provide an efficient fine-tuning and semantic segmentation method for large visual models of rail transit, in order to solve the problem of low accuracy in rail operating environment recognition.

[0007] To achieve the above objectives, this application provides the following solution.

[0008] This application provides an efficient fine-tuning and semantic segmentation method for large visual models of rail transit, including:

[0009] Automated annotation technology is used to annotate the collected images of the orbital operating environment to construct a semantic segmentation dataset of the orbital operating environment, and the semantic segmentation dataset of the orbital operating environment is used to train a large visual model; the large visual model includes a memory attention module.

[0010] During the training process of the large visual model, the mask autoencoder in the large visual model is used to encode the features of the track operation environment images in the track operation environment semantic segmentation dataset, determine the encoded environment images, and extract multi-scale high-dimensional features from the encoded environment images; the multi-scale high-dimensional features include the memory of the cue image and the memory of the target pointer.

[0011] The multi-scale high-dimensional features are input into the memory attention module, and the memory attention module utilizes the... The block performs cross-attention calculation on the memory of the prompt image and the memory of the target pointer to obtain the cross-attention calculation result.

[0012] Based on the cross-attention calculation results, and according to the prompt encoding, the mask decoder in the large visual model is used to perform mask decoding on the encoded environment image to determine the object pointer list.

[0013] Based on the object pointer list, the weights of the mask decoder are initialized, and the visual big model is adjusted using knowledge of the image segmentation task in the orbital running environment, incorporating prompt images and prompt text.

[0014] The operating environment of the track under test is determined by identifying the image of the track under test based on the adjusted visual large model.

[0015] According to the specific embodiments provided in this application, this application has the following technical effects.

[0016] This application provides an efficient fine-tuning and semantic segmentation method for a large-scale visual model for rail transit. It utilizes automated annotation technology to label collected rail transit environment images, constructing a rail transit environment semantic segmentation dataset. Noise in the dataset is removed using Gaussian filtering, and image enhancement techniques are applied to improve the contrast and clarity, resulting in a preprocessed dataset. This improves annotation efficiency and ensures consistency and accuracy, providing high-quality labeled data for model training. The large-scale visual model is trained using this dataset. The large-scale visual model includes a memory-attention module. During training, a mask autoencoder within the large-scale visual model encodes the features of the rail transit environment images in the semantic segmentation dataset, determining the encoded environment image and extracting multi-scale high-dimensional features from it. These multi-scale high-dimensional features are then input into the memory-attention module, and the memory-attention module utilizes its... The block performs cross-attention calculation on the memory of the cue image and the memory of the target pointer, obtaining the cross-attention calculation results. These features not only contain local details of the image but also integrate global contextual information, providing strong support for subsequent image analysis and understanding. Based on the cross-attention calculation results, according to the cue encoding, the mask decoder in the visual big model is used to perform mask decoding on the encoded environment image to determine the object pointer list. Based on the object pointer list, the weights of the mask decoder are initialized, and the visual big model is adjusted by injecting knowledge of the orbital running environment image segmentation task using the cue image and cue text. Based on the adjusted visual big model, the target orbital running environment image is identified, and the target object pointer list is determined. By defining the operating environment of the track to be tested and fusing embedded images with predicted mask information, the input information of the model can be enriched, helping the model learn more robust feature representations. This fusion method enables the model to better generalize and apply previously learned knowledge when faced with new and unseen video data. The pre-trained visual model is adjusted using an efficient adaptive method to obtain an adjusted visual model. By constructing a track operating environment dataset and fine-tuning the visual model using an efficient adaptive method, this application fully considers the environmental data, making the visual model more suitable for track operating environment-related tasks, better performing semantic segmentation of images, and improving the accuracy of track operating environment recognition. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart illustrating an efficient fine-tuning and semantic segmentation method for a large visual model of rail transit in one embodiment of this application.

[0019] Figure 2 This is a schematic diagram illustrating the working principle of a method for efficient fine-tuning and semantic segmentation of a large visual model for rail transit, provided in an embodiment of this application. Detailed Implementation

[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0021] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0022] This application discloses an efficient fine-tuning and semantic segmentation method for large-scale visual models in rail transit, aiming to achieve efficient fine-tuning of large-scale visual models and accurate semantic segmentation of rail transit operation environment datasets. An automated annotation pipeline for rail transit operation environment data is constructed to achieve large-scale, high-quality data annotation and construction: First, a rail transit operation environment dataset is created, incorporating multiple weather and lighting conditions. Then, preprocessing including denoising, enhancement, and distortion correction is employed, and automated annotation techniques ensure the quality and consistency of the dataset. Furthermore, a semantic segmentation framework for rail transit operation environments is proposed to achieve a general representation extraction and learning method: First, a Hiera image encoder pre-trained using Masked Autoencoders (MAE) extracts multi-scale high-dimensional feature information, and combines these features with historical image features through a memory attention module. Then, sparse and dense cues are processed through a prompt encoder and a mask decoder to generate spatially corresponding output masks. Subsequently, a memory mechanism is constructed to store the memory of recent images, thereby achieving accurate semantic segmentation.

[0023] This study proposes an efficient fine-tuning method for large visual models. It employs an efficient adaptive fine-tuning approach, which optimizes specific tasks by introducing a lightweight adapter module without altering the original large model architecture. This further improves the effectiveness of large visual models in orbital operating environments.

[0024] This application proposes an efficient fine-tuning and semantic segmentation method for large-scale visual models in rail transit. It constructs a high-quality, diverse dataset of rail transit operating environments through automated annotation, enabling the model to effectively address image degradation issues caused by complex lighting, motion blur, and extreme environments. Utilizing a memory attention mechanism and efficient adaptive fine-tuning techniques, the model's robustness and generalization ability under rare events or anomalies are enhanced, ensuring stable performance across different scenarios. Furthermore, a lightweight adapter module is introduced to optimize specific tasks, reducing the number of parameters and computational costs, thus lowering the demand for computational resources and significantly improving the accuracy and efficiency of semantic segmentation in rail transit operating environments.

[0025] like Figure 1 As shown in the embodiments of this application, an efficient fine-tuning and semantic segmentation method for large visual models of rail transit is provided.

[0026] S1: The collected images of the track operation environment are labeled using automated annotation technology to construct a semantic segmentation dataset of the track operation environment, and a large visual model is trained using the semantic segmentation dataset of the track operation environment; the large visual model includes a memory attention module.

[0027] S2: During the training process of the large visual model, the mask autoencoder in the large visual model is used to encode the features of the track operation environment images in the track operation environment semantic segmentation dataset, determine the encoded environment images, and extract multi-scale high-dimensional features from the encoded environment images; the multi-scale high-dimensional features include the memory of the cue image and the memory of the target pointer.

[0028] A cue image is one or more images input into a large visual model to guide a specific task, such as image segmentation. It provides information about how to process or understand another main image.

[0029] S3: Input the multi-scale high-dimensional features into the memory attention module, and utilize the features in the memory attention module. The block performs cross-attention calculation on the memory of the prompt image and the memory of the target pointer to obtain the cross-attention calculation result.

[0030] S4: Based on the cross-attention calculation results, according to the prompt encoding, the mask decoder in the large visual model is used to perform mask decoding on the encoded environment image to determine the object pointer list.

[0031] Cue encoding is the process or result of converting cue images into a form that the model can understand and process.

[0032] S5: Based on the object pointer list, initialize the weights of the mask decoder, and adjust the visual big model by incorporating knowledge of the image segmentation task of the orbital running environment using the prompt image and prompt text.

[0033] S6: Identify the image of the test track's operating environment based on the adjusted visual large model, and determine the test track's operating environment.

[0034] Existing semantic segmentation algorithms still have significant shortcomings in the context of railway operation environments. On the one hand, there is a lack of automated pipeline methods for constructing high-quality datasets for railway operation environments. On the other hand, these algorithms perform poorly in handling the complex lighting conditions unique to railways, blurring caused by high-speed motion, and image quality degradation under extreme environments. Furthermore, extreme environments can lead to a decline in the quality of sensor data, further affecting segmentation results. On the other hand, the robustness and generalization ability of existing models in the face of rare events or anomalies urgently need improvement. These situations may be extremely rare in the training data, causing the model to be unable to effectively identify and segment these key regions in practical applications. Moreover, the diversity and dynamism of railway tasks require models to have strong adaptability and maintain stable performance under different scenarios and conditions, and the performance of existing models in this regard still needs improvement. Therefore, this application provides an efficient fine-tuning and semantic segmentation method for large-scale visual models in railway transportation. It constructs a dataset for railway operation environments through automated annotation and uses an efficient adaptive method to fine-tune the large-scale visual model, solving the problems of poor segmentation results and model robustness in existing methods for railway operation environments.

[0035] Furthermore, in an exemplary embodiment, step S1 can be replaced by the following steps.

[0036] S101: Acquire images of the orbital operating environment from multi-source data acquisition and open-source orbital datasets.

[0037] In an efficient fine-tuning and semantic segmentation method for large visual models of rail transit, it is necessary to construct a high-quality, representative semantic segmentation dataset of the rail operating environment. This application employs automated annotation technology, collecting rich images containing various weather conditions, different lighting conditions, and diverse rail structures through multi-source data acquisition and open-source rail datasets.

[0038] S102: Remove noise from the track operation environment image based on Gaussian filtering to determine the denoised track operation environment image.

[0039] S103: Use automated annotation technology to annotate the denoised track operation environment image to determine the track operation environment image with annotations.

[0040] S104: Enhance the contrast and clarity of the labeled track operation environment image using image enhancement technology to obtain the enhanced track operation environment image and construct a track operation environment semantic segmentation dataset.

[0041] In the preprocessing stage, Gaussian filtering is used to remove noise, and image enhancement techniques are applied to improve contrast and clarity, making the features of the track target objects more prominent. For geometrically distorted images, perspective transformation is performed to correct the distortion, ensuring the accurate shape and proportion of the track elements. Subsequently, this application utilizes automated annotation technology to generate masks of the track target objects from the images and prompts, improving annotation efficiency and ensuring the consistency and accuracy of annotations, providing high-quality labeled data for model training.

[0042] Furthermore, in an exemplary embodiment, step S2 can be replaced by the following steps.

[0043] S201: Using a masked autoencoder pre-trained image encoder, extract image features from the track operation environment image to obtain multi-scale high-dimensional feature information of the image; the pre-trained image encoder includes a ViT-H / 16 model with a 14x14 window attention and four equally spaced global attention blocks.

[0044] For the constructed semantic segmentation dataset of the orbital operation environment, image features are first extracted using a Hiera image encoder pre-trained with Masked Autoencoders (MAE), thus obtaining multi-scale high-dimensional feature information of the current image. Based on this, memory attention is calculated on the extracted features to adjust the current image features, historical image features, prediction results, and new prompts.

[0045] S202: Based on the multi-scale high-dimensional feature information of the image, a pre-trained image encoder is used to perform feature encoding on the track operation environment image to determine the multi-scale high-dimensional feature representation of the track operation environment image.

[0046] Multi-scale image feature extraction network image For input, where, The height of the image; The width of the image is used; the Hiera image encoder pre-trained with MAE is used to encode features of the orbital operating environment image. The outputs of the third and fourth stages of the feature encoder are used as multi-scale high-dimensional feature representations of the image, i.e., the size is [missing information]. and The image feature maps are used to generate embeddings for each image. Features from the first and second stages with sampling strides of 4 and 8, respectively, are not used for the memory attention mechanism, but are added to the upsampling layer of the mask decoder to help generate high-resolution segmentation details.

[0047] The memory attention module takes the multi-scale feature representation extracted from the current image as input. This application employs four... Each block performs self-attention, then performs cross-attention calculations on the cue image and the memory of the target pointer, stores the information in a memory bank, and then performs feature mapping through a multilayer perceptron (MLP). In addition to sinusoidal absolute position embedding, the module uses two-dimensional spatial rotational position embedding in the self-attention layer and cross-attention layer, which can be specifically represented as follows:

[0048] .

[0049] .

[0050] .

[0051] in, The coordinate components of the two-dimensional spatial location of the image, such as rows and columns, are used to generate rotation angles, thereby injecting spatial location information into the feature representation; for dimensional vector; This is the grouping index along the feature dimension, with a value range of [value range missing]. ; The rotation angle; It is a hyperparameter; It is an encoded vector.

[0052] Furthermore, in an exemplary embodiment, step S4 can be replaced by the following steps.

[0053] S401: Based on the cross-attention calculation results, the sparse and dense cue texts are mapped to 256-dimensional vector embeddings using a cue encoder to obtain image embeddings; the 256-dimensional vector embeddings are text features.

[0054] S402: Using a mask decoder, the image embedding and the prompt encoding are mapped to an output mask to determine the mask information; the mask decoder includes a modified... A decoder block and a dynamic mask prediction head.

[0055] S403: Construct a list of object pointers based on the mask information.

[0056] The outputs of the third stage and the fourth stage of the feature encoder are used as the multi-scale high-dimensional feature representations of the image.

[0057] After extracting multi-scale image features and calculating memory attention, the model is further provided with cue encoding and mask decoding. First, the cue encoder maps sparse and dense cues to 256-dimensional vector embeddings. For sparse cues, the representation is obtained by adding a positional encoding to one of two learned embeddings, which indicate whether the point is foreground or background, respectively. To handle free-format text, this application uses a text encoder from the Contrastive Language-Image Pre-training (CLIP) model.

[0058] S411: Based on the image embedding, input the input mask into the large visual model, and reduce the resolution of the input mask using the first convolution operation.

[0059] The mask has a spatial correspondence with the image. Specifically, the input mask enters the model at a resolution 4 times lower than the input image, and is then further reduced by a factor of 4 through two 2×2 convolutions with a stride of 2, which have 4 and 16 output channels, respectively.

[0060] S412: The second convolution operation is used to map the dimension of the reduced input mask to 256-dimensional image features to obtain the mask embedding.

[0061] The channel dimensions are mapped to 256 dimensions using a 1×1 convolution, with Gaussian Error Linear Unit (GELU) activation functions and layer normalization inserted between each layer. The mask embedding is then element-wise added to the image embedding. If no mask hint is provided, a learned embedding representing "no mask" is added to each image embedding location.

[0062] S413: Add the mask embedding to the image embedding to obtain the output mask.

[0063] The output mask refers to the result obtained after processing by the mask decoder. The result identifies the location and extent of a specific object or region in the image, and is used to indicate which pixels in the image belong to a certain identified object or region of interest.

[0064] S414: Based on the output mask, according to the modified... The decoder performs a self-attention operation on the cue markers and a cross-attention operation from the cue markers to the image embedding, and determines the result of the cross-attention operation.

[0065] S415: Based on the cross-attention operation result, update each cue marker using pointwise MLP, perform cross-attention operation from image embedding to cue marker, update the image embedding, and obtain a new image embedding.

[0066] The decoder module maps the image embedding and a set of cue embeddings to an output mask. To combine these inputs, this application modifies the standard... Decoder. Before applying the decoder, a learned output tag embedding is inserted into the cue embedding set, which will be used in the decoder's output, similar to category tags.

[0067] S416: Predict mask information based on the spatial pointwise product between the new image embedding and the pointwise MLP output.

[0068] The process involves self-attention on the labels, followed by cross-attention from the labels (as queries) to the image embeddings. This is then followed by a pointwise MLP updating each label, and finally, cross-attention from the image embeddings (as queries) to the labels, updating the image embeddings with cue information. During cross-attention, the image embeddings are treated as a set of 64×64 256-dimensional vectors, each with a residual connection, layer normalization, and a dropout rate of 0.1 during training. The next decoder layer receives the updated labels and updated image embeddings from the previous layer.

[0069] To ensure the decoder can capture crucial geometric information, positional encoding is added to the image embedding whenever it participates in the attention layer. Furthermore, whenever the original cue tags participate in the attention layer, they are re-added to the updated tags. After running the decoder, the updated image embedding is upsampled by a factor of 4 using two transposed convolutional layers. Then, the updated output tag embedding is passed to a small 3-layer MLP that outputs a vector matching the channel dimensions of the upsampled image embedding. Finally, the mask is predicted using the spatial pointwise product between the upsampled image embedding and the MLP output.

[0070] Specifically, this application is in A 256-dimensional embedding is used, with an internal dimension of 2048 within the MLP module, but this is only applied to a relatively small number of cue tags. In the cross-attention layer, for computational efficiency, the channel dimensions of the query, key, and value are reduced by a factor of 2 to 128. All attention layers use 8 heads. The transposed convolutions used for embedding the upsampled output image are 2×2 with a stride of 2, resulting in output channel dimensions of 64 and 32 respectively, and are equipped with the GELU activation function and layer normalization separation.

[0071] S421: Based on the mask information, construct a memory bank; the memory bank includes information about the most recent image and information about the prompt image.

[0072] This application completes feature extraction, computational memory attention, cue encoding, and mask decoding for images of the orbital operating environment. Simultaneously, it constructs a memory bank mechanism. The memory bank operates in a first-in-first-out queue format and can store a maximum of [number missing] [items missing]. The most recent image and the most Information from each cue image is stored in the form of spatial feature maps. In image segmentation tasks, the initial mask serves as a key cue, and the memory always retains the initial image and... The memory stores information about the most recent images, which can be understood as information about the most recently processed or most recently added images to this memory. In addition to the spatial memory, there is also a list of object pointers constructed based on the output tags of each image mask decoder, carrying high-level semantic information about the objects. Memory attention cross-focuses on the spatial memory features and these object pointers. For temporal location information, it is only embedded in... The memory of the most recent images is helpful for the model to express the positional changes of the target object in the image sequence.

[0073] S422: Based on the memory encoder and the memory library, the new image embedding is used, and the mask information is fused to obtain the memory features.

[0074] In the collaboration between the memory encoder and the memory bank, the image embedding generated by the Hiera encoder is reused, and the predicted mask information is fused to generate memory features, so that the memory features can benefit from the powerful representational capabilities of the image encoder.

[0075] S423: Project the memory features and split the memory of the target pointer to obtain the memory mark of the target pointer.

[0076] S424: Based on the memory markers of the target pointer, perform cross-attention operations on the memory bank to determine the pre-trained large visual model.

[0077] The memory features in the memory bank are projected onto 64 dimensions, and the 256-dimensional object pointers are split into four 64-dimensional tags to perform cross-attention operations on the memory bank. For multi-object segmentation tasks in the same image, a strategy of independent reasoning for each object is adopted. All objects share the visual features of the image encoder, but each object runs its own memory bank and other model components such as the mask decoder.

[0078] Furthermore, in an exemplary embodiment, step S5 can be replaced by the following steps.

[0079] S501: During the training process of the large visual model, the mask decoder is adjusted to obtain a new mask decoder.

[0080] Based on the above steps, image feature extraction, memory attention calculation, and storage for the orbital operating environment are achieved. Based on this, this application employs an efficient adaptive method to fine-tune the pre-trained large-scale visual model, making it more suitable for segmenting orbital operating environment image data. Specifically, this application uses the MAE pre-trained Hiera image encoder as the backbone of the segmentation network. This image encoder is a ViT-H / 16 model with 14x14 window attention and four equally spaced global attention blocks. This application keeps the weights of the pre-trained image encoder unchanged. Furthermore, this application utilizes the mask decoder of the pre-trained large-scale visual model, which consists of a modified... The method consists of a decoder block and a dynamic mask prediction head. This application initializes the weights of the mask decoder in its method using the weights of a pre-trained large visual model, while simultaneously adjusting the mask decoder during training, and without providing any cues to the model's original mask decoder.

[0081] The learning process employs an efficient adaptive approach, incorporating knowledge from the orbital operating environment image segmentation task. i Efficient adaptive fine-tuning employs the concept of prompts, thereby leveraging a base model already trained on large-scale datasets. Furthermore, using appropriate prompts to introduce task-specific knowledge can enhance the model's generalization ability on downstream tasks, especially in situations where labeled data is scarce.

[0082] S502: Based on the new mask decoder, knowledge of the orbital operating environment image segmentation task is injected using prompt images and prompt text.

[0083] S503: Based on the aforementioned knowledge, using the formula The output prompt is received; among them, For the upper projection layer shared across all efficient adaptive blocks; For each attached to the SAM model Layer output prompts; For activation functions; A linear layer specifically designed for each adaptive block; Task-specific knowledge.

[0084] like Figure 2 As shown, specifically, this application uses an efficient adaptive block containing only two MLPs and one activation function, and initializes four different efficient adaptive blocks. These four efficient adaptive blocks are then inserted into different layers at each stage. Furthermore, the weights of the efficient adaptive blocks are shared at each stage, meaning each efficient adaptive block receives information. And get the prompt output:

[0085] .

[0086] in, It is a linear layer used to generate specific task hints for each efficient adaptive block. It is an upprojection layer shared across all efficient adaptive blocks. It adjusts the dimension of the features of the efficient adaptive blocks. Adaptive blocks are a type of parameter-efficient fine-tuning method. They are modules used in deep learning models for fine-tuning or transfer learning. They are usually inserted between the layers of a pre-trained model to adapt to new tasks or datasets without requiring large-scale retraining of the entire model. This indicates that each element is attached to the SAM model. Output prompts for the layer. The activation function is ,in, Here is the activation function; x is the input to... The variables in the activation function are the output values ​​of the neural network layers; The error function is defined as follows. A dataset of orbital operating environments is constructed, and a large visual model is fine-tuned using an efficient adaptive approach. This makes the large visual model more suitable for tasks related to orbital operating environments and improves its ability to perform semantic segmentation of images.

[0087] S504: Adjust the visual large model based on the output prompt.

[0088] Following the steps outlined above, this application presents a method for efficient fine-tuning and semantic segmentation of a large visual model for rail transit. This method constructs a high-quality, representative semantic segmentation dataset for the rail operating environment through automated annotation, and combines multi-source data acquisition with open-source rail datasets to ensure data diversity. This enables the model to better handle the complex lighting conditions unique to rail transit, blurring caused by high-speed motion, and image quality degradation in extreme environments. Furthermore, the application of a memory attention mechanism and an efficient adaptive fine-tuning method further enhances the model's robustness and generalization ability in the face of rare events or abnormal situations in the rail operating environment, thus maintaining stable performance under different scenarios and conditions. This application's method for efficient fine-tuning and semantic segmentation of a large visual model for rail transit employs a large visual model to improve the accuracy of semantic segmentation in the rail operating environment. Based on this, it uses an efficient adaptive fine-tuning approach, without changing the original large model architecture, to achieve task-specific optimization by introducing a lightweight adapter module. This significantly reduces the number of parameters and computational costs, lowers the demand for computational resources, and further improves the effectiveness of the large visual model in the rail operating environment.

[0089] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0090] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A visual large model fine-tuning and semantic segmentation method for rail transit, characterized in that, The rail transit-oriented visual large model efficient fine-tuning and semantic segmentation method comprises: The collected rail operation environment images are labeled by using an automatic labeling technology, a rail operation environment semantic segmentation dataset is constructed, and a visual large model is trained by using the rail operation environment semantic segmentation dataset; the visual large model comprises a memory attention module; In the training process of the visual large model, the rail operation environment images in the rail operation environment semantic segmentation dataset are feature-encoded by using a mask autoencoder in the visual large model, the encoded environment images are determined, and multi-scale high-dimensional features in the encoded environment images are extracted, specifically comprising: An image encoder pre-trained by a mask autoencoder is used to extract image features in the track running environment image, and image multi-scale high-dimensional feature information is obtained. A ViT-H / 16 model with window attention and four equally spaced global attention blocks. Based on the multi-scale high-dimensional feature information of the image, a pre-trained image encoder is used to encode the features of the track operation environment image to determine the multi-scale high-dimensional feature representation of the track operation environment image; wherein the MAE pre-trained Hiera image encoder is used to encode the features of the track operation environment image, and the outputs of the third and fourth stages of the feature encoder are used as the multi-scale high-dimensional feature representation of the image, i.e. the image feature map with a size of and The embedding is generated for each image; the features with sampling steps of 4 and 8 from the first and second stages are added to the up-sampling layer of the mask decoder; the multi-scale high-dimensional features include the memory of the prompt image and the memory of the target pointer; The multi-scale high-dimensional features are input into the memory attention module, and the memory attention module is used to The block performs cross-attention calculation on the memory of the prompt image and the memory of the target pointer to obtain a cross-attention calculation result. The memory attention module adopts 4 blocks, each block performs self-attention, then cross-attention calculation is performed on the memory of the prompt image and the target pointer and stored in the memory bank, and then feature mapping is performed through a multi-layer perception; In addition to the sine absolute position embedding in the module, a two-dimensional space rotation position embedding is used in the self-attention layer and the cross-attention layer, which is specifically represented as: ​ ; ; ; wherein, is a coordinate component of the image two-dimensional spatial position; is a vector of dimensionality; is a group index in the feature dimensionality, taking values in ; is a rotation angle; is a hyper-parameter; is the encoded vector; Based on the cross-attention calculation result, object pointer lists are determined by using a prompt encoder to perform mask decoding on the encoded environment images in the visual large model according to prompt encoding; Based on the object pointer lists, the weights of the mask decoder are initialized, and the visual large model is adjusted by using a prompt image and a prompt text to inject knowledge of a rail operation environment image segmentation task; The rail operation environment to be detected is identified according to the adjusted visual large model, and a rail operation environment to be detected is determined.

2. The rail transit-oriented visual large model fine-tuning and semantic segmentation method according to claim 1, characterized in that, The collected rail operation environment images are labeled by using an automatic labeling technology, a rail operation environment semantic segmentation dataset is constructed, specifically comprising: Rail operation environment images are collected from multi-source data collection and open-source rail data; Based on Gaussian filtering, the noise of the rail operation environment images is removed, and the rail operation environment images after noise reduction are determined; The rail operation environment images after noise reduction are labeled by using an automatic labeling technology, and the rail operation environment images with labels are determined; The contrast and clarity of the rail operation environment images with labels are enhanced by using an image enhancement technology, and the enhanced rail operation environment images are obtained, thereby constructing a rail operation environment semantic segmentation dataset.

3. The rail transit-oriented visual large model fine-tuning and semantic segmentation method according to claim 1, characterized in that, Based on the cross-attention calculation result, object pointer lists are determined by using a prompt encoder to perform mask decoding on the encoded environment images in the visual large model according to prompt encoding, specifically comprising: Based on the cross-attention calculation result, sparse prompt text and dense prompt text are mapped to 256-dimensional vector embedding by using a prompt encoder, and image embedding is obtained; the 256-dimensional vector embedding is a text feature; determining mask information by mapping the image embedding and the cue encoding to an output mask using a mask decoder; the mask decoder comprising a modified decoder block and a dynamic mask prediction head; According to the mask information, an object pointer list is constructed.

4. The rail transit-oriented visual large model efficient fine-tuning and semantic segmentation method according to claim 3, characterized in that, The image embedding and the prompt encoding are mapped to an output mask by using a mask decoder, and mask information is determined, specifically comprising: Based on the image embedding, an input mask is input into the visual large model, and the resolution of the input mask is reduced by using a first convolution operation; The dimension of the reduced input mask is mapped to 256-dimensional image features by using a second convolution operation, and a mask embedding is obtained; The mask embedding and the image embedding are added, and an output mask is obtained; based on the output mask, determining, according to the modified The decoder performs a self-attention operation on the prompt tokens and a cross-attention operation from the prompt tokens to the image embeddings to determine a cross-attention operation result; Based on the cross-attention operation result, each prompt token is updated by using a point-wise MLP, and cross-attention operation is performed from image embedding to prompt token, the image embedding is updated, and a new image embedding is obtained; Predicting mask information according to a spatial point-by-point product between a new image embedding and the point-by-point MLP output.

5. The rail transit oriented visual large model efficient fine-tuning and semantic segmentation method according to claim 4, characterized in that, Predicting mask information according to a spatial point-by-point product between a new image embedding and the point-by-point MLP output, and then further comprising: Based on the mask information, a memory bank is constructed; the memory bank includes information of a recent image and information of a prompt image; Based on the memory encoder and the memory bank, a memory feature is obtained by using the new image embedding and fusing the mask information; The memory feature is projected, and the memory of the target pointer is split to obtain a memory mark of the target pointer; Based on the memory mark of the target pointer, a cross-attention operation is performed on the memory bank to determine a pre-trained visual large model.

6. The rail transit-oriented visual large model efficient fine-tuning and semantic segmentation method according to claim 5, characterized in that, Based on the object pointer list, the weights of a mask decoder are initialized, and the knowledge of a track running environment image segmentation task is injected into the prompt image and the prompt text to adjust the visual large model, specifically including: In the training process of the visual large model, the mask decoder is adjusted to obtain a new mask decoder; Based on the new mask decoder, the knowledge of the track running environment image segmentation task is injected into the prompt image and the prompt text; Based on the knowledge, the output prompt is obtained using the formula wherein, is an upper projection layer shared across all efficient adaptive blocks; is an output prompt appended to each layer of the SAM model; is an activation function; is a linear layer designed specifically for each adaptive block; is task-specific knowledge; Based on the output prompt, the visual large model is adjusted.

Citation Information

Patent Citations

  • Steel rail surface defect segmentation method based on SAM2-CLIP model

    CN120107604A