A track defect detection method and device based on a multi-modal large model
By employing a multimodal large model and a multi-stage training strategy, the problems of single modality and interpretability in orbital defect detection are solved, enabling broader defect detection and human-readable detection results, thereby improving detection effectiveness and efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-23
- Publication Date
- 2026-03-03
AI Technical Summary
Existing deep learning-based track defect detection technologies suffer from limitations in single-modal detection capabilities, difficulty in achieving generalized detection, lack of multimodal interaction, and low interpretability.
A multimodal large model is used for track defect detection. The visual perception capability is enhanced by a diffusion denoising decoder. A multi-stage training strategy is designed, including post-training fine-tuning, cold-start supervised fine-tuning, and reinforcement fine-tuning. The GRPO algorithm is combined to optimize the model performance.
It achieves better generalization and interpretability, enabling the detection of multiple defect types in the rail transit field and generating detection results in a human-readable form, while reducing computational overhead and context length.
Smart Images

Figure CN120997487B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of track inspection technology, and in particular to a method and apparatus for detecting track defects based on a multimodal large model. Background Technology
[0002] In rail transit maintenance, track defect detection is a crucial component, primarily detecting defects in the track structure that could affect the normal operation of rail transit. The objects of track defect detection are diverse, encompassing critical components such as rails, ballast, and fasteners, and involving various defect types including rail surface scratches, rail spalling, fastener deformation, and ballast cracks. Traditional detection methods mainly utilize ultrasonic and electromagnetic sensors for non-destructive testing. With the development of artificial intelligence technology, track defect detection methods based on machine vision technology are increasingly being used due to their advantages over sensor-based methods, such as low cost and high efficiency. A vision-based track defect detection system typically involves three processes: hardware, data, and algorithms—a high-speed camera on a detection train collects grayscale image data of the track scene, and computer vision technology is used to complete defect detection.
[0003] Previous methods for track defect detection based on machine vision mainly fell into two categories. The first was based on digital image processing. These methods typically utilized the pixel differences between defect locations and background locations in track images, such as contrast-based foreground models and consistency-based background models. However, these methods relied on manually constructed features and lacked robustness in practical applications. The second category was based on deep learning computer vision techniques. These methods typically used object detection visual models such as convolutional neural networks, YOLO, and RetinaNet, trained on track defect datasets to achieve the track defect detection task.
[0004] With the rapid development of deep learning technology, deep learning-based track defect detection technology has gradually become the mainstream method in the field. However, in the context of large-scale models driving the intelligent development of the field, the following problems also arise:
[0005] 1. Previous deep learning-based track defect detection techniques only used single-modality computer vision techniques, which had many problems, such as single defect detection models being able to detect only one type of defect and being difficult to achieve generalization detection; while multi-defect detection models rely on a large amount of track defect detection dataset to achieve good generalization results and are prone to overfitting.
[0006] 2. Previous deep learning-based track defect detection technologies only used vision as domain learning information, without other modal prior knowledge to help achieve defect detection. They lacked multimodal interaction of defect information, making it difficult to achieve good detection results for both seen and unseen defects.
[0007] 3. Previous deep learning-based track defect detection technologies could only perform vision-related detection tasks, and could not provide human-readable detection reasoning processes, resulting in low interpretability and limited applicability. Summary of the Invention
[0008] To address the technical problems of overfitting and poor track defect detection in existing technologies, this invention provides a track defect detection method and apparatus based on a multimodal large model. The technical solution is as follows:
[0009] On the one hand, a method for detecting track defects based on a multimodal large model is provided. This method is implemented by a track defect detection device based on a multimodal large model, and includes:
[0010] S1. Obtain orbital training samples and a model to be trained. The model to be trained includes a basic multimodal large model to be trained, a diffusion denoising decoder, and a large model target detection module.
[0011] S2. Based on the orbital training samples, the basic multimodal large model to be trained, and the diffusion denoising decoder, the basic multimodal large model to be trained is subjected to post-training fine-tuning based on orbital knowledge to obtain the first basic multimodal large model.
[0012] S3. Based on the orbit training samples and the first basic multimodal large model, perform cold start supervised fine-tuning of the first basic multimodal large model based on the orbit multimodal target detection task to obtain the second basic multimodal large model;
[0013] S4. Based on the track training samples, the second basic multimodal large model and the GRPO algorithm, the second basic multimodal large model is enhanced and fine-tuned based on the track multimodal target detection task, and the trained basic multimodal large model is determined as the track defect detection model.
[0014] S5. Input the track image to be detected into the track defect detection model to obtain the corresponding defect detection results.
[0015] On the other hand, a track defect detection device based on a multimodal large model is provided. This device is applied to a track defect detection method based on a multimodal large model. The device includes:
[0016] The acquisition unit is used to acquire orbital training samples and a model to be trained. The model to be trained includes a basic multimodal large model to be trained, a diffusion denoising decoder, and a large model target detection module.
[0017] The post-training fine-tuning unit is used to perform post-training fine-tuning based on orbital knowledge on the basic multimodal large model to be trained, based on orbital training samples, the basic multimodal large model to be trained, and the diffusion denoising decoder, to obtain the first basic multimodal large model.
[0018] The cold start supervised fine-tuning unit is used to perform cold start supervised fine-tuning of the first basic multimodal large model based on the orbit training samples and the first basic multimodal large model, so as to obtain the second basic multimodal large model.
[0019] The enhancement and fine-tuning unit is used to enhance and fine-tune the second basic multimodal large model based on the track training samples, the second basic multimodal large model and the GRPO algorithm, and to determine the trained basic multimodal large model as the track defect detection model.
[0020] The detection unit is used to input the track image to be detected into the track defect detection model to obtain the corresponding defect detection results.
[0021] On the other hand, a track defect detection device based on a multimodal large model is provided. The track defect detection device based on a multimodal large model includes: a processor; a memory, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, any one of the above-described track defect detection methods based on a multimodal large model is implemented.
[0022] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored therein, the at least one instruction being loaded and executed by a processor to implement any of the above-described methods for track defect detection based on a multimodal large model.
[0023] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following:
[0024] This invention proposes a track defect detection method based on a multimodal large model. By designing a diffusion denoising decoder, the method enhances the large model's visual perception capability of track images during training. A corresponding multi-stage training strategy of fusion enhancement and fine-tuning is also designed. This allows for the use of text in contextual scenarios to perform visual task reasoning for track defects, resulting in better task interpretability. Furthermore, because a large model is used to implement tasks in the rail transit domain and corresponding domain data is constructed for training, the large model exhibits better generalization in vertical domain knowledge, enabling it to perform defect detection tasks with a wider range of categories. It can also interact with users in a human-readable format to complete domain instruction generation tasks, resulting in better detection performance. This invention designs a target detection module adapted to contextual scenarios, effectively reducing the context length during user interaction and lowering the computational cost of the large model. In addition, the multimodal large model designed in this invention undergoes sufficient multimodal target detection training, including track defect detection tasks. It can perceive track geometry and align with track-related text corpora, learn the potential relationship between the visual features of track defect locations and the linguistic input representing defects, and possess the ability to detect known defect types through specific linguistic instructions and the potential for open-vocabulary detection of unseen defect types. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 This is a flowchart of a track defect detection method based on a multimodal large model provided by an embodiment of the present invention;
[0027] Figure 2 This is a schematic diagram of the structure of a basic multimodal large model combined with a diffusion denoising decoder provided in an embodiment of the present invention;
[0028] Figure 3 This is a schematic diagram of an example of an input model's orbital image provided in an embodiment of the present invention;
[0029] Figure 4 This is a block diagram of a track defect detection device based on a multimodal large model provided in an embodiment of the present invention;
[0030] Figure 5 This is a schematic diagram of the structure of a track defect detection device based on a multimodal large model provided in an embodiment of the present invention. Detailed Implementation
[0031] The technical solution of the present invention will now be described with reference to the accompanying drawings.
[0032] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.
[0033] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.
[0034] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.
[0035] To make the technical problem to be solved, the technical solution, and the advantages of the present invention clearer, a detailed description will be provided below in conjunction with the accompanying drawings and specific embodiments. The following is a brief explanation of the terminology:
[0036] Diffusion Model: A state-of-the-art deep learning model used in image generation, natural language processing, and other fields. The basic idea is to iteratively add Gaussian noise to data (such as images) during the forward pass, eventually transforming the data into pure noise; the model then removes the noise progressively during the backward pass, learning to reconstruct the original data from the random noise. Trained diffusion models exhibit excellent performance in tasks such as image generation.
[0037] The diffusion denoising decoder (Denoiser) is a neural network model that uses the diffusion concept to implement the denoising process as a decoder.
[0038] Flow Matching: Unlike diffusion models, which define a fixed stochastic process and learn its reversal, flow matching directly learns a velocity field that can be transformed along a flexible path. Simply put, it reduces the multi-step denoising process to a single step. Flow matching can be considered to retain the core advantages of diffusion models while simplifying the technique by eliminating the limitations of the forward noise process.
[0039] Reinforcement learning is an important research area in deep learning. Its main idea is to gain rich experience in attempting to solve problems through extensive trial and error and corresponding environmental rewards, learning the probability of finding optimal solutions from this experience. Previous reinforcement learning was limited to learning agent behavior in simulated environments. ChatGPT proposed Human Feedback Reinforcement Learning (HFRL), which uses human evaluations of the quality of responses generated by a large language model to strengthen the model's ability to align with human preferences, inspiring subsequent applications of reinforcement learning in large models.
[0040] Large Language Models (LLMs): The vast majority of large language models are implemented based on the Transformer Decoder architecture, typically with over a billion parameters. Training of large language models usually involves self-supervised pre-training and instruction fine-tuning. Pre-training on massive amounts of data endows the large language models with rich textual knowledge, while instruction fine-tuning enables them to output relevant answers with human prompts. Classic large language models include ChatGPT, Tongyiqianwen, and DeepSeek.
[0041] Multimodal Large Language Models (MLLMs): These models build upon large language models by adding a pre-trained visual encoder to provide visual features. Typically, due to a misalignment in the feature space between the pre-trained visual encoder and the large language model, MLLMs usually require alignment training to enable the large language model to effectively process visual features from the visual encoder. Then, fine-tuning is performed to improve the model's response capabilities to input images.
[0042] Fine-tuning refers to further training a pre-trained model using a dataset from the target task to improve its performance on a specific task, which helps the model adapt to new tasks or domains.
[0043] Open-vocabulary refers to models with sufficient training data, enabling them to achieve good fitting ability on known data and good generalization ability on unknown data. In image-text multimodal data, open-vocabulary typically leverages the richness and diversity of language expression to achieve visually relevant tasks.
[0044] This invention provides a method for detecting track defects based on a multimodal large model. This method can be implemented using a track defect detection device based on a multimodal large model, which can be a terminal or a server. Figure 1 The flowchart shown is for a track defect detection method based on a multimodal large model. The processing flow of this method may include the following steps:
[0045] S1. Obtain orbital training samples and the model to be trained. The model to be trained includes the basic multimodal large model to be trained, the diffusion denoising decoder, and the large model target detection module.
[0046] The track training samples include track defect sample images and corresponding defect description sample texts.
[0047] The diffusion denoising decoder uses a Diffusion Transformer structure.
[0048] The basic multimodal large model includes a CLIP visual encoder, a projection layer, a large language model, and a decoder stack network.
[0049] In one feasible implementation, the basic multimodal large model is a Llava (Large Language and Vision Assistant) architecture, comprising three network modules: a CLIP (Contrastive Language–Image Pre-training) visual encoder, a projection layer, and a large language model. The CLIP visual encoder uses large-scale image-text data to train the visual portion of the pre-trained CLIP model. The projection layer is a linear layer that aligns the features extracted by the CLIP visual encoder with the word vector space of the large language model, thereby giving the large language model a certain degree of visual perception capability. The large language model is obtained by adjusting the weights through post-training based on the pre-trained Llava model.
[0050] The diffusion denoising decoder is a Diffusion Transformer architecture that adds conditional features to the original Transformer to control the denoising generation process. It uses the visual features output by the large language model as conditional features for denoising, guiding the learning of the original image features after adding noise to the fine-grained image features extracted from the VAE, thereby enhancing the large language model's ability to process visual features extracted by the CLIP visual encoder.
[0051] The large-scale object detection module, adapted to specific contexts, consists of two modules: a bounding box embedding layer and a bounding box regression layer. Both networks are composed of two-layer perceptual mechanisms. Because the word segmenter and generation process of the large language model split the coordinates of the detected object into individual characters, this excessively increases the context length, affecting the large model's understanding of the context and its generation performance. Therefore, this invention designs to use the content of the detected object as a token feature of the large language model, and uses the bounding box regression method (i.e., the bounding box regression layer) from previous computer vision object detection methods to calculate the final predicted bounding box coordinates. During the large model generation process, if the bounding box coordinates have already been predicted, they are used through the bounding box embedding layer. The bounding box embedding layer maps the normalized coordinates in the input text to the word vector space of the large language model, and inputs them along with other plain text data into the large language model to participate in the generation process.
[0052] S2. Based on the orbital training samples, the basic multimodal large model to be trained, and the diffusion denoising decoder, the basic multimodal large model to be trained is subjected to post-training fine-tuning based on orbital knowledge to obtain the first basic multimodal large model.
[0053] In one feasible implementation, to enhance the task of inferring track visual defect detection using text in a multimodal large model, this invention designs a multi-stage training strategy that integrates reinforcement fine-tuning. This strategy comprises three stages: post-training fine-tuning based on track knowledge, cold-start supervised fine-tuning based on the track multimodal target detection task, and reinforcement fine-tuning based on the track multimodal target detection task. The track multimodal target detection task includes target detection of track components and target detection of track defects.
[0054] To enhance the visual perception capability of large language models in track vision detection tasks and reduce the training time required for multimodal large models, this embodiment of the invention uses a flow matching method to implement a diffusion denoising decoder for post-training fine-tuning based on track knowledge. This achieves one-step denoising instead of multi-step denoising, thereby enhancing the image processing capability of the large language model within the multimodal large model. In the post-training fine-tuning stage based on track knowledge, domain-specific knowledge related to rail transit operation and maintenance, including defect types and attributes, is collected from the internet and official documentation. This knowledge is then used to generate question-and-answer format instruction data and supervised fine-tuning, allowing the multimodal large model to learn domain knowledge. This stage does not include track defect target detection tasks and corresponding datasets; it only involves training based on the diffusion decoder and supervised fine-tuning.
[0055] Optionally, such as Figure 2 As shown, the specific operations of S2 may include:
[0056] S21. Input the track defect sample image into the basic multimodal large model to be trained to obtain the predicted text response and predicted visual output.
[0057] Optionally, the specific operations of S21 may include:
[0058] S211. Input the track defect sample image into the CLIP visual encoder ViT, divide the track defect sample image into multiple pixel blocks, and extract the initial visual features through a multi-layer encoder module consisting of a self-attention layer and a perceptron.
[0059] In one feasible implementation, the input track defect sample image , and These represent the height and width of the image, respectively. The image is then divided into [variables] by the CLIP visual encoder (ViT). Visual features can be extracted from pixel blocks through forward propagation via a multi-layer encoder module consisting of self-attention layers and a perceptron. :
[0060] (1)
[0061] in , That is, the length of the image sequence. It is the size of the image feature dimension.
[0062] S212. The initial visual features are linearly mapped through the projection layer to obtain aligned visual features.
[0063] In one feasible implementation, to align the visual features with the word vector feature space of the large model, the visual features output by ViT need to be linearly mapped through a projection layer to obtain the aligned visual features. :
[0064] (2)
[0065] in , It is the size of the text feature dimension.
[0066] S213. Based on the large language model, the defect description sample text is segmented to obtain a sequence token. After passing through the embedding layer, text embedding features are generated based on the sequence token.
[0067] In one feasible implementation, the text prompt corresponds to the input. , This refers to the length of the input text. After the original text is segmented into words by the large language model to form sequence tokens, the embedding features of the text are first formed through the embedding layer. :
[0068] (3)
[0069] in Embedding() represents the embedding operation.
[0070] S214. Concatenate the aligned visual features with the text embedding features to obtain the fused feature input.
[0071] In one feasible implementation, the projected visual features and text embedding features Concatenating the features yields the feature inputs for subsequent modules of the large language model. :
[0072] (4)
[0073] S215. The fused feature input is propagated forward through the decoder stacked network, passing through a linear layer to predict the next token, resulting in the predicted text response and the predicted visual output.
[0074] In one feasible implementation, the large language model is forward-propagated through a stacked network of decoders consisting of self-attention modules and perceptrons, and finally passed through a linear layer to predict the next token. The final output of the large language model contains the text response. and visual output :
[0075] (5)
[0076] in LLM() represents the processing of large language models.
[0077] S22. Calculate the first loss function based on the predicted text response, the true label, and the cross-entropy loss function.
[0078] In one feasible implementation, during this fine-tuning phase, the large language model predicts the next token using only a masking method; that is, the text input here includes the user's question and the answer the model should learn. During fine-tuning, the masking method allows the large language model to predict the (n+1)th token based on the previous n tokens. The final loss function for generating the response is the predicted loss function. Cross-entropy loss between the actual label and the true label:
[0079] (6)
[0080] S23. Input the track defect sample image into the pre-trained VEA to obtain fine-grained visual features.
[0081] In one feasible implementation, to enable the large language model to more fully process and perceive the features extracted from the CLIP visual encoder, a diffusion denoising decoder is used to reconstruct the fine-grained visual features extracted by the noisy VAE, thereby achieving alignment between the visual features output by the large language model and the fine-grained features of the VAE. The specific process is as follows:
[0082] Extracting the input image using a pre-trained VAE. Fine-grained visual features :
[0083] (7)
[0084] in , It is the dimension of the features extracted by VAE.
[0085] S24. Based on the predicted visual output, fine-grained visual features, and diffusion denoising decoder, denoising features are obtained.
[0086] In one feasible implementation, the diffusion denoising decoder uses a Diffusion Transformer architecture, adding an adaptive normalized AdaLN layer to the encoder module of the Transformer to control the generation of diffusion based on specified "conditions." Here, the "conditions" refer to the visual features output by the large language model. ,pass Fine-grained features of control input Alignment through noise reduction process after noise addition and After t steps of noise addition, the diffusion denoising decoder outputs denoising features. :
[0087] (8)
[0088] It should be noted that the implementation of the denoising decoder can be other denoising models or general diffusion methods, not limited to Diffusion Model and stream matching, and can also be other conventional denoising decoder structures in the prior art.
[0089] S25. Calculate the second loss function based on the fine-grained visual features, the denoising features, and the mean squared error loss function.
[0090] In one feasible implementation, wherein A diffusion-based denoising decoder is used to train a large language model to output visual features. and fine-grained VAE features Alignment, using Mean Square Error (MSE) to compute predicted denoised features and and after adding noise The difference between them:
[0091] (9)
[0092] S26. Calculate the post-training fine-tuning loss function based on the first loss function and the second loss function. Perform post-training fine-tuning based on orbital knowledge based on the post-training fine-tuning loss function, and determine the trained basic multimodal large model as the first basic multimodal large model.
[0093] In one feasible implementation, during this fine-tuning phase, the total training loss function is:
[0094] (10)
[0095] S3. Based on the orbit training samples and the first basic multimodal large model, perform cold start supervised fine-tuning of the first basic multimodal large model based on the orbit multimodal target detection task to obtain the second basic multimodal large model.
[0096] In one feasible implementation, during the cold-start supervised fine-tuning stage of the orbital multimodal target detection task, the vocabulary content of the original multimodal large model needs to be changed. This requires retraining to adapt to the bounding box representation token "<|box_token|>" and to learn the parameters suitable for reinforcement fine-tuning. <think> ... < / think> <answer> ... < / answer> Formatted output, <think> ... < / think> The content in between refers to the reasoning process generated by the large model. <answer> ... < / answer> The content in between represents the direct result of the large model's answer. This stage uses general multimodal cold-start data for training, learning a new vocabulary distribution while maintaining the original capabilities of the large model. It aligns BoxHead and BoxEmbed with the word vector space of the large language model and learns the reasoning process for general visual tasks. This stage still relies on a diffusion decoder and supervised fine-tuning training.
[0097] Previous large-scale multimodal models primarily used plain text as the output format, which resulted in excessively long context when applied to visual object detection tasks. For example, coordinates [176, 106, 232, 160] contain 12 coordinate-related characters and 5 format characters. If previous large-scale multimodal models were used for object detection, predicting a bounding box for an object might require 17 characters. These characters are actually detrimental to other text inputs because they contain a relatively long information interval, requiring a longer token to represent a meaning. Furthermore, the increased sequence length leads to relatively low Softmax results in the context, affecting the attention weights of important semantic tokens and subsequent processing. Moreover, when an image contains many objects to be detected, the context length increases significantly, which is unfavorable for long-context dialogues.
[0098] This invention, referring to previous computer vision techniques for object detection, designs an object detection module adapted to the context to solve the aforementioned problems. It comprises two parts: a bounding box regression layer and a bounding box embedding layer.
[0099] Optionally, S3, based on the orbital training samples and the first basic multimodal large model, performs cold-start supervised fine-tuning of the first basic multimodal large model based on the orbital multimodal target detection task to obtain the second basic multimodal large model, including:
[0100] S31. When the next token predicted by the large language model is detected to be a coordinate, the token is input into the bounding box regression layer as the output feature of the decoder stacked network to obtain the predicted coordinates corresponding to the feature image.
[0101] In one feasible implementation, when the large language model predicts the next token, to reduce the contextual burden on predicted coordinates, the coordinates are represented by "<|box_token|>" and defined in the word segmentation table of the large model. When the large model generates the predicted next token, if this token is the coordinate "<|box_token|>", then the features output by the last decoder module in the large language model are input into the bounding box regression layer to predict the detection coordinates in the format xywh (the horizontal and vertical coordinates of the center point, and the length and width of the bounding box). The mathematical description process is as follows:
[0102] (11)
[0103] in These are output features from the last decoder module in the large language model regarding "<|box_token|>". It is a bounding box regression layer composed of two sensing layers. These are the predicted bounding box coordinates, consisting of four data points: the x-coordinate of the center point, ... and the y-coordinate of the bounding box. , center point ordinate Length of the border and the width of the border .
[0104] If the input text to the large language model contains detection coordinates, the coordinates are extracted, normalized, and replaced with "<|box_token|>" in their original positions. The bounding box coordinates are encoded into bounding box features through the bounding box embedding layer and placed in the large language model along with the text embedding features.
[0105] S32. Calculate the L1 loss function based on the average difference between the predicted coordinates and the true coordinate labels.
[0106] The training process involving bounding box regression uses GIoU and L1 to calculate the loss function.
[0107] (12)
[0108] S33. Calculate the IoU loss function based on the intersection area of the predicted coordinates and the true coordinate labels, and the union area of the predicted coordinates and the true coordinate labels.
[0109] In one feasible implementation, the IoU loss function is calculated according to the following equation (13):
[0110] (13)
[0111] in, This represents the overlap area between the border corresponding to the predicted coordinates and the border corresponding to the actual coordinate labels. This represents the total area of the border corresponding to the predicted coordinates and the border corresponding to the actual coordinate labels.
[0112] S34. Calculate the GIoU loss function based on the IoU loss function, predicted coordinates, and true coordinate labels.
[0113] In one feasible implementation, the GIoU loss function is calculated according to the following equation (14):
[0114] (14)
[0115] Here, Area represents the smallest rectangle size that can cover the borders corresponding to the predicted coordinates and the borders corresponding to the actual coordinate labels.
[0116] S35. Calculate the Box loss function based on the L1 loss function and the GIoU loss function.
[0117] In one feasible implementation, the Box loss function is calculated according to the following equations (15)-(16):
[0118] (15)
[0119] (16)
[0120] S36. Calculate the cold start supervised fine-tuning loss function based on the first loss function, the second loss function, and the Box loss function. Perform cold start supervised fine-tuning based on the orbital multimodal target detection task based on the cold start supervised fine-tuning loss function, and determine the trained model as the second basic multimodal large model.
[0121] In one feasible implementation, the total loss function in supervised fine-tuning of a multimodal large model is:
[0122] (17)
[0123] S4. Based on the track training samples, the second basic multimodal large model and the GRPO algorithm, the second basic multimodal large model is enhanced and fine-tuned based on the track multimodal target detection task, and the trained basic multimodal large model is determined as the track defect detection model.
[0124] In one feasible implementation, considering that while track images are easy to acquire, the corresponding text instructions and fine-grained generated responses are difficult to construct manually or synthesize using a general multimodal large model with low quality, this invention designs a track defect visual reasoning strategy based on large model reinforcement and fine-tuning. By leveraging the domain visual and textual knowledge acquired after previous fine-tuning training, and the visual reasoning training through large model reinforcement and fine-tuning, the multimodal large model can spontaneously learn the language reasoning process of track domain visual tasks, thereby realizing reasoning and prediction for track defect detection tasks.
[0125] Therefore, in the enhancement and fine-tuning stage of the track multimodal target detection task, track image data and track operation and maintenance knowledge data are processed into enhancement and fine-tuning data. <answer> ... < / answer> The format involves spontaneously optimizing the inference process and learning the answer content through reinforcement fine-tuning. The reinforcement fine-tuning algorithm used in this stage is GRPO, which uses the multimodal large model before reinforcement training as the reference policy and the multimodal large model to be updated as the current policy. The policy is updated by comparing the differences between the old and new policies and using the average reward generated by multiple samplings of the new policy.
[0126] During the fine-tuning phase, a large model with formatted reward constraints is used. <think> ... < / think> <answer> ... < / answer>The format input generates content; the IoU value between the predicted bounding box and the ground truth label is used as the bounding box regression reward, which enhances the object detection and visual reasoning process of the large model in terms of language form.
[0127] S5. Input the track image to be detected into the track defect detection model to obtain the corresponding defect detection results.
[0128] In one feasible implementation, during the application process, only the basic multimodal large model consisting of three network modules—visual encoder, projection layer, and large language model—is used to complete the track defect detection task.
[0129] The following is a detailed explanation of the process using examples:
[0130] (a) Given as Figure 3 The image shown is the trajectory image to be input, which is then processed and input into the model. The size is 336x336 pixels:
[0131] (b) Language commands can be set to enable track defect detection for large models. The question asks, "What defects can be observed in the track from this track image? Please provide relevant visual perceptions and output the final predicted type of defect."
[0132] (c) Extracting features from serialized images using a visual encoder .
[0133] (d) The image features are transformed into the word vector space of the large model through the projection layer. .
[0134] (e) Word vectors embedded after language instructions have been segmented. Image features output by the projection layer These features are concatenated to form multimodal features that include images and text, which are then input to a larger model. After forward propagation through the large model, the next token is predicted (this token, after decoding, is a "graph"). The word vectors embedded in this token are then merged into the original word vectors, resulting in a new word vector. The new word vectors are then fed back into the Big Oracle model to predict the next token, until a token indicating the end is predicted or the specified maximum generation length is reached.
[0135] If, during the autoregressive generation process at step t, a "<|box_token|>" representing the existence of a bounding box is predicted, then the feature corresponding to this token is... enter Perform bounding box prediction and output bounding box coordinates. Input Word vectors with coordinates obtained , and the total word vectors at step t splicing to obtain the first Step input word vectors of large model Continue with autoregressive generation.
[0136] (f) Following the above steps of multi-step autoregression until the termination state, the large model outputs the final generated content about the text. and image features Text generation The decoded text reads: "The surface of the rail in the center of the image appears relatively smooth, but there may be some wear marks, presumably due to external damage to the rail head surface." The "<|box_token|>" part will be... The predicted detection coordinates "[1046,1807,1116,1964]", "[1054,1312,1099,1566]", and "[1053,1602,1097,1723]" are replaced and processed into the end-user visible language output: "The surface of the rail in the center of the image appears relatively smooth, but there may be some wear marks, presumably <|box_start|>rail head surface damage [[1046,1807,1116,1964],[1054,1312,1099,1566],[1053,1602,1097,1723]]<|box_end|>". The content between "<|box_start|>" and "<|box_end|>" is the target and location detected by the large model.
[0137] This invention proposes a track defect detection method based on a multimodal large model. By designing a diffusion denoising decoder, the method enhances the large model's visual perception capability of track images during training. A corresponding multi-stage training strategy involving fusion enhancement and fine-tuning is also designed. This allows for the use of text in contextual scenarios to perform visual task reasoning for track defects, resulting in better task interpretability. Furthermore, because a large model is used to implement tasks in the rail transit domain and corresponding domain data is constructed for training, the large model exhibits better generalization in vertical domain knowledge, enabling it to perform defect detection tasks across a wider range of categories. It can also interact with users in a human-readable format to complete domain instruction generation tasks. This invention designs a target detection module adapted to contextual scenarios, effectively reducing the context length during user interaction and lowering the computational cost of the large model. In addition, the multimodal large model designed in this invention undergoes sufficient multimodal target detection training, including track defect detection tasks. It can perceive track geometry and align with track-related text corpora, learn the potential relationship between visual features of track defect locations and linguistic inputs representing defects, and possess the ability to detect known defect types through specific linguistic instructions and the potential for open-vocabulary detection of unseen defect types.
[0138] Figure 4 This is a block diagram of a track defect detection device based on a multimodal large model, provided by an embodiment of the present invention. This device is used in a track defect detection method based on a multimodal large model. (Refer to...) Figure 3 The device includes an acquisition unit 410, a post-training fine-tuning unit 420, a cold-start supervised fine-tuning unit 430, a reinforcement fine-tuning unit 440, and a detection unit 450. Wherein:
[0139] The acquisition unit 410 is used to acquire orbit training samples and a model to be trained. The model to be trained includes a basic multimodal large model to be trained, a diffusion denoising decoder, and a large model target detection module.
[0140] The post-training fine-tuning unit 420 is used to perform post-training fine-tuning based on orbital knowledge on the basic multimodal large model to be trained, based on orbital training samples, the basic multimodal large model to be trained, and the diffusion denoising decoder, to obtain the first basic multimodal large model.
[0141] The cold start supervised fine-tuning unit 430 is used to perform cold start supervised fine-tuning of the first basic multimodal large model based on the orbit training samples and the first basic multimodal large model, so as to obtain the second basic multimodal large model.
[0142] The enhancement and fine-tuning unit 440 is used to enhance and fine-tune the second basic multimodal large model based on the track training samples, the second basic multimodal large model and the GRPO algorithm, and determine the trained basic multimodal large model as the track defect detection model.
[0143] The detection unit 450 is used to input the track image to be detected into the track defect detection model to obtain the corresponding defect detection result.
[0144] This invention proposes a track defect detection device based on a multimodal large model. By designing a diffusion denoising decoder, the device enhances the large model's visual perception capability of track images during training. A corresponding multi-stage training strategy of fusion enhancement and fine-tuning is also designed. This allows for the use of text in contextual scenarios to perform visual task reasoning for track defects, resulting in better task interpretability. Furthermore, because a large model is used to implement tasks in the rail transit domain and corresponding domain data is constructed for training, the large model exhibits better generalization in vertical domain knowledge, enabling it to perform defect detection tasks with a wider range of categories. It can also interact with users in a human-readable format to complete domain instruction generation tasks. This invention designs a target detection module adapted to contextual scenarios, effectively reducing the context length during user interaction and lowering the computational cost of the large model. In addition, the multimodal large model designed in this invention undergoes sufficient multimodal target detection training, including track defect detection tasks. It can perceive track geometry and align with track-related text corpora, learn the potential relationship between the visual features of track defect locations and the language input representing defects, and possess the ability to detect known defect types through specific language instructions and the potential for open-vocabulary detection of unseen defect types.
[0145] Figure 5 This is a schematic diagram of the structure of a track defect detection device based on a multimodal large model provided in an embodiment of the present invention, as shown below. Figure 5 As shown, the track defect detection equipment based on a multimodal large model can include the above-mentioned... Figure 4 The illustrated track defect detection device is based on a multimodal large model. Optionally, the track defect detection device 510 based on the multimodal large model may include a first processor 2001.
[0146] Optionally, the track defect detection device 510 based on a multimodal large model may also include a memory 2002 and a transceiver 2003.
[0147] The first processor 2001, memory 2002, and transceiver 2003 can be connected via a communication bus.
[0148] The following is combined Figure 5A detailed introduction to each component of the track defect detection device 510 based on a multimodal large model is provided below:
[0149] The first processor 2001 is the control center of the multimodal large-scale model-based track defect detection device 510. It can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 can be one or more central processing units (CPUs), application-specific integrated circuits (ASICs), or one or more integrated circuits configured to implement embodiments of the present invention, such as one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs).
[0150] Optionally, the first processor 2001 can perform various functions of the track defect detection device 510 based on a multimodal large model by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.
[0151] In a specific implementation, as one example, the first processor 2001 may include one or more CPUs, for example... Figure 5 CPU0 and CPU1 are shown in the diagram.
[0152] In a specific implementation, as one example, the track defect detection device 510 based on a multimodal large model may also include multiple processors, for example... Figure 5 The first processor 2001 and the second processor 2004 are shown in the diagram. Each of these processors can be a single-core processor or a multi-core processor. Here, a processor can refer to one or more devices, circuits, and / or processing cores used to process data (such as computer program instructions).
[0153] The memory 2002 is used to store the software program that executes the present invention, and is controlled by the first processor 2001 to execute it. The specific implementation method can be referred to the above method embodiment, and will not be repeated here.
[0154] Optionally, the memory 2002 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. The memory 2002 may be integrated with the first processor 2001 or may exist independently, and may be connected via the interface circuit of the multimodal large model-based track defect detection device 510. Figure 5 (Not shown in the image) is coupled to the first processor 2001, and this embodiment of the invention does not specifically limit this.
[0155] The transceiver 2003 is used to communicate with network devices or with terminal devices.
[0156] Alternatively, transceiver 2003 may include a receiver and a transmitter. Figure 5 (Not shown separately). The receiver is used to implement the receiving function, and the transmitter is used to implement the transmitting function.
[0157] Optionally, the transceiver 2003 can be integrated with the first processor 2001, or it can exist independently, and it can be connected to the interface circuit of the track defect detection device 510 based on a multimodal large model. Figure 5 (Not shown in the image) is coupled to the first processor 2001, and this embodiment of the invention does not specifically limit this.
[0158] It should be noted that, Figure 5 The structure of the multimodal large model-based track defect detection device 510 shown in the figure does not constitute a limitation on the router. Actual multimodal large model-based track defect detection devices may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0159] Furthermore, the technical effects of the track defect detection device 510 based on the multimodal large model can be referred to the technical effects of the track defect detection method based on the multimodal large model described in the above method embodiments, and will not be repeated here.
[0160] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or it may be any conventional processor, etc.
[0161] It should also be understood that the memory in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0162] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0163] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.
[0164] In this invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single item or multiple items.
[0165] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0166] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0167] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0168] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0169] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0170] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0171] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0172] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for detecting track defects based on a multimodal large model, characterized in that, The method includes: S1. Obtain orbital training samples and a model to be trained. The model to be trained includes a basic multimodal large model to be trained, a diffusion denoising decoder, and a large model target detection module. S2. Based on the orbital training samples, the basic multimodal large model to be trained, and the diffusion denoising decoder, the basic multimodal large model to be trained is subjected to post-training fine-tuning based on orbital knowledge to obtain the first basic multimodal large model. The basic multimodal large model includes a CLIP visual encoder, a projection layer, a large language model, and a decoder stacked network. S2 includes: S21, inputting the track defect sample image into the basic multimodal large model to be trained to obtain the predicted text response and the predicted visual output; S211. Input the track defect sample image into the CLIP visual encoder ViT, divide the track defect sample image into multiple pixel blocks, and then propagate forward through a multi-layer encoder module composed of a self-attention layer and a perceptron to extract the initial visual features. S212. The initial visual features are linearly mapped through the projection layer to obtain aligned visual features; S213. Based on the large language model, the defect description sample text is segmented to obtain a sequence token. After passing through the embedding layer, text embedding features are generated based on the sequence token. S214. Concatenate the aligned visual features with the text embedding features to obtain the fused feature input; S215. The fused feature input is propagated forward through the decoder stacked network, and then through a linear layer to predict the next token, resulting in the predicted text response and the predicted visual output. S3. Based on the orbit training samples and the first basic multimodal large model, perform cold start supervised fine-tuning of the first basic multimodal large model based on the orbit multimodal target detection task to obtain the second basic multimodal large model; Specifically, S3, based on orbital training samples and the first basic multimodal large model, performs cold-start supervised fine-tuning of the first basic multimodal large model based on orbital multimodal target detection task to obtain the second basic multimodal large model, including: S31. When it is detected that the next token predicted by the large language model is a coordinate, the output feature of the token in the decoder stacking network is input into the bounding box regression layer to obtain the predicted coordinates corresponding to the feature image. S32. Calculate the L1 loss function based on the average difference between the predicted coordinates and the true coordinate labels; S33. Calculate the IoU loss function based on the intersection area of the predicted coordinates and the true coordinate labels and the union area of the predicted coordinates and the true coordinate labels. S34. Calculate the GIoU loss function based on the IoU loss function, predicted coordinates, and true coordinate labels; S35. Calculate the Box loss function based on the L1 loss function and the GIoU loss function; S36. Calculate the cold start supervised fine-tuning loss function based on the first loss function, the second loss function, and the Box loss function. Perform cold start supervised fine-tuning based on the cold start supervised fine-tuning loss function for the orbital multimodal target detection task. Determine the trained model as the second basic multimodal large model. S4. Based on the track training samples, the second basic multimodal large model and the GRPO algorithm, the second basic multimodal large model is enhanced and fine-tuned based on the track multimodal target detection task, and the trained basic multimodal large model is determined as the track defect detection model. S5. Input the track image to be detected into the track defect detection model to obtain the corresponding defect detection results.
2. The track defect detection method based on a multimodal large model according to claim 1, characterized in that, The track training samples include track defect sample images and corresponding defect description sample text; The diffusion denoising decoder is a Diffusion Transformer structure.
3. The track defect detection method based on a multimodal large model according to claim 2, characterized in that, The S2 method, based on orbital training samples, the basic multimodal large model to be trained, and a diffusion denoising decoder, performs post-training fine-tuning based on orbital knowledge on the basic multimodal large model to be trained, resulting in a first basic multimodal large model, including: S21. Input the track defect sample image into the basic multimodal large model to be trained to obtain the predicted text response and the predicted visual output; S22. Calculate the first loss function based on the predicted text response, the true label, and the cross-entropy loss function; S23. Input the track defect sample image into the pre-trained VEA to obtain fine-grained visual features; S24. Based on the predicted visual output, fine-grained visual features, and diffusion denoising decoder, the denoising features are obtained; S25. Calculate the second loss function based on the fine-grained visual features, the denoising features, and the mean squared error loss function; S26. Calculate the post-training fine-tuning loss function based on the first loss function and the second loss function. Perform post-training fine-tuning based on orbital knowledge based on the post-training fine-tuning loss function, and determine the trained basic multimodal large model as the first basic multimodal large model.
4. A track defect detection device based on a multimodal large model, wherein the track defect detection device based on the multimodal large model is used to implement the track defect detection method based on a multimodal large model as described in any one of claims 1-3, characterized in that, The device includes: The acquisition unit is used to acquire orbital training samples and a model to be trained. The model to be trained includes a basic multimodal large model to be trained, a diffusion denoising decoder, and a large model target detection module. The post-training fine-tuning unit is used to perform post-training fine-tuning based on orbital knowledge on the basic multimodal large model to be trained, based on orbital training samples, the basic multimodal large model to be trained, and the diffusion denoising decoder, to obtain the first basic multimodal large model. The cold start supervised fine-tuning unit is used to perform cold start supervised fine-tuning of the first basic multimodal large model based on the orbit training samples and the first basic multimodal large model, so as to obtain the second basic multimodal large model. The enhancement and fine-tuning unit is used to enhance and fine-tune the second basic multimodal large model based on the track training samples, the second basic multimodal large model and the GRPO algorithm, and to determine the trained basic multimodal large model as the track defect detection model. The detection unit is used to input the track image to be detected into the track defect detection model to obtain the corresponding defect detection results.
5. The track defect detection device based on a multimodal large model according to claim 4, characterized in that, The track training samples include track defect sample images and corresponding defect description sample text; The diffusion denoising decoder is a Diffusion Transformer structure.
6. The track defect detection device based on a multimodal large model according to claim 5, characterized in that, The post-training fine-tuning unit is used for: S21. Input the track defect sample image into the basic multimodal large model to be trained to obtain the predicted text response and the predicted visual output; S22. Calculate the first loss function based on the predicted text response, the true label, and the cross-entropy loss function; S23. Input the track defect sample image into the pre-trained VEA to obtain fine-grained visual features; S24. Based on the predicted visual output, fine-grained visual features, and diffusion denoising decoder, the denoising features are obtained; S25. Calculate the second loss function based on the fine-grained visual features, the denoising features, and the mean squared error loss function; S26. Calculate the post-training fine-tuning loss function based on the first loss function and the second loss function. Perform post-training fine-tuning based on orbital knowledge based on the post-training fine-tuning loss function, and determine the trained basic multimodal large model as the first basic multimodal large model.
7. A track defect detection device based on a multimodal large model, characterized in that, The track defect detection device based on a multimodal large model includes: processor; A memory storing computer-readable instructions that, when executed by the processor, implement the method as described in any one of claims 1 to 3.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium contains program code that can be invoked by a processor to execute the method as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Multi-modal track defect segmentation method based on large model
CN118887236A
Model training and information replying method and device, storage medium and program product
CN120725157A