Railway ballastless track slab crack multi-mode small sample detection method, medium and equipment

By constructing a multimodal small-sample detection method for cracks in railway ballastless track slabs, a visual model is trained using a very small amount of data and combined with a multimodal large model, achieving high-precision crack detection and automated decision-making. Detailed detection results and maintenance suggestions are output, solving the problems of low detection accuracy and insufficient decision-making in existing technologies.

CN122156128APending Publication Date: 2026-06-05CENT SOUTH UNIV +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CENT SOUTH UNIV
Filing Date
2026-03-03
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Existing computer vision-based track slab crack detection methods suffer from a lack of data samples and an imbalanced ratio, resulting in low detection accuracy, poor generalization ability, and a lack of automated maintenance decision support, failing to provide detailed semantic descriptions and maintenance suggestions.

Method used

A multimodal small-sample detection method for cracks in railway ballastless track slabs is constructed. By acquiring a very small amount of track slab image data, crack coordinate information is labeled and a YOLO dataset is constructed. The backbone network parameters are frozen, and the visual model is fine-tuned. Combined with low-rank adaptation technology, a multimodal large model is trained to output image scene descriptions, crack coordinate information, and maintenance decision suggestions.

Benefits of technology

Achieving high-precision detection under small sample conditions, outputting detailed crack geometry parameters and maintenance suggestions, improves the intelligence level of detection, solves the problems of data scarcity and insufficient automated decision support, and improves detection accuracy and objectivity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122156128A_ABST
    Figure CN122156128A_ABST
Patent Text Reader

Abstract

The present application relates to the field of rail transit technology, and in particular to a railway ballastless track slab crack multi-modal small sample detection method, medium and equipment, the method comprising: constructing a track slab crack visual dataset; obtaining a track slab crack detection visual model; constructing a multi-modal instruction fine-tuning dataset in the form of text and image pairs; constructing a multi-modal large model based on Qwen3-VL and a track slab crack detection multi-modal large model; composing a final track slab crack detection multi-modal small sample model; and outputting a text description containing image scene description, crack coordinate information, severity estimation and maintenance decision suggestions. The present application solves the problems of sample scarcity and unbalanced positive and negative sample ratio caused by privacy and long-tail distribution of railway inspection data; at the same time, the hallucination of the multi-modal large model is effectively suppressed by using coordinate prior, realizing the full-process automation from crack accurate positioning to intelligent maintenance decision support.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of rail transit technology, and in particular to a method, medium, and equipment for multimodal small-sample detection of cracks in railway ballastless track slabs. Background Technology

[0002] High-speed railway track slabs are a crucial component of ballastless track systems, supporting core components such as rails and fastening systems. They provide robust support for the track, fixing its position to ensure the safe and stable operation of high-speed trains. They also effectively transfer the weight and dynamic loads of the train to the underlying foundation structure, reducing impact wear. Furthermore, track slabs help maintain the geometric accuracy of the track, reducing deformation and wear, ensuring the safety and comfort of high-speed train operation. They also reduce track maintenance requirements, extend track lifespan, lower long-term operating costs, and improve track availability.

[0003] During long-term service, track slabs inevitably develop cracks due to multiple factors, including cyclic high-frequency loads from trains, alternating ambient temperatures, rain erosion, and natural aging of components. If these cracks are not detected and repaired promptly, they will continue to worsen, affecting the stability of the components supporting them, and larger cracks can even threaten train operation safety. Therefore, efficient and accurate detection and assessment of track slab crack conditions are of great significance for ensuring train safety and developing maintenance plans.

[0004] For a long time, track slab condition assessment has mainly relied on manual visual inspection. However, this method is labor-intensive, inefficient, and the results are easily affected by the subjective experience and fatigue of the inspectors, leading to missed detections or misjudgments. This makes it difficult to meet the maintenance and support requirements of high-speed railways with ultra-long mileage, all-weather, high-frequency operation, and millimeter-level precision. Against this backdrop, intelligent sensing methods integrating cutting-edge technologies such as computer vision, deep learning, and pattern recognition have emerged, providing a new solution for track defect detection. This technology uses high-performance industrial imaging equipment to acquire images of the track surface, and then applies advanced image processing and feature analysis algorithms to achieve non-contact automatic identification and quantitative assessment of minute defects such as cracks, significantly improving both the timeliness and objectivity of detection.

[0005] Although computer vision-based track slab crack detection has been widely applied, the following problems still exist: (1) Data samples are scarce and disproportionate. Due to the special nature of railway inspection data, there is still a lack of publicly available datasets of track slab cracks. Of the 1,000 images in the existing dataset, fewer than 40 contain crack data, which severely limits the upper limit of the detection model's accuracy. Although methods such as style transfer and digital twins are currently used to generate track slab crack data, the performance of models trained on synthetic images remains to be tested. Existing computer vision-based detection methods usually require large-scale labeled data for training, and are prone to overfitting with small sample data, resulting in low detection accuracy and poor generalization ability.

[0006] (2) Lack of automated maintenance decision support. Existing computer vision methods (such as Faster R-CNN, YOLO series, etc.) can usually only provide the coordinates and category of cracks, but cannot semantically describe the width, length, trend and severity of cracks, nor can they provide specific maintenance suggestions. Maintenance personnel still need to manually review the images based on the detection results to formulate plans, which fails to achieve true intelligent and automated maintenance.

[0007] In summary, the scarcity of high-quality datasets and the lack of automated decision support remain the biggest obstacles to achieving automated detection of track slab cracks. The railway industry urgently needs a detection method that can achieve high-precision detection under small sample conditions and provide detailed natural language evaluation and decision suggestions by combining visual coordinate information. Summary of the Invention

[0008] The main objective of this invention is to provide a method, medium, and equipment for multimodal small-sample detection of cracks in railway ballastless track slabs, aiming to solve the technical problems of insufficient and disproportionate data samples and lack of automated maintenance decision support in existing computer vision-based track slab crack detection methods.

[0009] To achieve the above objectives, this invention proposes a multimodal small-sample detection method for cracks in railway ballastless track slabs, comprising the following steps: S1. Obtain track slab image data; S2. For the image data obtained in S1, mark the crack coordinate information, organize the image data and crack coordinate information obtained in S1 into YOLO dataset format, and construct a visual dataset of track slab cracks. S3. For the image data obtained in S1, randomly select images and perform expert-level descriptive annotations, including scene descriptions, crack geometric features, and maintenance suggestions. S4. Construct a visual model based on YOLO v11x, freeze the parameters of the model backbone network, and fine-tune the training using the track slab crack visual dataset from S2 to obtain a track slab crack detection visual model. S5. Combining the crack coordinate information obtained in S2 and the description annotations obtained in S3, construct a multimodal instruction fine-tuning dataset in the form of image-text pairs; S6. Construct a multimodal large model based on Qwen3-VL, use the multimodal instruction fine-tuning dataset obtained in S5, and use low-rank adaptation technology for fine-tuning to obtain a multimodal large model for track slab crack detection. S7. Combine the visual model for track slab crack detection obtained in S4 and the large multimodal model for track slab crack detection obtained in S6 to form the final small sample model for track slab crack detection. S8. The multimodal small sample model for track slab crack detection outputs a text description that includes image scene description, crack coordinate information, severity estimation, and maintenance decision suggestions.

[0010] The method for detecting multimodal small-sample cracks in railway ballastless track slabs of the present invention is further improved in that, when marking the crack coordinate information, an open-source annotation tool is used to load the image and an annotation specification is formulated; the annotation categories include four types: rail, fastening system, track slab and crack; the annotation method adopts axis-aligned bounding boxes.

[0011] The method for detecting multimodal small-sample cracks in railway ballastless track slabs of the present invention is further improved by including the following steps when constructing the visual dataset of track slab cracks: The annotation results are exported in YOLO format for training the visual model; in YOLO format, each object is represented by a quadruple: ;in, and The x and y coordinates of the target center point are: and For the width and height of the target, all values ​​are normalized to the range [0,1]. The annotation results are also exported in COCO format for subsequent multimodal cue word construction; in COCO format, the target coordinates are represented as... ,in The coordinates of the top left corner of the object. The coordinates are the bottom right corner of the object.

[0012] The method for detecting multimodal small-sample cracks in railway ballastless track slabs of the present invention is further improved in that S4 specifically includes the following steps: Load the pre-trained weights of YOLO v11x, freeze the weights of the backbone and neck network of YOLO v11x, and only update the gradient of the detection head parameters to adapt to the features of small sample data. Training is performed using a combined loss function that includes both classification and regression loss; loss function The calculation formula is as follows: ; in: CIoU loss is used for bounding box regression. for The balance coefficient, Binary cross-entropy loss is used for category classification. for The balance coefficient, For the distribution focus loss, for The balance coefficient.

[0013] The method for detecting multimodal small-sample cracks in railway ballastless track slabs of the present invention is further improved in that S5 specifically includes the following steps: Constructing user prompts involves embedding the COCO-formatted coordinates obtained from S2 into the prompt template; The construction assistant's answer specifically refers to the scene description, crack geometry parameters (width, length), severity, and maintenance suggestions marked in S3. Organize the image path, user prompts, and assistant responses into JSON-formatted dialogue data.

[0014] The method for detecting multimodal small-sample cracks in railway ballastless track slabs of the present invention is further improved in that S6 specifically includes the following steps: Freeze all raw parameters of the visual encoder and language model of the Qwen3-VL model. ; Injecting a low-rank matrix in a linear layer bypass and Forward propagation is calculated as ,in, For input features, For incremental updates of the LoRA branch, The sum of the original parameters and the incremental updates of the LoRA branch will be used as the final parameters of this linear layer; The autoregressive language modeling loss function is used for optimization, and the calculation formula is as follows: ; in, For the predicted text terms, For the input image, The prompt word contains coordinate information. T This indicates the total length of the generated text sequence. To calculate the total value of the loss between the predicted text and the real text, This represents the index number of the word in the text sequence. For conditional probability, This indicates a prompt word given an input image and coordinates. and all previously generated text terms Under the condition that the model predicts the current number of... Each lexical element is a real label. The probability value; A large multimodal model for track slab crack detection was trained using the multimodal instruction fine-tuning dataset obtained from S5.

[0015] In addition, the present invention provides a readable storage medium storing a computer program adapted to be loaded by a processor and executed as described above for the multimodal small sample detection method for cracks in railway ballastless track slabs.

[0016] In addition, the present invention also provides a computer device, the computer device including a memory and a processor, the memory storing a computer program, and when the computer program is executed by the processor, running the multimodal small sample detection method for cracks in railway ballastless track slabs as described above.

[0017] The technical solution of the present invention has the following beneficial effects: This invention presents a multimodal small-sample detection method for cracks in railway ballastless track slabs. It constructs a dataset based on a very limited number of inspection images from real-world scenarios and fine-tunes the visual model using a transfer learning strategy with a frozen backbone network. This achieves high-precision localization of track slab cracks and key components even under conditions of extreme data scarcity. Simultaneously, this invention utilizes the coordinate information output by the visual model to construct location guidance prompts and fine-tunes the multimodal large model through low-rank adaptation technology, effectively suppressing the "illusion" phenomenon of general large models in industrial defect detection. This invention effectively solves the problems of data privacy and the scarcity of high-quality crack samples with long-tail distribution in the railway field, as well as the inability of traditional computer vision methods to provide semantic decision support. Furthermore, the text description of the detection results generated by this invention represents a leap from simple "defect perception" to "expert-level cognition" that includes geometric parameter evaluation and maintenance suggestions, significantly improving the intelligence level of railway track maintenance. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.

[0019] Figure 1 This is a flowchart of the multimodal small sample detection method for cracks in railway ballastless track slabs according to the present invention; Figure 2This is a schematic diagram of the training process of the multimodal small sample detection method for cracks in railway ballastless track slabs according to the present invention; Figure 3 This is an example diagram of the multimodal command fine-tuning dataset for the multimodal small sample detection method for cracks in railway ballastless track slabs of the present invention; Figure 4 This is a schematic diagram illustrating the reasoning process of the multimodal small-sample detection method for cracks in railway ballastless track slabs according to the present invention. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0021] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present invention are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indication will also change accordingly.

[0022] Furthermore, in this invention, descriptions involving "first," "second," etc., are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0023] In this invention, unless otherwise explicitly specified and limited, the terms "connection," "fixed," etc., should be interpreted broadly. For example, "fixed" can mean a fixed connection, a detachable connection, or an integral part; it can mean a mechanical connection or an electrical connection; it can mean a direct connection or an indirect connection through an intermediate medium; it can mean the internal communication of two components or the interaction between two components, unless otherwise explicitly limited. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0024] Furthermore, the technical solutions of the various embodiments of the present invention can be combined with each other, but only if they are feasible for those skilled in the art. If the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention.

[0025] likeFigures 1-4 As shown, this invention aims to construct a high-precision detection and evaluation model using a very small number of samples, and proposes a multi-modal small-sample detection method for cracks in railway ballastless track slabs, including the following steps: S1. Obtain real-world track slab image data; including positive sample images with cracks and negative sample images without cracks, to build a small sample original dataset (≤50 images). Specifically, this includes: In this embodiment, images of the track slab surface are acquired using an industrial line-scan camera mounted on a high-speed railway integrated inspection vehicle. To verify the small-sample learning capability of this method, only N=50 track slab images containing different lighting conditions (such as inside a tunnel, under strong open light, or with shadow coverage) and different background textures are collected and selected to establish a small-sample original dataset. Among them, there are 5 positive sample images containing cracks and 25 negative sample images without cracks.

[0026] S2. For the image data obtained in S1, mark the crack coordinate information, organize the image data and crack coordinate information obtained in S1 into YOLO dataset format, and construct a visual dataset of track slab cracks. For the 50 images acquired by S1, detailed manual annotation was performed. The specific steps are as follows: The image was loaded using the open-source annotation tool (Label Studio), and annotation specifications were defined. The annotation categories included four types: rails, fastening systems, track slabs, and cracks. The annotation method used was axis-aligned bounding boxes (AABB).

[0027] The annotation results are exported in YOLO format for training the visual model; in YOLO format, each object is represented by a quadruple: ;in, and The x and y coordinates of the target center point are: and For the width and height of the target, all values ​​are normalized to the range [0,1]. The annotation results are also exported in COCO format for subsequent multimodal cue word construction; in COCO format, the target coordinates are represented as... ,in The coordinates of the top left corner of the object. The coordinates are the bottom right corner of the object.

[0028] S3. Based on the image data obtained in S1, randomly select a portion of images (≤5 images) and perform expert-level detailed annotation, including scene description, crack geometric features, and maintenance suggestions; specifically including: To endow the model with expert-level cognitive capabilities, a very small number (five in this example) of typical crack images are randomly selected from the images acquired by S1. This section involves textual descriptions and annotations of these five images by railway engineering experts with over five years of experience. The descriptions are structured into three parts: Scenario description: For example, "The image shows a straight section of ballastless track with uniform lighting." Crack geometry features: estimated based on image pixels, for example, "the crack is located on the right side of the track slab, is longitudinally distributed, has a length of about 120 pixels, and an average width of about 3 pixels." Maintenance recommendations: According to the "Rules for Maintenance of High-Speed ​​Railway Lines" (Tiegongdian

[2023] No. 106), for example, "a crack width of 3mm is classified as Class I damage, and it is recommended to perform surface sealing treatment and continue to monitor it." S4. Construct a visual model based on YOLO v11x, freeze the backbone network parameters, and fine-tune the training using the track slab crack visual dataset from S2 to obtain a track slab crack detection visual model; specifically including: (1) Model construction and initialization: YOLO v11x was selected as the basic detection network, and its pre-trained weights on the COCO general dataset were loaded.

[0029] (2) Freezing strategy: To prevent overfitting on small samples, all weight parameters of the model's backbone and neck networks are frozen to maintain its general feature extraction capability. Gradient updates are performed only on the parameters of the head.

[0030] (3) Model training: Fine-tuning was performed using the visual dataset of track slab cracks constructed by S2. The training environment was based on the PyTorch framework (an open-source deep learning framework), using an NVIDIA RTX 4090 D (24G), with a batch size of 8, an initial learning rate of 0.0001, and the AdamW optimizer. The training was iteratively trained for 50 epochs (the model had traversed and learned all the training samples once).

[0031] (4) Loss function: A combined loss function is used for optimization during training. The calculation formula is as follows: ; in: CIoU loss is used for bounding box regression to improve localization accuracy. for The balance coefficient, Binary cross-entropy loss is used for category classification. for The balance coefficient, This is the distributed focus loss, used to optimize the distribution of bounding boxes; for The balance coefficient, , and In this embodiment, the values ​​are set to 2.5, 0.5, and 1.0, respectively.

[0032] After training, a visual model for detecting cracks in the track slab is obtained.

[0033] S5. Combining the crack coordinate information obtained in S2 and the description annotations obtained in S3, construct a multimodal instruction fine-tuning dataset in the form of image-text pairs; Construct user prompts: Develop a unified prompt template and embed the COCO format coordinates obtained from S2 into the prompt template; for example: "You are a track inspection expert. The following component coordinates have been detected in the image: rail [1996,0,2117,4096], fastening system [1257,761,1966,1529], [2195,739,2989,1514], [1239, 2452, 1967,3224], [2187, 2434, 2999, 3202], track slab [0, 0, 3133, 4096], crack [831, 63, 907,600]. Please analyze this image in detail." The build assistant replied: Use the scene description, crack geometry parameters (width, length), severity, and maintenance suggestions marked in S3 as the answer; Data formatting: Organize image paths, user prompts, and assistant responses into JSON format dialogue data as required by Qwen3-VL, forming a multimodal instruction fine-tuning dataset in the form of image-text pairs. Figure 3 This is an example of a dataset for fine-tuning multimodal instructions.

[0034] S6. Construct a large-scale multimodal model based on Qwen3-VL, fine-tune the dataset obtained in S5 using the multimodal command, and fine-tune it using Low-Rank Adaptation (LoRA) technology to obtain a large-scale multimodal model for track slab crack detection; specifically including: Model selection: The Qwen3-VL multimodal large model was selected as the base model; Fine-tuning technique selection: Low-rank adaptation (LoRA) was employed for efficient parameter fine-tuning. All original parameters of the visual encoder and language model of the Qwen3-VL model were frozen. Low-rank matrices are only injected as a bypass into the Query and Value projection matrices of the Attention layer of the language model. and Forward propagation is calculated as ; in, For input features, For incremental updates of the LoRA branch, The sum of the original parameters and the incremental updates of the LoRA branch will be used as the final parameters of this linear layer, and will be initialized during this process. Initialize using a Gaussian distribution. Initialize to 0; The autoregressive language modeling loss function is used for optimization, and the calculation formula is as follows: ; in, For predicted text tokens. For the input image, The prompt word contains coordinate information. T This represents the total length of the generated text sequence (i.e., the total number of tokens). To calculate the total value of the loss between the predicted text and the real text, The index number of the word in the text sequence (with a value from 1 to 1). T ), For conditional probability, This indicates a prompt word given an input image and coordinates. and all previously generated text terms Under the condition that the model predicts the current number of... Each lexical element is a real label. The probability value; A large multimodal model for track slab crack detection was trained using the multimodal instruction fine-tuning dataset obtained from S5.

[0035] S7. Combine the visual model for track slab crack detection obtained in S4 and the large multimodal model for track slab crack detection obtained in S6 to form the final small sample multimodal model for track slab crack detection; specifically including: The visual model for track slab crack detection obtained in S4 and the large multimodal model for track slab crack detection obtained in S6 are concatenated to form the final small sample multimodal model for track slab crack detection. For the image to be detected, the visual model for track slab crack detection first outputs the coordinates of the railway components and cracks. All coordinates are uniformly filled into the prompt word template, and finally input into the large multimodal model for track slab crack detection.

[0036] S8. The multimodal small-sample model for track slab crack detection outputs a textual description including image scene description, crack coordinate information, severity estimation, and maintenance decision suggestions. Specifically, it includes: A1. Image Input: Input an image of the track slab to be detected; A2. Visual Inspection: Input the image into the track slab crack detection visual model. The model quickly scans the entire image and outputs the category and coordinate information of the crack and surrounding components such as rails and rail supports; A3. Coordinate guidance: Fill the coordinate information output by A2 into the prompt word template described in S5 to construct prompt words containing precise location priors; A4. Multimodal Analysis: The original image and the constructed prompts are input into the multimodal large-scale model for track slab crack detection. Guided explicitly by coordinates, the large-scale model focuses on the crack region, combining visual features of the image; A5. Output Results: The model generates a text description through autoregression. This text not only includes the conclusion that "cracks were found," but also describes the width (e.g., "minor" or "wide"), length, and trend of the cracks in natural language. It also provides a severity estimate (e.g., "Level II damage") based on pre-trained expert knowledge and specific maintenance recommendations (e.g., "low-pressure glue injection repair recommended").

[0037] In addition, the present invention provides a readable storage medium storing a computer program adapted to be loaded by a processor and executed as described above for the multimodal small sample detection method for cracks in railway ballastless track slabs.

[0038] In addition, the present invention also provides a computer device, the computer device including a memory and a processor, the memory storing a computer program, and when the computer program is executed by the processor, running the multimodal small sample detection method for cracks in railway ballastless track slabs as described above.

[0039] This embodiment effectively solves the problems of low detection accuracy and poor generalization ability caused by the scarcity of samples in traditional computer vision methods by using the above method with only 50 training images. At the same time, it overcomes the defect of inaccurate localization of general multimodal large models in industrial scenarios and realizes full automation of detection and decision-making processes.

[0040] As shown in Table 1, compared to mainstream detection methods such as Grounding DINO, YOLO v11x, DETR, and Qwen3-VL, the method of this invention ranks first in all four detection performance evaluation metrics: Precision, Recall, F1 score, and mAP@50, demonstrating the effectiveness and advantages of this method. Specifically, compared to computer vision models (such as Grounding DINO, YOLO v11x, and DETR), the method of this invention not only has superior detection performance but also allows for interaction and output of detection results in natural language, making it easier for non-experts to use. Compared to multimodal large models (such as Qwen3-VL), the method of this invention also possesses natural language interaction and output capabilities, while exhibiting superior detection performance. This invention not only improves detection performance but also simplifies interactive operations, advancing the intelligent development of track slab crack detection.

[0041] Table 1. Performance Comparison of Other Mainstream Detection Methods of the Present Invention

[0042] The above description is only a preferred embodiment of the present invention and does not limit the scope of the present invention. All equivalent structural transformations made under the inventive concept of the present invention using the contents of the present invention specification and drawings, or direct / indirect applications in other related technical fields, are included within the protection scope of the present invention.

Claims

1. A method for detecting cracks in railway ballastless track slabs using a multimodal, small-sample approach, characterized in that... Includes the following steps: S1. Obtain track slab image data; S2. For the image data obtained in S1, mark the crack coordinate information, organize the image data and crack coordinate information obtained in S1 into YOLO dataset format, and construct a visual dataset of track slab cracks. S3. For the image data obtained in S1, randomly select images and perform expert-level descriptive annotations, including scene descriptions, crack geometric features, and maintenance suggestions. S4. Construct a visual model based on YOLO v11x, freeze the parameters of the model backbone network, and fine-tune the training using the track slab crack visual dataset from S2 to obtain a track slab crack detection visual model. S5. Combining the crack coordinate information obtained in S2 and the description annotations obtained in S3, construct a multimodal instruction fine-tuning dataset in the form of image-text pairs; S6. Construct a multimodal large model based on Qwen3-VL, use the multimodal instruction fine-tuning dataset obtained in S5, and use low-rank adaptation technology for fine-tuning to obtain a multimodal large model for track slab crack detection. S7. Combine the visual model for track slab crack detection obtained in S4 and the large multimodal model for track slab crack detection obtained in S6 to form the final small sample model for track slab crack detection. S8. The multimodal small sample model for track slab crack detection outputs a text description that includes image scene description, crack coordinate information, severity estimation, and maintenance decision suggestions.

2. The method for detecting cracks in railway ballastless track slabs using a multimodal small sample method as described in claim 1, characterized in that, When marking crack coordinate information, an open-source annotation tool is used to load the image and an annotation specification is established; the annotation categories include four types: rail, fastening system, track slab and crack; the annotation method adopts axis-aligned bounding boxes.

3. The method for detecting cracks in railway ballastless track slabs using a multimodal small sample method as described in claim 1, characterized in that, The following steps are included when constructing the visual dataset of track slab cracks: The annotation results are exported in YOLO format for training the visual model; in YOLO format, each object is represented by a quadruple: ;in, and The x and y coordinates of the target center point are: and For the width and height of the target, all values ​​are normalized to the range [0,1]. The annotation results are also exported in COCO format for subsequent multimodal cue word construction; in COCO format, the target coordinates are represented as... ,in The coordinates of the top left corner of the object. The coordinates are the bottom right corner of the object.

4. The method for detecting cracks in railway ballastless track slabs using a multimodal small sample method as described in claim 1, characterized in that, S4 specifically includes the following steps: Load the pre-trained weights of YOLO v11x, freeze the weights of the backbone and neck network of YOLO v11x, and only update the gradient of the detection head parameters to adapt to the features of small sample data. Training is performed using a combined loss function that includes both classification and regression loss; loss function The calculation formula is as follows: ; in: CIoU loss is used for bounding box regression. for The balance coefficient, Binary cross-entropy loss is used for category classification. for The balance coefficient, For the distribution focus loss, for The balance coefficient; after training, a visual model for detecting cracks in the track slab is obtained.

5. The method for detecting cracks in railway ballastless track slabs using a multimodal small sample method as described in claim 1, characterized in that, S5 specifically includes the following steps: Constructing user prompts involves embedding the COCO-formatted coordinates obtained from S2 into the prompt template; The construction assistant's answer specifically refers to the scene description, crack geometry parameters (width, length), severity, and maintenance suggestions marked in S3. Organize the image path, user prompts, and assistant responses into JSON-formatted dialogue data.

6. The method for detecting cracks in railway ballastless track slabs using a multimodal small sample method as described in claim 1, characterized in that, S6 specifically includes the following steps: Freeze all raw parameters of the visual encoder and language model of the Qwen3-VL model. ; Injecting a low-rank matrix in a linear layer bypass and Forward propagation is calculated as ,in, For input features, For incremental updates of the LoRA branch, The sum of the original parameters and the incremental updates of the LoRA branch will be used as the final parameters of this linear layer; The autoregressive language modeling loss function is used for optimization, and the calculation formula is as follows: ; in, For the predicted text terms, For the input image, The prompt word contains coordinate information. T This indicates the total length of the generated text sequence. To calculate the total value of the loss between the predicted text and the real text, This represents the index number of the word in the text sequence. For conditional probability, This indicates a prompt word given an input image and coordinates. and all previously generated text terms Under the condition that the model predicts the current number of... Each lexical element is a real label. The probability value; A large multimodal model for track slab crack detection was trained using the multimodal instruction fine-tuning dataset obtained from S5.

7. A readable storage medium, characterized in that, The readable storage medium stores a computer program that is adapted to be loaded by a processor and executed by the multimodal small sample detection method for cracks in railway ballastless track slabs according to any one of claims 1-6.

8. A computer device, characterized in that, The computer device includes a memory and a processor. The memory stores a computer program, which, when executed by the processor, performs the multimodal small sample detection method for cracks in railway ballastless track slabs according to any one of claims 1-6.