Steel rail B-display image damage detection method, electronic equipment and storage medium
By combining block coding with Transformer, combined with morphological processing and attention mechanism, the problems of high false alarm rate, high missed alarm rate and high computational complexity in rail flaw detection equipment are solved, and efficient identification and positioning of rail damage are achieved, which is suitable for existing ultrasonic flaw detection equipment.
Patent Information
- Application Number
- CN202510851942.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-09-16
AI Technical Summary
Existing technologies in rail flaw detection equipment have high false alarm rates, high missed alarm rates, and rely on manual analysis, making it difficult to accurately identify minor damage to the rails. In addition, the deep learning model has high computational complexity for high-resolution images, making it difficult to meet rapid detection requirements.
By combining block coding with Transformer, through morphological processing, connected domain extraction and feature filling technology, combined with multi-head local attention mechanism and deformable attention mechanism, a lightweight Transformer encoder and decoder are designed to achieve automatic extraction and fusion of damage features.
It significantly improves the efficiency and accuracy of damage detection in rail B-display images, reduces false alarm and missed alarm rates, and reduces computational complexity. It is suitable for rapid processing of high-resolution images and improves the intelligent level of detection.
Smart Images

Figure CN120655995A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of rail transit detection, and in particular relates to a rail B display image damage detection method, electronic equipment and storage medium. Background Art
[0002] Rails, the carrier of railway transportation, are a key factor affecting train safety. Currently, most rail flaw detection equipment, both domestically and internationally, relies on traditional piezoelectric ultrasonic principles. Based on factors such as speed, size, and cost, these devices can be categorized into the following types: hand-pushed rail flaw detectors, dual-track flaw detectors, road-rail vehicles, and large rail flaw detection vehicles. These ultrasonic flaw detection devices use ultrasonic waves to inspect rails for internal flaws, and the location and type of flaws are determined by playback and analysis of the rail's ultrasonic B-ray image.
[0003] The current flaw detection vehicle data analysis software suffers from high false alarm rates and missed detections. Accurate identification of rail damage still relies on manual analysis of B-display playback data. The damage detection rate of manual analysis is highly dependent on the analyst's experience and is subject to subjective errors. Furthermore, when analyzing rail damage over a long distance, the reviewer must spend extended periods watching the software, which can lead to fatigue, inadvertent errors, and even misjudgments or missed detections.
[0004] Traditional methods rely on manual feature extraction, making it difficult to capture damage details in B-scan images (e.g., weak signals from tiny nuclear lesions and pore fissures). Existing deep learning models (e.g., CNNs) lack the ability to model global context and fail to incorporate prior signal characteristics (e.g., scale range, channel type) to optimize feature representation. Using Transformers to directly process high-resolution images (e.g., 1280×1280) is computationally complex and cannot meet the requirements of rapid detection. Summary of the Invention
[0005] The technical solution of the present invention is used to solve the problem of how to achieve efficient identification and positioning of complex damages such as rail core damage and hole cracks.
[0006] The present invention solves the above technical problems through the following technical solutions: The present invention provides a rail B display image damage detection method, comprising: S1 divides the B-display image into blocks, normalizes each sub-block, and generates a one-dimensional sequence code in row and column order to preserve the spatial position information; S2 performs morphological processing and connected domain extraction on the B-display image processed in step S1; S3 extracts physical features from each connected domain, including geometric features, intensity features, and morphological features of the three RGB channels; S4 performs feature filling on each sub-block, maps the constructed multi-dimensional feature vector to a high-dimensional space through a fully connected layer, and adds position encoding as the input of the Transformer encoder; The S5 uses a three-layer Transformer encoder and employs a multi-head local attention mechanism in each layer to reduce computational complexity and achieve lightweight Transformer encoders. The S6 decoder is designed as a three-layer decoder. A classification detection head and regression operation are added to the last decoder layer. The classification detection head is used to output damage classification, thereby realizing damage identification, and the regression operation is used to realize damage location. At the same time, a deformable attention mechanism is used instead of the self-attention mechanism in the decoder structure of each layer to reduce computational complexity.
[0007] Furthermore, the method for extracting the connected domain is as follows: 1) Copy the RGB image three times, traverse the pixels of each image, filter out pixels of other colors, and only keep the data of the same color; 2) For each piece of data, traverse the pixels and set the non-zero pixels to 255 and the zero-value pixels to 0, thus performing binarization. 3) Use the Suzuki boundary tracing algorithm to extract all connected regions in the binary image; 4) Delete connected domains with less than 10 pixels to eliminate noise interference.
[0008] Furthermore, the rule for filling features for each sub-block is as follows: if the sub-block is completely inside a connected domain, all features of the connected domain are filled; if the sub-block spans multiple connected domains, features are weighted and fused according to area ratio, and feature vectors of undamaged sub-blocks are set to zero.
[0009] Furthermore, the multidimensional feature vector is designed as follows: each of the three RGB channels has 6 geometric features, namely area, perimeter, aspect ratio, width of the minimum circumscribed rectangle, height and rotation angle, totaling 18 for the three channels; each channel has 2 morphological features, namely circularity and number of skeleton branches, totaling 6 for the three channels; there are 2 global intensity features, namely grayscale mean and maximum value; together they constitute a 26-dimensional feature vector.
[0010] Furthermore, the process of the local attention mechanism is as follows: 1) Generate query (Q), key (K) and value (V) from the input sequence through linear transformation; 2) Calculate attention scores and attention weights; 3) Multiply the value vector by the attention weight to get the value tensor output; 4) Multi-head splicing of value tensors to output value tensor vectors; 5) Project the concatenated value tensor vector and calculate the local attention output.
[0011] Furthermore, the implementation process of the deformable attention mechanism is as follows: 1) The decoder receives a feature map from the encoder containing contextual information about the input image; 2) After the query vector is mapped from the feature map to the content feature space through a fully connected layer, semantic information features and position features are generated. Based on the coordinates of the position feature, random sampling points around the coordinates are focused on to avoid global calculations. 3) The offset is initialized to 0 and adaptively adjusted through linear layer training. For each query, the corresponding offset is calculated and used to dynamically update the sampling point position on the feature map; 4) Sample features from the feature map based on the calculated offset; 5) By calculating the attention score, the attention weight is generated to weight the sampled features; 6) Use attention weights to perform weighted aggregation on the sampled features to generate the final output features.
[0012] Furthermore, the calculation formula of the offset is as follows:
[0013] Among them, query is the query vector, △p is the offset, Linear offset It is a linear network; The formula for updating the sampling point position is as follows:
[0014] Among them, p k is the updated sampling point position, scale is the scaling factor, and reference_points is the coordinate in the query vector.
[0015] Furthermore, the formula of the sampling feature is as follows:
[0016] Among them, bilinear_sample means that features are obtained by bilinear interpolation of 4 pixels around the sampling point through the feature map, sampled_features means the feature vector of the sampling point, and feature_map means the feature map; The final output features of the generation are as follows:
[0017] Among them, output represents the generation of the final output features, weights k Represents the attention weight of the k-th sampling point, Linear V Represents the value projection matrix, Sampled_features k Represents the k-th sampling point feature vector.
[0018] The present invention also provides an electronic device, comprising a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the above-mentioned rail B display image damage detection method, and the processor is configured to execute the program stored in the memory.
[0019] The present invention also provides a storage medium having a computer program stored thereon. When the computer program is run by a processor, the steps of the above-mentioned rail B display image damage detection method are executed.
[0020] The beneficial effects of the present invention are as follows: 1) By combining block coding with the Transformer approach, the efficiency and accuracy of damage detection in rail B-display images were significantly improved, reducing the false positive and false negative rates of traditional manual analysis. The lightweight encoder design and local attention mechanism effectively reduced computational complexity while maintaining global context modeling capabilities, making it suitable for fast processing of high-resolution images. 2) Morphological processing, connected domain extraction, and feature filling techniques are used to automatically extract and fuse damage features, reducing reliance on manual experience and improving the objectivity of detection. The deformable attention mechanism and classification detection head design enable the model to adaptively focus on key areas, further enhancing the intelligent level of damage identification and localization. 3) Combining the advantages of block coding and Transformer, it solves the problem of insufficient global context modeling in traditional CNN models while avoiding the high computational complexity of directly processing high-resolution images. Through physical feature enhancement and dynamic position encoding, it optimizes damage representation capabilities, making it particularly suitable for detecting complex damage such as core damage and pore cracks. 4) The method can be directly applied to existing ultrasonic flaw detection equipment without the need for additional hardware support and can be easily integrated into actual detection systems. It avoids non-maximum suppression (NMS) post-processing, simplifies the process, and improves detection efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a flow chart of the method of the present invention; Figure 2 This is an example of a nuclear injury graphic; Figure 3 This is an example of an RGB three-channel B display image; Figure 4It is the structural design diagram of the encoder and decoder; Figure 5 It is an example diagram of the joint range in the prior features of the B-display image. DETAILED DESCRIPTION
[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0023] The technical solution of the present invention is further described below with reference to the accompanying drawings and specific embodiments: Example 1 like Figure 1 As shown, the rail B display image damage detection method according to the embodiment of the present invention includes the following steps: Step 1: B-display image segmentation and sequence encoding 1.1 Block rules In the B-display image, Figure 2 As shown, the size of the nuclear injury pattern is within the range of 80×40 pixels. Therefore, the 1280×1280 pixel B-display image is divided into sub-blocks of 80×40 pixels, generating a total of 512 sub-blocks (number of horizontal blocks = 1280 / 80 = 16, number of vertical blocks = 1280 / 40 = 32).
[0024] 1.2 Sequence Encoding After normalizing each sub-block, a one-dimensional sequence code (format: row number_column number) is generated in row and column order, preserving the spatial position information.
[0025] Normalization goal: Map row and column indices to the interval [0, 1] to preserve spatial location information.
[0026] Normalization formula: Normalized_row = row / 31, Normalized_col = col / 15 (row and column indexes start at 0); where Normalized_row represents the normalized row coordinate, Normalized_col represents the normalized column coordinate, row represents the row, and col represents the column.
[0027] For example, the sub-block at row 5 and column 10 is selected, and the normalized coordinates are Normalized_row=5 / 31≈0.1613, Normalized_col=10 / 15≈0.6667, and the final generated format is 0.1613_0.6667.
[0028] Step 2: Morphological processing and connected domain extraction 2.1. Morphological processing includes: noise suppression, dilation processing and corrosion processing.
[0029] (1) Noise suppression: Median filtering (kernel size 3×3) is performed on the B-display image to eliminate speckle noise.
[0030] Median filtering is a nonlinear filtering method that replaces the central pixel value with the median value of the pixels in the sliding window. The mathematical expression is:
[0031] Among them, I filtered The pixel value after median filtering at (x, y), I (x+i, y+j) represents the pixel value of the original image at coordinate (x, y), W represents the filtering window, the Median{} function represents the median of all pixels in the window, i represents the offset in the row direction, and j represents the offset in the column direction.
[0032] (2) Expansion processing: A rectangular structural element (radius = 3 pixels) is used for expansion to connect the broken damaged areas.
[0033] The dilation process is to assign the maximum value within the area covered by the structural element to the central pixel. The specific formula is as follows:
[0034] Where A is the input image set, B is the structure element set, (X, Y) is the current pixel coordinate, and (i, j) is the offset vector within the structure element.
[0035] (3) Corrosion treatment: Use the same structural elements for corrosion to eliminate isolated noise points and retain the main damage contours.
[0036] Erosion is the inverse operation of dilation, which assigns the minimum value within the area covered by the structural element to the central pixel. The specific formula is as follows:
[0037] 2.2 Connected Domain Extraction like Figure 3 As shown in the figure, each color stripe in the RGB three-channel B-display image represents an ultrasonic scanning channel. For a sub-block, there are at most three colors, representing the three channels of ultrasonic scanning. Therefore, connected domains are extracted for the three channels in each area. The specific steps are as follows: 1) Copy the RGB image three times, traverse the pixels of each image, filter out pixels of other colors, and only keep the data of the same color; Figure 2Taking data as an example, one piece of data only retains red pixels, one piece only retains yellow pixels, and one piece only retains white pixels; 2) For each piece of data, traverse the pixels and set the non-zero pixels to 255 and the zero-value pixels to 0, thus performing binarization. 3) Use the Suzuki boundary tracing algorithm to extract all connected regions in the binary image; 4) Delete connected domains with less than 10 pixels to eliminate noise interference.
[0038] The steps of the Suzuki boundary tracking algorithm are as follows: ① Traverse the entire image, scanning pixels from left to right and from top to bottom, and find foreground pixels with a value of 255 that have not been tracked; ② Start boundary tracking from the starting point, put the pixel into a contour queue, record it as visited, and start from this point to detect continuous foreground pixels along the neighborhood and move along the boundary; ③ Use neighborhood search (8-neighborhood), direction number (clockwise), the search order is to start from the next direction of the previous boundary point, and poll the 8 neighboring points clockwise; starting from the direction of the previous point, search the neighboring pixels clockwise. If a foreground point (value = 255) is found, it is used as the next contour point and continues to track along the boundary until it returns to the starting point (completely closed); ④After the tracking is completed, the generated contour uses a sequence of closed contour points to mark the contour area to avoid repeated tracking; ⑤Continue scanning the remaining pixels, skipping the tracked area until all foreground pixels are visited and tracked.
[0039] Step 3: Extract features inside connected domains The physical features of each connected domain are extracted, and the physical features include: geometric features and morphological features of three channels of the RGB image, and global intensity features.
[0040] (1) Geometric characteristics of the three RGB channels After dividing the area, we can see from the RGB image characteristics of the B-display image that there are at most three channels within the 40-pixel height range. The geometric features of each channel include: area, perimeter, aspect ratio, width and height of the minimum enclosing rectangle, and rotation angle.
[0041] The area is obtained by: the number of pixels in the connected domain; The perimeter is obtained by: the sum of the Euclidean distances between boundary pixels; The aspect ratio is obtained as follows: width of bounding rectangle / height of bounding rectangle; The width, height, and rotation angle of the minimum bounding rectangle are obtained by calculating the convex hull of the contour, using the rotating calculus algorithm to find the minimum area enclosing rectangle on the convex hull, and outputting the width, height, and rotation angle.
[0042] (2) The morphological characteristics of each of the three RGB channels include circularity and the number of skeleton branches, where the number of skeleton branches represents the complexity of the damage.
[0043] 1) The calculation method of circularity is: circularity = (4*π*area) / (circumference²) 2) The skeleton branch number extraction method is as follows: use the Zhang-Suen algorithm to obtain the connected domain skeleton; traverse the skeleton pixels and mark if there is only one adjacent skeleton point in the 8-neighborhood; the statistical count of all marked points is the skeleton branch number.
[0044] The process of the Zhang-Suen algorithm is as follows: ① Mark all pixels that meet the deletion conditions of the first iteration; ② Delete all marked pixels; ③Mark all pixels that meet the second iteration deletion conditions; ④ Delete all marked pixels; ⑤ Repeat steps ①-④ until no pixels are deleted.
[0045] The relevant definitions of the Zhang-Suen algorithm are as follows: For the central pixel P1, its 8 neighboring pixels are defined as follows: P9 P2 P3; P8 P1 P4; P7 P6 P5; The first iteration deletion conditions are as follows: 2 ≤ B(P1) ≤ 6; A(P1) = 1; P2 × P4 × P6 = 0; P4 × P6 × P8 = 0; The second iteration deletion condition is as follows: 2 ≤ B(P1) ≤ 6; A(P1) = 1; P2 × P4 × P8 = 0; P2 × P6 × P8 = 0; Where B(P1) is the number of foreground pixels in the 8-neighborhood of P1; B(P1) = P2 + P3 + P4 + P5 +P6 + P7 + P8 + P9; A(P1) is the number of 0→1 transitions from P2 to P9 in the 8-neighborhood of P1; A(P1) = the number of 0→1 transitions.
[0046] (3) Global intensity features include: grayscale mean and maximum value.
[0047] The grayscale mean is obtained by traversing the coordinate points of each connected area and extracting the RGB value in the original image according to the coordinates. The grayscale value is calculated according to the formula: value=0.299*R+0.587*G+0.114*B. All grayscale values are added up and divided by the number of pixels in the area: 80*40=3200, which is the grayscale mean.
[0048] The maximum value is obtained by taking the maximum value of the heights of the boundary rectangles of all connected domains.
[0049] Step 4: Feature filling and multi-dimensional feature vector design Feature filling is performed on each sub-block, and the constructed multi-dimensional feature vector is mapped to a high-dimensional space through a fully connected layer, and position encoding is added as the input of the Transformer encoder.
[0050] The rule for feature filling is: if a sub-block is completely within a connected domain, all features of that domain are filled. If a sub-block spans multiple connected domains, features are weighted and fused based on their area proportions, and the feature vectors of undamaged sub-blocks are set to zero. Feature filling is performed to ensure that each sub-block can be feature mapped and retain the features of the current sub-block. Without feature filling, sub-blocks will lose features and data alignment will be impossible.
[0051] Design of multi-dimensional feature vector: Each sub-block corresponds to a 26-dimensional feature vector, and the 26-dimensional feature vector is composed as follows: Each of the three RGB channels has 6 geometric features (area, perimeter, aspect ratio, width and height of the minimum enclosing rectangle, and rotation angle), for a total of 18. Each channel has 2 morphological features (circularity and number of skeleton branches), for a total of 6. There are also 2 global intensity features (grayscale mean and maximum), which together form a 26-dimensional feature vector.
[0052] The 26-dimensional feature vector is mapped to a high-dimensional space (embedding dimension = 256) through a fully connected layer, and position encoding (generated based on row and column numbers) is added. The purpose of input embedding is to add spatial position relationships to the high-dimensional information as the input of the Transformer encoder.
[0053] Step 5: Lightweight Transformer encoder design like Figure 4 As shown in Figure 1, a three-layer Transformer encoder is used, with each layer containing a 4-head local attention mechanism. Using a multi-head local attention mechanism in each layer of the Transformer encoder can reduce the amount of computation and achieve lightweight Transformer encoder. Since the joint range in the prior features of the B-display image is within 400 pixels, as shown in Figure 1, Figure 5 As shown in the figure, the calculation scope of the multi-head local attention mechanism is limited to adjacent 5×5 sub-blocks (25 sub-blocks in total) to reduce the amount of calculation.
[0054] The process of the local attention mechanism is as follows: 1) Generate query (Q), key (K) and value (V) from the input sequence through linear transformation; 2) Calculate attention scores and attention weights; 3) Multiply the value vector by the attention weight to get the value tensor output; 4) Multi-head splicing of value tensors to output value tensor vectors; 5) Project the concatenated value tensor vector and calculate the local attention output.
[0055] Step 6: Design a decoder with a classification detection head and regression operation The purpose of designing a decoder with a classification detection head and regression operation is to further process the encoder features through a deformable attention mechanism, add classification and regression detection heads, and finally efficiently identify damage and locations in the b-display image.
[0056] like Figure 4 As shown in the figure, the decoder is designed as three layers, with a classification detection head and regression operation added to the last decoder layer. The classification detection head is used to output the damage classification, thereby achieving efficient damage identification, and the regression operation is used to achieve damage location. At the same time, a deformable attention mechanism is used instead of a self-attention mechanism in each decoder layer. This replaces the global calculation of traditional attention with dynamic sparse sampling, focusing only on a small number of sampling points (such as 4-8) in the feature map rather than the global position, significantly reducing computational complexity. At the same time, the position offset of the sampling point is generated by learnable parameters, enabling the model to adaptively focus on key target areas (such as object edges or texture-dense areas).
[0057] The following is the implementation process of the deformable attention mechanism: 1) The decoder receives feature maps from the encoder, which contain contextual information of the input image.
[0058] 2) After the query vector is mapped from the feature map to the content feature space through the fully connected layer, semantic information features and position features are generated. Based on the coordinates of the position feature, a small number of sampling points (4) around the coordinates are randomly focused to avoid global calculations.
[0059] 3) The offset is initialized to 0 and adaptively adjusted through linear layer training. For each query, the model calculates the corresponding offset, which is used to dynamically update the sampling point position on the feature map. The offset calculation formula is as follows:
[0060] Among them, query is the query vector, △p is the offset, Linear offset It is a linear network.
[0061] The formula for updating the sampling point position is as follows:
[0062] Among them, p k is the updated sampling point position, scale is the scaling factor, and reference_points is the coordinate in the query vector.
[0063] 4) Based on the calculated offset, features are sampled from the feature map. These features are dynamically selected based on the context of the query, and the calculation formula is as follows:
[0064] Among them, bilinear_sample means obtaining features through bilinear interpolation of 4 pixels around the sampling point through the feature map, Sampled_features means the feature vector of the sampling point, and feature_map means the feature map.
[0065] 5) By calculating the attention scores, attention weights are generated, which are used to weight the sampled features.
[0066] 6) Use attention weights to perform weighted aggregation on the sampled features to generate the final output features.
[0067]
[0068] Among them, output represents the generation of the final output features, weights k Represents the attention weight of the k-th sampling point, Linear V Represents the value projection matrix, Sampled_features k Represents the k-th sampling point feature vector.
[0069] Example 2 An electronic device includes a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the rail B display image damage detection method in embodiment 1, and the processor is configured to execute the program stored in the memory.
[0070] Example 3 A storage medium stores a computer program, which, when executed by a processor, executes the steps of the rail B image damage detection method in embodiment 1.
[0071] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A rail B display image damage detection method, characterized in that: include: S1 divides the B-display image into blocks, normalizes each sub-block, and generates a one-dimensional sequence code in row and column order to preserve the spatial position information; S2 performs morphological processing and connected domain extraction on the B-display image processed in step S1; S3 extracts physical features from each connected domain, including geometric features, intensity features, and morphological features of the three RGB channels; S4 performs feature filling on each sub-block, maps the constructed multi-dimensional feature vector to a high-dimensional space through a fully connected layer, and adds position encoding as the input of the Transformer encoder; The S5 uses a three-layer Transformer encoder and employs a multi-head local attention mechanism in each layer to reduce computational complexity and achieve lightweight Transformer encoders. The S6 decoder is designed as a three-layer decoder. A classification detection head and regression operation are added to the last decoder layer. The classification detection head is used to output damage classification, thereby realizing damage identification, and the regression operation is used to realize damage location. At the same time, a deformable attention mechanism is used instead of the self-attention mechanism in the decoder structure of each layer to reduce computational complexity.
2. The rail B image damage detection method according to claim 1, characterized in that: The method for extracting the connected domain is as follows: 1) Copy the RGB image three times, traverse the pixels of each image, filter out pixels of other colors, and only keep the data of the same color; 2) For each piece of data, traverse the pixels and set the non-zero pixels to 255 and the zero-value pixels to 0, thus performing binarization. 3) Use the Suzuki boundary tracing algorithm to extract all connected regions in the binary image; 4) Delete connected domains with less than 10 pixels to eliminate noise interference.
3. The rail B image damage detection method according to claim 1, characterized in that: The rule for filling features for each sub-block is as follows: if the sub-block is completely inside a connected domain, all features of the connected domain are filled; if the sub-block spans multiple connected domains, the features are weighted and fused according to the area ratio, and the feature vectors of the undamaged sub-blocks are set to zero.
4. The rail B image damage detection method according to claim 1, characterized in that: The multidimensional feature vector is designed as follows: each of the three RGB channels has 6 geometric features, namely area, perimeter, aspect ratio, width of the minimum circumscribed rectangle, height and rotation angle, totaling 18 for the three channels; each channel has 2 morphological features, namely circularity and number of skeleton branches, totaling 6 for the three channels; there are 2 global intensity features, namely grayscale mean and maximum value; together they form a 26-dimensional feature vector.
5. The rail B image damage detection method according to claim 1, characterized in that: The process of the local attention mechanism is as follows: 1) Generate query (Q), key (K) and value (V) from the input sequence through linear transformation; 2) Calculate attention scores and attention weights; 3) Multiply the value vector by the attention weight to get the value tensor output; 4) Multi-head splicing of value tensors to output value tensor vectors; 5) Project the concatenated value tensor vector and calculate the local attention output.
6. The rail B image damage detection method according to claim 1, characterized in that: The implementation process of the deformable attention mechanism is as follows: 1) The decoder receives a feature map from the encoder containing contextual information about the input image; 2) After the query vector is mapped from the feature map to the content feature space through a fully connected layer, semantic information features and position features are generated. Based on the coordinates of the position feature, random sampling points around the coordinates are focused on to avoid global calculations. 3) The offset is initialized to 0 and adaptively adjusted through linear layer training. For each query, the corresponding offset is calculated and used to dynamically update the sampling point position on the feature map; 4) Sample features from the feature map based on the calculated offset; 5) By calculating the attention score, the attention weight is generated to weight the sampled features; 6) Use attention weights to perform weighted aggregation on the sampled features to generate the final output features.
7. The rail B image damage detection method according to claim 6, characterized in that: The calculation formula of the offset is as follows: Among them, query is the query vector, △p is the offset, Linear offset It is a linear network; The formula for updating the sampling point position is as follows: Among them, p k is the updated sampling point position, scale is the scaling factor, and reference_points is the coordinate in the query vector.
8. The rail B image damage detection method according to claim 6, characterized in that: The formula of the sampling feature is as follows: Among them, bilinear_sample means obtaining features through bilinear interpolation of 4 pixels around the sampling point through the feature map, sampled_features means the feature vector of the sampling point, and feature_map means the feature map; The final output features of the generation are as follows: Among them, output represents the generation of the final output features, weights k Represents the attention weight of the k-th sampling point, Linear V Represents the value projection matrix, Sampled_features k Represents the k-th sampling point feature vector.
9. An electronic device comprising a memory and a processor, characterized in that: The memory is used to store a program that supports the processor to execute the rail B image damage detection method according to any one of claims 1 to 8, and the processor is configured to execute the program stored in the memory.
10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the rail B image damage detection method according to any one of claims 1 to 8 are executed.