Guide rail surface defect detection method and system based on image segmentation model, and medium

The integration of RGB and polarimetric data with DeepLabv3+ and attention mechanisms improves rail surface defect detection precision and robustness by enhancing feature extraction and fusion, addressing lighting and reflection issues in single RGB image methods.

CN120318174APending Publication Date: 2025-07-15SOUTH CHINA AGRICULTURAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510385617.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-29
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing guide rail surface defect detection method based on RGB images has poor detection effect under the influence of light conditions and surface reflection, making it difficult to effectively identify complex and diverse defects.

Method used

The cross-modal data set combined with the image segmentation model is used to extract the light intensity and polarization characteristics of RGB data and polarization data through a multi-layer convolutional network, and train it using the CBAM attention mechanism and the improved DeepLabv3+ network to achieve image segmentation and improve defect detection accuracy.

Benefits of technology

It significantly improves the segmentation accuracy of small defects, enhances the robustness of the model in complex environments, reduces the computational complexity, and meets the needs of real-time industrial detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318174A_ABST
    Figure CN120318174A_ABST
Patent Text Reader

Abstract

The invention relates to the field of image processing, in particular to a guide rail surface defect detection method and system based on an image segmentation model and a medium. The method comprises the following steps: acquiring guide rail surface RGB data and guide rail surface polarization data, forming a cross-modal data set, constructing an image segmentation model, and segmenting a to-be-segmented image on the guide rail surface by using the image segmentation model reaching a convergence state to obtain a segmentation result; and based on the segmentation result, obtaining a detection result of the surface defect of the guide rail. Accurate detection of the surface defects of the guide rail is achieved, and the detection effect and efficiency are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and particularly to a method, system and medium for detecting rail surface defects based on an image segmentation model. Background Art

[0002] Image segmentation is a representative task that supports computer-aided image content analysis and has important applications in industrial defect detection. As a key component in industrial manufacturing, rails are widely used in fields such as automated equipment, machine tools, and robots. The surface quality of rails directly affects the accuracy and service life of equipment. As an important guiding and supporting component in a mechanical system, the surface quality of rails directly affects the operating accuracy and life of the mechanical system. Detecting defects on the rail surface is a key link to ensure the quality of rails. Most traditional detection methods rely on manual visual inspection, which is not only inefficient but also easily affected by human factors, resulting in missed or false detections.

[0003] With the development of computer vision and deep learning technologies, automated detection methods based on image processing have gradually become a research hotspot. Most existing methods for detecting rail surface defects based on image processing use only single RGB (Red, Green, Blue) image data as input and implement defect detection by constructing an image classification or object detection model. However, defects on the rail surface often have complex and diverse characteristics, such as scratches, rust, cracks, etc. The appearance of these defects in RGB images may be affected by factors such as lighting conditions and surface reflections, resulting in poor detection effects. Summary of the Invention

[0004] (1) Technical Problems to be Solved

[0005] The main purpose of the present invention is to provide a method, system and medium for detecting rail surface defects based on an image segmentation model to solve the above technical problems.

[0006] (2) Technical Solutions

[0007] To achieve the above object, the present invention provides a method for detecting rail surface defects based on an image segmentation model, including the steps of:

[0008] S1, obtaining RGB data of the rail surface and polarization data of the rail surface, and forming a cross-modal data set;

[0009] S2, constructing an image segmentation model; including:

[0010] S21. Extract the features of different modality data in the cross-modal dataset through a multi-layer convolutional network. The features include the light intensity features extracted from the RGB data and the polarization features extracted from the polarization data. Then, fuse the features using a multi-scale feature fusion module.

[0011] S22. Based on the fused features, introduce the CBAM attention mechanism to construct an improved DeepLabv3+ network model, realize the construction of the image segmentation model, and perform training iterations on the image segmentation model until the image segmentation model reaches a convergence state.

[0012] S3. Use the image segmentation model that has reached the convergence state to segment the image to be segmented on the guide rail surface to obtain a segmentation result. Based on the segmentation result, obtain the detection result of the defects on the guide rail surface.

[0013] Preferably, the step S21 includes:

[0014] Use MobileNetV3-CA as the backbone network architecture of the image segmentation model for network improvement.

[0015] Adopt the MobileNetV3-CA structure in the RGB data branch to extract the color and texture features of the guide rail surface in the RGB data.

[0016] Adopt the MobileNetV3-CA structure in the polarization data branch to extract the material and polarization features of the guide rail surface in the polarization data.

[0017] Use h-swish (Hard Swish) as the activation function of the image segmentation model; use the coordinate attention mechanism as the attention mechanism of the image segmentation model.

[0018] Preferably, the CBAM attention mechanism includes:

[0019] A channel attention module that extracts the global statistical information of each channel through global average pooling and global maximum pooling respectively, and learns the weights of each channel.

[0020] A spatial attention module that obtains the maximum value and average value of each spatial position through maximum pooling and average pooling respectively, and learns the weight distribution of each spatial position.

[0021] Preferably, the formula of the h-swish activation function is:

[0022]

[0023] ReLU6 = min(6, max(0, x))

[0024] Among them, x represents the input feature, min() represents finding the minimum value, and max() represents finding the maximum value.

[0025] Preferably, the coordinate attention mechanism includes:

[0026] Compress the spatial dimension (H×W) of the feature map into two independent one-dimensional vectors through global pooling operations, respectively representing the global information in the height and width directions;

[0027] Generate attention weights using the two one-dimensional vectors through convolution and non-linear transformation, and respectively act on the height and width dimensions of the feature map.

[0028] Preferably, the formulas for the two independent one-dimensional vectors are:

[0029]

[0030] Among them, X represents the feature map, H represents the height, W represents the width, and z h represents the average pooling of the feature map along the height direction, and z ω represents the average pooling of the feature map along the width direction.

[0031] Preferably, the coordinate attention mechanism is preferentially used in the shallow layer of the encoder, and the CBAM attention mechanism is adopted in the deep layer of the decoder. The coordinate attention mechanism and the CBAM attention mechanism are adaptively fused through a gating mechanism.

[0032] Preferably, the multi-scale feature fusion module in step S21 includes:

[0033] S211, fuse the feature maps from the MobileNetV3-CA network branch of the RGB data and the MobileNetV3-CA network branch of the polarization data in the way of pixel addition;

[0034] S212, input the fused feature map into the atrous spatial pyramid pooling module of DeepLabV3+;

[0035] S213, apply the CBAM attention mechanism again in the atrous spatial pyramid pooling module and the shallow feature extraction network of the encoder respectively.

[0036] The present invention also provides a guide rail surface defect detection system based on an image segmentation model, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the guide rail surface defect detection method based on an image segmentation model as described in any one of the above.

[0037] The present invention also provides a computer-readable storage medium storing a computer program, which when executed by a processor, implements the steps of the method for detecting rail surface defects based on an image segmentation model as described in any one of the above.

[0038] (III) Advantageous Effects

[0039] A method, system and medium for detecting rail surface defects based on an image segmentation model proposed in this application, by introducing polarization image information and combining a dual-input fusion mechanism of light intensity features and polarization features, can effectively enhance the contrast of micro-defects, solve the deficiencies of traditional RGB images in detecting low-contrast defects, and significantly improve the segmentation accuracy of micro-defects.

[0040] A method, system and medium for detecting rail surface defects based on an image segmentation model proposed in this application, combines the coordinate attention mechanism and the CBAM attention mechanism, fully extracts the texture information in the polarization image, avoids background noise interference, further enhances the model's attention to key features, improves the segmentation accuracy, ensures the efficient distinction between the defect area and the background, and improves the robustness of the model in complex industrial environments.

[0041] A method, system and medium for detecting rail surface defects based on an image segmentation model proposed in this application, based on the dual-input feature fusion design of the DeepLabv3+ network, while maintaining high segmentation accuracy, reduces the computational complexity by optimizing the feature extraction and fusion process, meeting the requirements of industrial real-time detection and large-scale applications. Description of the Drawings

[0042] Figure 1 is the main flowchart of the method for detecting rail surface defects based on an image segmentation model provided by an embodiment of the present invention;

[0043] Figure 2 is the network architecture diagram of the method for detecting rail surface defects based on an image segmentation model provided by an embodiment of the present invention;

[0044] Figure 3 is the backbone network bottleneck structure diagram of the method for detecting rail surface defects based on an image segmentation model provided by an embodiment of the present invention;

[0045] Figure 4 is the attention mechanism structure diagram of the method for detecting rail surface defects based on an image segmentation model provided by an embodiment of the present invention;

[0046] Figure 5 is the visualization diagram of the segmentation results corresponding to various image segmentation methods provided by an embodiment of the present invention;

[0047] Figure 6It is a block diagram of the module structure of the guide rail surface defect detection system provided by an embodiment of the present invention;

[0048] Figure 7 It is a schematic diagram of the hardware structure of the guide rail surface defect detection system provided by an embodiment of the present invention. Specific embodiments

[0049] To better explain the present invention for easy understanding, the present invention will be described in detail below with reference to the accompanying drawings through specific embodiments.

[0050] It should be noted that all directional indications (such as up, down, left, right, front, back...) in the embodiments of the present invention are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indications will also change accordingly.

[0051] In addition, in the present invention, descriptions such as "first" and "second" are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0052] In the present invention, unless otherwise clearly defined and limited, terms such as "connection" and "fixation" should be understood in a broad sense. For example, "fixation" can be a fixed connection, a detachable connection, or integrated; "connection" can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and can be the communication inside two components or the interaction relationship between two components, unless otherwise clearly limited. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0053] As Figure 1 shown, an embodiment of the present invention provides a method for detecting guide rail surface defects based on an image segmentation model. The method includes the steps:

[0054] S1. Obtain the RGB data and polarization data of the guide rail surface and form a cross-modal data set. In other embodiments, the RGB data can be, for example, RGB standard data, and the polarization data of the guide rail surface can be, for example, polarization optical defect data.

[0055] S2. Build an image segmentation model, including:

[0056] S21. Extract the features of different modalities of data in the cross-modal dataset through a multi-layer convolutional network (the different modalities of data refer to two different sources of data: RGB data on the guide rail surface and polarization data on the guide rail surface). The features include the light intensity features extracted from the RGB data and the polarization features extracted from the polarization data. Then, use a multi-scale feature fusion module to fuse the features, which can enhance the contrast of tiny defects in the detection image. In other embodiments, the features include high-level features and low-level features.

[0057] S22. Based on the fused features, introduce the CBAM attention mechanism to construct an improved DeepLabv3+ network model, realize the construction of the image segmentation model, and perform training iterations on the image segmentation model until the image segmentation model reaches a convergence state.

[0058] S3. Use the image segmentation model that has reached the convergence state to segment the image to be segmented on the guide rail surface, and obtain the segmentation result of the image. Based on the segmentation result, obtain the detection result of the defects on the guide rail surface of the image. As Figure 5 shown, a method for detecting guide rail surface defects based on an image segmentation model in this embodiment has achieved better performance in terms of the edge accuracy of the target area.

[0059] Optionally, in this embodiment, step S1 further includes:

[0060] S11. Use a high-frame-rate industrial camera integrated with a polarization filter (resolution ≥ 20MP) to synchronously obtain RGB three-channel light intensity data and Stokes vector polarization parameters on the XYZ three-axis guide rail motion platform. Among them, the Stokes vector polarization parameters are a set of parameters used to describe the polarization state of electromagnetic waves (especially light). Construct a multi-modal data pair with sub-pixel level alignment characteristics.

[0061] S12. Measure the defect depth (accuracy ±0.1μm) through a microhardness tester, and combine X-ray diffraction to analyze the surface stress distribution to establish a multi-dimensional label system of defect morphology - physical properties.

[0062] S13. Use the stratified random sampling method to construct a training set and a test set in a ratio of 4:1.

[0063] S14. Through data enhancement strategies such as adaptive illumination transformation, non-rigid deformation enhancement, multi-scale Gaussian blur, and composite noise injection, simulate complex illumination conditions and imaging interference factors in the industrial field on the training set.

[0064] Among them, the RGB data and the polarization data are complementary, and the polarization data is sensitive to the Brewster angle effect of surface scratches (the signal-to-noise ratio is increased by 4.2 dB); the RGB data retains the material color characteristics (the color difference tolerance ΔE < 2.5).

[0065] As Figure 2 shown, as a preferred embodiment of the present invention, the step S21 includes:

[0066] Using MobileNetV3-CA as the backbone network architecture of the image segmentation model for network improvement; among them, MobileNetV3 is a lightweight network model proposed by the Google team in 2019, which combines neural architecture search (NAS) and optimized convolutional structures, further improving the computational efficiency and inference speed of the network. The MobileNetV3-CA introduces the attention mechanism CA on the basis of MobileNetV3 to better adapt to specific tasks or datasets;

[0067] Adopting the MobileNetV3-CA structure in the RGB data branch to extract the color and texture features of the guide rail surface in the RGB data;

[0068] Adopting the MobileNetV3-CA structure in the polarization data branch to extract the material and polarization features of the guide rail surface in the polarization data;

[0069] Using h-swish as the activation function of the image segmentation model; using the coordinate attention mechanism as the attention mechanism of the image segmentation model, among which, the h-swish function is a new type of activation function proposed by the MNASNet team of Huawei in 2019, mainly optimizing the computational efficiency of the activation function.

[0070] Specifically, in this embodiment, the CBAM attention mechanism includes:

[0071] Channel attention module, which extracts the global statistical information of each channel through global average pooling and global maximum pooling respectively, and learns the weights of each channel;

[0072] Spatial attention module, which obtains the maximum value and average value of each spatial position through maximum pooling and average pooling respectively, and learns the weight distribution of each spatial position.

[0073] Furthermore, in this embodiment, the channel attention module includes:

[0074] Parallelly execute global average pooling (GAP) and global maximum pooling (GMP) on the input feature map to generate channel description vectors respectively:

[0075]

[0076] Among them, C represents the number of channels, H×W represents the height and width of the input feature map, c∈[1,C] represents the channel index, and F c represents the two-dimensional matrix of the c-th channel of the input feature map

[0077] Input the dual-path vector into a multi-layer perceptron (MLP) with shared weights. After the dimensionality reduction - dimensionality increase operation (compression ratio r = 16), generate the channel weight matrix through element-wise addition and Sigmoid activation Enable the model to adaptively enhance the response intensity of the defect-related channels:

[0078] M c = σ(MLP(v avg ) + MLP(v max ))

[0079] Among them, σ represents the Sigmoid function, MLP is the multi-layer perceptron, and r is the channel compression ratio of the MLP, which is responsible for controlling the number of parameters and the computational complexity.

[0080] Specifically, the spatial attention module in this embodiment includes:

[0081] Perform maximum pooling and average pooling respectively along the channel dimension on the channel attention output to generate a two-channel spatial feature map After splicing the two-channel features, generate the spatial weight matrix through a 7×7 convolutional layer (activation function is Sigmoid)

[0082] M s = σ(f 7×7 (S max ; S avg ))

[0083] Among them, σ represents the Sigmoid function, f 7×7 represents the 7×7 convolutional layer, and S max , S avg represent the spatial feature maps after maximum pooling and average pooling respectively. This operation focuses on the geometric distribution characteristics of the defect area and can increase the signal-to-noise ratio (SNR) of the target area by 4.7dB in a strong noise background. The final output feature is Realize the cascaded attention optimization in the channel - spatial dimension.

[0084] As a preferred embodiment of the present invention, the coordinate attention mechanism is preferentially used in the shallow layer of the encoder (resolution ≥ 128×128) to accurately locate the defect boundary coordinates, and the CBAM attention mechanism is adopted in the deep layer of the decoder (resolution ≤ 64×64) to enhance the channel sensitivity to minute defects. The coordinate attention mechanism and the CBAM attention mechanism are adaptively fused through a gating mechanism.

[0085] As Figure 3 shown, as a preferred embodiment of the present invention, the image segmentation model further includes: input feature map processing; the input feature map first undergoes 1×1 convolution for channel expansion to map low-dimensional features to a high-dimensional space to enhance the expression ability, and then the lightweight h-swish activation function is applied to reduce the computational overhead while retaining the non-linear characteristics; then depthwise separable convolution is performed, spatial features are extracted through per-channel convolution and the number of parameters is greatly reduced, and then the channel weights are dynamically calibrated through the coordinate attention module to strengthen the response of the key feature channels; finally, 1×1 convolution is used to compress the high-dimensional features to the target output channel number and a residual connection is made with the original input to form an inverted residual structure.

[0086] Shallow feature extraction; in the shallow feature extraction stage, by optimizing the inverted residual structure of MobileNetV3, a 5×5 depthwise separable convolution with an expansion factor of 2 is used to replace the standard 3×3 convolution,

[0087] the computational amount and receptive field, and the improved convolution kernel satisfies the following relational expression:

[0088]

[0089] where K = 5 is the convolution kernel size, C in is the number of channels of the input feature map, C out is the number of channels of the output feature map, H·W is the spatial resolution of the input feature map, G is the number of groups of the depthwise separable convolution. Compared with the standard 3×3 convolution, the computational amount is reduced by about 28% while the receptive field is increased from 3×3 to 5×5.

[0090] Layer normalization; layer normalization (LayerNorm) is introduced into the inverted residual structure, so that the fluctuation range of the gradient L2 norm during the training process decreases from ±1.7e -3 to ±0.4e -3 .

[0091] The h-swish activation function introduced in this embodiment in MobileNetV3 retains the smooth characteristics of the Swish activation function while reducing the computational complexity by introducing ReLU6 and retaining the gradient continuity characteristics in the zero neighborhood. Through the truncation characteristic of ReLU6 (x ∈ [0, 6]), the activation value can be mapped to the 8-bit integer range (0 - 255) during the model deployment stage, reducing the quantization error. h-swish retains non-zero gradients in the interval x ∈ [-3, 0]:

[0092]

[0093] where ReLU6 = min(6, max(0, x))

[0094] where x represents the input feature value (the input signal of the neuron), that is, the input variable of the activation function, min() represents finding the minimum value, and max() represents finding the maximum value. h-swish alleviates the neuron death problem compared to ReLU.

[0095] As Figure 4 shown, the coordinate attention mechanism includes: coordinate information embedding and coordinate attention generation. In the coordinate information embedding stage, the model compresses the spatial dimension (H×W) of the feature map into two independent one-dimensional vectors through global pooling operations, respectively representing the global information in the height and width directions. Assume the input feature map is X ∈ RC×H×W, where C is the number of channels, and H and W are the height and width respectively. The pooling operation in the height direction can be expressed as:

[0096]

[0097] The pooling operation in the width direction can be expressed as:

[0098]

[0099] where X represents the feature map, H represents the height, W represents the width, z h represents the average pooling of the feature map along the height direction, z ω represents the average pooling of the feature map along the width direction, and the two one-dimensional vectors obtained

[0100] In the coordinate attention generation stage, the two one-dimensional vectors generate attention weights through convolution and non-linear transformation, and act on the height and width dimensions of the feature map respectively. Specifically, first concatenate z h and z ω , and then generate attention weights through a 1×1 convolutional layer and a Sigmoid non-linear activation function:

[0101] A h = σ(Fh (z h ))),A ω =σ(F ω (z ω ))

[0102] where F h represents the convolution operation in the height direction, and F ω represents the convolution operation in the width direction, and σ is the Sigmoid function. Finally, the attention weights A h and A ω are respectively multiplied by the input feature map to generate the weighted feature map:

[0103] Y(h,ω) = X(h,ω)·A h (h)·A ω (ω).

[0104] In other embodiments, the coordinate attention mechanism includes: compressing the spatial dimension (H×W) of the feature map into two independent one-dimensional vectors through global pooling operations, respectively representing the global information in the height and width directions; generating attention weights by convolution and non-linear transformation using the two one-dimensional vectors, and respectively acting on the height and width dimensions of the feature map.

[0105] As a preferred embodiment of the present invention, the multi-scale feature fusion module in step S21 includes:

[0106] S211, fusing the feature maps from the RGB data MobileNetV3-CA network branch and the polarization data MobileNetV3-CA network branch by means of pixel addition;

[0107] S212, inputting the fused feature map into the atrous spatial pyramid pooling module (ASPP) of DeepLabV3+;

[0108] S213, applying the CBAM attention mechanism again in the atrous spatial pyramid pooling module and the encoder shallow feature extraction network respectively.

[0109] Furthermore, in this embodiment, before step S211, there is also an L2 normalization process for increasing the channel dimension to balance the contribution degrees of the bimodal features:

[0110]

[0111] where F RGB represents the feature map of the RGB image, F Pol represents the feature map of the polarization image, and F fusion represents the fused feature map.

[0112] Optionally, in this embodiment, step S212 includes: adopting a progressive combination of dilation rates (rates = [6, 12, 18]), and the calculation formula for its maximum receptive field is:

[0113]

[0114] where RF represents the receptive field size of the final layer, L represents the depth of the current layer in the neural network, K l represents the convolutional kernel size of the first layer, and S i represents the stride of the i-th layer. When L = 3, the diameter of the effectively covered defective area reaches 256 pixels.

[0115] The present invention also provides a guide rail surface defect detection system based on an image segmentation model, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the guide rail surface defect detection method based on an image segmentation model as described in any one of the above. In other embodiments, as Figure 6 shown, a guide rail surface defect detection system based on an image segmentation model includes a cross-modal data acquisition module, a model construction module, an iterative training module, and an image segmentation module.

[0116] Among them, the cross-modal data acquisition module is used to acquire RGB images and polarization images containing guide rail surface defects, and after strict screening by a professional quality inspection team based on dimensions such as image clarity, defect integrity, and annotation feasibility, a cross-modal dataset is obtained;

[0117] The model construction module is used to optimize the DeepLabv3+ network model, extract the depth information of the guide rail area, and segment the defective area of the guide rail area;

[0118] The iterative training module is used to perform a new round of iterative training on the image segmentation model according to the updated model parameters until the image segmentation model obtained from the last round of iterative training reaches a convergence state;

[0119] The image segmentation module is used to pass the image to be segmented through the trained image segmentation model to obtain the segmentation result of the image to be segmented.

[0120] The present invention also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it implements the steps of the guide rail surface defect detection method based on an image segmentation model as described in any one of the above.

[0121] Figure 7This is a schematic diagram of the hardware structure provided by an embodiment of the present invention for running a guide rail surface defect detection method based on an image segmentation model. As Figure 7 shown, this embodiment / computer 6 includes: a processor 60, a memory 61, and a computer program 62 stored in the memory 61 and executable on the processor 60, such as a program for running a guide rail surface defect detection method based on an image segmentation model. When the processor 60 executes the computer program 62, it implements the steps in each of the above embodiments of the guide rail surface defect detection method based on an image segmentation model. Alternatively, when the processor 60 executes the computer program 62, it implements the functions of each module / unit in each of the above device embodiments.

[0122] Exemplarily, the computer program 62 can be divided into one or more modules / units. The one or more modules / units are stored in the memory 61 and executed by the processor 60 to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 62 in the computer 6.

[0123] The computer 6 can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer 6 device may include, but is not limited to, a processor 60 and a memory 61. Those skilled in the art can understand that Figure 7 this is only an example of the computer 6 and does not constitute a limitation on the computer 6. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the computer 6 may further include input and output devices, network access devices, a bus, etc.

[0124] The so-called processor 60 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc.

[0125] The memory 61 may be an internal storage unit of the computer 6, such as the hard disk or memory of the computer 6. The memory 61 may also be an external storage device of the computer 6, such as a plug-in hard disk equipped on the terminal device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 61 may also include both the internal storage unit of the computer 6 and external storage devices. The memory 61 is used to store the computer program and other programs and data required by the terminal device. The memory 61 may also be used to temporarily store data that has been output or will be output.

[0126] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In practical applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment and will not be elaborated herein.

[0127] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0128] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0129] In the embodiments provided by the present invention, it should be understood that the disclosed device / terminal device and method can be implemented in other ways. For example, the device / terminal device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.

[0130] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0131] In addition, each functional unit in various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0132] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above method embodiments of the present invention, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0133] The above are only specific application examples of the present invention, which do not constitute any limitation to the protection scope of the present invention. In addition to the above embodiments, the present invention may also have other implementation manners. All technical solutions formed by equivalent substitution or equivalent transformation fall within the scope of protection required by the present invention.

Claims

1. A method for detecting rail surface defects based on an image segmentation model, characterized in that Including the steps: S1. Obtain the RGB data and polarization data of the guide rail surface, and form a cross-modal dataset; S2. Construct an image segmentation model, including: S21. Extract the features of different modal data in the cross-modal dataset through a multi-layer convolutional network. The features include the light intensity features extracted from the RGB data and the polarization features extracted from the polarization data; then use a multi-scale feature fusion module to fuse the features; S22. Based on the fused features, introduce the CBAM attention mechanism, construct an improved DeepLabv3+ network model to realize the construction of the image segmentation model, and perform training iteration on the image segmentation model until the image segmentation model reaches a convergence state; S3. Use the image segmentation model that has reached the convergence state to segment the image to be segmented on the guide rail surface to obtain a segmentation result; based on the segmentation result, obtain the detection result of the defects on the guide rail surface.

2. The method for detecting rail surface defects based on an image segmentation model according to claim 1, wherein The step S21 includes: Adopt MobileNetV3-CA as the backbone network architecture of the image segmentation model for network improvement; Adopt the MobileNetV3-CA structure in the RGB data branch to extract the color and texture features of the guide rail surface in the RGB data; Adopt the MobileNetV3-CA structure in the polarization data branch to extract the material and polarization features of the guide rail surface in the polarization data; Adopt h-swish as the activation function of the image segmentation model; adopt the coordinate attention mechanism as the attention mechanism of the image segmentation model.

3. The rail surface defect detection method based on an image segmentation model according to claim 2, wherein The CBAM attention mechanism includes: A channel attention module that extracts the global statistical information of each channel through global average pooling and global max pooling respectively, and learns the weights of each channel; A spatial attention module that obtains the maximum value and average value of each spatial position through max pooling and average pooling respectively, and learns the weight distribution of each spatial position.

4. The method for detecting guide rail surface defects based on an image segmentation model according to claim 2, wherein, The formula of the h-swish activation function is: ReLU6 = min(6, max(0, x)) where x represents the input feature, min() represents finding the minimum value, and max() represents finding the maximum value.

5. The method for detecting rail surface defects based on an image segmentation model according to claim 2, wherein, The coordinate attention mechanism includes: Compress the spatial dimension (H×W) of the feature map into two independent one-dimensional vectors through global pooling operations, respectively representing the global information in the height and width directions; Generate attention weights through convolution and non-linear transformation using the two one-dimensional vectors, and act on the height and width dimensions of the feature map respectively.

6. The method for detecting rail surface defects based on an image segmentation model according to claim 5, wherein, The formula of the two independent one-dimensional vectors is: Among them, X represents the feature map, H represents the height, W represents the width, and z h represents average pooling of the feature map along the height direction, and z ω represents average pooling of the feature map along the width direction.

7. A method for detecting rail surface defects based on an image segmentation model according to claim 2, characterized in that Use the coordinate attention mechanism preferentially in the shallow layer of the encoder and the CBAM attention mechanism in the deep layer of the decoder. The coordinate attention mechanism and the CBAM attention mechanism are adaptively fused through a gating mechanism.

8. A method for detecting rail surface defects based on an image segmentation model according to claim 2, wherein, The multi-scale feature fusion module in the step S21 includes: S211. Fuse the feature maps from the MobileNetV3-CA network branch of the RGB data and the MobileNetV3-CA network branch of the polarization data in a pixel addition manner; S212, the fused feature map is input into the atrous spatial pyramid pooling module of DeepLabV3+; S213, the CBAM attention mechanism is applied again to the atrous spatial pyramid pooling module and the encoder shallow feature extraction network respectively.

9. A guide rail surface defect detection system based on an image segmentation model, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the rail surface defect detection method based on the image segmentation model according to any one of claims 1 to 8 are implemented.

10. A computer-readable storage medium stores a computer program, characterized in that, When the computer program is executed by the processor, the steps of the rail surface defect detection method based on the image segmentation model according to any one of claims 1 to 8 are implemented.