A train body surface disease detection method based on multi-source data fusion
Patent Information
- Application Number
- CN202610785921.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-02
- Publication Date
- 2026-09-08
AI Technical Summary
[0005]为此,本发明提供一种多源数据融合的列车车体表面病害检测方法,用以克服现有技术中单一可见光模态易受光照变化及阴影遮挡影响,检测方法集中于轨道表面而缺乏对车体病害有效检测的问题
[0015] Compared with the prior art, the beneficial effects of the present invention are that by simultaneously acquiring visible light images and depth images to construct a multimodal dataset, and introducing complementary enhancements of depth geometric information and visible light semantic information, the present invention solves the problem that the detection accuracy is insufficient due to changes in illumination and shadow occlusion of a single visible light modality, thereby improving the detection accuracy and robustness under complex working conditions.
Smart Images

Figure CN122714331A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of rail transit operation and maintenance technology, and in particular to a method for detecting defects on the surface of train bodies through multi-source data fusion. Background Technology
[0002] As the core moving carrier of the rail transit system, the train body exhibits a highly integrated structural system, encompassing key components such as bogies, couplers, and body frames. With increasing service life, these critical components will gradually exhibit typical defects such as fatigue cracks and connection failures, directly jeopardizing operational safety. Traditional methods for detecting surface defects in train bodies rely primarily on manual inspection, which suffers from low efficiency, high labor intensity, and strong subjectivity, making it difficult to meet current inspection needs. Existing automated inspection technologies are mostly based on two-dimensional image analysis, using image processing and deep learning algorithms to identify and locate defects. However, in engineering applications, such as entrance / exit tunnels and nighttime operations, lighting conditions are complex and variable. Data acquired using a single visible light camera lacks stability in quality, and key defect features are easily interfered with, obscured, or even lost, thus weakening the reliability of the algorithm.
[0003] Chinese Patent Publication No. CN115866155A discloses a method and apparatus for processing high-speed rail maintenance data using a fusion algorithm. The method includes frame synchronization of infrared and visible light detection image data acquired during high-speed rail maintenance. After the two types of image data are matched, a fusion algorithm is used to fuse the multi-layered image data. This approach effectively avoids image fusion misalignment while preserving more texture features, edge and structural information from the source images, thus improving the image fusion effect.
[0004] The existing technology still has the following problems: Although the fusion processing of infrared images and visible light images has been achieved, the fusion method is image-level superposition and enhancement. It does not design a self-learning fusion network for the heterogeneous characteristics of visible light mode and depth mode in feature space. It is difficult to make full use of the spatial geometric information contained in the depth image to perform feature-level complementary enhancement of the visible light image. This results in insufficient feature extraction capability for defects under complex lighting conditions, thus leading to poor detection accuracy and robustness of defects on the surface of the train body. Summary of the Invention
[0005] To address this, the present invention provides a method for detecting defects on the surface of train bodies using multi-source data fusion, which overcomes the problems of existing technologies where single visible light modes are easily affected by changes in illumination and shadow occlusion, and where detection methods focus on the track surface but lack effective detection of defects in the train body.
[0006] To achieve the above objectives, the present invention provides a method for detecting defects on the surface of a train body based on multi-source data fusion, comprising: Visible light images and depth images of the train body surface are simultaneously acquired using a multimodal synchronous acquisition device installed on a gantry above the track, and a multimodal dataset of train body surface defects is constructed. A multimodal detection model for defects on the surface of a train body is constructed. The multimodal detection model includes a multimodal feature extraction network, a multimodal feature fusion network, and a head network. Based on the multimodal dataset of defects on the train body surface, the multimodal detection model is trained using a weighted composite equilibrium loss function; The trained multimodal detection model was used to detect defects on the surface of the train body, and the detection results were obtained.
[0007] Furthermore, the multimodal synchronous acquisition device consists of a structured light camera, a linear laser, and a support frame. The laser line emitted by the laser is perpendicularly projected onto the train body surface, and the structured light camera captures the laser line to obtain a laser cross-sectional image of the train body surface. The structured light camera acquires these laser cross-sectional images of the train body surface at equal intervals during train movement. The acquired cross-sectional images are then stitched together to generate visible light and depth images, respectively. The multimodal dataset of train body surface defects is divided into a training set and a validation set in an 8:2 ratio.
[0008] Furthermore, the multimodal feature extraction network includes a visible light feature extraction subnetwork and a depth feature extraction subnetwork with identical structures, used to independently extract multi-scale features from visible light images and depth images, respectively. Both the visible light feature extraction subnetwork and the depth feature extraction subnetwork consist of a convolutional module, four downsampling modules, and a pooling module cascaded sequentially. The downsampling module is composed of a convolutional submodule and a feature extraction submodule connected in series, and the feature extraction submodule consists of convolutional units and bottleneck units. The pooling module consists of a convolutional submodule, three max-pooling submodules, a channel concatenation submodule, and another convolutional module cascaded sequentially.
[0009] Furthermore, the multimodal feature fusion network includes three self-learning deep feature fusion modules, a high-level semantic transmission module, and a low-level detail enhancement module. The inputs to the three self-learning deep feature fusion modules are the visible light feature maps and depth feature maps of the shallow, middle, and deep layers of the multimodal feature extraction network, respectively, which are used to achieve multimodal feature complementarity and information gain, and the output is a fused feature map after gain at three different scales.
[0010] Furthermore, the processing flow of the self-learning deep feature fusion module includes: constructing a dual-branch extraction architecture for visible light features, including a variance branch and a main branch. The variance branch is used to statistically analyze the spatial variance of different channels of the visible light features to generate variance features. The main branch integrates adaptive max pooling and depthwise separable convolution to expand the receptive field while reducing spatial resolution, generating main branch features. The variance features and the main branch features are weighted and fused using learnable scaling weights to obtain hybrid visible light features. The hybrid visible light features are modulated by generating a modulation mask using ordinary convolution, GeLU activation function, and upsampling operation, and multiplied pixel-by-pixel with the original visible light image features to obtain representative visible light features. The original depth features are modeled with spatial hierarchical information and processed sequentially by depthwise separable convolution, ordinary convolution, GeLU activation function, and ordinary convolution to obtain representative depth features. The representative visible light feature and the representative depth feature are concatenated through channels and processed by ordinary convolution to obtain the first fused feature. The original visible light feature and the original depth feature are concatenated through channels and processed by ordinary convolution to obtain the second fused feature. The first fused feature and the second fused feature are added proportionally by adjusting the factor to obtain the final fused output.
[0011] Furthermore, the high-level semantic transmission module is located after the three self-learning deep feature fusion modules, and is composed of several convolutional sub-modules, upsampling sub-modules, channel concatenation sub-modules, and feature extraction sub-modules connected in series alternately. The high-level semantic transmission module performs progressive upsampling and channel-dimensional concatenation on the outputs of the three self-learning deep feature fusion modules from deep to shallow layers, realizing the step-by-step transmission of high-level semantic information.
[0012] Furthermore, the low-level detail enhancement module is located after the high-level semantic transmission module and is composed of several convolutional sub-modules, channel concatenation sub-modules, and feature extraction sub-modules connected in series alternately. The low-level detail enhancement module further enhances the features of the output of the high-level semantic transmission module by concatenating and fusing low-level detail information through channel-dimensional concatenation.
[0013] Furthermore, the weighted composite equilibrium loss function is a weighted combination of the quality focus loss function and the adaptive boundary loss function, used to simultaneously optimize the classification quality of continuous labels and the classification boundary of minority class samples. The quality focus loss function extends the traditional focus loss to the continuous label space, introducing the absolute difference between the predicted probability and the actual label as a dynamic modulation factor. The adaptive boundary loss function, for imbalanced datasets, assigns a larger classification boundary to the minority class. The quality focus loss function and the adaptive boundary loss function are fused through dynamic weight allocation, and the balancing hyperparameter of the dynamic weights is determined using a grid search method.
[0014] Furthermore, the model performance was evaluated using precision, recall, and mean precision during training. The training results were compared with a single-modal train body surface defect detection model, and the multimodal detection model outperformed the comparison model in all three metrics: precision, recall, and mean precision.
[0015] Compared with the prior art, the beneficial effects of the present invention are that by simultaneously acquiring visible light images and depth images to construct a multimodal dataset, and introducing complementary enhancements of depth geometric information and visible light semantic information, the present invention solves the problem that the detection accuracy is insufficient due to changes in illumination and shadow occlusion of a single visible light modality, thereby improving the detection accuracy and robustness under complex working conditions.
[0016] Furthermore, this invention performs dual-branch extraction and weighted fusion of visible light features and depth features through a self-learning deep feature fusion module, and introduces an adjustment factor to dynamically balance the contribution ratio of the original features and the fused features, thereby realizing deep coupling of heterogeneous modal information and significantly improving the model's ability to extract disease features.
[0017] Furthermore, this invention dynamically fuses the quality focus loss function and the adaptive boundary loss function through a weighted composite equilibrium loss function, while simultaneously optimizing the classification quality of continuous labels and the classification boundary of minority class samples. This solves the problem of insufficient sample attention in traditional loss functions, thereby significantly reducing the missed detection rate of defects while ensuring real-time inspection efficiency, and realizing high-precision and high-reliability automated detection of defects on train body surfaces. Attached Figure Description
[0018] Figure 1 This is a flowchart of a train body surface defect detection method based on multi-source data fusion according to an embodiment of the present invention; Figure 2 This is a structural diagram of the multimodal detection model according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the downsampling module according to an embodiment of the present invention; Figure 4 This is a comparison image showing the effect before and after generating a sample of missing small components of an urban rail train according to an embodiment of the present invention; Figure 5 This is a schematic diagram of the pooling module structure according to an embodiment of the present invention; Figure 6 This is a structural diagram of the self-learning deep feature fusion module in an embodiment of the present invention. Detailed Implementation
[0019] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.
[0020] Please see Figure 1 As shown, it is a flowchart of the train body surface defect detection method based on multi-source data fusion according to an embodiment of the present invention.
[0021] The method for detecting defects on the surface of a train body based on multi-source data fusion according to embodiments of the present invention includes: Step S1: Use a multimodal synchronous acquisition device installed on the gantry above the track to synchronously acquire visible light images and depth images of the train body surface to construct a multimodal dataset of train body surface defects. Step S2: Construct a multimodal detection model for defects on the train body surface. The multimodal detection model includes a multimodal feature extraction network, a multimodal feature fusion network, and a head network. Step S3: Based on the multimodal dataset of defects on the train body surface, train the multimodal detection model using a weighted composite equilibrium loss function; Step S4: Use the trained multimodal detection model to detect defects on the surface of the train body and obtain the detection results.
[0022] Specifically, the multimodal synchronous acquisition device consists of a structured light camera, a linear laser, and a support frame. The laser beam emitted by the laser is perpendicularly projected onto the surface of the train body, and the structured light camera captures a cross-sectional image of the laser beam on the train body surface. The structured light camera acquires these cross-sectional images at equal intervals along the train's movement, obtaining visible light and depth information about the surface. The acquired cross-sectional images are then stitched together to generate a visible light image reflecting two-dimensional features and a depth image reflecting three-dimensional features, with pixels in the visible light and depth images matched one-to-one.
[0023] The labeling tool Labeling was used to label the surface defects of the train body in each pair of images, with the label category being loose roof fasteners. A multimodal train body surface defect dataset was established, and the image pairs in the dataset were divided into training and validation sets in an 8:2 ratio.
[0024] Please see Figure 2 As shown, it is a network structure diagram of the multimodal detection model in an embodiment of the present invention.
[0025] Specifically, the multimodal feature extraction network includes a visible light feature extraction subnetwork and a depth feature extraction subnetwork with identical structures, used to independently extract multi-scale features from visible light images and depth images, respectively. Both the visible light feature extraction subnetwork and the depth feature extraction subnetwork consist of a convolutional module, four downsampling modules, and a pooling module cascaded sequentially.
[0026] Please see Figure 3 and Figure 4 As shown, Figure 3 This is a schematic diagram of the downsampling module structure according to an embodiment of the present invention. Figure 4 This is a schematic diagram of the feature extraction submodule structure in an embodiment of the present invention.
[0027] The downsampling module is used for downsampling and efficient feature extraction, and is composed of a convolutional submodule and a feature extraction submodule connected in series. The feature extraction submodule is used to increase the network depth and receptive field, thereby improving the feature extraction capability, and is composed of convolutional units and bottleneck units.
[0028] Please see Figure 5 As shown, it is a schematic diagram of the pooling module structure in an embodiment of the present invention.
[0029] The pooling module is used to achieve feature aggregation and enhance the network's ability to understand the global structure. It consists of a convolutional submodule, three max pooling submodules, a channel splicing submodule, and another convolutional submodule that are cascaded in sequence.
[0030] Visible light images and depth images are input into the visible light feature extraction subnetwork and the depth feature extraction subnetwork, respectively, and are processed by a convolutional module to output visible light feature maps. and depth feature map The visible light feature map and depth feature map are processed by four downsampling modules in their respective feature extraction networks, and the visible light feature map is output sequentially. , , , and depth feature map , , , Visible light characteristic map and depth feature map Each feature map is processed by a pooling module and then output separately. and .
[0031] Please see Figure 6 As shown, it is a structural diagram of the self-learning deep feature fusion module in an embodiment of the present invention.
[0032] Specifically, the multimodal feature fusion network includes three self-learning deep feature fusion modules, one high-level semantic transmission module, and one low-level detail enhancement module. The inputs to the three self-learning deep feature fusion modules are the visible light feature maps and depth feature maps of the shallow, middle, and deep layers of the multimodal feature extraction network, respectively, used to achieve multimodal feature complementarity and information gain. The outputs are fused feature maps after gain at three different scales. The visible light feature maps and depth feature maps of the shallow, middle, and deep layers are respectively... , and The fused feature maps after gain at the three different scales are as follows: , and .
[0033] Taking a shallow feature map as an example, the processing flow of the self-learning deep feature fusion module includes: constructing a dual-branch extraction architecture for visible light features, including a variance branch and a main branch. The variance branch is used to statistically analyze the spatial variance of different channels of the visible light features, quantify the dispersion of pixel values within a channel, and generate variance features. :
[0034] Where Var is the channel variance calculation and F1 is the shallow visible light feature map.
[0035] The main branch integrates adaptive max pooling and depthwise separable convolution, which expands the receptive field while reducing spatial resolution, and generates main branch features. :
[0036] in, For a depthwise separable convolution with a size of 3×3, This is for adaptive max pooling.
[0037] Learnable scaling weights and Dynamically balancing the contributions of the two branches, the variance feature and the main branch feature are weighted and fused to obtain a hybrid visible light feature. :
[0038] in, and It can be optimized through end-to-end training and adaptive updates.
[0039] The hybrid visible light features are modulated using ordinary convolution, GeLU activation function, and upsampling operation to generate a modulation mask M.
[0040] in, For upsampling operation, This is the GeLU activation function. This is a regular convolution with a size of 1×1.
[0041] The modulation mask M is multiplied pixel by pixel with the original visible light image features to obtain representative visible light features. :
[0042] Spatial hierarchical information modeling is performed on the original depth features, and then processed sequentially through depthwise separable convolution, ordinary convolution, GeLU activation function, and ordinary convolution to obtain representative depth features. :
[0043] The representative visible light features and the representative depth features are concatenated by channels and then processed by ordinary convolution to obtain the first fused feature. :
[0044] The original visible light features and the original depth features are concatenated by channels and then processed by ordinary convolution to obtain the second fused feature. :
[0045] in, For channel splicing.
[0046] By introducing an adjustment factor, the first fusion feature and the second fusion feature are added proportionally to obtain the final fusion output. :
[0047] Where λ is the adjustment factor.
[0048] Specifically, the high-level semantic transmission module is located after the three self-learning deep feature fusion modules and is used to execute the feature fusion strategy. The high-level semantic transmission module is composed of the following sub-modules connected in series: a first convolution sub-module, a first upsampling sub-module, a first channel concatenation sub-module, a first feature extraction sub-module, a second convolution sub-module, a second upsampling sub-module, a second channel concatenation sub-module, and a second feature extraction sub-module.
[0049] The first convolutional submodule Perform convolution processing to output feature maps. The first upsampling submodule for Upsampling is performed. The first channel splicing submodule will... Compared with the upsampling process The feature maps are concatenated along the channel dimension. The first feature extraction submodule processes the feature maps concatenated by the first channel concatenation submodule and outputs the feature map. .
[0050] The second convolutional submodule Perform convolution processing to output feature maps. The second upsampling submodule for Upsampling is performed. The second channel splicing submodule will... Compared with the upsampling process The feature maps are concatenated along the channel dimension. The second feature extraction submodule processes the feature maps concatenated by the second channel concatenation submodule and outputs the feature map. .
[0051] Specifically, the low-level detail enhancement module is located after the high-level semantic transmission module and is used to execute feature enhancement strategies. The low-level detail enhancement module is composed of the following sub-modules connected in series: a third convolution sub-module, a third channel splicing sub-module, a third feature extraction sub-module, a fourth convolution sub-module, a fourth channel splicing sub-module, and a fourth feature extraction sub-module.
[0052] The third convolutional submodule Perform convolution processing to output feature maps. The third channel splicing submodule will and The feature maps are concatenated along the channel dimension. The third feature extraction submodule processes the feature maps concatenated by the third channel concatenation submodule and outputs the feature map. .
[0053] The fourth convolutional submodule Perform convolution processing to output feature maps. The fourth channel splicing submodule will and The feature maps are concatenated along the channel dimension. The fourth feature extraction submodule processes the feature maps concatenated by the fourth channel concatenation submodule and outputs the feature map. .
[0054] Specifically, the head network includes three detection heads, each receiving... , and It is used to identify and locate surface defects on train bodies of three different sizes: large, medium, and small.
[0055] Specifically, the training method of the multimodal detection model is as follows: the training set in the multimodal dataset of defects on the train body surface is input into the multimodal detection model, and training is performed through the weighted composite equilibrium loss function.
[0056] During model training, precision (P), recall (R), and mean precision (mAP) are used to evaluate model performance. Precision (P) represents the proportion of correctly detected targets, and recall (R) represents the proportion of successfully detected targets among all targets. Mean precision (mAP) is the average of mean precision (AP), where AP represents the average precision with which the model identifies a specific target. This displays the mAP value, representing the average mAP at various intersection-union (IU) ratios (Intersection over Unions) ranging from 0.5 to 0.95 in increments of 0.05. The mAP is calculated as follows:
[0057]
[0058]
[0059]
[0060] Wherein, TP represents true positive, FP represents false positive, TN represents true negative, FN represents false negative, and n represents the category of the sample being tested.
[0061] Specifically, the weighted composite equilibrium loss function is a weighted combination of the quality focus loss function and the adaptive boundary loss function, used to simultaneously optimize the classification quality of continuous labels and the classification boundary of minority class samples.
[0062] Specifically, the quality focus loss function is calculated, extending the traditional focus loss to the continuous label space, and introducing the absolute difference between the predicted probability and the actual label as a dynamic modulation factor to reduce the loss contribution of easily classified samples:
[0063] in, Let y be the quality focus loss function, p be the prediction probability, and β be the modulation factor. Calculate the adaptive boundary loss function, which assigns a larger classification boundary to the class with fewer samples for imbalanced datasets, forcing the model to learn a more robust decision boundary:
[0064] in, Output a score for the model in category y. Let be the minimum boundary from all samples of class y to the decision boundary. For any other category, the predicted score is given. The calculation method is as follows:
[0065] Where C is a hyperparameter. y represents the number of samples for category y.
[0066] The quality focus loss function and the adaptive boundary loss function are fused using dynamic weight allocation to obtain the final weighted composite equilibrium loss function. :
[0067] Here, ε is a hyperparameter used to control the contribution ratio of the two loss functions to the total loss. The optimal value of the balancing hyperparameter ε is determined by a grid search method.
[0068] In this embodiment of the invention, during model training, the operating system used is Ubuntu 20.04, the GPU is an NVIDIA GeForce RTX 4090, and the Python version in the training environment is 3.8.20. The model training cycle is set to 200 times, and the batch size is set to 32. During training, the input images are uniformly scaled to 640×640 pixels, the momentum coefficient is set to 0.937, the weight decay coefficient is set to 0.0005, the initial learning rate is 0.01, and a warm-up learning rate strategy and SGD optimizer are used.
[0069] The training effect was verified using the validation set in the multimodal dataset of train body surface defects, resulting in the trained multimodal detection model. The training results were compared with those of a single-modal train body surface defect detection model, and the comparison results are shown in Table 1.
[0070] Table 1 Comparison of Model Detection Results
[0071] As shown in Table 1, the multimodal detection model outperforms the comparison model in terms of precision, recall, and mean accuracy in the task of detecting defects in roof fasteners. The multimodal detection model effectively balances the false alarm rate and false negative rate through the synergistic effect of a multimodal feature fusion mechanism and a weighted composite equilibrium loss function, exhibiting high detection accuracy and robustness under complex working conditions.
[0072] The trained multimodal detection model was used to detect defects on the surface of the train body, and the detection results were obtained. The detection results of the multimodal detection model were compared with those of other methods. The multimodal detection model, while maintaining high accuracy and recall, also showed better detection confidence than the comparison methods.
[0073] This invention constructs a multimodal feature extraction network to simultaneously extract multi-scale features from visible light and depth images. A self-learning deep feature fusion module performs targeted feature extraction and multi-round interactive fusion of data from different modalities, achieving complementary enhancement and deep coupling of heterogeneous modal information. Furthermore, a weighted composite equilibrium loss function dynamically adjusts the model optimization direction, effectively improving the detection performance of surface defects on train bodies. This invention provides a reliable technical solution for the intelligent inspection of surface defects on train bodies, which is of great significance for improving the operation and maintenance efficiency of rail transit systems and ensuring their safe operation.
[0074] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.
Claims
1. A method for detecting defects on the surface of a train body using multi-source data fusion, characterized in that, include: Visible light images and depth images of the train body surface are simultaneously acquired using a multimodal synchronous acquisition device installed on a gantry above the track, and a multimodal dataset of train body surface defects is constructed. A multimodal detection model for defects on the surface of a train body is constructed. The multimodal detection model includes a multimodal feature extraction network, a multimodal feature fusion network, and a head network. Based on the multimodal dataset of defects on the train body surface, the multimodal detection model is trained using a weighted composite equilibrium loss function; The trained multimodal detection model was used to detect defects on the surface of the train body, and the detection results were obtained.
2. The method for detecting defects on the surface of a train body using multi-source data fusion according to claim 1, characterized in that, The multimodal feature extraction network includes a visible light feature extraction subnetwork and a depth feature extraction subnetwork with identical structures. The visible light feature extraction subnetwork is used to independently extract multi-scale features of the visible light image, and the depth feature extraction subnetwork is used to independently extract multi-scale features of the depth image.
3. The method for detecting surface defects of train bodies using multi-source data fusion according to claim 1, characterized in that, The multimodal feature fusion network includes a self-learning deep feature fusion module, which is used to perform feature-level fusion of the visible light feature map and the depth feature map output by the multimodal feature extraction network, and output a fused feature map after gain.
4. The method for detecting defects on the surface of a train body using multi-source data fusion according to claim 3, characterized in that, The self-learning deep feature fusion module extracts the spatial variance features and main branch features of visible light features through a dual-branch architecture, performs weighted fusion of the dual-branch outputs to obtain mixed visible light features, and performs modulation processing to obtain representative visible light features. Representative depth features are obtained by modeling the spatial hierarchical information of the depth features; The representative visible light features and the representative depth features are fused together, and the results of the direct fusion of the original visible light features and the original depth features are weighted by an adjustment factor to obtain the fused output.
5. The method for detecting surface defects of train bodies using multi-source data fusion according to claim 3, characterized in that, The multimodal feature fusion network also includes a high-level semantic transmission module and a low-level detail enhancement module. The high-level semantic transmission module is used to perform progressive upsampling and channel splicing on the multi-scale fused feature map from deep to shallow layers, and the low-level detail enhancement module is used to enhance the features of the transmitted feature map.
6. The method for detecting surface defects of train bodies using multi-source data fusion according to claim 1, characterized in that, The weighted composite equilibrium loss function is composed of a quality focus loss function and an adaptive boundary loss function, which are dynamically weighted and combined.
7. The method for detecting surface defects of train bodies using multi-source data fusion according to claim 6, characterized in that, The quality focus loss function introduces the absolute difference between the predicted probability and the actual label as a dynamic modulation factor to reduce the loss contribution of easily classified samples.
8. The method for detecting surface defects of train bodies using multi-source data fusion according to claim 6, characterized in that, The adaptive boundary loss function is designed for imbalanced datasets, assigning larger classification boundaries to classes with fewer samples.
9. The method for detecting surface defects of train bodies using multi-source data fusion according to claim 6, characterized in that, The dynamic weights are determined using a grid search method.
10. The method for detecting surface defects of train bodies using multi-source data fusion according to claim 1, characterized in that, The head network includes multiple detection heads, which respectively receive feature maps of different scales output by the multimodal feature fusion network, for identifying and locating defects on the surface of train bodies of different sizes.
Citation Information
Patent Citations
Method and device for processing high-speed rail maintenance data by using fusion algorithm
CN115866155A