Multi-source data fusion steel rail surface damage detection method

Through the multi-source data fusion method, combined with the detection model of texture and depth feature extraction network, the misjudgment and reliability problems of rail surface damage detection in the prior art are solved, and the detection effect of high-precision and low false alarm rate is achieved.

CN119992157APending Publication Date: 2025-05-13BEIJING MASS TRANSIT RAILWAY OPERATION CORPORATION LIMITED
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411822483.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-11
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing rail surface damage detection methods are prone to misjudgment of stains, shadows, etc., and are not highly reliable.

Method used

Using the multi-source data fusion method, high-precision detection of rail surface damage is achieved by constructing a multi-source data set containing rail grayscale maps and depth maps, combining a symmetric texture feature extraction network and a deep feature extraction network, as well as detection models of multiple cross-modal feature fusion modules, multi-scale feature fusion modules and decoupled detection heads.

Benefits of technology

It effectively reduces the interference of shadows, stains, etc. on detection, improves the accuracy and reliability of detection, reduces false alarm rates, and improves detection efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992157A_ABST
    Figure CN119992157A_ABST
Patent Text Reader

Abstract

The invention discloses a steel rail surface damage detection method based on multi-source data fusion, and belongs to the technical field of machine vision, and the method comprises the following steps: constructing a steel rail multi-source data set; wherein the data set comprises a steel rail grey-scale map and a steel rail depth map; constructing a steel rail surface damage detection model; the model comprises a texture feature extraction network, a depth feature extraction network, a plurality of cross-modal feature fusion modules, a multi-scale feature fusion module and a plurality of decoupling detection heads, training the steel rail surface damage detection model by using the steel rail multi-source data set; and obtaining a grey-scale map and a depth map of a to-be-detected steel rail, and inputting the grey-scale map and the depth map of the to-be-detected steel rail into the trained steel rail surface damage detection model to obtain a surface damage detection result of the to-be-detected steel rail. According to the invention, through fusion, interaction and optimization of multi-source data feature levels, accurate detection of steel rail surface damage is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of machine vision, and in particular to a rail surface damage detection method based on multi-source data fusion. Background Art

[0002] Rails are one of the important components of tracks, which guide and carry trains. If the damage on the surface of the rails, such as cracks, peeling, abrasions, corrugation, etc., is not discovered and handled in time, it may lead to serious accidents such as train derailment and rollover, posing a huge threat to passengers. The damage on the surface of the rails will also affect the smooth operation of the train, increase the vibration and noise of the train, and reduce the comfort of passengers. Therefore, detecting the damage on the surface of the rails can timely detect and repair minor damage and prevent it from deteriorating into more serious damage, thereby extending the service life of the rails and preventing accidents.

[0003] With the rapid development of the railway transportation industry, more and more advanced intelligent detection technologies are being applied to rail surface damage detection. Among them, data-driven deep learning target detection methods are widely used because they can improve detection efficiency and accuracy. However, the target detection algorithm based on single-source image data is prone to misjudgment of stains and shadows on the rail surface due to the single data information and lack of three-dimensional depth information. In addition, the detection algorithm under single data is not reliable and has poor stability.

[0004] With the development of sensor technology, multi-source data fusion detection methods have gradually emerged. By integrating data from different sensors, more comprehensive and accurate information can be obtained, reducing the inaccuracy of target position estimation and category estimation. Even if a sensor fails or the data is abnormal, it can be supplemented and verified by data from other sensors to ensure stable operation of the system. However, it is very important to use multi-source data efficiently and scientifically, that is, to achieve effective information extraction and fusion, and to achieve detection tasks efficiently and quickly. Therefore, it is very necessary to develop an efficient multi-source data fusion high-precision rail surface damage detection method. Summary of the invention

[0005] The present invention provides a rail surface damage detection method with multi-source data fusion, so as to solve the technical problems that the existing detection method is prone to misjudgment of stains, shadows, etc. on the rail surface and has low reliability.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0007] In one aspect, the present invention provides a rail surface damage detection method using multi-source data fusion, comprising:

[0008] Constructing a rail multi-source data set; wherein the rail multi-source data set includes a grayscale image and a depth image of the rail; and the grayscale image of the rail carries rail surface damage annotation data;

[0009] Constructing a rail surface damage detection model; wherein the rail surface damage detection model includes: a symmetrical texture feature extraction network and a deep feature extraction network, as well as multiple cross-modal feature fusion modules, a multi-scale feature fusion module and multiple decoupled detection heads;

[0010] Using the rail multi-source data set to train the rail surface damage detection model;

[0011] The grayscale image and depth image of the rail to be detected are obtained, and the grayscale image and depth image of the rail to be detected are input into the trained rail surface damage detection model to obtain the surface damage detection result of the rail to be detected.

[0012] Furthermore, the construction of the rail multi-source dataset includes:

[0013] Deploy multi-source sensors around the rails and calibrate and register the deployed sensors;

[0014] Use the calibrated and registered sensors to collect the grayscale image and depth image of the rail;

[0015] The surface damage of the rails is annotated on the collected grayscale images, and a multi-source rail dataset is constructed using the depth map of the rails and the grayscale images of the annotated rails. The constructed multi-source rail dataset is divided into a training set, a validation set, and a test set according to a preset ratio; the training set is used to test the model, the validation set is used to verify the trained model, and the test set is used to test the trained model.

[0016] Furthermore, the input of the texture feature extraction network is a grayscale image of the rail, which is used to extract a grayscale feature map;

[0017] The input of the depth feature extraction network is the depth map of the rail, which is used to extract the depth feature map;

[0018] The texture feature extraction network has the same structure as the depth feature extraction network, both of which include: multiple convolutional layers, depth-separable convolutional layers, a C3 module, and a SimSPPF module.

[0019] Further, the grayscale feature map includes: a shallow grayscale feature map, a middle grayscale feature map and a deep grayscale feature map; the depth feature map includes: a shallow depth feature map, a middle depth feature map and a deep depth feature map;

[0020] The number of the cross-modal feature fusion modules is three, and the three cross-modal feature fusion modules are respectively used to: fuse the shallow grayscale feature map and the shallow depth feature map to obtain a shallow fusion feature map; fuse the middle grayscale feature map and the middle depth feature map to obtain a middle fusion feature map; and fuse the deep grayscale feature map and the deep depth feature map to obtain a deep fusion feature map.

[0021] Furthermore, the process of implementing feature map fusion by the cross-modal feature fusion module includes:

[0022] The grayscale feature map and the depth feature map with the same receptive field at the same level to be fused are input into the corresponding cross-modal feature fusion module; the cross-modal feature fusion module adds the input grayscale feature map and the depth feature map to obtain a first feature map, which are respectively input into four depth-wise separable convolutions and respectively average pooled to obtain a second feature map, a third feature map, a fourth feature map and a fifth feature map; wherein the convolution kernel sizes of the four depth-wise separable convolutions are 7×7, 9×9, 13×13 and 17×17 respectively;

[0023] The first feature map and the second feature map are added to obtain a sixth feature map; the third feature map is upsampled and added to the sixth feature map to obtain a seventh feature map; the fourth feature map is upsampled and added to the seventh feature map to obtain an eighth feature map; the fifth feature map is upsampled and added to the eighth feature map to obtain a ninth feature map; the sixth feature map, the seventh feature map, the eighth feature map and the ninth feature map are concatenated, and a convolution operation with a convolution kernel of 1×1 is performed to obtain a tenth feature map;

[0024] The tenth feature map is added to the first feature map to obtain a fused feature map of the corresponding level.

[0025] Furthermore, the input of the multi-scale feature fusion module is a shallow fusion feature map, a middle fusion feature map and a deep fusion feature map; the multi-scale feature fusion module includes: multiple depth-separable convolutional layers, a C3 module and an upsampling module.

[0026] Furthermore, the process of the multi-scale feature fusion module processing the input shallow fusion feature map, middle fusion feature map and deep fusion feature map includes:

[0027] After the deep fusion feature map passes through a C3 module, an eleventh feature map is obtained. After the eleventh feature map passes through a depth-separable convolution layer and an upsampling module in sequence, a twelfth feature map having the same resolution as the middle-layer fusion feature map is obtained. The twelfth feature map is spliced ​​with the middle-layer fusion feature map to obtain a thirteenth feature map. After the thirteenth feature map passes through a C3 module, a fourteenth feature map is obtained. After the fourteenth feature map passes through a depth-separable convolution layer and an upsampling module in sequence, a fifteenth feature map having the same resolution as the shallow-layer fusion feature map is obtained. The fifteenth feature map is spliced ​​with the shallow-layer fusion feature map to obtain a sixteenth feature map.

[0028] After the sixteenth feature map passes through a C3 module and a depth-wise separable convolutional layer in sequence, a seventeenth feature map having the same resolution as the fourteenth feature map is obtained. The seventeenth feature map is spliced ​​with the fourteenth feature map to obtain an eighteenth feature map. After the eighteenth feature map passes through a C3 module and a depth-wise separable convolutional layer in sequence, a nineteenth feature map having the same resolution as the eleventh feature map is obtained. The nineteenth feature map is spliced ​​with the eleventh feature map to obtain a twentieth feature map.

[0029] Furthermore, the number of the decoupling detection heads is three, which are respectively used to perform decoupling operations on the sixteenth feature map, the eighteenth feature map, and the twentieth feature map to predict targets of different scales;

[0030] The decoupled detection head includes a layer of C3 module, and the C3 module includes two branches, one of which includes two layers of depth-separable convolutional layers, and the other branch includes a layer of convolution. The outputs of the two branches are merged through a splicing operation to obtain the predicted category, predicted position and predicted confidence corresponding to the corresponding scale target.

[0031] On the other hand, the present invention further provides an electronic device, comprising a processor and a memory; wherein the memory stores at least one instruction, and the instruction is loaded and executed by the processor to implement the above method.

[0032] In yet another aspect, the present invention further provides a computer-readable storage medium, wherein at least one instruction is stored in the storage medium, and the instruction is loaded and executed by a processor to implement the above method.

[0033] The beneficial effects brought about by the technical solution provided by the present invention include at least:

[0034] The present invention can simultaneously and efficiently and scientifically process the texture edge information of the two-dimensional grayscale image and the spatial position information of the three-dimensional depth image, and has higher accuracy and faster detection speed for multiple categories of rail surface damage detection results, and basically does not require human intervention. In addition, the present invention performs fusion at the feature level, and the fusion detection process is relatively simple, which reduces the impact of data redundancy caused by multi-source information and improves detection efficiency. The use of three-dimensional depth information greatly reduces the interference of shadows, stains, etc. on surface damage detection, and reduces the false alarm rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0036] Figure 1 It is a schematic diagram of a rail surface damage detection method using multi-source data fusion provided by an embodiment of the present invention;

[0037] Figure 2 is a schematic diagram of the structure of a cross-modal fusion module provided in an embodiment of the present invention;

[0038] Figure 3 is a framework diagram of a rail surface damage detection model provided by an embodiment of the present invention;

[0039] Figure 4 It is a system block diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0040] In order to make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0041] First of all, it should be noted that in the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "exemplarily" is intended to present the concept in a concrete way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.

[0042] First embodiment

[0043] This embodiment provides a rail surface damage detection method using multi-source data fusion. The implementation principle of the rail surface damage detection method is as follows: Figure 1The method can be implemented by an electronic device, which can be a terminal or a server. Specifically, the execution process of the method includes the following steps:

[0044] S1, constructing a rail multi-source dataset; wherein the rail multi-source dataset includes a grayscale image and a depth image of the rail; and the grayscale image of the rail carries rail surface damage annotation data;

[0045] Specifically, in this embodiment, the process of constructing the rail multi-source dataset includes:

[0046] S11, deploying multi-source sensors around the rails, and calibrating and registering the deployed sensors; wherein the multi-source sensors include a grayscale image acquisition device and a depth image acquisition device;

[0047] S12, using the calibrated and registered sensor to collect a grayscale image and a depth image of the rail;

[0048] S13, annotating the surface damage of the rail in the collected grayscale image, constructing a rail multi-source dataset using the rail depth map and the annotated grayscale image, and dividing the constructed rail multi-source dataset into a training set, a validation set and a test set according to a preset ratio; wherein the training set is used to test the model, the validation set is used to verify the trained model, and the test set is used to test the trained model.

[0049] S2, constructing a rail surface damage detection model; wherein the rail surface damage detection model includes: a symmetrical texture feature extraction network and a deep feature extraction network, as well as multiple cross-modal feature fusion modules, a multi-scale feature fusion module and multiple decoupled detection heads;

[0050] The input of the texture feature extraction network is the grayscale image of the rail, which is used to extract the grayscale feature map; the input of the depth feature extraction network is the depth map of the rail, which is used to extract the depth feature map; the structure of the texture feature extraction network is the same as that of the depth feature extraction network, both of which include: multiple convolutional layers, depthwise separable convolutional layers, C3 modules, and simplified fast spatial pyramid pooling (SimSPPF) modules. The depthwise separable convolutional layer includes a layer of depthwise separable convolution, a layer of batch normalization, and a layer of Silu activation function in sequence; the C3 module includes multiple depthwise separable convolutional layers and multiple skip connections.

[0051] Specifically, the feature extraction network extracts features from an image to obtain feature maps of multiple different levels, including: inputting an image, sequentially passing through a convolutional layer, two depth-wise separable convolutional layers and a C3 module to obtain a shallow feature map, then passing through a depth-wise separable convolutional layer and a C3 module to obtain a middle feature map, and then passing through a depth-wise separable convolutional layer and a SimSPPF module to obtain a deep feature map.

[0052] Furthermore, in this embodiment, the grayscale feature map extracted by the texture feature extraction network includes: a shallow grayscale feature map, a middle grayscale feature map, and a deep grayscale feature map; the depth feature map extracted by the depth feature extraction network includes: a shallow depth feature map, a middle depth feature map, and a deep depth feature map.

[0053] The cross-modal feature fusion module is used to fuse the grayscale feature map and the depth feature map with the same receptive field at the same level; specifically, in this embodiment, there are three cross-modal feature fusion modules, which are respectively located in the shallow, middle and deep layers of the feature extraction network. The three cross-modal feature fusion modules are used to: fuse the shallow grayscale feature map and the shallow depth feature map to obtain the shallow fusion feature map F1; fuse the middle grayscale feature map and the middle depth feature map to obtain the middle fusion feature map F2; and fuse the deep grayscale feature map and the deep depth feature map to obtain the deep fusion feature map F3.

[0054] The structure of the cross-modal feature fusion module is as follows: Figure 2 As shown, the execution process is as follows:

[0055] The grayscale feature map is added to the depth feature map to obtain the feature map f1, which is input into four depth-wise separable convolutions for average pooling operations. The convolution kernel sizes of the four depth-wise separable convolutions are 7×7, 9×9, 13×13, and 17×17, respectively. Thus, feature maps f2, f3, f4, and f5 are obtained, respectively. The expression of this process is:

[0056] f2=DWConv(f1,k=(7,7))f3=DWConv(f1,k=(9,9))

[0057] f4=DWConv(f1,k=(13,13))f5=DWConv(f1,k=(17,17))

[0058] Among them, DWConv represents the depth-wise separable convolutional layer; k is the convolution kernel size.

[0059] Feature map f1 is added to feature map f2 to obtain feature map f6. Feature map f3 is upsampled and added to feature map f6 to obtain feature map f7. Feature map f4 is upsampled and added to feature map f7 to obtain feature map f8. Feature map f5 is upsampled and added to feature map f8 to obtain feature map f9. Feature maps f6, f7, f8 and f9 are concatenated and convolution operation with convolution kernel of 1×1 is performed to obtain feature map f 10 , the expression of this process is;

[0060] f6=f1+f2 f7=upsample(f3)+f6

[0061] f8=upsample(f4)+f7 f9=upsample(f5)+f8

[0062] f 10 =Conv(torch.cat(f6,f7,f8,f9),k=(1,1))

[0063] Among them, upsample is upsampling; Conv is the convolution operation; torch.cat is the concatenation operation.

[0064] Feature map f 10 After adding it to the feature map f1, the final fused feature map is output. Based on this operation, the shallow, middle and deep fused feature maps can be obtained respectively through three cross-modal feature fusion modules.

[0065] The input of the multi-scale feature fusion module is the shallow fusion feature map, the middle fusion feature map and the deep fusion feature map; the module consists of multiple depth-separable convolutional layers, C3 modules and upsampling modules. The process of processing the input shallow fusion feature map, middle fusion feature map and deep fusion feature map includes:

[0066] The deep fusion feature map F3 passes through a C3 module to obtain the feature map F4, and then passes through the depth separable convolution layer and the upsampling module in sequence. The resolution of the obtained feature map is the same as that of the middle-layer fusion feature map F2. The obtained feature map is spliced ​​with the middle-layer fusion feature map F2 to obtain the feature map F5. The feature map F5 passes through a C3 module to obtain the feature map F6, and then passes through the depth separable convolution layer and upsampling in sequence. The resolution of the obtained feature map is the same as that of the shallow-layer fusion feature map F1. The obtained feature map is spliced ​​with the shallow-layer fusion feature map F1 to obtain the feature map F7. The formula is expressed as follows:

[0067] F4=C3(F3)F5=torch.cat(F2,upsample(DWConv(F4)))

[0068] F6=C3(F5)F7=torch.cat(F1,upsample(DWConv(F6)))

[0069] After the feature map F7 passes through a C3 module and a depth-separable convolutional layer, the resolution of the feature map obtained is the same as that of the feature map F6. The obtained feature map is concatenated with the feature map F6 to obtain the feature map F8. After the feature map F8 passes through a C3 module and a depth-separable convolutional layer, the resolution of the feature map obtained is the same as that of the feature map F4. The obtained feature map is concatenated with the feature map F4 to obtain the feature map F9. The formula is expressed as:

[0070] F8=torch.cat(DWConv(C3(F7)),F6)

[0071] F9=torch.cat(DWConv(C3(F8)),F4)

[0072] Furthermore, there are three decoupled detection heads, which are used to perform decoupling operations on feature maps F7, F8, and F9 to predict objects of different scales; each decoupled detection head performs decoupling operations on a feature map; a feature map corresponds to a decoupled detection head; and the three decoupled detection heads have the same structure. The decoupling operation includes a layer of C3 module, which is then divided into two branches, including two layers of depth-separable convolutional layers and a layer of convolution, which are then merged through a splicing operation to obtain the predicted category, position, and confidence. The three decoupled detection heads predict objects of different scales.

[0073] The framework of the rail surface damage detection model finally constructed is as follows: Figure 3 shown.

[0074] S3, train the rail surface damage detection model using rail multi-source dataset;

[0075] Specifically, the above S3 includes: inputting the training set and the validation set into the rail surface damage detection model for training to obtain the optimal weight, and inputting the test set into the model for testing to obtain the rail surface damage detection result.

[0076] S4, obtaining a grayscale image and a depth image of the rail to be detected, inputting the grayscale image and the depth image of the rail to be detected into a trained rail surface damage detection model, and obtaining a surface damage detection result of the rail to be detected.

[0077] In summary, this embodiment provides a multi-source data fusion rail surface damage detection method, constructs a symmetrical fusion rail surface damage detection model, efficiently utilizes the two-dimensional grayscale image and three-dimensional depth image collected by multi-source sensors, performs multiple feature fusions and exchanges at the feature level, bidirectionally guides the symmetrical texture and depth feature extraction networks to extract their respective effective features, and uses the multi-level fused feature map as the source of the detection head to obtain high-precision position, category and confidence prediction results, reduces interference such as shadows and stains, improves the detection accuracy of various damage categories such as rail cracks, corrugation, peeling, and abrasions, and improves the reliability of the detection results.

[0078] Second embodiment

[0079] This embodiment provides an electronic device, such as Figure 4As shown, the electronic device includes: a processor and a memory; wherein the processor and the memory can be connected via a communication bus; the memory stores at least one instruction, and the instruction is loaded and executed by the processor to implement the method of the first embodiment. In addition, the electronic device may also include a transceiver, the processor and the transceiver can be connected via a communication bus, and the transceiver is used to communicate with other devices.

[0080] Next, combine Figure 4 The following is a detailed introduction to the various components of the electronic device:

[0081] Among them, the processor is the control center of the electronic device, and the electronic device may include multiple processors, each of which may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here may be a processor or a general term for multiple processing elements. For example, the processor is one or more central processing units (CPUs), or other general-purpose processors, application specific integrated circuits (ASICs), or one or more integrated circuits configured to implement an embodiment of the present invention, such as one or more microprocessors (digital signal processors, DSPs), or one or more field programmable gate arrays (field programmable gate arrays, FPGAs), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor may execute various functions of the electronic device by running or executing software programs stored in the memory and calling data stored in the memory.

[0082] In a specific implementation, as an embodiment, the processor may include one or more CPUs, such as Figure 4 The CPU0 and CPU1 shown in the figure are, of course, only exemplary.

[0083] The memory is used to store the software program for executing the solution of the present invention, and the execution is controlled by the processor. The specific implementation method can refer to the above method embodiment, which will not be repeated here.

[0084] Optionally, the memory may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory may be integrated with the processor or exist independently and accessed through the interface circuit ( Figure 4 (not shown) is coupled to the processor, which is not specifically limited in this embodiment of the present invention.

[0085] The transceiver may include a receiver and a transmitter ( Figure 4 The receiver is used to implement the receiving function, and the transmitter is used to implement the sending function. The transceiver can be integrated with the processor or exist independently and communicate with the electronic device through the interface circuit ( Figure 4 (not shown) is coupled to the processor, which is not specifically limited in this embodiment of the present invention.

[0086] In addition, it should be noted that Figure 4 The structure of the electronic device shown in the figure does not constitute a limitation on the device, and the actual device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. In addition, the technical effects achieved by the electronic device when executing the method of the first embodiment above can refer to the technical effects described in the first embodiment above, so they are not repeated here.

[0087] Third embodiment

[0088] This embodiment provides a computer-readable storage medium, which stores at least one instruction, and the instruction is loaded and executed by a processor to implement the method of the first embodiment. The computer-readable storage medium may be a ROM, a random access memory, a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc. The instructions stored therein may be loaded by a processor in a terminal to execute the method.

[0089] In addition, it should be noted that the present invention can be provided as a method, an apparatus or a computer program product. Therefore, the embodiment of the present invention can be in the form of a full or partial hardware embodiment, a full or partial software embodiment or an embodiment combining software and hardware. Moreover, when implemented using software, the embodiment of the present invention can be in the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program codes. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from one website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center containing one or more available media sets. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a DVD), or a semiconductor medium. The semiconductor medium may be a solid state hard disk.

[0090] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0091] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable terminal device provide for implementing the process in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0092] It should also be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of more restrictions, the elements defined by the sentence "including one..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements. In addition, the term "and / or" is only an association relationship describing the associated objects, indicating that there can be three relationships, for example, A and / or B, which can represent: A exists alone, A and B exist at the same time, and B exists alone, wherein A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding. "At least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can be represented by: a, b, c, ab, ac, bc or abc, where a, b, c can be single or plural.

[0093] In addition, it can be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0094] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0095] In several embodiments provided by the present invention, it should be understood that the disclosed equipment, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of functional modules / units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms. The unit described as a separate component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it may be located in one place, or it may be distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, each functional unit in each embodiment of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0096] If the method is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0097] Finally, it should be noted that the above is only a preferred embodiment of the present invention. It should be pointed out that although the preferred embodiment of the present invention has been described, for ordinary technicians in this technical field, once the basic creative concept of the present invention is known, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the protection scope of the present invention. Therefore, the attached claims are intended to be interpreted as including the preferred embodiment and all changes and modifications that fall within the scope of the embodiments of the present invention.

Claims

1. A rail surface damage detection method based on multi-source data fusion, characterized in that: include: Constructing a rail multi-source data set; wherein the rail multi-source data set includes a grayscale image and a depth image of the rail; and the grayscale image of the rail carries rail surface damage annotation data; Constructing a rail surface damage detection model; wherein the rail surface damage detection model includes: a symmetrical texture feature extraction network and a deep feature extraction network, as well as multiple cross-modal feature fusion modules, a multi-scale feature fusion module and multiple decoupled detection heads; Using the rail multi-source data set to train the rail surface damage detection model; The grayscale image and depth image of the rail to be detected are obtained, and the grayscale image and depth image of the rail to be detected are input into the trained rail surface damage detection model to obtain the surface damage detection result of the rail to be detected.

2. The rail surface damage detection method based on multi-source data fusion according to claim 1, characterized in that: The construction of a rail multi-source dataset includes: Deploy multi-source sensors around the rails and calibrate and register the deployed sensors; Use the calibrated and registered sensors to collect the grayscale image and depth image of the rail; The surface damage of the rails is annotated on the collected grayscale images, and a multi-source rail dataset is constructed using the depth map of the rails and the grayscale images of the annotated rails. The constructed multi-source rail dataset is divided into a training set, a validation set, and a test set according to a preset ratio; the training set is used to test the model, the validation set is used to verify the trained model, and the test set is used to test the trained model.

3. The rail surface damage detection method based on multi-source data fusion according to claim 1, characterized in that: The input of the texture feature extraction network is a grayscale image of the rail, which is used to extract a grayscale feature image; The input of the depth feature extraction network is the depth map of the rail, which is used to extract the depth feature map; The texture feature extraction network has the same structure as the depth feature extraction network, both of which include: multiple convolutional layers, depth-separable convolutional layers, a C3 module, and a SimSPPF module.

4. The rail surface damage detection method based on multi-source data fusion as claimed in claim 3, characterized in that: The grayscale feature map includes: a shallow grayscale feature map, a middle grayscale feature map and a deep grayscale feature map; the depth feature map includes: a shallow depth feature map, a middle depth feature map and a deep depth feature map; The number of the cross-modal feature fusion modules is three, and the three cross-modal feature fusion modules are respectively used to: fuse the shallow grayscale feature map and the shallow depth feature map to obtain a shallow fusion feature map; fuse the middle grayscale feature map and the middle depth feature map to obtain a middle fusion feature map; and fuse the deep grayscale feature map and the deep depth feature map to obtain a deep fusion feature map.

5. The rail surface damage detection method based on multi-source data fusion as claimed in claim 4, characterized in that: The process of implementing feature map fusion by the cross-modal feature fusion module includes: The grayscale feature map and the depth feature map with the same receptive field at the same level to be fused are input into the corresponding cross-modal feature fusion module; the cross-modal feature fusion module adds the input grayscale feature map and the depth feature map to obtain a first feature map, which are respectively input into four depth-wise separable convolutions and respectively average pooled to obtain a second feature map, a third feature map, a fourth feature map and a fifth feature map; wherein the convolution kernel sizes of the four depth-wise separable convolutions are 7×7, 9×9, 13×13 and 17×17 respectively; The first feature map and the second feature map are added to obtain a sixth feature map; the third feature map is upsampled and added to the sixth feature map to obtain a seventh feature map; the fourth feature map is upsampled and added to the seventh feature map to obtain an eighth feature map; the fifth feature map is upsampled and added to the eighth feature map to obtain a ninth feature map; the sixth feature map, the seventh feature map, the eighth feature map and the ninth feature map are concatenated, and a convolution operation with a convolution kernel of 1×1 is performed to obtain a tenth feature map; The tenth feature map is added to the first feature map to obtain a fused feature map of the corresponding level.

6. The rail surface damage detection method based on multi-source data fusion as claimed in claim 4, characterized in that: The input of the multi-scale feature fusion module is the shallow fusion feature map, the middle fusion feature map and the deep fusion feature map; The multi-scale feature fusion module includes: multiple depth-wise separable convolutional layers, a C3 module, and an upsampling module.

7. The rail surface damage detection method based on multi-source data fusion as claimed in claim 6, characterized in that: The process of the multi-scale feature fusion module processing the input shallow fusion feature map, middle fusion feature map and deep fusion feature map includes: After the deep fusion feature map passes through a C3 module, an eleventh feature map is obtained. After the eleventh feature map passes through a depth-separable convolution layer and an upsampling module in sequence, a twelfth feature map having the same resolution as the middle-layer fusion feature map is obtained. The twelfth feature map is spliced ​​with the middle-layer fusion feature map to obtain a thirteenth feature map. After the thirteenth feature map passes through a C3 module, a fourteenth feature map is obtained. After the fourteenth feature map passes through a depth-separable convolution layer and an upsampling module in sequence, a fifteenth feature map having the same resolution as the shallow-layer fusion feature map is obtained. The fifteenth feature map is spliced ​​with the shallow-layer fusion feature map to obtain a sixteenth feature map. After the sixteenth feature map passes through a C3 module and a depth-wise separable convolutional layer in sequence, a seventeenth feature map having the same resolution as the fourteenth feature map is obtained. The seventeenth feature map is spliced ​​with the fourteenth feature map to obtain an eighteenth feature map. After the eighteenth feature map passes through a C3 module and a depth-wise separable convolutional layer in sequence, a nineteenth feature map having the same resolution as the eleventh feature map is obtained. The nineteenth feature map is spliced ​​with the eleventh feature map to obtain a twentieth feature map.

8. The rail surface damage detection method based on multi-source data fusion as claimed in claim 7, characterized in that: The number of the decoupling detection heads is three, which are respectively used to perform decoupling operations on the sixteenth feature map, the eighteenth feature map, and the twentieth feature map to predict targets of different scales; The decoupled detection head includes a layer of C3 module, and the C3 module includes two branches, one of which includes two layers of depth-separable convolutional layers, and the other branch includes a layer of convolution. The outputs of the two branches are merged through a splicing operation to obtain the predicted category, predicted position and predicted confidence corresponding to the corresponding scale target.