Method for detecting unknown surface damage of steel rail based on multi-source data fusion

Through the combination of multi-source data fusion and unsupervised detection models, the problems of incomplete and inaccurate surface damage detection in the existing technology are solved, and accurate detection and positioning of more damage categories are achieved, reducing the cost of data annotation.

CN119992156APending Publication Date: 2025-05-13BEIJING MASS TRANSIT RAILWAY OPERATION CORPORATION LIMITED
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411817959.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-11
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing rail surface damage detection methods are incomplete and have insufficient accuracy, especially when rail damage is high in diversity and complexity, it is difficult to effectively detect more damage categories.

Method used

The multi-source data fusion method is adopted to construct a multi-source data set of rails, including a registered two-dimensional grayscale map and three-dimensional point cloud, and an unsupervised detection model for rail surface damage, including a data fusion module, a density multi-scale residual module, a multi-scale feature fusion module and a damage discrimination module. Through the combination of these modules, the detection of rail surface damage is realized.

Benefits of technology

This method can detect more damage categories without relying on a large amount of damage data, improve the comprehensiveness and accuracy of the detection content, reduce the manpower and material consumption of data labeling, and avoid the problem of algorithm invalidity caused by the lack of damage data of a certain category.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992156A_ABST
    Figure CN119992156A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-source data fusion steel rail unknown surface damage detection method, which belongs to the technical field of machine vision, and comprises the following steps: constructing a steel rail multi-source data set; wherein the data set comprises a steel rail two-dimensional grey-scale map and a steel rail three-dimensional point cloud which are registered; a steel rail surface damage unsupervised detection model is constructed and trained; the model comprises a data fusion module, a density multi-scale residual module, a multi-scale feature fusion module and a damage discrimination module. And acquiring a two-dimensional grey-scale map and a three-dimensional point cloud of the to-be-detected steel rail, inputting the acquired two-dimensional grey-scale map and the three-dimensional point cloud into the trained model, and acquiring a surface damage detection result of the to-be-detected steel rail based on a model output result. According to the method, fusion of multi-source information is realized, the unknown surface damage detection model of the steel rail is obtained by training under the condition that only the normal steel rail image is utilized, and detection and positioning of the unknown type of surface damage of the steel rail are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of machine vision, and in particular to a method for detecting unknown rail surface damage by fusion of multi-source data. Background Art

[0002] Rails are the main components of railway tracks and carry high-speed trains for a long time. As the service life increases, the service status of rails will change, and there will be damage on their surface, such as cracks and wear, which may affect the safety of train operation. In serious cases, it may cause train derailment and rollover, resulting in casualties and property losses. Therefore, regular inspection of rail surface damage is an important measure to ensure the safety of train operation.

[0003] In recent years, deep learning-based target detection and semantic segmentation algorithms have been gradually applied to rail surface damage detection. A large amount of damage data is used to train the model, thereby realizing the automatic identification of rail surface damage. However, due to the diversity and complexity of rail damage, the types of rail surface damage contained in the established data set and the accuracy of damage data annotation have greatly affected the comprehensiveness of the algorithm detection content and the detection accuracy. In reality, the amount of damage data accounts for a very small proportion of the total data set compared to the amount of normal data, and data annotation requires a lot of manpower and material resources. Therefore, it is necessary to develop an intelligent rail surface damage detection algorithm that does not rely on a large amount of damage data and can detect more damage categories. Summary of the invention

[0004] The present invention provides a multi-source data fusion rail surface damage detection method to solve the technical problems of the existing rail surface damage detection method in that the comprehensiveness of detection content and detection accuracy are not ideal.

[0005] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0006] In one aspect, the present invention provides a method for detecting unknown rail surface damage by multi-source data fusion, the method comprising:

[0007] Constructing a rail multi-source dataset; wherein the rail multi-source dataset includes a registered two-dimensional grayscale image and a three-dimensional point cloud of the rail;

[0008] Constructing an unsupervised detection model for rail surface damage; wherein the unsupervised detection model for rail surface damage includes a data fusion module, a density multi-scale residual module, a multi-scale feature fusion module and a damage discrimination module;

[0009] Using the rail multi-source data set to train the rail surface damage unsupervised detection model;

[0010] A two-dimensional grayscale image and a three-dimensional point cloud of the rail to be inspected are obtained, and the two-dimensional grayscale image and the three-dimensional point cloud of the rail to be inspected are input into a trained unsupervised detection model for rail surface damage. Based on the output result of the trained unsupervised detection model for rail surface damage, a surface damage detection result of the rail to be inspected is obtained.

[0011] Furthermore, the construction of the rail multi-source dataset includes:

[0012] Deploy multi-source sensors around the rails;

[0013] The two-dimensional grayscale image and three-dimensional point cloud of the normal rail surface are collected by using the deployed multi-source sensors;

[0014] The 2D grayscale image and 3D point cloud of the same rail surface are registered, and a multi-source rail dataset is constructed using the registered 2D grayscale image and 3D point cloud of the rail surface.

[0015] Furthermore, the input of the data fusion module is a two-dimensional grayscale image and a three-dimensional point cloud of the rail;

[0016] The process of the data fusion module processing the input two-dimensional grayscale image and three-dimensional point cloud includes:

[0017] Convert the 3D point cloud into a depth map and normalize the depth map to the same order of magnitude as the 2D grayscale image;

[0018] The pixel points at the same position in the two pixel matrices corresponding to the two-dimensional grayscale image and the normalized depth image are calculated by pixel weighted average to obtain the rail fusion image.

[0019] Furthermore, when performing pixel weighted averaging calculation on the pixels at the same position in the two pixel matrices corresponding to the two-dimensional grayscale image and the normalized depth image, the weight coefficients of the pixels at the same position in the two pixel matrices are both 0.5.

[0020] Furthermore, converting the three-dimensional point cloud into a depth map includes:

[0021] The 3D point cloud is converted into a depth map by using the principle of reverse mapping of the point cloud to the pixel coordinate position.

[0022] Furthermore, the input of the density multi-scale residual module is the rail fusion image output by the data fusion module;

[0023] The dense multi-scale residual module includes: two 3×3 convolutional layers and three dense multi-scale residual blocks; wherein the stride of the two 3×3 convolutional layers is 2 and the number of channels is 64; except for the number of convolutional channels, the three dense multi-scale residual blocks have the same structure, and all include: a layer of depth-separable convolutional layer with a convolution kernel of 1×1, a first branch, a second branch, and a layer of convolutional layer with a convolution kernel of 1×1; the first branch includes: a layer of depth-separable convolutional layer with a convolution kernel of 1×5 and a layer of depth-separable convolutional layer with a convolution kernel of 5×1; the second branch includes: a layer of depth-separable convolutional layer with a convolution kernel of 1×9 and a layer of depth-separable convolutional layer with a convolution kernel of 9×1;

[0024] The number of convolution channels of the three dense multi-scale residual blocks are: 64, 128 and 256 respectively.

[0025] Furthermore, the process of processing the input rail fusion image by the density multi-scale residual module includes:

[0026] Use two 3×3 convolutional layers to operate on the rail fusion image in sequence to obtain the first feature map;

[0027] Inputting the first feature map into three dense multi-scale residual blocks respectively, extracting shallow detail information and deep semantic information of the first feature map respectively through the three dense multi-scale residual blocks, so as to obtain a second feature map with a resolution of 256×256, a third feature map with a resolution of 128×128, and a fourth feature map with a resolution of 64×64 respectively through the three dense multi-scale residual blocks;

[0028] The process of processing the first feature map by the dense multi-scale residual block includes:

[0029] Input the first feature map into a depth-separable convolution layer with a convolution kernel of 1×1 to obtain the fifth feature map;

[0030] Input the fifth feature map into the first branch and the second branch respectively, obtain a sixth feature map from the first branch, and obtain a seventh feature map from the second branch;

[0031] Adding the sixth feature map and the seventh feature map, and inputting the addition result into a convolution layer with a convolution kernel of 1×1 to obtain an eighth feature map;

[0032] The fifth feature map is input into a convolution layer with a convolution kernel of 3×3, and multiplied with the eighth feature map to obtain a feature map of corresponding resolution.

[0033] Furthermore, the input of the multi-scale feature fusion module is the second feature map, the third feature map and the fourth feature map output by the density multi-scale residual module; the process of the multi-scale feature fusion module processing the input second feature map, the third feature map and the fourth feature map includes:

[0034] Downsampling the second feature map by using a maximum pooling operation, and unifying the resolution to 64×64, to obtain a ninth feature map; downsampling the third feature map by using a maximum pooling operation, and unifying the resolution to 64×64, to obtain a tenth feature map;

[0035] Selecting the first 64 layers of the channel of the ninth feature map to obtain an eleventh feature map; selecting the first 64 layers of the channel of the fourth feature map to obtain a twelfth feature map;

[0036] The ninth feature map, the eleventh feature map and the twelfth feature map are channel-spliced ​​to obtain a multi-scale feature block; wherein the number of channels during channel splicing is 192.

[0037] Furthermore, the input of the damage identification module is the multi-scale feature block output by the multi-scale feature fusion module; the process of the damage identification module processing the input multi-scale feature block includes:

[0038] The multi-scale feature blocks are fitted using multivariate Gaussian distribution to obtain distribution coefficients corresponding to model inputs; wherein, during the model training process, the damage discrimination module obtains distribution coefficients corresponding to normal rail images; and during the inspection of the rails to be inspected, the damage discrimination module obtains distribution coefficients corresponding to the rail images to be inspected.

[0039] Furthermore, the output result of the trained rail surface damage unsupervised detection model is used to obtain the surface damage detection result of the rail to be detected, including:

[0040] The cosine similarity function is used to determine the spatial distance between the distribution coefficient corresponding to the image of the rail to be detected and the distribution coefficient corresponding to the normal rail image, and the surface damage detection result of the rail to be detected is obtained based on the spatial distance.

[0041] On the other hand, the present invention further provides an electronic device, comprising a processor and a memory; wherein the memory stores at least one instruction, and the instruction is loaded and executed by the processor to implement the above method.

[0042] In yet another aspect, the present invention further provides a computer-readable storage medium, wherein at least one instruction is stored in the storage medium, and the instruction is loaded and executed by a processor to implement the above method.

[0043] The beneficial effects brought about by the technical solution provided by the present invention include at least:

[0044] The multi-source data fusion rail unknown surface damage detection method provided by the present invention first utilizes the texture and spatial features in the multi-source data of the rail to enrich the rail damage features, and selects fusion at the data layer to maximize the fusion of multi-source information and the retention of effective information. Secondly, this algorithm only uses normal rail images for training to obtain a rail unknown surface damage detection model, and uses an unsupervised method to achieve accurate detection and positioning of unknown types of rail surface damage, which greatly reduces the human and material consumption caused by data annotation, and the detection content is more comprehensive, which can avoid the problem of algorithm invalidation caused by lack of training of a certain type of damage data. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0046] Figure 1 It is a schematic diagram of a method for detecting unknown rail surface damage by multi-source data fusion provided by an embodiment of the present invention;

[0047] Figure 2 It is a system block diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0048] To make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0049] First of all, it should be noted that in the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "exemplarily" is intended to present the concept in a concrete way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.

[0050] First embodiment

[0051] This embodiment provides a method for detecting unknown rail surface damage by multi-source data fusion. The implementation principle of the method for detecting unknown rail surface damage is as follows: Figure 1 As shown, the method can be implemented by an electronic device, which can be a terminal or a server. Specifically, the execution process of the method includes the following steps:

[0052] S1, constructing a rail multi-source dataset; wherein the rail multi-source dataset includes a registered two-dimensional grayscale image and a three-dimensional point cloud of the rail;

[0053] Specifically, in this embodiment, the implementation process of the above S1 includes:

[0054] S11, deploying a multi-source sensor around the rail; wherein the multi-source sensor includes: a two-dimensional grayscale image acquisition device and a three-dimensional point cloud acquisition device;

[0055] S12, using the deployed multi-source sensors to collect a two-dimensional grayscale image and a three-dimensional point cloud of the normal rail surface;

[0056] S13, registering the two-dimensional grayscale image and the three-dimensional point cloud of the same rail surface, and constructing a rail multi-source dataset using the registered two-dimensional grayscale image and the three-dimensional point cloud of the rail surface.

[0057] S2, constructing an unsupervised detection model for rail surface damage; wherein the unsupervised detection model for rail surface damage includes a data fusion module, a density multi-scale residual module, a multi-scale feature fusion module and a damage discrimination module;

[0058] The input of the data fusion module is the two-dimensional grayscale image and the three-dimensional point cloud of the rail. In this regard, the process of processing the input two-dimensional grayscale image and the three-dimensional point cloud by the data fusion module includes:

[0059] According to the camera intrinsic parameters, the 3D point cloud is converted into a depth map by using the principle of reverse mapping of the point cloud to the pixel coordinate position, and the depth map is normalized to the same order of magnitude as the 2D grayscale image, that is, [0,255]. The formula is:

[0060]

[0061] Among them, Image Depth is the depth map; Image Depth (i,j) is the pixel with coordinates (i,j) in the depth map.

[0062] The pixel points at the same position in the two pixel matrices corresponding to the two-dimensional grayscale image and the normalized depth image are calculated by weighted average, with the weight set to 0.5, to obtain the rail fusion image. The formula is:

[0063] Image fused (i,j)=Image Depth (i,j)*ω1+Image Grayscale (i,j)*ω2

[0064] Among them, Image fused(i,j) represents the pixel with coordinates (i,j) in the rail fusion image; Image Depth (i,j) represents the pixel with coordinates (i,j) in the depth map; Image Grayscale (i, j) represents the pixel with coordinates (i, j) in the two-dimensional grayscale image; ω1 and ω2 are weight coefficients, both of which are 0.5.

[0065] The input of the dense multi-scale residual module is the rail fusion image output by the data fusion module; the dense multi-scale residual module includes: two 3×3 convolutional layers and three dense multi-scale residual blocks; among them, the stride of the two 3×3 convolutional layers is 2, and the number of channels is 64; except for the number of convolution channels, the three dense multi-scale residual blocks have exactly the same structure, including: a layer of depth-separable convolutional layer with a convolution kernel of 1×1, the first branch, the second branch, and a layer of convolutional layer with a convolution kernel of 1×1; the first branch includes: a layer of depth-separable convolutional layer with a convolution kernel of 1×5 and a layer of depth-separable convolutional layer with a convolution kernel of 5×1; the second branch includes: a layer of depth-separable convolutional layer with a convolution kernel of 1×9 and a layer of depth-separable convolutional layer with a convolution kernel of 9×1; the number of convolutional channels of the three dense multi-scale residual blocks are 64, 128 and 256 respectively.

[0066] Based on the above, the process of processing the input rail fusion image by the density multi-scale residual module includes:

[0067] Use two 3×3 convolutional layers to operate on the rail fusion image to obtain the feature map F1. The formula is:

[0068] F1=nn.Conv2d(nn.Conv2d(64,3,3,2),3,3,2)

[0069] The feature map F1 is input into three series-connected dense multi-scale residual blocks; wherein the dense multi-scale residual block processes the feature map F1 in the following process:

[0070] Input the feature map F1 into a depth-separable convolution layer with a convolution kernel of 1×1 to obtain the feature map f1;

[0071] The feature map f1 is input into the first branch and the second branch for operation, and the feature map f2 is obtained from the first branch, and the feature map f3 is obtained from the second branch;

[0072] Add feature map f2 and feature map f3, and input the addition result into a convolution layer with a convolution kernel of 1×1 to obtain feature map f4;

[0073] The feature map f1 is input into a convolution layer with a convolution kernel of 3×3, and multiplied with the feature map f4 to obtain a feature map of the corresponding resolution. Based on this, the present embodiment uses three dense multi-scale residual blocks to respectively extract shallow detail information and deep semantic information of the feature map F1, so as to respectively obtain a feature map F2 with a resolution of 256×256, a feature map F3 with a resolution of 128×128, and a feature map F4 with a resolution of 64×64 through three dense multi-scale residual blocks.

[0074] The input of the multi-scale feature fusion module is the feature map F2, feature map F3 and feature map F4 output by the density multi-scale residual module; the process of the multi-scale feature fusion module processing the input feature map includes:

[0075] The feature maps F2 and F3 are downsampled by using the maximum pooling operation, and the resolution is unified to 64×64 to obtain feature maps F2′ and F3′;

[0076] Select the first 64 layers of the feature map F3' and the feature map F4 channels respectively to obtain the feature map F3" and the feature map F4';

[0077] The channel splicing feature map F2′, feature map F3″ and feature map F4′ obtains a multi-scale feature block Block, where the number of channels is 192; the formula of this process is expressed as:

[0078] Block=torch.cat(F2′,F3″,F4′)

[0079] Among them, torch.cat represents the concatenation operation.

[0080] The input of the damage identification module is the multi-scale feature block output by the multi-scale feature fusion module; the process of the damage identification module processing the input multi-scale feature block includes:

[0081] The multi-scale feature blocks are fitted using multivariate Gaussian distribution to obtain the distribution coefficient corresponding to the model input; during the training process, the damage discrimination module obtains the distribution coefficient corresponding to the normal rail image; during the detection process, the damage discrimination module obtains the distribution coefficient corresponding to the rail image to be detected.

[0082] S3, using rail multi-source datasets to train an unsupervised detection model for rail surface damage;

[0083] S4, obtaining a two-dimensional grayscale image and a three-dimensional point cloud of the rail to be inspected, inputting the two-dimensional grayscale image and the three-dimensional point cloud of the rail to be inspected into a trained unsupervised detection model for rail surface damage, and obtaining a surface damage detection result of the rail to be inspected based on an output result of the trained unsupervised detection model for rail surface damage;

[0084] Specifically, in this embodiment, based on the output result of the trained rail surface damage unsupervised detection model, the surface damage detection result of the rail to be detected is obtained by using the cosine similarity function to determine the spatial distance between the multi-scale feature block of the rail image to be detected and the distribution coefficient of the normal rail image, and obtaining the surface damage detection result of the rail to be detected based on the calculated spatial distance, specifically: after obtaining the spatial distance, the abnormal score value of each pixel is calculated using the cosine similarity function, and the area with a larger score value is determined as an abnormal pixel. The cosine similarity function is expressed as:

[0085]

[0086] Wherein, μ is the distribution coefficient of the normal rail image; x is the multi-scale feature block of the rail image to be detected.

[0087] In summary, this embodiment provides a method for detecting unknown rail surface damage by fusion of multi-source data. The method only uses normal rail images to train an unsupervised detection model for rail surface damage, accurately realizing the detection of surface damage of unknown rail categories, improving the problem of low algorithm detection accuracy due to the scarcity of abnormal samples, and solving the problem of incomplete detection content due to limited damage data categories. In addition, the use of multi-source rail data enriches effective information at the data level, effectively improves the detection accuracy, and the steps are relatively simple, requiring less computing resources, and has high universality.

[0088] Second embodiment

[0089] This embodiment provides an electronic device, such as Figure 2 As shown, the electronic device includes: a processor and a memory; wherein the processor and the memory can be connected via a communication bus; the memory stores at least one instruction, and the instruction is loaded and executed by the processor to implement the method of the first embodiment. In addition, the electronic device may also include a transceiver, the processor and the transceiver can be connected via a communication bus, and the transceiver is used to communicate with other devices.

[0090] Next, combine Figure 2 The following is a detailed introduction to the various components of the electronic device:

[0091] Among them, the processor is the control center of the electronic device, and the electronic device may include multiple processors, each of which may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here may be a processor or a general term for multiple processing elements. For example, the processor is one or more central processing units (CPUs), or other general-purpose processors, application specific integrated circuits (ASICs), or one or more integrated circuits configured to implement an embodiment of the present invention, such as one or more microprocessors (digital signal processors, DSPs), or one or more field programmable gate arrays (field programmable gate arrays, FPGAs), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor may execute various functions of the electronic device by running or executing software programs stored in the memory and calling data stored in the memory.

[0092] In a specific implementation, as an embodiment, the processor may include one or more CPUs, such as Figure 2 The CPU0 and CPU1 shown in the figure are, of course, only exemplary.

[0093] The memory is used to store the software program for executing the solution of the present invention, and the execution is controlled by the processor. The specific implementation method can refer to the above method embodiment, which will not be repeated here.

[0094] Optionally, the memory may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory may be integrated with the processor or exist independently and accessed through the interface circuit ( Figure 2 (not shown) is coupled to the processor, which is not specifically limited in this embodiment of the present invention.

[0095] The transceiver may include a receiver and a transmitter ( Figure 2 The receiver is used to implement the receiving function, and the transmitter is used to implement the sending function. The transceiver can be integrated with the processor or exist independently and communicate with the electronic device through the interface circuit ( Figure 2 (not shown) is coupled to the processor, which is not specifically limited in this embodiment of the present invention.

[0096] In addition, it should be noted that Figure 2 The structure of the electronic device shown in the figure does not constitute a limitation on the device, and the actual device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. In addition, the technical effects achieved by the electronic device when executing the method of the first embodiment above can refer to the technical effects described in the first embodiment above, so they are not repeated here.

[0097] Third embodiment

[0098] This embodiment provides a computer-readable storage medium, which stores at least one instruction, and the instruction is loaded and executed by a processor to implement the method of the first embodiment. The computer-readable storage medium may be a ROM, a random access memory, a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc. The instructions stored therein may be loaded by a processor in a terminal to execute the method.

[0099] In addition, it should be noted that the present invention can be provided as a method, an apparatus or a computer program product. Therefore, the embodiment of the present invention can be in the form of a full or partial hardware embodiment, a full or partial software embodiment or an embodiment combining software and hardware. Moreover, when implemented using software, the embodiment of the present invention can be in the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program codes. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from one website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center containing one or more available media sets. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a DVD), or a semiconductor medium. The semiconductor medium may be a solid state hard disk.

[0100] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0101] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable terminal device provide for implementing the process in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0102] It should also be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of more restrictions, the elements defined by the sentence "including one..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements. In addition, the term "and / or" is only an association relationship describing the associated objects, indicating that there can be three relationships, for example, A and / or B, which can represent: A exists alone, A and B exist at the same time, and B exists alone, wherein A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding. "At least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can be represented by: a, b, c, ab, ac, bc or abc, where a, b, c can be single or plural.

[0103] In addition, it can be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0104] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0105] In several embodiments provided by the present invention, it should be understood that the disclosed equipment, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of functional modules / units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms. The unit described as a separate component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it may be located in one place, or it may be distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, each functional unit in each embodiment of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0106] If the method is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0107] Finally, it should be noted that the above is only a preferred embodiment of the present invention. It should be pointed out that although the preferred embodiment of the present invention has been described, for ordinary technicians in this technical field, once the basic creative concept of the present invention is known, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the protection scope of the present invention. Therefore, the attached claims are intended to be interpreted as including the preferred embodiment and all changes and modifications that fall within the scope of the embodiments of the present invention.

Claims

1. A rail unknown surface damage detection method based on multi-source data fusion, characterized in that: include: Constructing a rail multi-source dataset; wherein the rail multi-source dataset includes a registered two-dimensional grayscale image and a three-dimensional point cloud of the rail; Constructing an unsupervised detection model for rail surface damage; wherein the unsupervised detection model for rail surface damage includes a data fusion module, a density multi-scale residual module, a multi-scale feature fusion module and a damage discrimination module; Using the rail multi-source data set to train the rail surface damage unsupervised detection model; A two-dimensional grayscale image and a three-dimensional point cloud of the rail to be inspected are obtained, and the two-dimensional grayscale image and the three-dimensional point cloud of the rail to be inspected are input into a trained unsupervised detection model for rail surface damage. Based on the output result of the trained unsupervised detection model for rail surface damage, a surface damage detection result of the rail to be inspected is obtained.

2. The rail unknown surface damage detection method based on multi-source data fusion as claimed in claim 1, characterized in that: The construction of a rail multi-source dataset includes: Deploy multi-source sensors around the rails; The two-dimensional grayscale image and three-dimensional point cloud of the normal rail surface are collected by using the deployed multi-source sensors; The 2D grayscale image and 3D point cloud of the same rail surface are registered, and a multi-source rail dataset is constructed using the registered 2D grayscale image and 3D point cloud of the rail surface.

3. The rail unknown surface damage detection method based on multi-source data fusion as claimed in claim 1, characterized in that: The input of the data fusion module is a two-dimensional grayscale image and a three-dimensional point cloud of the rail; The process of the data fusion module processing the input two-dimensional grayscale image and three-dimensional point cloud includes: Convert the 3D point cloud into a depth map and normalize the depth map to the same order of magnitude as the 2D grayscale image; The pixel points at the same position in the two pixel matrices corresponding to the two-dimensional grayscale image and the normalized depth image are calculated by pixel weighted average to obtain the rail fusion image.

4. The rail unknown surface damage detection method based on multi-source data fusion as claimed in claim 3, characterized in that: When performing pixel weighted averaging calculation on the pixels at the same position in the two pixel matrices corresponding to the two-dimensional grayscale image and the normalized depth image, the weight coefficients of the pixels at the same position in the two pixel matrices are both 0.

5.

5. The rail unknown surface damage detection method based on multi-source data fusion as claimed in claim 3, characterized in that: The step of converting a three-dimensional point cloud into a depth map comprises: The 3D point cloud is converted into a depth map by using the principle of reverse mapping of the point cloud to the pixel coordinate position.

6. The rail unknown surface damage detection method based on multi-source data fusion as claimed in claim 3, characterized in that: The input of the density multi-scale residual module is the rail fusion image output by the data fusion module; The dense multi-scale residual module includes: two 3×3 convolutional layers and three dense multi-scale residual blocks; wherein the stride of the two 3×3 convolutional layers is 2 and the number of channels is 64; except for the number of convolutional channels, the three dense multi-scale residual blocks have the same structure, and all include: a layer of depth-separable convolutional layer with a convolution kernel of 1×1, a first branch, a second branch, and a layer of convolutional layer with a convolution kernel of 1×1; the first branch includes: a layer of depth-separable convolutional layer with a convolution kernel of 1×5 and a layer of depth-separable convolutional layer with a convolution kernel of 5×1; the second branch includes: a layer of depth-separable convolutional layer with a convolution kernel of 1×9 and a layer of depth-separable convolutional layer with a convolution kernel of 9×1; The number of convolution channels of the three dense multi-scale residual blocks are: 64, 128 and 256 respectively.

7. The rail unknown surface damage detection method based on multi-source data fusion as claimed in claim 6, characterized in that: The process of the density multi-scale residual module processing the input rail fusion image includes: Use two 3×3 convolutional layers to operate on the rail fusion image in sequence to obtain the first feature map; Inputting the first feature map into three dense multi-scale residual blocks respectively, extracting shallow detail information and deep semantic information of the first feature map respectively through the three dense multi-scale residual blocks, so as to obtain a second feature map with a resolution of 256×256, a third feature map with a resolution of 128×128, and a fourth feature map with a resolution of 64×64 respectively through the three dense multi-scale residual blocks; The process of processing the first feature map by the dense multi-scale residual block includes: Input the first feature map into a depth-separable convolution layer with a convolution kernel of 1×1 to obtain the fifth feature map; Input the fifth feature map into the first branch and the second branch respectively, obtain a sixth feature map from the first branch, and obtain a seventh feature map from the second branch; Adding the sixth feature map and the seventh feature map, and inputting the addition result into a convolution layer with a convolution kernel of 1×1 to obtain an eighth feature map; The fifth feature map is input into a convolution layer with a convolution kernel of 3×3, and multiplied with the eighth feature map to obtain a feature map of corresponding resolution.

8. The rail unknown surface damage detection method based on multi-source data fusion as claimed in claim 7, characterized in that: The input of the multi-scale feature fusion module is the second feature map, the third feature map and the fourth feature map output by the density multi-scale residual module; the process of the multi-scale feature fusion module processing the input second feature map, the third feature map and the fourth feature map includes: Downsampling the second feature map by using a maximum pooling operation, and unifying the resolution to 64×64, to obtain a ninth feature map; downsampling the third feature map by using a maximum pooling operation, and unifying the resolution to 64×64, to obtain a tenth feature map; Selecting the first 64 layers of the channel of the ninth feature map to obtain an eleventh feature map; selecting the first 64 layers of the channel of the fourth feature map to obtain a twelfth feature map; The ninth feature map, the eleventh feature map and the twelfth feature map are channel-spliced ​​to obtain a multi-scale feature block; wherein the number of channels during channel splicing is 192.

9. The rail unknown surface damage detection method based on multi-source data fusion as claimed in claim 8, characterized in that: The input of the damage identification module is the multi-scale feature block output by the multi-scale feature fusion module; The process of the damage identification module processing the input multi-scale feature block includes: The multi-scale feature blocks are fitted using multivariate Gaussian distribution to obtain distribution coefficients corresponding to model inputs; wherein, during the model training process, the damage discrimination module obtains distribution coefficients corresponding to normal rail images; and during the inspection of the rails to be inspected, the damage discrimination module obtains distribution coefficients corresponding to the rail images to be inspected.

10. The rail unknown surface damage detection method based on multi-source data fusion as claimed in claim 9, characterized in that: The output result of the trained rail surface damage unsupervised detection model is used to obtain the surface damage detection result of the rail to be detected, including: The cosine similarity function is used to determine the spatial distance between the distribution coefficient corresponding to the image of the rail to be detected and the distribution coefficient corresponding to the normal rail image, and the surface damage detection result of the rail to be detected is obtained based on the spatial distance.