Railway scene microscopic disease detection method and device

CN122841818APending Publication Date: 2026-09-29山东铁投智能科技有限公司 +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610845264.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-11
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0005]本发明提供一种铁路场景微观病害检测方法及装置,用以解决现有技术中人工检测方案主观性强且覆盖范围有限,专用检测车存在检测盲区,而传统人工智能方案因依赖单一感知模态导致识别精度不足,且检测出的隐患仍需依赖人工进行二次评估、分级和工单派发的缺陷

Benefits of technology

[0020]本发明提供的铁路场景微观病害检测方法及装置,基于对可见光数据、红外数据和激光雷达点云数据进行多源融合,并基于融合后的多维度场景特征提取出目标缺陷的物理量化参数和空间距离,由此可以全面、客观地量化表征目标缺陷的严重程度与空间位置风险,进而能够基于这些硬件采集的客观数据在铁路运维知识库中进行自动匹配,实现从缺陷识别、风险研判到决策指令生成的自动化闭环,提高对隐蔽性、复杂性病害的检测精度和覆盖全面性,避免因单一数据模态导致的信息缺失和人工研判的主观性、滞后性问题,保障铁路运维响应的及时性和处置的精准性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122841818A_ABST
    Figure CN122841818A_ABST
Patent Text Reader

Abstract

This invention provides a method and apparatus for detecting microscopic defects in railway scenarios. The method includes: fusing visible light data, infrared data, and lidar point cloud data to obtain multi-dimensional scene features; identifying target defects in the railway scene to be detected based on the multi-dimensional scene features; extracting physical quantitative parameters of the target defects and the spatial distance between the target defects and a preset track based on the multi-dimensional scene features; and matching the physical quantitative parameters and spatial distance in a preset railway operation and maintenance knowledge base to obtain the risk level of the target defects and generate early warning instructions or maintenance work orders. This invention extracts physical quantitative parameters and spatial distance of target defects based on the fused multi-dimensional scene features, thereby comprehensively quantifying the severity and spatial location risk of target defects. It can automatically match these objective data collected by hardware in the railway operation and maintenance knowledge base, ensuring the timeliness of railway operation and maintenance response.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of railway infrastructure operation and maintenance and safety monitoring technology, and in particular to a method and device for detecting microscopic defects in railway scenarios. Background Technology

[0002] With the continuous growth of railway operating mileage and the constant increase in train operating speed, high-precision, high-efficiency, and routine monitoring of the operating status of infrastructure such as railway subgrade, bridges, tunnels, and overhead contact lines, as well as their external environment along the line, to achieve early identification and precise control of hidden defects and sudden risks, has become an urgent technical requirement for improving railway operation and maintenance and ensuring transportation safety.

[0003] To meet these needs, existing technologies mainly employ three approaches. The first is manual inspection, relying on technicians to conduct on-the-ground inspections and use simple tools for assessment. The second is dedicated inspection vehicle solutions, utilizing track inspection vehicles, overhead contact line inspection vehicles, or similar vehicles equipped with specific sensors to perform periodic inspections along the track. The third is traditional artificial intelligence inspection, which primarily uses visible light cameras mounted on mobile or fixed platforms to acquire images and employs trained visual recognition models to automatically identify abnormal targets within the images.

[0004] However, manual inspection is highly subjective, inefficient, and cannot effectively cover dangerous or inaccessible areas such as high altitudes, tunnel arches, and track boundaries. The inspection range of dedicated inspection vehicles is limited by their track operation mode, making it difficult to comprehensively scan and continuously monitor the sides of bridge piers, the entire cross-section of tunnel linings, details of power supply equipment, and the complex external environment along the track, resulting in blind spots. Traditional AI-based inspection solutions, which typically rely solely on visible light images, often lead to missed or false detections. Furthermore, detected hazards still require secondary assessment, grading, and work order assignment by humans, resulting in fragmented processes, slow responses, and an inability to meet the control requirements for real-time early warning and rapid response to sudden risks. Summary of the Invention

[0005] This invention provides a method and device for detecting microscopic defects in railway scenarios, which solves the problems of existing technologies such as the high subjectivity and limited coverage of manual detection schemes, the blind spots of special detection vehicles, and the insufficient recognition accuracy of traditional artificial intelligence schemes due to their reliance on a single perception mode, and the fact that detected hidden dangers still need to be evaluated, graded and assigned by humans.

[0006] This invention provides a method for detecting microscopic defects in railway scenarios, comprising the following steps.

[0007] Acquire visible light data, infrared data, and lidar point cloud data of the railway scene to be detected; The visible light data, the infrared data, and the lidar point cloud data are fused from multiple sources to obtain multi-dimensional scene features. Based on the multi-dimensional scene features, target defects in the railway scene to be detected are identified. Based on the multi-dimensional scene features, the physical quantification parameters of the target defects and the spatial distance between the target defects and the preset line are extracted. Based on the physical quantification parameters and the spatial distance, a matching is performed in a preset railway operation and maintenance knowledge base to obtain the risk level of the target defect; Generate early warning instructions or maintenance work orders corresponding to the risk level.

[0008] According to the present invention, a method for detecting microscopic defects in railway scenes includes multi-source fusion processing of visible light data, infrared data, and lidar point cloud data to obtain multi-dimensional scene features, including: Extract visible light features from the visible light data, extract infrared features from the infrared data, and extract geometric features from the lidar point cloud data; The visible light feature, the infrared feature, and the geometric feature are spliced ​​together to obtain the spliced ​​feature; Based on the concatenation features, a query matrix, a key matrix, and a value matrix are generated; Based on the dot product of the query matrix and the key matrix, attention weights are determined, and the feature vectors at each position in the value matrix are weighted and summed based on the attention weights to obtain the multi-dimensional scene features.

[0009] According to the present invention, a method for detecting microscopic defects in a railway scene includes identifying target defects in the railway scene to be detected based on the multi-dimensional scene features, comprising: The multi-dimensional scene features are input into a semantic segmentation network to obtain a binary mask image of the target defect output by the semantic segmentation network; the binary mask image is used to characterize the position and outline of the target defect in the railway scene to be detected.

[0010] According to the present invention, a method for detecting microscopic defects in railway scenarios includes extracting the physical quantification parameters of the target defect, comprising: In the case where the target defect is a crack, the skeleton of the binary mask image is extracted to obtain the centerline skeleton representing the direction of the target defect; The distance from each crack pixel in the binary mask image to the centerline skeleton is determined, and the maximum value among all distances is determined as the maximum width of the target defect. The maximum width is used as the physical quantization parameter.

[0011] According to the method for detecting microscopic defects in railway scenes provided by the present invention, the semantic segmentation network is obtained by iteratively executing the following steps until a preset iteration termination condition is met: Acquire training samples and corresponding target defect labels; the training samples include visible light data, infrared data, and lidar point cloud data. The sample visible light data, the sample infrared data, and the sample lidar point cloud data are subjected to multi-source fusion processing to obtain multi-dimensional scene features of the sample. The multi-dimensional scene features of the sample are input into the initial semantic segmentation network to obtain the predicted binary mask image output by the initial semantic segmentation network. Based on the difference between the predicted binary mask image and the target defect label, the target loss is determined, and the model parameters of the initial semantic segmentation network are updated based on the target loss.

[0012] According to the present invention, a method for detecting microscopic defects in railway scenarios includes extracting the physical quantification parameters of the target defect, comprising: In the case that the target defect is a water seepage defect, the defect area in the railway scene to be detected is identified based on the visible light data, and the temperature distribution in the railway scene to be detected is identified based on the infrared data; Based on the affected area and the temperature distribution, the seepage range of the seepage defect can be located; The physical quantification parameters are determined based on the area and / or contour dimensions of the seepage range.

[0013] According to the present invention, a method for detecting microscopic defects in railway scenarios includes extracting the physical quantification parameters of the target defect, comprising: When the target defect is a corrosion defect, the color and texture features of the corrosion area in the railway scene to be detected are extracted based on the visible light data, and the temperature distribution in the railway scene to be detected is identified based on the infrared data. Based on the color features, texture features, and temperature distribution, the corrosion area of ​​the target defect is determined; Based on the corrosion area, the physical quantification parameters are determined.

[0014] According to the present invention, a method for detecting microscopic defects in railway scenarios includes extracting the physical quantification parameters of the target defect, comprising: In the case where the target defect is a contact wire defect, the contact wire is fitted based on the lidar point cloud data to determine the contact wire's guide height, pull-out value, and offset. Based on the infrared data, the first temperature rise of the insulator in the contact network system to which the contact wire belongs, and the second temperature rise of the box-type substation located in the power supply area of ​​the contact network system are determined. The physical quantization parameters are determined based on the guide height, the pull-out value, the offset, the first temperature rise, and the second temperature rise.

[0015] According to the present invention, a method for detecting microscopic defects in railway scenarios further includes: The lidar point cloud data of the railway scene to be detected is acquired at multiple time points. Based on the lidar point cloud data of each time phase, a three-dimensional digital twin model of the corresponding period is constructed. The three-dimensional digital twin models of different time periods are registered using the same reference coordinate system, and the registered three-dimensional digital twin models are compared differentially to obtain the roadbed settlement and / or slope displacement vector as change quantification parameters. Based on the aforementioned change quantification parameters, the risk of gradual deformation of the roadbed or slope in the railway scenario to be detected is assessed.

[0016] The present invention also provides a microscopic defect detection device for railway scenarios, comprising the following units: The acquisition unit is used to acquire visible light data, infrared data, and lidar point cloud data of the railway scene to be detected. The extraction unit is used to perform multi-source fusion processing on the visible light data, the infrared data and the lidar point cloud data to obtain multi-dimensional scene features, identify target defects in the railway scene to be detected based on the multi-dimensional scene features, and extract the physical quantification parameters of the target defects and the spatial distance between the target defects and the preset line based on the multi-dimensional scene features. A matching unit is used to perform a matching operation in a preset railway operation and maintenance knowledge base based on the physical quantification parameters and the spatial distance to obtain the risk level of the target defect. The generation unit is used to generate early warning instructions or maintenance work orders corresponding to the risk level.

[0017] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the railway scene micro-defect detection method as described above.

[0018] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for detecting microscopic defects in railway scenarios as described above.

[0019] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the railway scene micro-defect detection method as described above.

[0020] The railway scene micro-defect detection method and device provided by this invention is based on multi-source fusion of visible light data, infrared data, and lidar point cloud data. Based on the multi-dimensional scene features after fusion, the physical quantitative parameters and spatial distance of the target defect are extracted. This allows for a comprehensive and objective quantitative characterization of the severity and spatial location risk of the target defect. Furthermore, based on the objective data collected by these hardware devices, automatic matching can be performed in the railway operation and maintenance knowledge base to achieve an automated closed loop from defect identification and risk assessment to decision command generation. This improves the detection accuracy and comprehensiveness of hidden and complex defects, avoids information loss caused by single data modalities and the subjectivity and lag of manual judgment, and ensures the timeliness of railway operation and maintenance response and the accuracy of handling. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0022] Figure 1 This is one of the flowcharts of the method for detecting microscopic defects in railway scenarios provided by the present invention.

[0023] Figure 2 This is a flowchart illustrating the training steps of the semantic segmentation network provided by the present invention.

[0024] Figure 3 This is the second flowchart of the method for detecting microscopic defects in railway scenarios provided by the present invention.

[0025] Figure 4 This is a schematic diagram of the microscopic defect detection device for railway scenarios provided by the present invention.

[0026] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0028] This invention provides a method for detecting microscopic defects in railway scenarios, aiming to solve the technical problems of incomplete coverage, high false negative rate, inability to quantify physical parameters of defects, and disconnect between detection results and actual operation and maintenance decisions in the existing single-modal detection technology. Figure 1 This is one of the flowcharts illustrating the method for detecting microscopic defects in railway scenarios provided by this invention, such as... Figure 1 As shown, the method includes the following: Step 110: Obtain visible light data, infrared data, and lidar point cloud data of the railway scene to be detected.

[0029] Specifically, firstly, visible light data, infrared data, and lidar point cloud data of the railway scene to be detected can be acquired.

[0030] The railway scene to be detected can include various infrastructures along the railway line and their surrounding environment. For example, the railway scene to be detected can be the railway subgrade, bridges, tunnels, catenary system, power supply equipment, or the environment along the railway safety protection zone, such as slopes, vegetation, surrounding buildings or structures. The power supply equipment can be box-type substations, insulators, etc., and this embodiment of the invention does not specifically limit the scope of the detection.

[0031] Here, visible light data is used to reflect surface information such as color and texture of objects in the railway scene to be inspected. Visible light data can be two-dimensional images or video streams captured by a visible light camera. For example, visible light data can be used to visually identify the crack morphology of concrete surfaces, the rust color and texture of rail surfaces, and the type and growth status of vegetation.

[0032] Here, infrared data is used to reflect the surface temperature distribution information of objects in the railway scene to be inspected. Infrared data can be thermal images or video streams acquired by an infrared thermal imager. For example, infrared data can be used to detect abnormal temperature areas on the surface of concrete structures caused by internal water seepage, or to detect abnormal temperature rise points in overhead contact lines, power supply equipment, etc., caused by poor contact or overload. This information is usually invisible in visible light data.

[0033] LiDAR point cloud data is used to reflect the three-dimensional geometric structure and spatial location information of objects in the railway scene to be detected. LiDAR point cloud data is a collection of a large number of three-dimensional coordinate points (X, Y, Z) obtained by a LiDAR (Light Detection and Ranging) system. Each three-dimensional coordinate point precisely represents a position on the object's surface. For example, LiDAR point cloud data can be used to accurately construct digital three-dimensional models of bridges and tunnels, measure the depth and volume of concrete spalling, or calculate the precise spatial distance between foreign objects beside the railway line and the railway track.

[0034] In a preferred embodiment, a high-pixel visible light camera, a high-resolution infrared thermal imager, and a high-frequency lidar can be integrated into the same acquisition platform, such as an industrial drone, inspection robot, or a dedicated inspection vehicle, to scan or photograph the railway scene to be inspected, thereby simultaneously acquiring visible light data, infrared data, and lidar point cloud data.

[0035] Specifically, during the data collection process, this embodiment can set differentiated collection parameters and flight modes according to the characteristics of different railway facilities and environments to ensure comprehensive and accurate data. Specifically, it employs an oblique photography mode for bridge piers or beams; a circular scanning mode for tunnel linings; a fixed-point hovering shooting mode for overhead contact lines; a close-up shooting mode for power supply equipment; and a cruise shooting mode for the railway line environment.

[0036] To ensure the accuracy of subsequent fusion processing, the acquisition of visible light data, infrared data, and lidar point cloud data can include spatiotemporal alignment of the multi-source data. Specifically, in the spatial dimension, a combination of checkerboard calibration and 3D target calibration methods can be used to complete the extrinsic parameter calibration of the visible light camera, infrared thermal imager, and lidar, obtaining the rotation matrix R and translation vector T. The spatial transformation formula is as follows: ; In the formula, The coordinates of a point in the visible light camera coordinate system. These are the coordinates of a point in the LiDAR coordinate system of the lidar. T is the rotation matrix for transforming the visible light camera coordinate system to the lidar coordinate system, and T is the translation vector. In the time dimension, this embodiment adopts a method of hardware-triggered synchronization as the main approach and software interpolation compensation as a supplement, combined with inertial measurement unit (IMU) data compensation for minor time deviations, to ensure accurate correspondence and consistency of multi-source data in timestamps.

[0037] Step 120: Perform multi-source fusion processing on the visible light data, the infrared data, and the lidar point cloud data to obtain multi-dimensional scene features. Based on the multi-dimensional scene features, identify the target defect in the railway scene to be detected. Based on the multi-dimensional scene features, extract the physical quantification parameters of the target defect and the spatial distance between the target defect and the preset line.

[0038] Specifically, after acquiring visible light data, infrared data, and lidar point cloud data, multi-source fusion processing can be performed on the visible light data, infrared data, and lidar point cloud data to obtain multi-dimensional scene features. Based on the multi-dimensional scene features, target defects in the railway scene to be detected can be identified, and based on the multi-dimensional scene features, physical quantitative parameters of the target defects and the spatial distance between the target defects and the preset line can be extracted.

[0039] Here, multi-source fusion processing aims to deeply integrate the texture and color features contained in visible light data, the temperature features contained in infrared data, and the three-dimensional geometric features contained in lidar point cloud data. Through multi-source fusion processing, the shortcomings of single-modal data in detecting complex scenes can be effectively overcome. For example, for water seepage defects inside concrete with blurred boundaries, visible light images alone are insufficient for identification. However, by fusing the temperature anomaly information from infrared data and the minute deformation information from lidar point cloud data, the detection rate and recognition accuracy will be significantly improved.

[0040] Among them, multi-dimensional scene features are the output results obtained from multi-source fusion processing. Multi-dimensional scene features are a high-dimensional feature representation, in which each feature vector simultaneously contains texture, temperature, and three-dimensional geometric information of a specific location in the scene.

[0041] Here, identifying target defects in a railway scene based on multi-dimensional scene features means taking multi-dimensional scene features as input and using one or more preset recognition models to determine whether a target defect exists in the railway scene to be detected and to determine its location and range.

[0042] The target defects may include microscopic defects in railway infrastructure, such as cracks, spalling, spalling, water seepage, and corrosion in concrete structures; cracks and wear in rails; missing parts and conductor deviation in the overhead contact system; abnormal temperature rise in power supply equipment; and macroscopic environmental risks that affect line safety, such as illegal construction, foreign object encroachment, and vegetation encroachment. This embodiment of the invention does not specifically limit these.

[0043] It should be noted that the types of physical quantification parameters depend on the type of target defect. For example, if the target defect is a concrete crack, the physical quantification parameters may include the crack's length, maximum width, and average width. If the target defect is concrete spalling, the physical quantification parameters may include the area, maximum depth, and volume of the spalled area. If the target defect is abnormal temperature rise in power supply equipment, the physical quantification parameters may include the highest temperature, average temperature, and the difference between the abnormal area and the ambient temperature. This extraction process fully utilizes the advantages of multi-dimensional scene features. For example, the length and width of the crack can be calculated from the segmentation results that fuse visible light and point cloud information, while the spalling depth is directly derived from LiDAR point cloud data.

[0044] Extracting the spatial distance between the target defect and the preset line refers to calculating the real-world three-dimensional spatial distance between the identified target defect, especially the hidden dangers in the external environment of the line, and the critical railway line.

[0045] Here, the preset line can be the center line of the rail, the position of the contact wire, or the boundary line of the railway safety protection zone, etc., and the embodiments of the present invention do not specifically limit it.

[0046] In one specific embodiment, firstly, a subset of 3D coordinates of the LiDAR point cloud corresponding to the target defect is parsed from the multi-dimensional scene features. Based on this subset, the spatial location representation points of the target defect are determined, such as the geometric center point of the defect area or the feature point closest to the preset route. Simultaneously, the 3D spatial equation of the preset route in the LiDAR coordinate system is parsed from the multi-dimensional scene features. Then, the shortest Euclidean distance from the spatial location representation point of the target defect to the 3D spatial equation of the preset route is calculated, or the distance from the spatial location of the target defect to each point in the discrete point set of the preset route is calculated and the minimum value is taken. The result is the spatial distance between the target defect and the preset route. In this process, the multi-dimensional scene features provide a correlation mapping between the defect identification result and the LiDAR point cloud coordinates, allowing the calculation of the spatial distance to be directly based on the fused feature representation without repeatedly retrieving the original point cloud data.

[0047] Step 130: Based on the physical quantification parameters and the spatial distance, match them in the preset railway operation and maintenance knowledge base to obtain the risk level of the target defect.

[0048] Specifically, after obtaining the physical quantification parameters and spatial distance, the risk level of the target defect can be obtained by matching the physical quantification parameters and spatial distance in a preset railway operation and maintenance knowledge base.

[0049] The pre-defined railway operation and maintenance knowledge base is a structured information repository that stores professional knowledge, norms, and standards in the field of railway operation and maintenance. This knowledge base can be constructed based on industry standards and on-site operation and maintenance details. It includes a series of evaluation rules that clearly and quantifiably correlate the physical parameters and spatial distances of different types of target defects with specific risk levels. For example, the pre-defined knowledge base might include rules such as: "For bridge piers, when a crack with a maximum width greater than 0.4 mm is detected, the risk level is determined to be emergency; when large illegal construction machinery is detected within 30 meters of the track, the risk level is determined to be emergency."

[0050] Here, the risk level of a target defect refers to the classification of the potential harm caused by that defect. For example, the risk level of a target defect can be divided into multiple levels such as general, relatively serious, severe, and urgent, or corresponding to warning levels such as Level 1 and Level 2. The risk level of the target defect directly determines the subsequent response measures to be taken.

[0051] In one specific embodiment, when automatically determining the risk level, a multi-level warning line is preset within a pre-defined railway operation and maintenance knowledge base. Specifically, the system automatically matches the spatial distance, the type of the target defect, and physical quantitative parameters with the pre-defined railway operation and maintenance knowledge base. If the target defect is located within the preset first-level warning line range, meaning the characteristics of the target defect are determined to endanger train operation safety, the system will automatically determine it as the corresponding highest risk level and immediately trigger a first-level warning instruction to be pushed to the operation and maintenance emergency system to prompt relevant departments to take immediate action. If the target defect is located within the preset second-level warning line range, meaning the characteristics of the target defect are determined to affect the normal use of equipment, the system will automatically determine it as the corresponding second-highest risk level and trigger a second-level warning instruction to be pushed to relevant operation and maintenance personnel to prompt them to complete the handling within a preset time limit.

[0052] It should be noted that the risk level determination and early warning triggering process mentioned above relies entirely on hardware inputs such as spatial distance and multimodal detection results obtained by data collection terminals like drones. The generation and delivery of instructions are automatically completed through algorithmic logic, without relying on real-time human intervention or secondary assessment. This mechanism, deeply integrated with hardware-quantified data and a pre-set railway operation and maintenance knowledge base, achieves an automated decision-making closed loop from target identification to risk perception, significantly shortening the response cycle for hazard handling.

[0053] Step 140: Generate an early warning instruction or maintenance work order corresponding to the risk level.

[0054] Specifically, after obtaining the risk level of the target defect, a warning instruction or maintenance work order corresponding to the risk level can be generated.

[0055] Warning commands typically target high-risk defects that require immediate or time-limited attention and resolution. A warning command is an alert that can be automatically pushed to relevant operations and maintenance emergency systems or the mobile devices of responsible personnel. For example, for a defect classified as urgent, the system will automatically generate a Level 1 warning command, prompting immediate action.

[0056] Here, a maintenance work order is a standardized work task order. When a target defect requires repair, the system can automatically generate a maintenance work order. In a specific embodiment, the maintenance work order may include the precise location of the target defect, defect type, risk level, detailed physical quantitative parameters, spatial distance, visible light photographs and infrared thermal images of the site, recommended repair plan, designated responsible department or person, and required completion deadline, etc., which are not specifically limited in this embodiment of the invention. After the maintenance work order is generated, it can be automatically pushed to the railway's operation and maintenance management system for subsequent task assignment, execution, and tracking.

[0057] The method provided in this invention is based on multi-source fusion of visible light data, infrared data, and lidar point cloud data. Based on the multi-dimensional scene features after fusion, the physical quantitative parameters and spatial distance of the target defect are extracted. This allows for a comprehensive and objective quantitative characterization of the severity and spatial location risk of the target defect. Furthermore, it enables automatic matching in the railway operation and maintenance knowledge base based on these objective data collected by hardware, realizing an automated closed loop from defect identification and risk assessment to decision command generation. This improves the detection accuracy and comprehensiveness of hidden and complex defects, avoids information loss due to a single data modality and the subjectivity and lag of manual judgment, and ensures the timeliness of railway operation and maintenance response and the accuracy of handling.

[0058] Based on the above embodiments, step 120, which involves multi-source fusion processing of the visible light data, the infrared data, and the lidar point cloud data to obtain multi-dimensional scene features, includes: Step 1201: Extract the visible light features from the visible light data, extract the infrared features from the infrared data, and extract the geometric features from the lidar point cloud data; Step 1202: The visible light feature, the infrared feature, and the geometric feature are spliced ​​together to obtain the spliced ​​feature; Step 1203: Based on the splicing features, generate a query matrix, a key matrix, and a value matrix; Step 1204: Based on the dot product of the query matrix and the key matrix, determine the attention weights, and based on the attention weights, perform a weighted summation of the feature vectors at each position in the value matrix to obtain the multi-dimensional scene features.

[0059] Specifically, firstly, visible light features are extracted from visible light data, infrared features are extracted from infrared data, and geometric features are extracted from lidar point cloud data.

[0060] Visible light features are used to characterize information related to the color and texture of an object's surface. Here, visible light data can be input into a pre-defined backbone network, such as a convolutional neural network, to obtain visible light features.

[0061] Infrared features are used to characterize information related to the temperature distribution pattern on the object's surface. Here, infrared data can be input into a backbone network adapted to thermal imaging to obtain infrared features output by the backbone network adapted to thermal imaging.

[0062] Here, geometric features are used to characterize information related to the object's three-dimensional spatial shape, structure, and position. Due to the unstructured nature of point cloud data, specialized networks for processing point clouds can be used; for example, the PointNet series of networks can extract geometric features from LiDAR point cloud data.

[0063] Then, the visible light features, infrared features, and geometric features can be stitched together to obtain the stitched feature. The stitched feature is obtained by connecting and combining the visible light features, infrared features, and geometric features along the channel dimension. For example, if the dimensions of the visible light features, infrared features, and geometric features are all H×W×C, where H is the height, W is the width, and C is the number of channels, then the dimension of the stitched feature is H×W×3C.

[0064] Next, based on the concatenated features, a query matrix, a key matrix, and a value matrix are generated. These matrices are obtained by linearly transforming the concatenated features. Specifically, three independent, learnable linear projection weight matrices can be used, for example, by linearly mapping the concatenated features through a 1×1 convolutional layer to generate the query matrix, key matrix, and value matrix.

[0065] In a specific implementation, let the visible light characteristics be... Geometric features Infrared features The splicing features are obtained by splicing along the channel dimension. ; Generate Query, Key, and Value through a 1×1 convolutional layer: ; In the formula, , , It is a learnable linear projection weight matrix.

[0066] Finally, attention weights are determined based on the dot product of the query matrix and the key matrix. These attention weights quantify the correlation between different features. The attention weights are calculated by performing a scaling operation and a softmax activation function on the dot product of the query matrix and the key matrix.

[0067] Here, the formula for attention weights is as follows: ; in, Indicates attention weights, Represents the query matrix. Represents the key matrix, This represents the dimension of the key matrix.

[0068] Each value in the attention weight matrix reflects a specific location in the railway scene to be detected, indicating how much attention its features should give to features at other locations.

[0069] Here, the feature vectors at each position in the value matrix are weighted and summed based on attention weights to obtain the formula for multi-dimensional scene features as follows: ; in, Represents multi-dimensional scene features. Represents the query matrix. Represents the key matrix, Represents a value matrix, Indicates splicing characteristics, This represents the attention weight.

[0070] The method provided in this invention employs an attention-based fusion processing approach, enabling the model to automatically learn and quantify the correlation between visible light features, infrared features, and geometric features, and dynamically adjust the attention weights based on this correlation. Therefore, it can adaptively focus on the most critical feature information for different types of target defects, achieving deep and selective feature fusion. Compared to simple feature stitching, the generated multi-dimensional scene features have stronger representational capabilities and anti-interference properties.

[0071] Based on the above embodiments, step 120, which involves identifying the target defect in the railway scene to be detected based on the multi-dimensional scene features, includes: The multi-dimensional scene features are input into a semantic segmentation network to obtain a binary mask image of the target defect output by the semantic segmentation network; the binary mask image is used to characterize the position and outline of the target defect in the railway scene to be detected.

[0072] Specifically, after obtaining the multi-dimensional scene features, the multi-dimensional scene features can be input into the semantic segmentation network to obtain the binary mask image of the target defect output by the semantic segmentation network. The binary mask image is used to represent the position and outline of the target defect in the railway scene to be detected.

[0073] Here, a semantic segmentation network is a deep learning model whose task is to assign a class label to each pixel in the multi-dimensional scene features of the input. Semantic segmentation networks can employ network architectures such as U-Net or DeepLabv3+.

[0074] The binary mask is the output of the semantic segmentation network. A binary mask is a binary image in which pixels identified by the semantic segmentation network as defects in the target are assigned a specific value, such as 1, while all pixels identified as background are assigned another value, such as 0.

[0075] Therefore, binary masks accurately depict the location and outline of target defects in the railway scene to be detected through pixel-level annotations. For example, for a concrete crack, a binary mask will clearly outline the precise shape, direction, and coverage of the crack in the visible light data image using a set of pixels with a value of 1.

[0076] Based on the above embodiments, the step 120 of extracting the physical quantization parameters of the target defect includes: Step 120-1: In the case that the target defect is a crack, the skeleton of the binary mask image is extracted to obtain the centerline skeleton representing the direction of the target defect. Step 120-2: Determine the distance from each crack pixel in the binary mask image to the centerline skeleton, and determine the maximum value among all distances as the maximum width of the target defect, and use the maximum width as the physical quantization parameter.

[0077] Specifically, firstly, when the target defect is a crack, the skeleton is extracted from the binary mask image to obtain the centerline skeleton representing the direction of the target defect. Here, crack specifically refers to a linear or network-like fracture-type defect that appears on the surface of structures such as concrete and rails.

[0078] Here, the centerline skeleton is the result of processing the binary mask image of the crack using a morphological skeletonization algorithm. The centerline skeleton is a single-pixel-width curve that accurately reflects the topological structure and main extension direction of the crack in the two-dimensional plane, i.e., the direction of the crack.

[0079] Then, the distance from each crack pixel in the binary mask image to the centerline skeleton is determined, and the maximum value among all distances is determined as the maximum width of the target defect. The maximum width is used as a physical quantization parameter.

[0080] Here, each crack pixel refers to the set of all pixels in the binary mask image that have a value of 1, i.e., are identified as cracks. This step can be implemented as follows: The maximum width of the crack is calculated using a distance transformation algorithm, where the maximum width... The calculation formula is: ; In the formula, For each crack pixel Distance to the centerline skeleton, The set of crack pixels. The maximum width of the crack can be determined by calculating the Euclidean distance from each crack pixel to the nearest point on the centerline skeleton and finding the maximum value among all distances.

[0081] In addition, the average width of the crack can be calculated using a distance transformation algorithm, by measuring the pixels of each crack. The distance to the centerline skeleton is obtained by averaging.

[0082] Based on the above embodiments, the semantic segmentation network is obtained by iteratively executing the following steps until a preset iteration termination condition is met: Step 210: Obtain training samples and target defect labels corresponding to the training samples; the training samples include sample visible light data, sample infrared data, and sample lidar point cloud data; Step 220: Perform multi-source fusion processing on the sample visible light data, the sample infrared data, and the sample lidar point cloud data to obtain multi-dimensional scene features of the sample; Step 230: Input the multi-dimensional scene features of the sample into the initial semantic segmentation network to obtain the predicted binary mask image output by the initial semantic segmentation network; Step 240: Based on the difference between the predicted binary mask image and the target defect label, determine the target loss, and update the model parameters of the initial semantic segmentation network based on the target loss.

[0083] Specifically, Figure 2 This is a flowchart illustrating the training steps of the semantic segmentation network provided by the present invention, as shown below. Figure 2 As shown, firstly, training samples and their corresponding target defect labels are obtained. The training samples are the dataset used to train the initial semantic segmentation network. Each training sample contains a set of multimodal raw data, namely, sample visible light data, sample infrared data, and sample LiDAR point cloud data.

[0084] Here, the target defect label is the standard answer corresponding to each training sample. For semantic segmentation tasks, the target defect label is a binary mask image that is manually and precisely annotated by domain experts. The target defect label accurately marks the true location and outline of the target defect in the training sample.

[0085] In a specific embodiment of the present invention, in order to solve the technical problem of scarce samples of serious railway defects and special environmental hazards and poor model generalization ability, a generative adversarial network architecture can be used to perform style transfer, morphological transformation, illumination transformation or background replacement on real defect and hazard images, thereby generating a large number of highly realistic virtual defect samples to expand the training samples and solve the problem of insufficient model generalization caused by small samples.

[0086] Furthermore, a transfer learning strategy can be introduced during the training of the initial semantic segmentation network. This involves transferring the model parameters pre-trained on a general dataset to the railway defect and hazard detection task. By freezing the parameters of the first few layers of the backbone network and only fine-tuning the subsequent classification and segmentation layers, high-quality model training can be completed using a small number of real samples.

[0087] Then, the sample visible light data, sample infrared data and sample lidar point cloud data are fused from multiple sources to obtain multi-dimensional scene features of the samples. The multi-dimensional scene information of the samples is then input into the initial semantic segmentation network to obtain the predicted binary mask map output by the initial semantic segmentation network.

[0088] The parameters of the initial semantic segmentation network can be preset or randomly generated; this embodiment of the invention does not impose specific limitations on this.

[0089] Finally, based on the difference between the predicted binary mask image and the target defect label, the target loss is determined, and the model parameters of the initial semantic segmentation network are updated based on the target loss, thus obtaining the semantic segmentation network.

[0090] Here, the target loss can be a combination of cross-entropy loss and Dice loss to address the imbalance problem of disease samples. The formula for the target loss is: ; In the formula, Indicates target loss. For cross-entropy loss, This is a loss for Dice.

[0091] Subsequently, the backpropagation algorithm is used to calculate the gradient of the target loss with respect to all trainable parameters in the initial semantic segmentation network, and an optimizer, such as the Adam optimizer, is used to fine-tune and update the network's model parameters based on this gradient, thereby reducing the target loss.

[0092] The preset iteration termination condition may be that the performance of the initial semantic segmentation network no longer improves on the validation set or reaches the preset number of training rounds. This embodiment of the invention does not specifically limit this.

[0093] Furthermore, to achieve continuous evolution of the initial semantic segmentation network's model capabilities, this embodiment also constructs a continuous learning framework and establishes a data closed loop. Specifically, the trained semantic segmentation network can be deployed on a cloud control platform, and a human-machine collaborative review interface can be designed for experts to quickly review the alarm results output by the model. For all confirmed missed and false alarm cases, the system automatically marks them as high-value difficult samples and sends them to the training database. Through an incremental learning strategy, the model is periodically retrained and fine-tuned using these difficult samples, avoiding catastrophic forgetting without retraining the entire model, thereby achieving continuous evolution and self-optimization of the semantic segmentation network's model capabilities.

[0094] The method provided in this invention iteratively trains on training samples of multimodal data and their corresponding precise labels, enabling the semantic segmentation network to learn the complex mapping relationship between complex features that integrate visible light, infrared, and geometric features and the precise contour of target defects. This ensures that the semantic segmentation network can fully utilize the complementary advantages of multimodal data to accurately identify and segment various target defects, maintaining high segmentation accuracy even in complex backgrounds and when defect features are not obvious.

[0095] Based on the above embodiments, the step 120 of extracting the physical quantization parameters of the target defect includes: Step 310: If the target defect is a water seepage defect, identify the defect area in the railway scene to be detected based on the visible light data, and identify the temperature distribution in the railway scene to be detected based on the infrared data. Step 320: Based on the affected area and the temperature distribution, locate the seepage range of the seepage defect; Step 330: Determine the physical quantification parameters based on the area and / or contour dimensions of the seepage range.

[0096] Specifically, when the target defect is a water seepage defect, the defect area in the railway scene to be inspected is identified based on visible light data, and the temperature distribution in the railway scene to be inspected is identified based on infrared data.

[0097] Among them, water seepage defects are used to reflect the damage to railway concrete structures, such as tunnel linings and bridge box girders, caused by the failure of the waterproof layer or structural cracks, resulting in internal moisture seeping to the surface. Here, the damaged area refers to the visually abnormal area that appears in visible light data as darkening of color, the formation of water stains or patches, or the presence of alkaline corrosion crystals.

[0098] Temperature distribution is used to reflect the heat distribution on the surface of an object. Since water evaporation absorbs heat, areas with water seepage defects are usually represented in infrared data as low-temperature anomalies or specific low-temperature gradients where the local temperature is lower than the surrounding dry area.

[0099] Then, based on the affected area and temperature distribution, the seepage range of the seepage defect can be located. Here, the seepage range refers to the actual physical boundary of the seepage defect on the structural surface. By performing multimodal spatial alignment and feature mapping between visual patches identified by visible light data and low-temperature anomaly areas identified by infrared data, interference from simple surface stains can be eliminated, and the true boundary of water penetration can be accurately defined.

[0100] Finally, physical quantification parameters can be determined based on the area and / or contour dimensions of the seepage range. The area and / or contour dimensions reflect the macroscopic geometric scale of the seepage defect, including the square of the seepage expansion, the maximum major or minor axis, etc. The system converts the seepage range into specific physical values ​​through pixel statistics or point cloud projection calculations.

[0101] The method provided in this invention, by combining the defect area from visible light images with the temperature distribution from infrared thermography, can effectively identify water seepage within concrete structures—a weak-feature defect—and accurately locate its seepage range. This enables an objective quantitative evaluation of the severity of seepage, overcoming the technical challenge of distinguishing between surface stains and actual seepage using a single visible light modality, and significantly improving the detection accuracy of hidden defects in tunnels and bridges.

[0102] Based on the above embodiments, the step 120 of extracting the physical quantization parameters of the target defect includes: Step 41: If the target defect is a rust defect, extract the color and texture features of the rust area in the railway scene to be detected based on the visible light data, and identify the temperature distribution in the railway scene to be detected based on the infrared data. Step 42: Based on the color features, texture features, and temperature distribution, determine the corrosion area of ​​the target defect; Step 43: Determine the physical quantification parameters based on the corrosion area.

[0103] Specifically, when the target defect is a rust defect, the color and texture features of the rust area in the railway scene to be detected are extracted based on visible light data, and the temperature distribution in the railway scene to be detected is identified based on infrared data.

[0104] Here, corrosion defects are used to reflect the damage to the metal substrate caused by oxidation reactions in railway rails, metal fasteners, or steel bridge structures. Among them, color and texture characteristics are used to reflect the inherent visual properties of corrosion defects, such as the typical yellowish-brown or reddish-brown distribution, and the rough, irregular surface texture that accompanies metal peeling.

[0105] Then, the corrosion area of ​​the target defect can be determined based on color features, texture features, and temperature distribution.

[0106] Here, the rust area refers to the total connected areas on the metal surface that have undergone oxidation and exhibit abnormal characteristics. Because rusted areas differ from intact metal surfaces in terms of infrared emissivity and thermal conductivity, combining this with temperature distribution analysis can further eliminate false detections caused by surface dirt or reflected light, thus allowing for a more accurate determination of the rust area.

[0107] Finally, physical quantification parameters can be determined based on the rust area, that is, the rust area can be used as the physical quantification parameter.

[0108] The method provided in this invention identifies corrosion defects through multimodal fusion, uses visual features for preliminary localization, and supplements this with infrared features to verify the attributes of the defects, thereby achieving accurate extraction of the corrosion area. This provides a scientific and quantitative basis for the anti-corrosion maintenance of rails and steel structures, avoiding the problems of strong subjectivity and inability to quantitatively measure corrosion levels in traditional manual inspections.

[0109] Based on the above embodiments, the step 120 of extracting the physical quantization parameters of the target defect includes: Step 51: In the case that the target defect is a contact wire defect, the contact wire is fitted based on the lidar point cloud data to determine the contact wire's guide height, pull-out value, and offset. Step 52: Based on the infrared data, determine the first temperature rise of the insulator in the contact network system to which the contact wire belongs, and the second temperature rise of the box-type substation located in the power supply area of ​​the contact network system; Step 53: Determine the physical quantization parameters based on the guide height, the pull-out value, the offset, the first temperature rise, and the second temperature rise.

[0110] Specifically, when the target defect is a contact wire defect, the contact wire is fitted based on lidar point cloud data to determine the contact wire's guide height, pull-out value, and offset.

[0111] Here, "contact wire defects" refers to out-of-limit geometric parameters or performance degradation of key components in the contact wire system of a railway power supply system. Among these, "contact wire height" refers to the vertical height of the contact wire relative to the rail surface. "Pull-out value" refers to the deviation of the contact wire from the track centerline on the horizontal plane.

[0112] The offset is used to reflect the degree to which the contact wire deviates from the preset standard position due to environmental factors or structural loosening.

[0113] In one specific embodiment, after acquiring the lidar point cloud data of the contact wire section, noise points and non-target points are first removed using a point cloud filtering algorithm. Then, a segmentation method based on elevation threshold or reflection intensity threshold is used to extract a candidate point set belonging to the contact line from the point cloud. Subsequently, a random sampling consensus algorithm is used to iteratively fit the candidate point set, eliminating outliers to obtain a set of three-dimensional coordinate points representing the spatial position of the contact line. Based on the three-dimensional coordinate point set, the least squares method is used to fit the spatial straight line equation or cubic spline curve equation of the contact line, thereby obtaining a precise geometric description of the contact line in three-dimensional space. On this basis, according to the known reference surface information of the track plane, the vertical distance from each point on the contact line to the track surface is calculated, and the minimum value or the corresponding value at the design position is taken as the guide height; the lateral distance from the projection point of the contact line on the horizontal plane to the centerline of the track is calculated as the pull-out value; the measured contact line position is spatially compared with the design standard position, and the deviation vector between the two in three-dimensional space is calculated, and the magnitude or specific direction component of the deviation vector is taken as the offset.

[0114] Then, based on infrared data, the first temperature rise of the insulator in the contact network system to which the contact wire belongs, and the second temperature rise of the box-type substation located in the power supply area of ​​the contact network system can be determined.

[0115] The first temperature rise reflects the difference between the surface temperature of power supply equipment such as insulators and the ambient temperature under operating conditions. The second temperature rise reflects the difference between the surface temperature of power supply equipment such as prefabricated substations and the ambient temperature under operating conditions. By capturing the thermal radiation characteristics of these components using infrared data, localized abnormal overheating phenomena caused by poor contact, insulation degradation, or overload operation can be identified.

[0116] Finally, the physical quantization parameters can be determined based on the guide height, pull-out value, offset, first temperature rise, and second temperature rise. That is, the guide height, pull-out value, offset, first temperature rise, and second temperature rise are all used as physical quantization parameters.

[0117] The method provided in this invention integrates the geometric measurement capabilities of lidar point cloud data with the thermal diagnostic capabilities of infrared data to achieve parallel detection of structural deformation defects and electrical heating defects in the overhead contact system. This allows for a comprehensive assessment of the health status of the overhead contact system from both spatial structural safety and electrical operational safety perspectives.

[0118] Based on the above embodiments, the method further includes: Step 61: Acquire lidar point cloud data of the railway scene to be detected at multiple time points; Step 62: Based on the lidar point cloud data of each time phase, construct a three-dimensional digital twin model of the corresponding period of the time phase; Step 63: Register the three-dimensional digital twin models of different time periods using the same reference coordinate system, and perform differential comparison on the registered three-dimensional digital twin models to obtain the roadbed settlement and / or slope displacement vector as change quantification parameters. Step 64: Based on the change quantification parameters, assess the risk of gradual deformation of the roadbed or slope in the railway scenario to be detected.

[0119] Specifically, LiDAR point cloud data of the railway scene to be detected at different time points are acquired in multiple time phases. Then, based on the LiDAR point cloud data of each time phase, a three-dimensional digital twin model of the corresponding period is constructed.

[0120] Among them, the three-dimensional digital twin model refers to a three-dimensional digital base that is highly consistent with the physical railway infrastructure in terms of geometry and spatial location, constructed in virtual space using lidar point cloud data and / or oblique photography data.

[0121] Furthermore, three-dimensional digital twin models of different time periods can be registered using the same reference coordinate system, and the registered three-dimensional digital twin models can be compared differentially to obtain the roadbed settlement and / or slope displacement vector as quantification parameters of change.

[0122] Among them, the change quantification parameter is used to reflect the deformation trend of railway infrastructure over time. The roadbed settlement is used to reflect the vertical displacement of the roadbed; the slope displacement vector is used to reflect the direction and magnitude of the slope displacement in three-dimensional space.

[0123] Finally, based on the quantification parameters of change, the risk of gradual deformation of the subgrade or slope in the railway scenario under test can be assessed. This step uses a high-precision cross-temporal point cloud registration algorithm to capture minute displacements that are difficult to detect with the naked eye, enabling early identification of potential risks such as landslides and subsidence.

[0124] The method provided in this invention achieves quantitative monitoring of gradually changing risks such as roadbed settlement and slope displacement by constructing three-dimensional digital twin models corresponding to different time periods and performing cross-temporal phase difference comparisons. This overcomes the limitation that a single inspection is insufficient to detect evolving defects and enables early warning of geological disaster risks.

[0125] Based on any of the above embodiments Figure 3 This is the second flowchart of the method for detecting microscopic defects in railway scenarios provided by the present invention, as shown below. Figure 3 As shown, firstly, visible light data, infrared data, and lidar point cloud data of the railway scene to be detected are acquired through sensor terminals. The acquired multi-source data undergoes spatiotemporal calibration and standardization to ensure accurate alignment of different modal information in the spatiotemporal dimensions and uniformity of data format. Then, a cross-modal attention network is used to fuse three modal features, achieving deep interaction and enhancement of texture, temperature, and geometric feature information. Based on this, the system enters a parallel processing stage, where a semantic segmentation network performs pixel-level segmentation of defects to extract detailed defect contour information. Simultaneously, an adaptive detection model is used to detect environmental hazards. The system first detects and identifies macro-environmental risk targets. Then, based on the extracted physical quantification parameters and spatial distance, it matches them in a pre-set railway operation and maintenance knowledge base to automatically obtain the risk level of the target defect. Based on the risk assessment results, it generates warning instructions or maintenance work orders corresponding to the risk level, and simultaneously completes result output, data storage, and integration with third-party platforms, realizing a closed-loop business process for detection and operation and maintenance. Finally, by constructing a three-dimensional digital twin model, the system further conducts slow-deformation risk detection and visualization, thereby achieving comprehensive and multi-dimensional intelligent monitoring and operation and maintenance decision support for railway infrastructure.

[0126] The railway scene micro-defect detection device provided by the present invention is described below. The railway scene micro-defect detection device described below and the railway scene micro-defect detection method described above can be referred to in correspondence.

[0127] Based on any of the above embodiments, the present invention provides a device for detecting microscopic defects in railway scenarios. Figure 4 This is a schematic diagram of the microscopic defect detection device for railway scenarios provided by the present invention, as shown below. Figure 4 As shown, the device includes: The acquisition unit 410 is used to acquire visible light data, infrared data and lidar point cloud data of the railway scene to be detected; Extraction unit 420 is used to perform multi-source fusion processing on the visible light data, the infrared data and the lidar point cloud data to obtain multi-dimensional scene features, identify target defects in the railway scene to be detected based on the multi-dimensional scene features, and extract the physical quantification parameters of the target defects and the spatial distance between the target defects and the preset line based on the multi-dimensional scene features. The matching unit 430 is used to perform a matching operation in a preset railway operation and maintenance knowledge base based on the physical quantification parameters and the spatial distance to obtain the risk level of the target defect. The generation unit 440 is used to generate early warning instructions or maintenance work orders corresponding to the risk level.

[0128] The device provided in this invention is based on multi-source fusion of visible light data, infrared data, and lidar point cloud data. Based on the multi-dimensional scene features after fusion, it extracts the physical quantitative parameters and spatial distance of the target defect. This allows for a comprehensive and objective quantitative characterization of the severity and spatial location risk of the target defect. Furthermore, it enables automatic matching of the objective data collected by these hardware components in the railway operation and maintenance knowledge base, achieving an automated closed loop from defect identification and risk assessment to decision command generation. This improves the detection accuracy and comprehensiveness of hidden and complex defects, avoids information loss due to a single data modality and the subjectivity and lag of manual judgment, and ensures the timeliness of railway operation and maintenance response and the accuracy of handling.

[0129] Based on any of the above embodiments, the extraction unit 420 is specifically used for: Extract visible light features from the visible light data, extract infrared features from the infrared data, and extract geometric features from the lidar point cloud data; The visible light feature, the infrared feature, and the geometric feature are spliced ​​together to obtain the spliced ​​feature; Based on the concatenation features, a query matrix, a key matrix, and a value matrix are generated; Based on the dot product of the query matrix and the key matrix, attention weights are determined, and the feature vectors at each position in the value matrix are weighted and summed based on the attention weights to obtain the multi-dimensional scene features.

[0130] Based on any of the above embodiments, the extraction unit 420 is specifically used for: The multi-dimensional scene features are input into a semantic segmentation network to obtain a binary mask image of the target defect output by the semantic segmentation network; the binary mask image is used to characterize the position and outline of the target defect in the railway scene to be detected.

[0131] Based on any of the above embodiments, the extraction unit 420 is specifically used for: In the case where the target defect is a crack, the skeleton of the binary mask image is extracted to obtain the centerline skeleton representing the direction of the target defect; The distance from each crack pixel in the binary mask image to the centerline skeleton is determined, and the maximum value among all distances is determined as the maximum width of the target defect. The maximum width is used as the physical quantization parameter.

[0132] Based on any of the above embodiments, a training unit is further included, wherein the training unit is specifically used for: Acquire training samples and corresponding target defect labels; the training samples include visible light data, infrared data, and lidar point cloud data. The sample visible light data, the sample infrared data, and the sample lidar point cloud data are subjected to multi-source fusion processing to obtain multi-dimensional scene features of the sample. The multi-dimensional scene features of the sample are input into the initial semantic segmentation network to obtain the predicted binary mask image output by the initial semantic segmentation network. Based on the difference between the predicted binary mask image and the target defect label, the target loss is determined, and the model parameters of the initial semantic segmentation network are updated based on the target loss.

[0133] Based on any of the above embodiments, the extraction unit 420 is specifically used for: In the case that the target defect is a water seepage defect, the defect area in the railway scene to be detected is identified based on the visible light data, and the temperature distribution in the railway scene to be detected is identified based on the infrared data; Based on the affected area and the temperature distribution, the seepage range of the seepage defect can be located; The physical quantification parameters are determined based on the area and / or contour dimensions of the seepage range.

[0134] Based on any of the above embodiments, the extraction unit 420 is specifically used for: When the target defect is a corrosion defect, the color and texture features of the corrosion area in the railway scene to be detected are extracted based on the visible light data, and the temperature distribution in the railway scene to be detected is identified based on the infrared data. Based on the color features, texture features, and temperature distribution, the corrosion area of ​​the target defect is determined; Based on the corrosion area, the physical quantification parameters are determined.

[0135] Based on any of the above embodiments, the extraction unit 420 is specifically used for: In the case where the target defect is a contact wire defect, the contact wire is fitted based on the lidar point cloud data to determine the contact wire's guide height, pull-out value, and offset. Based on the infrared data, the first temperature rise of the insulator in the contact network system to which the contact wire belongs, and the second temperature rise of the box-type substation located in the power supply area of ​​the contact network system are determined. The physical quantization parameters are determined based on the guide height, the pull-out value, the offset, the first temperature rise, and the second temperature rise.

[0136] Based on any of the above embodiments, an evaluation unit is further included, wherein the evaluation unit is specifically used for: The lidar point cloud data of the railway scene to be detected is acquired at multiple time points. Based on the lidar point cloud data of each time phase, a three-dimensional digital twin model of the corresponding period is constructed. The three-dimensional digital twin models of different time periods are registered using the same reference coordinate system, and the registered three-dimensional digital twin models are compared differentially to obtain the roadbed settlement and / or slope displacement vector as change quantification parameters. Based on the aforementioned change quantification parameters, the risk of gradual deformation of the roadbed or slope in the railway scenario to be detected is assessed.

[0137] Figure 5 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 5 As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 can call logical instructions in the memory 530 to execute a method for detecting micro-defects in railway scenes. This method includes: acquiring visible light data, infrared data, and lidar point cloud data of the railway scene to be detected; performing multi-source fusion processing on the visible light data, the infrared data, and the lidar point cloud data to obtain multi-dimensional scene features; identifying target defects in the railway scene to be detected based on the multi-dimensional scene features; extracting physical quantification parameters of the target defects and the spatial distance between the target defects and a preset line based on the multi-dimensional scene features; matching the physical quantification parameters and the spatial distance in a preset railway operation and maintenance knowledge base to obtain the risk level of the target defects; and generating an early warning instruction or maintenance work order corresponding to the risk level.

[0138] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0139] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the railway scene micro-defect detection method provided by the above methods. The method includes: acquiring visible light data, infrared data, and lidar point cloud data of the railway scene to be detected; performing multi-source fusion processing on the visible light data, the infrared data, and the lidar point cloud data to obtain multi-dimensional scene features; identifying target defects in the railway scene to be detected based on the multi-dimensional scene features; extracting physical quantification parameters of the target defects and the spatial distance between the target defects and a preset line based on the multi-dimensional scene features; matching the physical quantification parameters and the spatial distance in a preset railway operation and maintenance knowledge base to obtain the risk level of the target defects; and generating an early warning instruction or maintenance work order corresponding to the risk level.

[0140] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the method for detecting microscopic defects in railway scenes provided by the above methods. The method includes: acquiring visible light data, infrared data, and lidar point cloud data of a railway scene to be detected; performing multi-source fusion processing on the visible light data, the infrared data, and the lidar point cloud data to obtain multi-dimensional scene features; identifying target defects in the railway scene to be detected based on the multi-dimensional scene features; extracting physical quantification parameters of the target defects and the spatial distance between the target defects and a preset line based on the multi-dimensional scene features; matching the physical quantification parameters and the spatial distance in a preset railway operation and maintenance knowledge base to obtain the risk level of the target defects; and generating an early warning instruction or maintenance work order corresponding to the risk level.

[0141] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0142] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0143] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for detecting microscopic defects in railway scenarios, characterized in that, include: Acquire visible light data, infrared data, and lidar point cloud data of the railway scene to be detected; The visible light data, the infrared data, and the lidar point cloud data are fused from multiple sources to obtain multi-dimensional scene features. Based on the multi-dimensional scene features, target defects in the railway scene to be detected are identified. Based on the multi-dimensional scene features, the physical quantification parameters of the target defects and the spatial distance between the target defects and the preset line are extracted. Based on the physical quantification parameters and the spatial distance, a matching is performed in a preset railway operation and maintenance knowledge base to obtain the risk level of the target defect; Generate early warning instructions or maintenance work orders corresponding to the risk level.

2. The method for detecting microscopic defects in railway scenarios according to claim 1, characterized in that, The process of multi-source fusion processing of the visible light data, the infrared data, and the lidar point cloud data to obtain multi-dimensional scene features includes: Extract visible light features from the visible light data, extract infrared features from the infrared data, and extract geometric features from the lidar point cloud data; The visible light feature, the infrared feature, and the geometric feature are spliced ​​together to obtain the spliced ​​feature; Based on the concatenation features, a query matrix, a key matrix, and a value matrix are generated; Based on the dot product of the query matrix and the key matrix, attention weights are determined, and the feature vectors at each position in the value matrix are weighted and summed based on the attention weights to obtain the multi-dimensional scene features.

3. The method for detecting microscopic defects in railway scenarios according to claim 1, characterized in that, The identification of target defects in the railway scene to be detected based on the multi-dimensional scene features includes: The multi-dimensional scene features are input into a semantic segmentation network to obtain a binary mask image of the target defect output by the semantic segmentation network; the binary mask image is used to characterize the position and outline of the target defect in the railway scene to be detected.

4. The method for detecting microscopic defects in railway scenarios according to claim 3, characterized in that, The extraction of physical quantization parameters of the target defect includes: In the case where the target defect is a crack, the skeleton of the binary mask image is extracted to obtain the centerline skeleton representing the direction of the target defect; The distance from each crack pixel in the binary mask image to the centerline skeleton is determined, and the maximum value among all distances is determined as the maximum width of the target defect. The maximum width is used as the physical quantization parameter.

5. The method for detecting microscopic defects in railway scenarios according to claim 3, characterized in that, The semantic segmentation network is obtained by iteratively executing the following steps until a preset iteration termination condition is met: Acquire training samples and corresponding target defect labels; the training samples include visible light data, infrared data, and lidar point cloud data. The sample visible light data, the sample infrared data, and the sample lidar point cloud data are subjected to multi-source fusion processing to obtain multi-dimensional scene features of the sample. The multi-dimensional scene features of the sample are input into the initial semantic segmentation network to obtain the predicted binary mask image output by the initial semantic segmentation network. Based on the difference between the predicted binary mask image and the target defect label, the target loss is determined, and the model parameters of the initial semantic segmentation network are updated based on the target loss.

6. The method for detecting microscopic defects in railway scenarios according to any one of claims 1 to 5, characterized in that, The extraction of physical quantization parameters of the target defect includes: In the case that the target defect is a water seepage defect, the defect area in the railway scene to be detected is identified based on the visible light data, and the temperature distribution in the railway scene to be detected is identified based on the infrared data; Based on the affected area and the temperature distribution, the seepage range of the seepage defect can be located; The physical quantification parameters are determined based on the area and / or contour dimensions of the seepage range.

7. The method for detecting microscopic defects in railway scenarios according to any one of claims 1 to 5, characterized in that, The extraction of physical quantization parameters of the target defect includes: When the target defect is a corrosion defect, the color and texture features of the corrosion area in the railway scene to be detected are extracted based on the visible light data, and the temperature distribution in the railway scene to be detected is identified based on the infrared data. Based on the color features, texture features, and temperature distribution, the corrosion area of ​​the target defect is determined; Based on the corrosion area, the physical quantification parameters are determined.

8. The method for detecting microscopic defects in railway scenarios according to any one of claims 1 to 5, characterized in that, The extraction of physical quantization parameters of the target defect includes: In the case where the target defect is a contact wire defect, the contact wire is fitted based on the lidar point cloud data to determine the contact wire's guide height, pull-out value, and offset. Based on the infrared data, the first temperature rise of the insulator in the contact network system to which the contact wire belongs, and the second temperature rise of the box-type substation located in the power supply area of ​​the contact network system are determined. The physical quantization parameters are determined based on the guide height, the pull-out value, the offset, the first temperature rise, and the second temperature rise.

9. The method for detecting microscopic defects in railway scenes according to any one of claims 1 to 5, characterized in that, The method further includes: The lidar point cloud data of the railway scene to be detected is acquired at multiple time points. Based on the lidar point cloud data of each time phase, a three-dimensional digital twin model of the corresponding period is constructed. The three-dimensional digital twin models of different time periods are registered using the same reference coordinate system, and the registered three-dimensional digital twin models are compared differentially to obtain the roadbed settlement and / or slope displacement vector as change quantification parameters. Based on the aforementioned change quantification parameters, the risk of gradual deformation of the roadbed or slope in the railway scenario to be detected is assessed.

10. A device for detecting microscopic defects in railway scenarios, characterized in that, include: The acquisition unit is used to acquire visible light data, infrared data, and lidar point cloud data of the railway scene to be detected. The extraction unit is used to perform multi-source fusion processing on the visible light data, the infrared data and the lidar point cloud data to obtain multi-dimensional scene features, identify target defects in the railway scene to be detected based on the multi-dimensional scene features, and extract the physical quantification parameters of the target defects and the spatial distance between the target defects and the preset line based on the multi-dimensional scene features. A matching unit is used to perform a matching operation in a preset railway operation and maintenance knowledge base based on the physical quantification parameters and the spatial distance to obtain the risk level of the target defect. The generation unit is used to generate early warning instructions or maintenance work orders corresponding to the risk level.