Track video inspection method and system based on deep learning
Through deep learning and image processing technology, the drone track inspection path and image enhancement are optimized, which solves the problems of track edge information loss and the influence of lighting changes, and realizes efficient and accurate track detection and damage identification.
Patent Information
- Application Number
- CN202510951440.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2045-07-10
AI Technical Summary
Existing drone track inspections fail to reasonably consider the impact of track width limitations and dynamic adjustment of drone flight altitude on image acquisition results, resulting in serious loss of track edge information. Light changes in complex environments affect fluctuations in video acquisition results. Traditional methods are unable to accurately detect track anomalies and complex geometric structures in high dynamic range scenarios.
Through a deep learning-based method, map information is obtained to analyze the effective shooting width of drone inspections, track boundary constraints are defined, and the flight trajectory path is optimized. Combined with dynamic adjustment of light intensity and infrared fill light, image enhancement and grayscale processing are performed, the track center point sequence is extracted, the Hough transform algorithm is used to detect edges, and the pre-trained model is combined to identify damaged areas.
It solves the problems of track edge information loss and the influence of lighting changes, realizes efficient and accurate track detection and damage identification in complex environments, and improves the operational efficiency and image quality of inspection tasks.
Smart Images

Figure CN120779990A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of video inspection, in particular to a track video inspection method and system based on deep learning. BACKGROUND
[0002] Track inspection mainly relies on manual detection, fixed-point equipment inspection and unmanned aerial vehicle inspection and other ways, with the introduction of unmanned aerial vehicle technology, a flexible and efficient way is provided for track inspection, especially the unmanned aerial vehicle shooting combined with high-resolution camera equipment, which can collect track area video data, facilitate subsequent image processing and analysis, in addition, with the evolution of deep learning technology, image segmentation and target detection algorithm has important application potential in track inspection, and part of the existing research uses image algorithm based on deep learning to realize track defect classification.
[0003] However, the existing unmanned aerial vehicle does not reasonably consider the influence of track width limitation and dynamic adjustment of unmanned aerial vehicle flight height on image acquisition effect in the track inspection operation, resulting in serious loss of track edge information, at the same time, the video acquisition effect is greatly fluctuated under complex environmental conditions affected by light intensity change, and the track area features in weak light environment cannot be accurately identified due to insufficient image quality, in addition, for the processing of inspection images, the traditional track damage identification method mainly relies on single scale image segmentation and simple edge detection algorithm, which cannot cope with the accurate detection of track abnormal features and complex geometric structure in high dynamic range scene. SUMMARY
[0004] In view of the above existing problems, the present application is proposed.
[0005] Therefore, the present application provides a track video inspection method and system based on deep learning to solve the problem that the influence of track width limitation and dynamic adjustment of unmanned aerial vehicle flight height on image acquisition effect cannot be reasonably considered, resulting in serious loss of track edge information, at the same time, the video acquisition effect is greatly fluctuated under complex environmental conditions affected by light intensity change, and the track area features in weak light environment cannot be accurately identified due to insufficient image quality, in addition, for the processing of inspection images, the traditional track damage identification method mainly relies on single scale image segmentation and simple edge detection algorithm, which cannot cope with the accurate detection of track abnormal features and complex geometric structure in high dynamic range scene.
[0006] To solve the above technical problems, the present application provides the following technical scheme:
[0007] In a first aspect, the present application provides a track video inspection method based on deep learning, which comprises:
[0008] The map information is acquired to analyze an effective shooting width of unmanned aerial vehicle inspection, a track boundary constraint is defined to analyze an effective width of track shooting, and the effective width is used to determine different flight track paths and evaluate fitness values of the flight track to determine an optimal path, the optimal path is analyzed for turning energy consumption and smooth flight energy consumption, and step-by-step coordinate points are added to determine track segments of the inspection task.
[0009] Based on the illumination intensity of the inspection time, the average illumination visibility of different segments is determined, the weak light collection task in the unmanned aerial vehicle inspection segment is screened, and the infrared light supplement range is calculated, and the video images collected by the unmanned aerial vehicle are dynamically image-enhanced according to different illumination visibilities to form an enhanced image sequence.
[0010] The enhanced image sequence is subjected to grayscale processing, the accumulated grayscale value is calculated to extract a track center point sequence, edge points are extracted based on a grayscale image through a Canny operator, the edge points of a non-track area are filtered to form an effective edge point image, mapping is performed through a Hough transform algorithm, a Hough straight line parameter set is screened through a Hough space cumulative matrix, the track center point sequence is classified according to the transverse position, the transverse coordinate classification position is calculated as a left boundary and a right boundary, and image cropping is performed.
[0011] Based on a pre-trained deep learning model output segmentation probability map, damage category area recognition is performed.
[0012] As a preferred scheme of the track video inspection method based on deep learning, wherein: the analysis of the optimal path turning energy consumption and smooth flight energy consumption, the addition of step-by-step coordinate points to determine the track segment of the inspection task, includes,
[0013] Based on the set of track center line coordinate point data, the track boundary constraint is defined according to the effective shooting width, and the initial coordinate point set is determined.
[0014] According to the coordinate point data and offset coordinate points of the same y direction coordinate, different y direction coordinates and corresponding selected x direction coordinates are searched to determine different flight track paths, and a local path candidate set is formed, and the path selection is evaluated by an adaptability function value of the unmanned aerial vehicle covering each sub-region, wherein the adaptability function value is determined according to the number of covered coordinate points and the total flight distance.
[0015] The path point set based on the highest adaptability function value is used as the optimal path, and the flight turning angle and smooth flight distance data between different coordinate points of the unmanned aerial vehicle in the optimal path are comprehensively considered to evaluate the total energy consumption of the flight smooth energy consumption and the turning energy consumption of the unmanned aerial vehicle.
[0016] If the total energy consumption exceeds the maximum flight energy consumption value of the unmanned aerial vehicle, a sub-step coordinate point is added to the midpoint of the coordinate point segment in the optimal path where the unmanned aerial vehicle needs to turn, and each coordinate point segment task set is taken as a track segment task set according to the added distribution points.
[0017] As a preferred scheme of the track video inspection method based on deep learning, the method comprises the steps of performing dynamic image enhancement to form an enhanced image sequence, and the dynamic image enhancement comprises the steps of,
[0018] The light intensity of the target track area is obtained based on weather forecast data, and the real-time light intensity is recorded by a light sensor.
[0019] The average light visibility in the path range is calculated by the track segment task set and the light intensity.
[0020] The average light visibility of different optimized track segment tasks is compared based on the minimum light visibility threshold, if the average light visibility is greater than or equal to the minimum light visibility threshold, the corresponding segment task is taken as a normal collection task, and if the average light visibility is less than the minimum light visibility threshold, the corresponding segment task is taken as a weak light collection task.
[0021] For the weak light collection task, the weakening factor is determined by analyzing the real-time light intensity and the minimum light intensity according to the corresponding track segment task, the height of the unmanned aerial vehicle is adjusted, and the infrared light compensation range is calculated.
[0022] Based on the video image data collected by the unmanned aerial vehicle in the weak light collection task, low-exposure pixel images and high-exposure pixel images are collected according to different exposure times respectively, and a high dynamic range (HDR) image is synthesized.
[0023] The synthesis weight is dynamically set according to the weakening factor of the track segment task, and the synthesized pixel value is subjected to multispectral fusion with the pixel value of the infrared video stream to form an enhanced image sequence.
[0024] As a preferred scheme of the track video inspection method based on deep learning, the method comprises the steps of calculating the cumulative gray value to extract a track center point sequence, and the method comprises the steps of,
[0025] The enhanced image is subjected to gray processing to form a single-channel gray image, and a degree normalization coefficient is introduced based on the height change of the weak light collection task to calculate the projection gray value of the gray image in the track direction.
[0026] The track center point sequence is extracted according to the peak value distribution of the cumulative gray value.
[0027] As a preferred solution of the track video inspection method based on deep learning described in the present invention, wherein: the Hough line parameter set is screened by the Hough space accumulation matrix, and classified according to the lateral position of the track center point sequence, including:
[0028] Use the Canny operator to extract edge points from the normalized single-channel grayscale image. Dynamically restrict edge points based on the track center point sequence. Determine the specific width based on the ground sampling resolution determined by the actual track width and the dynamic altitude of the drone. Then analyze whether the edge points belong to the valid area.
[0029] According to the edge points of the effective edge point area and the grayscale image, the edge points of the non-track area are filtered out, and only the edge points of the effective track area are retained to form a valid edge point map;
[0030] The effective edge point graph is mapped to the polar coordinate space through the Hough transform algorithm to construct the Hough space accumulation matrix;
[0031] Calculate the maximum value in the Hough space accumulation matrix, which represents the maximum linear density of the accumulated values, and determine the dynamic accumulation threshold;
[0032] Based on the cumulative value, only the lines with high support points are retained. The horizontal position of the track center point sequence is classified according to the Hough line parameter set. Each element is converted into a linear expression and the horizontal coordinate classification position is calculated;
[0033] In the set of straight lines selected in the Hough space, the left and right track boundaries are determined according to the horizontal coordinate classification position;
[0034] The left and right boundary lines are smoothed, and the track body is cropped.
[0035] As a preferred solution of the rail video inspection method based on deep learning of the present invention, wherein: the pre-trained deep learning model outputs a segmentation probability map to perform damage category area recognition, including:
[0036] Based on the cropped track area image data, normalization preprocessing is performed. Based on the pre-trained deep learning model, multi-scale feature extraction is performed on the damaged area in the track cropped image, and a segmentation probability map is output. The predicted damaged area is binarized and the area where the pixel points belong to the damage category is extracted.
[0037] The damage category output of the segmentation model is binarized, the mask is output and the connected domain is calculated, the circumscribed rectangular box of the damage area is extracted, and the rectangular box annotation results are classified.
[0038] As a preferred solution of the rail video inspection method based on deep learning of the present invention, the acquisition of map information to analyze the effective shooting width of the drone inspection includes:
[0039] Based on high-precision track map information, track center line coordinate point data and track basic width data are obtained;
[0040] Based on the unmanned aerial vehicle carrying video acquisition equipment, the flight height is determined according to the required ground sampling accuracy and sampling resolution and the camera angle, and the effective shooting width of the camera is calculated in cooperation with the track basic width data.
[0041] In a second aspect, the present application provides a track video inspection system based on deep learning, comprising,
[0042] The path planning module analyzes the optimal flight trajectory of the unmanned aerial vehicle through the map information, track boundary constraints and effective width, evaluates the energy consumption of turning and smooth flight, and determines the segmented tasks;
[0043] The task classification module calculates the average light visibility of the segments based on the light intensity of the inspection time, filters the weak light acquisition tasks and optimizes the infrared light compensation range;
[0044] The enhanced image processing module uses light dynamic adjustment to dynamically enhance the video images collected, and generates an enhanced image sequence with high contrast;
[0045] The track center detection module performs grayscale on the enhanced image sequence, extracts the cumulative grayscale value curve peak value, generates a track center point sequence to locate the track main body area;
[0046] The track edge extraction module extracts edge points based on the grayscale image through the Canny operator, filters background noise to form an effective edge point image, and locates the left and right boundary lines of the track through Hough transformation;
[0047] The image cropping module crops the image according to the track center point and boundary line data;
[0048] The damage area segmentation module generates a segmentation probability map for the cropped image based on a pre-trained deep learning model, identifies the track damage category and area position.
[0049] In a third aspect, the present application provides a computer device comprising a memory and a processor, the memory storing a computer program, wherein the computer program is executed by the processor to implement any step of the track video inspection method based on deep learning according to the first aspect of the present application.
[0050] In a fourth aspect, the present application provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by the processor to implement any step of the track video inspection method based on deep learning according to the first aspect of the present application.
[0051] The application has the beneficial effects that: by further optimizing the segmented task set, combining with real-time energy consumption, forming a track segmented task set, the effect of converting a long path into a small-scale inspection task of multiple sub-paths is achieved, through grading of the track segmented task, it is ensured that the track collection task can still be efficiently executed under insufficient light conditions, and the definition of the weak light collection task set provides clear guidance for subsequent high-level adjustment and infrared light supplement optimization, solves the problem of ambiguous task priority under complex changes of track light conditions, and through the screening of effective edge points and the Hough transform cumulative space analysis technology, the unstable problem of track boundary detection caused by weak light environment, dynamic height change and track width difference is completely solved, and through the linkage of the track center point and the classified position, the detection and cutting are more accurate, and the operation efficiency of the subsequent track main body cutting task is improved. BRIEF DESCRIPTION OF DRAWINGS
[0052] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed to be used in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0053] Fig. 1 The flowchart of the track video inspection method based on deep learning in embodiment 1.
[0054] Fig. 2 The structure diagram of the track video inspection system based on deep learning in embodiment 1. DETAILED DESCRIPTION
[0055] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the drawings of the specification.
[0056] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application, but the present application can also be implemented in other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of the present application, therefore the present application is not limited to the specific embodiments disclosed below.
[0057] Secondly, the "one embodiment" or "embodiment" referred to herein means that the specific features, structures or characteristics can be included in at least one implementation of the present application. "In one embodiment" appearing in different places in the specification does not mean the same embodiment, nor is it an independent or alternative embodiment that excludes other embodiments.
[0058] Embodiment 1, refer to Figs. 1-2For the first embodiment of the present application, the embodiment provides a track video inspection method based on deep learning, comprising the following steps:
[0059] S1, obtaining map information to analyze the effective shooting width of unmanned aerial vehicle inspection, defining track boundary constraints to analyze the effective width of track shooting, to determine different flight trajectory paths, and to evaluate the fitness value of the flight trajectory to determine the optimal path, to analyze the turning energy consumption and smooth flight energy consumption of the optimal path, and to add step coordinate points to determine the track segmentation of the inspection task;
[0060] Preferably, obtaining map information to analyze the effective shooting width of unmanned aerial vehicle inspection comprises,
[0061] Based on high-precision track map information, obtaining track center line coordinate point data and track basic width data;
[0062] Based on the unmanned aerial vehicle carrying video acquisition equipment, according to the required ground sampling accuracy and sampling resolution and camera angle, the flight height is determined, and the effective shooting width of the camera is calculated in combination with the track basic width data, which is represented as:
[0063]
[0064] Where h represents the flight height of the unmanned aerial vehicle, GSD represents the ground sampling accuracy, r vid represents the sampling resolution, θ represents the camera angle, and W represents the effective shooting width.
[0065] Further, analyzing the turning energy consumption and smooth flight energy consumption of the optimal path, and adding step coordinate points to determine the track segmentation of the inspection task, comprises,
[0066] Based on the set of track center line coordinate point data, the track boundary constraints are defined according to the effective shooting width, and the initial coordinate point set is determined, which is represented as:
[0067] P lef =(x-W,y);
[0068] P right =(x+W,y);
[0069] T0={P lef ,P,P right};
[0070] Where P lef represents the left offset coordinate point, x and y represent the coordinate data of the track center line coordinate point, P right represents the right offset coordinate point, P represents the track center line coordinate point, T0 represents the initialization coordinate point set, including coordinate point data at different positions and corresponding offset coordinate point data;
[0071] According to the coordinate point data of the same y direction coordinate and the offset coordinate point, different y direction coordinates and corresponding selected x direction coordinates are searched to determine different flight trajectory paths, and a local path candidate set is formed. Path selection is evaluated by an adaptability function value of the unmanned aerial vehicle covering each sub-region, wherein the adaptability function value is determined according to the number of covered coordinate points and the total flight distance, and is expressed as:
[0072]
[0073] wherein Fit(T i ) represents the adaptability function value of the i th path set, represents the number of covered coordinate points of the i th path set, represents the total flight distance of the i th path set.
[0074] The path point set based on the highest adaptability function value is the optimal path. Based on the flight turning angle and smooth flight distance data between different coordinate points of the unmanned aerial vehicle on the optimal path, the total energy consumption of the unmanned aerial vehicle is evaluated by comprehensively considering the flight smooth energy consumption and turning energy consumption, and is expressed as:
[0075]
[0076] E tot = E tun +E lin ;
[0077] wherein E lin represents linear flight energy consumption, D(T opt ) represents total flight distance, v represents flight speed, P ho represents hovering power consumption, E tun represents turning power consumption, m represents the total number of coordinate point segments of turning flight, P tun represents turning power consumption coefficient, which is marked by the equipment manufacturer, Δt k represents turning duration, and E tot represents the evaluation total energy consumption.
[0078] If the evaluation total energy consumption exceeds the maximum flight energy consumption value of the unmanned aerial vehicle, then step coordinate points are added to the midpoint of the coordinate point segment in which the unmanned aerial vehicle needs to turn in the optimal path. The number of distribution points can be determined by taking the integer part downward according to the ratio of the evaluation total energy consumption to the maximum flight energy consumption of the unmanned aerial vehicle, and each coordinate point segment task set is taken as an orbit segmentation task set according to the added distribution points.
[0079] By combining the track center line coordinates with the offset coordinate point data selection, searching different y direction coordinates, and corresponding x direction coordinates, the different flight trajectory paths are determined, the effect of constructing the local path candidate set is achieved, the discontinuity problem of track curves, complex terrain or inflection point area is solved, the unmanned aerial vehicle can select the path with the optimal adaptability function in the track along the line flight task, while balancing the coverage range and flight distance, avoiding the risk of redundant path planning and increasing energy consumption;
[0080] By introducing the adaptability function value calculation, the number of covered coordinate points and the total flight distance are combined to reduce the unmanned aerial vehicle flight time and improve the track area coverage rate, and the flight turning angle and smooth flight distance data of the unmanned aerial vehicle between different coordinate points of the optimal path are combined, the effect of optimizing the unmanned aerial vehicle energy consumption calculation is achieved, the turning energy consumption and linear flight energy consumption comprehensive evaluation is more scientific, the energy consumption of the whole path coverage of the track inspection task is made complete evaluation, by dynamically adding the step coordinate point in the turning point, the path flight task is more efficient and the energy consumption distribution is balanced, avoiding the problem of energy consumption overload or conflict with device limits;
[0081] By further optimizing the segmented task set and combining with real-time energy consumption, the track segmented task set is formed, the effect of converting a long path into a small-scale inspection task of multiple sub-paths is achieved, the unmanned aerial vehicle track coverage is more flexible, and the problem of flight power concentration in complex track areas is solved, making the inspection task easier to perform and adapt to different track conditions.
[0082] S2, based on the illumination intensity of the inspection time, determining the average illumination visibility of different segments, screening the weak light collection task in the unmanned aerial vehicle inspection segment, and calculating the infrared light compensation range, according to different illumination visibility, the video image collected by the unmanned aerial vehicle is dynamically image enhanced to form an enhanced image sequence;
[0083] Preferably, the dynamic image enhancement to form an enhanced image sequence comprises,
[0084] Based on the weather forecast data, the illumination intensity of the target track area is obtained, the illumination intensity is defined as the hourly watt per square meter illumination of the time schedule in a day, and the real-time illumination intensity is recorded by the light sensor;
[0085] The average illumination visibility in the path range is calculated by analyzing the track segmented task set and the illumination intensity, which is represented as:
[0086]
[0087] Where V i (t) represents the average illumination visibility of the i th track segmented task, t e and t s respectively represent the end time and the start time of the path video collection, Li (t) represents the real-time light intensity of the i th orbital segment task, d at (t) represents the atmospheric transmittance, μ represents the empirical coefficient of the sensor collecting the atmospheric transmittance, which is determined based on historical experience values;
[0088] The average light visibility of different optimized orbital segment tasks is compared based on the minimum light visibility threshold. If the average light visibility is greater than or equal to the minimum light visibility threshold, the corresponding segment task is regarded as a normal collection task. If the average light visibility is less than the minimum light visibility threshold, the corresponding segment task is regarded as a weak light collection task.
[0089] For a weak light collection task, a weakening factor is determined according to the corresponding orbital segment task by analyzing the real-time light intensity and the minimum light intensity, the height of the unmanned aerial vehicle is adjusted, and the infrared light compensation range is calculated, which is represented as:
[0090]
[0091] where hR i represents the adjusted height value of the i th orbital segment task, λ i represents the weakening factor of the i th orbital segment task, L day represents the minimum light intensity, θR represents the diffusion angle of the infrared light source, SR i represents the infrared light compensation width of the i th orbital segment task.
[0092] Based on the video image data collected by the unmanned aerial vehicle in the weak light collection task, low exposure pixel images and high exposure pixel images are collected according to different exposure times, respectively, different exposure times are determined based on historical experience, and a high dynamic range (HDR) image is synthesized, which is represented as:
[0093] R br = L i (t) · A 2 · k · t ex ;
[0094] R dr = L i (t) · A 2 · k · t br ;
[0095]
[0096] where R br represents the high exposure pixel image, R dr represents the low exposure pixel image, A represents the aperture value, k represents the conversion rate of the camera photosensitive sensor, and the unit value is defined by the device, t ex and t brrespectively represent long and short exposure time, R HDR represents a synthesized pixel image, and represents a minimum value to avoid a small value in the denominator when calculating;
[0097] According to the weakening factor of the track segment task, the synthesis weight is dynamically set, and the synthesized pixel value is multispectral fused with the pixel value of the infrared video stream to form an enhanced image sequence, represented as:
[0098] α = 1-λ i ;
[0099] R fu = α·R HDR +(1-α)·R IR ;
[0100] wherein α represents a synthesis weight, R fu represents a final enhanced image, and R IR represents an infrared light supplement image.
[0101] By analyzing the average light visibility in the path range through the track segment task set and the light intensity, and classifying each segment task in combination with the minimum light visibility threshold, the collection method of the track segment task can adapt to the environmental light conditions in real time, effectively reducing the video information quality decline caused by insufficient light, while ensuring the video collection efficiency in the area with sufficient light;
[0102] By dynamically comparing the real-time light intensity with the minimum light intensity, the weakening factor is determined and the height of the unmanned aerial vehicle is adjusted, so that the unmanned aerial vehicle can appropriately reduce the flight height according to the light condition to improve the ground sampling resolution. Through the result calculated by the weakening factor, in combination with the diffusion angle of the infrared light source, the infrared light supplement width is determined and the light supplement range is optimized, so that the video collection of the track segment task can cover every detail of the target area, avoiding the problem of distortion of important track areas caused by insufficient light supplement range;
[0103] By collecting low and high exposure video frames, and synthesizing high dynamic range (HDR) images after determining different exposure times based on historical experience values, the effect of improving image details is achieved, so that the track video can still maintain clarity in the case of large brightness range distribution. By dynamically setting the synthesis weight through the weakening factor, and multispectral fusing the synthesized HDR image with the pixel value of the infrared video stream, the effect of further enhancing the track image details and texture integrity is achieved;
[0104] By optimizing the track segment task in combination with the dynamic light information in the path, the collection method and video enhancement strategy are automatically adjusted, and finally the effect of accurately matching the collection conditions and ensuring the track video detail clarity with a more efficient image processing method is achieved;
[0105] By grading the track segment task, it is ensured that the track collection task can still be efficiently executed under insufficient light conditions, and the definition of the weak light collection task set provides clear guidance for subsequent height adjustment and infrared light optimization, solving the problem of ambiguous task priority under complex track light conditions. By using the weakening factor of the weak light collection task set for height optimization and infrared light range calculation, combined with the dynamic synthesis of the HDR image, the track task can dynamically adapt to the environment under insufficient light conditions, achieving the effects of spatial distribution optimization and energy utilization efficiency improvement, ensuring the consistency of track video collection quality under complex light conditions, and making the inspection task more balanced and reliable in wide-area track coverage and local feature detection.
[0106] S3, the enhanced image sequence is grayed, the accumulated gray value is calculated to extract the track center point sequence, the edge points are extracted based on the gray image through the Canny operator, the edge points of the non-track area are filtered, the effective edge point map is formed, the mapping is performed through the Hough transform algorithm, the Hough line parameter set is screened through the Hough space accumulation matrix, the horizontal position of the track center point sequence is classified, the horizontal coordinate classification position is calculated as the left boundary and the right boundary, and the image is cropped;
[0107] Preferably, the accumulated gray value is calculated to extract the track center point sequence, including,
[0108] The enhanced image is grayed to form a single-channel gray image, and a degree normalization coefficient is introduced based on the height change of the weak light collection task. The projection gray value of the gray image is calculated in the track direction, that is, the accumulated gray value of each column of pixels, which is represented as:
[0109]
[0110] where P pro (x) represents the cumulative gray value of the x-direction pixel of the gray image, H is the height of the image, represents the number of all y-direction pixels in the image, I gra (x, y) represents the normalized single-channel gray image of coordinates (x, y), h minIR represents the minimum flight height of the weak light task, h represents the current task flight height, I cw (x, y) represents the single-channel gray image of coordinates (x, y);
[0111] The track center point sequence is extracted according to the peak value distribution of the accumulated gray value, wherein the peak value position in the accumulated gray value curve, that is, the large gray value region, represents the center region of the column where the track is located. The gray value of the track center region is usually higher in brightness than the background region in the image. This high difference can accurately reflect the track center position through the change of the cumulative value.
[0112] By gray-scaling the enhanced image to form a single-channel grayscale image, the brightness distribution characteristics of the track area can be intuitively mapped after optimizing the low-light acquisition task, so that the light intensity information of the track area can be accurately reflected in the image data.
[0113] By introducing an altitude change normalization coefficient based on low-light acquisition tasks, the grayscale image is compensated for lighting effects caused by flight altitude adjustments, reducing the uneven grayscale distribution of the track image caused by differences in drone altitude. By calculating the cumulative grayscale value of each column of pixels and combining it with a unified projection of the track direction, a curve distribution that directly reflects the maximum grayscale value of the track center area is constructed, achieving the effect of locating the track center column using spatially distributed data. This calculation of cumulative grayscale values fully utilizes the steady-state characteristics of the track area brightness data in the longitudinal space of the image, thereby effectively eliminating background noise interference outside the track boundary. Compared with traditional binarization or edge detection algorithms, this technology for directly locating the track center based on grayscale value changes improves the robustness of track center identification and reduces the dependence of complex algorithms on computing resources.
[0114] The track center point sequence is extracted by accumulating the peak distribution of the grayscale value curve. Combined with the image characteristics of low-light acquisition tasks, it accurately reflects the area with large brightness values at the center of the track column, achieving the effect of directly describing the track direction with point sequences, greatly reducing track detection errors caused by complex lighting or background interference.
[0115] The integration of multispectral data under weak light conditions directly enables the dynamic prediction of the track center, making the target track positioning in inspection tasks more accurate and stable. At the same time, this extraction method based on dynamic normalization of grayscale data ensures the consistency of image data in complex lighting environments, thereby providing reliable data support and dynamic adaptation capabilities for the complete track analysis process.
[0116] Furthermore, the Hough line parameter set is screened by the Hough space accumulation matrix and classified according to the lateral position of the track center point sequence, including:
[0117] The Canny operator is used to extract the edge points of the normalized single-channel grayscale image. The edge points are dynamically restricted based on the track center point sequence. The specific width is determined based on the ground sampling resolution determined by the actual width of the track and the dynamic height of the drone. The edge points are then analyzed to see whether they belong to the valid area. This is expressed as:
[0118] W bd =GSD h ·δ;
[0119]
[0120] Among them GSD hGround sampling resolution of the height of the UAV, δ represents the actual width of the track, R(x, y) represents the effective area of the edge point, x c x represents the x-direction coordinate value of the track center point b x represents the x-direction coordinate value of the edge point coordinate
[0121] According to the edge point of the edge point effective area and the gray image, the edge points of the non-track area are filtered, only the edge points of the track effective area are reserved, and an effective edge point image is formed, which is represented as:
[0122] E res (x you ,y you )=R(x,y)·E(x,y);
[0123] Where E(x, y) represents the edge point image, E res (x you ,y you ) represents the effective edge point image
[0124] The effective edge point image is mapped to the polar coordinate space by the Hough transform algorithm to construct the Hough space cumulative matrix C ρ,θ , which is represented as:
[0125] ρ=x you cosθ+y you sinθ;
[0126] Where ρ represents the shortest distance from the effective edge point to the straight line, x you and y you represent the coordinate data of the effective edge point, θ represents the angle between the effective edge point and the coordinate axis, and C ρ,θ represents the cumulative value (the number of points on the straight line). The matrix represents the aggregation degree of all edge points in the motor area
[0127] The maximum value in the Hough space cumulative matrix is calculated, which represents the maximum straight line density of the cumulative value, and a dynamic cumulative threshold is determined, which is represented as:
[0128] T caa =β·C max ;
[0129] Where T caa represents the dynamic cumulative threshold, β represents the proportion coefficient, which is determined based on experimental analysis, and C max represents the maximum value of the Hough cumulative matrix
[0130] Based on the cumulative value, only the straight line with high support point number is reserved, and the constraint condition is represented as:
[0131] L hou ={(ρ,θ)|C ρ,θ≥T caa};
[0132] where L hou represents all the Hough line parameter sets satisfying the cumulative threshold, each element (p, q) represents a straight line;
[0133] According to the Hough line parameter set, the horizontal position of the track center point sequence is classified, and the horizontal coordinate classification position is converted into a straight line expression, which is represented as:
[0134]
[0135] where x lin represents the horizontal coordinate classification position;
[0136] In the straight line set screened in the Hough space, the left and right track boundary lines are determined according to the horizontal coordinate classification position, and the Hough line with a horizontal coordinate classification position less than x c is defined as the left boundary, and the Hough line with a horizontal coordinate classification position greater than x c is defined as the right boundary;
[0137] The left and right boundary lines are smoothed, and the track body is cropped.
[0138] By using the Canny operator to extract the edge points of the normalized single-channel grayscale image, and combining the track center point sequence for dynamic restriction, the accurate edge point extraction effect based on the track center point is achieved, avoiding the interference of background information on the track edge extraction, and improving the accuracy of the edge points;
[0139] By introducing the dynamic height of the unmanned aerial vehicle and the ground sampling resolution combined with the actual width of the track to calculate the effective area width, the edge points are strictly constrained whether they belong to the effective track range, not only solving the data redundancy problem that the traditional edge detection may expand to the area outside the track, but also realizing the dynamic adjustment of the track area, so that the track detection algorithm can adapt to the diversity needs of different flight tasks and lighting environments, while maintaining the consistency and dynamic maneuverability of the track detection data;
[0140] By performing Hough transform calculation on the effective edge point graph, it is successfully mapped to the polar coordinate space and the Hough space cumulative matrix is constructed, achieving the effect of analyzing the straight line aggregation area through spatial data, and converting the spatial distribution characteristics of the track edge points into high-density straight line data in the cumulative matrix in the Hough space. This spatial aggregation characteristic avoids the quality interference of image noise or discrete edge points on the straight line fitting, thereby improving the stability of the track boundary line fitting;
[0141] By calculating the maximum value in the Hough space accumulation matrix and determining the dynamic accumulation threshold, combined with the dynamic adjustment of the scale factor, the optimization process in track edge verification is completed, achieving the effect of automatically filtering out lines with low support points.
[0142] By classifying the line parameters based on the Hough line and the lateral position of the track center point sequence, the line is accurately divided into the left and right track boundaries on the horizontal axis. This achieves the effect of dynamically screening the boundaries based on the track center point. It can effectively adapt to track width differences, curved areas and background noise interference, while ensuring the geometric consistency of track detection and the accuracy of track boundary line classification.
[0143] Through the screening of effective edge point maps and Hough transform cumulative space analysis technology, the instability problem of track boundary detection caused by weak light environment, dynamic height changes and track width differences is completely solved. The linkage between the track center point and the classification position makes the detection and cropping more accurate, while improving the operational efficiency of subsequent track body cropping tasks, providing a highly robust technology chain from data collection to precise processing for track inspection tasks.
[0144] S4, based on the pre-trained deep learning model output segmentation probability map, performs damage category area recognition;
[0145] Preferably, the segmentation probability map output by the pre-trained deep learning model is used to identify the damage category area, including:
[0146] Based on the cropped track area image data, normalization preprocessing is performed. Based on the pre-trained deep learning model (PSPNet model), multi-scale feature extraction is performed on the damaged area in the track cropped image, and a segmentation probability map is output. The predicted damaged area is binarized and the area where the pixel points belong to the damage category is extracted.
[0147] The damage category output of the segmentation model is binarized, the mask is output and the connected domain is calculated, the circumscribed rectangular box of the damage area is extracted, and the rectangular box annotation results are classified.
[0148] By using the pre-trained PSPNet model to segment the cropped images and leveraging its multi-scale feature extraction capabilities, we successfully captured details of damaged areas of varying sizes and types during track damage detection. Because PSPNet focuses on multi-scale feature extraction and can extract both global and local features through a pyramid pooling module, it ensures that both small localized damage, such as cracks and wear, and large-scale track anomalies, such as wear and foreign object coverage, can be accurately identified.
[0149] By performing connected component calculation on the binary mask, the bounding rectangle of the damage area is extracted, and the connected component in the damage mask is converted into the bounding rectangle, which can directly describe the range and spatial position of the track damage.
[0150] By classifying and labeling the extracted rectangular frame, the detected damage can be directly classified into specific categories such as cracks and wear, and the position data of the bounding rectangle can be combined to provide complete and clear results for the inspection report.
[0151] The embodiment also provides a track video inspection system based on deep learning, comprising,
[0152] Path planning module: analyze the optimal flight trajectory of the unmanned aerial vehicle through map information, track boundary constraints and effective width, evaluate the energy consumption of turning and smooth flight, and determine the segmented tasks;
[0153] Task classification module: based on the light intensity of the inspection time, calculate the average light visibility of the segment, filter the weak light collection task and optimize the infrared light compensation range;
[0154] Enhanced image processing module: use light dynamic adjustment to dynamically enhance the collected video images, and generate enhanced image sequences with high contrast;
[0155] Track center detection module: grayscale the enhanced image sequence, extract the cumulative grayscale value curve peak, and generate the track center point sequence to locate the track main body area;
[0156] Track edge extraction module: extract edge points based on the grayscale image through the Canny operator, filter background noise to form effective edge point images, and locate the left and right boundary lines of the track through Hough transformation;
[0157] Image cropping module: crop the image according to the track center point and boundary line data;
[0158] Damage area segmentation module: based on the pre-trained deep learning model, generate a segmentation probability map for the cropped image, and identify the track damage category and area position.
[0159] The embodiment also provides a computer device suitable for the track video inspection method based on deep learning, comprising: a memory and a processor; the memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions to realize the track video inspection method based on deep learning as proposed in the above embodiment.
[0160] The computer device can be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected by a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is configured to perform wired or wireless communication with an external terminal. The wireless communication can be achieved by WIFI, a carrier network, NFC (Near Field Communication) or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, a trackball or a touchpad arranged on the shell of the computer device, or an external keyboard, a touchpad or a mouse.
[0161] The embodiment also provides a storage medium having a computer program stored thereon, and the computer program is executed by a processor to implement the track video inspection method based on deep learning provided in the above embodiment. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic storage, a flash memory, a magnetic disk or an optical disk.
[0162] To sum up, the application achieves the effect of converting a long path into a small-scale inspection task of multiple sub-paths by further optimizing the segmented task set in combination with real-time energy consumption to form a track segmented task set, guarantees that the track collection task can still be efficiently executed under insufficient light conditions through grading of the track segmented task, and provides clear guidance for subsequent high-degree adjustment and infrared light supplement optimization through definition of the weak light collection task set, solves the problem of ambiguous task priority under complex changes in track light conditions, completely solves the unstable problem of track boundary detection caused by weak light environment, dynamic height changes and track width differences through screening of effective edge point graphs and Hough transform cumulative space analysis technology, and makes detection and cropping more accurate through linkage of the track center point and the classified position, while improving the operation efficiency of the subsequent track main body cropping task.
[0163] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application but not limit the present application. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced equivalently without departing from the spirit and scope of the technical solutions of the present application, and all of them should be covered in the scope of the claims of the present application.
Claims
1. A track video inspection method based on deep learning, characterized in that: include: Obtain map information to analyze the effective shooting width of the drone inspection, define track boundary constraints to analyze the effective width of the track shooting, determine different flight trajectory paths, evaluate the fitness value of the flight trajectory to determine the optimal path, analyze the turning energy consumption and stable flight energy consumption of the optimal path, and add step-by-step coordinate points to determine the track segments of the inspection task; Based on the light intensity during the inspection time, the average light visibility of different segments is determined, the low-light acquisition tasks in the drone inspection segment are screened, and the infrared fill light range is calculated. The video images collected by the drone are dynamically enhanced according to different light visibility to form an enhanced image sequence; Grayscale the enhanced image sequence, calculate the accumulated grayscale value to extract the track center point sequence, extract edge points based on the grayscale image using the Canny operator, filter the edge points in the non-track area to form a valid edge point map, map it using the Hough transform algorithm, filter the Hough line parameter set using the Hough space accumulation matrix, classify the track center point sequence according to its horizontal position, calculate the horizontal coordinate classification position as the left and right boundaries, and perform image cropping; The pre-trained deep learning model outputs a segmentation probability map for damage category area recognition.
2. The track video inspection method based on deep learning according to claim 1, characterized in that: The above analysis of the turning energy consumption and stable flight energy consumption of the optimal path adds step-by-step coordinate points to determine the track segments of the inspection task. include, Based on the set of track centerline coordinate point data, track boundary constraints are defined according to the effective shooting width, and an initial coordinate point set is determined; Based on the coordinate point data of the same y-direction coordinate and the offset coordinate point, different y-direction coordinates and the corresponding selected x-direction coordinates are searched to determine different flight trajectory paths and form a local path candidate set. The path selection is evaluated by the fitness function value of each sub-area covered by the drone, where the fitness function value is determined according to the number of covered coordinate points and the total flight distance; The optimal path is determined based on the set of path points with the highest fitness function value. The total energy consumption of the drone is evaluated based on the flight turning angle and stable flight distance data between different coordinate points on the optimal path, taking into account the flight stability energy consumption and turning energy consumption of the drone. If the total energy consumption is evaluated to exceed the maximum flight energy consumption of the UAV, a step-by-step coordinate point is added to the midpoint of the coordinate point segment where the UAV needs to turn in the optimal path, and each coordinate point segment task set is used as a track segment task set based on the added distribution points.
3. The track video inspection method based on deep learning according to claim 2, characterized in that: The dynamic image enhancement is performed to form an enhanced image sequence, include, The light intensity of the target track area is known based on weather forecast data, and the light sensor records the real-time light intensity; The average light visibility within the path range is calculated through track segmentation task set and light intensity analysis; The average light visibility of different optimized track segmentation tasks is compared based on the minimum light visibility threshold. If the average light visibility is greater than or equal to the minimum light visibility threshold, the corresponding segmentation task is treated as a normal collection task. If the average light visibility is less than the minimum light visibility threshold, the corresponding segmentation task is treated as a low-light collection task. For low-light acquisition tasks, the system divides the tasks into corresponding track segments, determines the attenuation factor by analyzing the real-time light intensity and the minimum light intensity, adjusts the drone's altitude, and calculates the infrared fill light range. Based on the video image data collected by the UAV in the low-light acquisition mission, low-exposure pixel images and high-exposure pixel images are collected according to different exposure times, and high dynamic range HDR images are synthesized; The synthesis weight is dynamically set according to the attenuation factor of the track segmentation task, and the synthesized pixel value is multi-spectrally fused with the pixel value of the infrared video stream to form an enhanced image sequence.
4. The track video inspection method based on deep learning according to claim 3, characterized in that: The calculation of the accumulated grayscale value to extract the track center point sequence includes: The enhanced image is gray-scaled to form a single-channel grayscale image. A normalization coefficient is introduced based on the height variation of the low-light acquisition task, and the projection grayscale value of the grayscale image is calculated according to the track direction. The track center point sequence is extracted based on the peak distribution of the accumulated grayscale values.
5. The track video inspection method based on deep learning according to claim 4, characterized in that: The Hough line parameter set is screened by the Hough space accumulation matrix and classified according to the lateral position of the track center point sequence. include, Use the Canny operator to extract edge points from the normalized single-channel grayscale image. Dynamically restrict edge points based on the track center point sequence. Determine the specific width based on the ground sampling resolution determined by the actual track width and the dynamic altitude of the drone. Then analyze whether the edge points belong to the valid area. According to the edge points of the effective edge point area and the grayscale image, the edge points of the non-track area are filtered out, and only the edge points of the effective track area are retained to form a valid edge point map; The effective edge point graph is mapped to the polar coordinate space through the Hough transform algorithm to construct the Hough space accumulation matrix; Calculate the maximum value in the Hough space accumulation matrix, which represents the maximum linear density of the accumulated values, and determine the dynamic accumulation threshold; Based on the cumulative value, only the lines with high support points are retained. The horizontal position of the track center point sequence is classified according to the Hough line parameter set. Each element is converted into a linear expression and the horizontal coordinate classification position is calculated; In the set of straight lines selected in the Hough space, the left and right track boundaries are determined according to the horizontal coordinate classification position; The left and right boundary lines are smoothed, and the track body is cropped.
6. The track video inspection method based on deep learning according to claim 5, characterized in that: The pre-trained deep learning model outputs a segmentation probability map to identify damage category areas, including: Based on the cropped track area image data, normalization preprocessing is performed. Based on the pre-trained deep learning model, multi-scale feature extraction is performed on the damaged area in the track cropped image, and a segmentation probability map is output. The predicted damaged area is binarized and the area where the pixel points belong to the damage category is extracted. The damage category output of the segmentation model is binarized, the mask is output and the connected domain is calculated, the circumscribed rectangular box of the damage area is extracted, and the rectangular box annotation results are classified.
7. The track video inspection method based on deep learning according to claim 2, characterized in that: The map information is obtained to analyze the effective shooting width of the drone inspection. include, Based on high-precision track map information, obtain track centerline coordinate point data and track basic width data; Based on the video acquisition equipment carried by the drone, the flight altitude is determined according to the required ground sampling accuracy, sampling resolution and camera angle, and the effective shooting width of the camera is calculated in conjunction with the basic track width data.
8. A deep learning-based track video inspection system, based on the deep learning-based track video inspection method according to any one of claims 1 to 7, characterized in that: include, Path planning module: Analyzes the optimal flight trajectory of the UAV through map information, track boundary constraints and effective width, evaluates the energy consumption of turning and stable flight, and determines the step-by-step mission segmentation; Task classification module: Based on the light intensity during the inspection time, calculate the average light visibility of each segment, select low-light collection tasks, and optimize the infrared fill light range; Enhanced image processing module: uses dynamic illumination adjustment to perform dynamic image enhancement on the captured video images, generating enhanced image sequences with high contrast; Track center detection module: grayscales the enhanced image sequence, extracts the peak value of the accumulated grayscale value curve, and generates a track center point sequence to locate the track main area; Track edge extraction module: This module extracts edge points based on the grayscale image using the Canny operator, filters background noise to form a valid edge point map, and locates the left and right boundary lines of the track using the Hough transform. Image cropping module: crops the image according to the track center point and boundary line data; Damage area segmentation module: Based on the pre-trained deep learning model, it generates segmentation probability maps for cropped images and identifies track damage categories and area locations.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the track video inspection method based on deep learning are implemented in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the track video inspection method based on deep learning are implemented.
Citation Information
Patent Citations
Robot weak light environment grabbing detection method based on multi-task sharing network
CN112949452A
Steel logistics route planning method and system based on navigation positioning
CN114706397A
Weather phenomenon automatic identification method under dark light condition
CN117746374A
Lithology identification system and method based on cooperation of double mechanical arms
CN118288262A
Bridge deflection intelligent accurate measurement method based on artificial intelligence
CN119067967A
Cited By
Railway track abrasion line detection system and method
CN121019646A