Railway overhead line system unmanned aerial vehicle intelligent inspection method and system

By using a deep learning model on the drone payload and gimbal closed-loop control, efficient and automated inspection of railway catenary equipment has been achieved, solving the problems of low inspection efficiency and low automation under manual operation, and improving the safety and data quality of inspection.

CN121397191APending Publication Date: 2026-01-23CHINA RAILWAY DESIGN GRP CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511954889.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-23
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

The inspection of railway power supply contact network equipment relies on manual operation, has a low degree of automation, is difficult to detect defects, is inefficient, and lacks closed-loop control, resulting in slow inspection speed and poor data consistency.

Method used

The system employs a deep learning model on the drone payload for real-time detection and disease identification, combined with gimbal closed-loop control to achieve precise rotation, intelligent zoom, focus, and photography. It also integrates edge inference acceleration and video encoding/network streaming technology to achieve efficient and stable inspection operations.

Benefits of technology

It enables fully automated operation of drones in flight missions, reduces the intensity of manual operation and the risks of high-altitude operations, improves the level of automation and safety of operations, reduces the rate of missed detections and false detections, and improves the level of data intelligence and management capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121397191A_ABST
    Figure CN121397191A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle intelligent inspection method and system for a railway overhead line system. According to the method, real-time detection is carried out on a railway overhead line system facility target through an unmanned aerial vehicle load end by using a deep learning model, and precise rotation, intelligent zooming and focusing and centering reset are realized through holder closed-loop control; an edge reasoning acceleration framework and a video coding / network flow pushing module can be combined to improve the processing real-time performance and the data transmission efficiency. The scheme has the advantages of high integration level, high response speed, simplicity and convenience in operation and the like, and is suitable for railway overhead line system unmanned aerial vehicle intelligent inspection application scenes.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of unmanned aerial vehicle intelligent inspection and computer vision technology, and particularly relates to a railway catenary unmanned aerial vehicle intelligent inspection method and system. BACKGROUND

[0002] With the continuous expansion of China's railway network, the large-scale laying of high-speed, heavy-load, and high-cold, high-altitude complex lines puts higher requirements on the operation stability of railway power supply catenary equipment. As the core guarantee for safe train operation, its health status is directly related to train safety.

[0003] At present, the inspection of railway power supply catenary equipment is mainly manual pole climbing or manual control of unmanned aerial vehicles, and the disease is interpreted by visual means. This mode has the following outstanding problems: 1. Manual pole climbing needs to be carried out during the window time, and the online complexity is low.

[0004] 2. The manual control of unmanned aerial vehicles has low automation and relies on manual operation: the flight route, gimbal rotation and lens zoom of the unmanned aerial vehicle need to be controlled manually, and the heading angle and pitch angle of the gimbal and the imaging parameters cannot be automatically adjusted according to the device position. The efficiency is highly dependent on the operator's experience.

[0005] 3. Disease detection is difficult and inefficient: the disease of the catenary equipment is often small in size, diverse in shape and subtle in change. Due to these characteristics, the traditional method relying on manual interpretation not only has high recognition difficulty and high omission rate, but also has low efficiency and insufficient reliability when facing a large number of inspection images.

[0006] 4. Lack of closed-loop control: the existing operation scheme lacks a closed-loop automation process from "target detection-precise alignment-intelligent zoom focusing-photographing", and it is difficult to realize the integration of detection and imaging, resulting in slow inspection speed and poor data consistency. SUMMARY

[0007] Therefore, the purpose of the present application is to provide a railway catenary unmanned aerial vehicle intelligent inspection method and system, which aims at the technical difficulties of small feature size, complex shape and strong background interference of railway power supply catenary equipment disease, and realizes real-time detection and disease identification of catenary equipment by combining a depth learning model with an unmanned aerial vehicle payload end. Precise rotation, intelligent zoom, focusing, and photographing are completed through gimbal closed-loop control, and edge inference acceleration and video encoding / network streaming technology are integrated to realize efficient and stable inspection operation.

[0008] In order to achieve the above purpose, a railway catenary unmanned aerial vehicle intelligent inspection method of the present application comprises the following steps: Data acquisition: acquiring inspection images collected during the inspection process of the unmanned aerial vehicle according to the preset flight route; Multi-level target detection: At least three levels of target detection are performed on the inspection image, the first level of target detection is used to locate the main structure of the catenary, and based on the positioning result of the first level of target detection, small components are identified in the second level of target detection for the local region of interest; based on the positioning result of the second level of target detection, local disease identification is performed in the third level of target detection; Adaptive pan-tilt control: based on the results of the multi-level target detection, the offset of the target from the image center is calculated, and the pan-tilt is driven to rotate and / or zoom to center the target in the image and achieve a preset imaging size; Closed-loop imaging: After the pan-tilt completes alignment and zooming, autofocus and high-definition shooting are performed to complete the inspection of the target; Data upload and cloud collaboration: the images, detection results and UAV state data obtained during the inspection are uploaded to the cloud server through the wireless network, and the cloud server performs specific disease judgment and protection warning.

[0009] Further preferably, the preset flight path is laid out according to the following rules: The line connecting the catenary poles is used as the flight path; one detection point is set on each side of the post; the height of the detection point in front of the post is consistent with that of the detection point behind the post; The route of the UAV when performing forward inspection is used as the forward view angle flight path, which is connected by the detection points in front of each post; The route of the UAV when returning is used as the reverse view angle flight path; the reverse view angle flight path is connected by the detection points behind each post.

[0010] Further preferably, when the UAV reaches the detection point, adaptive pan-tilt control is automatically triggered to identify, align, zoom, focus and shoot the target.

[0011] Further preferably, the first level of target detection: a first deep learning model is used to analyze the image frame by frame, identify and locate the main structure of the catenary, and output the bounding box information of the main structure of the catenary; the second level of target detection: based on the output bounding box information, generate a region of interest according to a preset expansion coefficient, and use a second deep learning model optimized by multi-scale feature fusion in the region of interest to identify small components; the third level of target recognition: according to the identified small components, use a third deep learning model to perform defect detection and determine whether at least one of cracks, damage and missing exists, and output a detection box when it exists.

[0012] Further preferably, based on the results of the multi-level target detection, the pixel offset of the target from the image center is calculated, including: according to the center of the detection box output by the third level of target detection and the center of the image ; the pixel offset is calculated as ; is the horizontal resolution of the image, and is the vertical resolution.

[0013] Further preferably, when the gimbal is driven to rotate according to the yaw rotation amount and the pitch rotation amount calculated according to the following formula, the yaw rotation amount , and the pitch rotation amount ; is the horizontal focal length of the image, and is the vertical focal length of the image.

[0014] The absolute pose of the gimbal after rotation is read and compared with the original pose to obtain the actual yaw rotation amount and the actual pitch rotation amount ; the residual error is calculated as If both , exceed the preset threshold, the values of , are taken as new relative values to issue a rotation command again until the residual error is suppressed within the preset threshold or the maximum number of compensations is reached. After the initial rotation is completed, the updated frame is read and the deviation of the target center from the field center is re-detected. If there is still a residual error exceeding the threshold: the above angle calculation and rotation process are repeated until the deviation converges or the maximum number of iterations is reached, where R represents the residual error, represents the horizontal deviation of the target center from the field center, and represents the vertical deviation of the target center from the field center.

[0015] Further preferably, when the gimbal is driven to zoom, the target zoom factor is determined in the following way: Based on the ratio of the pixel size of the target in the image to the desired imaging size, an initial magnification is calculated by introducing a marginal coefficient; The initial magnification is constrained by the minimum and maximum zoom magnification of the gimbal camera to obtain the final target zoom factor.

[0016] Further preferably, in the adaptive gimbal control step, when multiple targets to be inspected are simultaneously identified in a single frame image, the system generates an imaging task sequence according to the horizontal spatial relationship thereof in the image. All detection boxes are sorted according to the horizontal coordinates of the center points in a continuous left-to-right alignment order, so that the gimbal rotation path is the shortest and the direction is consistent, thereby avoiding unnecessary back-and-forth movements. The generated target sequence can be represented as: wherein, is the horizontal coordinate of the center of the i-th target detection box.

[0017] When starting to process a certain target, first calculate the center offset of the target in the image, and map the offset to the pitch angle and yaw angle that need to be adjusted according to the current lens parameters; the gimbal performs a preliminary rotation to make the target enter the center area of the screen; after the initial rotation is completed, wait for the updated image of the next frame, detect the position of the current target in the new frame again, and recalculate the deviation between the target and the center of the screen to determine whether further compensation is needed.

[0018] Further preferably, in the data uploading and cloud coordination, there is also an end-to-end video push streaming step: During video transmission, the captured image frames are buffered using a frame queue; wherein, : the video frame queue length at the i-th frame moment, : the number of new frames enqueued at the i-th frame moment, : the number of frames dequeued and started to encode at the i-th frame moment, : the maximum number of frames that the queue can accommodate, i.e., when the number of frames exceeds the threshold , the oldest frame is discarded; : the threshold at which the oldest frame is discarded when the queue length exceeds this threshold; : the maximum number of frames that the queue can accommodate, i.e., when the number of frames exceeds the threshold , the oldest frame is discarded; : the threshold at which the oldest frame is discarded when the queue length exceeds this threshold; : the threshold at which the oldest frame is discarded when the queue length exceeds this threshold; The video stream captured by the onboard camera is encoded in real time, and the original frame needs to be scaled to the target resolution before video encoding, and then pushed to the cloud through the 5G network; The push streaming process includes at least one of pixel format conversion, H.264 / HEVC encoding, and code rate control.

[0019] The present application also provides a railway catenary unmanned aerial vehicle intelligent inspection system for implementing the above method, the system comprising: an unmanned aerial vehicle platform; including an inspection camera carried on the unmanned aerial vehicle and a controllable gimbal;​ An edge computing device, integrated on the drone, is used to run the multi-level target detection algorithm and gimbal control logic; The cloud server is used to receive, store, and analyze inspection data, and to run a high-precision disease identification model. A wireless communication module is used to establish a data transmission link between the edge computing device and the cloud server.

[0020] Beneficial effects This application discloses a method and system for intelligent inspection of railway catenary by unmanned aerial vehicle (UAV). The UAV payload combines a deep learning model to achieve real-time detection and defect identification of catenary equipment. The system completes precise rotation, intelligent zoom, focusing, and photography through gimbal closed-loop control. It also integrates edge inference acceleration and video encoding / network streaming technology to achieve efficient and stable inspection operations. 1. This invention achieves fully automated operation of UAVs during flight path missions by constructing a fully closed-loop inspection process based on visual detection and event-driven control. After the aircraft reaches the preset waypoint, the system can autonomously complete target detection, gimbal rotation, zooming, focusing, high-definition photography, and attitude reset without manual intervention. This method effectively reduces the intensity of manual operation and the risks of high-altitude operations, significantly shortens the inspection cycle, and improves the level of automation and safety of operations.

[0021] 2. Achieve high-precision target positioning and full-view coverage of components; This invention employs a combination of layered visual inspection and adaptive pan-tilt control to achieve high-precision identification and automatic alignment of contact wire poles and their auxiliary components. Through iterative pan-tilt fine-tuning and a closed-loop zoom focusing mechanism, it ensures stable imaging of the detected target at the image center and supports complete acquisition of both sides of the same component. This method effectively reduces the rates of missed and false detections, providing high-quality, comprehensive image data support for subsequent defect diagnosis and maintenance decisions.

[0022] 3. 5G real-time data transmission and cloud-based disease detection; This invention integrates a high-bandwidth, low-latency 5G communication module, enabling real-time uploading of drone location, gimbal attitude, detected target category, photos, and video streams to the cloud. The cloud system uses deep learning algorithms to perform defect identification and intelligent analysis on the uploaded data, and combines pose information to achieve anomaly localization and visualization. This design realizes a collaborative mechanism of "edge detection—cloud verification," balancing real-time detection with identification accuracy, significantly improving the intelligence level and data management capabilities of overhead contact line inspection. Attached Figure Description

[0023] Figure 1 A flowchart illustrating an intelligent unmanned aerial vehicle (UAV) inspection method for railway overhead contact lines provided by this invention; Figure 2 A schematic diagram of the overall structure of the railway catenary unmanned aerial vehicle intelligent inspection system provided by the present invention; Figure 3 This is a schematic diagram of the adaptive gimbal rotation and closed-loop zoom focusing and shooting process of the present invention. Detailed Implementation

[0024] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0025] like Figure 1 As shown in the figure, an embodiment of the present invention provides a method for intelligent inspection of railway catenary by unmanned aerial vehicle (UAV), which includes the following steps: S1. Data Acquisition: Acquire inspection images collected by the UAV during the inspection process along a preset route; the preset route is laid out according to the following rules: The line connecting the contact wire poles is used as the flight path; a detection point is set on both the front and rear sides of the column; the detection point on the front side of the column and the detection point on the rear side of the column are at the same height. The route taken by the drone during forward inspection is called the forward view route, which is formed by connecting the detection points located in front of each column. The route taken by the drone when it turns back is used as the reverse view route; the reverse view route is formed by connecting the detection points located behind each pillar.

[0026] S2. Multi-level target detection: At least three levels of target detection are performed on the inspection image. The first level of target detection is used to locate the main structure of the catenary. Based on the location results of the first level of target detection, small components are identified in the local region of interest in the second level of target detection. Based on the location results of the second level of target detection, local defects are identified in the third level of target detection. First-layer target detection: The first deep learning model is used to perform full-frame analysis of the image, identify and locate the main structure of the catenary, and output the bounding box information of the main structure of the catenary. The main structures of railway contact wire poles and crossarms are large and have clear textures, but the background is complex. This stage employs a lightweight YOLOv8 model for rapid detection of the entire frame image, achieving efficient recognition of large targets. The detection head directly predicts the center coordinates of the bounding box for each anchor point. Width and height Target confidence level and category probability vector ( (where the number of categories is ), the output result is: RGB images captured by the airborne camera are processed Normalization and scaling to adjust to the network input size. and channel permutation conversion (BGR→RGB) is performed: wherein, denotes the normalized image, denotes the channel index. I raw is the original pixel value, I norm is the pixel value after normalization, I resize is the pixel value after bilinear interpolation; is the image under the original BGR channel; , are the dimensions of the image input network, i.e., the height and width of the image, respectively.

[0027] The inference time can be approximately expressed as: wherein, is the YOLOv8 model calculation amount (FLOPs), is the number of GPU cores, is the single-core operation capability (GFLOPS), is the TensorRT optimization efficiency coefficient Second layer target detection: based on the output bounding box information, generate the region of interest according to the preset expansion coefficient, and use the second deep learning model optimized by multi-scale feature fusion in the region of interest to identify small components; The contact net pole region detected in the first stage is large, and directly performing small target detection may cause problems such as “mistaking the insulators on other pole bodies” or “detecting similar shapes of interference objects in the background”. Therefore, this stage generates a local region of interest (ROI) according to the output result of the first stage: wherein, is the expansion coefficient, is the horizontal expansion distance, is the vertical expansion distance; , are the horizontal and vertical coordinates of the first stage detection result, respectively. The purpose of ROI constraint is to limit the detection range to cover only the current pole body area, so as to avoid detecting the insulators on the adjacent pole bodies, and reduce the false judgment caused by the non-target objects such as wires and signal poles in the complex background.

[0028] In the ROI range, an improved small target detection model is adopted, and a multi-scale feature fusion module (BiFPN) is introduced to enhance the semantic perception ability of the small target feature layer. The structure adopts a bidirectional feature fusion method with learnable weights, which enhances the multi-scale information interaction ability while maintaining lightweight.

[0029] For the input features of the same scale , the fusion calculation form of BiFPN is: wherein, is a learnable weight, is a small constant to avoid zero denominator. This formula realizes weighted feature fusion, allowing the model to automatically adjust the importance of each layer feature according to training.

[0030] The fused output features are processed by depth separable convolution and batch normalization: Finally, the enhanced multi-scale features are output through a nonlinear activation function. This improved structure can effectively balance high-level semantics and low-level detail information, thereby improving the detection accuracy and stability of insulators in complex backgrounds.

[0031] Third layer target recognition: According to the identified small components, a third deep learning model is used for defect detection to determine whether at least one of cracks, damage, and missing exists, and when it exists, output the detection frame. Edge detection has real-time performance, but is limited by computing power and lacks the ability to recognize subtle disease characteristics such as cracks and notches. To further improve the accuracy and robustness of insulator disease recognition, the present invention deploys an improved MAP-YOLOv8 detection network in the cloud. The cloud uses the improved MAP-YOLOv8 network based on YOLOv8 to enhance feature expression, mainly including: GSConv module: combines group convolution and point-wise convolution to reduce computational complexity. Its basic calculation form is: SimSPPF module: based on spatial pyramid pooling, different scale pooling is used to enhance multi-scale information expression: ; wherein k1, k2, and k3 represent the number of pooling layers, respectively.

[0032] MAP-CA attention mechanism: the channel attention map is obtained by adding the global average pooling and maximum pooling: wherein, is a sigmoid function, are full connection layer parameters.

[0033] In the loss function design, the improved loss function is adopted, including the boundary box regression loss, the classification loss and the confidence loss. The SIoU loss is used in the boundary box regression part, considering the center point distance, shape and angle factors: wherein, IoU is the intersection over union of the predicted box and the real box, D, S and A are distance , shape and angle cost functions. The angle cost function evaluates the direction which needs to be adjusted most between the predicted box and the real box, and calculates the angle between the center line of the real box and the predicted box and the horizontal axis; The distance cost function calculates the distance between the predicted box and the real box after normalization and angle constraint, and when the angle cost is large, the weight of the distance cost becomes small, and the model will relatively pay more attention to the shape matching; the shape cost function measures the difference between the predicted box and the real box in the aspect ratio (shape), When the category of the detection box is "insulator" and the prediction probability exceeds the set threshold, the system further analyzes the target region. If there are abnormal features such as cracks and missing in the boundary box, it is determined as "damage disease", and the detection position, category and confidence are output.

[0034] S3, adaptive pan-tilt control: based on the result of the multi-level target detection, the offset of the target and the image center is calculated, the pan-tilt is driven to rotate and / or zoom to make the target centered in the image and reach the preset imaging size; In order to automatically place the detected target in the center of the picture, the accurate field of view angle and pixel focal length of the current lens are calculated first, then the center position of the target detection box output by the deep learning model is mapped into the rotation angle of the pan-tilt, and the instruction is issued. The position deviation of the target in the image is obtained through the deep learning model, the required pan-tilt rotation angle is calculated combined with the real-time zoom ratio and the camera calibration parameters, and multiple iterative corrections are supported to realize the target centered imaging.

[0035] The adaptive pan-tilt control specifically includes the following steps: S301, physical focal length and field of view angle calculation read the zoom ratio of the current pan-tilt , in the pre-defined zoom ratio-focal length table , do linear interpolation to get the physical focal length . According to the physical width and the physical height of the pan-tilt camera sensor, the horizontal field of view angle is calculated using the physical focal length . and vertical field of view : and combined with the horizontal resolution of the image and the vertical resolution of the image convert the physical focal length to pixel focal length: S302, calculate the pixel offset and gimbal rotation angle mapping take the target detection box center output by the model and the picture center , calculate the pixel offset: divide the pixel offset by the pixel focal length to get the corresponding yaw rotation amount and pitch rotation amount : wherein, is the horizontal focal length of the image; is the vertical focal length of the image.

[0036] S303, issue rotation instructions and compensate for residual errors In the gimbal control interface, take the current attitude as the reference, and issue the relative rotation amount , to the gimbal; after issuing the rotation instruction, wait for the gimbal mechanical movement to converge; read the absolute attitude of the gimbal after rotation and compare it with the original attitude to get the actual rotation amount , ; calculate the residual error: If , exceeds the preset threshold, take ( , ) as the new relative amount and issue the rotation command again until the residual error is suppressed within the threshold or the maximum compensation number is reached.

[0037] At the same time, after completing the initial rotation, read the updated frame and re-detect the deviation of the target center and the field of view center, if there is still a residual error exceeding the threshold: repeat the above angle calculation and rotation process until the deviation converges or the maximum iteration number is reached. Wherein, R represents the residual error, represents the horizontal deviation of the target center and the field of view center, represents the vertical deviation of the target center and the field of view center.

[0038] This closed-loop adjustment process ensures that the target center gradually coincides with the imaging center, thereby achieving high-precision automatic alignment.

[0039] S304, optical zoom and auto focus After target alignment, the system enters the automatic zoom and focus phase. First, read the camera model, select the corresponding minimum zoom factor and maximum zoom factor , and call the interface to get the current optical zoom factor .

[0040] According to the image resolution width , height and the detection frame width , height , calculate the target zoom factor; Where M is the marginal coefficient, used to reserve the edge space, this factor ensures that the picture after removing the edge space can be filled in both horizontal and vertical directions, and the smaller value is taken to prevent out-of-bound, thereby generating the target zoom factor That is, the target zoom factor issued to the gimbal.

[0041] Read the latest zoom every 100 ms , judge whether it is greater than the threshold value, if the condition is met, it is considered complete, a prompt "zoom complete" is given and the polling is exited, and the focus coordinates are uniformly set to the center of the picture.

[0042] S305, multi-target processing and state recovery When there are multiple detection targets in the picture, the present application adopts the order of spatial position to sequentially execute the "rotation-tweak-zoom-focus-photograph" operation until all targets are imaged. When starting to process a certain target, first calculate its center offset in the image, and map the offset to the pitch angle and yaw angle that need to be adjusted according to the current lens parameters. Subsequently, the gimbal performs a preliminary rotation to make the target basically enter the center area of the picture. After the initial rotation is completed, the system does not image immediately, but must wait for the updated image frame, detect the position of the current target in the new frame again, and recalculate the deviation between it and the center of the picture to determine whether further compensation is needed.

[0043] To improve the stability of small target detection in complex scenes, the system will automatically generate a local detection area around the current target after the initial rotation. The target box is centered on the local detection area, and a certain proportion is automatically expanded according to the actual size to exclude the interference of adjacent rods or background objects. When the compensation operation is performed, the system will automatically shrink the area range according to the detection effect, so that the detection is more focused. If the local area does not detect the target for several consecutive frames, the system will automatically fall back to whole-frame analysis and maintain this fallback state for several frames to avoid jitter caused by switching back and forth.

[0044] When the system detects multiple similar-shaped targets in a new frame, it uses the degree of overlap between the target and the previous frame detection box to determine whether it belongs to the same physical component. If the degree of overlap is low, the target closest to the center of the image is used as the current object to ensure that the compensation action is always performed on the same target. Through this inter-frame association mechanism, the system can maintain target consistency during continuous rotation compensation.

[0045] If the target in the new frame still does not meet the centering requirement, the system will continue to perform rotation compensation and enter the next round of "wait for new frame - re-detect - judge deviation - compensate" closed-loop process. The local detection area will also be updated synchronously during this period to further narrow the subsequent detection range. If the number of compensations reaches the upper limit, the system will force imaging and end this round of compensation, thereby avoiding long-term alignment stagnation in special scenarios.

[0046] When the target finally enters the acceptable deviation range, the system automatically performs target imaging, including zooming, focusing, and taking pictures. Considering that zooming will change the field of view of the picture, the system will clear the original target association information and local detection area after taking the picture to ensure that the positioning of the next target is not affected by the change in the viewing angle of the previous target.

[0047] For some scenes where there is occlusion or poor lighting that causes the target to be unrecognized, the system combines the structural rules of the catenary pole tower and uses a spatial relationship model pre-installed or learned online to infer the possible location of other components. In this case, the system will automatically generate a virtual detection area in the inferred area and perform directional compensation and imaging, so that the inspection process does not rely on the direct recognition result of a single frame and has a certain degree of active discovery capability.

[0048] S4, closed-loop imaging: after the gimbal completes alignment and zooming, automatic focusing and high-definition photography are performed to complete the inspection of the target; To achieve full-automatic inspection of unmanned aerial vehicle (UAV) in complex line environment, an intelligent control method based on route task and event-driven mechanism is proposed. According to the spatial distribution of railway overhead contact system (OCS) and the line center line, the system sets the route task manually, and sets two detection waypoints at each OCS pole position, which are located at a certain distance and height before and after the pole body respectively, to obtain the imaging data of front and back view, so as to ensure the all-around coverage of the target component. During the execution of the route task, the system uses the automatic control logic of event-driven mechanism to realize the cooperative switching of detection, rotation, imaging and flight process. When the UAV reaches the preset waypoint, the detection and gimbal control process is automatically triggered, and the following steps are executed: 1) Start the target detection module to identify and locate the OCS pole in the current field of view; 2) If the detection is successful, automatically drive the gimbal to rotate to align the pole body center with the image center; 3) Start small target detection after the target is centered to identify the detail components such as insulator and bolt on the pole body; 4) The detected small targets perform automatic zooming, focusing and high-definition photographing operations in turn; as shown in the figure, which specifically includes: applying for gimbal permission, setting gimbal rotation mode, calculating field of view angle, pixel focal length, calculating pixel offset, mapping gimbal angle offset, rotating angle compensation, zooming, focusing, photographing, and returning to initial position; Figure 3 5) After imaging is completed, the detection information is uploaded to the cloud, the gimbal is restored to the initial attitude, and the aircraft automatically turns to the next waypoint.

[0049] During flight, the system can realize automatic event response and process switching according to the task state: 1) When the task is interrupted (such as communication loss or flight control abnormality), the system immediately suspends detection and photographing, and keeps hovering state; 2) When the task is resumed, the system enters standby mode and waits for the next waypoint to trigger the detection task; 3) When the route is terminated or the inspection task is completed, the system automatically ends the detection process and returns to the initial position; 4) When the "photographing completed" event is triggered, the system records the detection state, image result and position information of the current pole body, and automatically switches to the next waypoint to continue the task.

[0050] Through the event-driven mechanism, the system can realize dynamic coordination of detection and flight tasks in complex high-speed rail line and electrified section environment, significantly improving the automation and safety of the inspection process. This method ensures the continuity of route task execution and the consistency of data acquisition, and can realize precise inspection and high-quality image acquisition without human intervention.

[0051] ​S5. Data Upload and Cloud Collaboration: Images, test results, and drone status data acquired during the inspection process are uploaded to the cloud server via wireless network, whereby the cloud server makes specific disease judgments and provides early warnings.

[0052] This system supports real-time encoding of video data captured by airborne cameras and direct push to the cloud, achieving low-latency, high-quality end-to-end video transmission. The streaming process includes steps such as frame queue management, image scaling and pixel format conversion, H.264 encoding, bitrate control, and buffer management.

[0053] (1) Frame queue management During video transmission, the system uses a frame queue to buffer the acquired image frames to adapt to fluctuations in network bandwidth and changes in encoding rate. The change in frame queue length over time can be expressed as: in, : No. The length of the video frame queue at a given frame time (number of frames). : No. The number of new frames enqueued at each frame time (frames / cycle). : No. The number of frames dequeued and started encoding at frame time (frames / cycle). The maximum number of frames a queue can hold; That is when Exceeding the threshold When discarding the earliest frame, The threshold for discarding the earliest frame when the queue length exceeds this value. (2) Scaling and pixel format conversion The original frames need to be scaled to the target resolution before video encoding: : Input the width and height (in pixels) of the frame.

[0054] : Encodes the width and height (in pixels) of the target frame.

[0055] : Scaling ratio, setting the output height The same applies to width.

[0056] Pixel format conversion uses the standard BGR→YUV420P formula: BGR→YUV420P formula: : Red, green, blue channel values (range 0–255) of a pixel in an image.

[0057] : Transformed luma (Y) and chroma (U, V) components for YUV420P format (3) H.264 macroblock transform and quantization Video encoding divides each frame into 16×16 macroblocks, which are further divided into 4×4 subblocks, representing the luma value of a pixel within a macroblock.

[0058] For each subblock, a discrete cosine transform is applied: : Discrete cosine (DCT) coefficients of a subblock in the frequency domain.

[0059] : DCT frequency coordinate index (values 0–3).

[0060] : Normalized coefficient : Quantization step size at position in the predefined quantization matrix.

[0061] : Quantized integer coefficient.

[0062] After arranging the quantized coefficients in scan order, adaptive binary arithmetic coding (CABAC) or variable length coding (CAVLC) is used to generate the final bitstream.

[0063] (4) Rate and buffer management The average bit budget per frame is: B = R / f where R is the target output bitrate (bits per second, bps), f is the video frame rate (frames per second, fps), B is the bit budget allocated to each frame (bits per frame).

[0064] The buffer occupancy is maintained at , satisfying where, ​​​Buffer bit margin at time t. Actual output bits of frame t.

[0065] Smooth the code stream output by adjusting the quantization step.

[0066] As Figure 2 The present application also provides a railway overhead line system unmanned aerial vehicle intelligent inspection system for implementing the above method, the system comprises: An unmanned aerial vehicle platform; An inspection camera and a controllable gimbal mounted on the unmanned aerial vehicle; An edge computing device integrated on the unmanned aerial vehicle for running the multi-level target detection algorithm and gimbal control logic; A cloud server for receiving, storing, analyzing inspection data, and running a high-precision disease identification model; A wireless communication module for establishing a data transmission link between the edge computing device and the cloud server.

[0067] Obviously, the above embodiments are only examples for clearly illustrating, not limiting the embodiments. For those skilled in the art, on the basis of the above description, other different forms of changes or variations can also be made. Here, it is not necessary and impossible to enumerate all the embodiments. The obvious changes or variations derived therefrom are still within the protection scope of the present application.

Claims

1. A railway catenary unmanned aerial vehicle intelligent inspection method, characterized in that, The method comprises the following steps: Data acquisition: acquiring inspection images collected by the UAV during the inspection process according to the preset flight path; Multi-level target detection: performing at least three levels of target detection on the inspection images, the first level of target detection is used to locate the main structure of the overhead contact system, and based on the positioning result of the first level of target detection, small components are identified in the local region of interest in the second level of target detection; based on the positioning result of the second level of target detection, local disease identification is performed in the third level of target detection; Adaptive pan-tilt control: based on the results of the multi-level target detection, the offset of the target from the center of the image is calculated, and the pan-tilt is driven to rotate and / or zoom to center the target in the image and achieve a preset imaging size; Closed-loop imaging: after the pan-tilt completes the alignment and zooming, autofocus and high-definition shooting are performed to complete the inspection of the target; Data upload and cloud collaboration: the images, detection results and UAV state data obtained during the inspection process are uploaded to the cloud server through a wireless network, and the cloud server performs specific disease judgment and protection warning. 2.The railway overhead line system unmanned aerial vehicle intelligent inspection method according to claim 1, characterized in that, The preset flight path is arranged according to the following rules: The line connecting the overhead contact system poles is used as the flight path; one detection point is arranged on each side of the pole; the height of the detection point in front of the pole is consistent with the height of the detection point behind the pole; The route of the UAV when performing forward inspection is used as the forward view angle flight path, and the forward view angle flight path is connected by the detection points in front of each pole; The route of the UAV when returning is used as the reverse view angle flight path; the reverse view angle flight path is connected by the detection points behind each pole.

3. The intelligent inspection method of the railway overhead contact system UAV according to claim 2, characterized in that: When the UAV reaches the detection point, adaptive pan-tilt control is automatically triggered to identify, align, zoom, focus and shoot the target.

4. The intelligent inspection method of the railway overhead contact system UAV according to claim 2, characterized in that: First level of target detection: a first deep learning model is used to analyze the image frame by frame, identify and locate the main structure of the overhead contact system, and output the bounding box information of the main structure of the overhead contact system; Second level of target detection: based on the output bounding box information, a region of interest is generated according to a preset expansion coefficient, and a second deep learning model optimized by multi-scale feature fusion is used to identify small components in the region of interest; Third level of target detection: based on the identified small components, a third deep learning model is used to detect defects and determine whether at least one of the cracks, damages and missing parts exists, and when at least one of the cracks, damages and missing parts exists, a detection box is output.

5. The railway catenary unmanned aerial vehicle intelligent inspection method according to claim 2, characterized in that, The pixel offset of the target from the center of the image is calculated based on the results of the multi-level target detection, which comprises: center of the detection frame output according to the third layer target detection and the picture center ; The pixel offset is calculated as ; is the image horizontal resolution and is the vertical resolution.

6. The railway catenary unmanned aerial vehicle intelligent inspection method according to claim 5, characterized in that, When the pan-tilt is driven to rotate according to the following formula to calculate the yaw angle rotation amount and the pitch angle rotation amount, yaw angle rotation amount ; pitch angle rotation amount ; is the horizontal focal length of the image; is the vertical focal length of the image; Read the absolute attitude of the rotated holder and compare it with the original attitude to get the actual yaw angle rotation and the actual pitch angle rotation ; calculate the residual: If , all exceed the preset threshold, the rotation command is issued again as a new relative quantity (step 3) until the residual is suppressed within the preset threshold or the maximum compensation number is reached. , ​ After the initial rotation is completed, the update frame is read and the target center is re-detected for deviation from the center of the field of view. If there is still a residual deviation that exceeds the threshold : Then the angle calculation and rotation process described above are repeated until the deviation converges or a maximum number of iterations is reached; wherein R represents a residual, represents a lateral deviation of the target center from the field of view center, represents a longitudinal deviation of the target center from the field of view center.

7. The railway catenary unmanned aerial vehicle intelligent inspection method according to claim 1, characterized in that, When the pan-tilt is driven to zoom, the target zoom factor is determined by the following method: Based on the ratio of the pixel size of the target in the image to the expected imaging size, and introducing a margin coefficient, an initial magnification is calculated; Combined with the minimum and maximum zoom magnification of the pan-tilt camera, the initial magnification is constrained to obtain the final target zoom factor.

8. The unmanned aerial vehicle intelligent inspection method for railway overhead line system according to claim 2, characterized in that, In the adaptive pan-tilt control step, further comprising: when multiple targets to be inspected are simultaneously identified in a single frame of image, generating an imaging task sequence according to the horizontal spatial relationship of the targets to be inspected in the image; the imaging task sequence is sorted according to the horizontal coordinates of the center points of all detection boxes, and a continuous left-to-right alignment order is adopted to make the rotation path of the pan-tilt shortest; the generated target sequence can be expressed as: T = {T1, T2, T3, T4, T5, T6, T7, T8, T9, T10, T11, T12, T13, T14, T15, T16, T17, T18, T19, T20, T21, T22, T23, T24, T25, T26, T27, T28, T29, T30, T31, T32, T33, T34, T35, T36, T37, T38, T39, T40, T41, T42, T43, T44, T45, T46, T47, T48, T49, T50, T51, T52, T53, T54, T55, T56, T57, T58, T59, T60, T61, T62, T63, T64, T65, T66, T67, T68, T69, T70, T71, T72, T73, T74, T75, T76, T77, T78, T79, T80, T81, T82, T83, T84, T85, T86, T87, T88, T89, T90, T91, T92, T93, T94, T95, T96, T97, T98, T99, T100, T101, T102, T103, T104, T105, T106, T107, T108, T109, T110, T111, T112, T113, T114, T115, T116, T117, T118, T119, T120, T121, T122, T123, T124, T125, T126, T127, T128, T129, T130, T131, T132, T133, T134, T135, T136, T137, T138, T139, T140, T141, T142, T143, T144, T145, T146, T147, T148, T149, T150, T151, T152, T153, T154, T155, T156, T157, T158, T159, T160, T161, T162, T163, T164, T165, T166, T167, T168, T169, T170, T171, T172, T173, T174, T175, T176, T177, T178, T179, T180, T181, T182, T183, T184, T185, T186, T187, T188, T189, T190, T191, T192, T193, T194, T195, T196, T197, T198, T199, T200, T201, T202, T203, T204, T205, T206, T207, T208, T209, T210, T211, T212, T213, T214, T215, T216, T217, T218, T219, T220, T221, T222, T223, T224, T225, T226, T227, T228, T229, T230, T231, T232, T233, T234, T235, T236, T237, T238, T239, T240, T241, T242, T243, T244, T245, T246, T247, T248, T249, T250, T251, T252, T253, T254, T255, T256, T257, T258, T259, T260, T261, T262, T263, T264, T265, T266, T267, T268, T269, T270, T271, T272, T273, T274 wherein, is the horizontal coordinate of the center of the th target detection frame ​ 9.The railway overhead line system unmanned aerial vehicle intelligent inspection method of claim 2, wherein, ​ ​ wherein, : the frame queue length at the frame time, : the number of new frames enqueued at the frame time, : the number of frames dequeued and started encoding at the frame time, : the upper limit of the maximum number of frames the queue can hold, when the threshold is exceeded, the oldest frame is discarded; the threshold for discarding the oldest frame when this queue length is exceeded; ​ ​ 10. A railway catenary unmanned aerial vehicle intelligent inspection system for implementing the method according to any one of claims 1-9, characterized in that, ​ ​ ​ ​ ​

Citation Information

Patent Citations

  • On-line character detection method based on machine vision and system thereof

    CN101576956A

  • Method and system for storing video data

    CN105430480A

  • Image intelligent acquisition system and method for power transmission line UAV patrol inspection

    CN107729808A

  • Lightweight multi-unmanned aerial vehicle power grid inspection fault identification method and system

    CN118691988A