Unmanned aerial vehicle fine inspection method and system for beam-pumping unit
By constructing a collaborative mechanism between a multi-point refined inspection model and a multimodal visual large model, the problem of lack of refined decomposition and differentiated identification in the UAV inspection method for beam pumping units was solved, achieving efficient and accurate defect identification and reducing the missed detection rate and false judgment rate.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-25
- Publication Date
- 2026-03-31
AI Technical Summary
Existing drone inspection methods for beam pumping units lack detailed decomposition of the overall structure, making it difficult to identify defects that are different in each component, especially in terms of the ability to identify small-sized, low-contrast, or rare defects.
A multi-point refined inspection model is constructed, combining a dedicated small-model defect detector and a multimodal visual large-scale model, and a three-level collaborative mechanism is used for defect identification. The multi-point refined inspection model automatically segments the pumping unit image, the dedicated small model performs preliminary analysis on specific components, and the multimodal visual large-scale model performs semantic-level re-examination, combining prior knowledge of component structure to improve recognition accuracy.
It significantly improves the automation level and discrimination accuracy of defect identification in beam pumping units, reduces the rate of missed detection and false judgment, and achieves highly stable, refined, and intelligent inspection.
Smart Images

Figure CN121767448A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent operation and maintenance technology for oil extraction equipment, specifically to a method and system for refined inspection of beam pumping units using unmanned aerial vehicles (UAVs). Background Technology
[0002] As the most widely used oil production equipment in onshore oilfields, the operating status of beam pumping units directly affects the safe production and recovery efficiency of oil wells. This equipment has a complex structure and numerous components, and is exposed to harsh environments such as wind, sand, rain, snow, extreme cold, or high temperatures for extended periods, making it prone to various defects such as loose bolts, structural cracks, seal failures, and transmission wear. Traditional inspections mainly rely on regular manual on-site checks, which suffers from high labor intensity, strong subjectivity, high missed detection rates, and difficulty in quantification. Furthermore, some high-level or dangerous areas (such as the pump head and connecting rods) are difficult to observe closely, hindering early detection and preventative maintenance.
[0003] In recent years, drone inspection technology has been gradually applied to oilfield equipment operation and maintenance due to its advantages such as flexibility, efficiency, and non-contact operation. Equipped with visible light cameras, drones can automatically fly around pumping units and collect images from multiple angles, significantly improving data acquisition efficiency. However, existing methods often use a single, general target detection model to perform end-to-end defect identification on the entire unit image, failing to fully consider the significant differences in morphology, function, and defect patterns among the various components of a beam pumping unit. This "one-size-fits-all" analysis often struggles to ensure the detection accuracy of all critical parts when faced with occlusion, scale variations, or complex backgrounds, especially limiting its ability to identify small-sized, low-contrast, or rare defects.
[0004] To improve recognition performance, some studies have attempted to introduce deep learning models and optimize performance in specific scenarios through data augmentation or model fine-tuning. Patent CN113837004B discloses a deep learning-based kinematic analysis method for beam pumping units. By applying this method in oilfields, it combines Yolov4 models and mathematical models to address the difficulty of monitoring the kinematic parameters of pumping units. Existing systems generally rely on training with a large number of labeled samples, have weak generalization ability for sparse defect types, and cannot effectively utilize domain prior knowledge for semantic-level reasoning and correction, making it difficult to meet the urgent needs of oilfields for highly reliable and refined intelligent inspection. Summary of the Invention
[0005] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method and system for refined inspection of beam pumping units by unmanned aerial vehicles (UAVs). This solves the problem that existing UAV inspection methods for beam pumping units lack a detailed breakdown of the overall structure and are unable to meet the different defect identification needs of various components.
[0006] The objective of this invention is achieved through the following technical solution: A method for refined inspection of beam pumping units using unmanned aerial vehicles (UAVs) includes the UAV automatically flying along a preset route to acquire images of the pumping unit, and then performing defect identification on the acquired images to output inspection results. The defect identification and inspection results output from the acquired oil pumping unit images include: Build and deploy a deep learning-based multi-point refined inspection model; through learning, this model can automatically identify, locate and segment component sub-images corresponding to eighteen predefined standard inspection points from overall images of the pumping unit taken from different angles and distances. Multiple component sub-images are input into a pre-trained, point-to-point-corresponding dedicated small-model defect detector for parallel analysis to obtain preliminary defect identification results; the output results include defect category and bounding box coordinates. Defects identified in the initial defect identification results with a confidence level below a set threshold or belonging to a sparse data category, along with their corresponding component sub-images, are sent to a multimodal visual large model for secondary semantic-level review. The multimodal visual large model combines the image content with the injected prior knowledge of the component structure to generate the final defect judgment and description.
[0007] Through a three-tiered collaborative mechanism of 18-point segmentation, initial screening using a dedicated small model, and re-examination using a multimodal large model, the automation level and discrimination accuracy of defect identification in beam pumping units have been significantly improved. On the one hand, a multi-point refined inspection model is used to achieve structured decomposition of the whole machine image into key components, avoiding manual inspection of each part. On the other hand, the dedicated small model is optimized for the defect characteristics of each component, ensuring efficient and accurate screening of known defects. At the same time, for suspected defects with low confidence or sparse data, a multimodal large model that integrates prior knowledge of the components is introduced for semantic-level verification, effectively improving the ability to identify fuzzy, rare, or complex defects, reducing the overall missed detection rate and false judgment rate, and achieving highly stable refined intelligent inspection.
[0008] As a preferred approach, the preliminary results of the dedicated small model defect detector are combined with the re-inspection results of the multimodal visual large model to generate a structured inspection report containing defect type, precise location, corresponding point number, confidence level, and semantic description. The drones are automatically scheduled to take off and be recovered by drone nests deployed at the oil field site, and perform fully automated inspections according to a precise repeat shooting route around the pumping unit; The dedicated small-model defect detector is a lightweight object detection network that supports real-time inference at the edge; the multimodal visual large model is a visual-language model with open vocabulary understanding and zero-shot inference capabilities.
[0009] As a preferred approach, the training of the dedicated small model defect detector adopts a hybrid data strategy, including real historical defect samples and simulated defect images synthesized through physical simulation; for minor defects such as loose bolts and peeling paint, edge enhancement maps extracted by traditional image processing are fused into the input, and SIoU Loss is used as the bounding box regression loss function to improve the localization accuracy.
[0010] As a preferred approach, the multimodal visual large model employs a dual-path inference mechanism during the secondary review: For known defect types, structured prior knowledge is injected through prompts to guide the model to focus on key areas; For unknown or rare defects, its zero-sample transfer capability is activated, and open vocabulary recognition and semantic description generation are performed based on the universal vision-language alignment space.
[0011] As a preferred method, the images acquired by the UAV include high-resolution visible light images and infrared thermal images. A multi-altitude layer circling flight strategy is adopted to ensure full coverage of the top, middle, bottom and moving connection parts of the pumping unit. The overlap rate between adjacent images is not less than 30% to ensure the continuity of subsequent image stitching and defect location.
[0012] As a preferred approach, to control the positioning deviation of defects in the image within a preset threshold, the UAV flight altitude H, camera parameters p,f, and adjacent image overlap rate R are set to satisfy the following relationship: in, This indicates the maximum permissible location deviation of the defect on the image plane (unit: mm); This indicates the drone's flight altitude relative to the pumping unit components (unit: mm). This indicates the physical size of a single pixel on the camera sensor (unit: mm / pixel). This indicates the focal length of the camera lens (unit: mm). This represents the typical pixel-level localization error of the defect identification model on the image plane. This represents the overlap rate (in %) between adjacent aerial images, and .
[0013] As a preferred approach, a continuous defect tracking and early warning mechanism is set up: when similar defects on the same component are identified in N consecutive (N≥2) inspections and their positioning deviation is less than a preset position threshold, they are automatically marked as "continuously deteriorating defects", the early warning level is increased and the warning is pushed out with priority.
[0014] As a preferred method, the eighteen standard inspection points are defined based on the core load-bearing structure, kinematic pairs, power transmission chain, and safety devices of the walking beam pumping unit, including: first walking beam, second walking beam, walking beam, connecting rod, tower base, suspension rope device, polished rod, insulation box, balancer, reducer, belt, motor, rope, base, fence, crossbeam, distribution box, and crank pin.
[0015] As a preferred approach, a drone-based precision inspection system for beam pumping units includes: a drone-based automatic inspection subsystem, an edge computing analysis subsystem, and a cloud-based intelligent verification and management subsystem; wherein: The unmanned aerial vehicle (UAV) automatic inspection subsystem controls the UAV nest to complete autonomous take-off and landing, orbiting flight and multimodal image acquisition; The edge computing analysis subsystem is deployed in the oilfield and runs a multi-point refined inspection model and a group of dedicated small model defect detectors to achieve real-time component disassembly and initial defect screening. The cloud-based intelligent review and management subsystem integrates a multimodal visual large model, a digital twin platform, a model training engine, and an inspection database. It is responsible for reviewing low-confidence tasks, generating reports, and performing 3D visualization playback.
[0016] As a preferred approach, the edge computing analysis subsystem adopts a containerized deployment of a dedicated small model defect detector group, equipped with a dynamic task scheduling and load balancing module, supports concurrent inference of multiple component sub-images, and ensures that the end-to-end processing latency is ≤ 2 seconds per pumping unit; The cloud-based intelligent review and management subsystem's multimodal large model module integrates a domain knowledge base and a prompt word template library, which can automatically select the optimal reasoning strategy based on the defect type, thereby improving the accuracy and interpretability of the review.
[0017] The present invention has at least the following beneficial effects: This invention significantly improves the automation level and discrimination accuracy of defect identification in beam pumping units by constructing a three-level collaborative analysis process: a multi-point refined inspection model, a dedicated small-model defect detector, and a multimodal visual large model. First, using a deep learning-based multi-point refined inspection model, it can automatically identify and accurately segment component sub-images corresponding to multiple predefined standard inspection points from overall images of the pumping unit taken from different angles and distances. This avoids the inefficient operation of manually searching each part individually, achieving a structured decomposition from the overall image to key components. Subsequently, each component sub-image is fed into a dedicated small-model defect detector corresponding to its specific point for parallel analysis. Each small model is optimized and trained for typical defect types of a particular component, efficiently outputting high-confidence defect categories and bounding box coordinates, ensuring rapid and accurate screening of known defects.
[0018] For defects identified in the initial recognition results with low confidence or belonging to sparse data categories, the system further introduces a multimodal visual large model for secondary semantic-level re-examination. This large model not only analyzes the image itself but also incorporates injected prior knowledge of component structure to deeply understand and correct the semantics of the defects, effectively improving the ability to identify ambiguous, rare, or complex defects. This two-level recognition mechanism maintains the efficient reasoning advantages of dedicated small models while leveraging the generalization and knowledge reasoning capabilities of the large model to compensate for its strong data dependence, thereby achieving a more refined defect recognition effect with lower false negative and false positive rates overall. Attached Figure Description
[0019] To reveal the technical details of the embodiments of the present invention, the accompanying drawings involved in the embodiments will be briefly described below. It should be emphasized that these drawings only present several embodiments of the present invention and should not be considered as defining the scope of the invention. For those skilled in the art, other related drawings can still be derived based on these drawings without inventive effort.
[0020] Figure 1 This is a flowchart illustrating a method for refined inspection of a beam pumping unit using a drone, as described in this embodiment. Figure 2 A schematic diagram illustrating the boundary box identification for refined inspection by drones; In the diagram, 1-First donkey head area, 2-Second donkey head area, 3-Walking beam area, 4-Connecting rod area, 5-Tower base area, 6-Suspension rope area, 7-Smooth pole area, 8-Insulation box area, 9-Balancer area, 10-Reducer area, 11-Belt area, 12-Motor area, 13-Rope area, 14-Base area, 15-Fence area, 16-Crossbeam area, 17-Distribution box area, 18-Crank pin area. Detailed Implementation
[0021] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings, but the scope of protection of the present invention is not limited to the following description.
[0022] In the following description, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. However, it should be understood that the present disclosure is not limited to the specific forms shown herein. Rather, it should be understood to encompass various variations, equivalents, and / or alternatives to the embodiments of the present disclosure. In illustrating the drawings, the same reference numerals will be used to denote similar components.
[0023] In the various embodiments of this disclosure, the terms "first," "second," "the first," or "the second" are intended to modify different components and not to indicate order and / or importance, nor do they constitute a limitation on the respective components. For example, a first user equipment and a second user equipment represent different user equipments, although they both fall under the category of user equipment. Similarly, a first component may be named a second component, and a second component may be named a first component, without changing their essential attributes within the scope of this disclosure.
[0024] In this disclosure, terminology is used to describe specific embodiments and does not constitute a limitation thereof. In this context, the use of the singular form also encompasses the plural form, unless otherwise expressly stated herein. In the course of description, terms such as “comprising” or “having” are intended to indicate the presence of features, quantities, steps, operations, structural components, parts, or combinations thereof, and do not preclude the possibility or addition of one or more other features, quantities, steps, operations, structural components, parts, or combinations thereof.
[0025] It should be clarified that while the following description provides detailed specific information to aid in a comprehensive understanding of the exemplary embodiments, those skilled in the art will recognize that the exemplary embodiments can be implemented even without these specific details. For example, the system may be illustrated using block diagrams to avoid excessive detail that could obscure the clarity of the example. In other cases, to maintain the clarity of the example, unnecessary details of well-known processes, structures, and techniques may be omitted.
[0026] like Figure 1 As shown, a method for refined inspection of beam pumping units using unmanned aerial vehicles (UAVs) includes the UAV automatically flying along a preset route to acquire images of the pumping unit (or flying according to the instructions of the UAV operator), and identifying defects in the acquired images to output inspection results. The defect identification and inspection results output from the acquired oil pumping unit images include: Build and deploy a deep learning-based multi-point refined inspection model (18-point refined inspection model); through learning, this model can automatically identify, locate and segment component sub-images corresponding to 18 predefined standard inspection points from the overall image of the pumping unit taken from different angles and distances; Multiple component sub-images (eighteen component sub-images) are input into pre-trained dedicated small model defect detectors that correspond one-to-one with the points for parallel analysis to obtain preliminary defect identification results; among them, each dedicated small model is optimized and trained for a specific defect type of its corresponding component, and the output results include defect category and bounding box coordinates; Defects identified in the initial defect identification results with a confidence level below a set threshold or belonging to a sparse data category, along with their corresponding component sub-images, are sent to a multimodal visual large model for secondary semantic-level review. The multimodal visual large model combines the image content with the injected prior knowledge of the component structure to generate the final defect judgment and description.
[0027] Through a three-level collaborative mechanism, this invention significantly improves the automation and accuracy of defect identification in beam pumping units. First, an 18-point refined inspection model based on deep learning automatically decomposes the overall unit image into predefined key component sub-images, achieving structured and comprehensive coverage of parts, eliminating the tedious manual location of each component. Second, each component sub-image is analyzed in parallel by a dedicated small-model defect detector corresponding to it, optimizing for typical defects in specific components and efficiently outputting high-confidence defect category and location information. Finally, for preliminary results with insufficient confidence or sparse data, the system calls a multimodal visual large model that integrates prior knowledge of component structure for semantic-level secondary review, effectively enhancing the ability to identify ambiguous, rare, or complex defects. Overall, while ensuring efficiency, it significantly reduces the missed detection rate and false positive rate, achieving stable, reliable, and refined intelligent inspection.
[0028] In a preferred embodiment, the preliminary results of the dedicated small model defect detector are fused with the re-inspection results of the multimodal visual large model to generate a structured inspection report containing defect type, precise location, location number, confidence level, and semantic description. The drones are automatically scheduled to take off and be recovered by drone nests deployed at the oil field site, and perform fully automated inspections according to a precise repeat shooting route around the pumping unit; The dedicated small-model defect detector is a lightweight object detection network that supports real-time inference at the edge; the multimodal visual large model is a visual-language model with open vocabulary understanding and zero-shot inference capabilities.
[0029] The workflow begins with a drone nest deployed at the oilfield site. This nest automatically schedules drones to take off along a pre-set circling route, acquiring multi-angle, high-overlap images of the target pumping unit. After completing the mission, the drones automatically return and are retrieved, achieving fully automated inspection without human intervention. The acquired images of the entire unit are first input into a deep learning model called the "18-Point Refined Inspection Model." This model automatically identifies and crops 18 predefined key component regions, generating corresponding sub-images. Subsequently, these sub-images are fed into dedicated small-model defect detectors that are matched one-to-one with each model. Each small model is a lightweight object detection network specifically trained for typical defects in a particular component (such as a gearbox or connecting rod), enabling rapid preliminary analysis on edge devices and outputting defect type, bounding box location, and confidence level. For preliminary results with low confidence or belonging to rare categories, the system does not discard them directly. Instead, it forwards them, along with the original sub-image, to a multimodal vision model in the cloud for secondary review. This model can not only "read the image" but also understand the injected prior knowledge of the component structure (e.g., "a cracked donkey head is a serious defect"), generating a more accurate and interpretable final judgment through semantic reasoning. Finally, the system merges the results of the two identifications and automatically generates a structured inspection report, which includes the defect type, precise location, corresponding point number, confidence level, and natural language description, providing maintenance personnel with clear, reliable, and traceable decision-making support.
[0030] In a preferred embodiment, the 18-point refined inspection model adopts an improved YOLOv8m architecture, embedding a deformable convolutional module (DCNv3) in the backbone network, using an enhanced weighted bidirectional feature pyramid network (EBiFPN) in the feature fusion layer, and introducing a depth-separable visual Transformer module (SepViT Block) in the prediction head to improve the reliability of component localization under occlusion, scale changes and complex backgrounds.
[0031] The refined inspection model for the 18 points is based on the highly efficient object detection architecture YOLOv8m and deeply optimized to more reliably locate the 18 key components of the pumping unit from complex field images. To address challenges commonly encountered in actual inspections, such as component occlusion, large size differences (e.g., the base in the foreground differing from the donkey's head in the distance), and cluttered backgrounds (e.g., intersecting pipelines and vegetation), the model has undergone targeted improvements in three key areas: First, a deformable convolutional module (DCNv3) is introduced into the backbone network, enabling it to adaptively adjust the receptive field shape and flexibly capture features of irregular or partially occluded components, rather than being limited to a fixed grid. Second, an enhanced weighted bidirectional feature pyramid network (EBiFPN) is used in the feature fusion stage, which more effectively integrates feature information from different scales through an intelligent weighting mechanism, thus simultaneously considering the overall structure of large components and the detailed texture of small components. Finally, a deep separable visual Transformer module (SepViT Block) is embedded in the prediction head, utilizing its global attention mechanism to understand the semantic relationships between different regions in the image, helping the model to more accurately focus on real components rather than distracting objects in complex backgrounds. These three improvements work synergistically to significantly enhance the model's stability in identifying and locating key components in real oilfield environments.
[0032] In a preferred embodiment, the training of the dedicated small model defect detector adopts a hybrid data strategy, including real historical defect samples and simulated defect images synthesized through physical simulation; for minor defects such as loose bolts and peeling paint, edge enhancement maps extracted by traditional image processing are fused into the input, and SIoU Loss is used as the bounding box regression loss function to improve the localization accuracy.
[0033] The training of the dedicated small-model defect detector employs a hybrid data strategy of "real data + simulation synthesis" to address the problem of scarce and difficult-to-collect samples of certain defects in oilfields (such as loose bolts and peeling paint). Specifically, in addition to using real defect images accumulated from historical inspections, the system also uses physically driven image simulation technology to artificially generate simulated defects that conform to actual shapes and lighting conditions on normal component images, thereby significantly expanding the diversity and coverage of training data. For subtle defects such as loose bolts and peeling paint, the original images alone are often insufficient to capture their weak features. Therefore, an edge enhancement map generated by traditional image processing algorithms (such as edge detection) is additionally integrated during the model input stage to highlight the component outline and abnormal textures, helping the model to more sensitively identify subtle changes. Simultaneously, during model training, SIoU Loss is used as the optimization objective for bounding box regression. This loss function not only considers positional deviations but also introduces constraints on the shape and angle of the bounding box, enabling the predicted box to more closely and accurately fit the actual contour of the defect, thereby significantly improving positioning accuracy and providing a reliable basis for subsequent accurate identification and repair.
[0034] In a preferred embodiment, the multimodal visual large model employs a dual-path inference mechanism during secondary review: For known defect types, structured prior knowledge is injected through prompts to guide the model to focus on key areas; For unknown or rare defects, its zero-sample transfer capability is activated, and open vocabulary recognition and semantic description generation are performed based on the universal vision-language alignment space.
[0035] The multimodal vision big data model employs a dual-path reasoning mechanism during secondary inspections to ensure high-precision identification of various defect types. For known defect types, the system injects structured prior knowledge into the model through carefully designed prompts, such as explicitly stating "This component is a gearbox; common defects include oil leaks and loose screws," guiding the model to focus on these key areas for detailed analysis. This prompt-based approach allows the model to fully utilize domain expert knowledge, improving the efficiency and accuracy of detecting common defects. For unknown or rare defects, the model activates its zero-shot transfer capability, leveraging its visual-language alignment space learned on large-scale general datasets to automatically understand and describe newly emerging anomalies. In this mode, the model not only analyzes the image itself but also generates corresponding semantic descriptions, such as "A suspected crack was found, resembling a thin, elongated line, located near the edge of the component," helping maintenance personnel quickly understand and take appropriate measures. Through these two complementary reasoning paths, the multimodal vision big data model can efficiently handle common defects while flexibly responding to complex, ambiguous, or unprecedented defect types, significantly improving the overall reliability of the inspection system.
[0036] In a preferred embodiment, the images acquired by the UAV include high-resolution visible light images and infrared thermal images. A multi-altitude layer circling flight strategy is adopted to ensure full coverage of the top, middle, bottom, and moving connection parts of the pumping unit. The overlap rate between adjacent images is not less than 30% to ensure the continuity of subsequent image stitching and defect location.
[0037] When performing inspection missions, drones simultaneously acquire high-resolution visible light images and infrared thermal images. The former is used to identify external defects (such as cracks, peeling, and corrosion), while the latter can capture abnormal heating of equipment (such as motor overload and bearing friction), achieving complementarity between visual and thermal information. To ensure that all critical parts of the pumping unit—including the top donkey head, the middle connecting rod crank, the bottom base, and all moving connections—are clearly and completely covered, the drone adopts a multi-altitude layered circling flight strategy, that is, flying around at different altitudes in layers and taking pictures layer by layer from multiple perspectives. At the same time, the system strictly controls the overlap rate between adjacent aerial images to be no less than 30%. This ensures sufficient matching features during image stitching or 3D reconstruction, avoiding the omission of details, and also ensures that the same defect is repeatedly captured in multiple images, improving the continuity and stability of subsequent AI model localization, and providing a high-quality, blind-spot-free image foundation for refined defect identification.
[0038] To further improve the reliability of defect identification, the system introduces a multimodal confidence fusion mechanism, jointly identifying texture anomalies in visible light images and thermal anomalies in infrared images. For any component sub-image region... Its final defect confidence level Calculated by the following formula: in, The final defect confidence level after fusion (value range [0,1]); The confidence level of the original defects is based on the output of the visible light image for a dedicated small model; The thermal gradient magnitude (unit: °C / mm) of this region in the infrared image is calculated using the Sobel operator; This is the thermistor gain coefficient (unit: mm / °C), used to adjust the contribution of thermal anomalies to the confidence level; a typical value is 0.5 mm / °C. Visible light weighting coefficient Dynamically adjust according to component type (e.g., motor components) Structural components ); The base of the natural logarithm .
[0039] This embodiment utilizes a thermal gradient. The degree of localized temperature abrupt changes, such as bearing overheating or motor winding short circuits, is often manifested as a significant thermal gradient. When As the exponent increases, the exponent term approaches 1, and the contribution of thermal anomalies to the confidence level increases. Using this formula, the system can automatically suppress false defects (such as shadows and stains) that are only suspected in visible light but have no thermal response, while simultaneously strengthening the identification of true defects that combine visual cracks and temperature rise, significantly improving false alarm filtering capabilities.
[0040] To control the location deviation of defects in the image within a preset threshold, the drone's flight altitude is set. Camera parameters , and the overlap rate of adjacent images The following relationship must be satisfied: in, This indicates the maximum permissible location deviation of the defect on the image plane (unit: mm); This indicates the drone's flight altitude relative to the pumping unit components (unit: mm). This indicates the physical size of a single pixel on the camera sensor (unit: mm / pixel). This indicates the focal length of the camera lens (unit: mm). The typical pixel-level localization error of the defect identification model on the image plane can be calibrated based on historical verification data; This represents the overlap rate (in %) between adjacent aerial images, and By controlling the flight altitude and overlap rate, the defect location error can be ensured under given camera parameters, thus meeting the requirements of refined inspection.
[0041] To ensure sufficiently accurate defect localization in images, the system constrains the drone's flight parameters and imaging conditions. Specifically, defect localization deviation is mainly affected by three factors: the higher the drone flies, the larger the actual size represented by each pixel in the image, resulting in coarser localization; the camera's performance (such as sensor pixel size and lens focal length) determines the image's fineness; furthermore, the degree of overlap between adjacent images is crucial, as insufficient overlap can lead to stitching gaps or feature matching failures, thus amplifying the localization error. Therefore, this solution, by reasonably controlling flight altitude and image overlap rate, strictly limits defect localization deviation to a preset millimeter-level threshold under given camera hardware conditions. For example, when high-precision identification of loose bolts or micro-cracks is required, the system automatically lowers the flight altitude and increases the aerial overlap rate (no less than 30%), ensuring that the AI model can not only "see" the defect but also "accurately locate" its position, truly meeting the technical requirements of refined inspection.
[0042] In a preferred embodiment, the structured inspection report is automatically associated with the oilfield geographic information system (GIS) and asset management system. The defect location is mapped to the surface of the corresponding component in the digital twin pumping unit 3D model and linked with historical maintenance records to support preventive maintenance decisions.
[0043] The structured inspection reports generated by the system not only include information such as defect type, location, and description, but also automatically interface with the oilfield's existing Geographic Information System (GIS) and asset management system to achieve seamless data integration. Specifically, each defect in the report is precisely mapped onto a 3D model of the digital twin pumping unit—a virtual mirror built based on the real equipment, which visually displays the location of the defect on the surface of specific components (such as gearboxes, connecting rods, or suspension cables). Simultaneously, the system automatically retrieves the pumping unit's historical maintenance records, past defect trends, and operating condition data, comparing and analyzing the current findings with historical information. For example, if similar problems occur repeatedly in a certain area, the system will mark it as a potential risk of deterioration and send early warning suggestions to maintenance personnel. This deep integration of real-time inspection results, 3D visualization, and historical maintenance data effectively supports the shift from "passive maintenance" to "predictive maintenance," significantly improving the intelligence level and decision-making efficiency of oilfield equipment management.
[0044] In a preferred embodiment, a closed-loop optimization mechanism for model performance is also included: after each inspection task is completed, the system automatically selects high-quality defect samples that have been manually verified, incrementally updates the training dataset of the eighteen-point refined inspection model and the dedicated small model defect detector, and periodically triggers model fine-tuning to achieve continuous evolution of defect recognition capabilities.
[0045] The system also incorporates a closed-loop optimization mechanism for model performance, enabling continuous improvement in intelligent recognition capabilities. After each inspection task, the system automatically filters out high-quality defect samples that have been manually verified from the inspection results. These samples include newly appearing crack morphologies or previously missed minor anomalies. This "real and valid" data is incrementally added to the training dataset of the 18-point refined inspection model and the dedicated small-model defect detector. Over time, these accumulating high-quality samples allow the model to gradually cover more defect types and complex scenarios. Based on this, the system periodically and automatically triggers model fine-tuning, specifically enhancing the learning effect on new defects or error-prone cases without compromising the original recognition capabilities. This closed-loop process of "inspection—feedback—learning—optimization" enables the entire system to self-evolve, with recognition accuracy and generalization performance continuously improving over time, truly achieving the goal of intelligent inspection that becomes smarter and more accurate with use.
[0046] In a preferred embodiment, a defect continuous tracking and early warning mechanism is set up: when similar defects on the same component are identified in N consecutive (N≥2) inspections and their positioning deviation is less than a preset position threshold, they are automatically marked as "continuously deteriorating defects", the early warning level is increased and the warning is pushed out with priority.
[0047] The system incorporates a continuous defect tracking and early warning mechanism to proactively identify the evolving trends of potential risks. Its working principle is as follows: During each inspection, the system not only records the currently discovered defects but also compares them with historical inspection data for the pumping unit. When similar types of defects appear on the same component (e.g., multiple detections of oil leaks near the gearbox or cracks in the connecting rod area), and the spatial deviation of these defects is less than a preset location threshold (i.e., they essentially appear in the same area), and they occur consecutively at least twice (N≥2), the system automatically marks the defect as a "continuously worsening defect." These defects are given a higher warning level and are prioritized for notification to maintenance personnel, indicating potential issues such as fatigue accumulation, material aging, or incomplete repairs. Through this dynamic tracking method based on spatiotemporal consistency, the system can upgrade from "single snapshot" detection to "long-term trend awareness," effectively supporting early intervention and preventative maintenance, preventing small problems from escalating into major failures.
[0048] To avoid the underreporting of high-risk defects or the proliferation of low-risk alarms due to the use of a uniform alarm threshold, this invention introduces a risk-weighted early warning mechanism based on the criticality of components and the severity of defects. The system provides warnings for each standard inspection point. Pre-assign a structural safety weight and for each type of defect Define a functional impact coefficient Both modulate the original confidence level output by the model to generate an effective confidence level for alarm decision-making. The calculation formula is as follows: in, The effective confidence level of the defect is used as a risk-weighted score for alarm judgment; This represents the original defect confidence level output by the dedicated small model or re-inspection model (value range [0, 1]); Indicates the first The structural safety weight of each standard inspection point reflects the criticality of the component in the overall load-bearing structure, kinematic pairs, or power transmission chain (e.g., core moving parts such as crank pins and donkey heads have values close to 1, such as 0.90-0.95; auxiliary devices such as fences and insulated boxes have lower values, such as 0.2-0.4). Indicates the first The functional impact coefficient of a defect type characterizes the potential hazard of the defect to the safe operation of equipment or the continuity of production (e.g., high-risk defects such as cracks, fractures, and severe oil leaks are considered). Meanwhile, appearance defects such as paint peeling and minor rust are taken into account. The system sets basic alarm thresholds. (Typical value is 0.6), when the following conditions are met: Only then is a formal early warning work order generated and pushed to the operation and maintenance platform. This solution ensures that even if the initial confidence level is slightly lower than the normal threshold, if the defect occurs in a high-weight component and is of a high-risk type (such as a micro-crack in a crank pin), a valid alarm can still be triggered; conversely, high-confidence but low-risk defects (such as partial paint peeling on a fence) can be reasonably suppressed, significantly improving alarm accuracy and operation and maintenance efficiency. The weighting... With coefficient The domain knowledge base stored in the cloud-based intelligent review and management subsystem supports dynamic updates based on historical oilfield fault statistics, expert rules, or safety specifications, enabling continuous optimization and adaptive evolution of early warning strategies.
[0049] In a preferred embodiment, the eighteen standard inspection points are defined based on the core load-bearing structure, kinematic pairs, power transmission chain, and safety devices of the beam pumping unit, including: first donkey head, second donkey head, walking beam, connecting rod, tower base, suspension rope device, polished rod, insulation box, balancer, reducer, belt, motor, rope, base, fence, crossbeam, distribution box, and crank pin.
[0050] The eighteen standard inspection points were not randomly selected, but systematically defined based on the mechanical structure and operating characteristics of the beam pumping unit. They primarily cover four key areas: first, the core load-bearing structure bearing the main loads (e.g., base, tower base, crossbeam); second, frequently moving kinematic components (e.g., walking beam, connecting rod, crank pin); third, the power transmission chain (e.g., motor, belt, reducer, balancer); and fourth, safety and auxiliary devices ensuring operational safety (e.g., distribution box, enclosure, insulation box). The walking beam, requiring imaging from both sides for comprehensive inspection of cracks and other defects, is divided into two independent points: "First Walking Beam" and "Second Walking Beam." Ropes, suspension ropes, and polished rods are critical components for wellhead connection and sealing, prone to breakage, detachment, or leakage. By decomposing the entire unit into these 18 standardized inspection units with clear engineering significance, the system ensures that the model analyzes each high-risk area specifically, avoiding omissions, thus achieving comprehensive, efficient, and structured refined inspection. See also... Figure 2Eighteen bounding box regions were identified, including the first donkey head region 1, the second donkey head region 2, the walking beam region 3, the connecting rod region 4, the tower base region 5, the suspension rope region 6, the bare pole region 7, the insulation box region 8, the balancer region 9, the reducer region 10, the belt region 11, the motor region 12, the rope region 13, the base region 14, the fence region 15, the crossbeam region 16, the distribution box region 17, and the crank pin region 18. It should be noted that rope region 13 includes suspension rope region 6 and bare pole region 7, serving as enhanced recognition. Of course, if enhanced recognition is not needed, a refined inspection model with 17 bounding box regions and 17 points can be set up; only the corresponding settings are required. In practice, because the components in the rope region are particularly fine, enhanced recognition was implemented to improve the accuracy of the recognition.
[0051] In a preferred embodiment, the UAV-based precision inspection system for beam pumping units includes: an automatic UAV inspection subsystem, an edge computing analysis subsystem, and a cloud-based intelligent verification and management subsystem; wherein: The unmanned aerial vehicle (UAV) automatic inspection subsystem controls the UAV nest to complete autonomous take-off and landing, orbiting flight and multimodal image acquisition; The edge computing analysis subsystem is deployed in the oilfield and runs an 18-point refined inspection model and a group of dedicated small model defect detectors to achieve real-time component disassembly and initial defect screening. The cloud-based intelligent review and management subsystem integrates a multimodal visual large model, a digital twin platform, a model training engine, and an inspection database. It is responsible for reviewing low-confidence tasks, generating reports, and performing 3D visualization playback.
[0052] The UAV-based precision inspection system adopts a three-tiered collaborative architecture of "end-edge-cloud" to achieve a closed-loop process from data acquisition to intelligent decision-making. First, the UAV automatic inspection subsystem, relying on a drone nest deployed at the oilfield site, autonomously takes off along a preset route, flies around the pumping unit at multiple angles, and simultaneously acquires visible light and infrared images. After the mission, it automatically returns and is recovered, requiring no manual intervention throughout the entire process. The acquired images are transmitted in real-time to the edge computing analysis subsystem. This subsystem, deployed on an edge server close to the work site, runs an 18-point precision inspection model and a dedicated small-model defect detector group. It can complete the overall image segmentation, location of 18 key components, and preliminary defect screening within seconds, ensuring high efficiency and low latency. For suspected defects with low confidence levels or those difficult to determine, the system uploads relevant data to a cloud-based intelligent review and management subsystem. This cloud platform integrates a multimodal visual large model, a digital twin pumping unit model, an inspection database, and a model training engine. It not only performs in-depth semantic review of difficult samples but also automatically generates structured reports and accurately maps defects onto a 3D digital twin for visual playback. The entire system, through the organic combination of rapid edge response and in-depth cloud analysis, significantly improves identification reliability while ensuring real-time performance, providing efficient, intelligent, and traceable technical support for oilfield operation and maintenance.
[0053] In a preferred embodiment, the edge computing analysis subsystem adopts a containerized deployment of a dedicated small model defect detector group, equipped with a dynamic task scheduling and load balancing module, supports concurrent inference of multiple component sub-images, and ensures that the end-to-end processing latency is ≤ 2 seconds per pumping unit. The cloud-based intelligent review and management subsystem's multimodal large model module integrates a domain knowledge base and a prompt word template library, which can automatically select the optimal reasoning strategy based on the defect type, thereby improving the accuracy and interpretability of the review.
[0054] The edge computing analysis subsystem employs containerization technology to deploy multiple dedicated small-model defect detectors. Each container independently runs a lightweight model optimized for a specific component (such as a reducer or suspension cable), ensuring environmental isolation and flexible updates. The system is equipped with a dynamic task scheduling and load balancing module. After the 18-point refined inspection model outputs 18 component sub-images, the scheduler intelligently allocates tasks based on the current computational load of each model, enabling concurrent inference of multiple images and avoiding processing bottlenecks. This keeps the end-to-end latency from image input to initial screening result output for the entire pumping unit within 2 seconds, meeting on-site real-time requirements. Meanwhile, the multimodal big model in the cloud-based intelligent review and management subsystem does not process all review requests in a "one-size-fits-all" manner. Instead, it integrates a domain knowledge base and a prompt word template library. For example, it calls up relevant knowledge about sealing structures for defects such as "oil leaks" and activates material fatigue priors for "cracks." The system can automatically match the optimal prompt words and reasoning strategies based on the initially identified defect types, guiding the big model to focus on key features. This not only improves the accuracy of review but also generates more professional and interpretable semantic descriptions, making the final judgment results more reliable and closer to the needs of on-site operation and maintenance.
[0055] In a preferred embodiment, the system includes a model lifecycle management module that automatically monitors the online performance metrics of each model (such as mAP and F1-score) and drives the full-process automation from data annotation, training, validation to A / B testing deployment, thereby enabling continuous iteration of AI models.
[0056] The system's built-in model lifecycle management module is responsible for the full-lifecycle automated operation and maintenance of all models (including the 18-point refined inspection model and various specialized small models). This module continuously monitors the online performance indicators of the models during actual inspections, such as detection accuracy (mAP), defect identification precision and recall (F1-score), etc. Once a performance degradation or frequent occurrence of new defect types is detected, an optimization process is automatically triggered: First, high-value samples are selected from the latest inspection data and handed over to the annotation tool for manual or semi-automatic annotation; then, the training pipeline is automatically started to complete model fine-tuning, verification and evaluation, and A / B testing is used to compare the new version with the current online model in parallel to ensure that the improvement is effective and stable; only the new model that passes the verification will be automatically deployed to the corresponding nodes at the edge or cloud. The entire process requires no manual intervention, realizing a closed-loop iteration from "problem discovery" to "model update", ensuring that the system's identification capability is always in the optimal state and adapting to the actual needs of oilfield equipment aging, environmental changes and the continuous emergence of new defects.
[0057] In a preferred embodiment, the digital twin mapping function is built based on the CAD model of the oil pumping unit or the real-world 3D reconstruction model. It can accurately project the 2D defect bounding box onto the surface of the 3D component through homography transformation and depth estimation, supporting immersive defect review and training in a VR / AR environment.
[0058] The system's digital twin mapping function first constructs a virtual mirror based on the CAD design model of the pumping unit or a high-precision 3D model generated through real-scene 3D reconstruction technology, realistically restoring the equipment's geometry and component layout. When AI identifies defects in a 2D inspection image and outputs a bounding box, the system combines camera pose, lens parameters, and scene depth information, using homography transformation and depth estimation algorithms to precisely "reverse-project" this 2D box onto the corresponding physical location in the 3D model. For example, it accurately maps the oil leak area on the polished rod to the polished rod surface of the digital twin. This allows maintenance personnel to not only visually view the spatial distribution of defects within the entire unit on a computer, but also conduct immersive debriefing in VR or AR environments—wearing a headset allows them to "walk into" the virtual pumping unit, closely observe defect details, and even overlay historical maintenance records for comparative analysis. This function not only improves the visualization accuracy of defect location but also provides an efficient and intuitive digital tool for remote diagnosis, technical training, and emergency drills.
[0059] This invention discloses a method and system for refined inspection of beam pumping units using unmanned aerial vehicles (UAVs). The method constructs an 18-point refined inspection model based on an improved deep learning architecture, automatically and accurately extracting 18 predefined key component sub-images from the overall machine image. Then, it employs a dedicated small-model defect detector for parallel and rapid preliminary defect identification. Innovatively, it introduces a multimodal visual large model that integrates prior knowledge of the domain to perform semantic-level secondary review of low-confidence and difficult samples, forming a three-level collaborative intelligent discrimination mechanism of "splitting-preliminary inspection-re-review." The system integrates an "edge-cloud" collaborative architecture of automatic UAV inspection, real-time edge analysis, and deep cloud verification, and possesses digital twin mapping, risk-weighted early warning, and model self-optimization capabilities. This invention effectively solves the problems of traditional inspections relying on manual labor, low efficiency, and high rates of missed detection and misjudgment. It achieves automated, high-precision, and interpretable intelligent identification and full lifecycle management of defects in the core components of beam pumping units, significantly improving the intelligence level and safety assurance capabilities of oilfield equipment operation and maintenance.
[0060] Although the present invention has been described in detail above with reference to specific embodiments, this is not intended to limit the invention. Those skilled in the art, upon understanding the core concept of the invention—the three-level collaborative mechanism of refined 18-point decomposition, parallel initial detection using dedicated models, and semantic re-examination of a multimodal large model—can make various modifications, equivalent substitutions, or adaptive improvements. Therefore, any obvious changes or derivative solutions based on the spirit and essence of the principles and architecture disclosed in this invention should fall within the scope of protection defined by the appended claims. The description of preferred embodiments in this specification is for illustrative and explanatory purposes only and should not be construed as the sole limitation on the scope of protection of this invention.
Claims
1. An unmanned aerial vehicle (UAV) fine inspection method for a beam pumping unit, comprising: the UAV flying along a flight path to obtain images of the pumping unit, and performing defect recognition on the obtained images to output an inspection result, characterized in that: the defect recognition on the obtained images of the pumping unit to output the inspection result comprises: constructing and deploying a deep learning-based multi-point fine inspection model; the model can automatically identify, locate and split component sub-images corresponding to eighteen standard inspection points by learning from overall images of the pumping unit taken from different angles and distances; the plurality of component sub-images are input into a pre-trained special-purpose small model defect detector corresponding to the points one by one for parallel analysis to obtain preliminary defect recognition results; the output results include defect categories and bounding box coordinates; defects with a confidence level lower than a set threshold or belonging to a data sparse category in the preliminary defect recognition results are determined and sent together with the corresponding component sub-images to a multi-modal visual large model for secondary semantic-level re-inspection; the multi-modal visual large model combines image content and injected component structure prior knowledge to generate final defect determination and description. fusing the preliminary results of the special-purpose small model defect detector and the re-inspection results of the multi-modal visual large model to generate a structured inspection report containing defect types, accurate positions, point number, confidence level and semantic description; the UAV is automatically dispatched to take off and land by a UAV nest deployed on the oilfield site, and performs fully automatic inspection according to a precise re-shooting flight path around the pumping unit; the special-purpose small model defect detector is a lightweight target detection network supporting edge real-time inference; the multi-modal visual large model is a visual-language model with open vocabulary understanding and zero-shot inference capability. the training of the special-purpose small model defect detector adopts a hybrid data strategy, including real historical defect samples and simulated defect images synthesized through physical simulation; for fine defects such as loose bolts and paint peeling, edge enhancement images extracted through traditional image processing are fused in the input, and SIoU Loss is used as the bounding box regression loss function to improve the positioning accuracy.
2. The unmanned aerial vehicle fine inspection method for a beam-pumping unit according to claim 1, characterized in that, the multi-modal visual large model adopts a double-path inference mechanism during secondary re-inspection: for known defect types, structured prior knowledge is injected through prompt words to guide the model to focus on key areas; for unknown or rare defects, the zero-shot transfer capability is activated, and open vocabulary recognition and semantic description generation are performed based on a general visual-language alignment space.
3. The unmanned aerial vehicle fine inspection method for the beam-pumping unit according to claim 1, characterized in that, the images collected by the UAV include high-resolution visible light images and infrared thermal imaging images, and a multi-height layer surrounding flight strategy is adopted to ensure full coverage of the top, middle, bottom and moving connection parts of the pumping unit; the overlap rate between adjacent images is not less than 30% to ensure the continuity of subsequent image stitching and defect positioning.
4. The unmanned aerial vehicle fine inspection method for the beam-pumping unit according to claim 1, characterized in that, to control the positioning deviation of defects in images within a preset threshold, the UAV flight height H, camera parameters p, f and adjacent image overlap rate R satisfy the following relationship: 5. The unmanned aerial vehicle fine inspection method for the beam-pumping unit according to claim 1, characterized in that, 6. The unmanned aerial vehicle fine inspection method for a beam-pumping unit according to claim 1, characterized in that, wherein, represents the maximum allowed positioning deviation of the defect on the image plane; represents the flight height of the drone relative to the pumping unit component; represents the physical size of a single pixel of the camera sensor; represents the focal length of the camera lens; is the typical pixel-level positioning error of the defect recognition model on the image plane; represents the overlap ratio between adjacent aerial images, and .
7. The unmanned aerial vehicle fine inspection method for a beam-pumping unit according to claim 6, characterized in that, Set up a defect continuous tracking early warning mechanism: when similar defects on the same component are identified in consecutive N times (N≥2) inspections, and the positioning deviation is less than the preset position threshold, it is automatically marked as "continuous deterioration defect", the warning level is raised and it is pushed preferentially.
8. The unmanned aerial vehicle fine inspection method for a beam-pumping unit according to claim 1, characterized in that, The eighteen standard inspection points are defined according to the core bearing structure, kinematic pair, power transmission chain and safety device of the beam pumping unit, and include: first horse head, second horse head, beam, connecting rod, tower seat, rope hanger, polished rod, heat preservation box, balancer, speed reducer, belt, motor, rope, base, fence, cross beam, distribution box, crank pin.
9. An unmanned aerial vehicle fine inspection system for a beam pumping unit implementing the method of any one of claims 1-8, characterized in that, The system comprises: an unmanned aerial vehicle automatic inspection subsystem, an edge computing analysis subsystem, a cloud intelligent review and management subsystem; wherein: The unmanned aerial vehicle automatic inspection subsystem controls the nest to complete autonomous take-off and landing, circumferential flight and multi-modal image acquisition; The edge computing analysis subsystem is deployed at the oilfield site, runs a multi-point fine inspection model and a group of special small model defect detectors, realizes real-time component splitting and defect preliminary screening; The cloud intelligent review and management subsystem integrates a multi-modal visual large model, a digital twin platform, a model training engine and an inspection database, is responsible for low confidence task review, report generation and three-dimensional visualization playback.
10. The unmanned aerial fine inspection system for beam pumping unit according to claim 9, characterized in that, The edge computing analysis subsystem adopts containerized deployment of a group of special small model defect detectors, is equipped with a dynamic task scheduling and load balancing module, supports concurrent inference of multiple component sub-images, and ensures that the end-to-end processing delay is ≤ 2 seconds per pumping unit; The multi-modal large model module of the cloud intelligent review and management subsystem integrates a domain knowledge base and a prompt word template library, can automatically select the optimal inference strategy according to the defect type, and improves the review accuracy and interpretability.
Citation Information
Patent Citations
A kinematic analysis method for beam pumping units based on deep learning
CN113837004B
Cited By
Power distribution network inspection analysis system and platform
CN122052303A