A drilling target recognition system for visual feature extraction inside oil well casing

By employing multimodal image acquisition and adaptive preprocessing, multi-scale feature extraction, and multi-task collaborative recognition and tracking modules, combined with edge-cloud collaborative computing, the problems of poor image quality and low target recognition accuracy inside oil well pipes have been solved, achieving high-precision, real-time drilling target recognition.

CN122493377APending Publication Date: 2026-07-31XI'AN PETROLEUM UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XI'AN PETROLEUM UNIVERSITY
Filing Date
2026-05-12
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies suffer from poor image quality and low recognition accuracy within oil well casings, making it difficult to process both static and dynamic targets simultaneously. Furthermore, existing models involve large computational loads or insufficient recognition accuracy, failing to meet the real-time processing requirements of downhole systems.

Method used

By employing a multimodal image acquisition device combined with an adaptive preprocessing module, a multi-scale hierarchical feature extraction network, and a multi-task collaborative recognition and tracking module, along with an edge-cloud collaborative computing architecture, image quality improvement and target recognition accuracy are achieved.

Benefits of technology

Acquiring high-quality images in complex environments significantly improves the recognition accuracy of drilling targets at different scales, enabling static target detection and dynamic target tracking, and ensuring the real-time performance and intelligence level of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493377A_ABST
    Figure CN122493377A_ABST
Patent Text Reader

Abstract

This invention discloses a drilling target recognition system for visual feature extraction inside oil well casings, belonging to the interdisciplinary field of petroleum drilling and computer vision. It includes a multimodal image acquisition device, an adaptive preprocessing module, a multi-task collaborative recognition and tracking module, an edge-cloud collaborative computing module, and a depth calibration and positioning module. The multimodal image acquisition device is used to acquire visible light, infrared, and structured light images inside the oil well casing. The adaptive preprocessing module is connected to the multimodal image acquisition device and is used to preprocess the acquired multimodal images. By employing a multimodal image acquisition device that fuses visible light, infrared, and structured light modal data, high-quality image information can be acquired in complex environments such as high dust, low light, and uneven lighting. The adaptive preprocessing module is designed to further improve image quality through adaptive illumination compensation, multimodal fusion dehazing and noise reduction, and image super-resolution reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of interdisciplinary technology of oil drilling and computer vision, and more specifically, to a drilling target recognition system for extracting visual features inside oil well casings. Background Technology

[0002] In oil drilling, timely and accurate acquisition of downhole information is crucial for ensuring drilling safety, improving drilling efficiency, and reducing drilling costs. Traditional downhole information acquisition methods mainly rely on logging techniques, such as sonic logging, resistivity logging, and radioactive logging. While these methods can provide some downhole geological and engineering information, they suffer from drawbacks such as low resolution, poor real-time performance, and high cost. In recent years, with the rapid development of computer vision technology, vision-based downhole inspection technology has been increasingly widely used. By deploying image acquisition devices downhole, image information inside the well casing can be directly acquired. Then, through image processing and pattern recognition technology, various drilling targets such as tubing couplings, perforations, fractures, cuttings, and fluids can be identified, providing intuitive and detailed downhole information for the drilling process.

[0003] However, the environment inside oil well tubing is extremely complex, with problems such as high dust levels, water mist, low light levels, and uneven lighting, resulting in poor image quality, high noise, and low contrast, severely impacting the accuracy of subsequent target recognition. Some drilling targets inside the tubing, such as small fractures, small perforations, and tiny rock cuttings, are small in size and occupy a low percentage of pixels in the image. Existing models struggle to effectively extract features from these small targets, leading to a high false negative rate. Different types of drilling targets vary greatly in scale; for example, the diameter of a tubing coupling can reach tens of centimeters, while the width of a small fracture may only be a few millimeters. Existing models struggle to effectively handle targets of different scales simultaneously, resulting in incomplete recognition of large targets and missed detection of small targets. High-precision deep learning models typically involve large computational loads and numerous parameters, making real-time processing on downhole edge devices difficult; while lightweight models offer fast computation speeds, their recognition accuracy often falls short of requirements.

[0004] During drilling, targets such as cuttings and fluids are dynamically moving. Most existing systems can only perform static target detection and cannot continuously track and analyze the state of dynamic targets, making it difficult to obtain the target's trajectory and velocity information.

[0005] In view of this, the present invention is proposed to solve the above-mentioned technical problems. Summary of the Invention

[0006] The purpose of this invention is to provide a drilling target recognition system for visual feature extraction inside oil well pipes, in order to solve the technical problems of poor image quality, low recognition accuracy, poor real-time performance, and inconvenience in processing static and dynamic targets simultaneously in the complex environment inside oil well pipes of existing drilling target recognition systems.

[0007] To achieve the above objectives, the present invention provides the following technical solution: A drilling target recognition system for visual feature extraction inside oil well casing includes: A multimodal image acquisition device for acquiring visible light, infrared, and structured light images inside an oil well pipe; An adaptive preprocessing module, connected to a multimodal image acquisition device, is used to preprocess the acquired multimodal images, including adaptive illumination compensation, multimodal fusion dehazing and denoising, and image super-resolution reconstruction. A multi-scale hierarchical feature extraction network, connected to an adaptive preprocessing module, is used to extract multi-scale, multi-level visual features from preprocessed images. The multi-task collaborative identification and tracking module is connected to a multi-scale hierarchical feature extraction network to simultaneously achieve drilling target detection, semantic segmentation, instance segmentation, and continuous tracking of dynamic targets. The edge-cloud collaborative computing module is connected to the adaptive preprocessing module, the multi-scale hierarchical feature extraction network, and the multi-task collaborative recognition and tracking module, respectively, to realize the collaborative work of downhole edge computing and cloud deep learning; The depth calibration and positioning module, connected to the multi-task collaborative identification and tracking module, is used to combine visual odometer and coupling identification information to achieve high-precision depth calibration and three-dimensional positioning of the drilling target.

[0008] Furthermore, the multimodal image acquisition device includes: High-definition visible light camera, used to capture visible light images inside oil well pipes; Infrared thermal imaging camera, used to acquire infrared thermal images of the inside of oil well pipes; Line structured light emitter, used to project line structured light onto the inner wall of oil well pipe; Ring-shaped LED lighting source for providing uniform visible light illumination; Infrared illumination source, used to provide infrared illumination; The explosion-proof housing is used to house all the above components and is adapted to the high temperature, high pressure and high corrosion environment downhole.

[0009] Furthermore, the adaptive preprocessing module includes: The adaptive illumination compensation unit is used to enhance images with low illumination and uneven illumination by adopting an adaptive illumination compensation algorithm based on Retinex theory, according to the local brightness information of the image. The multimodal fusion dehazing and denoising unit is used to fuse information from visible light images, infrared images, and structured light images. It employs a dehazing algorithm based on multi-scale guided filtering and an adaptive median filtering algorithm to remove dust, water mist, and noise interference from the image. The image super-resolution reconstruction unit is used to reconstruct a high-resolution image from a low-resolution image using a super-resolution reconstruction algorithm based on a residual dense feature aggregation network.

[0010] Furthermore, the multi-scale hierarchical feature extraction network adopts a CNN-Transformer dual-branch collaborative architecture, including: The CNN backbone branch is used to extract local detail features of the image. It uses ResNet50 as the base network and introduces the Res2Net multi-scale feature representation module into it. The Transformer auxiliary branch is used to extract global contextual features of the image. It uses the Swing Transformer as the base network and introduces a spatial attention mechanism. The cross-stage feature interaction module is used to realize the interaction and fusion of features from different stages between the CNN backbone branch and the Transformer auxiliary branch; A multi-scale feature pyramid structure is used to fuse feature maps of different scales to generate multi-scale feature representations.

[0011] Furthermore, the multi-task collaborative identification and tracking module includes: The target detection unit is used to detect drilling targets in the image, including couplings, perforations, fractures, cuttings, and fluids; Semantic segmentation unit, used to perform pixel-level semantic segmentation of images to distinguish different types of drilling targets and background; The instance segmentation unit is used to distinguish different instances of the same category, enabling the individual identification of targets such as multiple rock cuttings and multiple perforations; The dynamic target tracking unit is used to continuously track dynamic targets using an improved DeepSort algorithm, and to acquire the target's motion trajectory and velocity information.

[0012] Furthermore, the target detection unit adopts an improved YOLOv8 model, introducing a bidirectional feature pyramid network and an adaptive feature fusion mechanism in its neck network, and a decoupled head structure in its head network.

[0013] Furthermore, the edge-cloud collaborative computing module includes: The downhole edge computing unit is deployed inside the multimodal image acquisition device to perform image preprocessing, lightweight target detection, and data compression and transmission tasks. Cloud computing units, deployed in ground data centers, are used to perform complex tasks such as model training, model updating, deep data analysis, and result visualization. The data transmission unit is used to realize data transmission between the downhole edge computing unit and the cloud computing unit. It adopts an adaptive data compression algorithm and a breakpoint resume mechanism to reduce bandwidth consumption.

[0014] Furthermore, the depth calibration and positioning module includes: The coupling recognition unit is used to identify oil pipe couplings in the image and record the depth information of the couplings; The visual odometry unit is used to calculate the motion distance and speed of the image acquisition device based on feature matching between consecutive image frames; The depth calibration unit is used to calibrate the depth error of traditional depth measurement systems by combining the depth information from the coupling recognition and the motion information from the visual odometry. The 3D reconstruction unit is used to reconstruct the 3D morphology of the inner wall of the oil well pipe based on structured light images and depth information.

[0015] Furthermore, it also includes a drilling target knowledge base, which is connected to a multi-task collaborative identification and tracking module to store the feature information, attribute information and historical identification data of various drilling targets, and supports intelligent classification and attribute analysis of target types.

[0016] Furthermore, it also includes an early warning module, which is connected to the multi-task collaborative identification and tracking module and the depth calibration and positioning module, to issue an early warning signal and display the target's location and attribute information when an abnormal target is identified.

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. Employing a multimodal image acquisition device that integrates visible light, infrared, and structured light data, it can acquire high-quality image information in complex environments such as high dust, low light, and uneven lighting. Simultaneously, an adaptive preprocessing module is designed, employing adaptive illumination compensation, multimodal fusion dehazing and noise reduction, and image super-resolution reconstruction techniques to further improve image quality, providing a reliable foundation for subsequent target recognition.

[0018] 2. Employing a CNN-Transformer dual-branch collaborative architecture and a multi-scale feature pyramid structure, this approach simultaneously extracts local detail features and global contextual features from images, effectively fusing feature information from different scales. This significantly improves the feature extraction capability and recognition accuracy for drilling targets at different scales. Particularly for small and dense targets, the improved YOLOv8 model further enhances detection accuracy.

[0019] 3. It integrates a multi-task collaborative recognition and tracking framework, enabling simultaneous target detection, semantic segmentation, instance segmentation, and continuous tracking of dynamic targets. It can not only identify static targets (such as couplings, perforations, and fractures) but also track dynamic targets (such as rock cuttings and fluids), acquiring their trajectory and velocity information to provide more comprehensive downhole information for the drilling process.

[0020] 4. An edge-cloud collaborative computing architecture is adopted, which rationally distributes computing tasks between downhole edge devices and cloud servers. Downhole edge devices are responsible for performing tasks with high real-time requirements, such as image preprocessing and lightweight object detection; cloud servers are responsible for performing computationally intensive tasks, such as model training and deep data analysis. This architecture improves the system's intelligence level while ensuring real-time performance.

[0021] 5. By using a depth calibration and positioning module, combined with visual odometry and coupling recognition information, the depth error of the traditional depth sounding system is calibrated, achieving high-precision depth positioning of the drilling target. Simultaneously, the three-dimensional morphology of the well casing's inner wall is reconstructed using structured light imaging, enabling three-dimensional positioning of the target. Attached Figure Description

[0022] The accompanying drawings, which form part of this application, are used to provide a further understanding of the invention. The illustrative embodiments and descriptions of the invention are used to explain the invention, but do not constitute an undue limitation of the invention. Obviously, the drawings described below are merely some embodiments, and those skilled in the art can obtain other drawings based on these drawings without creative effort. In the drawings: Figure 1 A framework diagram of a drilling target recognition system for extracting visual features inside an oil well pipe, provided in an embodiment of this application. Detailed Implementation

[0023] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0024] See Figure 1As shown, a drilling target recognition system for visual feature extraction inside an oil well casing includes a multimodal image acquisition device, an adaptive preprocessing module, a multi-scale hierarchical feature extraction network, a multi-task collaborative recognition and tracking module, an edge-cloud collaborative computing module, and a depth calibration and positioning module. The multimodal image acquisition device is used to acquire visible light images, infrared images, and structured light images inside the oil well casing. The multimodal image acquisition device is connected to the ground control system via a cable and can move up and down inside the oil well casing to acquire image data at different depths. The adaptive preprocessing module is connected to the multimodal image acquisition device and is used to preprocess the acquired multimodal images, including adaptive illumination compensation, multimodal fusion dehazing and denoising, and image super-resolution reconstruction. A multi-scale hierarchical feature extraction network is connected to an adaptive preprocessing module to extract multi-scale, multi-level visual features from preprocessed images. A multi-task collaborative recognition and tracking module is connected to the multi-scale hierarchical feature extraction network to simultaneously achieve drilling target detection, semantic segmentation, instance segmentation, and continuous tracking of dynamic targets. An edge-cloud collaborative computing module is connected to the adaptive preprocessing module, the multi-scale hierarchical feature extraction network, and the multi-task collaborative recognition and tracking module to achieve collaborative work between downhole edge computing and cloud-based deep learning. A depth calibration and positioning module is connected to the multi-task collaborative recognition and tracking module to combine visual odometry and coupling recognition information to achieve high-precision depth calibration and three-dimensional positioning of drilling targets.

[0025] The multimodal image acquisition device includes: a high-definition visible light camera, an infrared thermal imaging camera, a line structured light emitter, a ring LED illumination source, an infrared illumination source, and an explosion-proof housing. The high-definition visible light camera uses a 2-megapixel industrial-grade CMOS camera with a frame rate of 30fps and a resolution of 1920×1080, used to acquire visible light images inside oil well pipes. The infrared thermal imaging camera uses an uncooled infrared focal plane detector with a resolution of 640×512, a frame rate of 25fps, and a temperature measurement range of -20℃ to +150℃, used to acquire infrared thermal images inside oil well pipes. The line structured light emitter… The device employs a 650nm semiconductor laser to project a thin, bright line structured light beam onto the inner wall of the well casing. A ring-shaped LED illumination source, composed of 36 high-brightness LEDs, is arranged in a ring around the visible light camera to provide uniform visible light illumination. An infrared illumination source, consisting of 16 infrared LEDs, is arranged in a ring around the infrared thermal imaging camera to provide infrared illumination and enhance the brightness of the infrared image. The explosion-proof housing is made of stainless steel, capable of withstanding pressures of 100MPa and temperatures of 150℃, and exhibits excellent corrosion resistance. A transparent observation window made of sapphire glass is located at the front of the housing, offering high light transmittance and high wear resistance, housing all the aforementioned components and adapting to the high-temperature, high-pressure, and highly corrosive environment of the well.

[0026] The adaptive preprocessing module includes: The adaptive illumination compensation unit enhances images with low illumination and uneven lighting by employing an adaptive illumination compensation algorithm based on Retinex theory, using local brightness information. First, the input image is converted to the HSV color space, and the luminance component V is separated. Then, Gaussian filtering is applied to the luminance component V to obtain the illumination component L. Next, an adaptive gain coefficient is calculated based on the local statistical information of the illumination component L. Finally, the gain coefficient is applied to the luminance component V, and the image is converted back to the RGB color space to obtain the illumination-compensated image. This algorithm can automatically adjust the gain coefficient based on the local brightness information of the image, effectively solving the problems of low illumination and uneven lighting.

[0027] The multimodal fusion dehazing and denoising unit fuses information from visible light, infrared, and structured light images. It employs a multi-scale guided filtering-based dehazing algorithm and an adaptive median filtering algorithm to remove dust, water mist, and noise interference from the images. First, the visible light image is dehazed using a multi-scale guided filtering-based algorithm. This algorithm utilizes edge information from the infrared image as a guide, better preserving edge details. Then, an adaptive median filtering algorithm is used to denoise the dehazed image. This algorithm automatically adjusts the size of the filtering window based on the local noise level of the image, effectively protecting image details while removing noise. Finally, the processed visible light image is fused with the infrared and structured light images to obtain the final dehazed and denoised image.

[0028] The image super-resolution reconstruction unit is used to reconstruct low-resolution images into high-resolution images using a super-resolution reconstruction algorithm based on a residual dense feature aggregation network (SR-RDFAN-LOG), thereby improving the image's detail clarity. The algorithm employs a residual dense feature aggregation network (SR-RDFAN-LOG)-based super-resolution reconstruction algorithm. This network introduces a Local Implicit Image Function (LIIF) and includes two stages: encoding and decoding. In the encoding stage, a local feature extraction module is first designed to integrate contextual features; secondly, a residual dense module is introduced to extract deep features; and finally, a spatial attention module is constructed to focus on high-frequency details. In the decoding stage, a multilayer perceptron is used to decode the feature map, achieving super-resolution reconstruction of well logging images at any scale. This algorithm can reconstruct low-resolution images into high-resolution images, improving image detail clarity, especially for small targets such as small fractures and small perforations, where the detail enhancement effect is significant.

[0029] The multi-scale hierarchical feature extraction network adopts a CNN-Transformer dual-branch collaborative architecture, including: The CNN backbone is used to extract local detail features from images. It uses ResNet50 as the base network and introduces a Res2Net multi-scale feature representation module. The Res2Net module increases the receptive field of the network and improves the multi-scale feature representation capability without increasing the computational cost by grouping the channels of the convolutional layers and establishing residual connections between the groups. The CNN backbone outputs feature maps at four different scales: 1 / 4, 1 / 8, 1 / 16, and 1 / 32 resolution. The Transformer auxiliary branch extracts global contextual features from the image. It uses the Swin Transformer as the base network and introduces a spatial attention mechanism. The Swin Transformer effectively reduces computational complexity by dividing the image into non-overlapping windows and performing self-attention computation within each window. The introduced spatial attention mechanism adaptively adjusts feature weights at different spatial locations, enhancing the ability to extract features from key regions. The Transformer auxiliary branch outputs four feature maps at different scales, corresponding to the feature map scales of the CNN backbone branches. The cross-stage feature interaction module enables the interaction and fusion of features from different stages between the CNN backbone branch and the Transformer auxiliary branch. For each scale of feature map, the feature maps from the CNN branch and the Transformer branch are concatenated, and then fused through a 1×1 convolutional layer to obtain the fused feature map. This cross-stage feature interaction mechanism can fully utilize the local detail features of the CNN branch and the global context features of the Transformer branch, thereby improving feature representation capabilities. A multi-scale feature pyramid structure is used to fuse feature maps of different scales to generate multi-scale feature representations. A bidirectional feature pyramid network (BiFPN) structure is employed to further fuse the four feature maps of different scales after fusion. BiFPN effectively fuses feature information of different scales through top-down and bottom-up bidirectional connections, as well as skip connections, generating multi-scale feature representations. Simultaneously, an adaptive feature fusion mechanism is introduced, which automatically adjusts the fusion weights according to the importance of features at different scales, further improving the feature fusion effect.

[0030] The multi-task collaborative identification and tracking module includes: The target detection unit is used to detect drilling targets in images, including couplings, perforations, fractures, cuttings, and fluids; it employs an improved YOLOv8 model. The bidirectional feature pyramid network and adaptive feature fusion mechanism mentioned above are introduced into its neck network to improve multi-scale feature fusion capabilities. A decoupled head structure is introduced into its head network to separate classification and regression tasks, improving detection accuracy. Furthermore, considering the prevalence of small targets within the well casing, a small target detection head is added specifically for detecting targets smaller than 32×32 pixels. The improved YOLOv8 model can accurately detect various drilling targets within the well casing, including couplings, perforations, fractures, cuttings, and fluids. The semantic segmentation unit performs pixel-level semantic segmentation of the image, distinguishing different types of drilling targets from the background. It uses the U-Net++ model as the base network, incorporating the aforementioned multi-scale hierarchical feature extraction network into its encoder. U-Net++, through nested skip connections and deep supervision, effectively fuses feature information at different scales, improving semantic segmentation accuracy. This semantic segmentation unit enables pixel-level semantic segmentation of the image, distinguishing different types of drilling targets from the background, providing a foundation for subsequent instance segmentation and target analysis. The instance segmentation unit is used to distinguish different instances of the same category, enabling the individual identification of targets such as multiple rock cuttings and multiple perforations. The Mask R-CNN model is used as the base network, and its feature extraction network is replaced by the aforementioned multi-scale hierarchical feature extraction network. Mask R-CNN can simultaneously achieve target detection and instance segmentation, and can generate different masks for different instances of the same category, enabling the individual identification of targets such as multiple rock cuttings and multiple perforations. The dynamic target tracking unit employs an improved DeepSort algorithm to continuously track dynamic targets, acquiring their trajectory and velocity information. In its feature extraction section, the aforementioned multi-scale hierarchical feature extraction network is introduced to enhance the representation capability of target features. In its data association section, the Hungarian algorithm and Kalman filtering are incorporated to achieve continuous target tracking. Furthermore, considering the high speed and frequent obstruction of targets within the oil well casing, a target re-identification mechanism is added. When a target reappears after being obscured, it can be accurately re-identified and tracking can continue. The dynamic target tracking unit can continuously track dynamic targets such as rock cuttings and fluids, acquiring their trajectory and velocity information.

[0031] The target detection unit adopts an improved YOLOv8 model, introducing a bidirectional feature pyramid network and an adaptive feature fusion mechanism in its neck network, and a decoupled head structure in its head network to improve the detection accuracy of small and dense targets.

[0032] The edge-cloud collaborative computing module includes: The downhole edge computing unit, deployed inside the multimodal image acquisition device, performs image preprocessing, lightweight target detection, and data compression and transmission tasks. It utilizes the NVIDIA Jetson Xavier NX computing platform. This platform boasts powerful AI computing capabilities, enabling real-time execution of image preprocessing, lightweight target detection, and data compression and transmission. The downhole edge computing unit first preprocesses the acquired multimodal images, then performs preliminary target detection using a lightweight target detection model. After compressing the detected target regions and the original image, the data is transmitted to a cloud server for further analysis. The cloud computing unit, deployed in the ground data center, performs complex tasks such as model training, model updating, deep data analysis, and result visualization. It employs a cluster of multiple high-performance GPU servers. The cloud computing unit receives data from the downhole edge computing unit, uses a high-precision deep learning model for further object detection, semantic segmentation, instance segmentation, and dynamic object tracking, then visualizes the analysis results and stores them in a database. Simultaneously, the cloud computing unit is also responsible for periodically updating and optimizing the model and distributing the updated lightweight model to the downhole edge computing unit. The data transmission unit is used to realize data transmission between the downhole edge computing unit and the cloud computing unit. It adopts an adaptive data compression algorithm and a breakpoint resume mechanism, automatically adjusting the data compression ratio according to network bandwidth conditions to reduce bandwidth consumption. When the network is interrupted and reconnected, it can resume data transmission from the point of interruption to avoid data loss. The data transmission unit also supports encrypted transmission to ensure data security.

[0033] The depth calibration and positioning module includes: The coupling recognition unit identifies pipeline couplings in an image and records their depth information. The target detection unit identifies the pipeline couplings in the image. Since the spacing between pipeline couplings is fixed (typically 9.6 meters), the approximate depth of the image acquisition device can be determined by recognizing the couplings. The coupling recognition unit records the depth information of each coupling and uses it as a reference point for depth calibration. The visual odometry unit employs a feature-based visual odometry calculation method. First, ORB feature points are extracted from consecutive image frames. Then, the relative motion between adjacent frames is calculated through feature matching. Finally, the motion distance and velocity of the image acquisition device are calculated through integration. The visual odometry unit can provide high-precision relative motion information, but it suffers from accumulated errors. It is used to calculate the motion distance and velocity of the image acquisition device based on feature matching between consecutive image frames. The depth calibration unit combines depth information from coupling identification and motion information from the visual odometry to calibrate the depth error of the traditional depth sounding system. When a coupling is identified, its known depth is used as a reference to correct the cumulative error of the visual odometry. Simultaneously, the calibrated depth information is fused with the depth information from the traditional depth sounding system to obtain the final high-precision depth information. The 3D reconstruction unit is used to reconstruct the 3D topography of the inner wall of the oil well casing based on structured light images and depth information, thereby achieving 3D positioning of the drilling target. First, based on the triangulation principle of line structured light, the 3D coordinates of each point on the inner wall of the oil well casing are calculated. Then, the 3D point clouds at different depths are stitched together to obtain a complete 3D point cloud model of the inner wall of the oil well casing. Finally, based on the results of target detection and semantic segmentation, the position and attribute information of each drilling target are marked in the 3D point cloud model, achieving 3D positioning of the target.

[0034] In some possible implementations, the drilling target identification system for extracting visual features inside the well casing provided in this application also includes a drilling target knowledge base, which is connected to a multi-task collaborative identification and tracking module to store feature information, attribute information and historical identification data of various drilling targets, and supports intelligent classification and attribute analysis of target types.

[0035] It should be noted that the drilling target knowledge base includes: a target feature library, a target attribute library, and a historical identification database. The target feature library stores the visual features of various drilling targets, such as color, texture, shape, and size; the target attribute library stores the attribute information of various drilling targets, such as the model of the coupling, the parameters of the perforation, and the hazard level of the fracture; the historical identification database stores all target data identified by the system in the past, including the target type, location, time, and image.

[0036] The drilling target knowledge base supports intelligent classification and attribute analysis of target types. When the system identifies a target, it compares its features with those in the target feature database to determine the target type and retrieves the target's attribute information from the target attribute database. Simultaneously, the system stores the identification results in a historical identification database, providing data support for subsequent model training and data analysis.

[0037] In some possible implementations, the drilling target identification system for extracting visual features inside the well casing provided in this application also includes an early warning module, which is connected to a multi-task collaborative identification and tracking module and a depth calibration and positioning module. When an abnormal target (such as a severe crack, casing deformation, or falling object) is identified, an early warning signal is issued in a timely manner, and the location and attribute information of the target are displayed.

[0038] The early warning module includes: Anomaly Target Definition: Defines the types and warning levels of various anomalies, such as severe cracks (Level 1 warning), casing deformation (Level 1 warning), falling objects (Level 2 warning), and minor cracks (Level 3 warning).

[0039] Early warning judgment: When the system identifies a target, it will determine whether it is an abnormal target and the warning level based on the target's type and attribute information.

[0040] Warning Output: When an abnormal target is identified, the system will promptly issue warning signals, including audible alarms, visual alarms, and on-screen display alarms. Simultaneously, it will display the target's location, type, attribute information, and warning level, reminding operators to take appropriate measures.

[0041] Working principle 1. Multimodal downhole image synchronous acquisition stage The multimodal image acquisition device at the front end of the system is lowered into the well casing along with the drilling tools or logging instruments. Under the auxiliary illumination of a ring-shaped LED light source and an infrared light source, it simultaneously acquires three types of complementary data through three independent sensors: The texture, color, shape, and other appearance features of the inner surface of the acquisition tube of a high-definition visible light camera; Infrared thermal imaging cameras collect thermal radiation characteristics of targets of different materials and temperatures (unaffected by dust or water mist). A line structured light emitter projects laser stripes onto the pipe wall to collect the three-dimensional geometric contour features of the pipe wall.

[0042] All data is transmitted to the downhole edge computing unit via a high-speed interface inside the explosion-proof housing, while retaining the original timestamps and depth markers to provide a benchmark for subsequent multimodal fusion and positioning.

[0043] 2. Adaptive Image Preprocessing Stage The adaptive preprocessing module of the downhole edge computing unit performs real-time enhancement on the original multimodal images, solving the problem of downhole image quality degradation. Adaptive illumination compensation: Based on Retinex theory, the illumination component and reflection component of the image are separated, and the gain is automatically adjusted according to local brightness statistics to eliminate dark areas and overexposed areas caused by low light and uneven illumination. Multimodal fusion dehazing and noise reduction: Guided by the clear edges of the infrared image, multi-scale guided filtering is used to remove dust and water mist interference in the visible light image, and adaptive median filtering is used to suppress salt-and-pepper noise and Gaussian noise while preserving the edge details of the target. Image super-resolution reconstruction: A residual dense feature aggregation network is used to reconstruct low-resolution images into high-resolution images, focusing on enhancing the texture details of small targets such as fine cracks and micro-perforations, laying the foundation for subsequent feature extraction.

[0044] 3. Multi-scale hierarchical feature extraction stage The preprocessed image is input into a CNN-Transformer dual-branch collaborative feature extraction network, which simultaneously extracts local details and global contextual features. The CNN backbone branch extracts local detail features such as edges, corners, and textures of the image through a ResNet50 network with embedded Res2Net modules, and outputs four feature maps at different resolutions (1 / 4, 1 / 8, 1 / 16, and 1 / 32). Transformer Auxiliary Branch: Through the window self-attention mechanism of the Swing Transformer, the global contextual dependencies of the image are extracted, the overall structural features of large-scale targets are captured, and the feature map corresponding to the scale of the CNN branch is output. Cross-stage feature interaction and fusion: The feature maps of the same scale in the two branches are spliced ​​and fused through a 1×1 convolutional layer, and then multi-scale feature transfer from top to bottom and from bottom to top is achieved through a bidirectional feature pyramid network (BiFPN). Combined with an adaptive weight mechanism, multi-scale feature representations that take into account both local details and global semantics are generated.

[0045] 4. Multi-task collaborative identification and tracking stage The fused multi-scale feature input is used in a multi-task collaborative recognition and tracking module to simultaneously perform static target detection, pixel-level segmentation, and dynamic target tracking. Target detection: The improved YOLOv8 model handles classification and regression tasks separately by decoupling the head structure. A new small target detection head is added to specifically identify targets smaller than 32×32 pixels, and outputs bounding boxes and class confidence scores for targets such as couplings, perforations, cracks, rock cuttings, and fluids. Semantic and instance segmentation: A semantic segmentation network based on U-Net++ achieves pixel-level target classification, distinguishing targets from the background; an instance segmentation network based on Mask R-CNN generates independent masks for each target of the same type, enabling the individual differentiation of multiple rock fragments and multiple perforations. Dynamic target tracking: The improved DeepSort algorithm uses the high-dimensional features extracted by the dual-branch network to perform target matching, combines Kalman filtering to predict the target motion state, solves the occlusion problem through the target re-identification mechanism, and outputs the motion trajectory and velocity information of rock debris and fluid in real time.

[0046] 5. Edge-Cloud Collaborative Computing Scheduling Phase The system balances real-time performance and accuracy requirements through an edge-cloud collaborative computing architecture: Downhole edge: Prioritize tasks with high real-time requirements (image preprocessing, lightweight target detection, data compression), and only transmit the detected target area image and key feature data to the cloud via cable, significantly reducing bandwidth consumption; Ground-to-Cloud: After receiving data from the edge, the high-precision model is called to complete semantic segmentation, instance segmentation and deep data analysis, while historical data is used to continuously train and optimize the model; Model iteration and update: The cloud periodically distributes the optimized lightweight model to the downhole edge unit, realizing online upgrades of the system's recognition capabilities without the need to pull out the drill string and replace equipment.

[0047] 6. High-precision depth calibration and 3D positioning stage The depth calibration and positioning module combines visual information with structured light data to solve the problem of large cumulative errors in traditional depth sounding systems. Reference point calibration: When the coupling identification unit detects the oil pipe coupling, it uses the known fixed spacing of the coupling (usually 9.6 meters) as the absolute reference to correct the cumulative error of the visual odometer; Relative motion calculation: The visual odometry unit calculates the real-time motion distance and speed of the acquisition device by matching ORB features between consecutive frames; 3D Reconstruction and Positioning: Based on the principle of line structured light triangulation, the 3D coordinates of each point on the pipe wall are calculated, and point cloud data at different depths are stitched together to generate a complete 3D model of the inner wall of the pipe. The identified drilling target is then projected into the 3D model to achieve accurate positioning of the target's 3D coordinates and depth.

[0048] 7. Knowledge Base Assistance and Anomaly Warning Stage Drilling target knowledge base: Stores standard features, attribute parameters and historical identification data of various targets. The system will compare the target features extracted in real time with the knowledge base and automatically supplement the target's type, hazard level and other attribute information. Anomaly Warning: When anomalies such as severe cracks, casing deformation, or falling objects are detected, the warning module will immediately issue sound, light, and screen alarms according to the preset warning level. At the same time, it will display the precise location, type, and severity of the target, providing a basis for on-site decision-making.

[0049] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A drilling target recognition system for extracting visual features inside oil well casings, characterized in that, include: A multimodal image acquisition device for acquiring visible light, infrared, and structured light images inside an oil well pipe; An adaptive preprocessing module, connected to the multimodal image acquisition device, is used to preprocess the acquired multimodal images, including adaptive illumination compensation, multimodal fusion dehazing and denoising, and image super-resolution reconstruction. A multi-scale hierarchical feature extraction network, connected to the adaptive preprocessing module, is used to extract multi-scale, multi-level visual features from the preprocessed image. The multi-task collaborative identification and tracking module is connected to the multi-scale hierarchical feature extraction network to simultaneously realize the detection of drilling targets, semantic segmentation, instance segmentation, and continuous tracking of dynamic targets. The edge-cloud collaborative computing module is connected to the adaptive preprocessing module, the multi-scale hierarchical feature extraction network, and the multi-task collaborative recognition and tracking module, respectively, to realize the collaborative work of downhole edge computing and cloud deep learning; The depth calibration and positioning module, connected to the multi-task collaborative identification and tracking module, is used to combine visual odometer and coupling identification information to achieve high-precision depth calibration and three-dimensional positioning of the drilling target.

2. The drilling target recognition system for extracting visual features inside oil well casing according to claim 1, characterized in that, The multimodal image acquisition device includes: High-definition visible light camera, used to capture visible light images inside oil well pipes; Infrared thermal imaging camera, used to acquire infrared thermal images of the inside of oil well pipes; Line structured light emitter, used to project line structured light onto the inner wall of oil well pipe; Ring-shaped LED lighting source for providing uniform visible light illumination; Infrared illumination source, used to provide infrared illumination; The explosion-proof housing is used to house all the above components and is adapted to the high temperature, high pressure and high corrosion environment downhole.

3. The drilling target recognition system for extracting visual features inside oil well casing according to claim 2, characterized in that, The adaptive preprocessing module includes: The adaptive illumination compensation unit is used to enhance images with low illumination and uneven illumination by adopting an adaptive illumination compensation algorithm based on Retinex theory, according to the local brightness information of the image. The multimodal fusion dehazing and denoising unit is used to fuse information from visible light images, infrared images, and structured light images. It employs a dehazing algorithm based on multi-scale guided filtering and an adaptive median filtering algorithm to remove dust, water mist, and noise interference from the image. The image super-resolution reconstruction unit is used to reconstruct a high-resolution image from a low-resolution image using a super-resolution reconstruction algorithm based on a residual dense feature aggregation network.

4. The drilling target recognition system for extracting visual features inside oil well casing according to claim 3, characterized in that, The multi-scale hierarchical feature extraction network adopts a CNN-Transformer dual-branch collaborative architecture, including: The CNN backbone branch is used to extract local detail features of the image. It uses ResNet50 as the base network and introduces the Res2Net multi-scale feature representation module into it. The Transformer auxiliary branch is used to extract global contextual features of the image. It uses the Swing Transformer as the base network and introduces a spatial attention mechanism. The cross-stage feature interaction module is used to realize the interaction and fusion of features from different stages between the CNN backbone branch and the Transformer auxiliary branch; A multi-scale feature pyramid structure is used to fuse feature maps of different scales to generate multi-scale feature representations.

5. The drilling target recognition system for visual feature extraction inside oil well casing according to claim 4, characterized in that, The multi-task collaborative identification and tracking module includes: The target detection unit is used to detect drilling targets in the image, including couplings, perforations, fractures, cuttings, and fluids; Semantic segmentation unit, used to perform pixel-level semantic segmentation of images to distinguish different types of drilling targets and background; The instance segmentation unit is used to distinguish different instances of the same category, enabling the individual identification of targets such as multiple rock cuttings and multiple perforations; The dynamic target tracking unit is used to continuously track dynamic targets using an improved DeepSort algorithm, and to acquire the target's motion trajectory and velocity information.

6. The drilling target identification system for extracting visual features inside oil well casing according to claim 5, characterized in that, The target detection unit adopts an improved YOLOv8 model, introducing a bidirectional feature pyramid network and an adaptive feature fusion mechanism in its neck network, and a decoupled head structure in its head network.

7. The drilling target recognition system for extracting visual features inside oil well casing according to claim 6, characterized in that, The edge-cloud collaborative computing module includes: The downhole edge computing unit is deployed inside the multimodal image acquisition device to perform image preprocessing, lightweight target detection, and data compression and transmission tasks. Cloud computing units, deployed in ground data centers, are used to perform complex tasks such as model training, model updating, deep data analysis, and result visualization. The data transmission unit is used to realize data transmission between the downhole edge computing unit and the cloud computing unit. It adopts an adaptive data compression algorithm and a breakpoint resume mechanism to reduce bandwidth consumption.

8. The drilling target recognition system for extracting visual features inside oil well casing according to claim 7, characterized in that, The depth calibration and positioning module includes: The coupling recognition unit is used to identify oil pipe couplings in the image and record the depth information of the couplings; The visual odometry unit is used to calculate the motion distance and speed of the image acquisition device based on feature matching between consecutive image frames; The depth calibration unit is used to calibrate the depth error of traditional depth measurement systems by combining the depth information from the coupling recognition and the motion information from the visual odometry. The 3D reconstruction unit is used to reconstruct the 3D morphology of the inner wall of the oil well pipe based on structured light images and depth information.

9. The drilling target recognition system for visual feature extraction inside oil well casing according to claim 8, characterized in that, It also includes a drilling target knowledge base, which is connected to the multi-task collaborative identification and tracking module to store the feature information, attribute information and historical identification data of various drilling targets, and supports intelligent classification and attribute analysis of target types.

10. The drilling target recognition system for extracting visual features inside oil well casing according to claim 9, characterized in that, It also includes an early warning module, which is connected to the multi-task collaborative identification and tracking module and the depth calibration and positioning module, and is used to issue an early warning signal and display the target's location and attribute information when an abnormal target is identified.