Endoscope based on dual-camera fusion imaging

By using dual-camera fusion imaging technology, which combines near-focus and far-focus cameras with structured light projection, the problem of existing endoscopes being unable to simultaneously and clearly image both near and far views has been solved, thus improving the comprehensiveness and accuracy of images in aero-engine inspection.

CN122043723APending Publication Date: 2026-05-15BEIJING YICHEN TIMES TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING YICHEN TIMES TECH CO LTD
Filing Date
2026-03-02
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing industrial endoscopes cannot simultaneously and clearly image both near and far views in aero-engine inspection, resulting in insufficient comprehensiveness and accuracy of image acquisition.

Method used

An endoscope based on dual-camera fusion imaging is used, combining a near-focus camera and a far-focus camera. A structured light pattern is projected through a structured light projection module to perform image distortion correction and epipolar alignment, extract feature points for stereo matching, generate a disparity map, and perform weighted fusion.

Benefits of technology

It achieves clear imaging of the interior of aero engines, both close-up and distant, improving the comprehensiveness and accuracy of image acquisition, especially in low-texture environments where it can accurately identify feature points and perform image fusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122043723A_ABST
    Figure CN122043723A_ABST
Patent Text Reader

Abstract

According to the endoscope based on double-camera fusion imaging, through cooperation of the near-focus camera and the far-focus camera, close-shot and long-shot images of a detected target can be obtained respectively. When the texture features of the to-be-detected area are insufficient, the texture feature detection module can detect the situation that the texture feature values are lower than the standard value in time, at the moment, the image processing module controls the structured light projection module to project the preset structured light pattern, and recognizable feature points are added for the low-texture area. After the near-focus camera and the far-focus camera synchronously collect an image with a structured light pattern, the image processing module performs distortion correction and epipolar alignment on a near-focus original image and a far-focus original image and extracts feature points of the structured light pattern for stereo matching, so that a parallax value can be accurately calculated and a distance map can be generated; therefore, a reasonable fusion weight is obtained, accurate fusion of the near-focus image and the far-focus image is realized, and comprehensiveness and accuracy of image acquisition can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of industrial endoscope inspection, specifically to an endoscope based on dual-camera fusion imaging. Background Technology

[0002] Industrial endoscopes are widely used for non-destructive testing of internal structures of equipment, playing a particularly important role in the inspection of aero engines. Critical components inside aero engines, such as turbine blades, combustion chambers, and compressors, require regular inspection. By inserting an endoscope deep into the engine, defects such as wear, cracks, and corrosion of internal parts can be directly observed without disassembling the engine, thereby assessing the engine's health and ensuring flight safety.

[0003] Existing industrial endoscopes, when acquiring images, are limited by the fixed focal length and depth of field of the camera, and can only obtain clear images within a specific distance range. However, the internal space of an aero-engine is complex, and the distance distribution of the targets to be detected is wide. It is necessary to observe both close-range local details such as crack morphology and long-range overall structures such as blade arrangement. When both close-range and long-range targets exist in the detection scene, the camera cannot simultaneously capture clear images of both, resulting in blurry areas and affecting the comprehensiveness and accuracy of image acquisition. Summary of the Invention

[0004] This application provides an endoscope based on dual-camera fusion imaging, which can improve the comprehensiveness and accuracy of image acquisition.

[0005] This application provides an endoscope based on dual-camera fusion imaging, including an endoscope body, wherein an observation cavity is provided at the front end of the endoscope body, characterized in that it further includes: A close-focus camera is installed inside the observation cavity, and the optical axis of the close-focus camera is parallel to the axis of the endoscope body; A telephoto camera is installed inside the observation cavity and is spaced apart from the near-focus camera. The optical axis of the telephoto camera is parallel to the optical axis of the near-focus camera, and the fields of view of the near-focus camera and the telephoto camera overlap. A structured light projection module is installed inside the observation cavity, located between the near-focus camera and the far-focus camera. The structured light projection module includes a light source component and a pattern generation component. The pattern generation component is used to modulate the light beam emitted by the light source component into a preset structured light pattern and project it onto the area to be detected. The texture feature detection module is used to detect the texture feature values ​​of the region to be detected; The image processing module is used to compare the texture feature value with a preset texture standard value, and when the texture feature value is less than the texture standard value, control the structured light projection module to project a structured light pattern onto the area to be detected. When the structured light projection module projects the structured light pattern, the image processing module is also used to control the near-focus camera and the far-focus camera to simultaneously acquire images, obtain the near-focus original image and the far-focus original image, perform distortion correction and epipolar alignment on the near-focus original image and the far-focus original image to obtain the near-focus image and the far-focus image, extract the feature points of the structured light pattern and perform stereo matching to obtain the disparity value, generate a disparity map based on the disparity value, convert the disparity map into a distance map, calculate a weight map based on the distance value of each pixel position in the distance map, and perform weighted fusion of the near-focus image and the far-focus image based on the weight map to obtain a full depth fused image.

[0006] By adopting the above technical solution, the endoscope provided in this application can acquire near-field and far-field images of the target to be detected through the cooperation of a near-focus camera and a far-focus camera. When the texture features of the area to be detected are insufficient, the texture feature detection module can detect in time that the texture feature value is lower than the standard value. At this time, the image processing module controls the structured light projection module to project a preset structured light pattern, adding identifiable feature points to the low-texture area. After the near-focus camera and the far-focus camera synchronously acquire images with structured light patterns, the image processing module performs distortion correction and epipolar alignment on the original near-focus and far-focus images, extracts feature points of the structured light pattern for stereo matching, accurately calculates the disparity value and generates a distance map, and then obtains a reasonable fusion weight to achieve accurate fusion of near-focus and far-focus images, effectively improving the comprehensiveness and accuracy of image acquisition.

[0007] Optionally, the texture feature detection module is used to detect the texture feature values ​​of the region to be detected, specifically including: The texture feature detection module is used to pre-capture a pre-captured image of the area to be detected by the near-focus camera or the far-focus camera, and divide the pre-captured image into multiple sub-regions. The texture feature detection module is also used to calculate the gradient response intensity and corner response density of each sub-region, and determine the texture richness of each sub-region based on the gradient response intensity and the corner response density; The texture feature detection module is also used to count the proportion of sub-regions whose texture richness meets the preset requirements as the texture feature value of the region to be detected.

[0008] By adopting the above technical solution, the texture feature detection module can evaluate texture richness by dividing the pre-acquired image into multiple sub-regions and calculating the gradient response intensity and corner response density of each sub-region, thereby more accurately reflecting the texture feature distribution of the area to be detected. By statistically analyzing the proportion of sub-regions that meet the texture richness requirements as texture feature values, the texture feature situation of the entire area to be detected can be comprehensively evaluated, providing a more reliable basis for the activation of the structured light projection module and further improving the accuracy of image acquisition in low-texture environments.

[0009] Optionally, the texture feature detection module is further configured to calculate the gradient response intensity and corner response density of each sub-region, and determine the texture richness of each sub-region based on the gradient response intensity and the corner response density, specifically including: The texture feature detection module is also used to perform multi-directional gradient calculation on each pixel of each sub-region to obtain the gradient magnitude and gradient direction of each pixel in the sub-region, and to count the proportion of the number of pixels with gradient magnitude greater than a preset gradient threshold to the total number of pixels in the sub-region as the gradient response intensity of the sub-region. The texture feature detection module is also used to perform corner detection on each of the sub-regions, obtain the corner position and corner response value, count the number of corners whose corner response value is greater than a preset corner threshold, and calculate the ratio of the number of corners to the area of ​​the corresponding sub-region as the corner response density. The texture feature detection module is further configured to determine a target sub-region where the gradient response intensity is greater than a first preset threshold and the corner response density of the sub-region is greater than a second preset threshold, and to calculate the texture richness of the target sub-region based on the gradient response intensity and the corner response density.

[0010] By employing the aforementioned technical solution, the texture feature detection module can quantitatively evaluate texture features from two dimensions: gradient features and corner features, through multi-directional gradient calculation and corner detection of sub-regions. By statistically analyzing the proportion of pixels with gradient amplitudes exceeding a threshold as gradient response intensity and calculating the effective corner density as corner response density, the texture complexity of the sub-region can be comprehensively reflected. Based on this, texture richness is calculated only for target sub-regions that simultaneously meet the requirements of gradient response intensity and corner response density, allowing for more accurate identification of regions with significant texture features. This provides more reliable evaluation results for subsequent texture feature value calculations, further improving the accuracy of structured light projection in low-texture environments.

[0011] Optionally, the pattern generating component is a diffractive optical element, a microlens array, or a grating assembly, and the structured light pattern is a grid pattern, a random spot pattern, or a sinusoidal stripe pattern.

[0012] By employing the above technical solutions and using diffractive optical elements, microlens arrays, or grating assemblies as pattern generation components, the light beam emitted by the light source can be modulated into different types of structured light patterns, such as grid patterns, random spot patterns, or sinusoidal stripe patterns. These structured light patterns have obvious spatial and geometric features, and after being projected onto the area to be detected, they can form easily identifiable and matching feature points. This is beneficial for improving the accuracy of stereo matching in low-texture environments, thereby improving the precision of parallax calculation and image fusion.

[0013] Optionally, the light source component is an infrared laser or a visible light LED light source, and the light source component is coupled to the image processing module, which controls the turning the light source component on and off.

[0014] By adopting the above technical solution, using an infrared laser or visible light LED as the light source component, and controlling its on / off state by the image processing module, structured light can be projected on demand based on the results of texture feature detection. The light source is turned off when texture features are sufficient to save energy, and is promptly turned on to project structured light patterns when low-texture areas are detected, ensuring both image acquisition quality and improving the system's energy efficiency.

[0015] Optionally, the feature point density detection module is coupled to the image processing module and is used to detect the density values ​​of identifiable feature points in the near-focus original image and the far-focus original image after the structured light projection module is turned on.

[0016] By adopting the above technical solution, the feature point density detection module can detect the density values ​​of identifiable feature points in the near-focus and far-focus original images after structured light projection in real time, provide timely feedback on the effect of structured light projection, provide a feasibility assessment of feature point matching for the image processing module, and help ensure the accuracy of subsequent stereo matching and image fusion.

[0017] Optionally, the image processing module is further configured to calculate the average brightness value of the near-focus original image and the far-focus original image; When the average brightness value is less than a preset brightness threshold and the density value is less than a preset density threshold, the image processing module adjusts the projection intensity of the structured light projection module according to the continuity evaluation index and the proportion of disparity change regions of the currently generated disparity map until the disparity map meets the preset quality requirements.

[0018] By adopting the above technical solution, the image processing module can promptly detect image quality deficiencies by monitoring the average brightness and feature point density of the original image and combining this with the continuity and abrupt changes in the disparity map. By dynamically adjusting the projection intensity of the structured light projection module until the disparity map meets preset quality requirements, adaptive optimization of the structured light projection intensity can be achieved in complex environments with low illumination and low texture, ensuring the acquisition of high-quality disparity information and thus improving the accuracy of image fusion.

[0019] Optionally, the image processing module adjusts the projection intensity of the structured light projection module based on the continuity evaluation index and the proportion of disparity abrupt change regions of the currently generated disparity map until the disparity map meets preset quality requirements, specifically including: The image processing module is used to perform two-dimensional gradient calculation on the currently generated disparity map to obtain the gradient magnitude of each pixel, and mark the pixels with gradient magnitude greater than a preset gradient threshold as disparity jump points. The image processing module is further configured to determine multiple effective jump regions based on each of the disparity jump points, and calculate the proportion of the total area of ​​the effective jump regions to the total area of ​​the disparity map as the proportion of the disparity jump regions; The image processing module is further configured to divide the disparity map into multiple rectangular evaluation blocks according to spatial location, calculate the standard deviation of the disparity values ​​within each evaluation block and the cumulative sum of the disparity differences between adjacent pixels, mark the evaluation blocks whose standard deviation is less than a preset standard deviation threshold and whose cumulative sum is less than a preset cumulative sum threshold as continuous qualified blocks, and count the proportion of the number of continuous qualified blocks to the total number of evaluation blocks as the continuity evaluation index. When the continuity evaluation index is less than the preset quality threshold and the proportion of the disparity abrupt change region is greater than the preset abrupt change threshold, the image processing module determines the projection intensity increase based on the spatial distribution characteristics of the effective jump region, controls the structured light projection module to increase the projection intensity according to the projection intensity increase, controls the near-focus camera and the far-focus camera to re-acquire the target image after increasing the projection intensity, and the image processing module generates an updated disparity map based on the target image until the updated disparity map meets the preset quality requirements.

[0020] By adopting the above technical solution, the image processing module identifies disparity jump points and determines effective jump regions through two-dimensional gradient calculation. Combined with the continuity analysis of disparity values ​​within the rectangular evaluation block, the quality of the disparity map can be accurately evaluated from two dimensions: disparity abrupt changes and local continuity. When the continuity evaluation index and the proportion of disparity abrupt change regions do not meet the requirements, the structured light projection intensity is dynamically adjusted based on the spatial distribution characteristics of the effective jump regions, and continuously optimized through iteration until a disparity map that meets the quality requirements is obtained. This ensures the reliability of stereo matching results in complex detection environments and provides an accurate disparity information foundation for subsequent image fusion.

[0021] Optionally, the endoscope body also includes a calibration database storage module, which is coupled to the image processing module. The calibration database storage module stores the sharpness distance function of the near-focus camera and the sharpness distance function of the far-focus camera. The image processing module calculates a weight map based on the distance values ​​of each pixel position in the distance map, specifically including: The image processing module queries the corresponding near-focus sharpness value and far-focus sharpness value in the calibration database storage module based on the distance value of each pixel position in the distance map, calculates the sharpness dominance, and generates a weight map. The near-focus sharpness value and far-focus sharpness value are calculated by the sharpness distance function of the near-focus camera and the sharpness distance function of the far-focus camera, respectively.

[0022] By employing the above technical solution, the image processing module pre-stores the sharpness distance functions of the near-focus and far-focus cameras through a calibration database storage module. The image processing module can then directly query the corresponding sharpness value based on the distance value in the distance map and generate an accurate weight map based on the sharpness dominance. This weight calculation method based on pre-calibrated sharpness characteristics fully considers the differences in imaging quality between cameras of different focal lengths at various distances, resulting in a more reasonable allocation of fusion weights and contributing to obtaining a panoramic depth-fusion image with a natural sharpness transition.

[0023] Optionally, the image processing module performs weighted fusion of the near-focus image and the far-focus image based on the weight map to obtain a panoramic depth-of-field fused image, specifically including: The image processing module is further configured to perform a convolution operation on the weight map to obtain a smoothed weight map; and to determine the near-focus weight and far-focus weight of each pixel position based on the smoothed weight map. The image processing module is further configured to extract the near-focus pixel value and far-focus pixel value of each corresponding pixel position in the near-focus image and the far-focus image, multiply the near-focus pixel value by the corresponding near-focus weight to obtain the near-focus weighted value, multiply the far-focus pixel value by the corresponding far-focus weight to obtain the far-focus weighted value, add the near-focus weighted value and the far-focus weighted value to obtain the fused pixel value of the corresponding pixel position; traverse all pixel positions to complete the weighted fusion calculation, and obtain the full depth fused image.

[0024] By employing the above technical solution, the image processing module performs a convolution operation on the weight map to obtain a smoothed weight map, which can eliminate abrupt changes in weight distribution and avoid edge artifacts in the fused image. By calculating the near-focus weight value and the far-focus weight value pixel by pixel and superimposing them to obtain the fused pixel value, a smooth transition fusion of the near-focus image and the far-focus image is achieved, ensuring a natural transition between different distance regions in the panoramic depth-of-field fused image and improving the visual quality of the fused image. Attached Figure Description

[0025] Figure 1 This is a schematic diagram of an endoscope front end structure based on dual-camera fusion imaging provided in an embodiment of this application; Figure 2 This is a system hardware control architecture diagram of an endoscope based on dual-camera fusion imaging provided in an embodiment of this application.

[0026] Explanation of reference numerals in the attached figures: 1. Endoscope body; 2. Observation cavity; 3. Near-focus camera; 4. Far-focus camera; 5. Structured light projection module; 51. Light source assembly; 52. Pattern generation assembly; 6. Texture feature detection module; 7. Image processing module; 8. Feature point density detection module; 9. Calibration database storage module. Detailed Implementation

[0027] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.

[0028] In the description of the embodiments of this application, the words "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design that is described as "for example" or "for instance" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design options. Rather, the use of the words "for example" or "for instance" is intended to present the relevant concepts in a specific manner.

[0029] In the description of the embodiments of this application, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0030] The present application will be further described in detail below with reference to the accompanying drawings.

[0031] This application discloses an endoscope based on dual-camera fusion imaging.

[0032] See Figure 1 , Figure 1 This is a schematic diagram of the front end structure of an endoscope based on dual-camera fusion imaging provided in this application embodiment. It includes an endoscope body 1, an observation cavity 2 provided at the front end of the endoscope body 1, and the endoscope also includes: a near-focus camera 3, a far-focus camera 4, and a structured light projection module 5.

[0033] The near-focus camera 3 is installed on the left side of the observation cavity 2, and the optical axis of the near-focus camera 3 is parallel to the axis of the endoscope body 1. The far-focus camera 4 is installed on the right side of the observation cavity 2, spaced apart from the near-focus camera 3. The optical axis of the far-focus camera 4 is parallel to the optical axis of the near-focus camera 3, and there is an overlapping area in the fields of view of the near-focus camera 3 and the far-focus camera 4.

[0034] The structured light projection module 5 is installed inside the observation cavity 2, located centrally between the near-focus camera 3 and the far-focus camera 4. The structured light projection module 5 includes a light source assembly 51 and a pattern generation assembly 52. ​​The light source assembly 51 uses an infrared laser with a power range of 50mW to 200mW and a wavelength of 850nm. It is coupled to the image processing module 7, which controls the on and off states of the light source assembly 51. The pattern generation assembly 52 uses a diffractive optical element (DOE) to modulate the laser beam emitted by the light source assembly 51 into a preset random spot pattern and project it onto the area to be detected.

[0035] Figure 2 This application provides a system hardware control architecture diagram for an endoscope based on dual-camera fusion imaging, including a near-focus camera 3, a far-focus camera 4, a structured light projection module 5, a texture feature detection module 6, an image processing module 7, a feature point density detection module 8, and a calibration database storage module 9.

[0036] The texture feature detection module 6, using an embedded image processing chip, is coupled to the image processing module 7 and is used to detect the texture feature values ​​of the area to be detected. The image processing module 7, using a high-performance DSP processor, compares the texture feature values ​​with internally preset texture standard values. When the texture feature value is less than the texture standard value, it controls the structured light projection module 5 to project a structured light pattern onto the area to be detected.

[0037] Specifically, when inspectors insert an endoscope deep into the engine for inspection, the texture feature detection module 6 first pre-captures a pre-image of the area to be inspected using the near-focus camera 3, and divides the pre-captured image into multiple sub-regions, for example, into 8×8 sub-regions totaling 64 sub-regions. The texture feature detection module 6 then calculates the gradient response intensity and corner response density of each sub-region.

[0038] For calculating the gradient response intensity, the texture feature detection module 6 performs multi-directional gradient calculations on each pixel in each sub-region, using the Sobel operator to calculate the gradients in the horizontal and vertical directions respectively, obtaining the gradient magnitude and gradient direction of each pixel in the sub-region. The proportion of pixels with gradient magnitudes greater than a preset gradient threshold (e.g., 30) to the total number of pixels in the sub-region is counted as the gradient response intensity of that sub-region.

[0039] For calculating the corner response density, the texture feature detection module 6 performs corner detection on each sub-region using the Harris corner detection algorithm to obtain the corner positions and corner response values. The number of corners with response values ​​greater than a preset threshold (e.g., 0.01) is counted, and the ratio of this number to the area of ​​the corresponding sub-region is calculated as the corner response density. For example, if a sub-region has an area of ​​100×100 pixels and 50 valid corners are detected, the corner response density is 50 / 10000 = 0.005.

[0040] The texture feature detection module 6 identifies target sub-regions where the gradient response intensity is greater than a first preset threshold (e.g., 0.3) and the corner response density is greater than a second preset threshold (e.g., 0.003). For each target sub-region, the texture richness T is calculated according to the gradient response intensity G and the corner response density C, using the weighted formula T = 0.6 × G + 0.4 × C. The proportion of sub-regions with a texture richness greater than a preset requirement (e.g., 0.5) is used as the texture feature value of the region to be detected.

[0041] After acquiring the texture feature value, the image processing module 7 compares it with the texture standard value (e.g., 0.6). If the texture feature value is 0.45, which is less than the texture standard value of 0.6, it indicates that the area to be detected is a low-texture environment such as a smooth metal surface. The image processing module 7 then controls the structured light projection module 5 to turn on, and the light source component 51 emits an infrared laser. After being modulated by the pattern generation component 52, a random spot pattern is projected onto the area to be detected, adding identifiable feature points to the low-texture surface. When the texture features of the detection area are rich, three-dimensional reconstruction is performed directly using the binocular image (without projecting structured light).

[0042] The feature point density detection module 8 is coupled to the image processing module 7 and is used to detect the density values ​​of identifiable feature points in the near-focus and far-focus original images after the structured light projection module 5 is turned on. The feature point density detection module 8 uses the FAST feature point detection algorithm to detect the number of feature points in the image after the structured light is projected in real time.

[0043] When the structured light projection module 5 projects the structured light pattern, the image processing module 7 controls the near-focus camera 3 and the far-focus camera 4 to simultaneously acquire images, obtaining the near-focus original image and the far-focus original image. The image processing module 7 also calculates the average brightness value of the near-focus original image and the far-focus original image. When the average brightness value is less than a preset brightness threshold (e.g., 50, grayscale range 0-255), and the density value detected by the feature point density detection module 8 is less than a preset density threshold (e.g., 10 feature points per 1000 pixels), the image processing module 7 adjusts the projection intensity of the structured light projection module 5 according to the continuity evaluation index of the currently generated disparity map and the proportion of disparity abrupt change regions.

[0044] Specifically, image processing module 7 performs two-dimensional gradient calculation on the currently generated disparity map, using the Sobel operator to obtain the gradient magnitude of each pixel. Pixels with gradient magnitudes greater than a preset gradient threshold (e.g., disparity value changes greater than 5 pixels) are marked as disparity jump points. Based on each disparity jump point, image processing module 7 uses connected component analysis to determine multiple effective jump regions, and calculates the proportion of the total area of ​​the effective jump regions to the total area of ​​the disparity map as the proportion of disparity abrupt change regions.

[0045] Image processing module 7 divides the disparity map into multiple rectangular evaluation blocks according to spatial location, for example, into 256 evaluation blocks of 16×16. For each evaluation block, it calculates the standard deviation of its internal disparity values ​​and the cumulative sum of the disparity differences between adjacent pixels. Evaluation blocks with a standard deviation less than a preset standard deviation threshold (e.g., 2.0) and a cumulative sum less than a preset cumulative sum threshold (e.g., 100) are marked as continuous qualified blocks. The proportion of continuous qualified blocks to the total number of evaluation blocks is used as a continuity evaluation index.

[0046] When the continuity evaluation index is less than a preset quality threshold (e.g., 0.7) and the proportion of disparity abrupt change regions is greater than a preset abrupt change threshold (e.g., 0.15), the image processing module 7 determines the projection intensity increase based on the spatial distribution characteristics of the effective abrupt change regions. For example, if the abrupt change regions are mainly concentrated in the center of the image, the increase is set to 20% of the current projection intensity; if the abrupt change regions are scattered, the increase is set to 30%. The image processing module 7 controls the structured light projection module 5 to increase the projection intensity according to the projection intensity increase, controls the near-focus camera 3 and the far-focus camera 4 to re-acquire the target image, and generates an updated disparity map based on the target image until the continuity evaluation index of the updated disparity map is greater than 0.7 and the proportion of disparity abrupt change regions is less than 0.15, thus meeting the preset quality requirements.

[0047] Image processing module 7 performs distortion correction and epipolar alignment on the near-focus and far-focus original images to obtain the near-focus and far-focus images. Distortion correction uses lens calibration parameters to correct radial and tangential distortion. Epipolar alignment uses stereo calibration parameters to ensure that corresponding points in the near-focus and far-focus images are located on the same horizontal scan line.

[0048] Image processing module 7 extracts feature points from the structured light pattern and performs stereo matching. SIFT or ORB feature point extraction algorithms are used to extract feature points in both the near-focus and far-focus images. A feature descriptor-based matching method, combined with epipolar constraints, is used to perform stereo matching on the feature points in the near-focus and far-focus images, resulting in matching point pairs. The disparity value is calculated based on the difference in the horizontal position of the matching point pairs in the two images. For a configuration where the baseline distance between the near-focus camera 3 and the far-focus camera 4 is 5mm, and the focal lengths are 3mm and 6mm respectively, if the coordinates of a matching point pair are (100, 50) in the near-focus image and (90, 50) in the far-focus image, then the disparity value of that point is 10 pixels.

[0049] Image processing module 7 generates a disparity map based on the disparity values ​​of all matching point pairs. For pixels in the image that are not directly matched, disparity values ​​are filled using interpolation methods. Commonly used interpolation methods include bilinear interpolation or dense disparity estimation based on semi-global matching (SGM). The generated disparity map is a grayscale image, where the grayscale value of each pixel represents the disparity value at that location.

[0050] Image processing module 7 converts the disparity map into a distance map. According to the principle of binocular stereo vision, the distance value D is calculated using the formula: D = (f × B) / d, where f is the focal length, B is the baseline distance, and d is the disparity value. For the near-focus camera 3, the focal length f = 3mm and the baseline distance B = 5mm. If the disparity value d of a certain pixel is 10 pixels (pixel size is 5μm), then the actual distance of that point is D = (3 × 5) / (10 × 0.005) = 300mm. The distance map is generated by iterating through all pixels in the disparity map and calculating the corresponding distance value according to the above formula.

[0051] The endoscope body 1 also houses a calibration database storage module 9, which uses a flash memory chip and is coupled to the image processing module 7. The calibration database storage module 9 stores the sharpness distance function for the near-focus camera 3 and the sharpness distance function for the telephoto camera 4. These functions are obtained through pre-calibration experiments. Standard test images (such as ISO 12233 resolution test charts) are captured at different distances, and the modulation transfer function (MTF) value of the image is calculated as a sharpness index. A functional relationship between sharpness and distance is then fitted to obtain this relationship.

[0052] For the near-focus camera 3, its resolution distance function has a high MTF value (MTF value greater than 0.5) in the range of 50mm to 150mm, and the resolution gradually decreases beyond 150mm. For the far-focus camera 4, its resolution distance function has a high MTF value in the range of 100mm to 500mm. The image processing module 7 queries the calibration database storage module 9 for the corresponding near-focus resolution value S_near and far-focus resolution value S_far based on the distance values ​​of each pixel position in the distance map.

[0053] Image processing module 7 calculates the sharpness dominance and generates a weight map. For each pixel position in the distance map, based on the queried near-focus sharpness value S_near and far-focus sharpness value S_far, the near-focus weight W_near = S_near / (S_near + S_far) and the far-focus weight W_far = S_far / (S_near + S_far) are calculated. For example, for a pixel with a distance value of 100mm, if the queried values ​​are S_near = 0.6 and S_far = 0.7, then W_near = 0.6 / (0.6+0.7) ≈ 0.46 and W_far = 0.7 / (0.6+0.7) ≈ 0.54. All pixel positions in the distance map are traversed to generate two weight maps: a near-focus weight map and a far-focus weight map.

[0054] Image processing module 7 performs weighted fusion of the near-focus and far-focus images based on the weight map to obtain a full-depth fused image. First, image processing module 7 performs convolution operations on the near-focus weight map and the far-focus weight map respectively, and then performs smoothing filtering using a Gaussian kernel (e.g., 5×5, standard deviation σ=1.0) to obtain smoothed near-focus weight maps and far-focus weight maps. The smoothing operation can eliminate abrupt changes in the weight distribution and avoid obvious seams and edge artifacts in the fused image.

[0055] Based on the smoothed weight map, image processing module 7 extracts the near-focus pixel value I_near and the far-focus pixel value I_far for each corresponding pixel position in the near-focus and far-focus images. For color images, the RGB channels are processed separately. The near-focus pixel value is multiplied by its corresponding near-focus weight to obtain the near-focus weighted value V_near = I_near × W_near, and the far-focus pixel value is multiplied by its corresponding far-focus weight to obtain the far-focus weighted value V_far = I_far × W_far. The near-focus weighted value and the far-focus weighted value are added together to obtain the fused pixel value I_fused = V_near + V_far for the corresponding pixel position.

[0056] For example, for a pixel at coordinates (200, 300) in an image, the RGB values ​​at that location are (120, 100, 80) in the near-focus image and (130, 110, 90) in the far-focus image. The smoothed near-focus weight is 0.3, and the far-focus weight is 0.7. Therefore, the fused value for the pixel's R channel is: 120 × 0.3 + 130 × 0.7 = 36 + 91 = 127; the fused value for the G channel is: 100 × 0.3 + 110 × 0.7 = 30 + 77 = 107; and the fused value for the B channel is: 80 × 0.3 + 90 × 0.7 = 24 + 63 = 87. The final fused RGB value for this pixel is (127, 107, 87).

[0057] Image processing module 7 iterates through all pixel positions to complete weighted fusion calculations, resulting in a full-depth fused image. The fused image is mainly contributed by the near-focus image in the near-distance region, and mainly by the far-focus image in the far-distance region. In the intermediate distance transition region, the information from both is smoothly fused, thus achieving clear imaging across the entire depth range from near to far.

[0058] In practical applications, if the texture feature detection module 6 detects that the area to be detected has rich texture (texture feature value of 0.75, which is greater than the standard texture value of 0.6), the image processing module 7 will not turn on the structured light projection module 5, and will directly use the natural texture for stereo matching and image fusion, thereby saving energy and extending the service life of the equipment.

[0059] If the detection environment is completely dark or in extremely low light, the image processing module 7 will automatically turn on the auxiliary illumination LED light source (white LED, adjustable from 10W to 20W) equipped on the endoscope body 1 to provide sufficient ambient lighting. With the auxiliary illumination on, the texture feature detection module 6 can more accurately evaluate the texture features of the area to be detected, and the image processing module 7 will decide whether the structured light projection module 5 still needs to be turned on based on the actual texture conditions.

[0060] For highly reflective metal surfaces (such as the mirror-polished surface of turbine blades), overexposed areas may appear in the image. Image processing module 7 detects overexposed areas (areas with pixel values ​​> 240) during the image preprocessing stage and marks these areas. During stereo matching and image fusion, the reliability of overexposed areas is low, and their weights are appropriately reduced or they are directly excluded from the fusion calculation to avoid erroneous parallax in overexposed areas affecting the overall fusion effect.

[0061] The implementation principle of an endoscope based on dual-camera fusion imaging provided in this application embodiment is as follows: The inspector inserts the endoscope into the internal area to be inspected within an aircraft engine, gas turbine, or other equipment. The texture feature detection module 6 pre-acquires images and divides them into sub-regions, calculates the gradient response intensity and corner response density of each sub-region, comprehensively evaluates texture richness, and statistically analyzes texture feature values. If the texture feature value is less than the texture standard value, it indicates that the area to be inspected is a low-texture environment. The image processing module 7 activates the structured light projection module 5 to project random spot patterns, adding identifiable feature points to the surface. The near-focus camera 3 and the far-focus camera 4 simultaneously acquire images with structured light. The feature point density detection module 8 detects the feature point density in the image. The image processing module 7 assesses whether the structured light projection intensity needs to be adjusted based on the image brightness and feature point density, and further evaluates the results using the continuity evaluation index of the disparity map and the proportion of disparity abrupt change regions. The image processing module 7 performs quality assessment and adaptively adjusts the projection intensity until the disparity map meets the quality requirements, achieving closed-loop optimization. It corrects distortion in the acquired near-focus and far-focus original images to eliminate lens distortion, performs epipolar alignment to ensure corresponding points are on the same horizontal scan line, extracts ORB feature points and performs stereo matching to obtain matching point pairs, calculates disparity values, and generates a dense disparity map using the SGM algorithm. Based on the principle of binocular stereo vision, it converts the disparity map into a distance map. The image processing module 7 then queries the calibration database for the corresponding near-focus and far-focus sharpness values ​​based on the distance values ​​of each pixel in the distance map, calculates near-focus and far-focus weights, generates a weight map, performs Gaussian filtering to smooth the weight map, and finally performs pixel-by-pixel weighted fusion of the near-focus and far-focus images according to the weight map, resulting in a clear fused image across the entire depth range from near to far distances, effectively improving the comprehensiveness and accuracy of image acquisition.

[0062] The foregoing description is merely an exemplary embodiment of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Other embodiments of this disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described herein. The specification and embodiments are to be considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.

Claims

1. An endoscope based on dual-camera fusion imaging, comprising an endoscope body (1), wherein an observation cavity (2) is provided at the front end of the endoscope body (1), characterized in that, Also includes: A close-focus camera (3) is installed inside the observation cavity (2), and the optical axis of the close-focus camera (3) is parallel to the axis of the endoscope body (1); A telephoto camera (4) is installed inside the observation cavity (2) and spaced apart from the near-focus camera (3). The optical axis of the telephoto camera (4) is parallel to the optical axis of the near-focus camera (3), and the fields of view of the near-focus camera (3) and the telephoto camera (4) overlap. The structured light projection module (5) is installed in the observation cavity (2) and located between the near-focus camera (3) and the far-focus camera (4). The structured light projection module (5) includes a light source component (51) and a pattern generation component (52). The pattern generation component (52) is used to modulate the light beam emitted by the light source component (51) into a preset structured light pattern and project it onto the area to be detected. Texture feature detection module (6) is used to detect the texture feature values ​​of the area to be detected; The image processing module (7) is used to compare the texture feature value with the preset texture standard value, and when the texture feature value is less than the texture standard value, control the structured light projection module (5) to project a structured light pattern onto the area to be detected; When the structured light projection module (5) projects the structured light pattern, the image processing module (7) is also used to control the near-focus camera (3) and the far-focus camera (4) to simultaneously acquire images, obtain the near-focus original image and the far-focus original image, perform distortion correction and epipolar alignment on the near-focus original image and the far-focus original image to obtain the near-focus image and the far-focus image, extract the feature points of the structured light pattern and perform stereo matching to obtain the disparity value, generate a disparity map based on the disparity value, convert the disparity map into a distance map, calculate the weight map based on the distance value of each pixel position in the distance map, and perform weighted fusion on the near-focus image and the far-focus image based on the weight map to obtain a full-depth fused image.

2. An endoscope based on dual-camera fusion imaging according to claim 1, characterized in that, The texture feature detection module (6) is used to detect the texture feature values ​​of the region to be detected, specifically including: The texture feature detection module (6) is used to pre-capture the pre-captured image of the area to be detected by the near-focus camera (3) or the far-focus camera (4), and divide the pre-captured image into multiple sub-regions; The texture feature detection module (6) is also used to calculate the gradient response intensity and corner response density of each sub-region, and to determine the texture richness of each sub-region based on the gradient response intensity and the corner response density; The texture feature detection module (6) is also used to count the proportion of sub-regions whose texture richness meets the preset requirements as the texture feature value of the region to be detected.

3. An endoscope based on dual-camera fusion imaging according to claim 2, characterized in that, The texture feature detection module (6) is further configured to calculate the gradient response intensity and corner response density of each sub-region, and determine the texture richness of each sub-region based on the gradient response intensity and the corner response density, specifically including: The texture feature detection module (6) is also used to perform multi-directional gradient calculation on each pixel of each sub-region to obtain the gradient magnitude and gradient direction of each pixel in the sub-region, and to count the proportion of the number of pixels with gradient magnitude greater than a preset gradient threshold to the total number of pixels in the sub-region as the gradient response intensity of the sub-region. The texture feature detection module (6) is also used to perform corner detection on each sub-region, obtain the corner position and corner response value, count the number of corners whose corner response value is greater than a preset corner threshold, and calculate the ratio of the number of corners to the area of ​​the corresponding sub-region as the corner response density. The texture feature detection module (6) is also used to determine the target sub-region where the gradient response intensity is greater than the first preset threshold and the corner response density of the sub-region is greater than the second preset threshold, and to calculate the texture richness of the target sub-region based on the gradient response intensity and the corner response density.

4. An endoscope based on dual-camera fusion imaging according to claim 1, characterized in that, The pattern generation component (52) is a diffractive optical element, a microlens array, or a grating component, and the structured light pattern is a grid pattern, a random spot pattern, or a sinusoidal stripe pattern.

5. An endoscope based on dual-camera fusion imaging according to claim 1, characterized in that, The light source component (51) is an infrared laser or a visible light LED light source. The light source component (51) is coupled to the image processing module (7), and the image processing module (7) controls the opening and closing of the light source component (51).

6. An endoscope based on dual-camera fusion imaging according to claim 1 further includes a feature point density detection module (8), which is coupled to the image processing module (7) and is used to detect the density values ​​of identifiable feature points in the near-focus original image and the far-focus original image after the structured light projection module (5) is turned on.

7. An endoscope based on dual-camera fusion imaging according to claim 6, characterized in that, The image processing module (7) is also used to calculate the average brightness value of the near-focus original image and the far-focus original image; When the average brightness value is less than the preset brightness threshold and the density value is less than the preset density threshold, the image processing module (7) adjusts the projection intensity of the structured light projection module (5) according to the continuity evaluation index and the proportion of disparity change region of the currently generated disparity map until the disparity map meets the preset quality requirements.

8. An endoscope based on dual-camera fusion imaging according to claim 7, characterized in that, The image processing module (7) adjusts the projection intensity of the structured light projection module (5) according to the continuity evaluation index and the proportion of disparity abrupt change regions of the currently generated disparity map until the disparity map meets the preset quality requirements, specifically including: The image processing module (7) is used to perform two-dimensional gradient calculation on the currently generated disparity map to obtain the gradient magnitude of each pixel, and to mark the pixel with a gradient magnitude greater than a preset gradient threshold as a disparity jump point. The image processing module (7) is further configured to determine multiple effective jump regions based on each of the disparity jump points, and calculate the ratio of the total area of ​​the effective jump regions to the total area of ​​the disparity map as the proportion of the disparity jump regions. The image processing module (7) is further configured to divide the disparity map into multiple rectangular evaluation blocks according to spatial location, calculate the standard deviation of the disparity value within each evaluation block and the cumulative sum of the disparity differences between adjacent pixels, mark the evaluation blocks whose standard deviation is less than a preset standard deviation threshold and whose cumulative sum is less than a preset cumulative sum threshold as continuous qualified blocks, and count the proportion of the number of continuous qualified blocks to the total number of evaluation blocks as the continuity evaluation index. When the continuity evaluation index is less than the preset quality threshold and the proportion of the disparity abrupt change region is greater than the preset abrupt change threshold, the image processing module (7) determines the projection intensity increase based on the spatial distribution characteristics of the effective jump region, controls the structured light projection module (5) to increase the projection intensity according to the projection intensity increase, controls the near-focus camera (3) and the far-focus camera (4) to re-acquire the target image after increasing the projection intensity, and the image processing module (7) generates an updated disparity map based on the target image until the updated disparity map meets the preset quality requirements.

9. An endoscope based on dual-camera fusion imaging according to claim 1, characterized in that, The endoscope body (1) is further provided with a calibration database storage module (9), which is coupled to the image processing module (7). The calibration database storage module (9) stores the sharpness distance function of the near-focus camera (3) and the sharpness distance function of the far-focus camera (4). The image processing module (7) calculates a weight map based on the distance values ​​of each pixel position in the distance map, specifically including: The image processing module (7) queries the corresponding near-focus sharpness value and far-focus sharpness value in the calibration database storage module (9) according to the distance value of each pixel position in the distance map, calculates the sharpness dominance and generates a weight map. The near-focus sharpness value and far-focus sharpness value are calculated by the sharpness distance function of the near-focus camera (3) and the sharpness distance function of the far-focus camera (4), respectively.

10. An endoscope based on dual-camera fusion imaging according to claim 1, characterized in that, The image processing module (7) performs weighted fusion of the near-focus image and the far-focus image according to the weight map to obtain a panoramic depth-fusion image, specifically including: The image processing module (7) is further configured to perform a convolution operation on the weight map to obtain a smoothed weight map; and to determine the near-focus weight and far-focus weight of each pixel position based on the smoothed weight map. The image processing module (7) is further configured to extract the near-focus pixel value and far-focus pixel value of each corresponding pixel position in the near-focus image and the far-focus image, multiply the near-focus pixel value by the corresponding near-focus weight to obtain the near-focus weighted value, multiply the far-focus pixel value by the corresponding far-focus weight to obtain the far-focus weighted value, add the near-focus weighted value and the far-focus weighted value to obtain the fused pixel value of the corresponding pixel position; traverse all pixel positions to complete the weighted fusion calculation and obtain the full depth fused image.