Asymmetric binocular industrial endoscope depth of field extension method, device, equipment and medium

By using stereo matching of asymmetric binocular endoscopes and a fusion weight mask guided by physical depth maps, the problems of ghosting caused by parallax and unreliable fusion decisions were solved, achieving the generation of full-depth images without ghosting or distortion, thus improving the accuracy and reliability of industrial inspection.

CN122155967APending Publication Date: 2026-06-05BEIJING YICHEN TIMES TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING YICHEN TIMES TECH CO LTD
Filing Date
2026-03-03
Publication Date
2026-06-05

Smart Images

  • Figure CN122155967A_ABST
    Figure CN122155967A_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an asymmetric binocular industrial endoscope depth of field extension method, device, equipment and medium, relate to the industrial nondestructive testing technical field, the method comprises: by obtaining the target close focus image and target far focus image after correction, then the target close focus image and target far focus image are stereoscopic matching, obtain corresponding first parallax diagram, and the first parallax diagram is converted into corresponding physical depth diagram, and the fusion weight mask film corresponding to the physical depth diagram is generated, then according to the first parallax diagram, the pixel point re-projection is carried out to target far focus image, the view angle of target far focus image is transformed to the view angle of target close focus image, obtain the synthetic far focus image that is aligned with target close focus image in space, finally, according to the fusion weight mask film, target close focus image and synthetic far focus image are weighted fusion, obtain corresponding panoramic depth image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial nondestructive testing technology, and in particular to an asymmetric binocular industrial endoscope depth-of-field extension method, an asymmetric binocular industrial endoscope depth-of-field extension device, an electronic device, and a computer-readable storage medium. Background Technology

[0002] Industrial endoscopes are indispensable non-destructive testing tools in aerospace, energy and power, and precision manufacturing industries. They are mainly used for internal inspection of narrow, enclosed spaces such as engine blades, combustion chambers, and precision pipes. In actual industrial inspection, inspectors often face the contradiction of balancing "microscopic details" and "macroscopic field of view": on the one hand, they need to clearly observe tiny cracks, coating peeling, or pitting on the blade surface at extremely close distances (such as 5mm-20mm), requiring high magnification; on the other hand, they need to observe the overall structure of components, locate foreign objects, or check blade arrangement at greater distances (such as 50mm-200mm), requiring a large depth of field and a wide field of view.

[0003] Due to the limitations of optical physics, the depth of field of a single lens is finite, making it difficult to simultaneously meet the dual requirements of near-field macroscopic detection and far-field macroscopic observation. To address this issue, a hardware architecture of "asymmetric binocular endoscope" has been proposed, which integrates two lenses at the probe end, set as near-focus and far-focus respectively, attempting to obtain a full depth-of-field image by fusing the images from the two lenses.

[0004] However, in practical applications, this hardware architecture, when combined with simple image fusion processing, still faces insurmountable technical problems: One issue is image ghosting caused by parallax. Because the working distance of the endoscope is extremely close, the physical baseline between the binocular lenses introduces significant parallax, meaning that the position of the same object differs considerably between the two images. If fusion is performed directly, it will result in severe ghosting and misalignment in the synthesized image.

[0005] On the other hand, there is the issue of the reliability of the fusion decision-making criteria. Most existing fusion algorithms make decisions based on "image sharpness," assuming that areas with richer textures and sharper edges are clearer and should be preserved. However, in industrial inspection scenarios, the surfaces of workpieces such as engine blades and precision pipes are often smooth and lack texture, with minimal difference between clear and blurred areas, making it difficult for algorithms to effectively distinguish them. Furthermore, metal surfaces often exhibit strong specular reflections, easily misjudging bright, reflective edges as clear textures and preserving them, thus introducing artifacts.

[0006] The aforementioned problems make it difficult for existing asymmetric binocular endoscopes to generate full depth-of-field images that are free of ghosting and distortion, while maintaining macroscopic details and providing a large depth of field. This affects the accuracy and reliability of industrial inspection. Summary of the Invention

[0007] This invention provides a method, apparatus, electronic device, and computer-readable storage medium for extending the depth of field of an asymmetric binocular industrial endoscope, in order to solve or partially solve the problems of image ghosting caused by parallax and unreliable fusion decisions in asymmetric endoscope image fusion.

[0008] This invention discloses a method for extending the depth of field of an asymmetric binocular industrial endoscope, comprising: Acquire the corrected near-focus and far-focus images of the target; Stereo matching is performed on the near-focus image of the target and the far-focus image of the target to obtain the corresponding first disparity map, and the first disparity map is converted into the corresponding physical depth map, and a fusion weight mask corresponding to the physical depth map is generated. The target telephoto image is reprojected pixel by pixel according to the first disparity map, and the viewpoint of the target telephoto image is transformed to the viewpoint of the target near-focus image to obtain a synthetic telephoto image that is spatially aligned with the target near-focus image. The target near-focus image and the synthesized far-focus image are weighted and fused according to the fusion weight mask to obtain the corresponding full-depth image.

[0009] In some feasible implementations, the step of performing stereo matching between the near-focus image of the target and the far-focus image of the target to obtain the corresponding first disparity map includes: The near-focus image of the target and the far-focus image of the target are respectively input into a pre-trained deep learning stereo matching network, and the first feature map of the near-focus image of the target and the second feature map of the far-focus image of the target are extracted by the deep learning stereo matching network. The correlation between the first feature map and the second feature map is calculated to obtain the correlation cost body; The relevant cost body is iteratively optimized in multiple rounds, and a dense disparity map corresponding to each pixel in the near-focus image of the target is output as the first disparity map.

[0010] In some feasible implementations, converting the first disparity map into a corresponding physical depth map includes: The virtual focal length and binocular baseline length used when correcting the near-focus image and the far-focus image of the target are obtained, and the disparity value corresponding to each pixel in the first disparity map is obtained. For each pixel in the first disparity map, the physical depth value corresponding to the pixel is calculated using the virtual focal length, the binocular baseline length, and the disparity value, based on the triangulation principle, to generate a physical depth map corresponding to the first disparity map.

[0011] In some feasible implementations, generating the fusion weight mask corresponding to the physical depth map includes: Obtain the depth value corresponding to each pixel in the physical depth map; Based on the matching result between the depth value and the preset mapping function, the initial weights corresponding to each pixel in the physical depth map are determined; Occlusion detection is performed on the near-focus image and the far-focus image of the target to identify the occlusion area; Based on the occlusion region, target pixels are extracted from the physical depth map, and the initial weights of the target pixels are corrected to generate a fusion weight mask with the same resolution as the near-focus image of the target.

[0012] In some feasible implementations, the preset mapping function consists of a preset range and a weighting function. The preset range includes at least an upper limit for the sharpness range of the near-focus optical path and a lower limit for the sharpness range of the far-focus optical path, wherein the lower limit is greater than the upper limit. The determination of the initial weights corresponding to each pixel in the physical depth map based on the matching result between the depth value and the preset mapping function includes: Pixels with depth values ​​less than the upper limit value are assigned a preset first weight; Pixels with depth values ​​greater than the lower limit are assigned a preset second weight. For pixels whose depth value is greater than or equal to the upper limit value and less than or equal to the lower limit value, the corresponding third weight is obtained by calculating using the depth value, the upper limit value, and the lower limit value.

[0013] In some feasible implementations, the occlusion detection of the near-focus image and the far-focus image of the target, and the identification of the occlusion region, includes: With the geometric center of the image as the origin, rotate the near-focus image and the far-focus image of the target by the target angle respectively to obtain the rotated near-focus image and the rotated far-focus image; The rotated far-focus image is used as the reference image, and the rotated near-focus image is used as the matching image. The images are input into a deep learning stereo matching network for disparity inference to obtain the corresponding target disparity map. Rotate the target angle with the image geometric center of the target disparity map as the origin to obtain a second disparity map that is aligned with the pixels of the target telephoto image; Consistency verification is performed on the pixels representing the same position in the first disparity map and the second disparity map, and the pixels with disparity values ​​greater than a preset threshold are extracted as occlusion points. Based on the occlusion point, the occlusion area between the near-focus image of the target and the far-focus image of the target is determined.

[0014] In some feasible implementations, the step of extracting target pixels from the physical depth map based on the occlusion region and correcting the initial weights of the target pixels to generate a fusion weight mask with the same resolution as the near-focus image of the target includes: The occluded area is matched with the target close-focus image to extract the first target pixel that is occluded in the view of the target close-focus image, and the initial weight of the first target pixel is corrected to 0. The occluded area is matched with the target telephoto image to extract the second target pixel that is occluded in the view of the target telephoto image, and the initial weight of the second target pixel is corrected to 1. For the third target pixel that is not occluded, retain its corresponding initial weight; Based on the corrected weights, a fusion weight mask with the same resolution as the near-focus image of the target is generated.

[0015] In some feasible implementations, the step of reprojecting pixels of the target telephoto image according to the first disparity map, transforming the viewpoint of the target telephoto image to the viewpoint of the target near-focus image, and obtaining a synthetic telephoto image spatially aligned with the target near-focus image, includes: For each pixel in the near-focus image of the target, the sampling coordinates of the pixel in the far-focus image of the target are calculated according to the disparity value of the corresponding position in the first disparity map. Based on the sampling coordinates, the target telephoto image is resampled using a bilinear interpolation algorithm to generate a synthetic telephoto image that is spatially aligned with the target near-focus image.

[0016] In some feasible implementations, the step of weighted fusing of the target near-focus image and the synthesized far-focus image according to the fusion weight mask to obtain the corresponding full-depth image includes: For each color channel, the pixel values ​​at corresponding positions in the target near-focus image and the synthesized far-focus image are linearly weighted and summed according to the weight value corresponding to the fusion weight mask to obtain the fused pixel value; For invalid edge regions in the synthesized telephoto image caused by parallax offset, the pixel values ​​at the corresponding positions in the target near-focus image are used to fill them; Based on the pixel values ​​after fusion of all pixels, a corresponding panoramic depth image is generated.

[0017] Among some feasible implementation methods are: Acquire raw near-focus and raw far-focus images simultaneously acquired by an asymmetric binocular industrial endoscope; Obtain the pre-calibrated intrinsic parameter matrix, distortion coefficients, and binocular extrinsic parameters; Based on the intrinsic parameter matrix and the distortion coefficients, distortion correction is performed on the original near-focus image and the original far-focus image respectively to obtain the distortion-free near-focus image and the distortion-free far-focus image. Based on the binocular extrinsic parameters, epipolar correction is performed on the distortion-free near-focus image and the distortion-free far-focus image to align the left and right image rows, thereby obtaining the corrected target near-focus image and target far-focus image.

[0018] In some feasible implementations, the step of performing distortion correction on the original near-focus image and the original far-focus image respectively based on the intrinsic parameter matrix and the distortion coefficients to obtain the distortion-free near-focus image and the distortion-free far-focus image includes: For the pixel coordinates of each target pixel in the original near-focus image and the original far-focus image, the pixel coordinates are converted into normalized planar coordinates according to the inverse of the intrinsic parameter matrix; Substitute the normalized planar coordinates into the distortion model that includes radial and tangential distortion coefficients to calculate the non-integer sampling coordinates corresponding to the target pixel; Based on the non-integer sampling coordinates, pixel values ​​are sampled in the original near-focus image or the original far-focus image using an interpolation algorithm, and assigned to the pixel coordinates of the target pixel to obtain the distortion-free near-focus image and the distortion-free far-focus image.

[0019] In some feasible implementations, the step of performing epipolar correction on the distortion-free near-focus image and the distortion-free far-focus image based on the binocular extrinsic parameters to align the left and right image rows, thereby obtaining the corrected target near-focus image and target far-focus image, includes: Using the rotation matrix and translation vector in the binocular extrinsic parameters, the left correction rotation matrix and right correction rotation matrix are calculated by the epipolar correction algorithm to rotate the left camera image plane and the right camera image plane in the binocular sensor to coplanar parallelism. Construct a virtual camera projection matrix, which is used to make the epipolar corrected left and right images have the same virtual focal length and the same vertical coordinate center. The left correction rotation matrix, the right correction rotation matrix, and the virtual camera projection matrix are used to perform a reprojection transformation on the distortion-free near-focus image and the distortion-free far-focus image to obtain a row-aligned target near-focus image and a target far-focus image.

[0020] This invention also discloses an asymmetric binocular industrial endoscope depth-of-field extension device, comprising: The image acquisition module is used to acquire the corrected near-focus image and far-focus image of the target; The depth processing module is used to perform stereo matching on the near-focus image of the target and the far-focus image of the target to obtain the corresponding first disparity map, convert the first disparity map into the corresponding physical depth map, and generate a fusion weight mask corresponding to the physical depth map. The reprojection module is used to reproject the target telephoto image pixel by pixel according to the first disparity map, transform the viewpoint of the target telephoto image to the viewpoint of the target near-focus image, and obtain a synthetic telephoto image that is spatially aligned with the target near-focus image. The fusion module is used to perform weighted fusion of the target near-focus image and the synthesized far-focus image according to the fusion weight mask to obtain the corresponding full-depth image.

[0021] In some feasible implementations, the depth processing module is specifically used for: The near-focus image of the target and the far-focus image of the target are respectively input into a pre-trained deep learning stereo matching network, and the first feature map of the near-focus image of the target and the second feature map of the far-focus image of the target are extracted by the deep learning stereo matching network. The correlation between the first feature map and the second feature map is calculated to obtain the correlation cost body; The relevant cost body is iteratively optimized in multiple rounds, and a dense disparity map corresponding to each pixel in the near-focus image of the target is output as the first disparity map.

[0022] In some feasible implementations, the depth processing module is specifically used for: The virtual focal length and binocular baseline length used when correcting the near-focus image and the far-focus image of the target are obtained, and the disparity value corresponding to each pixel in the first disparity map is obtained. For each pixel in the first disparity map, the physical depth value corresponding to the pixel is calculated using the virtual focal length, the binocular baseline length, and the disparity value, based on the triangulation principle, to generate a physical depth map corresponding to the first disparity map.

[0023] In some feasible implementations, the depth processing module is specifically used for: Obtain the depth value corresponding to each pixel in the physical depth map; Based on the matching result between the depth value and the preset mapping function, the initial weights corresponding to each pixel in the physical depth map are determined; Occlusion detection is performed on the near-focus image and the far-focus image of the target to identify the occlusion area; Based on the occlusion region, target pixels are extracted from the physical depth map, and the initial weights of the target pixels are corrected to generate a fusion weight mask with the same resolution as the near-focus image of the target.

[0024] In some feasible implementations, the preset mapping function consists of a preset range and a weighting function. The preset range includes at least an upper limit for the near-focus optical path's sharpness range and a lower limit for the far-focus optical path's sharpness range, with the lower limit being greater than the upper limit. The depth processing module is specifically used for: Pixels with depth values ​​less than the upper limit value are assigned a preset first weight; Pixels with depth values ​​greater than the lower limit are assigned a preset second weight. For pixels whose depth value is greater than or equal to the upper limit value and less than or equal to the lower limit value, the corresponding third weight is obtained by calculating using the depth value, the upper limit value, and the lower limit value.

[0025] In some feasible implementations, the depth processing module is specifically used for: With the geometric center of the image as the origin, rotate the near-focus image and the far-focus image of the target by the target angle respectively to obtain the rotated near-focus image and the rotated far-focus image; The rotated far-focus image is used as the reference image, and the rotated near-focus image is used as the matching image. The images are input into a deep learning stereo matching network for disparity inference to obtain the corresponding target disparity map. Rotate the target angle with the image geometric center of the target disparity map as the origin to obtain a second disparity map that is aligned with the pixels of the target telephoto image; Consistency verification is performed on the pixels representing the same position in the first disparity map and the second disparity map, and the pixels with disparity values ​​greater than a preset threshold are extracted as occlusion points. Based on the occlusion point, the occlusion area between the near-focus image of the target and the far-focus image of the target is determined.

[0026] In some feasible implementations, the depth processing module is specifically used for: The occluded area is matched with the target close-focus image to extract the first target pixel that is occluded in the view of the target close-focus image, and the initial weight of the first target pixel is corrected to 0. The occluded area is matched with the target telephoto image to extract the second target pixel that is occluded in the view of the target telephoto image, and the initial weight of the second target pixel is corrected to 1. For the third target pixel that is not occluded, retain its corresponding initial weight; Based on the corrected weights, a fusion weight mask with the same resolution as the near-focus image of the target is generated.

[0027] In some feasible implementations, the reprojection module is used for: For each pixel in the near-focus image of the target, the sampling coordinates of the pixel in the far-focus image of the target are calculated according to the disparity value of the corresponding position in the first disparity map. Based on the sampling coordinates, the target telephoto image is resampled using a bilinear interpolation algorithm to generate a synthetic telephoto image that is spatially aligned with the target near-focus image.

[0028] In some feasible implementations, the fusion module is used for: For each color channel, the pixel values ​​at corresponding positions in the target near-focus image and the synthesized far-focus image are linearly weighted and summed according to the weight value corresponding to the fusion weight mask to obtain the fused pixel value; For invalid edge regions in the synthesized telephoto image caused by parallax offset, the pixel values ​​at the corresponding positions in the target near-focus image are used to fill them; Based on the pixel values ​​after fusion of all pixels, a corresponding panoramic depth image is generated.

[0029] Among some feasible implementation methods are: The raw image acquisition module is used to acquire the raw near-focus image and the raw far-focus image synchronously acquired by the asymmetric binocular industrial endoscope. The parameter acquisition module is used to acquire the pre-calibrated intrinsic parameter matrix, distortion coefficients, and binocular extrinsic parameters; The distortion correction module is used to perform distortion correction on the original near-focus image and the original far-focus image according to the intrinsic parameter matrix and the distortion coefficient, respectively, to obtain the distortion-free near-focus image and the distortion-free far-focus image; The epipolar correction module is used to perform epipolar correction on the distortion-free near-focus image and the distortion-free far-focus image based on the binocular extrinsic parameters, so as to align the left and right image rows and obtain the corrected target near-focus image and target far-focus image.

[0030] In some feasible implementations, the distortion correction module is specifically used for: For the pixel coordinates of each target pixel in the original near-focus image and the original far-focus image, the pixel coordinates are converted into normalized planar coordinates according to the inverse of the intrinsic parameter matrix; Substitute the normalized planar coordinates into the distortion model that includes radial and tangential distortion coefficients to calculate the non-integer sampling coordinates corresponding to the target pixel; Based on the non-integer sampling coordinates, pixel values ​​are sampled in the original near-focus image or the original far-focus image using an interpolation algorithm, and assigned to the pixel coordinates of the target pixel to obtain the distortion-free near-focus image and the distortion-free far-focus image.

[0031] In some feasible implementations, the epipolar correction module is specifically used for: Using the rotation matrix and translation vector in the binocular extrinsic parameters, the left correction rotation matrix and right correction rotation matrix are calculated by the epipolar correction algorithm to rotate the left camera image plane and the right camera image plane in the binocular sensor to coplanar parallelism. Construct a virtual camera projection matrix, which is used to make the epipolar corrected left and right images have the same virtual focal length and the same vertical coordinate center. The left correction rotation matrix, the right correction rotation matrix, and the virtual camera projection matrix are used to perform a reprojection transformation on the distortion-free near-focus image and the distortion-free far-focus image to obtain a row-aligned target near-focus image and a target far-focus image.

[0032] This invention also discloses an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; The memory is used to store computer programs; When the processor executes a program stored in the memory, it implements the method described in the embodiments of the present invention.

[0033] This invention also discloses a computer-readable storage medium storing instructions that, when executed by one or more processors, cause the processors to perform the methods described in this invention.

[0034] The embodiments of the present invention have the following advantages: In this embodiment of the invention, a corrected near-focus image and a far-focus image of the target are acquired. Then, stereo matching is performed on the near-focus and far-focus images to obtain a corresponding first disparity map. This first disparity map is then converted into a corresponding physical depth map, and a fusion weight mask corresponding to the physical depth map is generated. Next, pixel reprojection is performed on the far-focus image based on the first disparity map, transforming the viewpoint of the far-focus image to that of the near-focus image, resulting in a synthesized far-focus image spatially aligned with the near-focus image. Finally, the near-focus image and the synthesized far-focus image are aligned according to the fusion weight mask. Weighted fusion of near-focus and far-focus images yields a corresponding panoramic depth image. During image fusion, a depth map between near-focus and far-focus images is constructed using binocular parallax. A fusion weight mask based on physical depth is introduced, using real physical distance as the criterion, which significantly improves the reliability of fusion decisions. Simultaneously, pixel reprojection of the far-focus image is performed based on the parallax map before fusion, eliminating pixel-level positional deviations caused by binocular parallax and avoiding ghosting and geometric distortion caused by fusion. This significantly improves the applicability of industrial endoscopes in complex inspection scenarios and the accuracy of inspection results. Attached Figure Description

[0035] Figure 1 This is a flowchart illustrating the steps of an asymmetric binocular industrial endoscope depth-of-field extension method provided in this embodiment of the invention. Figure 2 This is a structural block diagram of an asymmetric binocular industrial endoscope depth-of-field extension device provided in an embodiment of the present invention. Detailed Implementation

[0036] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0037] As examples, in the field of industrial non-destructive testing, industrial endoscopes are crucial tools for inspecting defects inside enclosed spaces such as engine blades, combustion chambers, and precision pipes. In practical inspection tasks, inspectors need to observe microscopic defects such as tiny cracks, corrosion spots, or coating peeling on the workpiece surface at extremely close distances (e.g., 5mm-20mm), and also assess the overall structure of components, locate scattered foreign objects, or inspect the macroscopic arrangement of blades at greater distances (e.g., 50mm-200mm). This contradiction between "microscopic detail" and "macroscopic field of view" poses a significant challenge to traditional single-optical-system endoscopes, as the depth of field of a single lens is very limited due to the laws of optical physics. To address this issue, a hardware architecture of "asymmetric binocular endoscopes" has been proposed, integrating two lenses at the probe end, set as near-focus and far-focus respectively, attempting to obtain a full depth-of-field image by fusing the images from both lenses. However, existing fusion algorithms have significant shortcomings in handling near-range binocular parallax and dealing with low-texture surfaces, which can easily lead to problems such as ghosting, geometric distortion, or fusion decision failure in the fused image.

[0038] To address this issue, this invention proposes a depth-guided asymmetric binocular industrial endoscope depth-of-field expansion method. This method introduces a "range-first, then fusion" approach, utilizing the binocular parallax principle to construct a high-precision depth map. Based on the actual physical distance, it guides pixel-level fusion of near-focus and far-focus images, and uses depth information to correct geometric distortion, ensuring no ghosting during fusion. After the endoscope system acquires near-focus and far-focus images, it first performs stereo matching to obtain a disparity map, then calculates the physical depth map. Next, a fusion weight mask is generated based on the depth map, and the far-focus image is spatially aligned with the near-focus image by performing a perspective transformation based on the disparity map. Finally, the two images are weighted and fused according to the weight mask to output a full depth-of-field image. This method enables the endoscope to simultaneously acquire near-field macroscopic details and far-field macroscopic views in a single probe, with no ghosting or geometric distortion in the fused image, significantly improving the accuracy and reliability of industrial inspection.

[0039] Reference Figure 1 The diagram illustrates a step-by-step flowchart of an asymmetric binocular industrial endoscope depth-of-field extension method provided in an embodiment of the present invention, which may specifically include the following steps: Step 101: Obtain the corrected near-focus image and far-focus image of the target; In this embodiment of the invention, the probe tip of the asymmetric binocular industrial endoscope can integrate two image sensors, namely a near-focus sensor (Sensor A) and a far-focus sensor (Sensor B), with the physical baseline distance between their optical centers being... BFixed (e.g., between 2mm and 6mm, depending on the probe diameter), the optical axes can be designed to be parallel or at a specific angle. Optionally, the two optical axes are designed to be inwardly tilted, so that they intersect at a preset depth-of-field distance. D trans The two points intersect at that angle. θ Satisfy tan( θ / 2)=( B / 2) / D trans This allows for the maximum binocular overlap field of view in the transition area between near and far focus, ensuring a smooth fusion transition.

[0040] In terms of optical design, the close-focus lens is optimized for near-field macro detection, with a preferred clear imaging range of 5mm-30mm. The resolution reaches its peak at an object distance of 15mm, ensuring that fine crack textures can be resolved within this range. The telephoto lens is optimized for distant viewing, with its physical focusing distance set at the hyperfocal distance (approximately 60mm-80mm in this embodiment). According to optical principles, when focusing at this distance, the image of the object extends from half of this distance (approximately 30mm) to infinity, and the image is within the allowable circle of confusion, thereby obtaining the maximum physical depth of field coverage.

[0041] Because endoscope lenses typically have a large field of view, the original acquired images suffer from severe geometric distortion, and there is parallax between binocular images due to the physical baseline. Therefore, before performing depth calculation, the original images need to be corrected, including distortion correction and epipolar correction, to obtain a row-aligned and distortion-free corrected image.

[0042] In some feasible implementations, for the original near-focus image and the original far-focus image synchronously acquired by an asymmetric binocular industrial endoscope, during the correction process of the original near-focus image and the original far-focus image, a pre-calibrated intrinsic parameter matrix, distortion coefficients, and binocular extrinsic parameters can be obtained first. Then, based on the intrinsic parameter matrix and the distortion coefficients, distortion correction is performed on the original near-focus image and the original far-focus image respectively to obtain the distortion-free near-focus image and the distortion-free far-focus image. Then, based on the binocular extrinsic parameters, epipolar correction is performed on the distortion-free near-focus image and the distortion-free far-focus image to align the left and right image rows, thereby obtaining the corrected target near-focus image and target far-focus image.

[0043] The calibration data can be acquired before the system leaves the factory or before use, using a precision calibration board (such as a checkerboard or dot array) and mature calibration algorithms (such as the Zhang Zhengyou calibration method) to perform offline calibration of the binocular probe. Intrinsic parameter matrix. K near and K farEach includes the focal length of a close-up camera and a telephoto camera respectively. f x , f y ) and principal point coordinates ( c x , c y ); distortion coefficient D near , D far Includes radial distortion coefficient ( k1 , k2 , k3 ) and tangential distortion coefficient ( p1 , p2 The binocular extrinsic parameters include a rotation matrix that describes the spatial positional relationship between the telephoto camera and the near-photo camera. R Translation vector T .

[0044] During real-time acquisition, the near-focus and far-focus sensors at the front end of the probe are exposed synchronously to acquire raw images. I raw_near and I raw_far Subsequently, the two original images were corrected.

[0045] In the specific implementation, during the distortion correction process, for the pixel coordinates of each target pixel in the original near-focus image and the original far-focus image, the pixel coordinates are converted into normalized planar coordinates according to the inverse matrix of the intrinsic parameter matrix. Then, the normalized planar coordinates are substituted into the distortion model containing radial distortion coefficients and tangential distortion coefficients to calculate the non-integer sampling coordinates corresponding to the target pixel. Then, based on the non-integer sampling coordinates, pixel values ​​are sampled in the original near-focus image or the original far-focus image through an interpolation algorithm and assigned to the pixel coordinates of the target pixel to obtain the distortion-free near-focus image and the distortion-free far-focus image.

[0046] For example, for each pixel coordinate in the original near-focus image and the original far-focus image ( u , v First, based on the intrinsic parameter matrix K The inverse matrix converts it to normalized planar coordinates. x , y ),Right now[ x , y ,1]^ T = K ^(-1)*[ u , v ,1]^ T Then, the normalized coordinates are substituted into the distortion model formula to calculate the corresponding non-integer coordinates on the original distorted image.u' , v' Optionally, the distortion model may include radial distortion and tangential distortion, and its calculation formula can be expressed as: r ^2= x ^2+ y ^2 x' = x *(1+ k1 * r ^2+ k2 * r ^4+ k3 * r ^6)+2* p1 * x * y + p2 *( r ^2+2* x ^2) y' = y *(1+ k1 * r ^2+ k2 * r ^4+ k3 * r ^6)+ p1 *( r ^2+2* y ^2)+2* p2 * x * y u' = f x * x' + c x v' = f y * y' + c y in, k1 , k2 , k3 The radial distortion coefficient is... p1 , p2 The tangential distortion coefficients are obtained from the non-integer coordinates. u' , v' After that, bilinear interpolation or bicubic interpolation algorithms can be used to interpolate the original near-focus image. I raw Medium sampling ( u' , v'The pixel grayscale value at position () is assigned to the target pixel (). u , v This process generates a distortion-corrected image. The above operation is performed on both the near-focus and far-focus images to obtain the distortion-corrected near-focus image. I undist_near and the distortion-corrected telephoto image I undist_far .

[0047] Furthermore, after obtaining the distortion-corrected near-focus and far-focus images, epipolar correction can be performed. By using the rotation matrix and translation vector in the binocular extrinsic parameters, the epipolar correction algorithm calculates the left and right correction rotation matrices used to rotate the left and right camera image planes in the binocular sensor to coplanar parallelism. At the same time, a virtual camera projection matrix is ​​constructed to ensure that the epipolar-corrected left and right images have the same virtual focal length and the same vertical coordinate center. Finally, the left and right correction rotation matrices and the virtual camera projection matrix are used to perform a reprojection transformation on the distortion-corrected near-focus and far-focus images to obtain the row-aligned target near-focus and target far-focus images.

[0048] In practical implementation, the purpose of epipolar correction is to make the epipolar lines of the left and right images parallel and aligned, thereby reducing the search space for stereo matching from a two-dimensional plane to a one-dimensional horizontal line. This is based on the binocular extrinsic parameters obtained from calibration. R For T, the correction transformation matrix can be calculated using the Bouguet algorithm or the Stereo Rectify algorithm. This algorithm rotates the left and right camera image planes to be coplanar and aligns the epipolar lines horizontally, thus obtaining the left correction rotation matrix. R rect_near And right correction rotation matrix R rect_far These are applied to the left and right images respectively. Simultaneously, a new virtual camera projection matrix is ​​constructed. P new The virtual camera projection matrix P new This ensures that the left and right views have the same focal length and vertical coordinate center, and that the optical axes are parallel. After reprojection transformation using the corrected rotation matrix and virtual projection matrix, the near-focus image of the target... I near any point in ( x , y ) and target telephoto image I far Corresponding points in ( x' , y ) will be in the same row (i.e. y (Same coordinates), parallax exists only in the horizontal direction. d =x - x' .

[0049] Among them, based on binocular extrinsic parameters (rotation matrix) R Translation vector T ) and the original intrinsic parameter matrices of the left and right cameras. K near and K far Constructing a virtual camera projection matrix P new Specifically, firstly, the optical axes of the left and right cameras can be rotated to a parallel state using an epipolar correction algorithm (such as the Bouguet algorithm), and the epipolar lines aligned horizontally. This process decomposes the rotation matrix R into matrices that rotate each camera by half. R l and R r This ensures that the left and right image planes are coplanar. Optionally, the calculation method is as follows:

[0050] in, Rl and Rr These represent the rotation matrices required to rotate the image planes of the left and right cameras to a coplanar direction, respectively. This decomposition ensures that the rotation angles of the corrected left and right views are minimized, thus preserving the effective area of ​​the original image to the maximum extent.

[0051] Next, based on the coplanarity, the epipolar lines can be rotated to the horizontal direction to construct a rotation matrix. Rrect This transforms the pole of the left camera to infinity, and the epipolar line becomes horizontal. This matrix is ​​typically obtained by translating a vector. T The calculation shows that:

[0052] The three row vectors are calculated as follows: e 1: Direction of the poles, take the translation vector T The normalized direction, i.e. e 1= T / ∥ T ∥.

[0053] e 2: With optical axis and e The directions that are all orthogonal are usually taken as the optical axis direction (i.e., the left camera principal optical axis vector (0,0,1)) and... e Cross product of 1 and normalized.

[0054] e 3: with e 1 and e The two directions are both orthogonal, that ise 3= e 1× e 2.

[0055] Multiplying the coplanar rotation matrix by the epipolar alignment matrix yields the final left correction rotation matrix. R rect_near And right correction rotation matrix R rect_far : R rect_near = R rect R l , R rect_far = R rect R r To ensure that the corrected left and right images have the same focal length and principal point coordinates, a unified virtual camera intrinsic parameter matrix needs to be constructed. Knew Optionally, the average of the original focal lengths of the left and right cameras can be used as the virtual focal length, and the x-coordinate of the principal point coordinates can be kept at its original value, while the y-coordinate can be uniformly set to half the image height (or the average of the y-coordinates of the left and right principal points) to ensure that the epipolar lines are strictly horizontally aligned. Knew It can be:

[0056] in, It can be a virtual focal length. It can be the center of the unified vertical coordinate.

[0057] In the specific implementation, the virtual camera projection matrix P new It can be a 3x4 matrix used to directly project 3D points in the world coordinate system onto the corrected image pixel coordinate system. Its general form is as follows: Pnew = Knew [ Rrect | tnew Among them, among them, tnew This is a translation vector, which can be zeroed out to indicate that the virtual camera coordinate system coincides with the world coordinate system. In practical applications of stereo calibration, it is more common to directly construct a remapping lookup table using the calibration rotation matrix and the virtual intrinsic parameter matrix, rather than explicitly storing the complete coordinate system. Pnew In mathematical terms, the role of the projection matrix can be understood as follows:

[0058] in, It can be the inverse of the original intrinsic parameter matrix, used to transform the pixel coordinates on the image to normalized coordinates in the original camera coordinate system (i.e., equivalent to the ray direction in the camera coordinate system). It can be the homogeneous coordinates of a pixel on the original image, and can be represented as [ u , v ,1] T This can be a point on the image before correction.

[0059] This mapping relationship is ultimately merged into a Union Lookup Table (LUT) for a one-time remapping of the original image. Through the above construction process, the left and right views obtain the same virtual focal length and vertical coordinate center, achieving strict epipolar horizontal alignment and providing ideal input data for subsequent stereo matching.

[0060] Optionally, in practical engineering implementations, to meet the real-time requirements of industrial inspection, distortion correction and epipolar correction can be combined into a single joint remapping process instead of being performed in separate steps. Specifically, an XY lookup table (LUT) containing the combined effects of distortion correction and epipolar correction can be pre-calculated. During real-time acquisition, this LUT is directly used to remapping the original near-focus image. I raw_near and the original telephoto image I raw_far Performing a single remapping operation allows you to directly output a row-aligned and distortion-free corrected image. I Near and I Far This significantly reduces computational latency.

[0061] Through the above correction process, a spatially aligned, geometrically distortion-free near-focus image of the target can be obtained. I Near and target telephoto image I Far This provides standardized input data for subsequent high-precision depth calculations and image fusion.

[0062] Step 102: Perform stereo matching on the near-focus image of the target and the far-focus image of the target to obtain the corresponding first disparity map, convert the first disparity map into the corresponding physical depth map, and generate a fusion weight mask corresponding to the physical depth map. In this embodiment of the invention, stereo matching aims to establish a pixel-level correspondence between a near-focus image and a far-focus image of the target, obtain a corresponding first disparity map, and calculate high-precision physical distance information, i.e., a physical depth map, based on the first disparity map, and further generate a fusion weight mask corresponding to the physical depth map. The fusion weight mask can be used to guide the contribution ratios of the near-focus image and the far-focus image during the image fusion process.

[0063] In some feasible implementations, for the stereo matching process, the target near-focus image and the target far-focus image can be respectively input into a pre-trained deep learning stereo matching network. The deep learning stereo matching network extracts a first feature map of the target near-focus image and a second feature map of the target far-focus image. Then, the correlation between the first feature map and the second feature map is calculated to obtain a correlation cost body. Then, the correlation cost body is iteratively optimized in multiple rounds to output a dense disparity map corresponding to each pixel in the target near-focus image, which serves as the first disparity map.

[0064] In practical implementation, deep learning stereo matching networks can adopt architectures such as CREStereo or RAFT-Stereo. A deep learning stereo matching network can at least include a feature extraction network, a relevance cost body construction module, and an iterative optimization module. The feature extraction network can be a convolutional neural network with shared weights, composed of multiple stacked residual blocks and dilated convolutional layers. On the one hand, the feature extraction network captures low-level edge and texture information as well as high-level semantic information of the image through convolution operations; on the other hand, it expands the receptive field through dilated convolutions, enabling the network to utilize a wider range of contextual information. This allows it to extract discriminative features even in textureless or weakly textured areas commonly found in industrial endoscopes (such as smooth metal surfaces), providing a reliable basis for subsequent matching. The feature extraction network typically includes a downsampling stage, reducing the feature map size to 1 / 4 or 1 / 8 of the original image to reduce subsequent computation while preserving sufficient spatial resolution.

[0065] For the correlation cost body construction module, a 4D all-pair correlation body or multi-scale pyramid cost body is constructed based on the left and right feature maps extracted by the convolutional neural network. Specifically, for each pixel in the left feature map, the inner product of its feature vectors with all candidate pixels is calculated along the epipolar direction (horizontal direction) in the right feature map to form a cost body. The dimension of the cost body can be H×W×W×1 (H and W are the height and width of the feature map). This cost body quantifies the matching similarity between pixels in the left and right feature maps. Optionally, in order to handle large disparity ranges and reduce computational burden, a pyramid structure can be used to construct cost bodies at different resolutions, etc., which is not limited in this invention.

[0066] Furthermore, the iterative optimization module can be a cascaded refinement network based on a recurrent neural network. It can take the relevance cost volume and the current disparity estimate as input, and gradually correct the disparity value through multiple rounds of iterative optimization. In each iteration, the recurrent neural network uses the matching information provided by the cost volume and image context features to generate a disparity update, which is then added to the current disparity to obtain a more accurate disparity. Optionally, the iterative optimization module can also integrate a context encoder to extract structural information of the image (such as edges, occlusion boundaries, etc.), guiding the update process to focus more on object boundaries and detailed regions. After several iterations (usually 4-12), the network outputs a dense, sub-pixel precision disparity map. Thus, through multiple rounds of iterative optimization, ambiguities and noise in the initial matching can be gradually eliminated, resulting in a smooth and accurate disparity map.

[0067] Based on a deep learning-based stereo matching network, the stereo matching process first involves correcting the target near-focus image. I Near and target telephoto image I Far The input is a convolutional neural network (Feature Extractor) with shared weights, which extracts multi-scale image feature maps to obtain the corresponding first and second feature maps, which can be denoted as the left feature map. F near and right feature map F far Its size can be 1 / 4 or 1 / 8 of the original image.

[0068] Next, the correlation cost volume construction module can calculate the correlation between the extracted left and right feature maps and construct a 4D All-pairs Correlation Volume or pyramid cost volume to measure the similarity between pixels in the left image and candidate matching pixels in the right image. Specifically, for each pixel in the left image feature map, its dot product with the feature vector of the candidate pixel can be calculated along the horizontal epipolar direction in the right image feature map (after epipolar correction, the search range is limited to the horizontal line), forming a cost volume. The size of this cost volume can be H×W×W×1 (assuming the feature map height is H and the width is W), representing the matching cost of each pixel in the left image with all horizontally positioned pixels in the right image.

[0069] Finally, a recurrent neural network is used to perform multiple iterative updates based on the relevance cost volume. Specifically, the recurrent neural network learns geometric context information to continuously correct the disparity prediction values, ultimately outputting a dense disparity map. d .in, d ( u , v )express I NearThe median coordinate is ( u , v The pixels in I Far The horizontal position offset of the corresponding point in the map. The resolution of this disparity map is... I Near Consistent, the value of each pixel represents the disparity of that point, in pixels.

[0070] After obtaining the disparity map, it can be converted into a depth map with actual physical meaning. Even though the system contains two physical lenses and each has independent initial intrinsic parameters, after epipolar correction in the aforementioned process, the binocular system has been mathematically transformed into a standard parallel binocular model. Therefore, the depth conversion process can be based on the corrected virtual stereo parameters.

[0071] In some feasible implementations, during the depth conversion process, the virtual focal length and binocular baseline length used when correcting the near-focus and far-focus images of the target can be obtained first, and the disparity value corresponding to each pixel in the first disparity map can be obtained. Then, for each pixel in the first disparity map, the physical depth value corresponding to the pixel is calculated using the virtual focal length, binocular baseline length, and disparity value through the triangulation principle to generate a physical depth map corresponding to the first disparity map.

[0072] In a practical implementation, the virtual camera projection matrix generated in the aforementioned process can be obtained. P new Virtual focal length in f rect (Unit: pixels), and the translation vector from the initial extrinsic parameters. T Calculated physical baseline length B (Unit: millimeters). At this time, f rect The left and right views are consistent. B This represents the horizontal distance between the two optical centers. Based on the standard parallel binocular vision model, for each non-zero effective pixel in the disparity map ( u , v Perform the following calculations to determine the vertical depth of the point from the camera's optical center. Z : D ( u , v )= Z =( f rect * B ) / d ( u , v ) in, D ( u ,v This represents the physical depth value of that pixel, which can be expressed in millimeters. By traversing all pixels in the disparity map, a physical depth map with the same resolution as the near-focus image of the target can be generated. D .

[0073] After obtaining the physical depth map, in this embodiment of the invention, a fusion weight mask can be further generated based on the physical depth map to know the pixel-level fusion between the subsequent near-focus image and the far-focus image. Thus, in the process of image fusion, based on the real physical distance, the problem of fusion decision-making efficiency in low-texture, high-reflectivity scenes can be effectively avoided.

[0074] In some feasible implementations, the process of generating the fusion weight mask can be as follows: first, obtain the depth value corresponding to each pixel in the physical depth map; then, based on the matching result between the depth value and the preset mapping function, determine the initial weight corresponding to each pixel in the physical depth map; then, perform occlusion detection on the target near-focus image and the target far-focus image to identify the occlusion area; then, extract the target pixel from the physical depth map based on the occlusion area, and correct the initial weight of the target pixel to generate a fusion weight mask with the same resolution as the target near-focus image.

[0075] The preset mapping function consists of a preset range and a weight value function. The preset range includes at least the upper limit of the sharp range of the near-focus optical path and the lower limit of the sharp range of the far-focus optical path. If the lower limit is greater than the upper limit, then in the process of determining the initial weight, pixels with a depth value less than the upper limit can be set as the preset first weight, pixels with a depth value greater than the lower limit can be set as the preset second weight, and for pixels with a depth value greater than or equal to the upper limit and less than or equal to the lower limit, the depth value, the upper limit, and the lower limit are used to calculate and obtain the corresponding third weight.

[0076] As examples, let's set an upper threshold for the sharpness range of the near-focus optical path. d 1 represents the lower threshold of the sharpness range of a telephoto optical path at 25mm. d 2 is 35mm. For the physical depth map... D Each pixel in the array has a depth value. d Then the initial weights W init (Near-focus image weights) can be determined according to the following logic: like d < d 1 (Near-field high weight area): This area is determined to be sharp only in close-up lenses. W init ≈1; like d > d2 (Far-field high-weight area): This area is determined to be clear only in telephoto lenses, and is set... W init ≈0; like d 1≤ d ≤ d 2 (Transition Zone): To avoid obvious stitching artifacts in the merged images, linear interpolation can be used to generate gradual weights, for example... W init =( d 2- d ) / ( d 2- d 1).

[0077] Through the above mapping, each pixel obtains an initial weight between 0 and 1, representing the contribution ratio of the near-focus image during fusion, and the contribution ratio of the far-focus image is 1- W init .

[0078] Furthermore, due to the difference in viewing angles between binocular cameras, field-of-view occlusion is inevitable; that is, some areas are visible from one viewpoint but obscured from another. Directly fusing the occluded areas would result in severe artifacts. To address this, in this embodiment of the invention, occlusion detection is performed on the near-focus image and the far-focus image of the target to identify the occluded areas, thereby correcting the weights of the occluded areas.

[0079] In some feasible implementations, during the process of identifying occlusion regions, the target near-focus image and the target far-focus image are first rotated by a target angle with the geometric center of the image as the origin, respectively, to obtain the rotated near-focus image and the rotated far-focus image. The rotated far-focus image is used as the reference image, and the rotated near-focus image is used as the matching image. These are then input into a deep learning stereo matching network for disparity inference to obtain the corresponding target disparity map. Next, the target angle is rotated with the geometric center of the target disparity map as the origin to obtain a second disparity map that is aligned with the pixels of the target far-focus image. Then, the consistency of pixels representing the same position in the first disparity map and the second disparity map is checked, and pixels with disparity values ​​greater than a preset threshold are extracted as occlusion points. Finally, based on the occlusion points, the occlusion region between the target near-focus image and the target far-focus image is determined.

[0080] It should be noted that in the process of identifying the occluded area described above, a standard stereo matching network can be designed only to search for the positive disparity (i.e., the disparity value is positive) from the left image to the right image. If the left and right image inputs are directly swapped to calculate the reverse disparity, the disparity will become negative, exceeding the network's search space. Therefore, in this embodiment of the invention, a "rotation-inference-inverse inference" strategy is used to obtain the right-view disparity.

[0081] In the specific implementation, the corrected image is first taken as the origin, with the geometric center of the image as the origin. I Near and I Far Rotate each angle by 180 degrees (i.e., the target angle is 180 degrees) to obtain I Near_rot and I Far_rot This operation, while preserving the correspondence of image content, flips the original "left-right" geometric relationship to "right-left," thereby transforming the reverse disparity search problem into a forward disparity search problem acceptable to the network. Next, the rotated image is input into the same deep learning stereo matching network as described in the previous embodiment, but the input order needs to be reversed. I Far_rot As a reference image (front view) I Near_rot As the matching image (target view), the deep learning stereo matching network outputs a disparity map in a rotated coordinate system. d rot Then, d rot Rotating the image 180 degrees around its geometric center will yield a telephoto image of the target. I Far Pixel-aligned second disparity map d Right (i.e., the parallax diagram on the right) This second parallax diagram can represent the correspondence of each pixel in the near-focus image when the far-focus image is used as a reference.

[0082] Obtain the disparity map in the left image. d Left (i.e., the first disparity map obtained from the aforementioned process) and the right-hand disparity map d Right Next, the disparity map on the left can be analyzed. d Left Parallax diagram (right) d Right Perform a consistency check. Specifically, based on the epipolar geometry principle, the disparity of corresponding points should be equal in magnitude. The check criteria can be as follows: | d Left ( x , y )- d Right ( x - d Left ( x , y ), y )|> δ in, δThe preset threshold is preferably 1.0 pixel. If the absolute value of the difference at a certain pixel is greater than... δ If a pixel is found to be occluded, then that pixel is considered an occluded point. The set of all occluded points constitutes the occlusion region.

[0083] Furthermore, based on the identified occlusion region, the occlusion region can be matched with the target near-focus image to extract the first target pixel that is occluded in the view of the target near-focus image. The initial weight of the first target pixel is corrected to 0. The occlusion region can also be matched with the target far-focus image to extract the second target pixel that is occluded in the view of the target far-focus image. The initial weight of the second target pixel is corrected to 1. Meanwhile, for the unoccluded third target pixel, the corresponding initial weight is retained. Finally, based on the corrected weights, a fusion weight mask with the same resolution as the target near-focus image is generated.

[0084] In practical implementation, after identifying occlusion points through consistency checks, it's possible to further determine from which viewpoint each occluded point is occluded. For example, if a pixel is invisible from a telephoto viewpoint (i.e., occluded in the telephoto image), then during fusion, it should fully utilize information from the near-focus image, and a near-focus weight can be forcibly set for that pixel. W =1; conversely, if a pixel is occluded at close-up angles, its weight can be forcibly set. W =0. For non-occluded pixels, the initial depth-based weights calculated in the previous process are retained. After the above correction, a single-channel floating-point weight map with the same resolution as the input image is finally generated. W This is used to guide subsequent pixel-level fusion. The value of each pixel in this weight map is between 0 and 1, representing the contribution ratio of the near-focus image.

[0085] Step 103: Reproject the target telephoto image pixel by pixel according to the first disparity map, transform the viewpoint of the target telephoto image to the viewpoint of the target near-focus image, and obtain a composite telephoto image that is spatially aligned with the target near-focus image. In this embodiment of the invention, in order to eliminate the parallax caused by the physical baseline of the binocular camera, the target telephoto image can be reprojected pixel by pixel according to the first parallax map, and the viewing angle of the target telephoto image can be transformed to the viewing angle of the target near-focus image to obtain a synthetic telephoto image that is spatially aligned with the target near-focus image, thereby achieving pixel-level spatial alignment.

[0086] In some feasible implementations, for the pixel reprojection process, for the pixel coordinates corresponding to each pixel in the target near-focus image, the sampling coordinates of the pixel in the target far-focus image are calculated according to the disparity value of the corresponding position in the first disparity map. Then, based on the sampling coordinates, the target far-focus image is resampled by a bilinear interpolation algorithm to generate a synthetic far-focus image that is spatially aligned with the target near-focus image.

[0087] In the specific implementation, since epipolar correction has been completed in the aforementioned embodiments, the left and right images are aligned in the vertical direction (i.e., (Coordinates are consistent), parallax exists only in the horizontal direction. To avoid the hole problem caused by forward projection, this embodiment adopts a backward mapping strategy. Specifically, for each pixel coordinate in the reference viewpoint (i.e., near-focus viewpoint) image... Using the first disparity map Calculate its position in the source image (i.e., the telephoto image). The corresponding sampling coordinates in ) :

[0088] Among them, disparity value Since they are floating-point numbers, the calculated source coordinates It is usually a non-integer.

[0089] Furthermore, using the calculated non-integer coordinates For target telephoto images Sampling is performed. Optionally, to ensure image quality and prevent jagged edges, a bilinear interpolation algorithm can be used: set up The integer part is The decimal part is Then the sampled value is passed through exist and The pixel values ​​at that location are calculated using weighted averages.

[0090] By iterating through all the pixels, a completely new composite telephoto image can be generated. ,thereby Although the image content comes from a telephoto lens (possessing clarity for distant scenes), its geometry, perspective, and object positions have been completely distorted to the perspective of a close-up lens, meaning that... and Pixel-level overlap was achieved in space, laying the foundation for geometric alignment for subsequent weighted fusion.

[0091] Step 104: Perform weighted fusion of the target near-focus image and the synthesized far-focus image according to the fusion weight mask to obtain the corresponding full-depth image.

[0092] In this embodiment of the invention, the fusion weight mask generated by the aforementioned process can be used to weightedly synthesize the high-frequency details of the near-focus image with the reprojection-corrected far-focus image, outputting the final panoramic depth image. Thus, during the image fusion process, a depth map between the near-focus and far-focus images is constructed using binocular parallax. By introducing a fusion weight mask based on physical depth and using the actual physical distance as the criterion, the reliability of the fusion decision is significantly improved. At the same time, pixel reprojection of the far-focus image is performed based on the parallax map before fusion, eliminating pixel-level positional deviations caused by binocular parallax and avoiding ghosting and geometric distortion caused by fusion. This significantly improves the applicability of industrial endoscopes in complex inspection scenarios and the accuracy of inspection results.

[0093] In some feasible implementations, for each color channel, the pixel values ​​at corresponding positions in the target near-focus image and the synthesized far-focus image are linearly weighted and summed according to the weight value corresponding to the fusion weight mask to obtain the fused pixel value. At the same time, for the invalid edge regions in the synthesized far-focus image caused by parallax offset, the pixel values ​​at corresponding positions in the target near-focus image are used to fill them. Finally, based on the fused pixel values ​​of all pixels, the corresponding full-depth image is generated.

[0094] In the specific implementation, for each channel (R, G, B) of the image, a linear weighted fusion operation is performed based on the generated floating-point weight map W:

[0095] The logic of this operation can be summarized as follows: in the near-field region (W≈1), the output image is directly retained. I Near The original pixels are used to ensure that even tiny cracks are clearly visible; in the far-field region (W≈0), the output image uses... The pixels, due to It has undergone geometric correction and can seamlessly fill in distant background information without producing ghosting; in occluded or transitional areas, the smooth transition or occlusion processing of W eliminates splicing gaps.

[0096] Due to parallax shift, the reprojected image Blank areas (invalid pixels) with no information may appear on the left edge. During the fusion stage, this invention employs a reference image supplementation strategy for these boundary areas, that is, directly calling the near-focus image. I NearThe corresponding pixel data is filled in. This strategy ensures the integrity of the output image's field of view (keeping it consistent with the field of view of a close-up camera) while avoiding image cropping-induced loss of image size or artifacts caused by interpolation.

[0097] Output the synthesized image This image combines the high-magnification macro detail of the near-focus optical path (5mm-30mm) with the large depth of field macro view of the telephoto optical path (30mm-infinity), and eliminates the misalignment and ghosting commonly found in traditional binocular fusion, achieving full-view depth-of-field non-destructive detection in a single detection.

[0098] It should be noted that the embodiments of the present invention include, but are not limited to, the examples described above. It is understood that those skilled in the art can make further settings according to actual needs under the guidance of the ideas in the embodiments of the present invention, and the present invention does not limit such settings.

[0099] In this embodiment of the invention, a corrected near-focus image and a far-focus image of the target are acquired. Then, stereo matching is performed on the near-focus and far-focus images to obtain a corresponding first disparity map. This first disparity map is then converted into a corresponding physical depth map, and a fusion weight mask corresponding to the physical depth map is generated. Next, pixel reprojection is performed on the far-focus image based on the first disparity map, transforming the viewpoint of the far-focus image to that of the near-focus image, resulting in a synthesized far-focus image spatially aligned with the near-focus image. Finally, the near-focus image and the synthesized far-focus image are aligned according to the fusion weight mask. Weighted fusion of near-focus and far-focus images yields a corresponding panoramic depth image. During image fusion, a depth map between near-focus and far-focus images is constructed using binocular parallax. A fusion weight mask based on physical depth is introduced, using real physical distance as the criterion, which significantly improves the reliability of fusion decisions. Simultaneously, pixel reprojection of the far-focus image is performed based on the parallax map before fusion, eliminating pixel-level positional deviations caused by binocular parallax and avoiding ghosting and geometric distortion caused by fusion. This significantly improves the applicability of industrial endoscopes in complex inspection scenarios and the accuracy of inspection results.

[0100] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.

[0101] Reference Figure 2The diagram illustrates a structural block diagram of an asymmetric binocular industrial endoscope depth-of-field extension device provided in an embodiment of the present invention, which may specifically include the following modules: Image acquisition module 201 is used to acquire the corrected near-focus image and far-focus image of the target; The depth processing module 202 is used to perform stereo matching on the near-focus image of the target and the far-focus image of the target to obtain the corresponding first disparity map, convert the first disparity map into the corresponding physical depth map, and generate a fusion weight mask corresponding to the physical depth map. The reprojection module 203 is used to reproject the target telephoto image pixel by pixel according to the first disparity map, transform the viewing angle of the target telephoto image to the viewing angle of the target near-focus image, and obtain a composite telephoto image that is spatially aligned with the target near-focus image. The fusion module 204 is used to perform weighted fusion of the target near-focus image and the synthesized far-focus image according to the fusion weight mask to obtain the corresponding panoramic depth image.

[0102] In some feasible implementations, the depth processing module 202 is specifically used for: The near-focus image of the target and the far-focus image of the target are respectively input into a pre-trained deep learning stereo matching network, and the first feature map of the near-focus image of the target and the second feature map of the far-focus image of the target are extracted by the deep learning stereo matching network. The correlation between the first feature map and the second feature map is calculated to obtain the correlation cost body; The relevant cost body is iteratively optimized in multiple rounds, and a dense disparity map corresponding to each pixel in the near-focus image of the target is output as the first disparity map.

[0103] In some feasible implementations, the depth processing module 202 is specifically used for: The virtual focal length and binocular baseline length used when correcting the near-focus image and the far-focus image of the target are obtained, and the disparity value corresponding to each pixel in the first disparity map is obtained. For each pixel in the first disparity map, the physical depth value corresponding to the pixel is calculated using the virtual focal length, the binocular baseline length, and the disparity value, based on the triangulation principle, to generate a physical depth map corresponding to the first disparity map.

[0104] In some feasible implementations, the depth processing module 202 is specifically used for: Obtain the depth value corresponding to each pixel in the physical depth map; Based on the matching result between the depth value and the preset mapping function, the initial weights corresponding to each pixel in the physical depth map are determined; Occlusion detection is performed on the near-focus image and the far-focus image of the target to identify the occlusion area; Based on the occlusion region, target pixels are extracted from the physical depth map, and the initial weights of the target pixels are corrected to generate a fusion weight mask with the same resolution as the near-focus image of the target.

[0105] In some feasible implementations, the preset mapping function consists of a preset range and a weighting function. The preset range includes at least an upper limit for the near-focus optical path sharpness range and a lower limit for the far-focus optical path sharpness range, wherein the lower limit is greater than the upper limit. The depth processing module 202 is specifically used for: Pixels with depth values ​​less than the upper limit value are assigned a preset first weight; Pixels with depth values ​​greater than the lower limit are assigned a preset second weight. For pixels whose depth value is greater than or equal to the upper limit value and less than or equal to the lower limit value, the corresponding third weight is obtained by calculating using the depth value, the upper limit value, and the lower limit value.

[0106] In some feasible implementations, the depth processing module 202 is specifically used for: With the geometric center of the image as the origin, rotate the near-focus image and the far-focus image of the target by the target angle respectively to obtain the rotated near-focus image and the rotated far-focus image; The rotated far-focus image is used as the reference image, and the rotated near-focus image is used as the matching image. The images are input into a deep learning stereo matching network for disparity inference to obtain the corresponding target disparity map. Rotate the target angle with the image geometric center of the target disparity map as the origin to obtain a second disparity map that is aligned with the pixels of the target telephoto image; Consistency verification is performed on the pixels representing the same position in the first disparity map and the second disparity map, and the pixels with disparity values ​​greater than a preset threshold are extracted as occlusion points. Based on the occlusion point, the occlusion area between the near-focus image of the target and the far-focus image of the target is determined.

[0107] In some feasible implementations, the depth processing module 202 is specifically used for: The occluded area is matched with the target close-focus image to extract the first target pixel that is occluded in the view of the target close-focus image, and the initial weight of the first target pixel is corrected to 0. The occluded area is matched with the target telephoto image to extract the second target pixel that is occluded in the view of the target telephoto image, and the initial weight of the second target pixel is corrected to 1. For the third target pixel that is not occluded, retain its corresponding initial weight; Based on the corrected weights, a fusion weight mask with the same resolution as the near-focus image of the target is generated.

[0108] In some feasible implementations, the reprojection module 203 is used for: For each pixel in the near-focus image of the target, the sampling coordinates of the pixel in the far-focus image of the target are calculated according to the disparity value of the corresponding position in the first disparity map. Based on the sampling coordinates, the target telephoto image is resampled using a bilinear interpolation algorithm to generate a synthetic telephoto image that is spatially aligned with the target near-focus image.

[0109] In some feasible implementations, the fusion module 204 is used for: For each color channel, the pixel values ​​at corresponding positions in the target near-focus image and the synthesized far-focus image are linearly weighted and summed according to the weight value corresponding to the fusion weight mask to obtain the fused pixel value; For invalid edge regions in the synthesized telephoto image caused by parallax offset, the pixel values ​​at the corresponding positions in the target near-focus image are used to fill them; Based on the pixel values ​​after fusion of all pixels, a corresponding panoramic depth image is generated.

[0110] Among some feasible implementation methods are: The original image acquisition module 201 is used to acquire the original near-focus image and the original far-focus image synchronously acquired by the asymmetric binocular industrial endoscope. The parameter acquisition module is used to acquire the pre-calibrated intrinsic parameter matrix, distortion coefficients, and binocular extrinsic parameters; The distortion correction module is used to perform distortion correction on the original near-focus image and the original far-focus image according to the intrinsic parameter matrix and the distortion coefficient, respectively, to obtain the distortion-free near-focus image and the distortion-free far-focus image; The epipolar correction module is used to perform epipolar correction on the distortion-free near-focus image and the distortion-free far-focus image based on the binocular extrinsic parameters, so as to align the left and right image rows and obtain the corrected target near-focus image and target far-focus image.

[0111] In some feasible implementations, the distortion correction module is specifically used for: For the pixel coordinates of each target pixel in the original near-focus image and the original far-focus image, the pixel coordinates are converted into normalized planar coordinates according to the inverse of the intrinsic parameter matrix; Substitute the normalized planar coordinates into the distortion model that includes radial and tangential distortion coefficients to calculate the non-integer sampling coordinates corresponding to the target pixel; Based on the non-integer sampling coordinates, pixel values ​​are sampled in the original near-focus image or the original far-focus image using an interpolation algorithm, and assigned to the pixel coordinates of the target pixel to obtain the distortion-free near-focus image and the distortion-free far-focus image.

[0112] In some feasible implementations, the epipolar correction module is specifically used for: Using the rotation matrix and translation vector in the binocular extrinsic parameters, the left correction rotation matrix and right correction rotation matrix are calculated by the epipolar correction algorithm to rotate the left camera image plane and the right camera image plane in the binocular sensor to coplanar parallelism. Construct a virtual camera projection matrix, which is used to make the epipolar corrected left and right images have the same virtual focal length and the same vertical coordinate center. The left correction rotation matrix, the right correction rotation matrix, and the virtual camera projection matrix are used to perform a reprojection transformation on the distortion-free near-focus image and the distortion-free far-focus image to obtain a row-aligned target near-focus image and a target far-focus image.

[0113] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.

[0114] In addition, this invention also provides an electronic device, including: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the various processes of the above-described asymmetric binocular industrial endoscope depth-of-field extension method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0115] This invention also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the various processes of the aforementioned asymmetric binocular industrial endoscope depth-of-field extension method embodiment, achieving the same technical effect. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0116] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0117] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, embodiments of the present invention can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, EEPROM, Flash, and eMMC, etc.) containing computer-usable program code.

[0118] Embodiments of the present invention are described with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0119] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0120] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0121] Although preferred embodiments of the present invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention.

[0122] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0123] The foregoing has provided a detailed description of an asymmetric binocular industrial endoscope depth-of-field extension method and an asymmetric binocular industrial endoscope depth-of-field extension device. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for extending the depth of field in an asymmetric binocular industrial endoscope, characterized in that, include: Acquire the corrected near-focus and far-focus images of the target; Stereo matching is performed on the near-focus image of the target and the far-focus image of the target to obtain the corresponding first disparity map, and the first disparity map is converted into the corresponding physical depth map, and a fusion weight mask corresponding to the physical depth map is generated. The target telephoto image is reprojected pixel by pixel according to the first disparity map, and the viewpoint of the target telephoto image is transformed to the viewpoint of the target near-focus image to obtain a synthetic telephoto image that is spatially aligned with the target near-focus image. The target near-focus image and the synthesized far-focus image are weighted and fused according to the fusion weight mask to obtain the corresponding full-depth image.

2. The method according to claim 1, characterized in that, The step of performing stereo matching between the near-focus image and the far-focus image of the target to obtain the corresponding first disparity map includes: The near-focus image of the target and the far-focus image of the target are respectively input into a pre-trained deep learning stereo matching network, and the first feature map of the near-focus image of the target and the second feature map of the far-focus image of the target are extracted by the deep learning stereo matching network. The correlation between the first feature map and the second feature map is calculated to obtain the correlation cost body; The relevant cost body is iteratively optimized in multiple rounds to output a dense disparity map corresponding to each pixel in the near-focus image of the target, which serves as the first disparity map.

3. The method according to claim 1 or 2, characterized in that, The step of converting the first disparity map into a corresponding physical depth map includes: The virtual focal length and binocular baseline length used when correcting the near-focus image and the far-focus image of the target are obtained, and the disparity value corresponding to each pixel in the first disparity map is obtained. For each pixel in the first disparity map, the physical depth value corresponding to the pixel is calculated using the virtual focal length, the binocular baseline length, and the disparity value, based on the triangulation principle, to generate a physical depth map corresponding to the first disparity map.

4. The method according to claim 1 or 2, characterized in that, The generation of the fusion weight mask corresponding to the physical depth map includes: Obtain the depth value corresponding to each pixel in the physical depth map; Based on the matching result between the depth value and the preset mapping function, the initial weights corresponding to each pixel in the physical depth map are determined; Occlusion detection is performed on the near-focus image and the far-focus image of the target to identify the occlusion area; Based on the occlusion region, target pixels are extracted from the physical depth map, and the initial weights of the target pixels are corrected to generate a fusion weight mask with the same resolution as the near-focus image of the target.

5. The method according to claim 1, characterized in that, The step of reprojecting pixels of the target telephoto image according to the first disparity map, transforming the viewpoint of the target telephoto image to the viewpoint of the target near-focus image, and obtaining a composite telephoto image spatially aligned with the target near-focus image, includes: For each pixel in the near-focus image of the target, the sampling coordinates of the pixel in the far-focus image of the target are calculated according to the disparity value of the corresponding position in the first disparity map. Based on the sampling coordinates, the target telephoto image is resampled using a bilinear interpolation algorithm to generate a synthetic telephoto image that is spatially aligned with the target near-focus image.

6. The method according to claim 1, characterized in that, The step of weighted fusing the target near-focus image and the synthesized far-focus image according to the fusion weight mask to obtain the corresponding full-depth image includes: For each color channel, the pixel values ​​at corresponding positions in the target near-focus image and the synthesized far-focus image are linearly weighted and summed according to the weight value corresponding to the fusion weight mask to obtain the fused pixel value; For invalid edge regions in the synthesized telephoto image caused by parallax offset, the pixel values ​​at the corresponding positions in the target near-focus image are used to fill them; Based on the pixel values ​​after fusion of all pixels, a corresponding panoramic depth image is generated.

7. The method according to claim 1, characterized in that, Also includes: Acquire raw near-focus and raw far-focus images simultaneously acquired by an asymmetric binocular industrial endoscope; Obtain the pre-calibrated intrinsic parameter matrix, distortion coefficients, and binocular extrinsic parameters; Based on the intrinsic parameter matrix and the distortion coefficients, distortion correction is performed on the original near-focus image and the original far-focus image respectively to obtain the distortion-free near-focus image and the distortion-free far-focus image. Based on the binocular extrinsic parameters, epipolar correction is performed on the distortion-free near-focus image and the distortion-free far-focus image to align the left and right image rows, thereby obtaining the corrected target near-focus image and target far-focus image.

8. An asymmetric binocular industrial endoscope depth-of-field extension device, characterized in that, include: The image acquisition module is used to acquire the corrected near-focus image and far-focus image of the target; The depth processing module is used to perform stereo matching on the near-focus image of the target and the far-focus image of the target to obtain the corresponding first disparity map, convert the first disparity map into the corresponding physical depth map, and generate a fusion weight mask corresponding to the physical depth map. The reprojection module is used to reproject the target telephoto image pixel by pixel according to the first disparity map, transform the viewpoint of the target telephoto image to the viewpoint of the target near-focus image, and obtain a composite telephoto image that is spatially aligned with the target near-focus image. The fusion module is used to perform weighted fusion of the target near-focus image and the synthesized far-focus image according to the fusion weight mask to obtain the corresponding full-depth image.

9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; The memory is used to store computer programs; When the processor executes a program stored in the memory, it implements the method as described in any one of claims 1-7.

10. A computer-readable storage medium having instructions stored thereon that, when executed by one or more processors, cause the processors to perform the method as described in any one of claims 1-7.