A deep learning-based full-focus imaging method and related components

By building a full-focus imaging system and deep learning algorithms, combined with Gray code structured light and a motorized focusing lens camera, the problems of poor image restoration quality and high time cost in full-focus imaging were solved, achieving efficient full-focus imaging.

CN116528044BActive Publication Date: 2026-03-10TSINGHUA UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-18
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing all-focus imaging methods suffer from poor image recovery quality and long shooting time, resulting in excessive time costs.

Method used

A deep learning-based all-focus imaging method is adopted. By building an all-focus imaging system, using Gray code structured light and a motorized focusing lens camera, and combining an all-focus perception algorithm with a CNN+Transformer hybrid architecture, the all-focus scanning image is downsampled, transformed, and upsampled to generate an all-focus image.

Benefits of technology

It significantly improves the quality and accuracy of all-focus images, reduces shooting time costs, and enables all-focus imaging in deep-field scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116528044B_ABST
    Figure CN116528044B_ABST
Patent Text Reader

Abstract

The application provides a kind of full focus imaging method and related components based on deep learning, which comprises: establishing a full focus imaging system;Based on the full focus imaging system, the focus scanning image of the object to be measured is obtained;Based on the full focus perception algorithm, the full focus image of the focus scanning image is obtained, so as to carry out full focus imaging according to the full focus image, the method is simple, the operation efficiency is high, the time cost can be significantly reduced, the full focus imaging in large longitudinal depth scene can be realized, and the quality and precision of the full focus image are significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of total focusing imaging technology, and in particular to a deep learning-based total focusing imaging method and related components. Background Technology

[0002] In traditional camera imaging, only targets at a specific depth within the scene can be focused, resulting in a relatively clear image. Other targets outside this depth range appear out of focus and blurry. According to imaging principles, the range of movable depths within which an object can remain in sharp focus is called the camera's depth of field. Traditional cameras, limited by depth of field, struggle to capture all information about a large scene in a single image. Total focusing imaging methods overcome this limitation, enabling full-focus imaging of the scene and resulting in clearer coded fringe textures at different depths. Existing total focusing imaging methods can be categorized into focused scanning imaging methods and multi-image fusion methods.

[0003] The focused scanning imaging method moves the sensor at a constant speed during the camera's exposure time, generating an image with approximately spatially invariant blur (i.e., a focused scanning image). A fully focused image can then be obtained by deblurring the focused scanning image. However, due to the inconsistent spatial blur of the focused scanning image, deconvolving the focused scanning image using a pre-estimated single blur kernel results in relatively poor quality of the restored fully focused image.

[0004] Multi-image fusion methods require capturing a sequence of images of a scene using different focal lengths. Since these images are focused at different depths within the scene, each image contains both sharp and blurred portions. After capturing the image sequence, an image fusion algorithm extracts the sharp portions from each image and combines them to obtain a fully focused image of the entire scene. While this method can produce high-quality fully focused images, the requirement to capture multiple images results in a long capture time and high time costs.

[0005] Therefore, designing a full-focus imaging method that achieves high-quality image restoration and low time cost is an urgent problem to be solved. Summary of the Invention

[0006] This invention provides a deep learning-based all-focus imaging method and related components to address the shortcomings of existing all-focus images, such as poor recovery quality, long shooting time, and high time cost. The method is simple, computationally efficient, and can significantly reduce time costs. It can achieve all-focus imaging in deep-field scenes and significantly improve the quality and accuracy of all-focus images.

[0007] This invention provides a deep learning-based all-focus imaging method, comprising: establishing an all-focus imaging system; acquiring a focused scan image of the object under test based on the all-focus imaging system; acquiring an all-focus image of the focused scan image based on an all-focus perception algorithm, so as to perform all-focus imaging based on the all-focus image.

[0008] According to a deep learning-based all-focusing imaging method provided by the present invention, the establishment of an all-focusing imaging system includes: constructing a projector capable of projecting Gray code structured light and a camera including an electrically adjustable focusing lens; and calibrating the all-focusing imaging system composed of the projector and the camera using the Gray code calibration method.

[0009] According to a deep learning-based all-focusing imaging method provided by the present invention, the method for acquiring a focused scan image of an object under test based on the all-focusing imaging system includes: controlling the projector to project Gray code structured light onto the object under test; setting the current change period of the motorized focusing lens of the camera to be equal to the exposure time of the camera, and controlling the camera to capture the object under test to obtain a focused scan image of the object under test.

[0010] According to a deep learning-based all-focus imaging method provided by the present invention, the method for obtaining an all-focus image of the focused scan image based on the all-focus perception algorithm includes: using an encoder module to downsample the focused scan image to obtain image extraction features; using a Transformer module to perform Gaussian axial attention and external complementary attention transformation on the image extraction features to obtain Transformer-transformed image features; and using a decoder module to upsample the Transformer-transformed image features to obtain an all-focus image.

[0011] According to a deep learning-based full-focus imaging method provided by the present invention, the step of downsampling the focused scan image using an encoder module to obtain image extraction features includes: the focused scan image is downsampled through a convolutional layer and a residual module to obtain the image extraction features.

[0012] According to a deep learning-based all-focusing imaging method provided by the present invention, the step of using a Transformer module to perform Gaussian axial attention and external complementary attention transformation on the image extracted features to obtain Transformer-transformed image features includes: calculating the correlation within each feature window of the image extracted features using a local self-attention mechanism to obtain a local self-attention calculation result; performing feature processing on the local self-attention calculation result using a global self-attention mechanism to obtain a global self-attention calculation result; and weighting the global self-attention calculation result using a shared memory vector to obtain the Transformer-transformed image features.

[0013] According to a deep learning-based full-focus imaging method provided by the present invention, the step of upsampling the image features after the Transformer transformation using a decoder module to obtain a full-focus image includes: the image features after the Transformer transformation are upsampled through a stitching layer, a residual module, and a deconvolution module to obtain the full-focus image.

[0014] According to the deep learning-based full-focus imaging method provided by the present invention, the method further includes: decoding the full-focus image and calculating the three-dimensional coordinates of the object under test according to the triangulation method, so as to perform full-focus three-dimensional imaging based on the three-dimensional coordinates.

[0015] The present invention also provides a deep learning-based all-focusing imaging device, comprising: an establishment module for establishing an all-focusing imaging system; a first acquisition module for acquiring a focused scan image of an object under test based on the all-focusing imaging system; and a second acquisition module for acquiring an all-focusing image of the focused scan image based on an all-focusing perception algorithm, so as to perform all-focusing imaging based on the all-focusing image.

[0016] The present invention also provides a system employing the above-described deep learning-based all-focusing imaging method, comprising a projector capable of projecting Gray code structured light and a camera including an electrically adjustable focusing lens.

[0017] This invention provides a deep learning-based all-focus imaging method and related components. The method includes: establishing an all-focus imaging system; acquiring a focused scan image of the object under test based on the all-focus imaging system; acquiring an all-focus image of the focused scan image based on an all-focus perception algorithm, and performing all-focus imaging based on the all-focus image. The method is simple, has high computational efficiency, can significantly reduce time costs, and can achieve all-focus imaging in deep scenes with significantly improved quality and accuracy of the all-focus image. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating a deep learning-based all-focusing imaging method provided by the present invention.

[0020] Figure 2 This is a schematic diagram of the principle of a deep learning-based all-focusing imaging system provided by the present invention;

[0021] Figure 3 This is a schematic diagram of the all-focusing sensing algorithm framework provided by the present invention;

[0022] Figure 4 This is a diagram of the pinhole imaging model of the structured light three-dimensional imaging system provided by the present invention;

[0023] Figure 5 This is a schematic diagram of the structure of a deep learning-based all-focusing imaging device provided by the present invention. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0025] The following is combined with Figures 1-5 This invention describes a deep learning-based all-focusing imaging method and related components.

[0026] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating a deep learning-based all-focusing imaging method provided by the present invention.

[0027] Please refer to Figure 2 , Figure 2 This is a schematic diagram of a deep learning-based all-focusing imaging system provided by the present invention.

[0028] This invention provides a deep learning-based all-focusing imaging method, comprising:

[0029] 101: Establish a full-focus imaging system;

[0030] 102: Acquiring focused scan images of the object under test based on a total focusing imaging system;

[0031] 103: Obtain a full-focus image of a full-focus scan image based on a full-focus sensing algorithm, and perform full-focus imaging based on the full-focus image.

[0032] Out-of-focus blur is a common problem in traditional camera photography. Only targets at a specific depth within a scene can be focused, resulting in a sharp image, while other targets outside that depth range appear blurred. According to imaging principles, the range of depths within which an object can remain in sharp focus is called the camera's depth of field. Traditional cameras, limited by depth of field, struggle to capture all information about a large scene in a single image. All-focus imaging technology overcomes this limitation, allowing for full-focus imaging of the entire scene. However, existing all-focus imaging methods suffer from poor image quality and long shooting times, resulting in high time costs.

[0033] To address the shortcomings of existing technologies, this invention provides a deep learning-based full-focus imaging method that enables the generation of full-focus images from focused scanning images in an end-to-end manner, thereby significantly enhancing the reconstruction quality of full-focus images.

[0034] Specifically, firstly, an all-focus imaging system is established. For example, the system could consist of an industrial camera and a projector. The camera lens could be an electronically focused liquid lens, which can shorten the time required to capture the focused scanning image. After the all-focus imaging system is built, it can be calibrated to facilitate all-focus imaging in deep-field scenes. Then, the established all-focus imaging system is used to capture a focused scanning image of the object under test. This focused scanning image is referred to as the "focused scanning image." However, the focused scanning image acquired by the focused scanning imaging system may have some degree of defocus blur. Therefore, a deep learning-based all-focus imaging algorithm, specifically a CNN+Transformer hybrid architecture all-focus perception algorithm, is needed to obtain an all-focus image from the focused scanning image, enabling all-focus imaging based on this image.

[0035] Furthermore, the overall network architecture based on the full-focus perception algorithm can be, but is not limited to, an encoder module, a Transformer module, and a decoder module.

[0036] Of course, in order to further realize full-focus perception of coded stripe patterns at different depth positions in deep scenes, thereby expanding the depth of field of the structured light 3D imaging system and significantly improving the accuracy of 3D imaging, 3D coordinate calculation can be performed on the full-focus image to make the 3D imaging results in deep scenes more accurate and robust. This application does not impose any special limitations here.

[0037] In summary, the deep learning-based all-focus imaging method of the present invention is computationally simple and efficient, which can significantly reduce time costs, realize all-focus imaging in deep-field scenes, and significantly improve the quality and accuracy of all-focus images.

[0038] Based on the above embodiments:

[0039] As a preferred embodiment, a full-focus imaging system is established, including: setting up a projector capable of projecting Gray code structured light and a camera including an electrically adjustable focusing lens; and calibrating the full-focus imaging system composed of the projector and the camera using the Gray code calibration method.

[0040] In this embodiment, the all-focus imaging system can be constructed from an industrial camera and a DLP projector, with an additional motorized focusing lens added to the camera lens for capturing focused scan images.

[0041] Specifically, the projector projects a pre-set encoding pattern onto the scene to be reconstructed, the camera captures a focused scan image of the scene and generates a full-focus image, then decodes the full-focus image, and calculates the depth of the scene based on the decoding result and the system calibration result.

[0042] Gray code calibration can be used to calibrate a full-focus imaging system, with a dot plot serving as the reference calibration board. First, the reference calibration board is positioned in multiple poses and orientations. Then, Gray code structured light is projected onto the scene. Upon completion of calibration, the camera's intrinsic and extrinsic parameters, distortion coefficients, projector's intrinsic and extrinsic parameters, distortion coefficients, and the relative extrinsic parameters of the camera and projector will be obtained.

[0043] For example, two fixed focal length positions, 500mm and 800mm, were set up for calibration to obtain independent parameters. The corresponding parameters were then applied to images acquired at different focal lengths for reconstruction. Finally, the results from different focal positions were fused through image region segmentation to generate a 3D reconstruction of the entire scene.

[0044] In addition, the industrial camera model can be, but is not limited to, PointGreyGrasshopperGS3-U3-32S4M, the DLP projector model can be, but is not limited to, AODIN M8 projector, and the current-driven focus-adjusting lens model can be, but is not limited to, Optotune EL-10-30-Ci.

[0045] As a preferred embodiment, acquiring a focused scan image of the object under test based on a full-focus imaging system includes: controlling a projector to project Gray code structured light onto the object under test; setting the current change period of the camera's motorized focusing lens to be equal to the camera's exposure time, and controlling the camera to capture the object under test to obtain a focused scan image of the object under test.

[0046] In this embodiment, acquiring a focused scan image of the object under test based on a full-focus imaging system involves projecting Gray code structured light onto the object using a projector, and then capturing the focused scan image of the scene using a camera. To ensure the integrity of the focused scan, the current variation period of the motorized focusing lens needs to be set to be equal to the camera's exposure time. The driving current variation range is, for example, 30mA-50mA, and a triangular wave periodic signal can be used. The current variation period is, for example, 0.05s, and an 11-bit Gray code method can be used, divided into 11 horizontal codes and 11 vertical codes. A total of 22 focused scan images need to be acquired.

[0047] Please refer to Figure 3 , Figure 3 A schematic diagram of the all-focusing sensing algorithm framework provided by the present invention.

[0048] As a preferred embodiment, obtaining a fully focused image of a focused scan image based on a fully focused sensing algorithm includes: using an encoder module to downsample the focused scan image to obtain image extraction features; using a Transformer module to perform Gaussian axial attention and external complementary attention transformations on the image extraction features to obtain Transformer-transformed image features; and using a decoder module to upsample the Transformer-transformed image features to obtain a fully focused image.

[0049] Considering that the acquired focused scan image may have a certain degree of defocus blur, this embodiment uses a full-focus perception algorithm based on a CNN+Transformer hybrid architecture to obtain the full-focus image of the focused scan image.

[0050] Specifically, the full-focus perception network structure based on the CNN+Transformer hybrid architecture follows the encoder-decoder paradigm and consists of three parts: an encoder module, a Transformer module, and a decoder module.

[0051] In the encoder stage, a residual module similar to ResNet was designed to improve the performance of feature extraction at each level of the input image. Unlike the method of extracting visual features of the input image by simply stacking convolutional layers, the introduction of the residual module effectively solves the problems of gradient explosion and gradient vanishing, and greatly enhances the network's ability to extract image features at different levels.

[0052] The core module of the all-focusing perception algorithm is the Transformer module, whose main function is to further mine the features extracted by the encoder, thereby improving the all-focusing perception performance of the entire network. The Transformer module mainly consists of two parts: a Gaussian axial attention module and an external complementary attention module.

[0053] Generally, the basic structure of the decoder module is similar to that of the encoder. By introducing deconvolution upsampling operations corresponding to the convolution downsampling operations in the encoder, it ensures that the size of the fully focused image output by the network remains consistent with the original input image. In the network structure of this invention, in order to better realize the network's fully focused deblurring function, the decoder module introduces a context-based inference attention convolution operation to address the problem of complex and diverse blurred regions and degrees of blurring. This enhances the ability to perceive global information and further enriches the model's ability to express visual features through visual semantic reasoning.

[0054] As a preferred embodiment, an encoder module is used to downsample the focused scan image to obtain image extraction features, including: the focused scan image is downsampled through a convolutional layer and a residual module to obtain image extraction features.

[0055] Specifically, in the encoder stage, downsampling operations are performed through convolutional layers and residual modules to obtain multi-level visual feature information of the image. The convolutional layer contains a 3×3 convolutional layer, and the residual module calculates by adding the original input to the output of the convolutional layer, batch normalization layer, and activation function layer, and then passing it through a triple attention module to obtain the final encoded result.

[0056] As a preferred embodiment, the Transformer module is used to perform Gaussian axial attention and external complementary attention transformations on the extracted image features to obtain the Transformer-transformed image features. This includes: calculating the correlation within each feature window of the extracted image features using a local self-attention mechanism to obtain the local self-attention calculation result; performing feature processing on the local self-attention calculation result using a global self-attention mechanism to obtain the global self-attention calculation result; and weighting the global self-attention calculation result using a shared memory vector to obtain the Transformer-transformed image features.

[0057] Specifically, the Transformer module's main function is to further mine the features extracted by the encoder, thereby improving the overall focusing perception performance of the network. This module consists of two parts: a Gaussian axial attention module and an external complementary attention module. The Gaussian axial attention module uses a local self-attention mechanism to calculate the correlation within each feature window of the extracted image features, obtaining the local self-attention calculation result; then, a global self-attention mechanism is used to perform a linear mapping feature transformation on the local self-attention calculation result, obtaining the global self-attention calculation result.

[0058] Let Sim(Q,K) represent the similarity between the query matrix and the key matrix, and let the variance be σ. 2 The standard normal distribution of is defined as follows, and its corresponding probability density function can be expressed as:

[0059]

[0060] Where d is the distance between matrices Q, K, and V.

[0061] The final self-attention score can then be expressed as:

[0062]

[0063] in As a relative position bias, it can help the network strengthen its ability to perceive relative position information, thereby improving the performance of the model.

[0064] The external complementary attention module leverages the correlation between samples to enhance representation learning capabilities. During training, all samples share two memory units, thus recording the most core features of the entire dataset. Furthermore, this module adaptively recalibrates channel feature responses by explicitly modeling the interdependencies between channels. The external complementary attention module ensures that the input tensor X∈R for any training sample... H×W×C With memory vector M K and M V Performing matrix multiplication, the resulting structure is added to the original input X to obtain the tensor Y∈R. H×W×C .

[0065]

[0066] For tensor Y∈R H×W×C Z is first obtained through average pooling. P ∈R 1×1×C Then, a 1×1 convolutional layer is used to mine the correlation between channels, thereby obtaining the corresponding weight distribution of each channel. Finally, the output vector is normalized using the sigmoid activation function to obtain the weight vector V∈R. 1×1×CThe module's final output Z∈R H×W×C It can be represented as:

[0067]

[0068] in This represents matrix multiplication, and σ represents the Sigmoid activation function. This indicates a convolution operation with a kernel size of 1×1.

[0069] As a preferred embodiment, the image features after Transformer transformation are upsampled using a decoder module to obtain a fully focused image. This includes upsampling the image features after Transformer transformation through a stitching layer, a residual module, and a deconvolution module to obtain a fully focused image.

[0070] Specifically, the decoder module consists of a concatenation layer, a residual module, and a deconvolution module. During decoding, it restores the resolution through upsampling until it matches the resolution of the input image. The concatenation layer feeds the corresponding encoded features from the encoder into the corresponding layer of the decoder, concatenating them with the corresponding output as the input for the next layer of decoding. The convolution in the residual module employs context-based inference attention, which can adaptively adjust the weights of the convolution kernel to perceive information in different receptive domains.

[0071] Please refer to Figure 4 , Figure 4 A diagram of a pinhole imaging model for the structured light three-dimensional imaging system provided by this invention.

[0072] As a preferred embodiment, it further includes: decoding the full-focus image and calculating the three-dimensional coordinates of the object under test according to triangulation, so as to perform full-focus three-dimensional imaging based on the three-dimensional coordinates.

[0073] Specifically, to obtain Gray code information from a fully focused image, the Gray code pattern in the image needs to be decoded first to determine the coordinates of the projection plane. Assume the point S to be measured in the scene has coordinates (u... c v c The corresponding coordinates in the projection plane are (u p v p Based on the position of point S in the Gray code pattern, its encoded value can be determined as (C). u C v ), will (C u C v Converting ) to decimal will give you (u) p v p ).

[0074] The known calibration parameters of the system include A cA p The intrinsic parameters, referred to as camera and projection parameters, (R c ,T c ) and (R p ,T p ) are the extrinsic parameters of the camera and the projection, respectively. Let the camera coordinate system be (o). c ;x c ,y c ,z c ), projected coordinate system (o p ;x p ,y p ,z p ), World coordinate system (o ω ;x ω ,y ω ,z ω ).

[0075] Coordinates in the world coordinate system (x) ω ,y ω ,z ω ) and camera plane coordinates (u c v c ) and projection plane coordinates (u p v p The relationship is:

[0076]

[0077]

[0078] Where s c and s p These are fixed constant coefficients.

[0079] A c A p The intrinsic parameters, referred to as those of the camera and the projection, can be expressed as:

[0080]

[0081]

[0082] in These represent the equivalent focal lengths of the camera and projector along two axes, respectively. These represent the optical center points of the imaging and illumination planes, respectively.

[0083] Then, by combining the distortion coefficients of the camera and projector, distortion correction is performed to obtain the distortion-corrected pixel coordinates (u'). c ,v' c ), (u' p ,v' p The following system of linear equations can be established:

[0084] f1(x ω ,y ω ,z ω ,u′ c ) = 0

[0085] f2(x ω ,y ω ,z ω ,v′ c ) = 0

[0086] f3(x ω ,y ω ,z ω ,u′ p ) = 0

[0087] f4(x ω ,y ω ,z ω ,v′ p ) = 0

[0088] After the system's geometric parameters are calibrated, the parameters to be solved for the above linear equation system are (x... ω ,y ω ,z ω The world coordinates of the target point can be obtained by solving for the coordinates of the target point.

[0089] The fully focused image containing the Gray code pattern is decoded accordingly. Due to the change in the surface shape of the measured object, the stripes of the projected pattern are bent and deformed. The extracted stripe boundaries are calculated using triangulation, and then the three-dimensional point cloud coordinates (x, y, z) are obtained by combining the system calibration parameters. ω ,y ω ,z ω This yields a fully clear 3D imaging result.

[0090] Please refer to Figure 5 , Figure 5 This is a schematic diagram of a deep learning-based all-focusing imaging device provided by the present invention.

[0091] The present invention also provides a deep learning-based all-focus imaging device, comprising: a setup module 501 for setting up an all-focus imaging system; a first acquisition module 502 for acquiring a focused scan image of the object under test based on the all-focus imaging system; and a second acquisition module 503 for acquiring an all-focus image of the focused scan image based on an all-focus perception algorithm, so as to perform all-focus imaging based on the all-focus image.

[0092] For a description of the deep learning-based all-focusing imaging device provided by the present invention, please refer to the above method embodiments; the present invention will not be described again here.

[0093] The present invention also provides a system employing the above-described deep learning-based all-focusing imaging method, comprising a projector capable of projecting Gray code structured light and a camera including an electrically adjustable focusing lens.

[0094] For an introduction to the deep learning-based all-focusing imaging system provided by this invention, please refer to the above method embodiments; the invention itself will not be described in detail here.

[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A deep learning based full focus imaging method, characterized in that, include: Establish a full-focus imaging system; The focused scan image of the object under test is acquired based on the total focusing imaging system. A full-focus image of the focused scan image is obtained based on the full-focus sensing algorithm, so as to perform full-focus imaging based on the full-focus image; The process of acquiring the full-focus image of the focused scan image based on the full-focus sensing algorithm includes: The focused scan image is downsampled using an encoder module to obtain image extraction features; The image features are extracted using the Transformer module and transformed by Gaussian axial attention and external complementary attention to obtain the image features after Transformer transformation. The image features after the Transformer transformation are upsampled using a decoder module to obtain the fully focused image. The step of using the Transformer module to perform Gaussian axial attention and external complementary attention transformations on the image features to obtain the Transformer-transformed image features includes: The correlation within each feature window of the image feature extraction is calculated using a local self-attention mechanism to obtain the local self-attention calculation result; The local self-attention calculation results are processed using a global self-attention mechanism to obtain the global self-attention calculation results. The global self-attention calculation results are weighted by a shared memory vector to obtain the image features after the Transformer transformation.

2. The deep learning based full-focus imaging method of claim 1, wherein, The establishment of the total focusing imaging system includes: Set up a projector capable of projecting Gray code structured light and a camera including an electrically adjustable focus lens; The Gray code calibration method is used to calibrate the total focusing imaging system consisting of the projector and the camera.

3. The deep learning based full focus imaging method of claim 2, wherein, The acquisition of a focused scan image of the object under test based on the total focusing imaging system includes: Control the projector to project Gray code structured light onto the object under test; The current change period of the motorized focusing lens of the camera is set to be equal to the exposure time of the camera, and the camera is controlled to capture the object under test to obtain a focused scan image of the object under test. 4.The deep learning based full-focus imaging method of claim 1, wherein, The step of downsampling the focused scan image using an encoder module to obtain image extraction features includes: The focused scan image is downsampled through a convolutional layer and a residual module to obtain the image extraction features.

5. The deep learning based full-focus imaging method of claim 1, wherein, The step of upsampling the image features after the Transformer transformation using a decoder module to obtain a fully focused image includes: The image features transformed by the Transformer are upsampled through a stitching layer, a residual module, and a deconvolution module to obtain the fully focused image.

6. The deep learning based full-focus imaging method of any one of claims 1 to 5, wherein, Also includes: The full-focus image is decoded and the three-dimensional coordinates of the object under test are calculated using triangulation, so as to perform full-focus three-dimensional imaging based on the three-dimensional coordinates.

7. A deep learning based full-focus imaging apparatus, characterized by, include: Establish a module for building a full-focus imaging system; The first acquisition module is used to acquire a focused scan image of the object under test based on the total focusing imaging system; The second acquisition module is configured to acquire a full-focus image of the focus scanning image based on a full-focus perception algorithm, and perform full-focus imaging according to the full-focus image. The full-focus image of the focus scanning image is acquired based on a full-focus perception algorithm, and the full-focus imaging is performed according to the full-focus image. The encoder module is adopted to perform down-sampling processing on the focus scanning image to obtain image extraction features; The image extraction features are transformed by the Transformer module in a Gaussian axial attention and external complementary attention manner to obtain image features after Transformer transformation; The decoder module is adopted to perform up-sampling processing on the image features after Transformer transformation to obtain the full-focus image; The image extraction features are transformed by the Transformer module in a Gaussian axial attention and external complementary attention manner to obtain image features after Transformer transformation, and the transformation includes: The local self-attention mechanism is adopted to calculate the correlation inside each feature window of the image extraction features to obtain local self-attention calculation results; The global self-attention mechanism is adopted to perform feature processing on the local self-attention calculation results to obtain global self-attention calculation results; The global self-attention calculation results are weighted by the shared memory vector to obtain the image features after Transformer transformation.

8. A system employing the deep learning based full focus imaging method of any one of claims 1 to 6, characterized in that, A projector including a projectable Gray code structured light and a camera including a motorized focus lens.

Citation Information

Patent Citations

  • Large-depth-of-field fringe projection three-dimensional measurement method based on point spread function solution

    CN115546285A