A three-dimensional image reconstruction method and system based on coordinate attention mechanism and Unet
By combining deep learning and the coordinate attention mechanism (COAM) of a 3D image reconstruction system based on the coordinate attention mechanism and Unet, the problem of unsatisfactory modeling accuracy of complex and irregular objects is solved, and efficient 3D image reconstruction is achieved, which is applicable to biomedicine and cultural relic protection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-12
- Publication Date
- 2026-04-07
AI Technical Summary
In the process of processing, traditional geometric modeling techniques suffer from poor modeling efficiency and accuracy in the 3D image reconstruction of complex and irregular objects. Existing technologies are not ideal in processing complex and irregular objects, and the modeling cycle is long.
A 3D image reconstruction method based on coordinate attention mechanism and Unet is adopted. Through a system composed of photoelectric switches, projectors, cameras and motion controllers, combined with deep learning and coordinate attention mechanism (COAM), efficient 3D image reconstruction of complex and irregular objects is achieved.
It improves the accuracy and efficiency of 3D image reconstruction of complex and irregular objects, and can accurately extract the contour and depth information of the object, making it suitable for fields such as biomedicine and cultural relic protection.
Smart Images

Figure CN115937407B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of three-dimensional image reconstruction technology, and more specifically, to a three-dimensional image reconstruction method and system based on coordinate attention mechanism and Unet. Background Technology
[0002] 3D image reconstruction technology is a virtual reality technology that represents the three-dimensional world in a computer. It can create 3D mathematical models of 3D objects that are easy for computers to process and apply, allowing 2D images to be reverse-engineered into 3D objects. Therefore, 3D image reconstruction technology has wide applications in biomedicine, cultural relic preservation, and other fields.
[0003] Among numerous 3D image reconstruction systems, traditional geometric modeling techniques are mature and offer many advantages, such as the ability to accurately model 3D objects. However, geometric modeling techniques are less effective for reconstructing complex and irregular objects. Furthermore, geometric modeling is only suitable for offline object modeling, resulting in long modeling cycles. Summary of the Invention
[0004] To address the shortcomings of existing technologies, the present invention aims to provide a three-dimensional image reconstruction method and system based on coordinate attention mechanism and Unet, which has the advantages of improving the efficiency of three-dimensional image reconstruction and improving the modeling accuracy of complex and irregular objects.
[0005] The above-mentioned technical objective of this invention is achieved through the following technical solution: a three-dimensional image reconstruction method based on coordinate attention mechanism and Unet, comprising the following steps:
[0006] S1. A 3D image reconstruction system based on coordinate attention mechanism and Unet;
[0007] S2. On the three-dimensional image reconstruction system, a three-dimensional geometric image of the target is obtained according to a deep learning three-dimensional image reconstruction method.
[0008] In one embodiment, step S1 specifically involves:
[0009] A fixed conveyor belt is used, and a photoelectric switch is installed at the feed end of the conveyor belt, with the photoelectric switch being perpendicular to the conveyor belt.
[0010] The projector is mounted on a high-precision angle adjuster and is referred to as the first assembly. The first assembly is placed on one side of the conveyor belt.
[0011] The camera is installed at the exit end of the conveyor belt, and the light-blocking plate is installed on one side of the conveyor belt, facing the mirror surface of the camera.
[0012] When installing the photoelectric switch, the first assembly, and the camera, ensure that the projector, the camera, and the target are on the same horizontal plane.
[0013] In one embodiment, step S1 further includes:
[0014] The motion controller is connected to the conveyor belt. The photoelectric switch, the conveyor belt, and the motion controller together form a motion control device for identifying the target object and accurately transporting the target object to the shooting position.
[0015] Connect the computer to the motion controller, the projector, and the camera respectively;
[0016] Rotate the high-precision angle adjuster clockwise so that the horizontal angle between the mirror of the projector and the mirror of the camera is α. The projector projects stripe features onto the target object, and the camera acquires an image of the target object with the stripe features.
[0017] This led to the construction of a 3D image reconstruction system based on coordinate attention mechanism and Unet.
[0018] In one embodiment, step S2 includes the following steps:
[0019] S21. Use the three-dimensional image reconstruction system to capture a target image with stripes, and input the target image into a neural network;
[0020] S22. Introduce the Unet deep learning network to obtain the three-dimensional geometric image of the target object;
[0021] The image I∈R of the target C×H×W The input is fed into a neural network, and the output on the computer is a three-dimensional geometric image O∈R. 3×H×W Its mathematical model is expressed as follows:
[0022] O = O(x,y) = η(I,Θ)
[0023] In the above formula, O(x,y) represents a three-dimensional geometric image, η(·) is the mathematical expression of the deep neural network model, Θ is the network parameter of the deep neural network, and (x,y) represents the pixel coordinates.
[0024] In one embodiment, step S2 further includes the following steps:
[0025] S23. Introduce the Coordinate Attention (COAM) mechanism into the Unet network;
[0026] Given a C a The feature map F = [f1,L,f] of ×H×W C ], Ca The feature map has a size of H×W;
[0027] S24. Input the feature map F into the coordinate attention mechanism (COAM) to obtain a feature map U with complex and irregular regions. Its mathematical model is as follows:
[0028] U = COAM(F) = F·υ(F)
[0029] In the above formula, υ(·) is the feature localization and extraction function for the horizontal and vertical axes;
[0030] The aforementioned horizontal and vertical axis feature localization refers to performing average pooling Avg on the feature map F along the vertical and horizontal axes respectively, and obtaining the localization feature vectors Z∈R for each channel's horizontal and vertical axes through shared convolution. C / r×1×(H+W) Its mathematical model can be represented as follows:
[0031] Z = [Z h Z w ]=σ(W a [Avg x (F), Avg y (F)])
[0032] In the above formula, W a These are the weights of the convolutional layer, C / r is the number of feature channels in the output, σ(·) is the sigmoid modulo function, and Avg... x (·) represents average pooling along the vertical axis, Avg y (·) represents the average pooling process along the horizontal axis, r is the decay rate of the feature channel, and Z is the average pooling rate along the horizontal axis. h Z is the eigenvector along the vertical axis. w The eigenvectors along the vertical axis;
[0033] Then, feature extraction is performed on the feature vectors located along the horizontal and vertical axes using convolutional layers. The mathematical model is as follows:
[0034] υ(F)=[σ(W b Z h ),σ(W c Z w )]
[0035] In the above formula, W b W c are the weights of the convolutional layer, and the number of feature channels in the output is C. σ(·) is the sigmoid function.
[0036] In the training process of deep neural networks, the SGD function is used to optimize the loss function L(Θ), and the process is as follows:
[0037]
[0038] In the above formula, M represents the total number of training iterations, and I... j For the j-th image set being tested, G j Let θ be the j-th 3D geometric image, and Θ be the neural network parameters.
[0039] In one embodiment, step S2 further includes the following steps:
[0040] S24. Generate training dataset The image set in;
[0041] The three-dimensional image reconstruction system captures an image of the target object with stripes. j In a virtual environment, the dimensional relationships of the 3D image reconstruction system are restored according to the scaling factor β, and the measured target image I is... j A model is built, resulting in a three-dimensional geometric image G. j Finally, the complete training set A can be obtained;
[0042] After M training iterations, the optimized parameters can be obtained.
[0043] For the target image I captured by the three-dimensional image reconstruction system g It can be done Obtain a 3D geometric image.
[0044] The aforementioned 3D image reconstruction method based on coordinate attention mechanism and Unet has the following beneficial effects:
[0045] Firstly, Unet was used to accurately extract the stripe features of the image under test, mapping out the contour and depth information of the target image under test.
[0046] Secondly, based on Unet, the Coordinate Attention (COAM) mechanism is introduced, which enables the neural network model to accurately locate and extract features of complex and irregular regions, improves the accuracy of 3D image reconstruction of complex and irregular objects, and is beneficial to the application research of 3D image reconstruction technology.
[0047] A 3D image reconstruction system based on coordinate attention mechanism and Unet includes a computer, projector, motion controller, photoelectric switch, conveyor belt, light shield, camera, high-precision angle adjuster, and target object;
[0048] The computer is connected to the projector, the camera, and the motion controller. The motion controller is connected to the photoelectric switch and the conveyor belt. The photoelectric switch is located at the inlet of the conveyor belt, and the camera is located at the outlet of the conveyor belt. The camera is mounted on the high-precision angle adjuster, which is located on one side of the conveyor belt and between the camera and the photoelectric switch. The light-blocking plate is located on the other side of the conveyor belt and opposite the camera.
[0049] The aforementioned 3D image reconstruction system based on coordinate attention mechanism and Unet has the following beneficial effects:
[0050] The motion control device, consisting of a controller, photoelectric switch, and conveyor belt, can identify the target object and accurately transport it to the designated position. The projector can map stripe features onto the target object, facilitating the extraction of the target object's three-dimensional information by the neural network. The light-blocking plate prevents external light from interfering with the camera's acquisition of the target image, thus improving the system's stability. Attached Figure Description
[0051] Figure 1 This is a schematic diagram of the composition of the three-dimensional image reconstruction system in this embodiment;
[0052] Figure 2 This is a schematic diagram of the combination of the projector and the high-precision angle adjuster in this embodiment;
[0053] Figure 3 This is a schematic diagram of the neural network architecture in this embodiment;
[0054] Figure 4 yes Figure 3 A schematic diagram of the components of the Coordinate Attention Mechanism (COAM).
[0055] In the diagram: 101, Computer; 102, Projector; 103, Motion Controller; 104, Photoelectric Switch; 105, Conveyor Belt; 106, Light Blocking Plate; 107, Target Object; 108, Camera; 109, High-Precision Angle Adjuster. Detailed Implementation
[0056] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0057] like Figure 1 As shown, a 3D image reconstruction method based on coordinate attention mechanism and Unet includes the following steps:
[0058] S1 provides a 3D image reconstruction system based on coordinate attention mechanism and Unet.
[0059] Specifically, step S1 also includes:
[0060] Fixed conveyor belt 105, photoelectric switch 104 is installed at L1 in front of conveyor belt 105, L1 = 50cm, photoelectric switch 104 is set perpendicular to conveyor belt 105;
[0061] The projector 102 is mounted on the high-precision angle adjuster 109 and is referred to as the first assembly. The first assembly is set at L2 in front of the conveyor belt 105, where L2 = 60cm.
[0062] The camera 108 is installed at L3 in front of the conveyor belt 105, where L3 = 60cm. The light-blocking plate 106 is installed on one side of the conveyor belt 105 and is positioned opposite to the mirror surface of the camera 108. The light-blocking plate 106 prevents the camera 108 from being interfered with by external light when taking pictures.
[0063] When installing the photoelectric switch 104, the first assembly and the camera 108, ensure that the projector 102, the camera 108 and the target object 107 are on the same horizontal plane.
[0064] Specifically, step S1 also includes:
[0065] The motion controller 103 is connected to the conveyor belt 105. The photoelectric switch 104, the conveyor belt 105 and the motion controller 103 form a motion control device, which is used to identify the target object 107 and accurately transport the target object 107 to the shooting position.
[0066] Connect the computer 101 to the motion controller 103, the projector 102, and the camera 108 respectively;
[0067] like Figure 2 As shown, the high-precision angle adjuster 109 is rotated clockwise so that the horizontal angle between the mirror of the projector 102 and the mirror of the camera 108 is α. When α = 30°, the projector 102 projects the stripe feature onto the target object 107. The camera 108 can acquire the image of the target object 107 with the stripe feature and upload the image of the target object 107 with the stripe feature to the neural network.
[0068] This led to the construction of a 3D image reconstruction system based on coordinate attention mechanism and Unet.
[0069] S2. In a 3D image reconstruction system, a 3D geometric image of the target is obtained based on a deep learning 3D image reconstruction method.
[0070] Specifically, step S2 also includes the following steps:
[0071] S21. Use a 3D image reconstruction system to capture a target image with stripes and input it into a neural network;
[0072] pass Acquire target image I g ;
[0073] S22. Introduce the Unet deep learning network to obtain the three-dimensional geometric image of the target object 107;
[0074] Image I∈R of the target C×H×W The input is fed into the neural network, and the output on the computer 101 is a three-dimensional geometric image O∈R. 3×H×W Its mathematical model is expressed as follows:
[0075] O = O(x,y) = η(I,Θ)
[0076] In the above formula, O(x,y) represents a three-dimensional geometric image, η(·) is the mathematical expression of the deep neural network model, Θ is the network parameter of the deep neural network, and (x,y) represents the pixel coordinates.
[0077] S221. Introduce the Coordinate Attention (COAM) mechanism into the Unet network;
[0078] Given a C a The feature map F = [f1,L,f] of ×H×W C ], C a The feature map has a size of H×W;
[0079] S222. Input the feature map F into the coordinate attention mechanism COAM to obtain a feature map U with complex and irregular regions. Its mathematical model is as follows:
[0080] U = COAM(F) = F·υ(F)
[0081] In the above formula, υ(·) is the feature localization and extraction function for the horizontal and vertical axes;
[0082] Horizontal and vertical axis feature localization refers to performing average pooling Avg along the vertical and horizontal axes on the feature map F respectively, and obtaining the localization feature vectors Z∈R for each channel along the horizontal and vertical axes through shared convolution. C / r×1×(H+W) Its mathematical model can be represented as follows:
[0083] Z = [Z h Z w ]=σ(W a [Avg x (F), Avg y (F)])
[0084] In the above formula, W a These are the weights of the convolutional layer, C / r is the number of feature channels in the output, σ(·) is the sigmoid modulo function, and Avg... x (·) represents average pooling along the vertical axis, Avg y(·) represents the average pooling process along the horizontal axis, r is the decay rate of the feature channel, and Z is the average pooling rate along the horizontal axis. h Z is the eigenvector along the vertical axis. w The eigenvectors along the vertical axis;
[0085] Then, feature extraction is performed on the feature vectors located along the horizontal and vertical axes using convolutional layers. The mathematical model is as follows:
[0086] υ(F)=[σ(W b Z h ),σ(W c Z w )]
[0087] In the above formula, W b W c are the weights of the convolutional layer, and the number of feature channels in the output is C. σ(·) is the sigmoid function.
[0088] In the training process of deep neural networks, the SGD function is used to optimize the loss function L(Θ), and the process is as follows:
[0089]
[0090] In the above formula, M represents the total number of training iterations, and I... j For the j-th image set being tested, G j Let θ be the j-th 3D geometric image, and Θ be the neural network parameters.
[0091] S23. Generate training dataset The image set in;
[0092] The three-dimensional image reconstruction system captures an image of the target object with stripes. j In a virtual environment, the dimensional relationships of the 3D image reconstruction system are restored according to the scaling factor β, and the measured target image I is... j A model is built, resulting in a three-dimensional geometric image G. j Finally, the complete training set A can be obtained;
[0093] After M training iterations, the optimized parameters can be obtained.
[0094] For the target object image I captured by the 3D image reconstruction system g It can be done Obtain a 3D geometric image.
[0095] like Figure 3 As shown, as the only optional embodiment, the target image I∈R is measured. C×H×W=1×512×512 The input is fed into a neural network, and the output is a three-dimensional geometric image O∈R. 3×H×W=3×512×512Its mathematical model can be represented as follows:
[0096] O = O(x,y) = η(I,Θ)
[0097] In the above formula, O(x,y) represents a three-dimensional geometric image, η(·) is the mathematical expression of the deep neural network model, Θ is the network parameter of the deep neural network, and (x,y) represents the pixel coordinates.
[0098] To improve the accuracy of 3D image reconstruction of complex and irregular objects, the Coordinate Attention (COAM) mechanism was introduced into the Unet network, which enables the neural network model to accurately locate and extract features of complex and irregular regions.
[0099] like Figure 4 As shown, given a C a The feature map F = [f1,L,f] = 128×512×512 ×H×W. C ], C a The feature map size is H×W = 512×512. It is input into the coordinate attention mechanism (COAM) to obtain a feature map U with complex and irregular regions. Its mathematical model can be represented as follows:
[0100] U = COAM(F) = F·υ(F)
[0101] Where υ(·) is the feature localization extraction function for the horizontal and vertical axes. Specifically, the horizontal and vertical axis feature localization refers to performing average pooling Avg on the feature map F along the vertical and horizontal axes respectively, and obtaining the localization feature vectors Z∈R for each channel through shared convolution. C / r×1×(H+W)=128 / 32×1×(512+512) Its mathematical model can be represented as follows:
[0102] Z = [Z h Z w ]=σ(W a [Avg x (F), Avg y (F)])
[0103] In the above expression, W a These are the weights of the convolutional layer, the number of output feature channels is C / r = 128 / 32, σ(·) is the sigmoid modulo function, and Avg x (·) represents average pooling along the vertical axis, Avg y (·) represents the average pooling process along the horizontal axis, and r = 32 is the decay rate of the characteristic channel. Z h Z is the eigenvector along the vertical axis. w The feature vectors are located along the vertical axis. Then, convolutional layers are used to extract features from the feature vectors located along the horizontal and vertical axes. The mathematical model can be represented as follows:
[0104] υ(F)=[σ(Wb Z h ),σ(W c Z w )]
[0105] In the above expression, W b W c σ is the weight of the convolutional layer, and the number of feature channels output is C = 128. σ(·) is the sigmoid function.
[0106] In the training process of deep neural networks, the SGD function is used to optimize the loss function L(Θ), and the process is as follows:
[0107]
[0108] In the above formula, M represents the total number of training iterations, J represents the number of samples in the training set, and I... j For the j-th image set being tested, G j Let θ be the j-th 3D geometric image. Θ represents the neural network parameters.
[0109] Training dataset The process of generating the image set in the image set is as follows:
[0110] The three-dimensional image reconstruction system captures an image of the target object with stripes. j In a virtual environment, the dimensional relationships of the 3D image reconstruction system are restored according to a scaling factor β = 0.5, and the measured target image I is... j A model is built, resulting in a three-dimensional geometric image G. j Finally, the complete training set A can be obtained.
[0111] After M=200 training iterations, the optimized parameters can be obtained.
[0112] For the target object image I captured by the 3D image reconstruction system g It can be done Obtain a 3D geometric image.
[0113] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.
Claims
1. A three-dimensional image reconstruction method based on coordinate attention mechanism and Unet, characterized in that, Includes the following steps: S1. A 3D image reconstruction system based on coordinate attention mechanism and Unet; S2. On the three-dimensional image reconstruction system, a three-dimensional geometric image of the target is obtained according to a deep learning three-dimensional image reconstruction method; Step S2 includes the following steps: S21. Use the three-dimensional image reconstruction system to capture a target image with stripes, and input the target image into a neural network; S22. Introduce the Unet deep learning network to obtain the three-dimensional geometric image of the target object; The image of the target The input is fed into a neural network, and the output on the computer is a three-dimensional geometric image. Its mathematical model is expressed as follows: In the above formula Represents a three-dimensional geometric image. It is the mathematical expression of a deep neural network model. These are the network parameters of a deep neural network. Represents pixel coordinates; Step S22 further includes the following steps: S221. Introducing a coordinate attention mechanism To the Unet network; Given a Feature map , The size of the feature map is ; S222, the feature map Input to coordinate attention mechanism In the process, feature maps with complex and irregular regions are obtained. Its mathematical model is expressed as follows: In the above formula Functions for locating and extracting features along the horizontal and vertical axes; The horizontal and vertical axis feature localization refers to the localization of feature maps. Average pooling was performed along the vertical and horizontal axes respectively. And the localization feature vectors of the horizontal and vertical axes of each channel are obtained through shared convolution. Its mathematical model can be represented as follows: In the above formula These are the weights of the convolutional layer. The number of feature channels in the output. yes function, It is an average pooling process along the vertical axis. It is an average pooling process along the horizontal axis. The decay rate of the characteristic channel. The eigenvectors along the vertical axis, The eigenvectors along the vertical axis; Then, feature extraction is performed on the feature vectors located along the horizontal and vertical axes using convolutional layers. The mathematical model is as follows: In the above formula , These are the weights of the convolutional layer, and the number of output feature channels is... , yes function; In the training process of deep neural networks, the SGD function is used to adjust the loss function. The optimization process is as follows: In the above formula Total number of training sessions For the first A set of images to be tested, For the first A three-dimensional geometric image, These are the network parameters of a deep neural network.
2. The three-dimensional image reconstruction method based on coordinate attention mechanism and Unet according to claim 1, characterized in that, Step S1 specifically involves: A fixed conveyor belt is used, and a photoelectric switch is installed at the feed end of the conveyor belt, with the photoelectric switch being perpendicular to the conveyor belt. The projector is mounted on a high-precision angle adjuster and is referred to as the first assembly. The first assembly is placed on one side of the conveyor belt. The camera is installed at the exit end of the conveyor belt, and the light-blocking plate is installed on one side of the conveyor belt, facing the mirror surface of the camera. When installing the photoelectric switch, the first assembly, and the camera, ensure that the projector, the camera, and the target are on the same horizontal plane.
3. The three-dimensional image reconstruction method based on coordinate attention mechanism and Unet according to claim 2, characterized in that, Step S1 further includes: The motion controller is connected to the conveyor belt. The photoelectric switch, the conveyor belt, and the motion controller together form a motion control device for identifying the target object and accurately transporting the target object to the shooting position. Connect the computer to the motion controller, the projector, and the camera respectively; Rotate the high-precision angle adjuster clockwise so that the horizontal angle between the projector's mirror and the camera's mirror is [value missing]. The projector projects stripe features onto the target object, and the camera acquires an image of the target object with the stripe features. This led to the construction of a 3D image reconstruction system based on coordinate attention mechanism and Unet.
4. The three-dimensional image reconstruction method based on coordinate attention mechanism and Unet according to claim 1, characterized in that, Step S2 further includes the following steps: S23. Generate training dataset The image set in; The three-dimensional image reconstruction system captures images of the target object with stripes. In a virtual environment, according to the scaling factor To restore the dimensional relationships of a 3D image reconstruction system and to analyze the measured target image. A model is built, resulting in a three-dimensional geometric image. Finally, the complete training set was obtained. ; go through After training, the optimized parameters are obtained. ; For the target object image captured by the three-dimensional image reconstruction system ,pass Obtain a 3D geometric image.
Citation Information
Patent Citations
IP-FSRGAN-CA face image super-resolution reconstruction algorithm based on coordinate attention mechanism
CN113674148A
Implicit function three-dimensional reconstruction method based on image and three-dimensional input
CN113763539A