Accurate positioning method and system for single polymer growth based on hybrid network
Through the dual-stage collaborative processing mechanism of the hybrid network, high-precision focused images are generated and three-level processing is performed, which solves the problems of nanometer-level precision and real-time performance in single polymer growth monitoring, achieves nanometer-level axial positioning accuracy and real-time processing at a high frame rate, and supports in-situ dynamic monitoring of polymer materials.
Patent Information
- Application Number
- CN202511202730.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-08-27
AI Technical Summary
Existing technologies in single polymer growth monitoring have problems such as dynamic blur interference, poor low-texture robustness, low computational efficiency, insufficient feature coupling of deep learning models, and lack of long-range dependencies, making it difficult to achieve the unity of nanometer-level precision and real-time processing.
A hybrid network-based method is adopted, combining a physical-driven and data-driven two-stage collaborative processing mechanism. High-precision focused images are generated through the PreFocusNet model, and a CNN-ViT hybrid model is used for three-level processing to achieve the unification of nanometer-level axial positioning accuracy and real-time processing.
It achieves nanometer-level axial positioning accuracy of ±10 nanometers and real-time processing of 120 frames per second, meeting the needs of in-situ monitoring of polymer materials and improving positioning accuracy and computing efficiency.
Smart Images

Figure CN120708219A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of deep learning, computer vision and optical measurement technology, and in particular to a method and system for accurately positioning the growth of a single polymer based on a hybrid network. Background Art
[0002] In situ dynamic monitoring of polymer materials at the nanoscale is a key challenge in fields such as materials science and biomedical engineering. Precise, real-time positioning of the growth height of single polymers is crucial for a deeper understanding of molecular self-assembly mechanisms and precise control of material properties. However, conventional optical microscopy is limited by the optical diffraction limit, with axial resolution typically reaching only a few hundred nanometers, making it difficult to meet the precision requirements for real-time monitoring of nanoscale growth processes. To overcome this limitation, the Depth from Defocus (DFD) method was proposed. This method inverts the height of an object by analyzing image blur, providing a potential path to super-resolution positioning.
[0003] Despite this, existing DFD technology still faces significant technical bottlenecks when applied to single polymer growth monitoring. First, polymer growth is a continuous dynamic process, and the static defocus model relied upon by traditional DFD is difficult to adapt to rapid deformation, resulting in the accumulation of motion artifacts and affecting positioning accuracy. Second, polymer chain surfaces often exhibit weak optical texture characteristics. In low-texture areas, the defocus difference signal is weak and difficult to effectively capture, and existing algorithms (such as gradient detection methods based on the Laplacian operator) are extremely sensitive to this type of noise. Third, traditional computational methods that rely on iterative optimization (such as blind deconvolution) are highly complex and cannot meet the real-time processing requirements of more than 100 fps required for in-situ monitoring.
[0004] In recent years, deep learning models (such as U-Net and ResNet) have been introduced into the DFD field to improve performance. These models still have obvious shortcomings when dealing with the specific scenario of single polymer growth. Existing end-to-end networks usually directly regress the defocus depth value, ignoring the inherent physical correlation between the focused image and the defocused image, resulting in the loss of high-frequency detail information and affecting nanometer-level precision. At the same time, due to its inherent local receptive field characteristics, convolutional neural networks (CNNs) have difficulty in effectively modeling the long-range spatial dependence of the polymer growth direction and are prone to directional positioning deviations. In addition, the multi-scale diffraction features generated during the growth process (such as primary diffraction spots and secondary interference rings) are prone to attenuation during transmission in deep networks. The existing cross-layer connection mechanism has failed to fully solve the problem of effective retention and fusion of cross-scale nanometer-level signals.
[0005] Therefore, there is an urgent need for an innovative solution that can effectively solve problems such as dynamic blur interference, poor low-texture robustness, low computational efficiency, insufficient feature coupling of existing deep learning models, lack of long-range dependencies, and cross-scale feature fractures. Summary of the Invention
[0006] The problem to be solved by the present invention is to provide a method and system for precise positioning of single polymer growth based on a hybrid network. Through an innovative design that integrates physical drive and data drive, a two-stage collaborative processing mechanism for focused image generation and defocus depth calculation is constructed to achieve the unification of nanoscale axial positioning accuracy and real-time processing.
[0007] The present invention adopts the following technical solution: a method for accurately positioning the growth of a single polymer based on a hybrid network, comprising the following steps: Pre-focusing stage: S101, obtaining the original defocused image DF_map of the single polymer growth, and obtaining the corresponding real focused image GT_map through the microscope focus adjustment structure; S102, training based on the PreFocusNet model, using the CS_Diff layer to obtain a structural similarity map DFOut_map between the original defocused image DF_map and the generated focused image Out_map, as well as a structural similarity map DFGT_map between the original defocused image DF_map and the true focused image GT_map; S103, calculating structural difference features based on the obtained structural similarity maps DFOut_map and DFGT_map; S104, optimizing the PreFocusNet model using the structural difference features to generate a high-precision focused image Out_map1; Positioning phase: S201, fix the PreFocusNet model parameters, and use the CS_Diff layer again to obtain the structural similarity map DFOut_map1 of the original defocused image DF_map and the high-precision focused image Out_map1; S202, fusing the original defocused image DF_map and the structural similarity map DFOut_map1 to form a dual-channel fusion feature map; S203: Input the dual-channel fusion feature map into the CNN-ViT hybrid model and perform three-level processing to obtain the nanometer-scale defocus depth value of the single polymer growth position.
[0008] Preferably, in step S103, pixel-by-pixel difference analysis is performed on DFOut_map and DFGT_map to extract structural difference features. , which is used to guide the PreFocusNet model to minimize the perceptual difference between the generated focused image Out_map and the true focused image GT_map during training.
[0009] Preferably, in step S104, a high-precision focused image is generated, and a perceptual loss is defined, by minimizing the loss Perform back propagation and update the PreFocusNet parameters until convergence to obtain a high-precision focusing model.
[0010] Preferably, in step S202, the fusion mode is channel stitching, and the original defocused image DF_map and the structural similarity map DFOut_map1 are channel stitched to generate a dual-channel fusion input with an input size of ,in, Represent the height and width of the fused feature map respectively.
[0011] Preferably, in step S203, the dual-channel fusion input is fed into the CNN-ViT hybrid model for three-level processing, as follows: First-level processing: The first part of the CNN-ViT hybrid model includes two sets of Double Conv layers and MaxPool layers to extract primary diffraction features. The convolution kernel size is 3×3. Each set of Conv is followed by BatchNorm and ReLU activation. The pooling kernel size is 2×2 with a stride of 2. Second-level processing: The second part of the CNN-ViT hybrid model combines the linear embedding layer and fast mapping mechanism of the visual Transformer module to analyze the long-range spatial correlation characteristics of the growth direction of a single polymer; Third-level processing: The third part of the CNN-ViT hybrid model introduces a cross-layer skip connection mechanism to fuse the intermediate features of the first-level processing and the second-level processing, and outputs the predicted nanoscale defocus depth value of the single polymer growth position.
[0012] Preferably, in the second stage of processing, the long-range spatial correlation characteristics of the growth direction of a single polymer are analyzed as follows: The linear embedding layer of the Transformer module is used to divide the input convolutional feature map into patches and flatten it into a one-dimensional vector. It is then linearly mapped to a d-dimensional space through a fully connected layer to obtain a token sequence. The token sequence is input into the attention module to construct query, key, and value vectors, and the attention matrix is calculated based on relative position weights. The attention outputs of different heads are spliced and linearly mapped through the multi-head self-attention mechanism to obtain an updated feature sequence with global spatial dependencies. The updated feature sequence is inversely transformed into a two-dimensional feature map format according to the Patch partitioning rule, and feature fusion is performed. A fast mapping mechanism is used to capture the long-range spatial correlation features of the growth direction of a single polymer to obtain a long-range spatial correlation feature map.
[0013] Preferably, in the third-level processing, the cross-layer jump connection mechanism directly connects the primary diffraction features output by the first-level processing to the input end of the third-level processing, and fuses them with the long-range spatial correlation features output by the second-level processing in the channel dimension, thereby enhancing feature reuse and gradient transfer, suppressing deep network feature attenuation, and obtaining the final fused feature map.
[0014] Preferably, the positioning method has an axial positioning accuracy of ±10 nanometers and a processing speed of 120 frames per second or above.
[0015] The technical solution of the present invention further provides: a single polymer growth precise positioning system based on a hybrid network, used to implement any of the above-mentioned single polymer growth precise positioning methods, comprising: an image acquisition module, a pre-focusing processing module, and a positioning processing module; The image acquisition module is used to obtain the original defocused image DF_map and the corresponding real focused image GT_map of the single polymer growth; The pre-focusing processing module is connected to the image acquisition module and is used to perform a pre-focusing operation, including: CS_Diff unit: extracts the structural similarity map DFOut_map between the original defocused image DF_map and the generated focused image Out_map, as well as the structural similarity map DFGT_map between the original defocused image DF_map and the true focused image GT_map, and calculates the structural difference features between the generated focused image and the true focused image; PreFocusNet unit: receives the original defocused image DF_map, optimizes it using the structural difference features calculated by the CS_Diff unit, and outputs a high-precision focused image Out_map1.
[0016] The positioning processing module is connected to the pre-focusing processing module and is used to perform defocus depth calculation, including: Structural similarity map generation unit: uses the CS_Diff unit to calculate the difference between the high-precision focused image Out_map1 and the original defocused image DF_map to generate a structural similarity map DFOut_map1; Fusion unit: used to fuse the original defocused image and the structural similarity map DFOut_map1 to form a dual-channel fusion feature map; CNN-ViT hybrid processing unit: used to receive the dual-channel fusion feature map, and perform three-level processing in sequence to output the nanometer-scale defocus depth value of the single polymer growth position.
[0017] Preferably, the CNN-ViT hybrid processing unit includes: The first-level processing subunit: extracts primary diffraction features through the Double Conv layer and the MaxPool layer; The second-level processing sub-unit combines the Linear Embedding layer and Faster Mapping mechanism of the ViT module to analyze the long-range spatial correlation characteristics of the growth direction; The third-level processing subunit introduces cross-layer jump connections, fuses features from the first-level processing subunit and the second-level processing subunit, enhances feature reuse, and outputs the nanometer-scale defocus depth value of the single polymer growth position.
[0018] Compared with the prior art, the present invention adopts the above technical solution and has the following technical effects: 1. Physics-data dual-driven architecture: The method of the present invention integrates the physical prior of SSIM into deep learning through the CS_Diff layer to enhance the robustness of low-texture scenes.
[0019] 2. Dual-channel information complementarity: The method of the present invention preserves the blur information and focus difference while fusing the original defocused image with the Diff-Map, breaking through the bottleneck of feature coupling.
[0020] 3. Multi-scale feature collaboration: The CNN-ViT hybrid model of the present invention combines local diffraction feature extraction, global long-range dependency modeling and cross-layer feature reuse to achieve full-scale capture of nanoscale signals.
[0021] 4. Real-time high-performance output: The processing flow of the method of the present invention adopts a lightweight design, which supports 120 fps real-time processing while ensuring ±10 nanometer axial positioning accuracy, meeting the stringent requirements of in-situ monitoring of polymer materials. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 A flow chart showing the precise positioning of single polymer growth according to the present invention; Figure 2 A diagram illustrating the processing of a training data set obtained according to an embodiment of the present invention; Figure 3 This is a model structure diagram of PreFocusNet of the present invention; Figure 4 A set of defocused-focused image pairs and a structural similarity graph according to an embodiment of the present invention; Figure 5 This is the structural diagram of the CNN-ViT hybrid model of the present invention; Figure 6 This is a schematic diagram of the test results of a simulated data set according to an embodiment of the present invention; Figure 7Schematic diagram of the test results of a single polymer growth dataset according to an embodiment of the present invention. DETAILED DESCRIPTION
[0023] In order to make the purpose, technical solutions and advantages of the present invention clearer, the technical solutions of the application are further elaborated in detail below with reference to the accompanying drawings. The described embodiments are only a part of the embodiments involved in the present invention. All non-innovative embodiments of other researchers in this field on this embodiment fall within the scope of protection of the present invention. At the same time, the step numbers in the embodiments of the present invention are only set for the convenience of explanation and description, and the order between the steps is not limited in any way. The execution order of each step in the embodiment can be adaptively adjusted according to the understanding of those skilled in the art.
[0024] Example 1
[0025] A single polymer growth precise positioning system based on hybrid network is proposed, e.g. Figure 1 As shown, it includes: an image acquisition module, a pre-focusing processing module and a positioning processing module, and realizes algorithm pipeline optimization through hardware acceleration.
[0026] An image acquisition module is configured to acquire a raw defocused image of single polymer growth.
[0027] The pre-focus processing module is connected to the image acquisition module and is configured to perform a pre-focus operation. The module integrates a CS_Diff unit and a PreFocusNet unit, which are used to extract a structural similarity graph and optimize the focused image, respectively.
[0028] Among them, the CS_Diff unit is configured to extract the structural similarity map DFOut_map between the original defocused image DF_map and the generated focused image Out_map, as well as the structural similarity map DFGT_map between DF_map and the true focused image GT_map, and based on this, calculate and generate the difference features between the focused image Out_map and the true focused image GT_map.
[0029] PreFocusNet unit: configured to receive the original defocused image, optimize it using the difference features calculated by the CS_Diff unit, and output a high-precision focused image.
[0030] Positioning processing module: connected to the pre-focusing processing module, configured to perform defocus depth calculation, including: a difference map generation unit, a dual-channel fusion unit, a CNN-ViT hybrid processing unit and a depth output unit.
[0031] Among them, the difference image generation unit is configured to use the CS_Diff unit to calculate the difference between the high-precision focused image generated by the PreFocusNet unit and the original defocused image DF_map, and generate a difference map DFOut_map1.
[0032] Fusion unit: configured to fuse the original defocused image and the difference map DFOut_map1 to form a dual-channel input.
[0033] CNN-ViT hybrid processing unit: configured to receive the dual-channel input, internally set up and sequentially execute the first-level sub-unit, the second-level sub-unit and the third-level sub-unit: The first-level processing subunit: extracts primary diffraction features through the Double Conv layer and the MaxPool layer; The second-level processing sub-unit combines the Linear Embedding layer and Faster Mapping mechanism of the ViT module to analyze the long-range spatial correlation characteristics of the growth direction; The third-level processing subunit: introduces cross-layer jump connections, introduces the output of the first-level processing subunit into its own input end through cross-layer jump connections, fuses the features from the first-level processing subunit and the second-level processing subunit, enhances feature reuse, and outputs the nanometer-scale defocus depth value of the single polymer growth position.
[0034] Example 2
[0035] A method for precise positioning of single polymer growth based on hybrid networks is proposed, and a two-stage processing flow is designed, including a pre-focusing stage and a positioning stage.
[0036] In this implementation, the experimental device and environment are as follows: an image acquisition system based on an optical microscope + CMOS camera is used, with a camera resolution of 1024×768 pixels and a pixel size of 3.45μm×3.45μm; the microscope objective lens uses a 60× high numerical aperture (NA=0.85) oil immersion objective lens.
[0037] In this embodiment, the pre-focusing stage is specifically processed as follows: Step S101: The image acquisition module uses a microscope focus adjustment structure and a piezoelectric ceramic (PZT) drive system to Scan along the axial direction (z direction) with the step size, such as Figure 2 As shown in the figure, the multi-plane defocused image sequence DF_map corresponding to the polystyrene microspheres in the single polymer growth area is obtained, which is expressed as: ; in, is the initial reference focal plane position, are the defocused images at different depths, Corresponding to the number of scanning layers within the range of ±200nm.
[0038] Afterwards, the defocused images at different depths , manually locate the center of a single polystyrene microsphere, cut the microsphere into small blocks of 128×128 pixels, and obtain 500 sets of sub-image blocks for subsequent network training: ; in, Represents the sequence of small blocks after cropping, ( ) represents the center coordinate of a single located microsphere, Indicates ( ) is used as the center to extract a 128×128 image block. express Corresponding to the width, height, and depth dimensions of the image.
[0039] Step S102: Cut the small block sequence Input PreFocusNet (pre-focus network) model for training. The structure of PreFocusNet is as follows Figure 3 shown.
[0040] In this embodiment, the PreFocusNet model uses an encoder-decoder architecture to process the cropped 128×128 convolution (ReLU activation) and 2×2 maximum pooling operations, gradually increasing the number of channels to 1024 and reducing the feature map size to 8×8; in the decoding stage, the size is restored to 128×128 through 2×2 upsampling and residual connections, and 1×1 convolution is used to refine the features. Finally, the output layer generates a high-dimensional feature map.
[0041] The PreFocusNet model training process uses the CS_Diff layer to calculate and generate the structural similarity map DFOut_map between the focused image Out_map and the DF sequence, as well as the structural similarity map DFGT_map between the real focused image GT_map and the DF sequence, such as Figure 4 The specific formula is as follows: ; ; in, Represents the original defocused image and generate focused images The structural similarity diagram of Represents the original defocused image With real focus image Structural similarity diagram.
[0042] Specifically, the CS_Diff layer is used to generate a structural similarity graph as follows: ; in, is the input image of the CS_Diff layer, express The covariance of and Respectively The variance of is a constant.
[0043] In this embodiment, preferably, represents a small constant usually, and , is used for numerical stability, and its function is to prevent the denominator from being zero.
[0044] Step S103: Perform pixel-by-pixel difference analysis on DFOut_map and DFGT_map to extract structural difference features: ;
[0045] Step S104: define the perceptual loss as: ; in, represents the pre-trained feature map of the PreFocusNet model, 、 They represent the generated focused image and the real focused image respectively.
[0046] By minimizing losses Perform back propagation and update the PreFocusNet parameters until convergence to obtain a high-precision focusing model.
[0047] Furthermore, in this embodiment, the positioning stage is specifically processed as follows: Step S201: Freeze the parameters of the trained PreFocusNet model, and use the CS_Diff layer again to calculate the structural similarity graph DFOut_map between DF_map and Out_map.
[0048] Step S202: Concatenate the original defocused image DF_map and the DFOut_map to generate a dual-channel fusion feature map. The input image size is 2×H×W, where: Represent the height and width of the fused feature map respectively.
[0049] Step S203: Input the dual-channel fusion feature map into the CNN-ViT hybrid model for three-level processing. The overall structure is as follows: Figure 5 shown.
[0050] Step S203a (first-level processing): The first part of the CNN-ViT hybrid model consists of two sets of double convolutional layers (DoubleConv) and a maximum pooling layer (MaxPool), which are used to extract primary diffraction features.
[0051] In this embodiment, preferably, the convolution kernel size is 3×3, each group of Conv is followed by BatchNorm and ReLU activation, the pooling kernel size is 2×2, and the stride is 2.
[0052] The initial input size is H=W=128, and the size is halved after each pooling. The output feature map after three rounds of pooling is: .
[0053] Preferably, =64, It is the primary diffraction feature map output by the first-level processing.
[0054] Step S203b (second-level processing): Combining the linear embedding layer (LinearEmbedding) and the fast mapping mechanism (Faster Mapping) of the Visual Transformer (ViT) module to analyze the long-range spatial correlation characteristics of the polymer growth direction and overcome the modeling limitations of the local receptive field of traditional CNN.
[0055] In this embodiment, the dual-channel fusion feature map of step S202 is linearly mapped into a ViT Token sequence through the Linear Embedding layer, and a Faster Mapping mechanism (i.e., multi-head self-attention combined with relative position encoding) is used to capture the long-range spatial correlation features of the growth direction of a single polymer.
[0056] Specifically, in order to adapt the feature map to the Transformer module, the convolutional feature map is first patch-divided and flattened through the Linear Embedding layer, and linearly mapped into a token sequence.
[0057] The input feature map is divided into several non-overlapping patch areas. Assuming the patch size is P×P, for example, P=4, it can be divided into Patches; flatten each patch into a one-dimensional vector and linearly map it to the d-dimensional space through a fully connected layer to obtain a token sequence .
[0058] In an optional solution, the embedding dimension d is set to 64.
[0059] To enhance the expressiveness of features in the global spatial dimension, the token sequence is further input into an improved Transformer module, which includes a multi-head self-attention mechanism for modeling spatial dependencies in parallel in different subspaces and a relative positional encoding mechanism to enhance the model's perception of spatial structure.
[0060] Preferably, the processing of the attention module includes: constructing query, key and value vectors; calculating the attention matrix based on relative position weights; concatenating and linearly mapping the attention outputs of different heads. Through this mechanism, an updated feature sequence containing global spatial dependencies can be obtained. .
[0061] The updated feature sequence It can be inversely transformed into a two-dimensional feature map format according to the Patch division rule for subsequent feature fusion processing. That is, it can be reconstructed as: ; in, Represents the output feature map that incorporates long-range spatial dependency information.
[0062] Step S203c (third-level processing): A cross-layer skip connection mechanism is introduced to connect the primary diffraction features output by the first-level processing directly to the third-level processing input. The features are then fused with the long-range correlation features output by the second-level processing in the channel dimension to enhance feature reuse and gradient transfer, effectively suppressing the feature attenuation problem in the deep network. Finally, a fused feature map is obtained: ; in, The operator merges two feature maps along the channel dimension (dimension 0).
[0063] Step S204: The final output of the CNN-ViT hybrid model is output through the fully connected layer to obtain the nanometer-scale defocus depth value of the single polymer growth position, with a positioning accuracy of up to ±10 nanometers.
[0064] Furthermore, based on the method of the present invention, a data set was built in the laboratory. After 100 epochs of training, the test frame rate after convergence can reach more than 120 frames per second (fps), meeting the real-time requirements. Figure 6 As shown in Figure 2, the axial positioning accuracy evaluation results obtained using the simulated data set test show that the average error is less than 8 nanometers and the standard deviation is less than 5 nanometers. The trained model is applied to in-situ growth, as shown in Figure 2. Figure 7As shown, the depolymerization and polymerization processes during the growth process can be accurately reproduced.
[0065] In summary, the hybrid network-based precise positioning method and system for single polymer growth proposed in the present invention innovatively integrates the two-stage collaborative processing flow of focused image generation and defocused depth calculation. At the same time, combined with the physically inspired CS_Diff layer design and the multi-level feature resolution mechanism of the CNN-ViT hybrid model, it can achieve a breakthrough balance between nanoscale axial positioning accuracy (±10 nm) and high frame rate real-time processing (120 fps), thereby providing strong technical support for the in situ dynamic monitoring of polymer materials.
[0066] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A method for precise positioning of single polymer growth based on a hybrid network, characterized in that: The steps include: Pre-focusing stage: S101, obtaining the original defocused image DF_map of the single polymer growth, and obtaining the corresponding real focused image GT_map through the microscope focus adjustment structure; S102, training based on the PreFocusNet model, using the CS_Diff layer to obtain a structural similarity map DFOut_map between the original defocused image DF_map and the generated focused image Out_map, as well as a structural similarity map DFGT_map between the original defocused image DF_map and the true focused image GT_map; S103, calculating structural difference features based on the obtained structural similarity maps DFOut_map and DFGT_map; S104, optimizing the PreFocusNet model using the structural difference features to generate a high-precision focused image Out_map1; Positioning phase: S201, fix the PreFocusNet model parameters, and use the CS_Diff layer again to obtain the structural similarity map DFOut_map1 of the original defocused image DF_map and the high-precision focused image Out_map1; S202, fusing the original defocused image DF_map and the structural similarity map DFOut_map1 to form a dual-channel fusion feature map; S203: Input the dual-channel fusion feature map into the CNN-ViT hybrid model and perform three-level processing to obtain the nanometer-scale defocus depth value of the single polymer growth position.
2. The method for precise positioning of single polymer growth based on hybrid networks according to claim 1, characterized in that: In step S102, a structural similarity graph is generated using the CS_Diff layer as follows: ; in, is the input image of the CS_Diff layer, express covariance of and Respectively The variance of is a constant; ; ; in, Represents the original defocused image and generate focused images The structural similarity diagram of Represents the original defocused image With real focus image Structural similarity diagram.
3. The method for precise positioning of single polymer growth based on hybrid networks according to claim 2, characterized in that: In step S103, pixel-by-pixel difference analysis is performed on DFOut_map and DFGT_map to extract structural difference features, which are calculated as follows: ; in, The obtained structural difference features are used to guide the PreFocusNet model to minimize the perceptual difference between the generated focused image Out_map and the true focused image GT_map during training.
4. The method for precise positioning of single polymer growth based on hybrid networks according to claim 1, characterized in that: In step S104, a high-precision focused image is generated as follows: Define perceptual loss, expressed as: ; in, represents the pre-trained feature map of the PreFocusNet model, 、 Represent the generated focused image and the real focused image respectively; By minimizing losses Perform back propagation and update the PreFocusNet parameters until convergence to obtain a high-precision focusing model.
5. The method for precise positioning of single polymer growth based on hybrid network according to claim 1, characterized in that: In step S202, the fusion method is channel stitching, the original defocused image DF_map and the structural similarity map DFOut_map1 are channel stitched to generate a dual-channel fusion input with an input size of ,in, Represent the height and width of the fused feature map respectively.
6. The method for precise positioning of single polymer growth based on hybrid networks according to claim 5, characterized in that: In step S203, the dual-channel fusion input is fed into the CNN-ViT hybrid model for three-level processing as follows: First-level processing: The first part of the CNN-ViT hybrid model includes two sets of Double Conv layers and MaxPool layers, which are used to extract primary diffraction features. The convolution kernel size is 3×3. Each set of Conv is followed by BatchNorm and ReLU activation. The pooling kernel size is 2×2 with a stride of 2, and the output is the primary diffraction feature map. Second-level processing: The second part of the CNN-ViT hybrid model combines the linear embedding layer and fast mapping mechanism of the visual Transformer module to analyze the long-range spatial correlation characteristics of the growth direction of a single polymer; Third-level processing: The third part of the CNN-ViT hybrid model introduces a cross-layer skip connection mechanism to fuse the intermediate features of the first-level processing and the second-level processing, and outputs the predicted nanoscale defocus depth value of the single polymer growth position.
7. The method for precise positioning of single polymer growth based on hybrid networks according to claim 5, characterized in that: In the second stage of processing, the long-range spatial correlation characteristics of the growth direction of a single polymer are analyzed as follows: The linear embedding layer of the Transformer module is used to divide the input convolutional feature map into patches and flatten it into a one-dimensional vector. It is then linearly mapped to a d-dimensional space through a fully connected layer to obtain a token sequence. The token sequence is input into the attention module to construct query, key, and value vectors, and the attention matrix is calculated based on relative position weights. The attention outputs of different heads are spliced and linearly mapped through the multi-head self-attention mechanism to obtain an updated feature sequence with global spatial dependencies. The updated feature sequence is inversely transformed into a two-dimensional feature map format according to the Patch partitioning rule, and feature fusion is performed. A fast mapping mechanism is used to capture the long-range spatial correlation features of the growth direction of a single polymer to obtain a long-range spatial correlation feature map.
8. The method for precise positioning of single polymer growth based on hybrid networks according to claim 7, characterized in that: In the third-level processing, the cross-layer jump connection mechanism directly connects the primary diffraction features output by the first-level processing to the input end of the third-level processing, and fuses them with the long-range spatial correlation features output by the second-level processing in the channel dimension, thereby enhancing feature reuse and gradient transfer, suppressing the attenuation of deep network features, and obtaining the final fused feature map.
9. The method for precise positioning of single polymer growth based on hybrid networks according to claim 5, characterized in that: The CNN-ViT hybrid model outputs the predicted nanometer-scale defocus depth value of the single polymer growth position through a fully connected layer, with an axial positioning accuracy of ±10 nanometers and a processing speed of 120 frames per second or more.
10. A single polymer growth precise positioning system based on a hybrid network, used to implement the single polymer growth precise positioning method according to any one of claims 1 to 8, characterized in that: include: Image acquisition module, pre-focusing processing module, positioning processing module; The image acquisition module is used to obtain the original defocused image DF_map and the corresponding real focused image GT_map of the single polymer growth; The pre-focusing processing module is connected to the image acquisition module and is used to perform a pre-focusing operation, including: CS_Diff unit: extracts the structural similarity map DFOut_map between the original defocused image DF_map and the generated focused image Out_map, as well as the structural similarity map DFGT_map between the original defocused image DF_map and the true focused image GT_map, and calculates the structural difference features between the generated focused image and the true focused image; PreFocusNet unit: receives the original defocused image DF_map, optimizes it using the structural difference features calculated by the CS_Diff unit, and outputs a high-precision focused image Out_map1; The positioning processing module is connected to the pre-focusing processing module and is used to perform defocus depth calculation, including: Structural similarity map generation unit: uses the CS_Diff unit to calculate the difference between the high-precision focused image Out_map1 and the original defocused image DF_map to generate a structural similarity map DFOut_map1; Fusion unit: used to fuse the original defocused image and the structural similarity map DFOut_map1 to form a dual-channel fusion feature map; CNN-ViT hybrid processing unit: used to receive the dual-channel fusion feature map and perform three-level processing in sequence, including: The first-level processing subunit: extracts primary diffraction features through the Double Conv layer and the MaxPool layer; The second-level processing sub-unit combines the Linear Embedding layer and Faster Mapping mechanism of the ViT module to analyze the long-range spatial correlation characteristics of the growth direction; The third-level processing subunit introduces cross-layer jump connections, fuses features from the first-level processing subunit and the second-level processing subunit, enhances feature reuse, and outputs the nanometer-scale defocus depth value of the single polymer growth position.
Citation Information
Patent Citations
Multi-focus image fusion method based on multi-scale feature interaction network
CN113705675A
Multi-focal-length image fusion model, method and device for compound eye camera
CN115439376A
Multi-focus image fusion method and system based on structural similarity and region segmentation
CN115829895A
Method for fusing infrared light and visible light images
WO2025103079A1