A full-automatic coupling method of micro-optical devices based on multi-source visual feedback
Patent Information
- Application Number
- CN202611327975.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-31
- Publication Date
- 2026-09-29
AI Technical Summary
[0004]然而,在实际耦合应用场景中,系统需实时处理高分辨率微距图像流,导致视觉Transformer等深度学习模型内部的计算负荷极高,严重制约了耦合系统的实时响应能力
本发明通过获取微光学器件耦合区域实时图像,对图像Token进行高频特征显著性分析,并结合视觉Transformer模型提取和解析图像特征,提高了耦合光斑定位的准确性。在模型计算过程中,利用归一化全局注意力熵自适应确定空间先验权重的衰减率,配合调制注意力分数与高频显著性对Token进行双阈值筛选,将非关键Token加权融合至保留Token中,从而在降低模型冗余计算、提升图像处理实时响应效率的同时,保留了关键空间位置信息与局部高频细节特征。
Smart Images

Figure CN122841508A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of visual feedback technology. More specifically, this invention relates to a fully automated coupling method for micro-optical devices based on multi-source visual feedback. Background Technology
[0002] In the packaging and assembly manufacturing process of micro-optical devices, the optical coupling alignment between devices directly affects the optical transmission efficiency and overall performance of the product. Since the mode field size of micro-optical devices is usually on the order of micrometers or even submicrometers, it places extremely high demands on the accuracy and efficiency of automated coupling alignment.
[0003] Currently, traditional automated coupling methods mostly employ purely mechanical blind scanning of optical power to find the optimal coupling position. This method is not only time-consuming and inefficient, but also prone to getting trapped in local optima in complex mode field distributions. Chinese patent document CN110058355B, entitled "An Automatic Coupling Device and Automatic Coupling Method," discloses a scheme that utilizes a lens observation system to provide visual support and employs image acquisition and mechanical adjustment components to observe channel alignment. This method, which combines machine vision to acquire real-time images of the coupling area with optical power feedback to achieve closed-loop control, is expected to efficiently complete image feature extraction and pose prediction during the visual guidance stage using deep learning models such as the Visual Transformer, thereby significantly shortening the overall alignment time.
[0004] However, in real-world coupled applications, the system needs to process high-resolution macro image streams in real time, resulting in extremely high computational loads within deep learning models such as the Visual Transformer, severely limiting the real-time response capability of the coupled system. If conventional feature dimensionality reduction or sequence simplification strategies are directly adopted to improve computational speed, it is difficult to effectively distinguish and balance key local high-frequency details and spatial distribution prior information in the image. This can easily weaken or even lose the representational ability of the true geometric center of the coupled spot during feature compression, leading to serious deviations in pose prediction results and failing to meet the stringent requirements of high-precision alignment. Summary of the Invention
[0005] To address the technical challenges of computationally intensive visual models and the loss of key light spot features due to compression, this invention proposes a fully automated coupling method for micro-optical devices based on multi-source visual feedback. This method can balance real-time system response and achieve high-precision pose prediction and alignment.
[0006] In a first aspect, the present invention provides a fully automated coupling method for micro-optical devices based on multi-source visual feedback, comprising: S1, acquiring a real-time image of the coupling region, dividing it into image tokens, obtaining an initial token sequence through linear embedding and positional encoding, calculating the saliency of high-frequency features of the image tokens, and initializing the predicted geometric center of the coupling spot; S2, iteratively calculating the initial attention score matrix and the normalized global attention entropy of the current layer, and determining the decay rate of the spatial prior weight vector based on the updated predicted geometric center of the coupling spot; performing a Hadamard product operation on the initial attention score matrix and the spatial prior weight vector and re-normalizing it row-wise to obtain a modulation attention score matrix, multiplying it with the value matrix to generate the output features of the current layer, and extracting class token pairs. S3, the attention assignment value of each image token is used as the modulation attention score; S4, the image tokens with modulation attention scores not lower than the first preset threshold, or lower than the first preset threshold but with high frequency feature saliency higher than the second preset threshold are marked as retained tokens, and the rest are marked as tokens to be merged. They are weighted and fused into the retained token with the closest Euclidean distance through a shared learnable feature transformation network. The retained tokens maintain the original spatial center coordinates and high frequency feature saliency; S5, the pose adjustment amount of the token-like output of each layer is used to update the predicted geometric center of the coupled light spot. The position deviation between the predicted geometric center of the coupled light spot and the preset ideal alignment center is generated by a pre-calibrated mapping to generate a control signal for combining optical power feedback to drive the actuator.
[0007] By adopting the above technical solution, the real-time image of the coupling region is divided into image tokens and high-frequency feature saliency analysis is performed. At the same time, the global context features are extracted by combining the visual Transformer model, which improves the accuracy of coupling spot localization. The decay rate of spatial prior weights is adaptively determined by normalized global attention entropy, and non-critical tokens are weighted and fused into retained tokens by a dual threshold screening mechanism of modulated attention score and high-frequency saliency. This effectively reduces redundant model calculations, improves real-time response efficiency, and preserves key spatial location information and local high-frequency detail features. Furthermore, the pose adjustment amount is predicted layer by layer by each type of token and the predicted geometric center of the coupling spot is updated. The deviation between the spot center and the ideal alignment center is mapped into the control signal of the actuator. Combined with optical power feedback, closed-loop precise alignment is achieved. The synergistic fusion of visual feature guidance and optical power feedback is realized, which significantly improves the overall alignment efficiency and device assembly yield of fully automatic coupling of micro-optical devices.
[0008] Preferably, the step of acquiring a real-time image of the coupling region, dividing it into image tokens, and obtaining an initial token sequence through linear embedding and positional encoding includes: cropping the real-time image into multiple non-overlapping two-dimensional image blocks according to a preset window size; flattening each two-dimensional image block into a one-dimensional vector in the spatial dimension, and projecting the one-dimensional vector onto a preset feature dimension through a fully connected layer to obtain an image block embedding vector; generating a two-dimensional absolute positional encoding corresponding to the spatial position of the two-dimensional image block, and adding the two-dimensional absolute positional encoding to the corresponding image block embedding vector, and concatenating the tokens at the beginning of the sequence to generate the initial token sequence.
[0009] By adopting the above technical solution, the generation method of the initial token sequence is further restricted. Through non-overlapping clipping of preset windows, spatial dimension flattening, and linear projection of fully connected layers, two-dimensional image patches are converted into fixed-dimensional embedding vectors, and absolute position codes corresponding to spatial locations are superimposed. While reducing dimensionality, the spatial topological distribution information of the image patches is completely preserved. After concatenating learnable tokens at the beginning of the sequence, the initial token sequence is generated, which enables the model to effectively aggregate global context information and provide an ordered and structured input representation for subsequent self-attention feature interactions, thereby ensuring the accuracy of feature extraction and the stability of the subsequent localization process.
[0010] Preferably, the iterative calculation of the initial attention score matrix and normalized global attention entropy of the current layer includes: obtaining the query matrix, key matrix, and value matrix of the current token sequence through linear projection; calculating the matrix product of the transpose of the query matrix and the key matrix, dividing by a scaling factor, and normalizing it using the Softmax function to obtain the initial attention score matrix; calculating the information entropy of the attention distribution corresponding to each query token in the initial attention score matrix, normalizing the information entropy corresponding to each query token according to the number of key tokens in the initial attention score matrix of the current layer, averaging the normalized information entropy, and obtaining the normalized global attention entropy that represents the discreteness of the model's attention and does not drift independently with the change in the number of tokens in the current layer.
[0011] By adopting the above technical solution, the calculation methods of the initial attention score matrix and the normalized global attention entropy are further restricted. The query matrix, key matrix, and value matrix are obtained through linear projection. After obtaining the initial attention score matrix through scaling dot product and Softmax normalization, the information entropy of the attention distribution of each query token is calculated. Based on the number of key tokens in the current layer, normalization and averaging are performed to obtain the normalized global attention entropy. This index can stably characterize the dispersion of the model's attention, and its value does not drift independently with the change of the number of tokens in the current layer. It provides a reliable adaptive control signal for the subsequent dynamic adjustment of the decay of spatial prior weights, which is conducive to the rational allocation of attention resources under different network depths and token numbers.
[0012] Preferably, determining the decay rate of the spatial prior weight vector based on the updated predicted geometric center of the coupled spot includes: establishing a two-dimensional Gaussian distribution function with the coordinates of the predicted geometric center of the coupled spot as the mean; establishing a linear mapping relationship between the normalized global attention entropy and the variance of the two-dimensional Gaussian distribution function, using the calculated variance as a parameter to control the radial decay of the spatial prior weight vector, with a larger variance resulting in a smaller decay rate; substituting the spatial center coordinates of each image token into the two-dimensional Gaussian distribution function to calculate the probability density value, and after normalization and supplementing the class token with preset prior values, constructing the spatial prior weight vector.
[0013] By adopting the above technical solution, the determination method of the spatial prior weight vector decay rate is further restricted. By establishing a two-dimensional Gaussian distribution with the predicted geometric center of the coupled light spot as the mean, and establishing a direct proportional mapping between the normalized global attention entropy and the distribution variance, adaptive adjustment of the radial decay rate of the spatial prior weights is achieved: when the model attention is relatively dispersed, the variance automatically increases and the decay rate decreases, so that the spatial prior maintains a wide range of attention; when the attention is concentrated, the variance decreases and the decay rate increases, so that the spatial prior focuses on the neighborhood of the light spot center. The probability density value of the image token is calculated and normalized, and a preset prior value is added to the token class. The constructed spatial prior weight vector can be effectively combined with the initial attention score, guiding the model to dynamically focus on the key spatial region of the real coupled light spot, thereby improving the accuracy of pose prediction.
[0014] Preferably, the calculation of the high-frequency feature saliency of the image token includes: using a high-pass filter to filter the acquired real-time image to extract a high-frequency component image; calculating the sum of squares of pixel values in the image block region corresponding to the initial spatial position of each image token in the high-frequency component image, as the high-frequency feature saliency representing the edge and detail richness of the image token.
[0015] By adopting the above technical solution, the calculation method of high-frequency feature saliency of image tokens is further restricted. High-frequency component images are extracted by high-pass filtering, and the sum of squares of pixel values within the image block corresponding to each image token is calculated as the high-frequency feature saliency. This can objectively quantify the richness of edges, textures and optical details contained in the token. This saliency index provides an effective supplementary basis for subsequent token retention and merging decisions, ensuring that even if some tokens have low modulation attention scores, as long as they have prominent high-frequency structural features, they can still be identified as key tokens and retained. This prevents the loss of important local information such as spot contours during the fusion compression process, which is conducive to improving the reliability of coupled spot geometric center localization and pose estimation.
[0016] Preferably, the remaining tokens are to be merged, and are weighted and fused into the retained token with the smallest Euclidean distance through a shared learnable feature transformation network. The retained token maintains its original spatial center coordinates and high-frequency feature saliency. This includes: calculating the Euclidean distance between the feature vector of each to be merged token and the feature vectors of all retained tokens in the feature dimension; for each to be merged token, selecting the retained token with the smallest Euclidean distance as the target fusion node; inputting the feature vector of the to be merged token into a shared learnable feature transformation network composed of multilayer perceptrons, and outputting the transformed feature vector after dimensionality reduction and then dimensionality increase; performing a weighted summation of the transformed feature vector and the feature vector of the corresponding target fusion node to complete the feature update, while maintaining the spatial center coordinates and high-frequency feature saliency of the target fusion node unchanged.
[0017] Preferably, when the high-frequency feature significance of the target fusion node is higher than the second preset threshold, the weighted fusion further includes: determining whether the feature Euclidean distance between the token to be merged and the target fusion node is less than a preset fusion distance threshold, and whether the spatial center coordinate distance between the two is less than a preset spatial proximity threshold; if both are yes, then the weighted fusion is performed; if at least one is no, then the token to be merged is marked as a reserved token and passed to the next processing level.
[0018] Preferably, the step of updating the predicted coupled spot geometric center using the token-like prediction pose adjustment amount output by each layer includes: inputting the token-like prediction amount output by the current processing layer into the multilayer perceptron regression prediction head, and outputting a two-dimensional pose adjustment amount including lateral and longitudinal offsets; adding the lateral and longitudinal offsets of the predicted coupled spot geometric center used by the current processing layer to the lateral and longitudinal offsets of the two-dimensional pose adjustment amount, respectively, to obtain the updated geometric center coordinates; and passing the updated geometric center coordinates to the next processing layer as the reference point for constructing the spatial prior weight vector of the next layer.
[0019] Preferably, the generation of control signals for driving the actuator in conjunction with optical power feedback includes: acquiring optical power values in real time during the displacement adjustment process of the actuator, and iteratively optimizing the control signals based on the optical power values using an optimization algorithm; when any of the following preset convergence conditions are met, alignment is determined to be complete: in multiple consecutive iterations, the increase in optical power is less than a preset increase threshold; the displacement adjustment step size of the actuator is less than a preset displacement threshold; the real-time optical power reaches a preset qualified power threshold; and the number of iterations reaches a preset maximum number of iterations.
[0020] Preferably, the initialization of the predicted geometric center of the coupled spot includes: thresholding the single-channel grayscale image corresponding to the real-time image to generate a binary mask, and marking the connected components of the binary mask; calculating the sum of gray levels or the average brightness of the regions corresponding to each connected component, and selecting the connected component with the highest sum of gray levels or average brightness as the candidate coupled spot region; calculating the geometric centroid coordinates of the candidate coupled spot region using the image spatial moment algorithm, and using the geometric centroid coordinates as the initialized predicted geometric center of the coupled spot.
[0021] The present invention has the following beneficial effects: This invention improves the accuracy of coupled spot localization by acquiring real-time images of the coupling region of micro-optical devices, performing high-frequency feature saliency analysis on image tokens, and combining this with a visual Transformer model to extract and analyze image features. During model computation, the attenuation rate of spatial prior weights is adaptively determined using normalized global attention entropy. This, combined with modulation attention scores and high-frequency saliency, performs dual-threshold screening of tokens, weighting and fusing non-critical tokens into the retained tokens. This reduces redundant model computation, improves real-time image processing efficiency, and preserves key spatial location information and local high-frequency detail features.
[0022] Furthermore, this invention utilizes each layer of tokens to predict pose adjustment amounts layer by layer and updates the predicted geometric center of the coupled light spot in real time. The positional deviation between the updated light spot center and the ideal alignment center is mapped into a control signal for the actuator. Combined with optical power feedback, the actuator is driven to complete closed-loop precise alignment. This achieves the synergistic fusion of visual feature guidance information and optical power feedback information, improving the overall alignment efficiency of fully automatic coupling of micro-optical devices and the device assembly yield. Attached Figure Description
[0023] Figure 1 This is a flowchart of a fully automated coupling method for micro-optical devices based on multi-source visual feedback; Figure 2 It is a layer-by-layer pose update tracking map; Figure 3These are comparison images from ablation experiments. Detailed Implementation
[0024] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.
[0025] In this application, a fully automated coupling method for micro-optical devices based on multi-source visual feedback is disclosed, such as... Figure 1 As shown, the method includes: S1. Acquire real-time images of the coupling region of the micro-optical device, divide them into non-overlapping image tokens, perform linear embedding and position encoding to obtain the initial token sequence, calculate the high-frequency feature saliency of the image tokens, and initialize the predicted geometric center of the coupling spot.
[0026] The camera driver interface function calls the image acquisition device to obtain real-time RGB images of the coupled region, and synchronously converts the RGB real-time images into single-channel grayscale images. The RGB real-time images are used for subsequent image token embedding and visual Transformer feature extraction, while the single-channel grayscale images are used for high-frequency feature saliency calculation and initial spot geometric center extraction. Image patch embedding is performed using a 2D convolutional layer (Conv2d) of a deep learning framework, or by flattening image patches and then using a linear transformation layer (Linear). The kernel size and stride of the 2D convolutional layer (Conv2d) are set to the preset image patch size to achieve non-overlapping image token partitioning and linear embedding. Simultaneously, learnable absolute position embedding parameters consistent with the sequence length are generated, or positional features are generated using a sine / cosine position encoding algorithm and added element-wise to the embedding vector. During model training initialization, a learnable class token parameter is randomly initialized and updated through backpropagation during training. During model inference, the trained and fixed class token parameter is called and concatenated to the beginning of the image token sequence using the tensor concatenation function `cat` to obtain the initial token sequence.
[0027] For the calculation of high-frequency feature saliency, it is preferable to use a spatial domain high-pass filter to extract the high-frequency components of the grayscale image. When using the Fast Fourier Transform (FFT2) algorithm for frequency domain extraction, first perform FFT2 on the grayscale image or each image block and then perform spectrum centering. Subsequently, set the cutoff radius with the spectrum center as the origin. , radius greater than or equal to The frequency domain region is defined as the high-frequency region, and a high-frequency mask is constructed. , then calculate The sum of normalized absolute values is used as a scalar of high-frequency feature saliency. A binary mask is generated from the single-channel grayscale image using the Otsu thresholding algorithm, and connected component labeling is applied to the binary mask. Then, based on the original single-channel grayscale image, the sum of grayscale values or average brightness within each connected component is calculated. The connected component with the highest sum of grayscale values or average brightness is selected as the candidate coupled spot region. The two-dimensional coordinates of the geometric centroid of this candidate coupled spot region are calculated using the image spatial moment algorithm (moments), and these coordinates are used to initialize the predicted geometric center of the coupled spot.
[0028] In some embodiments, acquiring a real-time image of the coupling region of a micro-optical device, dividing it into non-overlapping image tokens, and performing linear embedding and position encoding to obtain an initial token sequence includes: cropping the real-time image into multiple non-overlapping two-dimensional image blocks according to a preset window size; flattening each two-dimensional image block into a one-dimensional vector in the spatial dimension, and projecting the one-dimensional vector onto a preset feature dimension through a fully connected layer to obtain an image block embedding vector; generating a two-dimensional absolute position code corresponding to the spatial position of the two-dimensional image block, and adding the two-dimensional absolute position code to the corresponding image block embedding vector, and concatenating the tokens at the beginning of the sequence to generate an initial token sequence.
[0029] Real-time RGB images of the coupling region of micro-optical devices are acquired using a high-resolution CMOS camera. The preset input resolution of the image is set to H×W×C, for example, with a size of 224×224×3, where C is the number of color channels. Before the image enters the visual Transformer, a 224×224×3 RGB tensor is maintained as the model input. Simultaneously, a 224×224×1 single-channel grayscale image is generated according to the grayscale conversion relationship Y=0.299R+0.587G+0.114B. This grayscale image is used for high-frequency saliency calculation, threshold segmentation, and geometric centroid initialization, avoiding the mixing of RGB channel data and grayscale processing.
[0030] The image cropping module, according to a preset window size P×P and with the preferred parameter set to 16×16 pixels, cuts the real-time image into N two-dimensional image blocks without overlap. The number of image blocks is... That is, output 196 image blocks.
[0031] For each image patch, a spatial dimension flattening operation is performed, unfolding the 16×16×3 three-dimensional image patch row by row into a one-dimensional feature vector of length 768. These flattened one-dimensional vectors are then input in parallel into a linear projection module for computation. This module is a neural network model with a network structure consisting of multiple fully connected layers. The input is a one-dimensional feature vector of length 768, and the output is an image patch embedding vector of dimension 1×768. The input vector is linearly mapped to a uniform preset feature dimension D via a weight matrix, preferably set to 768 dimensions, resulting in 196 image patch embedding vectors, each with a dimension of 1×768.
[0032] Based on this operation, an absolute position encoding generation mechanism based on sine and cosine trigonometric functions is adopted to generate a two-dimensional absolute position encoding vector with the same dimensions of 768, which is aligned with the spatial coordinates of each of the above image blocks, such as the region in the i-th row and j-th column of the original image. This position encoding vector is then added element by element to the corresponding image block embedding vector to preserve the spatial topological distribution features of the original image after dimensionality reduction.
[0033] This token with consistent feature dimensions is concatenated at the beginning of the sequence consisting of 196 embedded vectors with spatial location information, with index position 0, to construct an initial token sequence with a total dimension of 197×768. This initial token sequence is then passed as an input tensor to the visual Transformer network to perform feature extraction and interactive fusion.
[0034] In some embodiments, calculating the high-frequency feature saliency of an image token includes: using a high-pass filter to filter the acquired real-time image to extract a high-frequency component image; and calculating the sum of squares of pixel values in the image block region corresponding to the initial spatial position of each image token in the high-frequency component image as the high-frequency feature saliency representing the edge and detail richness of the image token.
[0035] In the front-end visual preprocessing channel, a spatial domain two-dimensional high-pass filter algorithm is invoked. A 3×3 Laplacian convolution operator with isotropic edge detection characteristics is preferred, with the center weight coefficient set to 8 and the weight coefficients of the surrounding eight neighborhoods uniformly set to -1. The control processor uses this convolution operator to perform discrete two-dimensional convolution filtering on a real-time image of a single-channel grayscale format with a resolution of 224×224. This operation can suppress and filter out low-frequency flat components in the image background and enhance the high-frequency signals of edge contours where grayscale transitions occur. The absolute value or squared response of the Laplacian convolution operation output is first taken to obtain a non-negative high-frequency response matrix. Then, the non-negative high-frequency response matrix is normalized to map its values to the range of 0 to 255, resulting in a high-frequency component grayscale image of the same resolution size of 224×224.
[0036] After caching the entire high-frequency component image, the process transitions to a token-level local feature quantization stage. Based on a sliding window segmentation mechanism with a fixed step size of 16 pixels, 196 pre-defined image tokens are iterated. For any image token during the traversal, for example, when it points to the first base image block that overlaps with the original image's horizontal and vertical coordinate ranges of 0 to 15, a 16×16 pixel sub-region matrix corresponding to the token's coordinate range is extracted from the aforementioned high-frequency component image through a slicing operation.
[0037] The processor extracts the high-frequency response grayscale values of each of the 256 pixels one by one. Where m and n iterate from 1 to 16, perform power-square operations on each grayscale value, and then sum them up. The calculation formula is expressed as follows: The total sum of squares scalar S calculated in this stage is set as a numerical metric representing the significance of high-frequency features in the current image Token, including edge texture, structural abrupt changes, and optical detail density.
[0038] The high-frequency saliency values corresponding to each image token obtained from the initial segmentation are compiled into the feature index table. When the size of the first-layer input image is 224×224 and the size of the image block is 16×16, there are 196 high-frequency saliency values. In order to prevent low-dimensional redundant background information from being fused with tokens with high saliency values during the fusion judgment, the fidelity of micro-optical device contour recognition is maintained.
[0039] S2, in the visual Transformer model, iteratively calculate the initial attention score matrix and normalized global attention entropy of the current layer. Based on the normalized global attention entropy, determine the decay rate of the spatial prior weight vector with the updated predicted coupled spot geometric center as the reference. Perform Hadamard product operation on the initial attention score matrix and the spatial prior weight vector and renormalize by row to obtain the modulation attention score matrix. Use the modulation attention score matrix to multiply with the value matrix to generate the output features of the current layer, and extract the attention assignment value of class token to each image token as the modulation attention score.
[0040] In the self-attention computation module of the visual Transformer model, three fully connected linear layers are used to generate query tensors, key tensors, and value tensors. The inner product of the query tensor and the transpose of the key tensor is calculated using the matrix multiplication algorithm `matmul`, divided by the square root of the feature dimension, and then input into the `Softmax` normalization function to obtain the initial attention score matrix. In the L-th layer of the visual Transformer computation, the number of image tokens currently participating in the computation is denoted as... The dimension of the current token sequence after concatenating the tokens is... The dimension of the initial attention score matrix of the current layer is The length of the spatial prior weight vector is ; where, the first layer input =196, D=768, corresponding to a first-level token sequence dimension of 197×768.
[0041] Based on information theory principles, the Shannon entropy distribution is calculated by summing the negative probability logarithmic products of the initial attention score matrix row by row, and then the upper limit of the logarithmic number of key tokens in the current layer is applied. The information entropy of each query token is normalized, and the average of the normalized token entropy values is calculated to obtain the normalized global attention entropy scalar. The variance parameter of the two-dimensional Gaussian space prior is calculated proportionally to the magnitude of the normalized global attention entropy. The radial decay degree of the spatial prior weights is then determined by the variance parameter. The Euclidean distance between the center coordinates of each image token and the geometric center coordinates of the currently updated predicted coupled spot is calculated, and this Euclidean distance is substituted into a two-dimensional Gaussian function with the variance as a parameter to generate a two-dimensional spatial prior weight vector.
[0042] The Hadamard product operation is performed by multiplying the initial attention score matrix and the spatial prior weight vector element-wise at corresponding positions using the deep learning tensor element-wise multiplication function `mul`, thus obtaining the modulation attention score matrix. Specifically, the length of `mul` is... The spatial prior weight vector is broadcast and copied along the key dimension to form a matrix identical to the initial attention score matrix. Spatial weight matrix ,in Instead of constructing a diagonal matrix with only the main diagonal non-zero elements from the spatial prior weight vector, it performs... This ensures that the attention allocation for each image token in row 0, where the token is located, is not cleared to zero by off-diagonal lines. The matrix is re-normalized row-wise to maintain the probability distribution properties of the modulation attention scores. Array slicing is then used to extract the token-like elements located in row 0, column 1 to column 2 from the beginning of the modulation attention score matrix. The elements of the column are used as the modulation attention score corresponding to each image token, with the probability value of the attention distribution generated by the class token to each image token being used as the value of the modulation attention score corresponding to each image token; the 0th column corresponds to the class token itself and does not participate in the image token selection.
[0043] In some embodiments, iteratively calculating the initial attention score matrix and normalized global attention entropy of the current layer in the visual Transformer model includes: obtaining the query matrix, key matrix, and value matrix of the current token sequence through linear projection; calculating the matrix product of the transpose of the query matrix and the key matrix, dividing by a scaling factor, and normalizing it using the Softmax function to obtain the initial attention score matrix; calculating the information entropy of the attention distribution corresponding to each query token in the initial attention score matrix, normalizing the information entropy corresponding to each query token according to the number of key tokens in the initial attention score matrix of the current layer, averaging the normalized information entropy, and obtaining the normalized global attention entropy that represents the discreteness of the model's attention and does not drift independently with the change in the number of tokens in the current layer.
[0044] In the feature extraction stage of each visual Transformer structure, the token sequence output from the forward propagation of the previous network layer is used. The input dimension is, for example, 197×768 in the first layer and 197×768 in the Lth layer. The input is fed into the multi-head self-attention calculation module, which is a neural network model. The network structure includes multi-head linear projection units and attention scaling dot product calculation units. The input is a dimensionless array. The token sequence is output as a fused feature matrix of the same dimension and an initial attention score matrix representing the association strength.
[0045] The input sequence is multiplied by three learnable weight matrices, each with a dimension of 768×768, to generate the corresponding query matrix Q, key matrix K, and value matrix V in parallel via linear projection. The generated Q, K, and V are then uniformly divided into h attention subheads along the 768 feature dimensions, with h preferably configured as 12 attention subheads. After the division, each attention subhead has a unique feature dimension d=64.
[0046] For each attention head, extract the matrix resulting from the transpose of its query matrix Q and key matrix K, perform matrix multiplication, and then divide each element of the resulting matrix by a set scaling factor. This means performing a division by 8 to prevent the inner product from becoming too large and causing the backpropagation gradient to vanish.
[0047] Normalization is performed along the row direction using the Softmax activation function, outputting the attention score matrix corresponding to each attention head. When a multi-head attention structure is used, the attention score matrix corresponding to each attention head is calculated separately; the arithmetic mean of the attention score matrices of each attention head is then obtained to obtain the average attention score matrix used for normalizing global attention entropy calculation and token selection scoring in the current layer. Simultaneously, for each attention head, its initial attention score matrix is multiplied element-wise by the spatial prior weight matrix and renormalized row-wise to obtain the modulation attention score matrix corresponding to that attention head. This modulation attention score matrix is then multiplied by the corresponding value matrix to generate the output feature of that attention head. After this step, based on the fundamental principles of information theory, the allocation values in the average attention score matrix used for normalizing global attention entropy calculation are extracted row-by-row to determine the attention probability distribution corresponding to each query token in the current layer sequence. Calculate its information entropy The specific quantitative calculation formula is defined as follows: In the formula, j iterates from 0 to... Or traverse all of the current layer Token key This represents the attention weight assigned to the i-th query token by the j-th key token, and the weights are in the same row. The sum equals 1.
[0048] In obtaining the included After creating an array of information entropy values, each information entropy... Divide by Normalization is performed, where This represents the number of key tokens participating in the attention calculation at the current layer; the normalized entropy value is between 0 and 1. Then, the arithmetic mean of all normalized entropy values in the array is performed to derive a scalar numerical form of the normalized global attention entropy. For example, when the normalized global attention entropy is close to 1, it indicates that the model's current attention distribution is uniformly discrete and unable to locate the target; while when the normalized global attention entropy converges to an empirical value below 0.3, it indicates that the model has focused on a few high-confidence tokens such as the center of a potential spot. This normalized global entropy value will be used as the control variable for the modulation spatial weight decay function of the next functional module.
[0049] In some embodiments, determining the decay rate of the spatial prior weight vector based on the updated predicted geometric center of the coupled spot according to the normalized global attention entropy includes: establishing a two-dimensional Gaussian distribution function with the coordinates of the predicted geometric center of the coupled spot as the mean; establishing a linear mapping relationship between the normalized global attention entropy and the variance of the two-dimensional Gaussian distribution function, using the calculated variance as a parameter to control the radial decay of the spatial prior weight vector, the larger the variance, the smaller the decay rate; substituting the spatial center coordinates of each image token into the two-dimensional Gaussian distribution function to calculate the probability density value, and after normalization processing and supplementing the class token with preset prior values, constructing a spatial prior weight vector.
[0050] To track and predict the pixel coordinates of the geometric center of the coupled spot of the micro-optical device , As the expected mean center, for example, the currently deduced spot center is located at position 112, 112 in the image coordinate system, a two-dimensional Gaussian probability density distribution function model is constructed, which is mathematically expressed as: The model aims to characterize the distribution characteristics of the radial exponential decay of energy in a real coupled spot.
[0051] Pre-establish the normalized global attention entropy Driving the variance of the two-dimensional Gaussian function The adjusted linear mapping relationship is mathematically defined as follows: In this governing equation, α and β are calibrated according to the pixel square scale in the image pixel coordinate system. The scaling factor α is preferably set in the range of 800 to 2000, and the basic bias constant β is preferably set in the range of 100 to 400. For example, when α=1200, β=200, and the normalized entropy value... When =0.8, =1160, corresponding to σ of approximately 34.1 pixels, which makes the two-dimensional Gaussian distribution surface have a wider spatial coverage range relative to the 16-pixel token spacing, thereby reducing the decay rate of the spatial prior weights decreasing from the expected center to the surrounding areas.
[0052] After establishing the aforementioned variance control parameters, the remaining tokens after the first one in the sequence will be excluded. The spatial center coordinates of each image token , For example, image token 1 represents the top left corner of the original image region, with center coordinates 8,8. This process is repeated row by row and column by column, substituting these values into the updated two-dimensional Gaussian distribution function. The result is obtained through calculation. The original probability density scalar value at each spatial location point.
[0053] The density values in this set are proportionally normalized to the maximum density value of the current layer, and then... ,in The density value obtained by substituting the i-th image token into the two-dimensional Gaussian distribution function is given, where max(f) is the maximum density value among all image tokens in the current layer. This method preserves the relative smoothness of the Gaussian surface under different variances, avoiding the artificial amplification of spatial weight differences under high entropy states caused by Min-Max stretching. A preset constant prior value, preferably a constant value of 1, is assigned to this type of token and does not participate in the gradient. This value is then added to the first index of the data sequence to assemble a sequence of length [length missing]. The complete spatial prior one-dimensional weight vector is obtained. This one-dimensional weight vector is broadcast and expanded according to the key token dimension into a spatial weight matrix of the same size as the initial attention score matrix, instead of being constructed as a diagonal matrix. It is used as a spatial regularization constraint to participate in the element-wise modulation operation of the initial attention score matrix, guiding network resources to focus on the real potential spot neighborhood.
[0054] S3, when the modulation attention score is not lower than the first preset threshold or is lower than the first preset threshold and the high frequency feature significance is higher than the second preset threshold, it is marked as a retained token. The remaining image tokens are marked as tokens to be merged. They are then weighted and fused into the retained token with the closest Euclidean distance through a shared learnable feature transformation network. The retained token maintains the original spatial center coordinates and high frequency feature significance.
[0055] All image tokens are traversed and classified using logical OR conditional statements. If the modulation attention score of an image token is greater than or equal to a first preset threshold, or if the modulation attention score is less than the first preset threshold but the significance of its corresponding high-frequency features is greater than a second preset threshold, then its feature Boolean mask is set to true and it is marked as a retained token. Otherwise, its feature Boolean mask is set to false and it is marked as a token to be merged. When all image tokens fail to meet the retention conditions, resulting in an empty set of retained tokens, at least one token is selected as a retained token according to its modulation attention score from high to low. If multiple tokens have the same modulation attention score, the token with higher significance of its high-frequency features is selected as the retained token to ensure that there is a target fusion node in the subsequent fusion process.
[0056] For each token to be merged, the Euclidean distance calculation algorithm is invoked to calculate the Euclidean distance in the feature dimension between the current feature vector of the token to be merged and the current feature vectors of all retained tokens. The index of the nearest target retained token is then found and a set of belonging tokens is constructed. When the high-frequency feature significance of the target retained token is higher than a second preset threshold, it is further determined whether the feature Euclidean distance between the token to be merged and the target retained token is less than a preset fusion distance threshold, and whether the spatial center coordinate distance between the two is less than a preset spatial proximity threshold. Only when both conditions are met simultaneously is the token to be merged allowed to be merged into the high-frequency retained token; otherwise, the token to be merged is reassigned to other retained tokens that meet the conditions, or it is passed to the next layer as an independent retained token. When the feature distances are equal or close, the two-dimensional spatial center coordinate distance can be used as a secondary sorting condition to balance semantic similarity and spatial proximity. By utilizing the fully connected layer of the constructed shared-parameter multilayer perceptron network, the feature vectors of all tokens to be merged and the token itself corresponding to the same retained token in the set are projected and mapped to obtain the fusion weight coefficients. The Softmax normalization function is used for normalization. The features of the tokens to be merged are weighted and fused into the feature vector of the corresponding target retained token through tensor dot product summation, thus completing the downsampling and merging process.
[0057] During this fusion downsampling process, the token is retained for information aggregation in the feature dimension, but the initial two-dimensional spatial center coordinate value of the token and the high-frequency feature saliency value obtained by the high-frequency feature extraction algorithm are retained through metadata caching so that they can be passed to the next layer of calculation.
[0058] In some embodiments, the remaining image tokens are marked as tokens to be merged, and are weighted and fused into the token with the smallest Euclidean distance through a shared learnable feature transformation network. The retained tokens maintain their original spatial center coordinates and high-frequency feature saliency. This includes: calculating the Euclidean distance between the feature vector of each token to be merged and the feature vectors of all retained tokens in the feature dimension; for each token to be merged, selecting the retained token with the smallest Euclidean distance as the target fusion node; inputting the feature vector of the token to be merged into a shared learnable feature transformation network composed of multilayer perceptrons, and outputting the transformed feature vector after dimensionality reduction and then dimensionality increase; performing a weighted summation of the transformed feature vector and the feature vector of the corresponding target fusion node to complete the feature update, while maintaining the spatial center coordinates and high-frequency feature saliency of the target fusion node unchanged.
[0059] After determining the ownership of image tokens using the decision strategy and outputting a high-priority set of reserved tokens and a low-priority set of tokens to be merged, each token marked as pending merging is read sequentially. If the set of reserved tokens is empty, at least one fallback reserved token is determined according to the principle of the highest modulation attention score, and this fallback reserved token is removed from the set of tokens to be merged before continuing to calculate the ownership of the tokens to be merged.
[0060] For each token to be merged, its current 768-dimensional feature vector is extracted, and the Euclidean distance between the feature vector and the feature vectors of all retained tokens in the aforementioned set in the 768-dimensional space is calculated in parallel using the tensor broadcast mechanism. For example, if the current batch retains 40 tokens, a measurement array containing 40 Euclidean distance results will be generated.
[0061] For the token to be merged, the Argmin algorithm is applied to the distance array to select the record with the smallest Euclidean distance. The token corresponding to this smallest distance is then designated as the target fusion node for the token to be merged. This mechanism ensures that image tokens with similar semantic relationships or texture features can undergo feature shrinkage, and spatial proximity is prioritized when feature distances are close. When the difference in feature Euclidean distance between two or more retained tokens is less than a preset tolerance, the two-dimensional spatial center coordinate distance between the token to be merged and the candidate retained tokens is further compared, and the one with the smaller spatial distance is selected as the target fusion node.
[0062] Once the target fusion link is established, the original 768-dimensional feature vectors of the tokens to be merged are pushed to a set of multilayer perceptrons that share network weights across all image tokens to perform feature denoising and alignment reconstruction. This multilayer perceptron is a feedforward neural network model with a two-stage bottleneck structure consisting of a first hidden fully connected layer and a second hidden fully connected layer. The input is the original 768-dimensional feature vector of the tokens to be merged, and the output is a transformed feature vector of the same 768-dimensional dimension. Since this network includes the GELU activation function, it is specifically described in this embodiment as a shared learnable feature transformation network; when the nonlinear activation function is omitted in the implementation, it can also degenerate into a shared learnable linear transformation.
[0063] The first hidden layer of this shared learnable feature transformation network performs feature dimensionality reduction and compression, mapping the input from 768 dimensions to a bottleneck low-dimensional space of 384 dimensions. It is supplemented by the GELU activation function to enhance nonlinear expressive power and suppresses redundant background perturbations through the nonlinear mapping learned during training. The second layer performs inverse feature dimensionality restoration, remapping the filtered 384-dimensional vector back to the normalized 768-dimensional interaction space, thereby outputting a transformed feature vector of 768 dimensions after filtering out background noise.
[0064] During the merging phase, for the set to which the same target retained token belongs, the feature vector of the target retained token itself and the transformed feature vectors of each token to be merged belonging to the target retained token are input into a shared scoring layer to obtain the corresponding fusion score. Softmax normalization is then performed on the fusion score to obtain the fusion weight of each token. Finally, the feature vector of the target retained token and the transformed feature vectors of each token to be merged are weighted and summed according to the fusion weight to obtain the updated feature vector of the target retained token. In this way, all tokens to be merged participate in the feature update of the target retained token with learnable weights, avoiding the problem of non-adaptive weights in different fusion scenarios caused by using a fixed global scalar.
[0065] During this unidirectional data flow superposition and fusion loop, the initial spatial center absolute coordinates of the target fusion node's token within the two-dimensional input grid and the associated high-frequency feature saliency values remain unchanged. This serves as a constant masking rewriting intervention, ensuring that the spatial positioning reference plane does not shift position during repeated tensor dimensionality reduction fusion.
[0066] S4. The pose adjustment amount is predicted by the token-like output of each layer, the geometric center of the predicted coupled spot is updated, and the position deviation between the geometric center of the predicted coupled spot and the preset ideal alignment center is used to generate a control signal through a pre-calibrated mapping. Combined with optical power feedback, the actuator is driven to complete the alignment.
[0067] The token-like feature vectors output from each layer of the visual Transformer model's feedforward neural network are extracted and input into an additional multilayer perceptron regression head network. A pose adjustment vector representing the current two-dimensional translational displacement of the coupling plane is generated using the ReLU activation function and fully connected layers. When used to update the predicted geometric center of the coupled spot, the pose adjustment is expressed as a pixel offset in the image coordinate system. If the regression head outputs a micrometer-level displacement in the physical coordinate system, it is first converted into a pixel offset using pre-calibrated pixel-to-micrometer conversion coefficients or a transformation matrix, and then added to the pixel coordinates of the predicted geometric center of the coupled spot. Through vector addition, the coordinates of the predicted geometric center of the coupled spot from the previous layer are added to the pose adjustment generated in this layer, achieving iterative updates to the two-dimensional position of the predicted geometric center of the coupled spot.
[0068] After all layers of the visual Transformer have completed feature iteration, the image coordinate deviation between the predicted geometric center coordinates of the coupled spot in the last layer and the ideal alignment center two-dimensional coordinates set in the initial system calibration is calculated using a coordinate difference algorithm. This image coordinate deviation is then converted into a displacement adjustment in the actuator's physical coordinate system by using a Jacobian kinematic mapping model or a polynomial fitting regression model (polyval) trained in advance using a Support Vector Regression (SVR) algorithm. This displacement is further mapped and transformed into a multi-channel voltage control signal sequence suitable for piezoelectric ceramic actuators.
[0069] Using a digital-to-analog converter (DAC) interface or an actuator drive control bus, multi-channel voltage control signals are sent to a six-degree-of-freedom high-precision parallel displacement stage actuator to drive it to perform sub-micron level displacement adjustment. During this alignment process, the connection library interface or communication protocol interface of the optical power meter is called to read the real-time optical power value of the current fiber coupling end face as the target feedback function. The Nelder-Mead simplex optimization algorithm is used for multi-round control signal adjustment optimization closed-loop control, and alignment is considered complete when any of the following conditions is met: the optical power increase in K consecutive iterations is less than a preset increase threshold. The displacement adjustment step size of the actuator is less than the preset displacement threshold. Real-time optical power reaches the preset qualified power threshold. Or the number of iterations reaches the preset maximum number of iterations. .
[0070] In some embodiments, the predicted coupled spot geometric center is updated using the token-like prediction pose adjustment amount output by each layer, including: inputting the token-like prediction amount output by the current processing layer into the multilayer perceptron regression prediction head, and outputting a two-dimensional pose adjustment amount including lateral and longitudinal offsets; adding the abscissa and ordinate of the predicted coupled spot geometric center used by the current processing layer to the lateral and longitudinal offsets of the two-dimensional pose adjustment amount, respectively, to obtain the updated geometric center coordinates; and passing the updated geometric center coordinates to the next processing layer as a reference point for constructing the spatial prior weight vector of the next layer.
[0071] In a deep feature extraction network composed of multi-layered cascaded visual Transformer neural networks, whenever a certain processing layer of the model, such as the Lth layer of the network depth, finishes a round of forward propagation tensor operation, the token-like tensor located at the first index position of the current output high-dimensional feature sequence, i.e., index position 0, is extracted.
[0072] Given that such tokens have internalized and aggregated the global geometric pose offset features of the entire coupled region image through the in-layer global self-attention interaction mechanism, their 768-dimensional one-dimensional feature vectors are input into the multilayer perceptron regression prediction head module customized for this localization task to perform prediction.
[0073] This module is a regression neural network model. Its network structure includes two cascaded hidden fully connected layers and a terminal directly connected fully connected layer without activation function. The input is a 768-dimensional one-dimensional feature vector, and the output is a two-dimensional vector containing horizontal offset adjustment increments and vertical offset adjustment increments.
[0074] The number of output nodes for each intermediate hidden layer neuron in this module's serial computation decreases progressively, from 256 nodes to 64 nodes. Each intermediate layer uses the ReLU activation function to extract displacement geometric transformation features. The terminal fully connected layer outputs a sequence containing two values, representing the lateral offset adjustment increment Δu and the longitudinal offset adjustment increment Δv, respectively. Δu and Δv are pixel offsets in the image coordinate system during the layer-by-layer geometric center update stage. For example, this module estimates the lateral correction offset increment Δu = 2.45 pixels and the longitudinal correction offset increment Δv = -1.38 pixels. Together, they form a complete two-dimensional image pose residual adjustment vector.
[0075] Following the above analysis process, the coordinate parameters of the predicted coupled spot geometric center used in the two-dimensional Gaussian probability density function decay evaluation calculation at the current calculation level L are obtained, denoted as... , Then, a scalar numerical addition operation is performed on the separated dimensions of the horizontal and vertical coordinate planes, which is to say, the cached data is read... Numerical values plus deduced Residual fine-tuning amount, and numerical superposition The derivation residual fine-tuning amount is used to calculate the updated absolute coordinate parameter value of the geometric center through this coordinate translation operation, denoted as . , ,because , Since Δu and Δv are all located in the image pixel coordinate system, the addition operation remains consistent in terms of unit dimensions.
[0076] The updated coordinates are passed to the (L+1)th subsequent network layer, serving as the reference anchor points for constructing the two-dimensional Gaussian distribution control function and the spatial prior weight vector required for building the next hidden network layer. The iterative convergence process of the X and Y pose offsets output by each sub-layer of the network is as follows: Figure 2 As shown, with increasing network depth, the prediction errors in both directions gradually converge to near zero. This residual error correction pose prediction adjustment scheme, based on the coherent progression of output levels within each depth layer, continuously iterates and fine-tunes the target anchor point of the local field of view through a closed-loop calibration process that performs a superimposed correction at each layer. This reduces the divergent attention dissipation effect in the later stages of the deep link and guides the six-degree-of-freedom precision adjustment mechanism to complete the light source alignment and coupling task between the micro-optical gratings.
[0077] This ablation experiment constructed a dataset based on 10,000 real-time images of the coupling region of micro-optical devices, divided into training and validation sets at an 8:2 ratio. The input image resolution was uniformly set to 224×224×3 pixels. The experiment was deployed on a deep learning server equipped with a graphics processing unit (GPU) and used a gradient descent optimizer for network parameter iteration, with an initial learning rate preset to 0.01%. The experiment selected the mean absolute error of spot center localization, device alignment coupling efficiency, and model inference frame rate as core evaluation metrics. Four network models were established: a basic visual Transformer model, a first variant with only spatial prior weights and a layer-by-layer geometric center update module, a second variant with only high-frequency feature saliency and a token fusion module, and the complete model of this application integrating all the above mechanisms. The performance comparison results of each model on the three core evaluation metrics of spot center localization error, coupling efficiency, and inference frame rate are as follows: Figure 3 As shown.
[0078] The experimental data and results show that the mean absolute error of spot center localization in the basic visual Transformer model is 3.45 μm, the alignment coupling efficiency is 82.3%, and the model inference frame rate is 45 frames per second. The first variant, which only adds spatial prior weights and layer-by-layer geometric center update modules, reduces the mean absolute error of center localization to 1.82 μm, improves the alignment coupling efficiency to 89.6%, and slightly reduces the inference frame rate to 42 frames per second. The second variant, which only adds high-frequency feature saliency and token fusion modules, achieves a mean absolute error of 2.65 μm in center localization, an alignment coupling efficiency of 85.4%, and an inference frame rate of 58 frames per second. The complete model, combining all core technology modules, achieves the best performance, with the lowest mean absolute error of center localization at 1.15 μm, the highest alignment coupling efficiency at 94.2%, and the model inference frame rate maintained at 55 frames per second.
[0079] Analysis of the above results shows that the first variant, by establishing a linear mapping between the normalized global attention entropy and the Gaussian distribution variance, and by using token-like predictions to fine-tune the absolute coordinates of the geometric center layer by layer, suppresses the attention divergence effect when deep networks process features, guiding network resources to focus on the real neighborhood of the spot and improving the center localization accuracy. The second variant, based on the high-frequency detail saliency of image tokens quantized by a spatial domain high-pass filter, reconstructs and weights low-priority redundant background tokens using a shared learnable feature transformation network and merges them into the retained nodes. This mechanism reduces the sequence length participating in multi-head attention calculations while filtering interference from flat optical substrates, thus significantly improving the model's inference speed. The complete model of this application combines spatial prior penalties with high-frequency feature dimensionality reduction, achieving both high-precision feature alignment and localization with micron-level errors and ensuring the real-time tracking speed required for micro / nano device packaging adjustments.
[0080] In the description of this specification, "multiple" or "several" means at least two, such as two, three or more, unless otherwise expressly and specifically defined.
Claims
1. A fully automated coupling method for micro-optical devices based on multi-source visual feedback, characterized in that, include: S1. Acquire real-time images of the coupling region, divide them into image tokens, obtain an initial token sequence through linear embedding and positional encoding, calculate the high-frequency feature saliency of the image tokens, and initialize the predicted geometric center of the coupling spot. S2. Iteratively calculate the initial attention score matrix and normalized global attention entropy of the current layer, and determine the decay rate of the spatial prior weight vector based on the updated predicted geometric center of the coupling spot. Perform Hadamard product operation on the initial attention score matrix and the spatial prior weight vector, and re-normalize by row to obtain the modulation attention score matrix. Multiply it with the value matrix to generate the output features of the current layer, and extract the attention allocation value of the class token to each image token as the modulation attention score. S3. Mark the image tokens with modulation attention scores not lower than the first preset threshold, or those lower than the first preset threshold but with high-frequency feature saliency higher than the second preset threshold, as retained tokens. Mark the rest as tokens to be merged. Through a shared learnable feature transformation network, the tokens are weighted and fused to the retained token with the closest Euclidean distance. The retained tokens maintain their original spatial center coordinates and high-frequency feature saliency. S4. Utilize the token-like prediction pose adjustment amount output from each layer to update the predicted coupled light spot geometric center. Then, generate a control signal for the actuator by mapping the position deviation between the predicted coupled light spot geometric center and the preset ideal alignment center through a pre-calibrated mapping.
2. The method according to claim 1, characterized in that, The process of acquiring a real-time image of the coupling region, dividing it into image tokens, and obtaining an initial token sequence through linear embedding and positional encoding includes: The real-time image is cropped into multiple non-overlapping two-dimensional image blocks according to the preset window size; Each two-dimensional image patch is flattened into a one-dimensional vector in the spatial dimension, and the one-dimensional vector is projected onto the preset feature dimension through a fully connected layer to obtain the image patch embedding vector. Generate a two-dimensional absolute position code corresponding to the spatial location of the two-dimensional image patch, add the two-dimensional absolute position code to the corresponding image patch embedding vector, and generate an initial token sequence by concatenating a token class at the beginning of the sequence.
3. The method according to claim 1, characterized in that, The iterative calculation of the initial attention score matrix and normalized global attention entropy of the current layer includes: Obtain the query matrix, key matrix, and value matrix of the current token sequence through linear projection; Calculate the matrix product of the query matrix and the transpose of the key matrix, divide by the scaling factor, and normalize using the Softmax function to obtain the initial attention score matrix; Calculate the information entropy of the attention distribution corresponding to each query token in the initial attention score matrix. Normalize the information entropy of each query token according to the number of key tokens in the current layer's initial attention score matrix. Calculate the average of the normalized information entropy to obtain the normalized global attention entropy, which represents the discreteness of the model's attention and does not drift independently with the change of the number of tokens in the current layer.
4. The method according to claim 3, characterized in that, The determination of the attenuation rate of the spatial prior weight vector based on the updated predicted geometric center of the coupled spot includes: A two-dimensional Gaussian distribution function with the mean of the predicted geometric center coordinates of the coupled light spot is established; Establish a linear mapping relationship between the normalized global attention entropy and the variance of the two-dimensional Gaussian distribution function. Use the calculated variance as a parameter to control the radial decay of the prior weight vector in the control space. The larger the variance, the smaller the decay rate. The spatial center coordinates of each image token are substituted into a two-dimensional Gaussian distribution function to calculate the probability density value. After normalization and supplementing the class token with preset prior values, a spatial prior weight vector is constructed.
5. The method according to claim 1 or 2, characterized in that, The calculation of the high-frequency feature saliency of the image token includes: High-frequency component images are obtained by filtering the acquired real-time images using a high-pass filter. The sum of squares of pixel values within the image block region corresponding to the initial spatial position of each image token in the high-frequency component image is calculated as the high-frequency feature saliency representing the edge and detail richness of the image token.
6. The method according to claim 1, characterized in that, The remaining tokens are designated as tokens to be merged. They are then weighted and fused through a shared learnable feature transformation network to the token with the closest Euclidean distance. The retained token maintains its original spatial center coordinates and high-frequency feature saliency, including: Calculate the Euclidean distance in the feature dimension between the feature vector of each token to be merged and the feature vectors of all retained tokens; for each token to be merged, select the retained token with the smallest Euclidean distance as the target fusion node; The feature vectors of the tokens to be merged are input into a shared learnable feature transformation network composed of multilayer perceptrons, and the output is a transformed feature vector after dimensionality reduction and then dimensionality increase. The feature update is completed by weighted summation of the transformed feature vector and the corresponding feature vector of the target fusion node, while maintaining the spatial center coordinates and high-frequency feature saliency of the target fusion node unchanged.
7. The method according to claim 6, characterized in that, When the high-frequency feature significance of the target fusion node is higher than a second preset threshold, the weighted fusion further includes: Determine whether the feature Euclidean distance between the Token to be merged and the target fusion node is less than a preset fusion distance threshold, and whether the spatial center coordinate distance between the two is less than a preset spatial proximity threshold; If both are yes, then perform the weighted fusion. If at least one of them is not true, the token to be merged is marked as a reserved token and passed to the next processing level.
8. The method according to claim 1, characterized in that, The step of using the token-like prediction pose adjustment amount output from each layer to update the predicted coupled spot geometric center includes: The token-like output from the current processing level is input into the multilayer perceptron regression prediction head, and the output includes a two-dimensional pose adjustment amount containing lateral and longitudinal offsets. The x and y coordinates of the predicted coupled spot geometric center used in the current processing level are added to the lateral and longitudinal offsets of the two-dimensional pose adjustment amount, respectively, to obtain the updated geometric center coordinates. The updated geometric center coordinates are passed to the next processing level as the reference point for constructing the spatial prior weight vector of the next level.
9. The method according to claim 1, characterized in that, The generation of control signals for combining optical power feedback to drive the actuator includes: During the displacement adjustment process of the driven actuator, the optical power value is acquired in real time, and the control signal is iteratively optimized based on the optical power value using an optimization algorithm. Alignment is considered complete when any of the following preset convergence conditions are met: In multiple consecutive iterations, the increase in optical power was less than the preset increase threshold. The displacement adjustment step of the actuator is less than the preset displacement threshold; The real-time optical power reaches the preset qualified power threshold. The number of iterations has reached the preset maximum number of iterations.
10. The method according to claim 1, characterized in that, The initialization of the predicted coupling spot geometric center includes: Threshold segmentation is performed on the single-channel grayscale image corresponding to the real-time image to generate a binary mask, and the connected component is labeled on the binary mask. Calculate the sum of gray levels or the average brightness within the corresponding regions of each connected component, and select the connected component with the highest sum of gray levels or the highest average brightness as the candidate coupled spot region; The geometric centroid coordinates of the candidate coupled spot region are calculated using the image spatial moment algorithm, and these geometric centroid coordinates are used as the initial predicted geometric center of the coupled spot.
Citation Information
Patent Citations
An automatic coupling device and automatic coupling method
CN110058355B