Micro-lens data acquisition device using optical bus transmission control and micro-lens packaging method

By using an optical bus control platform and a multi-level aggregated microlens detection network model, combined with a ceramic gripper structure, high-precision pose detection and stable clamping of microlenses were achieved, solving the problem of accurate identification and grasping in microlens packaging and improving the data transmission performance of high-speed optical modules.

CN120997464BActive Publication Date: 2025-12-26CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511534546.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-27
Publication Date
2025-12-26
Estimated Expiration
2045-10-27

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify and stably grasp microlenses, limiting the data transmission performance of high-speed optical modules, especially in miniaturized packaging where packaging quality is difficult to guarantee.

Method used

A microlens data acquisition device and packaging method using optical bus transmission and control, combined with an optical bus transmission and control platform, drive mechanism and camera, achieves high-precision pose detection and clamping of microlenses through a multi-level aggregated microlens detection network model, and uses a ceramic gripper structure for stable clamping.

Benefits of technology

It achieves efficient and precise grasping and packaging of microlenses, improving the reliability of microlens packaging and the stability of data transmission. Power fluctuation is less than 1.8%, extinction ratio fluctuation is less than 0.25dB, and posture recognition accuracy reaches ±0.15°.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997464B_ABST
    Figure CN120997464B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of optical microlens recognition and clamping, and specifically provides a microlens data acquisition device adopting optical bus transmission control and a microlens packaging method, which comprises the following steps: collecting the original image of the microlens by using the microlens data acquisition device adopting optical bus, and constructing a microlens data set based on the collected original image data; constructing a multi-level aggregated microlens detection network model, and detecting the pose of the microlens based on the multi-level aggregated microlens detection network model; clamping the microlens based on the pose detection result of the microlens; clamping and conveying the microlens to a specified position, and then performing closed-loop packaging on the microlens. The application realizes efficient detection of the microlens by constructing a multi-level aggregated network, and realizes an attitude recognition accuracy of ±0.15°.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of optical micro-lens recognition and clamping, and relates to a micro-lens data acquisition device adopting optical bus transmission control and a micro-lens packaging method. BACKGROUND

[0002] With the large-scale deployment of supercomputing and intelligent computing centers, high-speed optical devices have become the key nodes for controlling their high-performance communication. The role of the micro-lens is to shape the laser emitted light beam, and the shaped light beam is coupled into a single-mode optical fiber after long-distance propagation; by improving the accuracy of micro-lens posture detection and the stability of grabbing, the accurate packaging of the micro-lens position is realized, and finally the reliability of high-speed optical module data transmission is ensured. It can be seen that the packaging quality of the micro-lens in the high-speed optoelectronic device directly affects the data transmission performance of the optical module.

[0003] Optical modules play a role in high-speed data exchange in data centers, and their transmission power determines the data transmission distance and fidelity. The packaging quality of micro-optical devices in optical devices directly affects the transmission power of optical modules, especially under the trend of packaging miniaturization, micro-optical device packaging has become a major challenge. Optical micro-lenses used as beam shaping elements are difficult to achieve accurate posture detection due to their sub-millimeter size and high reflection characteristics. At the same time, the fragility of their glass material further limits the stability of grabbing. These problems limit the realization of high transmission optical power in high-speed optical devices. In fact, the accurate recognition and grabbing technology of micro-lenses in optoelectronic devices has become a research hotspot in recent years, but the current mastery of optical micro-lenses in optoelectronic devices mostly stays at the theoretical and simulation level, lacking in-depth research on practical applications. SUMMARY

[0004] The application provides a micro-lens data acquisition device adopting optical bus transmission control, comprising an industrial computer, an optical bus transmission control platform, a driving mechanism and a camera.

[0005] The driving mechanism comprises a base, a driving assembly and a controller.

[0006] The driving assembly comprises a first movement assembly for displacement along the X-axis, a second movement assembly for displacement along the Y-axis, a third movement assembly for displacement along the Z-axis, and a yawing assembly.

[0007] The fixed end of the first movement assembly is fixedly connected with the base, and the driving end of the first movement assembly is provided with the first camera.

[0008] The fixed end of the second movement assembly is fixedly connected with the driving end of the first movement assembly, and the driving end of the second movement assembly is provided with the second camera.

[0009] The fixed end of the third movement component is fixedly connected with the driving end of the second movement component, and a third camera is installed on the driving end of the third movement component;

[0010] The fixed end of the deflection component is fixedly connected with the driving end of the third movement component;

[0011] The controller is connected with the first movement component, the second movement component, the third movement component and the deflection component in signal mode, and is used for controlling the first movement component, the second movement component, the third movement component and the deflection component;

[0012] The optical bus includes an input end and an output end; the input end includes a camera terminal and a control terminal, the camera terminal is connected with the first camera, the second camera and the third camera at the same time, and the control terminal is connected with the controller; the output end is arranged as an optical head terminal and is used for being connected with an industrial computer; the original images captured by the first camera, the second camera and the third camera are transmitted to the industrial computer through the optical bus.

[0013] Further, the deflection component includes a connecting piece, a first deflection piece, a second deflection piece and a third deflection piece;

[0014] The connecting piece connects the fixed end of the first deflection piece and the driving end of the third movement component;

[0015] The fixed end of the second deflection piece is fixedly connected with the driving end of the first deflection piece;

[0016] The fixed end of the third deflection piece is fixedly connected with the driving end of the second deflection piece.

[0017] Further, the micro-lens data acquisition device adopting the optical bus transmission control further includes a clamping structure;

[0018] The clamping structure is installed on the driving end of the deflection component and is used for clamping and fixing the micro-lens.

[0019] Further, the clamping structure is arranged as a ceramic clamping jaw structure.

[0020] The application further provides a micro-lens packaging method adopting optical bus transmission control, including the following steps:

[0021] Step one, the original image of the micro-lens is acquired by using the micro-lens data acquisition device adopting the optical bus, and a micro-lens data set is constructed based on the acquired original image data;

[0022] Step two, a multi-level aggregated micro-lens detection network model is constructed, and the pose of the micro-lens is detected based on the multi-level aggregated micro-lens detection network model;

[0023] Step 3: Based on the pose detection results of the microlens, a clamping structure is used to clamp the microlens;

[0024] Step 4: The microlens, after being positioned and oriented, is gripped and transported to the designated position, and then the microlens is encapsulated in a closed loop.

[0025] Furthermore, the multi-level aggregated microlens detection network model includes a wavelet transform module, a multi-feature merging module, and a multi-source contrast driving module;

[0026] The multi-feature merging module has two layers;

[0027] The multi-source contrast driving module has two layers;

[0028] The wavelet transform module, two multi-feature merging modules, and two multi-source contrast driving modules are aggregated based on the deep learning model.

[0029] Furthermore, the specific operation process of the multi-level polymer microlens detection network model is as follows:

[0030] The original images from the microlens dataset were obtained during the downsampling stage of the multi-level aggregated microlens detection network model. Four downsampled feature maps of different sizes;

[0031] The original images from the microlens dataset were obtained during the upsampling stage of the multi-level aggregated microlens detection network model. Four feature maps of different sizes;

[0032] Among them, downsampled feature map This is obtained by downsampling the original image; downsampled feature map This is a low-frequency feature map processed using the wavelet transform module, and a downsampled feature map. It contains the main contour information and overall structural information of the original image; downsampled feature map For downsampled feature maps The feature map is obtained by downsampling. For downsampled feature maps Obtained by downsampling;

[0033] Upsampled feature map For downsampled feature maps Second-layer multi-source contrast driving module delivers features The connection yields the upsampled feature map. For upsampled feature maps Upsampling is performed to obtain the upsampled feature map. For upsampled feature maps Upsampling is performed to obtain the upsampled feature map. For the up-sampling feature map And the first layer multi-source contrast driving module delivers features Connection.

[0034] Further, the wavelet transform module is a 2D wavelet decomposition filter structure constructed based on a 1D wavelet transform filter;

[0035] The 1D wavelet transform filter includes a low-pass filter and a high-pass filter;

[0036] The 2D wavelet decomposition filter contains high-frequency detail information and low-frequency global information.

[0037] Further, the specific process of constructing the 2D wavelet decomposition filter structure is as follows:

[0038] Four 2D wavelet decomposition filters of low-frequency-low-frequency, low-frequency-high-frequency, high-frequency-low-frequency and high-frequency-high-frequency are constructed by using the outer product method; the overall structure information, horizontal edge, vertical edge and diagonal edge of the image are obtained by performing convolution operation on the original image using the four 2D wavelet decomposition filters as convolution kernels respectively, that is, the low-frequency image features , horizontal high-frequency image detail features , vertical high-frequency image detail features and diagonal high-frequency image detail features obtained after down-sampling by the wavelet transform module are obtained.

[0039] Further, the multi-feature merging module includes a channel information merging branch and a position information fusion branch;

[0040] The specific process of processing the three input feature maps of different scales by the channel information merging branch is as follows:

[0041] ①, map the input three different scale feature maps to adjust them uniformly to obtain adjusted feature map , adjusted feature map and adjusted feature map ;

[0042] ②, twice segment the adjusted feature map , adjusted feature map and adjusted feature map to obtain a feature map set​ ;

[0043] ③, the feature map set re-assembled in channel order, obtaining a feature map , simultaneously based on the Monte Carlo attention mechanism from the feature map extract key channel information;

[0044] The specific process of using the position information fusion branch to process three input feature maps of different proportions is as follows:

[0045] ;

[0046] ;

[0047] ;

[0048] ;

[0049] wherein, represents a non-overlapping spatial segmentation operation, indicating the bisection of the feature map and the feature map represents the recovery of the spatial structure of the feature map, represents local semantic compression and reconstruction, represents the merged feature map, represents the connection, represents the original feature map, represents the low-level feature map, represents the high-level feature map, represents the different pass feature map, represents the spatial segmentation operation, represents the output feature map of the position information fusion branch.

[0050] Further, the expression of the multi-source contrast driving module is as follows:

[0051] ;

[0052] ;

[0053] ;

[0054] ;

[0055] ;

[0056] ; ​​

[0057] ;

[0058] ;

[0059] wherein, denotes the size of the sliding sampling window, denotes the number of attention heads, denotes the number of windows obtained ; denotes the reconstructed feature Figure 1 , denotes the reconstructed feature Figure 2 , denotes the product of the features Figure 1 and 2 ; denotes the normalized feature map, denotes the interaction encoding 1, denotes the interaction encoding 2, denotes the addition of the encoding 1 and the encoding 2, denotes the output feature map, denotes the unfolding, denotes the max-pooling, denotes the encoding , denotes the encoding , denotes the encoding , denotes the encoding , denotes the encoding , denotes the encoding ;

[0060] When the fusion of shallow detail information and deep backbone features is completed, cross-scale bidirectional attention interaction is performed with the current feature ;

[0061] In the cross-comparison attention driver, denotes the cross-attention from to , denotes the cross-attention from to , denotes the operation of convolution , denotes the operation of convolution .

[0062] Compared with the prior art, the present application has the following beneficial effects: ​

[0063] (1) The micro-lens data acquisition device adopting optical bus transmission control provided by the application transmits the original image by adopting an optical bus, the optical bus communicates with an industrial computer through PCIE, and communicates with a plurality of optical terminals by using a beam splitter, and uses the optical terminals as relay stations to complete the conversion of photoelectric signals, realize the control of a driving mechanism and the image acquisition of a camera; the high-speed optical fiber communication with large bandwidth, low delay and high synchronization meets the characteristics of the micro-lens packaging platform for multi-camera monitoring and high-precision synchronization of motion axes.

[0064] (2) The micro-lens grabbing control system and the ceramic clamp designed by the application realize a micro-lens grabbing success rate of more than 95%, which is a reliable hardware system for micro-lens packaging.

[0065] (3) The application provides a micro-lens packaging method adopting optical bus transmission control, which realizes efficient detection of micro-lenses by constructing a multi-level aggregation network (MANet), and realizes an attitude recognition accuracy of ±0.15°; through testing of the packaged high-speed optical device, it is found that the method provided by the application has the characteristics of power fluctuation less than 1.8% and extinction ratio fluctuation less than 0.25dB.

[0066] (4) In order to fully extract detailed features and semantic information of micro-lenses, the wavelet transform module, two multi-feature merging modules and two multi-source contrast driving modules are aggregated on the basis of the deep learning model (U-Net). denotes four different size feature maps obtained during downsampling; denotes four different size feature maps aggregated by multi-scale during upsampling. In the downsampling stage, the low-frequency feature map processed using the wavelet transform module includes the main outline and overall structure of the image. In the multi-feature merging module, three high-frequency feature maps (H, V and T) containing horizontal, vertical and texture information of the image are further used to extract local details of the micro-lens attitude features; 、 、 The multi-feature merging module integrates multi-feature mapping to improve the performance of the deep learning model in complex tasks; through multi-feature fusion and local information supplement, the ability of the network to understand details and semantics is enhanced, and the accuracy of micro-lens detection is improved; the multi-source contrast driving module solves the cross-region dependency and the deficiency of global semantic modeling ability through attention interaction, realizes feature enhancement and improves the recovery of micro-lens edge feature.

[0067] (5) The multi-level aggregated microlens detection network model in the application fuses a wavelet transform module, a multi-feature merging module and a multi-source contrast driving module, constructs a cross-fusion multi-source feature integrated system from bottom to top, can effectively capture multi-level, multi-scale and multi-dimensional image information, and finally realizes accurate microlens posture detection.

[0068] (6) In the application, the low-frequency image features obtained are used for transmission of backbone network main information, horizontal high-frequency image detail features , vertical high-frequency image detail features and diagonal high-frequency image detail features are used as details to supplement fine feature extraction, so that the number of down-sampling layers is reduced, and the network efficiency is improved.

[0069] (7) When learning, a deep neural network tends to compress input information and retain overall information, which in turn leads to loss of detail information in a single-size feature map, therefore, it is necessary to effectively fuse features of different scales to make the output feature map improve semantic discrimination ability while maintaining spatial resolution. In order to enhance the interaction of detail information between feature maps of different levels, inspired by multi-feature fusion and multi-receiving field expansion mechanism, the application proposes a multi-feature merging module for merging information details. The multi-feature merging module not only realizes task-specific focusing, but also obtains semantic robustness and spatial accuracy through channel information merging and position information fusion.

[0070] In addition to the purposes, features and advantages described above, the application has other purposes, features and advantages. The application will be further described below with reference to the drawings. BRIEF DESCRIPTION OF DRAWINGS

[0071] The accompanying drawings, which form a part of this application, are included to provide a further understanding of the application, and are incorporated in and constitute a part of this application. The embodiments of the application illustrated in the drawings, and their description thereto, are presented to provide the applicant's best contemplation of the application, and are not intended to be an improper limitation of the application. In the drawings:

[0072] Figure 1 is a structure schematic view of a microlens data acquisition device using optical bus transmission control in embodiment 1 of the application;

[0073] Figure 2 is a flow schematic view of a microlens packaging method using optical bus transmission control in embodiment 2 of the application;

[0074] Figure 3 is a flow schematic view of a microlens data set construction in embodiment 2 of the application;

[0075] Figure 4 is a frame schematic view of a multi-level aggregated microlens detection network model in embodiment 2 of the application;

[0076] Figure 5 is a schematic diagram of the original image in embodiment 2 of the present application after the wavelet transform module;

[0077] Figure 6 is a schematic diagram of the framework of the multi-feature merging module in embodiment 2 of the present application;

[0078] Figure 7 is a schematic diagram of the framework of the multi-source contrast driving module in embodiment 2 of the present application;

[0079] Figure 8 is a schematic diagram of the process of clamping the micro-lens by the clamping structure in embodiment 2 of the present application;

[0080] Figure 9 is a schematic diagram of the visual assessment of the ablation experiment results in the experimental example of the present application;

[0081] Figure 10 is a schematic diagram of the visual comparison results of the six methods in the experimental example of the present application;

[0082] Figure 11 is a schematic diagram of the micro-lens posture detection accuracy results in the experimental example of the present application;

[0083] Figure 12(a) is a schematic diagram of the action synchronization test of the optical terminal 1 in the experimental example of the present application;

[0084] Figure 12(b) is a schematic diagram of the action synchronization test of the remaining two optical terminals in the experimental example of the present application;

[0085] Figure 12(c) is a schematic diagram of the test results of the stable grasping of the clamp in the experimental example of the present application;

[0086] Figure 13 is a schematic diagram of the micro-lens stable grasping results in the experimental example of the present application;

[0087] Figure 14 is a schematic diagram of the eye diagram performance test of the 4x25Gbps high-speed optical device in the experimental example of the present application.

[0088] wherein:

[0089] 1, base, 2, first movement assembly, 3, second movement assembly, 4, third movement assembly, 5, yaw assembly, 6, first camera, 7, second camera, 8, third camera, 9, clamping structure. DETAILED DESCRIPTION

[0090] In order to make the above-mentioned purposes, features and advantages of the present application more clear and easy to understand, the specific embodiments of the present application are described in detail below with reference to the drawings. It should be noted that the drawings of the present application are all simplified and use non-accurate proportions, and are only used to facilitate and clearly assist in describing the implementation of the present application; the number of several mentioned in the present application is not limited to the specific number in the example of the drawings.

[0091] Embodiment 1:

[0092] Referring to Figure 1 As shown in the figure, the micro-lens data acquisition device provided by the present application adopts optical bus transmission control, and comprises an industrial computer, an optical bus transmission control platform, a driving mechanism and a camera.

[0093] The driving mechanism comprises a base 1, a driving assembly and a controller.

[0094] The driving assembly comprises a first movement assembly 2 for displacement along the X axis, a second movement assembly 3 for displacement along the Y axis, a third movement assembly 4 for displacement along the Z axis and a yawing assembly 5.

[0095] The fixed end of the first movement assembly 2 is fixedly connected with the base 1, and the driving end of the first movement assembly 2 is provided with a first camera 6.

[0096] The fixed end of the second movement assembly 3 is fixedly connected with the driving end of the first movement assembly 2, and the driving end of the second movement assembly 3 is provided with a second camera 7.

[0097] The fixed end of the third movement assembly 4 is fixedly connected with the driving end of the second movement assembly 3, and the driving end of the third movement assembly 4 is provided with a third camera 8.

[0098] The fixed end of the yawing assembly 5 is fixedly connected with the driving end of the third movement assembly 4.

[0099] The controller is signal-connected with the first movement assembly 2, the second movement assembly 3, the third movement assembly 4 and the yawing assembly 5, and is used for controlling the first movement assembly 2, the second movement assembly 3, the third movement assembly 4 and the yawing assembly 5.

[0100] The optical bus comprises an input end and an output end; the input end comprises a camera terminal and a control terminal, the camera terminal is connected with the first camera 6, the second camera 7 and the third camera 8 at the same time, and the control terminal is connected with the controller; the output end is provided as an optical head terminal, and is used for being connected with the industrial computer; the original images captured by the first camera 6, the second camera 7 and the third camera 8 are transmitted to the industrial computer through the optical bus.

[0101] Preferably, the yawing assembly 5 comprises a connecting piece, a first yawing piece, a second yawing piece and a third yawing piece.

[0102] The connecting piece connects the fixed end of the first swing piece and the driving end of the third movement component 4.

[0103] The fixed end of the second swing piece is fixedly connected with the driving end of the first swing piece.

[0104] The fixed end of the third swing piece is fixedly connected with the driving end of the second swing piece.

[0105] Preferably, the first movement component 2, the second movement component 3 and the third movement component 4 are preferably any one of mechanical sliding table structure, motor ball screw pair structure or motor linear sliding rail structure.

[0106] Preferably, the first swing piece, the second swing piece and the third swing piece are preferably any one of rotary oil cylinder, rotary air cylinder or rotary motor.

[0107] As a further scheme of the embodiment, the micro-lens data acquisition device using optical bus transmission control further comprises a clamping structure 9.

[0108] The clamping structure 9 is installed on the driving end of the third swing piece and is used for clamping and fixing the micro-lens.

[0109] Preferably, the clamping structure 9 is a ceramic clamping jaw structure, which can maintain stable contact force to ensure the position accuracy of the micro-lens, and can effectively reduce damage or wear to the surface of the micro-lens during the grabbing process.

[0110] Preferably, the first camera 6, the second camera 7 and the third camera 8 are preferably CCD cameras using Ethernet transmission, which are connected to an optical terminal for high-definition image uploading, wherein the overhead camera obtains the planar image of the micro-lens and uses MANet processing to identify the tilt angle to realize grabbing, the monitoring camera is used for early warning of the distance between the chip and the lens, and the spatial angle of the lens is adjusted after the camera completes grabbing to realize the standardization of the lens posture.

[0111] Preferably, a tray assembly for placing the micro-lens is further arranged on the base 1, and the tray assembly comprises a tray and a tray motor shaft for driving the tray to rotate.

[0112] Embodiment 2:

[0113] Referring to Figure 2 The micro-lens packaging method using optical bus transmission control provided by the application comprises the following steps:

[0114] Step 1: The original image of the micro-lens is acquired by using the micro-lens data acquisition device using optical bus transmission control as described above, and a micro-lens data set is constructed based on the acquired original image data.

[0115] Step two, constructing a multi-level aggregated microlens detection network model (MA-Net), and detecting the pose of the microlens based on the multi-level aggregated microlens detection network model;

[0116] Step three, based on the pose detection result of the microlens, using a ceramic gripper to clamp the microlens; the ceramic gripper tip thickness is only 150 microns, which can realize stable clamping in a compact space, and based on the high precision synchronization of the optical bus, it can ensure the synchronous action of the two ends of the gripper to complete the controllable force clamping of the lens;

[0117] Step four, clamping and conveying the microlens after pose control to the designated position, and then performing closed-loop packaging on the microlens.

[0118] Further, the image collected by the first camera is the first original image, the image collected by the second camera is the second original image, and the image collected by the third camera is the third original image, then the original image of a single microlens includes the first original image, the second original image and the third original image; the microlens data set includes the original images of several microlenses.

[0119] Further, the specific process of constructing the microlens data set based on the collected original image data is as follows:

[0120] S1.1, drive the camera to the preset position by the driving mechanism (motion shaft);

[0121] S1.2, threshold presetting for image sharpness; specifically, three different threshold values are preset for image sharpness, and the three different threshold values are defined as first threshold value, second threshold value and third threshold value in order according to the value size;

[0122] S1.3, control the first camera, the second camera and the third camera to perform wide-range and large-step image connection collection respectively to obtain the first group of original images;

[0123] S1.4, extracting the image sharpness value of the first group of original images to obtain the first group of image sharpness values, and comparing and judging the first group of image sharpness values with the first threshold value to obtain the best sharpness image and save the best sharpness image;

[0124] S1.5, performing pixel-by-pixel mask marking on the best sharpness image to obtain a high-quality microlens data set.

[0125] Further, referring to Figure 3 the specific process of obtaining the best sharpness image is as follows:

[0126] ①, if the first group of image definition is greater than the first threshold value, then the middle range, middle step image connection collection mode is adopted to shoot the current microlens, and the second original image is obtained;

[0127] The image definition value of the second original image is extracted to obtain the second group of image definition values, and the second group of image definition values is compared with the preset second threshold value. If the second group of image definition values is greater than the second threshold value, then the small range, small step image connection collection mode is adopted to shoot the current microlens, and the third original image is obtained;

[0128] The image definition value of the third original image is extracted to obtain the third group of image definition values, and the third group of image definition values is compared with the preset third threshold value. If the third group of image definition values is greater than the third threshold value, then the third group of original images is output and saved as the best definition image;

[0129] ②, if the first group of image definition value is not greater than the first threshold value, then continue to judge whether the first group of image definition value is greater than the second threshold value;

[0130] If the first group of image definition value is greater than the second threshold value, then the small range, small step image connection collection mode is continued to be adopted to shoot the microlens, and the fourth original image is obtained;

[0131] The image definition value of the fourth original image is extracted to obtain the fourth group of image definition values, and the fourth group of image definition values is compared with the preset third threshold value. If the fourth group of image definition values is greater than the third threshold value, then the fourth original image is output and saved as the best definition image;

[0132] ③, if the first group of image definition is not greater than the second threshold value, then continue to judge whether the first group of image definition is greater than the third threshold value;

[0133] If the first group of image definition is greater than the third threshold value, then the first group of original images is output and saved as the best definition image;

[0134] If the first group of image definition is not greater than the third threshold value, then return to S1.3, and the first camera, the second camera and the third camera are adopted to shoot the microlens in the wide range, large step image connection collection mode to obtain the fifth group of original images;

[0135] ④, the first to third modes are adopted to judge the fifth group of original images until the best definition image is obtained and saved.

[0136] Further, referring to Figure 4As shown, the multi-level aggregated microlens detection network model (MA-Net) comprises a wavelet transform module (WTM), a multi feature merging module (MFMM) and a multi source contrast driving module (MCDM).

[0137] The multi feature merging module is provided with two layers.

[0138] The multi source contrast driving module is provided with two layers.

[0139] The wavelet transform module, the two multi feature merging modules and the two multi source contrast driving modules are aggregated on the basis of a deep learning model (U-Net).

[0140] The multi-level aggregated microlens detection network model has the following specific operation process:

[0141] The original image in the microlens dataset is obtained in the down-sampling stage of the multi-level aggregated microlens detection network model Four down-sampled feature maps of different sizes;

[0142] The original image in the microlens dataset is obtained in the up-sampling stage of the multi-level aggregated microlens detection network model Four feature maps of different sizes;

[0143] The down-sampled feature map is obtained based on the original image; the down-sampled feature map is a low-frequency feature map processed by the wavelet transform module, and the down-sampled feature map contains the main contour information and the overall structure information of the original image; the down-sampled feature map is obtained based on the down-sampled feature map ; the down-sampled feature map is obtained based on the down-sampled feature map ;

[0144] The up-sampled feature map is obtained by connecting the down-sampled feature map and the feature conveyed by the second layer multi source contrast driving module; the up-sampled feature map is obtained based on the up-sampled feature map ; the up-sampled feature map is obtained based on the up-sampled feature map ; the up-sampled feature map is obtained based on the up-sampled feature map For obtaining the first layer multi-source contrast driving module delivery feature and the first layer multi-source contrast driving module delivery feature is obtained.

[0145] Specifically, the first layer multi-source contrast driving module delivery feature is obtained as follows:

[0146] The down-sampling feature map is decomposed by the wavelet transform module into a horizontal high-frequency image detail feature map containing horizontal information of the image , a vertical high-frequency image detail feature map containing vertical information of the image , and a diagonal high-frequency image detail feature map containing texture information of the image The horizontal high-frequency image detail feature , the vertical high-frequency image detail feature , and the diagonal high-frequency image detail feature are input into the first multi-feature merging module for processing to extract local detail features of the microlens posture feature The local detail features are transmitted to the first multi-source contrast driving module for processing to obtain the first layer multi-source contrast driving module delivery feature .

[0147] Specifically, the second layer multi-source contrast driving module delivery feature is obtained as follows:

[0148] The down-sampling feature map is transmitted to the second multi-feature merging module to extract local detail features of the microlens posture feature ;

[0149] The down-sampling feature map and the local detail features are both transmitted to the second multi-source contrast driving module for processing to obtain .

[0150] Further, the down-sampling feature map is transmitted to the wavelet transform module through a transmission chain layer, and is decomposed by the wavelet transform module into a low-frequency image feature , a horizontal high-frequency image detail feature map containing horizontal information of the image , a vertical high-frequency image detail feature map containing vertical information of the image , and a diagonal high-frequency image detail feature map containing texture information of the image .

[0151] The wavelet transform module is often used in image processing, texture analysis, edge detection and other tasks to extract local frequency information, enhance texture or edge structure, and can effectively improve the accuracy of image detection. Further, in the present embodiment, the wavelet transform module is a 2D wavelet decomposition filter structure constructed based on a 1D wavelet transform filter for wavelet decomposition in a microlens image.

[0152] Preferably, the 1D wavelet transform filter includes a low-pass filter and a high-pass filter.

[0153] Low-pass filter of 1D wavelet transform The expression is:

[0154]

[0155] High-pass filter of 1D wavelet transform The expression is:

[0156]

[0157] wherein, represents the low-pass filter coefficient of 1D wavelet transform, and a Haar scale function is used in the present application; represents the high-pass filter coefficient of 1D wavelet transform, which is usually a function of wavelet details; represents the input signal; represents the translation parameter, represents the discrete translation parameter.

[0158] Preferably, the 2D wavelet decomposition filter contains high-frequency detail information and low-frequency global information.

[0159] Preferably, the low-frequency-low-frequency, low-frequency-high-frequency, high-frequency-low-frequency and high-frequency-high-frequency four 2D wavelet decomposition filters are constructed by using the outer product method; the overall structure information, horizontal edge, vertical edge and diagonal edge of the image are obtained by performing convolution operation on the original image using the four 2D wavelet decomposition filters as convolution kernels, respectively, i.e. the low-frequency image features , horizontal high-frequency image detail features , vertical high-frequency image detail features , and diagonal high-frequency image detail features obtained after down-sampling by the wavelet transform module are obtained; wherein:

[0160] Low-frequency image features retain the main information perceptible to human vision, focus on the shape profile of the microlens, but there are noise and background interference;

[0161] Horizontal high-frequency image detail features​​​ Capturing the top-to-bottom feature variation of the microlens image, preserving the horizontal edge details of the microlens, but at the same time preserving the feature of noise;

[0162] Vertical high-frequency image detail feature Focusing on the feature jump from left to right in the microlens image, highlighting the vertical edge gradient of the microlens, especially for the detail extraction of the microlens surface, providing key information for the subsequent accurate detection of the microlens posture;

[0163] Diagonal high-frequency image detail feature For detecting the diagonal information variation, for Figure 5 The microlens shown in the figure has a small tilt angle, and only the high-frequency gray background of the diagonal part is detected.

[0164] Further preferably, the 2D wavelet decomposition filter of low frequency-low frequency The expression is:

[0165] ;

[0166] 2D wavelet decomposition filter of low frequency-high frequency The expression is:

[0167] ;

[0168] 2D wavelet decomposition filter of high frequency-low frequency The expression is:

[0169] ;

[0170] 2D wavelet decomposition filter of high frequency-high frequency The expression is:

[0171] ;

[0172] wherein, is expressed as a low-frequency 2D wavelet transform, is expressed as a high-frequency 2D wavelet transform.

[0173] Referring to Figure 6As shown, the deep neural network tends to compress the input information and retain the overall information when learning, which in turn leads to the loss of detailed information in the single-size feature map; therefore, it is necessary to effectively fuse features of different scales to make the output feature map improve the semantic discrimination ability while maintaining the spatial resolution. In order to enhance the interaction of detailed information between feature maps of different levels, inspired by the multi-feature fusion and multi-receptive field expansion mechanism, the application provides a multi-feature merging module for merging information details, which not only realizes task-specific focusing, but also obtains semantic robustness and spatial accuracy through channel information merging and position information fusion.

[0174] The multi-feature merging module comprises a channel information merging branch and a position information fusion branch, which are used to process three input feature maps of different scales respectively by using the channel information merging branch and the position information fusion branch to obtain up-sampling feature maps .

[0175] Preferably, the channel information merging branch is used to process three input feature maps of different scales The specific process is as follows:

[0176] ①, the input three feature maps of different scales are mapped and uniformly adjusted into three optimized channel feature aggregation feature maps of different scales ;

[0177] ②, each obtained feature map is twice divided along the channel to provide multi-feature maps for subsequent channel information merging;

[0178] ③, the divided feature maps are reassembled in channel order , and key channel information is extracted from each feature map based on the Monte Carlo (MoCA) attention mechanism.

[0179] Further preferably, the input three feature maps of different scales are mapped and uniformly adjusted into The specific process is as follows:

[0180] ;

[0181] ;

[0182] ;

[0183] ;

[0184] ;

[0185] ;

[0186] ;

[0187] ;

[0188] ;

[0189] ;

[0190] , ;

[0191] wherein, denotes an adjusted feature map, denotes a feature map convolution, denotes three different scale feature maps equally divided by channel four times, denotes a combination operation, denotes feature attention and detail extraction using Monte Carlo attention mechanism, denotes the first convolution feature map after splitting along the channel, denotes any one convolution feature map after splitting along the channel, denotes an adjusted feature map, denotes the first convolution feature map after splitting along the channel, denotes any one convolution feature map after splitting along the channel, denotes an adjusted feature map, denotes the first convolution feature map after splitting along the channel, denotes any one convolution feature map after splitting along the channel, denotes a feature map after reassembly.

[0192] Further preferably, the principle is:

[0193] ;

[0194] ;

[0195] ;

[0196] ;

[0197] ;

[0198] wherein, denotes an element sum feature map of three quadratic homogenization channel feature maps; denotes the position and are randomly selected for Monte Carlo context-dependent sampling; denotes , denotes ; denotes element-wise multiplication; denotes a mapped feature map, denotes average pooling, denotes channel clipping, denotes a Monte Carlo feature, denotes reconstruction.

[0199] During the training process, a dynamic uncertain context agent vector is constructed by Monte Carlo scale and position random sampling, thereby improving the adaptability of the model to different spatial structures and small targets.

[0200] For the position information fusion branch, the merged feature map is position sliced to improve the information perception of the multi-encoding view. Feature map blocking not only improves the ability of the local network to represent fine-grained information, but also reduces the computational cost. Therefore, the position information fusion branch has the following representation:

[0201] ;

[0202] ;

[0203] ;

[0204] ;

[0205] wherein, denotes a non-overlapping spatial division operation, indicating a bisection of the feature map and the feature map ; denotes feature map spatial structure recovery, denotes local semantic compression and reconstruction, denotes a merged feature map, denotes concatenation, denotes an original feature map, denotes a low-level feature map, is expressed as a high-level feature map, is expressed as a different passing feature map, is expressed as a spatial segmentation operation, is expressed as an output feature map of the position information fusion branch.

[0206] Through position information fusion, the neglect of local information of a single feature map in channel feature interaction is made up, cross-region perception ability is improved, and target detection effect is improved.

[0207] In image segmentation, the semantic contrast correlation between shallow and deep layers always means key region discrimination information; however, most of the traditional attention mechanisms adopt the form of single-source input or self-attention, ignoring the important value of deep and shallow layer contrast information in region differentiation. Therefore, the application provides a multi-source contrast driving module, which forms a new feature enhancement strategy with multi-source guided perception by jointly modeling shallow, deep and current input features and introducing a cross-semantic perception mechanism.

[0208] Referring to Figure 7 , the input three feature maps are respectively from shallow, deep and uniform regions, and after multiple comparison and perception, fine-grained and context-aware feature reconstruction is realized. The expression of the multi-source contrast driving module is as follows:

[0209] ;

[0210] ;

[0211] ;

[0212] ;

[0213] ;

[0214] ;

[0215] ;

[0216] ;

[0217] wherein, represents a sliding sampling window size, represents the number of attention heads, represents the number of windows obtained ; ; is expressed as reconstructed feature Figure 1 , is expressed as reconstructed feature Figure 2 , is expressed as featureFigure 1 and 2 product, denoted as normalized feature map, denoted as interaction encoding 1, denoted as interaction encoding 2, denoted as encoding 1 and encoding 2 added, denoted as output feature map, denoted as unfolding, denoted as max pooling, denoted as encoding , denoted as encoding , denoted as encoding , denoted as encoding , denoted as encoding , denoted as encoding ;

[0218] When the fusion of shallow detail information and deep backbone features is completed, cross-scale bidirectional attention interaction is performed with the current feature ;

[0219] In the cross-comparison attention driver, denotes cross-attention from to , denotes cross-attention from to , denotes the operation of convolution , denotes the operation of convolution .

[0220] The multi-source contrast driven module takes the "shallow-deep contrast relationship" as the modeling starting point, introduces a multi-source cross-attention mechanism and a scale perception path, realizes the saliency reconstruction in the spatial domain through depth guidance and shallow compensation, and improves the feature adaptation capability through the combination of multi-source semantic fusion.

[0221] Further, as shown in Figure 8 , the specific process of clamping the microlens by the ceramic clamp jaw is as follows:

[0222] S4.1, based on the feedback result of the microlens pose, driving the deflection assembly to adjust the angle of the clamp jaw, so that the clamp jaw is aligned with the microlens to be clamped;

[0223] S4.2, the clamp jaw goes down to clamp the microlens.

[0224] Further preferably, the clamping jaw is used to set the thickness of the clamping part of the microlens to 150 microns to achieve stable clamping in a compact space.

[0225] Experimental example:

[0226] (I) Micro-lens segmentation experiment

[0227] A. Experimental preparation

[0228] The micro-lens segmentation experiment was performed on a 32G NVIDIA RTX 5060 GPU running PyTorch. The data set used for the experiment was a self-constructed micro-lens training model (SFMT) containing 2560 high-speed optical device micro-lens images under different backgrounds. For the MANet network, in order to verify its stability under different loss function training, four loss functions, BCELoss, DiceLoss, FocalLoss, and TverskyLoss, were selected for robustness verification. In addition, the most common numerical evaluation indicators (F1, mIOU, IOU) in the field of image detection were used in this application to comprehensively evaluate the detection effect of MANet.

[0229] B. Ablation study

[0230] In order to verify the effectiveness of all modules proposed in this application, a comprehensive ablation experiment was conducted.

[0231] Table 1 MANet ablation study

[0232]

[0233] Table 1 shows the numerical evaluation of the performance of different modules in the ablation experiment. From Table 1, it can be seen that when only the UNet detection network is used, the numerical evaluation indicators IOU, mIOU and F1 are 86.79%, 90.62% and 92.90% respectively. After introducing WTM, IOU, mIOU and F1 are increased by 0.57%, 0.59% and 0.33% respectively. The introduction of wavelet transform has a limited effect on the overall enhancement, mainly because the three extracted high-frequency feature maps are directly passed to the decoding process without further fusion learning. MFMM and MCDM, as independent modules fused into the UNet network, show almost the same enhancement effect, with the enhancement percentages of IOU, mIOU and F1 of MFMM being 1.79%, 1.79% and 1.05% respectively, and the enhancement percentages of IOU of MCDM being 1.29%, 1.82% and 0.53% respectively. From this result, it can be seen that the introduction of the multi-scale aggregation mechanism and the multi-source comparison driving mechanism can effectively improve the micro-lens detection accuracy and the overall performance of the network. WTM and MFMM are simultaneously fused into the UNet network, further enhancing the numerical evaluation results. Compared with the single-module ablation experiment results, the double-module results are increased by 0.43%, 1.38% and 0.26% in IOU, mIOU and F1 indicators respectively. Similarly, the combination of MFMM and MCDM increases at least 0.35%, 0.82% and 0.18% in IOU, mIOU and F1 indicators respectively. The results of the two-module fusion experiment show that the combination of specific mechanisms based on the baseline network can effectively improve the detection accuracy of micro-lenses. Further, the MANet composed of three modules performs best in the final numerical evaluation, with IOU, mIOU and F1 increased by at least 1.55%, 0.37% and 0.86% compared with the best results of the two modules, proving the effectiveness of the MANet network architecture in this application.

[0234] Figure 9 Further, the characteristics of each module are demonstrated from a visual perspective.

[0235] In this application, five representative micro-lens images are selected for ablation experiments, which have the following characteristics:

[0236] Sequence (1) contains micro-lenses placed at an angle, which seriously interferes with the segmentation of the micro-lens arcs;

[0237] Sequence (2) has micro-lenses placed horizontally, with their surface color almost the same as the background;

[0238] Sequence (3) is disturbed by a prominent background;

[0239] Sequence (4) does not contain micro-lenses, and the background is a continuous dotted interference;

[0240] Sequence (5) is almost completely submerged in the background due to the lack of light interference, and the direction of the circular arc is difficult to identify.

[0241] In sequence (1), due to the interference of the flipped microlens, when the fusion network of WTM, MFMM and MCDM is used alone, there will be obvious false alarm; when the baseline network fuses WTM and MFMM, the false alarm rate of the microlens is effectively improved. This result shows that by fusing the three obtained high-frequency feature maps with multi-scale features, the channel information and position information are effectively extracted, and the key feature extraction is realized. For the network that fuses MFMM and MCDM, the lack of input of low-frequency detail information and multi-layer contrast feature information will lead to multiple focusing of the network, thereby affecting the success rate of microlens segmentation.

[0242] The rotation angle of the microlens in sequence (2) is 0°, and for dim and weak targets, the baseline network lacking the input of receiving field and multi-scale contrast information cannot accurately locate the exact position of the microlens, so it cannot be accurately detected. Based on wavelet transform, MANet realizes the separation of detail information and backbone features, and through the fusion of multi-scale aggregation and multi-source contrast driving mechanism, it realizes accurate detail detection under a large receiving field.

[0243] In sequence (3), the microlens has a large curved surface, and the result of the ablation experiment mainly shows the difference in details. The combination of the baseline and WTM is not accurate enough for segmenting the inclined edge of the microlens, which will affect the subsequent detection of the lens tilt angle. On the other hand, MANet realizes accurate detection of the inclined surface of the microlens, while ensuring the accurate preservation of the curved surface detail information.

[0244] Sequence (4) has no microlens, but the baseline network and the combined WTM network will be disturbed by the complex background, resulting in a false detection scene. In the fluctuating background, it is necessary to expand the feeling field and enhance the context connection to avoid single feature map input that leads to network mismatch.

[0245] Sequence (5) shows a small curved microlens and is severely submerged in the background. Using WTM, MFMM and MCDM to fuse with the baseline network into a module, there is leakage and false detection of microlens pixels. However, MANet improves the ability to supplement details while focusing on the key segmentation area of the microlens, realizing high-precision detection of the microlens position.

[0246] The ablation experiment results show that WTM as a basic processing module can effectively separate the backbone information from the texture details, reducing the difficulty of subsequent network learning. MFMM aggregates multi-scale features, expands the receptive field of the network, and provides rich detailed information. MCDM enhances the semantic feature connection between contexts through contrast, realizing local focus of attention. MANet realizes local focus of attention by fusing the above three key modules, thus completing the accurate detection of micro-lenses.

[0247] C. Comparative study

[0248] The designed MANet is compared with the current most advanced target segmentation algorithm, and the MANet is comprehensively evaluated from the numerical and visual angles.

[0249] Among the five selected algorithms, HCFNet realizes accurate target extraction by integrating parallel patch perception attention module, dimension perception selective integration module and multi-dilution channel refiner; MRF3Net realizes target recognition in complex background by emphasizing the dual mechanism of multi-receptive field perception and effective feature fusion; MTUNet adopts a hybrid encoder combining visual transformer and convolutional neural network to extract multi-level features and establish long-range dependencies, effectively improving the performance of spatial infrared small target detection; RDIANNet constructs an acceptance domain and direction-induced attention mechanism to solve the imbalance between background and object, realizing the enhancement of object features; UIUNet enhances the extraction of local details and global semantic information in the small target detection task by introducing a residual network, realizing more effective feature extraction at multiple scales.

[0250] Table 2 Numerical evaluation and comparison with five latest methods

[0251]

[0252] Table 2 shows the comparison results between the five state-of-the-art methods and MANet from the perspective of numerical evaluation. According to the numerical results, HCFNet and MRF3Net achieve comparable performance with MANet network in terms of evaluation indicators such as IOU, mIOU and F1, indicating that the expansion of multi-dimensional receptive field and multi-feature fusion mechanism can effectively improve the accuracy of microlens detection. However, in terms of image processing efficiency, MANet leads all other networks, only requiring 4.05 seconds. MANet network can quickly perform image segmentation, which is of great significance in industrial automation production. The RDIANNet model performs poorly in the numerical evaluation of the network, with IOU, mIOU and F1 being 81.15%, 86.33% and 89.55% respectively. The difference from the highest value is 9.47%, 7.86% and 5.92% respectively. Although lightweight networks reduce computational complexity, they ignore the connection between contexts. When the background and foreground of the microlens cannot be clearly distinguished, the detection accuracy will be reduced. The performance of UIUNet is better than that of RDIANNet. The values of IOU, mIOU and F1 are 86.72%, 91.11% and 92.84% respectively. However, due to the use of a large number of U-shaped networks for nesting in UIUNet, a large number of parameters will be generated, which consumes a lot of memory in actual testing. According to the comprehensive comparison results, MANet leads in all four indicators except the number of parameters. The wavelet-guided multi-scale aggregation and context multi-source comparison mechanism integrated in microlens detection have positive significance.

[0253] Figure 10From the visual perspective, the six methods were further evaluated. In the listed 6 images, there was no significant difference in the detection results of HCFNet, MRF3Net and MANet. Only in sequence (3), the false positive rate of HCFNet was higher. The experimental results show that by using the multi-feature fusion mechanism to expand the receptive field of the target and enhance the contrast connection between the contexts, higher detection accuracy can be achieved in SFMT. For images without micro-lenses in the field of view, such as sequence (1), MTUNet, RDIANNet and UIUNet all show different degrees of false detection. When dealing with point-like background, a single attention mechanism cannot achieve global attention. It is necessary to strengthen the connection between shallow and deep features to reduce the occurrence of false positives. For micro-lenses with large inclination angles, such as the micro-lenses in sequences (3) and (5), all networks can correctly segment the overall shape of the micro-lenses. However, for small curved micro-lenses in sequence (5), UIUNet and RDIANNet still have a certain degree of missed detection. The lens body in sequences (2) and (4) is completely submerged in the background. Except for MANet, other networks have a large number of missed detections in the segmentation of circular arc surface lines, which affects the subsequent detection of lens tilt angles. There is interference of flipped micro-lenses in sequence (6), but due to the clear contrast between the target and the background, all networks except RADIANet can show accurate segmentation results. Based on the comparison of the above six typical image sequences, MANet achieves high-precision micro-lens detection in all results.

[0254] Table 3 Evaluation numerical results of different loss functions

[0255]

[0256] Table 3 shows the comparison of numerical evaluation results of MANet network trained using four different loss functions. From the comparison results, it can be seen that except for the focal loss, the three loss functions have almost the same performance in the numerical evaluation results. The focal loss performs poorly in SFMT, mainly due to the more balanced sample distribution and smaller error sample distribution, resulting in insufficient fitting during the network learning process. The comparison experiment results of the loss function show that MANet has stable detection performance and generalization ability, and is highly sensitive to difficult samples or local details.

[0257] MANet establishes a multi-scale aggregation and multi-source contrast-driven fusion network based on wavelet guidance, which achieves accurate detection of micro-lens pose by obtaining a larger receptive field while increasing the information contrast between contexts.

[0258] (II) Micro-lens grabbing and packaging experiment

[0259] A. Micro-lens pose extraction

[0260] After the microlens segmentation is completed, the tilt angle needs to be identified to provide angle feedback for the subsequent accurate gripper. Considering the high-precision segmentation effect of MANet on the curved surface shape of the microlens, in the subsequent microlens position identification, only the Hoff linear segmentation is used to detect the tilt angle of the microlens.

[0261] Figure 11 The recognition results of the microlens posture based on MANet detection are shown. The maximum difference between the position detection results of six randomly selected images and the standard results is 0.23°. The experimental results show that when the microlens tilt angle is randomly distributed within a wide range, the simple Hoff linear detection method based on MANet segmentation realizes high-precision lens posture recognition and provides accurate feedback for the subsequent microlens grabbing.

[0262] B. Steady-state grabbing system

[0263] The steady-state grabbing of glass and silicon-based microlenses has always been a problem for the efficiency and success of high-speed optical device packaging. The traditional cylindrical grabbing method cannot control the grabbing force, which easily causes the microlenses to break during the grabbing process. In order to realize controllable grabbing of microlenses, the present application adopts a 6-DOF microlens grabbing mechanism and ceramic clamps, and completes the online identification and grabbing action of microlenses through a high-performance synchronous optical bus control system.

[0264] Figures 12(a) and 12(b) show the synchronization test results of each optical terminal action in the optical bus control system, and the test results show that the action synchronization errors of optical terminal 1 and the other two optical terminals are 11.2 ns and 13.6 ns, respectively. The synchronization at the ns level not only ensures the synchronization accuracy of the motion trajectory of the motion axis, but also guarantees the control of the clamp to realize synchronous grabbing of the lens. On the other hand, Figure 12(c) shows the grabbing displacement control accuracy of the clamp. The test results show that when the displacement of the clamp is 4 μm, the maximum overshoot of the system is about 12.5%, the maximum overshoot error is ±0.3 μm, and the overshoot time is 0.16 s. This proves that the grabbing system we designed realizes stable control of the grabbing displacement, ensures the stable output of the microlens grabbing force, and avoids damage to the microlenses during the grabbing process.

[0265] Figure 13 The front view and side view of the microlens effect after grabbing the microlens are shown. The results of six randomly extracted images show that the ceramic clamp under the optical bus control system realizes the steady-state clamping of the microlens without causing damage to the microlens. In terms of the spatial posture of the microlens, in the six selected images, the clamp only clamps the upper 1 / 4 to 1 / 5 of the microlens, avoiding collision with the chip during the microlens packaging process. Overall, after long-term field verification, the microlens grabbing success rate has reached more than 95%.

[0266] C. High-speed optoelectronic device packaging performance test

[0267] After the identification and stable grasping of the microlenses are realized, the packaging of the microlenses in the optoelectronic device is completed based on the 6-DOF platform and the optical alignment algorithm. The influence of the proposed microlens identification and packaging method on the performance of the packaged device is further verified. Table 4 and Figure 14 The optical and electrical indicators of the 4x25 Gbps packaged device are tested respectively.

[0268] Table 4 4x25 Gbps device optical power results

[0269]

[0270] As shown in Table 4, three 4-channel optical devices are selected for microlens packaging experiments, and the optical power of each channel of the packaged device is tested. The experimental results show that the optical power fluctuation range of each device channel is less than 1.8%, which meets the optical power index requirements of 100Gbps high-speed optical device long-distance high-speed transmission. For optical signal transmission performance testing, the most common extinction ratio (ER) and eye diagram indicators are selected for evaluation. In Figure 12, the signal quality of the communication performance of each channel of the three optical devices is tested at a single-channel 25Gbps communication rate. In the acquired eye diagram, the eye diagram is more open, and the noise interference of the “0” and “1” levels is less. The maximum value of ER is 4.11dB, and the minimum value is 3.86dB, which meets the extinction ratio requirements of 25Gbps signal transmission. The experimental results show that the microlens detection and grasping strategy designed by us realizes the high-performance packaging of 100Gbps high-speed optical devices and ensures the high-fidelity signal transmission of the device.

[0271] This application proposes a new strategy of integrating deep learning for high-speed optoelectronic device microlens detection and stable packaging. Through wavelet guided multi-scale aggregation and multi-source feature comparison, a microlens pose recognition accuracy better than ±0.15° is realized. Based on the optical bus control system and the ceramic gripper, the success rate of microlens clamping is more than 95%. The experimental results of subsequent microlens packaging show that the strategy proposed by us realizes efficient and stable microlens packaging. However, at present, only single-frame image segmentation of microlenses is realized, and in the pursuit of efficiency in the automation industry, it is necessary to realize real-time acquisition and detection of high-definition video of microlenses. Subsequently, based on the advantages of optical bus large bandwidth transmission, the MANet network is applied to 3D video network to realize more efficient packaging of optical elements in high-speed optical devices.

[0272] The above merely provides the preferred embodiments of the present application, and is not used to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modifications, equivalent replacements, improvements, etc. made within the principles and technical scope of the present application shall fall into the scope of the present application.

Claims

1. A microlens data acquisition device employing optical bus transmission and control, characterized in that, This includes industrial control computers, optical bus transmission and control platforms, drive mechanisms, and cameras; The drive mechanism includes a base, a drive assembly, and a controller; The drive assembly includes a first motion assembly that displaces along the X-axis, a second motion assembly that displaces along the Y-axis, a third motion assembly that displaces along the Z-axis, and a yaw assembly. The fixed end of the first motion component is fixedly connected to the base, and the first camera is installed on the driving end of the first motion component; The fixed end of the second motion component is fixedly connected to the driving end of the first motion component, and a second camera is mounted on the driving end of the second motion component. The fixed end of the third motion component is fixedly connected to the driving end of the second motion component, and a third camera is mounted on the driving end of the third motion component. The fixed end of the yaw component is fixedly connected to the driving end of the third motion component; The controller is simultaneously connected to the first motion component, the second motion component, the third motion component, and the yaw component via signals, and is used to control the first motion component, the second motion component, the third motion component, and the yaw component; The optical bus includes an input terminal and an output terminal; the input terminal includes a camera terminal and a control terminal, the camera terminal is connected to the first camera, the second camera and the third camera simultaneously, and the control terminal is connected to a controller; the output terminal is set as an optical head terminal for connection to an industrial control computer; the original images captured by the first camera, the second camera and the third camera are transmitted to the industrial control computer via the optical bus; The specific process of constructing a microlens dataset based on the acquired raw image data is as follows: S1.1 The drive mechanism drives the camera to the preset position; S1.2, Preset a threshold for image sharpness; specifically, preset three different thresholds for image sharpness, and define the three different thresholds as the first threshold, the second threshold and the third threshold in order of their numerical values. S1.3 Control the first camera, second camera and third camera respectively to perform wide-range, large-step image connection acquisition to obtain the first set of original images; S1.4 Extract image sharpness values ​​from the first set of original images to obtain the first set of image sharpness values, compare and judge the first set of image sharpness values ​​with the first threshold to obtain the best sharpness image and save the best sharpness image; S1.

5. Perform pixel-by-pixel masking on the best-resolution image to obtain a high-quality microlens dataset; The specific process for obtaining the image with optimal resolution is as follows: ① If the clarity of the first set of images is greater than the first threshold, then the current microlens is captured using a medium-range, medium-step image connection acquisition method to obtain the second original image; The image sharpness values ​​of the second original image are extracted to obtain the second set of image sharpness values. The second set of image sharpness values ​​are compared with the preset second threshold. If the second set of image sharpness values ​​is greater than the second threshold, the current microlens is captured by a small-range, small-step image connection acquisition method to obtain the third original image. The image sharpness values ​​of the third original image are extracted to obtain the third set of image sharpness values. The third set of image sharpness values ​​are compared with the preset third threshold. If the third set of image sharpness values ​​is greater than the third threshold, the third set of original images is output and saved as the best sharpness image. ② If the image sharpness value of the first group is not greater than the first threshold, then continue to determine whether the image sharpness value of the first group is greater than the second threshold; If the image clarity value of the first group is greater than the second threshold, the microlens will continue to be captured using a small-range, small-step image connection acquisition method to obtain the fourth original image. The image sharpness values ​​of the fourth original image are extracted to obtain the fourth set of image sharpness values. The fourth set of image sharpness values ​​are compared with the preset third threshold. If the fourth set of image sharpness values ​​is greater than the third threshold, the fourth original image is output and saved as the image with the best sharpness. ③ If the image clarity of the first group is not greater than the second threshold, then continue to determine whether the image clarity of the first group is greater than the third threshold; If the clarity of the first set of images is greater than the third threshold, then the original first set of images will be output and saved as the images with the best clarity. If the clarity of the first set of images is not greater than the third threshold, then return to S1.3 and use the first camera, second camera and third camera to capture images of the microlens in a wide range and large step size image connection acquisition method to obtain the fifth set of original images; ④. Judge the fifth set of original images using the methods ①-③ until the image with the best clarity is obtained and saved.

2. The microlens data acquisition device using optical bus transmission and control according to claim 1, characterized in that, The yaw assembly includes a connector, a first yaw component, a second yaw component, and a third yaw component; The connector connects the fixed end of the first deflector and the drive end of the third motion component; The fixed end of the second deflector is fixedly connected to the driving end of the first deflector. The fixed end of the third deflector is fixedly connected to the driving end of the second deflector.

3. The microlens data acquisition device using optical bus transmission and control according to claim 1 or 2, characterized in that, It also includes a clamping structure; The clamping structure is installed on the drive end of the oscillation assembly and is used to clamp and fix the microlens.

4. The microlens data acquisition device using optical bus transmission and control according to claim 3, characterized in that, The clamping structure is configured as a ceramic gripper structure.

5. A microlens packaging method using optical bus control, characterized in that, Includes the following steps: Step 1: Use the microlens data acquisition device with optical bus as described in claim 4 to acquire the original image of the microlens, and construct a microlens dataset based on the acquired original image data. Step 2: Construct a multi-level converging microlens detection network model, and detect the pose of the microlenses based on the multi-level converging microlens detection network model; Step 3: Based on the pose detection results of the microlens, a clamping structure is used to clamp the microlens; Step 4: The microlens, after being positioned and oriented, is gripped and transported to the designated position, and then the microlens is encapsulated in a closed loop.

6. The microlens packaging method using optical bus control according to claim 5, characterized in that, The multi-level aggregation microlens detection network model includes a wavelet transform module, a multi-feature merging module, and a multi-source contrast driving module; The multi-feature merging module has two layers; The multi-source contrast driving module has two layers; The wavelet transform module, two multi-feature merging modules, and two multi-source contrast driving modules are aggregated based on the deep learning model.

7. The microlens packaging method using optical bus control according to claim 6, characterized in that, The specific operation process of the multi-level polymer microlens detection network model is as follows: The original images from the microlens dataset were obtained during the downsampling stage of the multi-level aggregated microlens detection network model. Four downsampled feature maps of different sizes; The original images from the microlens dataset were obtained during the upsampling stage of the multi-level aggregated microlens detection network model. Four feature maps of different sizes; Among them, downsampled feature map This is obtained by downsampling the original image; downsampled feature map This is a low-frequency feature map processed using the wavelet transform module, and a downsampled feature map. It contains the main contour information and overall structural information of the original image; downsampled feature map For downsampled feature maps The feature map is obtained by downsampling. For downsampled feature maps Obtained by downsampling; Upsampled feature map For downsampled feature maps Second-layer multi-source contrast driving module delivers features The connection yields the upsampled feature map. For upsampled feature maps Upsampling is performed to obtain the upsampled feature map. For upsampled feature maps Upsampling is performed to obtain the upsampled feature map. For upsampled feature maps and the first layer multi-source contrast driving module delivers features The connection was obtained.

8. The microlens packaging method using optical bus control according to claim 7, characterized in that, The wavelet transform module is a 2D wavelet decomposition filter structure built based on a 1D wavelet transform filter; The 1D wavelet transform filter includes a low-pass filter and a high-pass filter; The 2D wavelet decomposition filter contains high-frequency detail information and low-frequency global information.

9. The microlens packaging method using optical bus control according to claim 8, characterized in that, The specific process of constructing the 2D wavelet decomposition filter structure is as follows: Four 2D wavelet decomposition filters—low-frequency to low-frequency, low-frequency to high-frequency, high-frequency to low-frequency, and high-frequency to high-frequency—are constructed using the outer product method. The original image is then processed by convolution with each of these four 2D wavelet decomposition filters as a kernel. Perform convolution operations to obtain the overall structural information of the image, including horizontal edges, vertical edges, and diagonal edges; that is, obtain the low-frequency image features obtained after downsampling by the wavelet transform module. Horizontal high-frequency image detail features Vertical high-frequency image detail features and diagonal high-frequency image detail features .

10. The microlens packaging method using optical bus control according to claim 6, characterized in that, The multi-feature merging module includes a channel information merging branch and a location information fusion branch; Channel information is used to merge branches of three input feature maps with different proportions. The specific process for processing is as follows: ① Input three feature maps with different scales The mapping is adjusted uniformly to obtain the adjusted result. Feature map Adjusted Feature map And the adjusted Feature map ; ② Adjust the Feature map Adjusted Feature map And the adjusted Feature map Perform secondary segmentation to obtain the feature map set. ; ③ Add feature maps Reassemble according to channel order to obtain feature maps. Meanwhile, based on the Monte Carlo attention mechanism, from the feature map Extract key channel information; A location information fusion branch is used to process three input feature maps with different proportions. The specific process for processing is as follows: ; ; ; ; in, This represents a non-overlapping spatial segmentation operation, referring to the feature map. and feature map The two equal parts; This represents the reconstruction of the spatial structure of the feature map. This represents local semantic compression and reconstruction. This indicates the merging of feature maps. Indicates a connection. Represented as the original feature map, Represented as a low-level feature map, Represented as a high-level feature map, Represented as different feature maps. This is represented as a spatial partitioning operation. This is represented as the output feature map of the location information fusion branch.

11. The microlens packaging method using optical bus control according to claim 10, characterized in that, The expression for the multi-source contrast driving module is as follows: ; ; ; ; ; ; ; ; in, Indicates the size of the sliding sampling window. Indicates the number of heads of interest. Indicates obtaining The number of windows, ; Represented as the reconstructed feature map 1, Represented as the reconstructed feature map 2, Represented as the product of feature maps 1 and 2, Represented as a standardized feature map, Represented as interactive code 1, Represented as Interactive Coding 2, This is represented as the sum of code 1 and code 2. This is represented as the output feature map. Indicated as expanded, This is represented as max pooling. Represented as encoding , Represented as encoding , Represented as encoding , Represented as encoding , Represented as encoding , Represented as encoding ; When the fusion of shallow detail information and deep backbone features is complete, it is combined with the current features. Perform cross-scale bidirectional attention interaction; In cross-comparison attention-driven programs Indicates from arrive Cross attention, Indicates from arrive , Represents convolution The operation, Represents convolution The operation.

Citation Information

Patent Citations

  • Automatic coupling and packaging method for collimating lenses

    CN113786969A

  • Micro-lens segmentation method based on asymmetric convolution multi-level attention network

    CN117975020A