Microlens data acquisition device adopting optical bus for transmission and control and microlens packaging method
By using a microlens data acquisition device and packaging method controlled by an optical bus, combined with a multi-level aggregated detection network model and ceramic fixtures, the accuracy and stability issues in microlens packaging were solved, achieving efficient and reliable microlens packaging and attitude recognition.
Patent Information
- Application Number
- CN202511534546.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-27
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-10-27
AI Technical Summary
In the current technology, the packaging quality of microlenses is difficult to meet the high precision and stability requirements of high-speed optical modules. In particular, under the trend of miniaturization, there is a lack of practical application research on the accurate identification and grasping technology of microlenses.
A microlens data acquisition device and packaging method using optical bus transmission and control, combined with an optical bus transmission and control platform, drive mechanism and camera, achieves high-precision pose detection and clamping of microlenses through a multi-level aggregated microlens detection network model, and uses ceramic clamps for stable gripping.
It achieves efficient and reliable packaging of microlenses, with attitude recognition accuracy of ±0.15°, power fluctuation of less than 1.8%, extinction ratio fluctuation of less than 0.25dB, and grasping success rate of over 95%.
Smart Images

Figure CN120997464A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of optical microlens identification and clamping technology, and relates to a microlens data acquisition device and microlens packaging method using optical bus transmission control. Background Technology
[0002] With the large-scale deployment of supercomputing and intelligent computing centers, high-speed optical equipment has become a key node controlling their high-performance communication. The role of microlenses is to shape the laser beam, which then propagates over long distances and couples into a single-mode optical fiber. By improving the accuracy of microlens attitude detection and the stability of grasping, precise encapsulation of the microlens position is achieved, ultimately ensuring the reliability of data transmission in high-speed optical modules. Therefore, the encapsulation quality of microlenses in high-speed optoelectronic devices directly affects the data transmission performance of the optical module.
[0003] Optical modules play a crucial role in high-speed data exchange within data centers, and their emission power determines the transmission distance and fidelity. The packaging quality of micro-optical components directly impacts the emission power of optical modules, especially given the trend towards miniaturization, making micro-optical component packaging a significant challenge. Optical microlenses, used as beam shaping elements, face difficulties in achieving precise attitude detection due to their sub-millimeter size and high reflectivity. Furthermore, the fragility of their glass materials further limits the stability of grasping. These issues restrict the realization of high emission power in high-speed optical devices. In fact, the precise identification and grasping technology of microlenses in optoelectronic devices has become a research hotspot in recent years; however, current understanding of optical microlenses in optoelectronic devices largely remains at the theoretical and simulation levels, lacking in-depth research into practical applications. Summary of the Invention
[0004] This invention provides a microlens data acquisition device using optical bus transmission and control, comprising an industrial control computer, an optical bus transmission and control platform, a drive mechanism, and a camera; The drive mechanism includes a base, a drive assembly, and a controller; The drive assembly includes a first motion assembly that displaces along the X-axis, a second motion assembly that displaces along the Y-axis, a third motion assembly that displaces along the Z-axis, and a yaw assembly. The fixed end of the first motion component is fixedly connected to the base, and the first camera is installed on the driving end of the first motion component; The fixed end of the second motion component is fixedly connected to the driving end of the first motion component, and a second camera is mounted on the driving end of the second motion component. The fixed end of the third motion component is fixedly connected to the driving end of the second motion component, and a third camera is mounted on the driving end of the third motion component. The fixed end of the yaw component is fixedly connected to the driving end of the third motion component; The controller is simultaneously connected to the first motion component, the second motion component, the third motion component, and the yaw component via signals, and is used to control the first motion component, the second motion component, the third motion component, and the yaw component; The optical bus includes an input terminal and an output terminal; the input terminal includes a camera terminal and a control terminal, the camera terminal is connected to the first camera, the second camera and the third camera simultaneously, and the control terminal is connected to the controller; the output terminal is set as the optical head terminal, used to connect to the industrial control computer; the original images captured by the first camera, the second camera and the third camera are transmitted to the industrial control computer via the optical bus.
[0005] Furthermore, the yaw assembly includes a connector, a first yaw component, a second yaw component, and a third yaw component; The connector connects the fixed end of the first deflector and the drive end of the third motion component; The fixed end of the second deflector is fixedly connected to the driving end of the first deflector. The fixed end of the third deflector is fixedly connected to the driving end of the second deflector.
[0006] Furthermore, the microlens data acquisition device using optical bus transmission and control also includes a clamping structure; The clamping structure is installed on the drive end of the oscillation assembly and is used to clamp and fix the microlens.
[0007] Furthermore, the clamping structure is configured as a ceramic gripper structure.
[0008] This invention also provides a microlens packaging method using optical bus control, comprising the following steps: Step 1: Use the microlens data acquisition device with optical bus as described above to acquire the original image of the microlens, and construct a microlens dataset based on the acquired original image data. Step 2: Construct a multi-level converging microlens detection network model, and detect the pose of the microlenses based on the multi-level converging microlens detection network model; Step 3: Based on the pose detection results of the microlens, a clamping structure is used to clamp the microlens; Step 4: The microlens, after being positioned and oriented, is gripped and transported to the designated position, and then the microlens is encapsulated in a closed loop.
[0009] Furthermore, the multi-level aggregated microlens detection network model includes a wavelet transform module, a multi-feature merging module, and a multi-source contrast driving module; The multi-feature merging module has two layers; The multi-source contrast driving module has two layers; The wavelet transform module, two multi-feature merging modules, and two multi-source contrast driving modules are aggregated based on the deep learning model.
[0010] Furthermore, the specific operation process of the multi-level polymer microlens detection network model is as follows: The original images from the microlens dataset were obtained during the downsampling stage of the multi-level aggregated microlens detection network model. Four downsampled feature maps of different sizes; The original images from the microlens dataset were obtained during the upsampling stage of the multi-level aggregated microlens detection network model. Four feature maps of different sizes; Among them, downsampled feature map This is obtained by downsampling the original image; downsampled feature map This is a low-frequency feature map processed using the wavelet transform module, and a downsampled feature map. It contains the main contour information and overall structural information of the original image; downsampled feature map For downsampled feature maps The feature map is obtained by downsampling. For downsampled feature maps Obtained by downsampling; Upsampled feature map For downsampled feature maps Second-layer multi-source contrast driving module delivers features The connection yields the upsampled feature map. For upsampled feature maps Upsampling is performed to obtain the upsampled feature map. For upsampled feature maps Upsampling is performed to obtain the upsampled feature map. For upsampled feature maps and the first-layer multi-source contrast driving module delivers features The connection was obtained.
[0011] Furthermore, the wavelet transform module is a 2D wavelet decomposition filter structure built based on a 1D wavelet transform filter; The 1D wavelet transform filter includes a low-pass filter and a high-pass filter; The 2D wavelet decomposition filter contains high-frequency detail information and low-frequency global information.
[0012] Furthermore, the specific process of constructing the 2D wavelet decomposition filter structure is as follows: Four 2D wavelet decomposition filters—low-frequency to low-frequency, low-frequency to high-frequency, high-frequency to low-frequency, and high-frequency to high-frequency—are constructed using the outer product method. The original image is then processed by convolution with each of these four 2D wavelet decomposition filters as a kernel. Perform convolution operations to obtain the overall structural information of the image, including horizontal edges, vertical edges, and diagonal edges; that is, obtain the low-frequency image features obtained after downsampling by the wavelet transform module. Horizontal high-frequency image detail features Vertical high-frequency image detail features and diagonal high-frequency image detail features .
[0013] Furthermore, the multi-feature merging module includes a channel information merging branch and a location information fusion branch; Channel information is used to merge branches of three input feature maps with different proportions. The specific process for processing is as follows: ① Input three feature maps with different scales The mapping is adjusted uniformly to obtain the adjusted result. Feature map Adjusted Feature map And the adjusted Feature map ; ② Adjust the Feature map Adjusted Feature map And the adjusted Feature map Perform secondary segmentation to obtain the feature map set. ; ③ The feature map set Reassemble according to channel order to obtain the feature map. Meanwhile, based on the Monte Carlo attention mechanism, from the feature map Extract key channel information; A location information fusion branch is used to process three input feature maps with different proportions. The specific process for processing is as follows: ; ; ; ; in, This represents a non-overlapping spatial segmentation operation, referring to the feature map. and feature map The two equal parts; This represents the restoration of the spatial structure of the feature map. This represents local semantic compression and reconstruction. This indicates the merging of feature maps. Indicates a connection. Represented as the original feature map, Represented as a low-level feature map, Represented as a high-level feature map, Represented as different feature maps. This is represented as a spatial partitioning operation. This is represented as the output feature map of the location information fusion branch.
[0014] Furthermore, the expression for the multi-source contrast driving module is as follows: ; ; ; ; ; ; ; ; in, Indicates the size of the sliding sampling window. Indicates the number of heads of interest. Indicates obtaining The number of windows, ; Represented as reconstructed features Figure 1 , Represented as reconstructed features Figure 2 , Represented as features Figure 1 and 2 product, Represented as a standardized feature map, Represented as interactive code 1, Represented as Interactive Coding 2, This is represented as the sum of code 1 and code 2. This is represented as the output feature map. Indicated as expanded, This is represented as max pooling. Represented as encoding , Represented as encoding , Represented as encoding , Represented as encoding , Represented as encoding , Represented as encoding ; When the fusion of shallow detail information and deep backbone features is complete, it is combined with the current features. Perform cross-scale bidirectional attention interaction; In cross-comparison attention-driven programs Indicates from arrive Cross attention, Indicates from arrive , Represents convolution The operation, Represents convolution The operation.
[0015] Compared with the prior art, the present invention has the following beneficial effects: (1) The microlens data acquisition device with optical bus transmission and control provided by the present invention transmits the original image by using optical bus. The optical bus communicates with the industrial computer through PCIE and communicates with multiple optical terminals using a beam splitter. The optical terminals are used as relay stations to complete the conversion of photoelectric signals and realize the control of the drive mechanism and the image acquisition of the camera. The use of high bandwidth, low latency and high synchronization optical fiber communication meets the characteristics of microlens packaging platform for multi-camera monitoring and high-precision synchronization of motion axes.
[0016] (2) This invention has designed a reliable microlens gripping control system and ceramic fixture, achieving a microlens gripping success rate of over 95%, which is a reliable hardware system for microlens packaging.
[0017] (3) This application provides a microlens packaging method using optical bus transmission and control, which realizes efficient detection of microlenses by constructing a multi-level aggregation network (MANet) and achieves an attitude recognition accuracy of ±0.15°. Through testing of the packaged high-speed optical device, it was found that the method provided by this invention has the characteristics of power fluctuation of less than 1.8% and extinction ratio fluctuation of less than 0.25dB.
[0018] (4) In order to fully extract the detailed features and semantic information of the microlens, the wavelet transform module, two multi-feature merging modules and two multi-source contrast driving modules were aggregated on the basis of the deep learning model (U-Net). This represents four feature maps of different sizes obtained during downsampling; This represents four feature maps of different sizes aggregated through multi-scale aggregation during upsampling. During the downsampling phase, The low-frequency feature maps processed using the wavelet transform module include the main contours and overall structure of the image. The three high-frequency feature maps, containing horizontal, vertical, and texture information of the image, are further used in the multi-feature merging module. , , The module extracts local details of microlens pose features; the multi-feature merging module integrates multiple feature maps to improve the performance of deep learning models in complex tasks; through multi-feature fusion and local information supplementation, the network's ability to understand details and semantics is enhanced, improving the accuracy of microlens detection; the multi-source contrast-driven module solves the shortcomings of cross-regional dependency and global semantic modeling ability through attention interaction, realizing feature enhancement and improving microlens edge feature recovery.
[0019] (5) The multi-level aggregated microlens detection network model in this invention constructs a bottom-up cross-fusion multi-source feature integration system by integrating wavelet transform module, multi-feature merging module and multi-source contrast driving module, which can effectively capture multi-level, multi-scale and multi-dimensional image information, and finally achieve accurate microlens attitude detection.
[0020] (6) In this invention, the acquired low-frequency image features Used for transmitting key information in the backbone network, horizontal high-frequency image detail features Vertical high-frequency image detail features and diagonal high-frequency image detail features It serves as a detailed supplement for achieving fine feature extraction, thereby reducing the number of downsampling layers and improving network efficiency.
[0021] (7) Deep neural networks tend to compress input information and retain overall information during learning, which in turn leads to the loss of detailed information in single-size feature maps. Therefore, it is necessary to effectively fuse features of different scales so that the output feature map can maintain spatial resolution while improving semantic discriminative ability. To enhance the interaction of detailed information between feature maps of different levels, inspired by multi-feature fusion and multi-receptor field extension mechanisms, this application proposes a multi-feature fusion module for merging information details. The multi-feature fusion module not only achieves task-specific focusing, but also obtains semantic robustness and spatial accuracy through channel information fusion and location information fusion.
[0022] In addition to the objectives, features, and advantages described above, the present invention has other objectives, features, and advantages. The invention will now be described in further detail with reference to the figures. Attached Figure Description
[0023] The accompanying drawings, which form part of this application, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings: Figure 1 This is a schematic diagram of a microlens data acquisition device using optical bus transmission and control according to Embodiment 1 of the present invention; Figure 2 This is a flowchart illustrating a microlens packaging method using optical bus control in Embodiment 2 of the present invention. Figure 3 This is a schematic diagram of the process for constructing the microlens dataset in Embodiment 2 of the present invention; Figure 4 This is a schematic diagram of the framework of the multi-level polymer microlens detection network model in Embodiment 2 of the present invention; Figure 5 This is a schematic diagram of the result after the original image is processed by the wavelet transform module in Embodiment 2 of the present invention; Figure 6 This is a schematic diagram of the framework of the multi-feature merging module in Embodiment 2 of the present invention; Figure 7 This is a schematic diagram of the framework of the multi-source contrast driving module in Embodiment 2 of the present invention; Figure 8 This is a schematic diagram of the process of clamping the microlens by the clamping structure in Embodiment 2 of the present invention; Figure 9 This is a visual evaluation diagram of the ablation experiment results in the experimental examples of this invention; Figure 10 This is a schematic diagram showing the visual comparison results of the six methods in the experimental examples of this invention; Figure 11 This is a schematic diagram of the microlens attitude detection accuracy results in the experimental example of this invention; Figure 12(a) is a schematic diagram of the action synchronization test of optical terminal 1 in the experimental example of the present invention; Figure 12(b) is a schematic diagram of the synchronization test of the other two optical terminals in the experimental example of the present invention; Figure 12(c) is a schematic diagram of the steady-state gripping test results of the fixture in the experimental example of the present invention; Figure 13 This is a schematic diagram of the steady-state grasping result of the microlens in the experimental example of this invention; Figure 14 This is a schematic diagram of the eye diagram performance test of the 4×25Gbps high-speed optical device in the experimental example of this invention.
[0024] in: 1. Base; 2. First motion component; 3. Second motion component; 4. Third motion component; 5. Swing component; 6. First camera; 7. Second camera; 8. Third camera; 9. Clamping structure. Detailed Implementation
[0025] To make the above-mentioned objects, features, and advantages of the present invention clearer and easier to understand, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be noted that the accompanying drawings of the present invention are all in a simplified form and use non-precise proportions, and are only used to facilitate and clearly assist in illustrating the implementation of the present invention; the "several" mentioned in the present invention are not limited to the specific number shown in the examples in the accompanying drawings.
[0026] Example 1: See Figure 1 As shown, the microlens data acquisition device using optical bus transmission and control provided by the present invention includes an industrial control computer, an optical bus transmission and control platform, a drive mechanism, and a camera; The drive mechanism includes a base 1, a drive assembly, and a controller; The drive assembly includes a first motion assembly 2 that displaces along the X-axis, a second motion assembly 3 that displaces along the Y-axis, a third motion assembly 4 that displaces along the Z-axis, and a yaw assembly 5. The fixed end of the first motion component 2 is fixedly connected to the base 1, and the first camera 6 is installed on the driving end of the first motion component 2. The fixed end of the second motion component 3 is fixedly connected to the driving end of the first motion component 2, and the second camera 7 is mounted on the driving end of the second motion component 3. The fixed end of the third motion component 4 is fixedly connected to the driving end of the second motion component 3, and a third camera 8 is installed on the driving end of the third motion component 4. The fixed end of the yaw component 5 is fixedly connected to the driving end of the third motion component 4; The controller is simultaneously connected to the first motion component 2, the second motion component 3, the third motion component 4 and the yaw component 5 via signal connection, and is used to control the first motion component 2, the second motion component 3, the third motion component 4 and the yaw component 5; The optical bus includes an input end and an output end; the input end includes a camera terminal and a control terminal, the camera terminal is connected to the first camera 6, the second camera 7 and the third camera 8, and the control terminal is connected to a controller; the output end is set as an optical head end for connection to an industrial control computer; the original images captured by the first camera 6, the second camera 7 and the third camera 8 are transmitted to the industrial control computer via the optical bus.
[0027] Preferably, the yaw component 5 includes a connector, a first yaw component, a second yaw component, and a third yaw component; The connector connects the fixed end of the first deflector and the driving end of the third motion component 4; The fixed end of the second deflector is fixedly connected to the driving end of the first deflector. The fixed end of the third deflector is fixedly connected to the driving end of the second deflector.
[0028] Preferably, the first motion component 2, the second motion component 3, and the third motion component 4 are all preferably configured as any one of a mechanical slide structure, a motor ball screw pair structure, or a motor linear slide rail structure.
[0029] Preferably, the first, second, and third oscillating components are all preferably configured as any one of a rotary hydraulic cylinder, a rotary pneumatic cylinder, or a rotary motor.
[0030] As a further embodiment, the microlens data acquisition device using optical bus transmission and control also includes a clamping structure 9. The clamping structure 9 is installed on the drive end of the third deflector and is used to clamp and fix the microlens.
[0031] Preferably, the clamping structure 9 is a ceramic gripper structure to maintain a stable contact force to ensure the positional accuracy of the microlens. At the same time, the ceramic gripper can effectively reduce the damage or wear to the surface of the microlens during the gripping process.
[0032] Preferably, the first camera 6, the second camera 7, and the third camera 8 are all CCD cameras that use Ethernet for transmission. They are connected to an optical terminal via Ethernet for uploading high-definition images. The overhead camera acquires a planar image of the microlens and uses MANet to process and identify its tilt angle for grasping. The monitoring camera is used for early warning of the distance between the chip and the lens. The adjustment camera adjusts the spatial angle of the lens after the grasp is completed to achieve standardization of the lens posture.
[0033] Preferably, a tray assembly for placing microlenses is also provided on the base 1, the tray assembly including a tray and a tray motor shaft for driving the tray to rotate.
[0034] Example 2: See Figure 2 As shown, the microlens packaging method using optical bus control provided by the present invention includes the following steps: Step 1: Use the microlens data acquisition device with optical bus as described above to acquire the original image of the microlens, and construct a microlens dataset based on the acquired original image data. Step 2: Construct a multi-level converged microlens detection network model (MA-Net) and detect the pose of the microlenses based on the multi-level converged microlens detection network model; Step 3: Based on the pose detection results of the microlens, a ceramic gripper is used to grip the microlens. The tip thickness of the ceramic gripper is only 150 micrometers, which can achieve steady-state gripping in a compact space. At the same time, based on the high-precision synchronization of the optical bus, it can ensure that the two ends of the gripper move synchronously to complete the controllable force gripping of the lens. Step 4: The microlens, after being positioned and oriented, is gripped and transported to the designated position, and then the microlens is encapsulated in a closed loop.
[0035] Furthermore, assuming the image acquired by the first camera is the first original image, the image acquired by the second camera is the second original image, and the image acquired by the third camera is the third original image, then the original image of a single microlens includes the first original image, the second original image, and the third original image; the microlens dataset includes the original images of several microlenses.
[0036] Furthermore, the specific process of constructing a microlens dataset based on the acquired raw image data is as follows: S1.1 The drive mechanism (motion axis) drives the camera to the preset position; S1.2, Preset a threshold for image sharpness; specifically, preset three different thresholds for image sharpness, and define the three different thresholds as the first threshold, the second threshold and the third threshold in order of their numerical values. S1.3; Control the first camera, second camera and third camera respectively to perform wide-range, large-step image connection acquisition to obtain the first set of original images; S1.4 Extract image sharpness values from the first set of original images to obtain the first set of image sharpness values, compare and judge the first set of image sharpness values with the first threshold to obtain the best sharpness image and save the best sharpness image; S1.5. Perform pixel-by-pixel masking on the image with the best resolution to obtain a high-quality microlens dataset.
[0037] Further, see Figure 3 As shown, the specific process for obtaining the image with optimal sharpness is as follows: ① If the clarity of the first set of images is greater than the first threshold, then the current microlens is captured using a medium-range, medium-step image connection acquisition method to obtain the second original image; The image sharpness values of the second original image are extracted to obtain the second set of image sharpness values. The second set of image sharpness values are compared with the preset second threshold. If the second set of image sharpness values is greater than the second threshold, the current microlens is captured by a small-range, small-step image connection acquisition method to obtain the third original image. The image sharpness values of the third original image are extracted to obtain the third set of image sharpness values. The third set of image sharpness values are compared with the preset third threshold. If the third set of image sharpness values is greater than the third threshold, the third set of original images is output and saved as the best sharpness image. ② If the image sharpness value of the first group is not greater than the first threshold, then continue to determine whether the image sharpness value of the first group is greater than the second threshold; If the image clarity value of the first group is greater than the second threshold, the microlens will continue to be captured using a small-range, small-step image connection acquisition method to obtain the fourth original image. The image sharpness values of the fourth original image are extracted to obtain the fourth set of image sharpness values. The fourth set of image sharpness values are compared with the preset third threshold. If the fourth set of image sharpness values is greater than the third threshold, the fourth original image is output and saved as the image with the best sharpness. ③ If the image clarity of the first group is not greater than the second threshold, then continue to determine whether the image clarity of the first group is greater than the third threshold; If the clarity of the first set of images is greater than the third threshold, then the original first set of images will be output and saved as the images with the best clarity. If the clarity of the first set of images is not greater than the third threshold, then return to S1.3 and use the first camera, second camera and third camera to capture images of the microlens in a wide range and large step size image connection acquisition method to obtain the fifth set of original images; ④. Judge the fifth set of original images using the methods ①-③ until the image with the best clarity is obtained and saved.
[0038] Further, see Figure 4 As shown, the multi-level aggregated microlens detection network model (MA-Net) includes a wavelet transform module (WTM), a multi-feature merging module (MFMM), and a multi-source contrast driving module (MCDM). The multi-feature merging module has two layers; The multi-source contrast driving module has two layers; The wavelet transform module, two multi-feature merging modules, and two multi-source contrast driving modules are aggregated based on the deep learning model (U-Net).
[0039] The specific operation process of the multi-level polymer microlens detection network model is as follows: The original images from the microlens dataset were obtained during the downsampling stage of the multi-level aggregated microlens detection network model. Four downsampled feature maps of different sizes; The original images from the microlens dataset were obtained during the upsampling stage of the multi-level aggregated microlens detection network model. Four feature maps of different sizes; Among them, downsampled feature map This is obtained by downsampling the original image; downsampled feature map This is a low-frequency feature map processed using the wavelet transform module, and a downsampled feature map. It contains the main contour information and overall structural information of the original image; downsampled feature map For downsampled feature maps The feature map is obtained by downsampling. For downsampled feature maps Obtained by downsampling; Upsampled feature map For downsampled feature maps Second-layer multi-source contrast driving module delivers features The connection yields the upsampled feature map. For upsampled feature maps Upsampling is performed to obtain the upsampled feature map. For upsampled feature maps Upsampling is performed to obtain the upsampled feature map. For upsampled feature maps and the first-layer multi-source contrast driving module delivers features The connection was obtained.
[0040] Specifically, the first-layer multi-source contrast driving module delivers features. The process of obtaining it is as follows: Downsampled feature map The wavelet transform module decomposes the image into a horizontal high-frequency image detail feature map containing horizontal information. Vertical high-frequency image detail feature map containing vertical information of the image and diagonal high-frequency image detail feature maps containing image texture information Horizontal high-frequency image detail features Vertical high-frequency image detail features and diagonal high-frequency image detail features The input is fed into the first multi-feature merging module for processing to extract local detail features of the microlens pose. Local detail features The data is passed to the first multi-source contrast driving module for processing to obtain the first-layer multi-source contrast driving module delivery features. .
[0041] Specifically, the second-layer multi-source contrast driving module delivers features. The process of obtaining it is as follows: Downsampled feature map The data is then fed into a second multi-feature merging module to extract local detail features of the microlens' pose. ; Downsampled feature map and local details All data are passed to the second multi-source contrast driver module for processing, resulting in... .
[0042] Furthermore, downsampled feature maps The image is passed through a transmission chain to the wavelet transform module, where it is decomposed into low-frequency image features. Horizontal high-frequency image detail features containing horizontal image information Vertical high-frequency image detail features containing vertical information of the image and diagonal high-frequency image detail features containing image texture information .
[0043] The wavelet transform module is commonly used in tasks such as image processing, texture analysis, and edge detection to extract local frequency information and enhance texture or edge structures, effectively improving the accuracy of image detection. Furthermore, in this embodiment, the wavelet transform module is a 2D wavelet decomposition filter structure built based on a 1D wavelet transform filter, used for wavelet decomposition in microlens images.
[0044] Preferably, the 1D wavelet transform filter includes a low-pass filter and a high-pass filter; Low-pass filter of 1D wavelet transform The expression is: ; 1D wavelet transform high-pass filter The expression is: ; in, The coefficients of the low-pass filter in the 1D wavelet transform are represented by the Haar scaling function used in this invention. The high-pass filter coefficients representing the 1D wavelet transform are typically functions of wavelet details; Indicates the input signal; Represented as translation parameters, It is represented as a discrete translation parameter.
[0045] Preferably, the 2D wavelet decomposition filter includes high-frequency detail information and low-frequency global information.
[0046] Preferably, four 2D wavelet decomposition filters are constructed using the outer product method: low-frequency-low-frequency, low-frequency-high-frequency, high-frequency-low-frequency, and high-frequency-high-frequency. The original image is then processed by convolution with each of these four 2D wavelet decomposition filters as a kernel. Perform convolution operations to obtain the overall structural information of the image, including horizontal edges, vertical edges, and diagonal edges; that is, obtain the low-frequency image features obtained after downsampling by the wavelet transform module. Horizontal high-frequency image detail features Vertical high-frequency image detail features and diagonal high-frequency image detail features ;in: Low-frequency image features It retains the main information perceptible to human vision and focuses on the shape and contour of the microlens, but there is noise and background interference. Horizontal high-frequency image detail features It captures the feature variations of microlens images from top to bottom, preserving the horizontal edge details of the microlenses while also preserving the features of noise; Vertical high-frequency image detail features The study focuses on feature jumps from left to right in microlens images, highlighting the vertical edge gradient of microlenses, especially for the detailed extraction of microlens surfaces, providing key information for subsequent accurate detection of microlens pose. Diagonal high-frequency image detail features Used to detect changes in diagonal information, for Figure 5 The microlens shown, due to its small tilt angle, only detects the high-frequency gray background in the diagonal portion.
[0047] A further preferred low-frequency to low-frequency 2D wavelet decomposition filter The expression is: ; Low-frequency to high-frequency 2D wavelet decomposition filter The expression is: ; High-frequency to low-frequency 2D wavelet decomposition filter The expression is: ; High-frequency 2D wavelet decomposition filter The expression is: ; in, Represented as a low-frequency 2D wavelet transform, It is represented as a high-frequency 2D wavelet transform.
[0048] See Figure 6 As shown, deep neural networks tend to compress input information while retaining overall information during learning, which in turn leads to the loss of detailed information in single-size feature maps. Therefore, it is necessary to effectively fuse features at different scales so that the output feature map can maintain spatial resolution while improving semantic discriminative ability. To enhance the interaction of detailed information between feature maps at different levels, inspired by multi-feature fusion and multi-receptor field extension mechanisms, this application provides a multi-feature fusion module for merging information details. This module not only achieves task-specific focusing but also obtains semantic robustness and spatial accuracy through channel information fusion and location information fusion.
[0049] The multi-feature merging module includes a channel information merging branch and a location information fusion branch, used to merge three input feature maps with different proportions. The upsampled feature map is obtained by processing the data using a channel information merging branch and a location information fusion branch, respectively. .
[0050] Preferably, channel information is used to merge branches of three input feature maps with different proportions. The specific process for processing is as follows: ① Input three feature maps with different scales The mapping is uniformly adjusted to three optimized channel feature aggregation feature maps with different proportions. ; ② For each acquired feature map along the channel Secondary segmentation is performed to provide multiple feature maps for subsequent channel information merging; ③ Reassemble the segmented feature maps according to channel order. Meanwhile, key channel information is extracted from each feature map based on the Monte Carlo (MoCA) attention mechanism.
[0051] Further optimization involves inputting feature maps at three different scales. Mapping unified adjustment to The specific process is as follows: ; ; ; ; ; ; ; ; ; ; , ; in, Indicated as adjusted Feature map Represents feature map convolution. These represent feature maps at three different scales. Feature map set after dividing the channel into four equal parts express Combined operations, This indicates that the Monte Carlo attention mechanism is used for feature attention and detail extraction. Represented as the first segment after channel segmentation Convolutional feature maps Represented as any one after the channel is divided Convolutional feature maps Indicated as adjusted Feature map Represented as the first segment after channel segmentation Convolutional feature maps Represented as any one after the channel is divided Convolutional feature maps Represented as adjusted Feature map Represented as the first segment after channel segmentation Convolutional feature maps Represented as any one after the channel is divided Convolutional feature maps This is represented as the reassembled feature map.
[0052] Further preferred, The principle is as follows: ; ; ; ; ; in, This represents the element-sum feature map of three quadratic homogenized channel feature maps; Indicates position and It is randomly selected for Monte Carlo context-dependent sampling; express , express ; This indicates element-wise multiplication; Represents the mapping feature map, Indicates average pooling. Indicates channel interception. Indicates Monte Carlo characteristics, This indicates a refactoring.
[0053] During training, a dynamic and uncertain context proxy vector is constructed by using Monte Carlo-style random sampling of scale and location, thereby improving the model's adaptability to different spatial structures and small target responses.
[0054] For the location information fusion branch, the merged feature maps are sliced at different locations to improve information perception across multiple encoded views. Feature map segmentation not only enhances the local network's ability to represent fine-grained information but also reduces computational costs. Therefore, the location information fusion branch has the following representation: ; ; ; ; in, This represents a non-overlapping spatial segmentation operation, referring to the feature map. and feature map The two equal parts; This represents the reconstruction of the spatial structure of the feature map. This represents local semantic compression and reconstruction. This indicates the merging of feature maps. Indicates a connection. Represented as the original feature map, Represented as a low-level feature map, Represented as a high-level feature map, Represented as different feature maps. This is represented as a spatial partitioning operation. This is represented as the output feature map of the location information fusion branch.
[0055] By fusing location information, the neglect of local information in a single feature map during channel feature interaction is compensated for, thereby improving cross-regional perception capabilities and target detection performance.
[0056] In image segmentation, the semantic contrast correlation between shallow and deep layers always implies key region discrimination information; however, traditional attention mechanisms mostly adopt the form of single-source input or self-attention, neglecting the important value of shallow and deep contrast information in region differentiation. Therefore, this application provides a multi-source contrast-driven module, which forms a novel feature enhancement strategy with multi-source guided perception by jointly modeling shallow, deep, and current input features and introducing a cross-semantic perception mechanism.
[0057] See Figure 7 As shown, the three input feature maps come from shallow, deep, and uniform regions, respectively. Through multiple comparisons and perceptions, fine-grained and context-aware feature reconstruction is achieved. The expression of the multi-source contrast-driven module is as follows: ; ; ; ; ; ; ; ; in, Indicates the size of the sliding sampling window. Indicates the number of heads of interest. Indicates obtaining The number of windows, ; Represented as reconstructed features Figure 1 , Represented as reconstructed features Figure 2 , Represented as features Figure 1 and 2 product, Represented as a standardized feature map, Represented as interactive code 1, Represented as Interactive Coding 2, This is represented as the sum of code 1 and code 2. This is represented as the output feature map. Indicated as expanded, This is represented as max pooling. Represented as encoding , Represented as encoding , Represented as encoding , Represented as encoding , Represented as encoding , Represented as encoding ; When the fusion of shallow detail information and deep backbone features is complete, it is combined with the current features. Perform cross-scale bidirectional attention interaction; In cross-comparison attention-driven programs Indicates from arrive Cross attention, Indicates from arrive , Represents convolution The operation, Represents convolution The operation.
[0058] The multi-source contrast-driven module takes the "shallow-deep contrast relationship" as the modeling starting point, introduces a multi-source cross-attention mechanism and scale-aware path, realizes the saliency reconstruction of the spatial domain through depth guidance and shallow compensation, and improves the feature adaptation capability by combining multi-source semantic fusion.
[0059] Further, see Figure 8 As shown, the specific process of using ceramic grippers to grasp the microlens is as follows: S4.1 Based on the microlens pose feedback results, drive the yaw component to adjust the angle of the gripper so that the gripper is aligned with the microlens to be gripped. S4.2 The grippers descend to grasp the microlens.
[0060] More preferably, the thickness of the gripper used to hold the microlens is set to 150 micrometers to achieve stable gripping within a compact space.
[0061] Experimental example: (I) Microlens Segmentation Experiment A. Experimental Preparation Microlens segmentation experiments were conducted on a 32GB NVIDIA RTX 5060 GPU running PyTorch. The dataset used for the experiments was a self-built microlens training model (SFMT) containing 2560 images of microlenses from high-speed optics devices against different backgrounds. For the MANet network, to verify its stability under different loss functions, four loss functions—BCELoss, DiceLoss, FocalLoss, and TverskyLoss—were selected for robustness verification. Furthermore, this application employs the most common numerical evaluation metrics in image detection (F1, mIOU, IOU) to comprehensively evaluate the detection performance of MANet.
[0062] B. Ablation Studies To verify the effectiveness of all modules proposed in this application, comprehensive ablation experiments were conducted.
[0063] Table 1 MANet Ablation Study
[0064] Table 1 shows the numerical evaluation of the performance of different modules in the ablation experiments. As can be seen from Table 1, when only the UNet detection network is used, the numerical evaluation metrics IOU, mIOU, and F1 are 86.79%, 90.62%, and 92.90%, respectively. After introducing WTM, IOU, mIOU, and F1 improved by 0.57%, 0.59%, and 0.33%, respectively. The introduction of wavelet transform has a relatively limited impact on the overall enhancement, mainly because the three extracted high-frequency feature maps are directly passed to the decoding process without further fusion learning. MFMM and MCDM, as independent modules fused into the UNet network, show almost the same enhancement effect. The enhancement percentages of IOU, mIOU, and F1 for MFMM are 1.79%, 1.79%, and 1.05%, respectively, while the enhancement percentages of IOU for MCDM are 1.29%, 1.82%, and 0.53%. These results show that introducing multi-scale aggregation mechanisms and multi-source comparison-driven mechanisms can effectively improve the microlens detection accuracy and the overall performance of the network. The simultaneous fusion of WTM and MFMM into the UNet network further enhanced the numerical evaluation results. Compared with the single-module ablation experiments, the dual-module results improved the IOU, mIOU, and F1 scores by 0.43%, 1.38%, and 0.26%, respectively. Similarly, the combination of MFMM and MCDM improved the IOU, mIOU, and F1 scores by at least 0.35%, 0.82%, and 0.18%, respectively. The results of the two-module fusion experiments demonstrate that the combination of specific mechanisms based on the baseline network can effectively improve the detection accuracy of microlenses. Furthermore, MANet, composed of three modules, showed the best results in the final numerical evaluation, improving the IOU, mIOU, and F1 scores by at least 1.55%, 0.37%, and 0.86%, respectively, compared with the best results of the two modules, thus demonstrating the effectiveness of the MANet network architecture in this application.
[0065] Figure 9 The features of each module are further demonstrated from a visual perspective.
[0066] In this application, five representative microlens images were selected for ablation experiments, and they have the following characteristics: Sequence (1) contains tilted microlenses, which severely interfere with the segmentation of the microlens arc; Sequence (2) The microlenses are placed horizontally, and their surface color is almost the same as the background; Sequence (3) is disturbed by a prominent background; Sequence (4) does not contain microlenses, and the background is a continuous point interference; The microlens in sequence (5) is disturbed by the lack of light, resulting in it being almost completely submerged by the background, and its arc direction is difficult to identify.
[0067] In sequence (1), due to interference from the flipped microlenses, significant false alarms occur when using a fusion network of WTM, MFMM, and MCDM alone; however, when the baseline network fuses WTM and MFMM, the false alarm rate of the microlenses is effectively improved. This result demonstrates that by fusing the three acquired high-frequency feature maps with multi-scale features, channel and location information are effectively extracted, achieving key feature extraction. For networks fusing MFMM and MCDM, the lack of low-frequency detail information and multi-layer contrast feature information in the input leads to multiple focusing of the network, thus affecting the success rate of microlens segmentation.
[0068] The microlens rotation angle of sequence (2) is 0°. For faint targets, the baseline network, lacking input of receiving field and multi-scale contrast information, cannot accurately locate the exact position of the microlens, thus failing to perform accurate detection. Based on wavelet transform, MANet achieves the separation of detail information and backbone features, and through the fusion of multi-scale aggregation and multi-source contrast driving mechanisms, it achieves accurate detail detection under a large receiving field.
[0069] In sequence (3), the microlenses with large curved surfaces showed mainly subtle differences in ablation results. The combination of baseline and WTM was not accurate enough for segmenting the tilted edges of the microlenses, which affected the subsequent detection of the lens tilt angle. On the other hand, MANet achieved accurate detection of the microlens tilt surface while ensuring accurate preservation of surface detail information.
[0070] Sequence (4) lacks microlenses, but both the baseline network and the combined WTM network are subject to complex background interference, leading to incorrect scene detection. In undulating backgrounds, it is necessary to expand the receptive field and enhance contextual connectivity to avoid single feature map inputs that cause network mismatch.
[0071] Sequence (5) shows a small, curved microlens that is severely obscured by the background. Using WTM, MFMM, and MCDM fused with the baseline network into a single module resulted in microlens pixel leakage and false detections. However, MANet improved the ability to supplement details while focusing on the key segmented regions of the microlens, achieving high-precision detection of the microlens location.
[0072] Ablation experiments show that WTM, as a basic processing module, can effectively separate backbone information from texture details, reducing the difficulty of subsequent network learning. MFMM aggregates multi-scale features, expands the network's receptive domain, and provides rich detail information. MCDM enhances the semantic feature connections between contexts through contrast, achieving local focus of attention. MANet achieves local focus of attention by fusing these three key modules, thus completing the accurate detection of microlenses.
[0073] C. Comparative Study The designed MANet was compared and experimented with with the current state-of-the-art object segmentation algorithms, and a comprehensive evaluation of MANet was conducted from both numerical and visual perspectives.
[0074] Among the five selected algorithms, HCFNet achieves accurate target extraction by integrating a parallelized patch-aware attention module, a dimension-aware selective integration module, and a multi-dilution channel refiner; MRF3Net achieves target recognition in complex backgrounds by emphasizing a dual mechanism of multi-receptive field perception and effective feature fusion; MTUNet uses a hybrid encoder combining a visual transformer and a convolutional neural network to extract multi-level features and establish long-range dependencies, effectively improving the performance of small target detection in spatial infrared; RDIANNet constructs a receptive field and orientation-induced attention mechanism to address the imbalance between background and object, achieving enhanced diversity of object features; and UIUNet enhances the extraction of local details and global semantic information in small target detection tasks by introducing a residual network, achieving more effective feature extraction at multiple scales.
[0075] Table 2 Comparison of numerical evaluation with 5 state-of-the-art methods
[0076] Table 2 shows the comparison results between the five state-of-the-art methods and MANet from a numerical evaluation perspective. According to the numerical results, HCFNet and MRF3Net achieve performance comparable to the MANet network in evaluation metrics such as IOU, mIOU, and F1, indicating that the expansion of the multi-dimensional receptive field and the multi-feature fusion mechanism can effectively improve the accuracy of microlens detection. However, in terms of image processing efficiency, MANet leads all other networks, requiring only 4.05 seconds. The MANet network can perform image segmentation quickly, which is of great significance in industrial automated production. The RDIANNet model performs poorly in the numerical evaluation of the network, with IOU, mIOU, and F1 scores of 81.15%, 86.33%, and 89.55%, respectively. The differences from the highest values are 9.47%, 7.86%, and 5.92%, respectively. Although lightweight networks reduce computational complexity, they neglect connections between contexts. When the background and foreground of the microlens cannot be clearly distinguished, the detection accuracy will decrease. UIUNet outperforms RDIANNet. The IOU, mIOU, and F1 scores are 86.72%, 91.11%, and 92.84%, respectively. However, due to the extensive nesting of U-shaped networks, UIUNet generates a large number of parameters, which consumes significant memory in practical tests. Based on the comprehensive comparison results, MANet maintains its leading position in all four metrics except for the number of parameters. The integration of wavelet-guided multi-scale aggregation and contextual multi-source comparison mechanisms is of positive significance in microlens detection.
[0077] Figure 10The six methods were further evaluated from a visual perspective. Among the six images listed, there was no significant difference in detection results between HCFNet, MRF3Net, and MANet. Only in sequence (3) did HCFNet have a higher false positive rate. Experimental results show that by employing a multi-feature fusion mechanism to expand the receptive field of the target and enhance the contrast connection between contexts, higher detection accuracy can be achieved in SFMT. For images without microlenses in the field of view, such as sequence (1), MTUNet, RDIANNet, and UIUNet all exhibited varying degrees of false detection. When dealing with point-like backgrounds, a single attention mechanism cannot achieve global attention. It is necessary to strengthen the connection between shallow and deep features to reduce false positives. For microlenses with large tilt angles, such as those in sequences (3) and (5), all networks can correctly segment the overall shape of the microlens. However, for the small curved microlenses in sequence (5), UIUNet and RDIANNet still have a certain degree of missed detection. The lens bodies in sequences (2) and (4) are completely immersed in the background. Besides MANet, other networks exhibited significant missed detections in the segmentation of arc surface lines, affecting subsequent detection of lens tilt angles. Sequence (6) contained interference from flipped microlenses, but due to the clear contrast between the target and background, all networks except RADIANet could display accurate segmentation results. Based on the comparison of the above six typical image sequences, MANet achieved high-precision microlens detection in all results.
[0078] Table 3. Numerical evaluation results for different loss functions
[0079] Table 3 shows a comparison of the numerical evaluation results of the MANet network trained using four different loss functions. The comparison results show that, except for the focal loss function, the three loss functions perform almost identically in the numerical evaluation results. The focal loss function performs poorly in SFMT, mainly because the sample distribution is more balanced and the error sample distribution is smaller, leading to underfitting during network learning. The comparative experimental results of the loss functions indicate that MANet has stable detection performance and generalization ability, and is highly sensitive to difficult samples or local details.
[0080] MANet establishes a fusion network based on wavelet-guided multi-scale aggregation and multi-source contrast-driven architecture. By obtaining a larger sensory field and increasing the extraction of information contrast between contexts, it achieves accurate detection of microlens pose.
[0081] (II) Microlens grasping and packaging experiment A. Microlens pose extraction After completing the microlens segmentation, the tilt angle needs to be identified to provide angular feedback for subsequent precise gripping of the fixture. Considering MANet's high-precision segmentation effect on the curved surface shape of the microlens, only Hough linear segmentation is used to detect the tilt angle of the microlens in the subsequent microlens position identification.
[0082] Figure 11 The results of microlens pose recognition based on MANet detection are presented. The maximum difference between the position detection results of six randomly selected images and the standard results is 0.23°. Experimental results show that when the microlens tilt angle is randomly distributed over a wide range, the simple Hoff linear detection method based on MANet segmentation achieves high-precision lens pose recognition and provides accurate feedback for subsequent microlens grasping.
[0083] B. Steady-state capture system Steady-state gripping of glass and silicon-based microlenses has long been a challenge to the efficiency and success of high-speed optical device packaging. Traditional cylindrical gripping methods cannot control the gripping force, making microlenses prone to breakage during the gripping process. To achieve controllable gripping of microlenses, this application employs a 6-DOF microlens gripping mechanism and ceramic grippers, and completes the online identification and gripping of microlenses through a high-performance synchronous optical bus control system.
[0084] Figures 12(a) and 12(b) show the synchronization test results of the actions of each optical terminal in the optical bus control system. The test results show that the synchronization errors of optical terminal 1 and the other two optical terminals are 11.2 ns and 13.6 ns, respectively. This ns-level synchronization not only ensures the synchronization accuracy of the motion axis trajectory but also guarantees the control for the fixture to synchronously grasp the lens. On the other hand, Figure 12(c) shows the gripping displacement control accuracy of the fixture. The test results show that when the fixture displacement is 4 μm, the maximum overshoot of the system is approximately 12.5%, the maximum overshoot error is ±0.3 μm, and the overshoot time is 0.16 s. This proves that our designed gripping system achieves stable control of the gripping displacement, ensuring a stable output of the microlens gripping force and avoiding damage to the microlens during the gripping process.
[0085] Figure 13 The images show front and side views of the microlens after it has been grasped. Results from six randomly selected images demonstrate that the ceramic gripper under the optical bus control system achieves stable microlens grasping without damage. Regarding the microlens's spatial orientation, in the selected six images, the gripper only clamps the upper 1 / 4 to 1 / 5 of the microlens, preventing collisions with the chip during the microlens packaging process. Overall, after long-term field verification, the microlens grasping success rate has reached over 95%.
[0086] C. High-speed optoelectronic device packaging performance testing After achieving microlens identification and steady-state grasping, the microlens were packaged in optoelectronic devices based on a 6-DOF platform and optical alignment algorithm. This further verifies the impact of our proposed microlens identification and packaging method on the performance of the packaged device. Table 4 and... Figure 14 The optical and electrical properties of the 4×25 Gbps packaged device were tested.
[0087] Table 4. Results of emitted optical power of the 4×25Gbps device
[0088] As shown in Table 4, three 4-channel optical devices were selected for microlens packaging experiments, and the optical power of each channel of the packaged devices was tested. Experimental results show that the optical power fluctuation range of each device channel is less than 1.8%, meeting the optical power requirements for long-distance, high-speed transmission of 100Gbps high-speed optical devices. For optical signal transmission performance testing, the most common extinction ratio (ER) and eye diagram metrics were selected for evaluation. In Figure 12, the signal quality of the communication performance of each channel of the three optical devices was tested at a single-channel communication rate of 25Gbps. In the obtained eye diagrams, the eye diagrams are more open, with less noise interference at the "0" and "1" levels. The maximum ER value is 4.11dB, and the minimum value is 3.86dB, meeting the extinction ratio requirements for 25Gbps signal transmission. Experimental results show that our designed microlens detection and grasping strategy achieves high-performance packaging of 100Gbps high-speed optical devices and ensures high-fidelity signal transmission of the devices.
[0089] This application proposes a novel strategy for high-speed optoelectronic device microlens detection and steady-state packaging integrating deep learning. Through wavelet-guided multi-scale aggregation and multi-source feature comparison, a microlens attitude recognition accuracy better than ±0.15° is achieved. Based on an optical bus control system and a ceramic gripper, the microlens gripping success rate reaches over 95%. Subsequent experimental results on microlens packaging demonstrate that our proposed strategy achieves efficient and stable microlens packaging. However, currently only single-frame image segmentation of the microlens has been achieved. In efficiency-driven automation industries, it is necessary to realize real-time acquisition and detection of high-definition video of microlenses. Subsequently, leveraging the advantages of the high bandwidth transmission of the optical bus, the MANet network is applied to a 3D video network to achieve more efficient packaging of optical components in high-speed optical devices.
[0090] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A microlens data acquisition device employing optical bus transmission and control, characterized in that, This includes industrial control computers, optical bus transmission and control platforms, drive mechanisms, and cameras; The drive mechanism includes a base, a drive assembly, and a controller; The drive assembly includes a first motion assembly that displaces along the X-axis, a second motion assembly that displaces along the Y-axis, a third motion assembly that displaces along the Z-axis, and a yaw assembly. The fixed end of the first motion component is fixedly connected to the base, and the first camera is installed on the driving end of the first motion component; The fixed end of the second motion component is fixedly connected to the driving end of the first motion component, and a second camera is mounted on the driving end of the second motion component. The fixed end of the third motion component is fixedly connected to the driving end of the second motion component, and a third camera is mounted on the driving end of the third motion component. The fixed end of the yaw component is fixedly connected to the driving end of the third motion component; The controller is simultaneously connected to the first motion component, the second motion component, the third motion component, and the yaw component via signals, and is used to control the first motion component, the second motion component, the third motion component, and the yaw component; The optical bus includes an input terminal and an output terminal; the input terminal includes a camera terminal and a control terminal, the camera terminal is connected to the first camera, the second camera and the third camera simultaneously, and the control terminal is connected to the controller; the output terminal is set as the optical head terminal, used to connect to the industrial control computer; the original images captured by the first camera, the second camera and the third camera are transmitted to the industrial control computer via the optical bus.
2. The microlens data acquisition device using optical bus transmission and control according to claim 1, characterized in that, The yaw assembly includes a connector, a first yaw component, a second yaw component, and a third yaw component; The connector connects the fixed end of the first deflector and the drive end of the third motion component; The fixed end of the second deflector is fixedly connected to the driving end of the first deflector. The fixed end of the third deflector is fixedly connected to the driving end of the second deflector.
3. The microlens data acquisition device using optical bus transmission and control according to claim 1 or 2, characterized in that, It also includes a clamping structure; The clamping structure is installed on the drive end of the oscillation assembly and is used to clamp and fix the microlens.
4. The microlens data acquisition device using optical bus transmission and control according to claim 3, characterized in that, The clamping structure is configured as a ceramic gripper structure.
5. A microlens packaging method using optical bus control, characterized in that, Includes the following steps: Step 1: Use the microlens data acquisition device with optical bus as described in claim 4 to acquire the original image of the microlens, and construct a microlens dataset based on the acquired original image data. Step 2: Construct a multi-level converging microlens detection network model, and detect the pose of the microlenses based on the multi-level converging microlens detection network model; Step 3: Based on the pose detection results of the microlens, a clamping structure is used to clamp the microlens; Step 4: The microlens, after being positioned and oriented, is gripped and transported to the designated position, and then the microlens is encapsulated in a closed loop.
6. The microlens packaging method using optical bus control according to claim 5, characterized in that, The multi-level aggregation microlens detection network model includes a wavelet transform module, a multi-feature merging module, and a multi-source contrast driving module; The multi-feature merging module has two layers; The multi-source contrast driving module has two layers; The wavelet transform module, two multi-feature merging modules, and two multi-source contrast driving modules are aggregated based on the deep learning model.
7. The microlens packaging method using optical bus control according to claim 6, characterized in that, The specific operation process of the multi-level polymer microlens detection network model is as follows: The original images from the microlens dataset were obtained during the downsampling stage of the multi-level aggregated microlens detection network model. Four downsampled feature maps of different sizes; The original images from the microlens dataset were obtained during the upsampling stage of the multi-level aggregated microlens detection network model. Four feature maps of different sizes; Among them, downsampled feature map This is obtained by downsampling the original image; downsampled feature map This is a low-frequency feature map processed using the wavelet transform module, and a downsampled feature map. It contains the main contour information and overall structural information of the original image; downsampled feature map For downsampled feature maps The feature map is obtained by downsampling. For downsampled feature maps Obtained by downsampling; Upsampled feature map For downsampled feature maps Second-layer multi-source contrast driving module delivers features The connection yields the upsampled feature map. For upsampled feature maps Upsampling is performed to obtain the upsampled feature map. For upsampled feature maps Upsampling is performed to obtain the upsampled feature map. For upsampled feature maps and the first layer multi-source contrast driving module delivers features The connection was obtained.
8. The microlens packaging method using optical bus control according to claim 7, characterized in that, The wavelet transform module is a 2D wavelet decomposition filter structure built based on a 1D wavelet transform filter; The 1D wavelet transform filter includes a low-pass filter and a high-pass filter; The 2D wavelet decomposition filter contains high-frequency detail information and low-frequency global information.
9. The microlens packaging method using optical bus control according to claim 8, characterized in that, The specific process of constructing the 2D wavelet decomposition filter structure is as follows: Four 2D wavelet decomposition filters—low-frequency to low-frequency, low-frequency to high-frequency, high-frequency to low-frequency, and high-frequency to high-frequency—are constructed using the outer product method. The original image is then processed by convolution with each of these four 2D wavelet decomposition filters as a kernel. Perform convolution operations to obtain the overall structural information of the image, including horizontal edges, vertical edges, and diagonal edges; that is, obtain the low-frequency image features obtained after downsampling by the wavelet transform module. Horizontal high-frequency image detail features Vertical high-frequency image detail features and diagonal high-frequency image detail features .
10. The microlens packaging method using optical bus control according to claim 6, characterized in that, The multi-feature merging module includes a channel information merging branch and a location information fusion branch; Channel information is used to merge branches of three input feature maps with different proportions. The specific process for processing is as follows: ① Input three feature maps with different scales The mapping is adjusted uniformly to obtain the adjusted result. Feature map Adjusted Feature map And the adjusted Feature map ; ② Adjust the Feature map Adjusted Feature map And the adjusted Feature map Perform secondary segmentation to obtain the feature map set. ; ③ Add feature maps Reassemble according to channel order to obtain feature maps. Meanwhile, based on the Monte Carlo attention mechanism, from the feature map Extract key channel information; A location information fusion branch is used to process three input feature maps with different proportions. The specific process for processing is as follows: ; ; ; ; in, This represents a non-overlapping spatial segmentation operation, referring to the feature map. and feature map The two equal parts; This represents the reconstruction of the spatial structure of the feature map. This represents local semantic compression and reconstruction. This indicates the merging of feature maps. Indicates a connection. Represented as the original feature map, Represented as a low-level feature map, Represented as a high-level feature map, Represented as different feature maps. This is represented as a spatial partitioning operation. This is represented as the output feature map of the location information fusion branch.
11. The microlens packaging method using optical bus control according to claim 10, characterized in that, The expression for the multi-source contrast driving module is as follows: ; ; ; ; ; ; ; ; in, Indicates the size of the sliding sampling window. Indicates the number of heads of interest. Indicates obtaining The number of windows, ; Represented as the reconstructed feature map 1, Represented as the reconstructed feature map 2, Represented as the product of feature maps 1 and 2, Represented as a standardized feature map, Represented as interactive code 1, Represented as Interactive Coding 2, This is represented as the sum of code 1 and code 2. This is represented as the output feature map. Indicated as expanded, This is represented as max pooling. Represented as encoding , Represented as encoding , Represented as encoding , Represented as encoding , Represented as encoding , Represented as encoding ; When the fusion of shallow detail information and deep backbone features is complete, it is combined with the current features. Perform cross-scale bidirectional attention interaction; In cross-comparison attention-driven programs Indicates from arrive Cross attention, Indicates from arrive , Represents convolution The operation, Represents convolution The operation.
Citation Information
Patent Citations
Focusing method and device, computer readable storage medium and electronic equipment
CN110381261A
Lens coupling equipment driven by voice coil motor
CN111443438A
Automatic coupling and packaging method for collimating lenses
CN113786969A
Terahertz suppression fringe interference imaging method and system based on complex field circle construction method
CN114757921A
Defective chip repairing device and method based on femtosecond laser
CN117884757A