Method and system for quickly positioning patch element of optical communication device based on deep learning
Through the improved CSPDarknet-53 feature extraction backbone network and dynamically aligned feature rotation detection head module, the problem of inaccurate angle prediction and high learning threshold of optical communication device patch elements is solved, and high precision and real-time rotation target detection is achieved, which is suitable for optical communication device production lines.
Patent Information
- Application Number
- CN202510449106.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-18
AI Technical Summary
The existing rotary target detection networks have problems such as insufficient angle prediction and manual selection of template matching and positioning in optical communication devices, and traditional methods are difficult to deal with challenges unique to industrial scenarios such as small component size, high shape similarity, and surface reflection.
The improved CSPDarknet-53 feature extraction backbone network is adopted, combined with the dual-stage feature-focus diffusion fusion network and the rotation detection head module with dynamically aligned features, and the high-precision rotation target detection of the patch elements of the optical communication device is achieved through a square angle-sensitive loss function.
It improves the detection accuracy and angle prediction capabilities of the patch components of optical communication devices, ensures the real-time and robustness of the algorithm, and is suitable for the real-time detection requirements on the production line of optical communication devices.
Smart Images

Figure CN120339583A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and automated detection, and particularly relates to a method and system for rapid positioning of surface-mounted components of optical communication devices based on deep learning. Background Art
[0002] With the rapid development of Internet and industrial Internet of Things communication technologies, the market demand for communication network devices has also increased sharply. Optical communication devices, as key components, play a crucial role in the entire network system. Among them, the precise positioning of surface-mounted components is of great significance for assembly calibration, defect detection, and measuring the production quality of products.
[0003] Traditional vision positioning methods mainly rely on template matching algorithms. The basic idea is to match a pre-constructed template library with the area to be located to determine the position and rotation angle of the surface-mounted component. However, due to the diverse shapes of surface-mounted components and the large number of optical communication models, traditional template matching methods often require manual selection of templates of different models, and template creation is required for different types of optical communication devices, which increases the difficulty of getting started for some staff who are not familiar with the optical communication device models.
[0004] With the rapid development of deep learning, single-stage object detection algorithms such as the YOLO series have been continuously introduced in the field of industrial vision detection. Such algorithms have strong real-time performance and can achieve rapid positioning of different types of surface-mounted components through training specific network models. However, although traditional deep learning object detection algorithms have achieved remarkable success in many fields, they still face a key challenge in the precise positioning task of surface-mounted components of optical communication devices: mainstream object detection algorithms usually use horizontal rectangular boxes for object positioning, while surface-mounted components in optical communication devices are often placed at arbitrary angles, resulting in a large amount of background area being included in the horizontal bounding box and the inability to obtain pose information.
[0005] In recent years, rotation object detection networks have been widely used in the field of remote sensing image analysis, mainly for detecting objects with obvious directions such as airplanes, ships, and buildings. By predicting oriented bounding boxes instead of traditional horizontal bounding boxes, such networks can more accurately locate dense objects in remote sensing images. For the positioning of optical communication device patch components, the rotation object detection technology is also of great significance. The rotation object detection for optical communication device patch components can not only obtain the position information of various patch components, but also obtain their rotation angles, which are crucial for measuring the patching quality of optical communication devices. However, most of the current research on rotation object detection networks is applied to remote sensing images. In the rotation object detection of optical communication device patch components, many challenges need to be faced, such as problems unique to industrial scenarios like small component size, high shape similarity, and surface reflection.
[0006] Therefore, how to create a high-precision rotation object detection method applicable to optical communication device patch components, accurately obtain the position and rotation angle information of components while improving the detection accuracy, and at the same time maintain the real-time performance and robustness of the algorithm has become a key technical problem that needs to be solved urgently.
[0007] Through the above analysis, the problems and defects existing in the prior art are as follows:
[0008] Most of the current research on rotation object detection networks is applied to remote sensing images. In the rotation object detection of optical communication device patch components, many challenges need to be faced, such as problems unique to industrial scenarios like small component size, high shape similarity, and surface reflection. Summary of the Invention
[0009] Aiming at the problems existing in the prior art, the present invention provides a fast positioning method for optical communication device patch components based on deep learning. It aims to solve problems such as inaccurate angle prediction of different types of patch components of optical communication devices by existing rotation detection technologies, and the increase in learning threshold caused by manually selecting different models for template matching and positioning.
[0010] The present invention is implemented as follows. A fast positioning method for optical communication device patch components based on deep learning includes:
[0011] Step 1, obtain coaxial light source images of different types of optical communication devices;
[0012] Step 2, input images of different types of optical communication devices into a rotation object detection network based on a convolutional neural network, and output the position and angle information of different types of patch components.
[0013] Furthermore, the rotation target detection network includes: a feature extraction backbone network, a two-stage feature focus diffusion fusion network, and a rotation detection head module with feature dynamic alignment.
[0014] Furthermore, the feature extraction backbone network is an improved CSPDarknet-53 structure, whose structure includes a basic convolutional downsampling module and an improved C3k2 module for layer-by-layer feature extraction, and finally outputs five groups of feature maps with different scales;
[0015] The basic convolutional downsampling module includes convolution, normalization, and activation functions for low-level feature extraction;
[0016] The improved C3k2 module includes convolution and Bottleneck modules. It uses the CSP (Cross Stage Partial Network) structure to separate the feature map into two parts. One part is directly passed, and the other part is processed through multiple Bottleneck blocks and finally feature fused. This module aims to enhance the expressive ability of feature extraction while reducing the computational amount;
[0017] The five groups of feature maps with different scales can be regarded as a pyramid structure. The extraction sizes of the feature maps output by different pyramid layers are as follows: P1 is 64 dimensions × 320 pixels × 320 pixels, P2 is 128 dimensions × 160 pixels × 160 pixels, P3 is 256 dimensions × 80 pixels × 80 pixels, P4 is 512 dimensions × 40 pixels × 40 pixels, and P5 is 1024 dimensions × 20 pixels × 20 pixels.
[0018] Furthermore, the two-stage feature focus diffusion fusion network is used to perform two processes of feature focus and diffusion on P3 (low-level features), P4 (mid-level features) output by the feature extraction backbone network, and P5 (high-level features) passing through the multi-scale two-stage feature enhancement module, so as to enhance the perception ability of multi-scale features; The network is divided into a focus path and a diffusion path. Among them, first, in the first focus diffusion process, the feature map enhanced by the multi-scale two-stage feature enhancement module of P5 (high-level features) and the feature maps of P3 (low-level features) and P4 (mid-level features) are input into the feature focus module for feature focus, and a focus feature map with the size of the mid-level feature map is output. Then, it is diffused through upsampling and convolutional downsampling, and the diffused results are respectively spliced with the high-level and low-level feature maps, and then the C3k2 module is used to further refine the multi-scale features, so as to obtain three feature maps with different scales; Then, the second focus diffusion process is carried out, that is, the three-scale feature maps just obtained are focused and diffused again, spliced with the results of the first-round diffusion, and then the C3k2 module is used to extract features. Finally, the three-sized feature maps are beneficial to the positioning of optical communication device patch components with different sizes in this article; After two focus diffusions, the information flows interactively in the network multiple times;
[0019] The multi-scale two-stage feature enhancement module includes a fast spatial pyramid pooling module and a C2BRA module, which are used to enhance the high-level semantic features of the feature extraction backbone network so as to maximize the integration of multi-scale information;
[0020] The input of the multi-scale two-stage feature enhancement module is the feature map output by the highest layer of feature extraction. After the original optical communication device image is downsampled multiple times, the resolution is low but it contains highly abstract semantic information. Introducing this module can effectively integrate the information of the global context;
[0021] The fast spatial pyramid pooling module consists of multiple max-pooling layers to efficiently aggregate the context features of different receptive fields through a serial structure;
[0022] The C2BRA module is used to focus on the key target regions in the deep features, and also uses the CSP (Cross Stage Partial Network) structure. One part is directly passed, and the other part undergoes two-stage routing self-attention for feature selection. Finally, the two paths are concatenated to obtain a high-level feature map with enhanced semantics;
[0023] The described feature focusing module is used to perform information interaction and fusion on feature maps of multiple scales. First, different convolutions are used to adjust the scales of feature maps at different levels and then they are concatenated to obtain a focused feature map. Then, a residual structure is used to divide the feature map into two branches. One branch uses a multi-scale convolution kernel group to capture the context information of different receptive fields, and the other branch directly outputs the multi-scale feature map. Finally, the two are fused to enhance the multi-scale expression ability of the feature map, and finally the focused feature is output.
[0024] Furthermore, the rotation detection head module with feature dynamic alignment is used to perform classification and localization tasks on the feature maps of three scales obtained by the two-stage feature enhancement fusion network and achieve angle prediction;
[0025] The rotation detection head module with feature dynamic alignment includes a task-aware decomposition module, a feature space offset calculation module, a dynamic deformable convolution module, a class probability perception module, and a bounding box prediction decoding module;
[0026] The task-aware decomposition module is used to decompose the input feature map into classification task features and regression task features, and enhance the task discrimination ability of feature representation through adaptive average pooling and channel attention mechanism;
[0027] The feature space offset calculation module is used to generate spatial offset and mask values according to the input feature map, providing sampling position parameters for the dynamic deformable convolution;
[0028] The dynamic deformable convolution module is used to perform adaptive spatial deformability on the regression task features with the generated spatial offset and mask values;
[0029] The category probability perception module is used to generate a category probability weight map by weighting the features of the classification task.
[0030] Furthermore, the bounding box prediction and decoding module is used to convert the dynamically aligned regression features and weighted classification features into the final bounding box coordinates, dimensions, rotation angle parameters and category probability values, realizing the precise positioning and classification of the patch components of the optical communication device.
[0031] Another object of the present invention is to provide a fast positioning system for patch components of optical communication devices based on deep learning, including:
[0032] An image acquisition module, configured to acquire coaxial light source images of different types of optical communication devices;
[0033] An input / output module, configured to input images of different types of optical communication devices into a rotation target detection network based on a convolutional neural network, and output the position and angle information of different types of patch components.
[0034] Another object of the present invention is to provide a computer device, characterized in that the computer device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the fast positioning method for patch components of optical communication devices based on deep learning.
[0035] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the processor executes the steps of the fast positioning method for patch components of optical communication devices based on deep learning.
[0036] Another object of the present invention is to provide an information data processing terminal, and the information data processing terminal is used to implement the fast positioning system for patch components of optical communication devices based on deep learning.
[0037] Combined with the above technical solutions and solved technical problems, please analyze the advantages and positive effects of the technical solution to be protected by the present invention from the following aspects:
[0038] First, the present invention provides a rotation target detection network based on a convolutional neural network, which uses an improved CSPDarknet-53 as the feature extraction backbone network. While maintaining a high detection accuracy, this network structure replaces ordinary convolutions with the proposed dynamic local convolutions. Compared with traditional feature extraction backbone networks, this feature extraction backbone network still maintains a high feature extraction ability while reducing the number of model parameters, and is particularly suitable for the real-time detection requirements on the optical communication device production line.
[0039] The present invention provides a rotation target detection network based on a convolutional neural network, in which a dual-stage feature focusing and diffusion fusion network is designed. The network introduces a multi-scale dual-stage feature enhancement module applied to the top-level feature map of the feature extraction network, effectively solving the problem of difficult detection in traditional target detection networks when dealing with complex scenarios such as patch components of optical communication devices with various shapes, large size spans, and indefinite rotation angles. Moreover, the network innovatively designs a "focusing-diffusion" architecture, enabling high-level semantic information and low-level detail features to interact and fuse multiple times. Through two rounds of "focusing-diffusion" processes, the network can more effectively integrate multi-scale feature information, and this structure significantly enhances the network's ability to perceive global context information.
[0040] The present invention provides a rotation target detection network based on a convolutional neural network, in which the rotation detection head module with feature dynamic alignment can accurately capture the position, size, and rotation angle information of patch components of optical communication devices through task-aware feature decomposition and dynamic feature alignment mechanisms. The module realizes the alignment of class and localization information in the prediction target through dynamic deformable convolution technology and adaptive rotation of classification features, solving the problem of inaccurate angle prediction of traditional detection heads when dealing with patch components at arbitrary angles.
[0041] The present invention provides a rotation target detection network based on a convolutional neural network, in which a class-square angle-sensitive loss function is used in the loss function. By adding an angle penalty term for class-square targets, it effectively solves the problem of unstable angle prediction regression of the probability IOU rotation box loss for class-square targets, significantly improving the angle prediction ability of class-square patch components such as photosensitive components of optical communication devices.
[0042] Second, as the creative auxiliary evidence of the claims of the present invention, it is also reflected in the following important aspects:
[0043] (1) The expected benefits and commercial value after the transformation of the technical solution of the present invention are:
[0044] (2) The technical solution of the present invention fills the technical gaps at home and abroad in the industry:
[0045] (3) Whether the technical solution of the present invention solves the technical problems that people have been eager to solve but have never succeeded in:
[0046] (4) Whether the technical solution of the present invention overcomes technical prejudices: BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 It is a flowchart of a method for quickly positioning patch components of optical communication devices based on deep learning provided by an embodiment of the present invention.
[0048] Figure 2 It is a structural block diagram of a fast positioning system for optical communication device patch components based on deep learning provided by an embodiment of the present invention.
[0049] Figure 3 It is a schematic diagram of a rotation target detection network based on a convolutional neural network provided by an embodiment of the present invention;
[0050] Figure 4 It is a schematic diagram of dynamic local convolution provided by an embodiment of the present invention;
[0051] Figure 5 It is a schematic diagram of a two-stage feature focusing diffusion fusion network provided by an embodiment of the present invention;
[0052] Figure 6 It is a visualization diagram of the effect of a two-stage feature focusing diffusion fusion network provided by an embodiment of the present invention;
[0053] Figure 7 It is a schematic diagram of a multi-scale two-stage feature enhancement module provided by an embodiment of the present invention;
[0054] Figure 8 It is a schematic diagram of a feature focusing module provided by an embodiment of the present invention;
[0055] Figure 9 It is a schematic diagram of a rotation detection head with feature dynamic alignment provided by an embodiment of the present invention;
[0056] Figure 10 It is a schematic diagram of the detection result of optical communication device patch components by a rotation target detection network provided by an embodiment of the present invention. Detailed implementation manners
[0057] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0058] As Figure 1 shown, a method for fast positioning of optical communication device patch components based on deep learning provided by an embodiment of the present invention includes the following steps:
[0059] S101, obtaining coaxial light source images of different types of optical communication devices;
[0060] S102, inputting images of different types of optical communication devices into a rotation target detection network based on a convolutional neural network, and outputting the position and angle information of different types of patch components.
[0061] The rotation target detection network provided by the embodiment of the present invention includes: a feature extraction backbone network, a two-stage feature focusing diffusion fusion network, and a rotation detection head module with feature dynamic alignment.
[0062] The feature extraction backbone network provided by the embodiment of the present invention is an improved CSPDarknet-53 structure, and its structure includes a basic convolutional downsampling module and an improved C3k2 module for layer-by-layer feature extraction, and finally outputs five groups of feature maps with different scales;
[0063] The basic convolutional downsampling module includes convolution, normalization, and activation functions for low-level feature extraction;
[0064] The improved C3k2 module includes convolution and Bottleneck modules. It uses the CSP (Cross Stage Partial Network) structure to separate the feature map into two parts. One part is directly passed, and the other part is processed through multiple Bottleneck blocks and finally feature fusion is performed. This module aims to reduce the computational amount while enhancing the expression ability of feature extraction;
[0065] The five groups of feature maps with different scales can be regarded as a pyramid structure. The extraction sizes of the feature maps output by different pyramid layers are: P1 is 64-dimensional × 320 pixels × 320 pixels, P2 is 128-dimensional × 160 pixels × 160 pixels, P3 is 256-dimensional × 80 pixels × 80 pixels, P4 is 512-dimensional × 40 pixels × 40 pixels, and P5 is 1024-dimensional × 20 pixels × 20 pixels.
[0066] The dual - level feature focusing and diffusion fusion network provided by the embodiments of the present invention is used to perform the process of two - level feature focusing and diffusion on P3 (low - level features), P4 (mid - level features), and P5 (high - level features) output by the feature extraction backbone network, so as to enhance the perception ability of multi - scale features. The network is divided into a focusing path and a diffusion path. First, in the first focusing and diffusion process, the feature map enhanced by the multi - scale dual - level feature enhancement module for P5 (high - level features) and the feature maps of P3 (low - level features) and P4 (mid - level features) are input into the feature focusing module for feature focusing, and a focused feature map with the size of the mid - level feature map is output. Then, through upsampling and convolutional downsampling for diffusion, the diffused results are respectively spliced with the high - level and low - level feature maps, and then the C3k2 module is used to further refine the multi - scale features, so as to obtain three different - scale feature maps. Then, the second focusing and diffusion process is carried out, that is, the three - scale feature maps obtained just now are focused and diffused again, and then spliced with the results of the first - round diffusion, and then the C3k2 module is used to extract features. Finally, the three - size feature maps are beneficial to the positioning of optical communication device patch components of different sizes in this article. After two - level focusing and diffusion, the information flows interactively in the network multiple times, greatly enhancing the network's perception ability for optical communication device patch components of different sizes.
[0067] The multi - scale dual - level feature enhancement module includes a fast spatial pyramid pooling module and a C2BRA module, which are used to enhance the high - level semantic features of the feature extraction backbone network, so as to maximize the integration of multi - scale information.
[0068] The input of the multi - scale dual - level feature enhancement module is the feature map output by the highest layer of feature extraction. After the original optical communication device image is downsampled multiple times, the resolution is low but it contains highly abstract semantic information. Introducing this module can effectively integrate the information of the global context and enhance the model's semantic understanding of the positioning target.
[0069] The fast spatial pyramid pooling module is composed of multiple max - pooling layers, which efficiently aggregates the context features of different receptive fields through a serial structure, enhancing the model's robustness to target scale and deformation.
[0070] The C2BRA module is used to focus on the key target areas in the deep features, and also uses the CSP (Cross - Stage Partial Network) structure. One part is directly passed, and the other part undergoes feature selection through dual - level routing self - attention. Finally, the two paths are spliced to obtain a semantically enhanced high - level feature map.
[0071] The described feature focusing module is used to perform information interaction and fusion on feature maps of multiple scales. First, different convolutions are used to adjust the scales of feature maps at different levels and then they are concatenated to obtain a focused feature map. Then, a residual structure is used to divide the feature map into two branches. One branch uses a multi-scale convolution kernel group to capture context information of different receptive fields, and the other branch directly outputs the multi-scale feature map. Finally, the two are fused to enhance the multi-scale expression ability of the feature map and ultimately output the focused feature.
[0072] The Feature Dynamically Aligned Detection Head for Oriented Bounding Boxes (FDADH-OBB) module provided by the embodiment of the present invention is used to perform classification, localization tasks and realize angle prediction on feature maps of three scales obtained by the two-stage feature enhancement fusion network;
[0073] The feature dynamically aligned rotation detection head module includes a task-aware decomposition module, a feature space offset calculation module, a dynamic deformable convolution module, a class probability perception module, and a bounding box prediction decoding module;
[0074] The task-aware decomposition module is used to decompose the input feature map into classification task features and regression task features, and enhance the task discrimination ability of feature representation through adaptive average pooling and channel attention mechanism;
[0075] The feature space offset calculation module is used to generate spatial offset and mask values according to the input feature map, and provide sampling position parameters for the dynamic deformable convolution;
[0076] The dynamic deformable convolution module is used to perform adaptive spatial deformation on the regression task features with the generated spatial offset and mask values, and enhance the network's perception ability of the rotation angle of the patch component and the accuracy of target localization regression;
[0077] The class probability perception module is used to generate a class probability weight map, and improve the network's detection ability for small target patch components by weighting the classification task features.
[0078] The bounding box prediction decoding module provided by the embodiment of the present invention is used to convert the dynamically aligned regression features and weighted classification features into final bounding box coordinates, sizes, rotation angle parameters and class probability values, and realize the precise positioning and classification of the patch components of the optical communication device;
[0079] In view of the shape and rotation characteristics of the patch components of optical communication devices, the present invention designs a square-like angle-sensitive rotation bounding box loss function, which solves the problem of unstable angle regression prediction of the commonly used probabilistic IOU loss function in the rotation target detection network for square targets.
[0080] As Figure 2 shown, a fast positioning system for patch components of optical communication devices based on deep learning provided by an embodiment of the present invention includes:
[0081] An image acquisition module for acquiring coaxial light source images of different types of optical communication devices;
[0082] An input-output module for inputting images of different types of optical communication devices into a rotation target detection network based on a convolutional neural network and outputting the position and angle information of different types of patch components.
[0083] Another object of the present invention is to provide a computer device, characterized in that the computer device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method for quickly positioning patch components of optical communication devices based on deep learning.
[0084] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the processor executes the steps of the method for quickly positioning patch components of optical communication devices based on deep learning.
[0085] Another object of the present invention is to provide an information data processing terminal for implementing the fast positioning system for patch components of optical communication devices based on deep learning.
[0086] Specific implementation of the present invention:
[0087] Embodiment 1
[0088] A rotation target detection network for patch components of optical communication devices based on a convolutional neural network, including a feature extraction backbone network, a two-stage feature enhancement fusion network, and a rotation detection head module with feature dynamic alignment;
[0089] The feature extraction backbone network adopts an improved CSPDarknet-53 structure, which is composed of multiple basic convolutional downsampling modules and an improved C3k2 module. This network is used to extract multi-scale features and outputs three groups of feature maps with different scales for the detection of targets with different scales;
[0090] The two-stage feature focusing diffusion fusion network includes a multi-scale two-stage feature enhancement module and a feature focusing module.
[0091] The multi-scale two-stage feature enhancement module receives the highest-level feature map P5 output by the feature extraction backbone network, and enhances the feature expression ability through the fast spatial pyramid pooling module and the C2BRA module.
[0092] The feature focusing module receives three feature maps of different scales: P5' (high-level feature) output by the multi-scale two-stage feature enhancement module, P4 (middle-level feature) and P3 (low-level feature) output by the feature extraction backbone network, and fuses the three feature maps of different scales into a unified-scale focused feature.
[0093] The rotation detection head for feature dynamic alignment is used to align the target position localization and classification tasks for the three-scale feature maps obtained by the two-stage feature focusing diffusion fusion network, and simultaneously predict the angle of the target.
[0094] Based on the problems existing in the prior art, there is an urgent need to develop a rotation target detection method that can not only cope with the detection challenges of the variable rotation angles and large size differences of the patch components of optical communication devices, but also improve the generalization ability of the model for different types of patch components, so as to achieve the accurate positioning and angle prediction of various patch components of optical communication devices.
[0095] As Figure 3 shown, the embodiment of the present invention discloses a rotation target detection network based on a convolutional neural network, which is used to detect patch components with different shapes, large size differences and uncertain rotation angles in optical communication devices, and can simultaneously accurately locate the positions, sizes and rotation angles of the same type of patch components, ensure the positioning accuracy of production line defect detection and quality detection, and significantly improve the automation level and efficiency of optical communication device production. By adopting a rotation detection head with feature dynamic alignment, the network can align the localization task and the classification task, and improve the localization accuracy of the network prediction box. Among them, the difficulty of rotation target detection of patch components of optical communication devices lies in not only the diversity of scales and shapes, but also the need to accurately predict different angles of the target.
[0096] The feature extraction backbone network of the rotation target detection network is composed of an improved CSPDarknet-53 network. Among them, the scales of the three feature maps required for feature extraction are 256 dimensions × 80 pixels × 80 pixels, 512 dimensions × 40 pixels × 40 pixels, and 1024 dimensions × 20 pixels × 20 pixels respectively.
[0097] In the embodiment of this application, the CSPDarknet-53 network refers to the backbone extraction network in YOLO11. On this basis, the present invention improves the standard C3K2 structure therein, and proposes a Dynamic Partial Convolution (DP Conv) to replace the standard convolution in the C3k2 module. By replacing the standard convolution in C3k2 with the DPConv structure, the computational complexity of the model is significantly reduced. At the same time, to improve the feature extraction ability, the present invention retains the skip connection and residual structure of CSPDarknet-53 to ensure the effective transmission of information in the network.
[0098] In the embodiment of this application, the proposed Dynamic Partial Convolution DP Conv combines three core components: Partial Convolution (PConv), ReChannel, and PointWise Convolution (PWConv), as Figure 4 shown. It forms a more computationally efficient convolutional structure. The calculation formula of the dynamic partial convolution can be expressed as: Y = PWConv(Shuffle(PConv(X in ))). Where it means first passing through the partial convolution PConv, then through the Shuffle channel rearrangement, and finally performing the PointWise Convolution PWConv.
[0099] Specifically, the partial convolution is used to perform a directional convolution operation on the selected channels of the input feature map, and only applies a k×k convolution kernel (k = 3 in this embodiment) to c p channels (c = 1 / 4c in this embodiment, where c is the total number of channels), significantly reducing the computational complexity while retaining local spatial information. Assuming the size of the input image is (c, h, w), the size of the convolution kernel is k, and the output size is (c, h, w), the computational complexity of ordinary convolution is h×w×k p ×c 2 For partial convolution, it is 2 The channel rearrangement unit is used to reorganize the features in the channel dimension after partial convolution, rearrange the features according to a preset pattern, enhance the information interaction between channels, and provide a better feature representation for subsequent global feature fusion;
[0100] The global feature fusion unit is used to perform global information integration on the rearranged features through 1×1 pointwise convolution, restore the original channel dimension c, establish cross-channel dependencies, and output enhanced features with the same size as the input feature map. The running parameter quantity of this convolution is h×w×c
[0101] 2 .
[0102] The computational complexity of the dynamic partial convolution DP Conv is finally Among them The computational load is significantly reduced, effectively balancing the relationship between computational efficiency and detection accuracy.
[0103] Optionally, the dual-level feature focusing diffusion fusion network is as Figure 5 shown. This network adopts an innovative dual-wheel focusing diffusion architecture to achieve efficient interconnection and fusion of multi-scale features. First, a multi-scale dual-level feature enhancement module is applied to the high-level features extracted by the feature extraction backbone network, and then, together with P4 (mid-level features) and P3 (low-level features) output by the feature extraction backbone network, they are input into the feature focusing module. Through the dual-wheel cascading process of "focusing - diffusion - refocusing - rediffusion", the feature information fully interacts and flows between different scales, significantly enhancing the network's detection ability for patch components of different sizes in optical communication devices. The experimental visualization effect of this network is as Figure 6 shown.
[0104] The multi-scale dual-level feature enhancement module includes a multi-scale feature pooling module and a C2BRA module;
[0105] The multi-scale dual-level feature enhancement module is as Figure 7 shown. First, feature multi-scale perception and fusion are performed through the multi-scale pyramid pooling module SPFF to capture scene information under different receptive fields. Subsequently, the pooled features are restored to the original resolution through nearest neighbor upsampling and concatenated with the original features in the channel dimension to form a preliminary multi-scale feature representation. Then, the C2BRA module is used to selectively enhance the features to generate the final enhanced features.
[0106] The C2BRA module divides the input feature map into two parts: one part is directly passed to retain the original feature information, and the other part is processed through a dual-level routing attention mechanism to capture long-range dependencies. This module effectively integrates the CSP (Cross Stage Partial Network) idea and the dual-level routing attention mechanism, significantly enhancing the feature expression ability while reducing the computational load.
[0107] In the C2BRA module in the embodiments of the present invention, the features are first divided into two branches in the channel dimension: a direct transmission branch and an attention processing branch. The direct transmission branch retains the original feature information unchanged, while the attention processing branch is enhanced through n cascaded dual-level feature selection modules. Finally, the features of the two branches are concatenated and mapped back to the original number of channels through a 1×1 convolution. This CSP design significantly reduces the computational load and retains rich feature representations through cross-stage feature fusion.
[0108] Each double - level feature selection module contains a double - level routing attention module, and the principle of this module is as follows: The double - level routing attention mechanism includes attention calculations at two levels, namely the region - level and the token - level, which solves the computational efficiency problem of the traditional self - attention mechanism when dealing with high - resolution feature maps.
[0109] In addition, the double - level attention mechanism also introduces a local context enhancement mechanism. It captures local spatial context information through depth - separable convolutions (side - depth separable convolutions, with a convolution kernel size of 3 in this embodiment), adds it to the attention output, and further enhances the feature representation ability. Finally, an enhanced feature map is output through 1×1 convolution.
[0110] The Feature Focus Module is one of the core innovations of the present invention. As Figure 8 shown, it includes the following steps: Input the three different - scale feature maps output by the feature extraction backbone network and the multi - scale double - level feature enhancement module into the Feature Focus Module. The Feature Focus Module first adjusts the input high - level feature P5 (scale: 1024 - dimensional × 20 pixels × 20 pixels) to a feature map with a scale of 256 - dimensional × 40 pixels × 40 pixels through 2 - fold upsampling and 1×1 convolution operations, adjusts the middle - level feature P4 (scale: 512 - dimensional × 40 pixels × 40 pixels) to 256 - dimensional in terms of channel number through 1×1 convolution, and adjusts the low - level feature P3 (scale: 256 - dimensional × 80 pixels × 80 pixels) to a feature map with a scale of 256 - dimensional × 40 pixels × 40 pixels through an adaptive downsampling operation. Then, these three feature maps with unified scale and channel number are concatenated in the channel dimension to obtain a multi - scale fusion feature map with a scale of 768 - dimensional × 40 pixels × 40 pixels. Next, the Feature Focus Module processes the concatenated feature map through four parallel depth - separable convolution branches, using convolution kernels with sizes of 5×5, 7×7, 9×9, and 11×11 respectively to capture receptive field information at different scales. These multi - receptive - field features are fused with the original concatenated features, and then channel - by - channel information integration is performed through 1×1 point - wise convolution, and finally a feature focus result with a scale of 768 - dimensional × 40 pixels × 40 pixels is output. After being processed by the Feature Focus Module, the feature map effectively integrates the multi - scale context information and semantic information in the three - scale feature maps, especially achieving more accurate and comprehensive feature extraction for optical communication device patch components of different sizes, significantly enhancing the network's detection ability for small patch components and the angle prediction accuracy for large patch components.
[0111] Optionally, as Figure 9 The rotation detection head module for feature dynamic alignment includes a shared convolution module, a task - aware decomposition module, a feature space offset calculation module, a dynamic deformable convolution module, a class feature selection module, and a prediction regression module;
[0112] The shared convolution module includes two 3×3 convolution layers with group normalization, which are used to extract the basic semantic information of the feature map. The outputs of each convolution layer are concatenated in the channel dimension through the residual connection method to form enhanced features;
[0113] The task-aware decomposition module is used to decompose the enhanced features into classification features and localization regression features. First, global adaptive average pooling is performed on the enhanced features to obtain global context information, and then feature enhancement weights for classification and regression tasks are generated through a task-specific attention subnetwork and applied to the original features respectively, so as to obtain task-specific feature representations;
[0114] The feature space offset calculation module is used to process the enhanced features through 3×3×3 convolution to generate spatial offset amounts and mask values for dynamic convolution. The offset amount is 2×3×3 dimensions (representing the x and y direction offsets of 9 sampling points), and the mask value is 1×3×3 dimensions (representing the weights of 9 sampling points), providing adaptive sampling positions for dynamic deformable convolution;
[0115] The dynamic deformable convolution module performs spatial adaptive transformation on the regression features based on the generated spatial offset amounts and mask values, dynamically adjusts the sampling positions of the convolution kernels according to the content of the input features, enables the network to better adapt to the geometric characteristics of the rotated patch components, and improves the prediction accuracy of the bounding box coordinates and rotation angles;
[0116] The class feature selection module consists of two convolution layers. First, the enhanced features are reduced to 1 / 4 of the original channel number through 1×1 convolution, and then a single-channel probability map is generated through 3×3 convolution and processed by the sigmoid function to obtain weight values in the range of 0 to 1, which are used to spatially weight the classification features to obtain probability-weighted classification features, making the network pay more attention to the regions that may contain patch components;
[0117] The prediction regression module includes two parallel 1×1 convolutional layers, which are used for bounding box prediction and class prediction respectively. The bounding box prediction branch processes the dynamically aligned regression features using a parameter-learnable scaling layer and outputs features with 4×reg_max channels, where reg_max represents a multiple of the number of feature map channels and determines the representation ability and detection performance of each localization box (in this embodiment, reg_max = 16), representing the distribution representation of the central point coordinates, width, and height parameters of the bounding box; the class prediction branch applies a 1×1 convolution to the probability-weighted classification features and outputs features with nc channels, which are used to represent the probability distributions of nc classes; the rotation angle prediction is achieved through an additional convolutional branch, mapping the angle parameter to the range of [-π / 4, 3π / 4] through the sigmoid function, and finally obtaining the predicted rotated detection box as (x, y, w, h, θ), realizing the accurate prediction of the rotation angle of the patch components of the optical communication device.
[0118] In this specific example, the following steps are included:
[0119] Obtain the rotation target detection dataset of the patch components of the optical communication device. In this embodiment, for the image acquisition stage in the rotation detection process of the patch components of the optical communication device, 8 types of patch component samples on 5 different types of optical communication devices are collected, with a total of 213 sample images. 3 images of each model are taken as the test set, and then the remaining 198 sample images are subjected to data augmentation (including random rotation, scaling, cropping, etc.) to make the dataset size reach 1188 images. The sample image format is BMP, and the original size of each sample image is 4608 pixels × 3288 pixels. At the same time, each patch component is labeled, and finally, the training set and test set in the dataset are allocated according to the ratio of 8:2.
[0120] Build a rotation detection model for the patch components of the optical communication device based on a convolutional neural network. The model adopted in this embodiment includes a feature extraction backbone network (improved CSPDarknet-53 network), a two-stage feature focusing diffusion fusion network, and a rotation detection head with feature dynamic alignment.
[0121] For the rotation target detection task, the network structure of this embodiment proposes a loss function AF ProbIOU Loss (Angle Focused ProbIOU Loss) that is sensitive to the class square angle. An angle-sensitive penalty term is added to the widely used ProbIOU loss function in current rotation target detection. The calculation of this loss function mainly includes three modules: class square detection, angle difference calculation, and loss fusion.
[0122] The quasi-square detection module is used to identify patch components close to a square. By calculating the aspect ratios r1 = w1 / h1 and r2 = w2 / h2 of the predicted bounding box and the ground truth bounding box, and judging whether the conditions of 0.9 < r1 < 1.1 and 0.9 < r2 < 1.1 are satisfied simultaneously, a binary mask M(r1, r2) is generated. When the patch component is close to a square, M(r1, r2) = 1, otherwise M(r1, r2) = 0, achieving precise screening of quasi-square targets;
[0123] The angle difference calculation module is used to solve the periodic confusion problem that the shape of quasi-square patch components hardly changes after a 90° rotation. By calculating the difference |θ1 - θ2| between the predicted angle θ2 and the ground truth angle θ1, and using the remainder operation to map it to the range of [0, π / 2], the normalized angle difference |θ1 - θ2| mod(π / 2) is obtained. Then through the formula the angle difference is converted into a non-linear angle penalty term, where α is an angle sensitivity factor that controls the intensity of the penalty;
[0124] The loss fusion module is used to combine the angle penalty term with the ProbIoU loss, and apply the angle penalty only to quasi-square targets. Through the formula L AF-PROBIOU = 1 - ProbIoU×(1 - δ(θ1, θ2)×M(r1, r2)) to achieve selective application of the penalty. When the detected target is quasi-square (M = 1), the loss function will increase the penalty according to the degree of angle difference; when the target is not quasi-square (M = 0), the loss function degrades to the Prob IoU loss, maintaining the detection accuracy for non-square targets.
[0125] In this embodiment, the stochastic gradient descent (SGD) algorithm is used as the model optimizer. The patch component rotation detection network is trained for 100 epochs, with an initial learning rate of 0.01. The cosine annealing strategy is adopted for learning rate adjustment, with a weight decay coefficient of 0.0005 and a momentum coefficient of 0.937.
[0126] The backbone network of the patch component rotation detection model for optical communication devices designed in this embodiment is the improved CSPDarknet-53 network, in which the standard convolution part is replaced by dynamic local convolution.
[0127] The patch component rotation detection model for optical communication devices designed in this embodiment also uses a two-stage feature focus fusion network, in which a multi-scale two-stage enhancement module and a focus module are used to achieve the fusion and enhancement of multi-scale information
[0128] The patch component rotation detection model for optical communication devices designed in this embodiment also uses a rotation detection head with feature dynamic alignment. By adopting task decomposition and dynamic deformable convolution technologies, the accuracy of bounding box localization is significantly improved.
[0129] Write a Python program to randomly sort the training samples and then evenly distribute them. The batch size for batch training is 4 (unit: piece). Use the training set in the optical communication device chip component dataset to train the rotation detection network, and save the weight parameters of the optimized network model.
[0130] Specifically, in this embodiment, the training images are input into the model batch by batch. After the sum of the gradient descents of all samples in a batch is calculated, a weight update is performed once until all batches are updated. Use the validation set to evaluate the trained model, and calculate the mean average precision (mAP@0.5:0.95) of the network on the validation set. If the mAP is greater than the existing maximum mAP, store the weight parameters of the network model in this round and perform the next iteration; if it is less than the maximum mAP, directly perform the next iteration. At the same time, set the number of iterations of the network model to 100 times, and end the training after completing the number of iterations. Obtain the optimal network model and name it A2RDet_best.
[0131] Load the network weights corresponding to A2RDet_best, test the network performance on the test set, and evaluate the positioning accuracy and angle prediction accuracy of the rotation detection model on the validation set. Finally, the mAP@0.5:0.95 of the test set is 0.984, and the FPS is 105.84, realizing efficient and accurate rotation target positioning of optical communication device chip components, as Figure 10 shown
[0132] In the embodiment of the present invention, by inputting the multi-scale feature maps obtained through the feature extraction backbone network into the rotation detection network, using the multi-scale two-stage feature enhancement module to extract the context information with rich high-level semantics and long-distance dependence relationships, enabling the information to fully interact and flow between different scales through the feature focus diffusion fusion network, and finally accurately predicting the position and angle parameters of the chip components through the rotation detection head with feature dynamic alignment. It solves the technical problem that the existing detection technology has low detection accuracy in the case of variable angles, different sizes, and dense arrangements of optical communication device chip components, resulting in inaccurate positioning and affecting the subsequent defect detection and quality assessment of optical communication devices. It realizes the accurate detection of various rotating chip components and accurate angle prediction, greatly improves the detection accuracy of optical communication device chip components, realizes the precise and rapid detection of optical communication device chip components in complex backgrounds, and provides reliable technical support for the quality control of optical communication devices.
[0133] Embodiment 2
[0134] A method for detecting the rotation of patch components of an optical communication device based on feature enhancement is applied to the network structure described in any one of the first embodiments and includes: S1. Obtain an image of the optical communication device, and after preprocessing, obtain a to-be-detected image with a standardized size of 640 pixels × 640 pixels; S2. Input the to-be-detected image into a rotation detection network for patch components of an optical communication device based on a convolutional neural network, and output the position coordinates, size, category, and rotation angle of the patch components.
[0135] An embodiment of the present invention provides a method for detecting the rotation of patch components of an optical communication device based on a convolutional neural network, which is applied to execute a rotation detection network structure for patch components of an optical communication device based on a convolutional neural network, and has the same functions and beneficial effects. It can solve the problems of low detection accuracy and large angular error caused by diverse component types, variable angles, dense arrangements, and complex backgrounds in the rotation detection task of patch components of optical communication devices, as well as the problem of accuracy loss caused by the inability of traditional non-rotation detection methods to accurately identify patch components in a specific direction.
[0136] I. The specific application field or related products of the present invention.
[0137] II. Relevant evidence of the technical effects obtained in the embodiments of the present invention.
[0138] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated designed hardware. Those of ordinary skill in the art can understand that the above devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, such code is provided on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuits of programmable hardware devices such as very large scale integrated circuits or gate arrays, semiconductors such as logic chips and transistors, or field programmable gate arrays and programmable logic devices, can also be implemented by software executed by various types of processors, or can be implemented by a combination of the above hardware circuits and software, such as firmware.
[0139] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any modification, equivalent replacement, and improvement made within the spirit and principle of the present invention by those skilled in the art within the technical scope disclosed by the present invention shall be covered by the protection scope of the present invention.
Claims
1. A fast positioning method for optical communication device patch components based on deep learning, characterized in that, It includes the following steps: Step 1: Obtain coaxial light source images of different types of optical communication devices; Step 2: Input the images of different types of optical communication devices into a rotation target detection network based on a convolutional neural network, and output the position and angle information of different types of patch components.
2. The rapid positioning method of the optical communication device patch element based on deep learning according to claim 1, characterized in that The rotation target detection network includes: a feature extraction backbone network, a two-stage feature focusing and diffusion fusion network, and a rotation detection head module with feature dynamic alignment.
3. The method for quickly positioning the patch components of the optical communication device based on deep learning according to claim 2, wherein The feature extraction backbone network is an improved CSPDarknet-53 structure, and its structure includes a basic convolutional downsampling module and an improved C3k2 module for layer-by-layer feature extraction, and finally outputs five groups of feature maps with different scales; The basic convolutional downsampling module includes convolution, normalization, and activation functions for low-level feature extraction; The improved C3k2 module includes convolution and Bottleneck modules. It uses the CSP (Cross Stage Partial Network) structure to separate the feature map into two parts. One part is directly passed, and the other part is processed through multiple Bottleneck blocks and finally feature fused. This module aims to reduce the computational amount while enhancing the expression ability of feature extraction; The five groups of feature maps with different scales can be regarded as a pyramid structure. The extraction sizes of the feature maps output by different pyramid layers are: P1 is 64 dimensions × 320 pixels × 320 pixels, P2 is 128 dimensions × 160 pixels × 160 pixels, P3 is 256 dimensions × 80 pixels × 80 pixels, P4 is 512 dimensions × 40 pixels × 40 pixels, and P5 is 1024 dimensions × 20 pixels × 20 pixels.
4. The method for quickly positioning patch components of an optical communication device based on deep learning according to claim 2, wherein The two-stage feature focusing and diffusion fusion network is used to perform two processes of feature focusing and diffusion on P3 (low-level features), P4 (mid-level features) output by the feature extraction backbone network, and P5 (high-level features) passing through a multi-scale two-stage feature enhancement module, so as to enhance the perception ability of multi-scale features; The network is divided into a focusing path and a diffusion path. Among them, first, in the first focusing and diffusion process, the feature map obtained by enhancing the P5 (high-level features) through the multi-scale two-stage feature enhancement module and the P3 (low-level features) and P4 (mid-level features) feature maps are input into the feature focusing module for feature focusing, and a focused feature map with the size of the mid-level feature map is output. Then, it is diffused through upsampling and convolutional downsampling, and the diffused results are respectively concatenated with the high-level and low-level feature maps, and then the C3k2 module is used to further refine the multi-scale features, so as to obtain three feature maps with different scales; Then, the second focusing and diffusion process is carried out, that is, the three-scale feature maps just obtained are focused and diffused again, and then concatenated with the results of the first-round diffusion, and then the C3k2 module is used to extract features. Finally, the three-sized feature maps are beneficial to the positioning of patch components of different-sized optical communication devices in this article; After two rounds of focusing and diffusion, the information flows interactively in the network multiple times; The multi-scale two-stage feature enhancement module includes a fast spatial pyramid pooling module and a C2BRA module, which are used to enhance the high-level semantic features of the feature extraction backbone network so as to maximize the integration of multi-scale information; The input of the multi-scale two-stage feature enhancement module is the feature map output by the highest layer of feature extraction. After the original optical communication device image is downsampled multiple times, the resolution is low but it contains highly abstract semantic information. Introducing this module can effectively integrate the information of the global context; The fast spatial pyramid pooling module is composed of multiple max-pooling layers to efficiently aggregate the context features of different receptive fields through a serial structure; The C2BRA module is used to focus on the key target areas in the deep features, and also uses the CSP (Cross Stage Partial Network) structure. One part is directly passed, and the other part undergoes two-stage routing self-attention for feature selection. Finally, the two paths are concatenated to obtain a high-level feature map with enhanced semantics; The feature focusing module is used to perform information interaction and fusion on feature maps of multiple scales. First, different convolutions are used to adjust the scales of feature maps at different levels and then they are concatenated to obtain a focused feature map. Then, a residual structure is used to divide the feature map into two branches. One branch uses a multi-scale convolution kernel group to capture the context information of different receptive fields, and the other branch directly outputs the multi-scale feature map. Finally, the two are fused to enhance the multi-scale expression ability of the feature map, and finally the focused feature is output.
5. The method for quickly positioning the patch components of the optical communication device based on deep learning according to claim 2, wherein The feature dynamic alignment rotation detection head module is used to perform classification and localization tasks on the feature maps of three scales obtained by the two-stage feature enhancement fusion network and achieve angle prediction; The feature dynamic alignment rotation detection head module includes a task-aware decomposition module, a feature space offset calculation module, a dynamic deformable convolution module, a class probability perception module, and a bounding box prediction decoding module; The task-aware decomposition module is used to decompose the input feature map into classification task features and regression task features, and enhance the task discrimination ability of feature representation through adaptive average pooling and channel attention mechanism; The feature space offset calculation module is used to generate spatial offset and mask values according to the input feature map, and provide sampling position parameters for the dynamic deformable convolution; The dynamic deformable convolution module is used to perform adaptive spatial deformability on the regression task features with the generated spatial offset and mask values; The class probability perception module is used to generate a class probability weight map by weighting the classification task features.
6. The method for rapid positioning of optical communication device patch components based on deep learning according to claim 5, characterized in that, The bounding box prediction decoding module is used to convert the dynamically aligned regression features and weighted classification features into the final bounding box coordinates, dimensions, rotation angle parameters, and class probability values, so as to achieve precise positioning and classification of the patch components of the optical communication device.
7. A rapid positioning system for optical communication device patch components based on deep learning, which implements the rapid positioning method for optical communication device patch components based on deep learning according to any one of claims 1-6, characterized in that, The fast positioning system for patch components of optical communication devices based on deep learning includes: An image acquisition module, which is used to acquire coaxial light source images of different types of optical communication devices; An input-output module, which is used to input images of different types of optical communication devices into a rotation target detection network based on a convolutional neural network, and output the position and angle information of different types of patch components.
8. A computer device, characterized in that, The computer device includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor is caused to execute the steps of the method for rapid positioning of patch components of an optical communication device based on deep learning as described in any one of claims 1-6.
9. A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the processor is caused to execute the steps of the method for rapid positioning of patch components of an optical communication device based on deep learning as described in any one of claims 1-6.
10. An information data processing terminal, characterized in that, The information data processing terminal is used to implement the system for rapid positioning of patch components of an optical communication device based on deep learning as described in claim 7.
Citation Information
Cited By
Chip patch positioning method and system based on machine vision
CN121190314A
Chip pasting positioning method and system based on machine vision
CN121190314B
Visual positioning method for position and posture of micro device facing semiconductor substrate
CN122192290A