Unsupervised domain adaptation connector small target detection method and system

CN122821101APending Publication Date: 2026-09-25山东航空学院
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611109881.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-24
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

该类方法能够降低目标域人工标注依赖,但在装配体连接件小目标检测任务中仍存在如下不足:一是连接件在整幅图像中的像素占比小,纹理和边缘信息较弱,训练过程中卷积下采样后局部特征容易丢失;二是在无监督领域自适应方法中,为解决当前跨域迁移特征能力不足的问题,可采用图像等分的混合裁剪策略将源域图像和目标域图像内容混合并参与训练,但该类方法易产生目标结构不完整、标签污染或混合区域遮挡已有连接件的问题;三是现有基于对抗学习跨域特征对齐方法多偏向图像级或实例级整体对齐,难以充分关注前景小目标及迁移困难区域

Benefits of technology

[0024]上述无监督领域自适应连接件小目标检测方法及系统,通过获取包含连接件标签的源域装配体图像和无标签的目标域装配体图像并分别进行强弱数据增强,能够为跨域连接件检测任务构建具有标注信息且视觉分布差异可控的多分支训练数据基础;通过构建包含学生网络和教师网络的领域自适应检测框架并对弱增强目标域数据进行推理以生成连接件伪标签,能够为无标签目标域提供可靠的监督信号以降低对人工标注的依赖;基于源域连接件标签和目标域伪标签通过交并比约束的自适应混合裁剪策略对源域和目标域图像进行裁剪粘贴以生成跨域混合图像和跨域混合图像对应的混合标签,能够促进源域与目标域之间的内容交互和视觉分布融合,并确保混合图像中连接件目标结构的完整性和标签的语义一致性;通过嵌入C3CBAM特征增强模块的YOLOv5L骨干网络对源域数据、强增强目标域数据和跨域混合图像进行前向推理,能够增强小尺寸连接件的局部纹理和边缘特征表达,并抑制复杂背景干扰;通过根据连接件伪标签和目标域特征图生成多尺度注意力图并对目标域特征图进行前景区域加权聚合以得到前景特征对齐损失,利用图像级和实例级域判别器对源域特征图和目标域特征图进行对抗学习对齐以得到图像级和实例级对抗对齐损失,能够从前景细粒度、全局场景和局部实例多个层级同步缩小源域与目标域之间的特征分布差异;通过构建包含多分支监督损失和多重对齐损失的总损失函数并联合优化学生网络参数,同时以指数滑动平均策略更新教师网络参数,能够实现对检测主任务和域适应任务的协同训练并获得高精度跨域检测模型;在测试阶段通过训练完成的目标检测模型对待检测目标域装配体图像进行推理,输出连接件小目标检测结果,显著提升无标注目标域场景下装配体连接件小目标检测的准确性和稳定性。采用本方法能够在无需目标域人工标注的条件下实现跨域连接件检测知识的有效迁移,有效降低装配过程监测中人工标注成本,并提升对光照、材质、拍摄角度和背景分布等域偏移因素的抗干扰能力,支撑高效稳健的装配体连接件实时检测与装配质量监控。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821101A_ABST
    Figure CN122821101A_ABST
Patent Text Reader

Abstract

The application relates to a kind of unsupervised field self-adapting connector small target detection method and system, the method comprises: obtaining labeled source domain and unlabeled target domain assembly image and performing differential data enhancement;Teacher-student network framework is constructed, and target domain connector pseudo-label is generated by teacher network;Cross-domain mixed samples are generated by adaptive mixed clipping of intersection over union constraint;Features are extracted and detected by embedding C3CBAM in YOLOv5L backbone network;Feature domain self-adaptation is realized by combining multi-scale attention and multi-level confrontation alignment;Iterative optimization is obtained by combining multiple losses to obtain a detection model, and target domain small target inference detection is completed. Using the method can reduce the labeling cost, improve the small target detection precision and generalization ability, and adapt to the assembly process monitoring demand.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of machine vision inspection and artificial intelligence technology, and in particular relates to an unsupervised domain adaptive method and system for detecting small targets in connectors. Background Technology

[0002] In manual assembly lines for mechanical products, bolts, nuts, and other connectors are typically numerous, small in size, and installed in dispersed locations that are easily obscured by surrounding structures. During assembly, missing, incorrect, or improperly installed connectors directly impact product quality and safety. Current workshop quality management methods largely rely on manual verification or random checks after assembly, which are significantly lagging and make it difficult to promptly detect and correct abnormal operations during assembly.

[0003] Object detection technology can output the location and category of objects in images in real time, which can be used for assembly process monitoring. However, most existing deep learning-based assembly connector detection methods use supervised training, requiring the annotation of target boxes and category labels on a large number of real-world assembly images. For production scenarios of small-batch personalized products, manual annotation is costly, and different workshop environments, shooting angles, lighting conditions, and material surface conditions can cause significant domain shifts, resulting in insufficient cross-scene generalization ability of the model.

[0004] Unsupervised adaptive object detection methods typically utilize jointly trained labeled source domain data and unlabeled target domain data to enable the model to learn domain-invariant features and adapt to the target domain distribution, thus achieving the detection task for unlabeled target domain data. While these methods reduce reliance on manual annotation of the target domain, they still have the following shortcomings in small object detection tasks involving assembly connectors: First, connectors account for a small percentage of pixels in the overall image, with weak texture and edge information, making it easy to lose local features after convolutional downsampling during training. Second, in unsupervised adaptive methods, to address the current insufficient cross-domain feature transfer capability, a hybrid cropping strategy involving equal image division can be used to mix the source and target domain images for training; however, this approach is prone to problems such as incomplete target structure, label contamination, or occlusion of existing connectors by the mixed region. Third, existing cross-domain feature alignment methods based on adversarial learning tend to focus on image-level or instance-level overall alignment, making it difficult to adequately address small foreground objects and regions difficult to transfer. Summary of the Invention

[0005] Therefore, it is necessary to address the aforementioned technical problems by providing an unsupervised, adaptive method and system for detecting small targets such as bolts and nuts in complex assembly scenarios. This method and system can reduce the cost of manual annotation of the target domain and improve the detection accuracy of small-sized connectors such as bolts and nuts in complex assembly scenarios by utilizing both labeled source domain assembly images and unlabeled target domain assembly images.

[0006] Firstly, this application provides an unsupervised, domain-adaptive small target detection method for connectors, including:

[0007] S1. Obtain the source domain assembly image containing connector labels and the target domain assembly image without labels. Perform strong data augmentation on the source domain assembly image to obtain the augmented source domain data. Perform weak data augmentation and strong data augmentation on the target domain assembly image to obtain weakly augmented target domain data and strongly augmented target domain data.

[0008] S2. Construct a domain adaptive detection framework that includes student and teacher networks. Input weakly enhanced target domain data into the teacher network in the domain adaptive detection framework for inference and generate pseudo-labels for connectors in the target domain assembly image.

[0009] S3. Based on the connector labels and connector pseudo-labels of the source domain assembly image, the source domain assembly image and the target domain assembly image are cropped and pasted using an adaptive hybrid cropping strategy based on the intersection-union ratio constraint to generate a cross-domain hybrid image and the corresponding hybrid label.

[0010] S4. Input the enhanced source domain data, strongly enhanced target domain data, and cross-domain hybrid image into the student network of the domain adaptive detection framework. The student network performs forward inference on the enhanced source domain data, strongly enhanced target domain data, and cross-domain hybrid image through the YOLOv5L backbone network with embedded C3CBAM feature enhancement module, to obtain source domain feature map, target domain feature map, and hybrid image feature map, as well as source domain detection prediction results corresponding to the source domain feature map, strongly enhanced target domain detection prediction results corresponding to the target domain feature map, and hybrid image detection prediction results corresponding to the hybrid image feature map.

[0011] S5. Based on the connector pseudo-labels and the target domain feature map, generate a multi-scale attention map. Based on the multi-scale attention map, perform foreground region weighted aggregation on the target domain feature map to obtain the foreground feature alignment loss. Then, perform adversarial learning alignment on the global features of the source domain feature map and the global features of the target domain feature map through an image-level domain discriminator to obtain the image-level adversarial alignment loss. Finally, perform adversarial learning alignment on the instance features of the source domain feature map and the instance features of the target domain feature map through an instance-level domain discriminator to obtain the instance-level adversarial alignment loss.

[0012] S6. Based on the connector labels, hybrid labels, connector pseudo labels, image-level adversarial alignment loss, foreground feature alignment loss, instance-level adversarial alignment loss, source domain detection prediction results, strongly enhanced target domain detection prediction results, and hybrid image detection prediction results of the source domain assembly image, a total loss function is constructed. The parameters of the student network are updated by optimizing the total loss function, and the parameters of the teacher network are updated by the exponential moving average strategy to obtain the trained target detection model.

[0013] S7. During the testing phase, acquire the image of the target domain assembly of the connector to be detected, and use the trained target detection model to detect the image of the target domain assembly, and output the small target detection results of the connector.

[0014] Secondly, this application also provides an unsupervised domain adaptive connector small target detection system, comprising:

[0015] The data acquisition and enhancement module is used to acquire source domain assembly images containing connector labels and target domain assembly images without labels, perform strong data enhancement on the source domain assembly images to obtain enhanced source domain data, and perform weak data enhancement and strong data enhancement on the target domain assembly images to obtain weakly enhanced target domain data and strongly enhanced target domain data.

[0016] The pseudo-label generation module is used to construct a domain adaptive detection framework that includes student and teacher networks. It inputs weakly enhanced target domain data into the teacher network in the domain adaptive detection framework for inference and generates pseudo-labels for connectors in the target domain assembly image.

[0017] The adaptive hybrid cropping module is used to crop and paste the source domain assembly image and the target domain assembly image based on the connector labels and connector pseudo-labels of the source domain assembly image through an adaptive hybrid cropping strategy based on the intersection-union ratio constraint, generating a cross-domain hybrid image and the corresponding hybrid label of the cross-domain hybrid image.

[0018] The feature extraction and detection prediction module is used to input the enhanced source domain data, strongly enhanced target domain data, and cross-domain hybrid image into the student network of the domain adaptive detection framework. The student network performs forward inference on the enhanced source domain data, strongly enhanced target domain data, and cross-domain hybrid image through the YOLOv5L backbone network with embedded C3CBAM feature enhancement module, to obtain source domain feature map, target domain feature map, and hybrid image feature map, as well as source domain detection prediction results corresponding to the source domain feature map, strongly enhanced target domain detection prediction results corresponding to the target domain feature map, and hybrid image detection prediction results corresponding to the hybrid image feature map.

[0019] The feature alignment module generates a multi-scale attention map based on the connector pseudo-labels and the target domain feature map. Based on the multi-scale attention map, it performs foreground region weighted aggregation on the target domain feature map to obtain the foreground feature alignment loss. It then performs adversarial learning alignment on the global features of the source domain feature map and the global features of the target domain feature map through an image-level domain discriminator to obtain the image-level adversarial alignment loss. Finally, it performs adversarial learning alignment on the instance features of the source domain feature map and the instance features of the target domain feature map through an instance-level domain discriminator to obtain the instance-level adversarial alignment loss.

[0020] The joint optimization training module is used to construct the total loss function based on the connector labels, hybrid labels, connector pseudo labels, image-level adversarial alignment loss, foreground feature alignment loss, instance-level adversarial alignment loss, source domain detection prediction results, strongly enhanced target domain detection prediction results, and hybrid image detection prediction results of the source domain assembly image. The total loss function is optimized to update the parameters of the student network, and the parameters of the teacher network are updated through an exponential moving average strategy to obtain the trained target detection model.

[0021] The test inference module is used to acquire the image of the target domain assembly of the connector during the testing phase, and to detect the target domain assembly image through the trained target detection model, and output the small target detection results of the connector.

[0022] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of any of the methods described above.

[0023] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the methods described above.

[0024] The aforementioned unsupervised domain adaptive connector small target detection method and system, by acquiring source domain assembly images containing connector labels and unlabeled target domain assembly images and performing strong and weak data augmentation respectively, can construct a multi-branch training data foundation with labeled information and controllable visual distribution differences for cross-domain connector detection tasks. By constructing a domain adaptive detection framework including student and teacher networks and inferring from weakly augmented target domain data to generate connector pseudo-labels, it can provide reliable supervision signals for unlabeled target domains, reducing reliance on manual annotation. Based on source domain connector labels and target domain pseudo-labels, an adaptive hybrid cropping strategy with intersection-union ratio constraints is used to crop and paste source and target domain images to generate cross-domain hybrid images and corresponding hybrid labels, which can promote content interaction and visual distribution fusion between the source and target domains, and ensure the integrity of connector target structures and semantic consistency of labels in hybrid images. By embedding a C3CBAM feature enhancement module into the YOLOv5L backbone network, source domain data, strongly augmented target domain data, and cross-domain hybrid images are processed. Forward inference enhances the local texture and edge feature representation of small connectors and suppresses interference from complex backgrounds. Multi-scale attention maps are generated based on connector pseudo-labels and target domain feature maps, and foreground region weighted aggregation is performed on the target domain feature maps to obtain foreground feature alignment loss. Image-level and instance-level domain discriminators are used to perform adversarial learning alignment of source and target domain feature maps to obtain image-level and instance-level adversarial alignment losses. This simultaneously reduces the feature distribution differences between the source and target domains at multiple levels: fine-grained foreground, global scene, and local instance. By constructing a total loss function containing multi-branch supervision loss and multiple alignment loss and jointly optimizing student network parameters while updating teacher network parameters using an exponential moving average strategy, collaborative training of the main detection task and domain adaptation task is achieved, resulting in a high-precision cross-domain detection model. During the testing phase, the trained target detection model performs inference on the assembly image of the target domain to be detected, outputting small target detection results for connectors, significantly improving the accuracy and stability of small target detection for assembly connectors in unlabeled target domain scenarios. This method enables the effective transfer of cross-domain connector detection knowledge without the need for manual annotation of the target domain. It effectively reduces the cost of manual annotation in assembly process monitoring and improves the resistance to interference from domain offset factors such as lighting, material, shooting angle and background distribution, supporting efficient and robust real-time detection of assembly connectors and assembly quality monitoring. Attached Figure Description

[0025] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 A flowchart illustrating an unsupervised, domain-adaptive small target detection method for connectors, provided as an exemplary embodiment of the present invention;

[0027] Figure 2 A schematic diagram of the overall framework of an unsupervised domain adaptive connector small target detection method provided as an exemplary embodiment of the present invention;

[0028] Figure 3 A schematic diagram illustrating the principle of an adaptive hybrid pruning strategy based on intersection-union ratio constraints, provided as an exemplary embodiment of the present invention;

[0029] Figure 4 A schematic diagram of the structure of a YOLOv5L backbone network with an embedded C3CBAM feature enhancement module is provided as an exemplary embodiment of the present invention;

[0030] Figure 5 A schematic diagram of a connector dataset is provided as an exemplary embodiment of the present invention;

[0031] Figure 6 A schematic diagram of a connector detection result provided as an exemplary embodiment of the present invention;

[0032] Figure 7 This is a schematic diagram of an unsupervised domain adaptive connector small target detection system provided as an exemplary embodiment of the present invention. Detailed Implementation

[0033] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0034] In one embodiment, such as Figure 1 As shown, an unsupervised, domain-adaptive, small-target detection method for connectors is provided. This embodiment illustrates the application of this method to a detection terminal. It is understood that this method can also be applied to a target detection server, and further to a system including both a detection terminal and a detection server, implemented through interaction between the two. The overall architecture of the unsupervised, domain-adaptive, small-target detection method for connectors described in this embodiment can be found in [reference needed]. Figure 2 The method includes the following steps:

[0035] S1. Obtain the source domain assembly image containing connector labels and the target domain assembly image without labels. Perform strong data augmentation on the source domain assembly image to obtain the augmented source domain data. Perform weak data augmentation and strong data augmentation on the target domain assembly image to obtain weakly augmented target domain data and strongly augmented target domain data.

[0036] Specifically, the inspection terminal can acquire source domain assembly images containing connector labels and target domain assembly images without labels through an authorized industrial data interface. Samples of source domain assembly images with connector labels and target domain assembly images without labels can be found here. Figure 5 The detection terminal can perform strong data augmentation on the source domain assembly image, expanding the sample feature distribution through geometric transformation, illumination perturbation, and random occlusion operations to obtain the augmented source domain data. The detection terminal can also perform weak and strong data augmentation on the target domain assembly image. Weakly augmented target domain data preserves the original image's semantic distribution and is used to generate reliable pseudo-label supervision signals. Strongly augmented target domain data employs the same augmentation strategy as the source domain and is used for cross-domain adaptive training.

[0037] Optionally, the source domain assembly image containing connector labels can be generated by rendering a 3D assembly digital model, and the corresponding label file records the category identifier and normalized bounding box coordinates of each connector.

[0038] Optionally, the unlabeled target domain assembly image can be a real-world image of the physical assembly captured by an industrial camera at the actual assembly station. During the training phase of the unlabeled target domain assembly image, the manually labeled information of the source domain assembly image containing connector labels is not used. Labeled data is only introduced during the verification and testing phase for performance evaluation and to reflect the lighting conditions, material surface condition, shooting angle and background distribution characteristics in the real workshop environment.

[0039] Optionally, the enhanced source domain data can be a set of images obtained by strong data augmentation of the source domain assembly image. The connector labels corresponding to the enhanced source domain data are updated in coordinate mapping synchronously with the geometric transformation of the image, which is used to provide complete supervised training samples for the student network and support the basic learning process of source domain connector detection features.

[0040] Optionally, weakly enhanced target domain data can be used to characterize the image data obtained after slight transformations such as scale normalization and pixel normalization of the target domain assembly image. The spatial position and semantic structure of the connectors in the image are consistent with the original target domain image. This data is used to input the teacher network to perform forward inference and generate stable connector pseudo-labels.

[0041] Optionally, the strongly enhanced target domain data can be used to characterize the image data obtained after the target domain assembly image has undergone strongly enhanced operations such as geometric transformation, illumination perturbation, and random occlusion consistent with the source domain. This data can then be input into the student network to participate in unsupervised training of the target domain, thereby improving its adaptability to diverse visual distributions of the target domain.

[0042] S2. Construct a domain-adaptive detection framework that includes student and teacher networks. Input the weakly enhanced target domain data into the teacher network in the domain-adaptive detection framework for inference and generate pseudo-labels for connectors in the target domain assembly image.

[0043] Specifically, the detection terminal can construct a domain-adaptive detection framework comprising a Student Network and a Teacher Network. Both networks employ a consistent detection network structure. The Teacher Network parameters are initialized to the parameters of a pre-trained model in the source domain and do not directly participate in backpropagation updates during the training phase. The detection terminal can input weakly enhanced target domain data into the Teacher Network to perform forward inference, obtaining initial prediction results including category, bounding box, and confidence level. Based on a pre-set confidence threshold, high-confidence results are filtered to generate pseudo-labels for connectors in the target domain assembly image.

[0044] Preferably, pseudo-label filtering can be performed using the following calculation formula:

[0045]

[0046] In the formula, Indicates the first The pseudo-label output for each candidate target. Indicates the first Normalized bounding box coordinates of each candidate target. The category identifier representing the candidate target. Indicates the first Confidence scores of each candidate target This indicates a pre-set confidence threshold.

[0047] Optionally, the domain-adaptive detection framework can be a dual-network architecture consisting of a student network and a teacher network with identical structures, used to achieve cross-domain transfer of connector detection knowledge on unlabeled target domain data.

[0048] Optionally, the pseudo-labels of connectors in the target domain assembly image can be used to characterize the category attributes, spatial location boundaries, and corresponding confidence information of connector targets in the target domain assembly image, serving as alternative supervision signals in the unsupervised training phase of the target domain, and constraining the target domain detection and prediction process of the student network.

[0049] S3. Based on the connector labels and pseudo-labels of the source domain assembly image, the source domain assembly image and the target domain assembly image are cropped and pasted using an adaptive hybrid cropping strategy based on the intersection-union ratio constraint to generate a cross-domain hybrid image and the corresponding hybrid label.

[0050] Specifically, such as Figure 3 The diagram illustrates the principle of an adaptive hybrid cropping strategy based on intersection-over-union (IoU) constraints. The detection terminal can acquire connector labels from the source domain assembly image and connector pseudo-labels from the target domain assembly image, performing an adaptive hybrid cropping operation based on IoU constraints. The detection terminal can locate the target region based on the source domain connector labels and adaptively expand the cropping range outwards from the target bounding box, ensuring the cropped area covers the connector body and adjacent assembly background information. The detection terminal can map the cropped source domain region to candidate pasting positions in the target domain assembly image, calculate the IoU between the candidate pasting region and the bounding boxes of each connector pseudo-label, and complete the image pasting operation when the maximum IoU satisfies a preset constraint. The source domain labels are then mapped to coordinates and combined with the target domain pseudo-labels to generate a cross-domain hybrid image and corresponding hybrid labels.

[0051] Optionally, the adaptive hybrid cropping strategy based on the intersection-union ratio constraint can be a cross-domain training sample construction method with the connector target as the core, using the real labels of the connectors in the source domain as the cropping benchmark for adaptive expansion of the region, and constraining the degree of overlap between the pasting position and the existing connectors in the target domain by the intersection-union ratio threshold, so as to maintain the integrity of the target structure and label information while fusing cross-domain visual features.

[0052] Optionally, the cross-domain hybrid image can be used to characterize the hybrid training image generated by fusing the local region of the source domain connector with the background of the target domain assembly, while carrying the visual distribution features of the source and target domains, in order to reduce the visual distribution difference between the source and target domains and reduce the difficulty of cross-domain feature transfer.

[0053] Optionally, the hybrid label corresponding to the cross-domain hybrid image can be a set of labels composed of the real labels of the source domain connectors after coordinate mapping and the pseudo labels of the original connectors in the target domain. The label coordinates correspond one-to-one with the target positions in the hybrid image, which is used to provide supervision for the supervised training of the hybrid image branches.

[0054] S4. The enhanced source domain data, strongly enhanced target domain data, and cross-domain hybrid image are input into the student network of the domain adaptive detection framework. The student network performs forward inference on the enhanced source domain data, strongly enhanced target domain data, and cross-domain hybrid image through the YOLOv5L backbone network with embedded C3CBAM feature enhancement module, to obtain source domain feature map, target domain feature map, and hybrid image feature map, as well as source domain detection prediction results corresponding to the source domain feature map, strongly enhanced target domain detection prediction results corresponding to the target domain feature map, and hybrid image detection prediction results corresponding to the hybrid image feature map.

[0055] Specifically, such as Figure 4 The diagram illustrates the structure of a YOLOv5L backbone network embedded with a C3CBAM feature enhancement module. The detection terminal can input enhanced source domain data, strongly enhanced target domain data, and cross-domain mixed images into the student network. The student network uses YOLOv5L (You Only Look Once v5 Large, version 5 large-size single-stage object detection network) as its backbone network, embedding a C3CBAM feature enhancement structure. The C3CBAM feature enhancement structure is composed of a C3 structure and a CBAM (Convolutional Block Attention Module) concatenated. Basic feature extraction is performed first, followed by feature recalibration through channel attention and spatial attention. The detection terminal can use the backbone network to perform convolutional downsampling and feature enhancement on the enhanced source domain data, strongly enhanced target domain data, and cross-domain mixed images to obtain corresponding multi-scale feature maps, and then output the mixed image detection prediction results through the detection head.

[0056] Optionally, the source domain feature map can be used to characterize the multi-scale feature tensor output by the enhanced source domain data after layer-by-layer convolution, downsampling and feature enhancement of the student network backbone network. It carries the low-level texture, mid-level structure and high-level semantic features of the source domain assembly image, providing a feature basis for subsequent detection prediction and cross-domain feature alignment.

[0057] Optionally, the target domain feature map can be used to characterize the multi-scale feature tensor obtained after the strongly enhanced target domain data is extracted and enhanced by the student network backbone network, including but not limited to the connector and background feature information in the target domain assembly scenario, to support subsequent cross-domain feature alignment and unsupervised detection training in the target domain.

[0058] Optionally, the hybrid image feature map can be used to characterize the multi-scale feature tensor output by the student network backbone network after feature extraction and enhancement of the cross-domain hybrid image. At the same time, it integrates the feature distributions of the source domain target and the target domain background to support the supervised training of the hybrid image branch and promote cross-domain feature transfer.

[0059] Optionally, the source domain detection prediction result corresponding to the source domain feature map can be a set of prediction results obtained by decoding the source domain feature map by the detection head, including but not limited to the class probability, bounding box coordinate offset and confidence score of each candidate connector, which is used to calculate supervised loss with the source domain connector label.

[0060] Optionally, the strongly enhanced target domain detection prediction result corresponding to the target domain feature map can be a set of prediction results output by the detection head decoding the target domain feature map, including but not limited to the category, position and confidence information of the candidate connectors in the target domain under the strongly enhanced condition, which are used to calculate the unsupervised loss of the target domain with the connector pseudo-labels.

[0061] Optionally, the mixed image detection prediction result corresponding to the mixed image feature map can be a set of prediction results obtained by the detection head decoding the mixed image feature map, including but not limited to the category, location and confidence prediction information of all candidate connectors in the mixed image, which is used to calculate the supervised loss of the mixed image with the mixed label.

[0062] S5. Generate a multi-scale attention map based on the connector pseudo-labels and the target domain feature map. Based on the multi-scale attention map, perform foreground region weighted aggregation on the target domain feature map to obtain the foreground feature alignment loss. Then, perform adversarial learning alignment on the global features of the source domain feature map and the global features of the target domain feature map through an image-level domain discriminator to obtain the image-level adversarial alignment loss. Finally, perform adversarial learning alignment on the instance features of the source domain feature map and the instance features of the target domain feature map through an instance-level domain discriminator to obtain the instance-level adversarial alignment loss.

[0063] Specifically, the detection terminal can map foreground region masks at various levels onto the multi-scale target domain feature map based on connector pseudo-labels. After scale normalization of the foreground region masks at different levels, they are fused to generate a multi-scale attention map. The detection terminal can then perform element-wise weighting on the target domain feature map using the multi-scale attention map to obtain a reweighted target domain feature map. The detection terminal can then perform foreground region pooling on both the reweighted source and target domain feature maps to obtain source and target domain foreground feature vectors, respectively. Finally, the foreground feature alignment loss is calculated by minimizing the cosine distance between the source and target domain foreground feature vectors.

[0064] Furthermore, the detection terminal can invoke an image-level domain discriminator, inputting the global features of the source domain feature map and the weighted target domain feature map. Through adversarial learning, the global feature distributions of the source and target domains are constrained to tend towards consistency, resulting in an image-level adversarial alignment loss. The detection terminal can also invoke an instance-level domain discriminator, extracting instance features based on the source domain connector labels and target domain pseudo-labels, and inputting these features into the instance-level discriminator. Through adversarial learning, the distribution alignment of local instance features is achieved, resulting in an instance-level adversarial alignment loss.

[0065] Optionally, the multi-scale attention map can be a foreground weight mask generated by mapping the connector pseudo-labels onto the target domain feature maps at different levels, and a multi-scale weight tensor obtained after scale alignment and fusion. Each scale weight tensor corresponds to a feature level of a different receptive field and is used to weight the target domain feature map to enhance the feature response of the connector foreground region.

[0066] Optionally, the foreground feature alignment loss can be used to characterize the degree of difference in the distribution of foreground features of the source and target domain connectors in the high-dimensional feature space. The smaller the value of the foreground feature alignment loss, the higher the consistency of the distribution of foreground features between the two domains. This is used to constrain the model to focus on the core target region to achieve fine-grained feature alignment.

[0067] Optionally, the image-level domain discriminator can be a binary classification network consisting of multiple convolutional and fully connected layers. The input is the global features of the feature map, and the output is the classification probability of the domain to which the feature belongs. This is used to reduce the difference in global feature distribution between the source domain and the target domain through adversarial training.

[0068] Optionally, the instance-level domain discriminator can be a binary classification network that performs domain classification based on local instance features. The input is the single-connector instance features obtained by cropping based on the label, and the output is the classification probability of the domain to which the instance belongs, which is used to achieve fine-grained feature distribution alignment at the connector instance level.

[0069] Optionally, the image-level adversarial alignment loss can be the loss value generated during the adversarial training of the image-level domain discriminator and the backbone feature extraction network, used to measure the degree of domain alignment of global features and guide the backbone network to generate global semantic features with domain invariance.

[0070] Optionally, the instance-level adversarial alignment loss can be the loss value generated during the adversarial training of the instance-level domain discriminator and the backbone feature extraction network, used to measure the difference in domain distribution of instance features and guide the target detection model to extract connector instance features with domain invariance.

[0071] S6. Based on the connector labels, hybrid labels, connector pseudo labels, image-level adversarial alignment loss, foreground feature alignment loss, instance-level adversarial alignment loss, source domain detection prediction results, strongly enhanced target domain detection prediction results, and hybrid image detection prediction results of the source domain assembly image, a total loss function is constructed. The parameters of the student network are updated by optimizing the total loss function, and the parameters of the teacher network are updated by an exponential moving average strategy to obtain the trained target detection model.

[0072] Specifically, the detection terminal can calculate the supervised detection loss of the source domain based on the connector labels of the source domain assembly image and the source domain detection prediction results. The detection terminal can calculate the unsupervised detection loss of the target domain based on the connector pseudo-labels and the strongly enhanced target domain detection prediction results. The detection terminal can calculate the mixed image detection loss based on the mixed label and mixed image detection prediction results.

[0073] Furthermore, the detection terminal can combine image-level adversarial alignment loss, foreground feature alignment loss, and instance-level adversarial alignment loss. After assigning preset balanced weights to these losses, the terminals are weighted and summed to construct the total loss function. The detection terminal can optimize the total loss function using gradient descent, backpropagate to update the trainable parameters of the student network, and employ exponential moving average (EMA) parameter update strategy to weighted smooth the current parameters of the student network and the historical parameters of the teacher network. This iterative update of the teacher network parameters continues until training converges, resulting in a trained target detection model.

[0074] Optionally, the total loss function can be used to characterize the comprehensive optimization objective of the object detection model, integrating multi-branch detection supervision loss and multi-level domain adaptive alignment loss to guide the iterative update of student network parameters.

[0075] S7. During the testing phase, acquire the image of the target domain assembly of the connector to be detected, and use the trained target detection model to detect the image of the target domain assembly, and output the small target detection results of the connector.

[0076] Specifically, during the testing phase, the detection terminal can acquire images of the connector assembly to be detected through an industrial image acquisition link. The detection terminal can first perform pixel normalization and scale normalization preprocessing on the connector assembly image, consistent with the training phase, to ensure that the input data distribution remains consistent with the training data. The detection terminal can then input the preprocessed connector assembly image into the trained target detection model. After feature extraction and detection head decoding, an initial set of candidate detection boxes is obtained. After eliminating overlapping and redundant boxes using non-maximum suppression (NMS), valid targets are selected according to a preset test confidence threshold, and the small target detection results for the connector are output. The format of the small target detection results for the connector can be found in [reference needed]. Figure 6 .

[0077] Optionally, the image of the target domain assembly to be detected for the connector can be a real-time image of the target domain assembly captured by an industrial camera on an actual assembly line, which has not participated in model training, and can be used to verify the actual detection performance and domain generalization ability of the target detection model.

[0078] Optionally, the small target detection results of connectors can be used to characterize the category attributes, spatial boundary locations, and confidence information of all connectors in the assembly image of the target domain to be detected.

[0079] In the aforementioned unsupervised domain adaptive connector small target detection method, the detection terminal acquires source domain assembly images containing connector labels and unlabeled target domain assembly images, and performs strong and weak data augmentation on them respectively. This enables the construction of a multi-branch training data foundation with labeled information and controllable visual distribution differences for cross-domain connector detection tasks. By constructing a domain adaptive detection framework containing student and teacher networks and inferring from weakly augmented target domain data to generate connector pseudo-labels, reliable supervision signals can be provided for unlabeled target domains to reduce reliance on manual annotation. Based on source domain connector labels and target domain pseudo-labels, an adaptive hybrid cropping strategy with intersection-union ratio constraints is used to crop and paste source and target domain images to generate cross-domain hybrid images and corresponding hybrid labels. This promotes content interaction and visual distribution fusion between the source and target domains, and ensures the integrity of connector target structures and semantic consistency of labels in the hybrid images. Finally, a YOLOv5L backbone network with embedded C3CBAM feature enhancement modules is used to process source domain data, strongly augmented target domain data, and cross-domain hybrid images. Forward inference enhances the local texture and edge feature representation of small connectors and suppresses interference from complex backgrounds. Multi-scale attention maps are generated based on connector pseudo-labels and target domain feature maps, and foreground region weighted aggregation is performed on the target domain feature maps to obtain foreground feature alignment loss. Image-level and instance-level domain discriminators are used to perform adversarial learning alignment of source and target domain feature maps to obtain image-level and instance-level adversarial alignment losses. This simultaneously reduces the feature distribution differences between the source and target domains at multiple levels: fine-grained foreground, global scene, and local instance. By constructing a total loss function containing multi-branch supervision loss and multiple alignment loss and jointly optimizing student network parameters while updating teacher network parameters using an exponential moving average strategy, collaborative training of the main detection task and domain adaptation task is achieved, resulting in a high-precision cross-domain detection model. During the testing phase, the trained target detection model performs inference on the assembly image of the target domain to be detected, outputting small target detection results for connectors, significantly improving the accuracy and stability of small target detection for assembly connectors in unlabeled target domain scenarios. This method enables the effective transfer of cross-domain connector detection knowledge without the need for manual annotation of the target domain. It effectively reduces the cost of manual annotation in assembly process monitoring and improves the resistance to interference from domain offset factors such as lighting, material, shooting angle and background distribution, supporting efficient and robust real-time detection of assembly connectors and assembly quality monitoring.

[0080] In one embodiment, generating a multi-scale attention map based on the connector pseudo-labels and the target domain feature map in step S5 may include:

[0081] S51. Based on the feature maps of each level in the target domain feature map, obtain the target box coordinates and size information of the connector pseudo-label, and perform coordinate mapping and Gaussian response calculation on the target box coordinates and size information to obtain the Gaussian space prior map corresponding to each level feature map.

[0082] For example, the detection terminal can proportionally map the coordinates and size information of the connector pseudo-label target box in the original image coordinate system to the feature map coordinate system of the corresponding level, based on the downsampling ratio corresponding to each level of feature map, to obtain the target bounding box and center coordinates in the feature map dimension. The detection terminal uses the center of the mapped target box as the center of the Gaussian kernel, and uses the corresponding ratio of the target box width and height as the bandwidth parameter of the Gaussian kernel. It calculates the Gaussian response value pixel by pixel in the entire feature map space. The Gaussian response value decays smoothly from the center to the surrounding area in a Gaussian distribution, generating a Gaussian space prior map that perfectly matches the spatial size of the hierarchical feature map.

[0083] S52. Perform convolution processing and activation mapping on the feature maps of each level to generate the feature response maps corresponding to each level.

[0084] For example, the detection terminal can perform channel-dimensional compression and weighted fusion of multi-channel features through convolution operations with a kernel size of 1×1 on the target domain feature map of each level, and integrate the semantic information of different channels to obtain a single-channel feature activation tensor. The detection terminal can use a sigmoid function to nonlinearly map all activation values ​​to the [0,1] interval to generate a feature response map with the same spatial size as the feature map of the current level.

[0085] Optionally, the feature response maps corresponding to each level of feature maps can be used to characterize the degree of association between the features at each spatial location in the feature maps of different levels of the target domain and the semantics of the connector target, so as to provide a data-driven basis for semantic weight calibration for spatial priors.

[0086] Preferably, the expression for the feature response map can be:

[0087] ;

[0088] In the formula, Indicates the first The feature response map corresponding to the hierarchical feature map. The YOLOv5L backbone network representing the student network is... The target domain feature map output by the layer. This indicates that the input is mapped to... interval Type activation function, Indicates the kernel size as The convolution operation.

[0089] S53. Merge and normalize each Gaussian space prior map and the corresponding feature response map element by element to obtain the initial attention map of each level of feature map.

[0090] For example, the detection terminal can perform element-wise weighted fusion of the Gaussian space prior map and the feature response map at the same level, and control the contribution ratio of the semantic response in the fusion weight by a preset feature response adjustment coefficient to obtain the fused weight map. The detection terminal can perform maximum value normalization on the fused weight map and linearly map all weight values ​​to the [0,1] interval to obtain the initial attention map of the corresponding level.

[0091] Preferably, the expression for the initial attention map can be:

[0092] ;

[0093] In the formula, Indicates the first The initial attention map corresponding to the hierarchical feature map. This indicates a normalization operation. Indicates the first Gaussian space prior map of hierarchical feature maps This represents element-wise multiplication. This represents the preset characteristic response adjustment coefficient. Indicates the first The feature response map corresponding to the hierarchical feature map.

[0094] S54. Adjust each initial attention map to a uniform spatial scale and merge and normalize them to obtain a multi-scale attention map.

[0095] For example, the detection terminal can perform bilinear interpolation upsampling or downsampling operations on initial attention maps at different levels to adjust all initial attention maps to a uniform spatial resolution. The detection terminal can also perform element-wise weighted summation and fusion of multiple attention maps of the same scale, and then perform maximum value normalization to constrain the global weight range to the [0,1] interval, thus obtaining a multi-scale attention map covering multi-scale receptive field information.

[0096] In this embodiment, the detection terminal can generate a Gaussian spatial prior map by performing coordinate mapping and Gaussian response calculation on the target domain feature maps at each level based on the connector pseudo-label. Then, it combines the feature response maps obtained by convolution and activation mapping of each level feature map with the feature response maps and performs element-wise fusion and normalization. Finally, it adjusts the initial attention maps of each level to a unified spatial scale and fuses and normalizes them to obtain a multi-scale attention map. This allows the attention map to be simultaneously constrained by the pseudo-label position prior and guided by the feature response distribution, achieving accurate representation of the prior and feature response intensity of targets at different spatial locations. This provides accurate multi-scale attention weights for subsequent foreground region weighted aggregation, effectively improving the attention capability for small targets in the foreground of the connector and the cross-domain feature alignment accuracy.

[0097] In one embodiment, the Gaussian response value of the Gaussian space prior map can be calculated using the following formula:

[0098] ;

[0099] ;

[0100] ;

[0101] ;

[0102] ;

[0103] In the formula, Indicates the first Hierarchical feature maps in spatial location Gaussian response value at that point, Indicates the first Hierarchical feature map Indicates the first The x-coordinate of any spatial location on the hierarchical feature map. Indicates the first The ordinate of any spatial location on the hierarchical feature map. Indicates the first pseudo-label of the connector The target center of the connector is mapped to the first... The horizontal axis on the hierarchical feature map Indicates the first pseudo-label of the connector The target center of the connector is mapped to the first... The ordinate on the hierarchical feature map, Indicates the first pseudo-label of the connector The normalized x-coordinate of the center of each connector target. Indicates the first Spatial width of the hierarchical feature map Indicates the first pseudo-label of the connector The normalized center ordinate of each connector target, Indicates the first Spatial height of hierarchical feature maps Indicates the first The connector target is in the first The horizontal Gaussian response range of the hierarchical feature map. Indicates the first pseudo-label of the connector The normalized width of each connector target. This represents the preset scaling factor. Indicates the first The connector target is in the first The range of Gaussian response in the vertical direction of the hierarchical feature map. Indicates the first pseudo-label of the connector The normalized height of each connector target. This function represents the maximum value among the values ​​within the parentheses. Represented by natural constant An exponential function with base 0.

[0104] For example, when calculating the Gaussian response value of the Gaussian space prior map, the detection terminal can multiply the normalized center coordinates of each target in the connector pseudo-label with the spatial size of the corresponding level feature map, completing the mapping from the normalized coordinate system of the original image to the pixel coordinate system of the feature map, thus obtaining the target center coordinates in the feature map dimension. The detection terminal can use the normalized target width and height as a basis, combined with the feature map size and a preset scale adjustment coefficient, to calculate the Gaussian response bandwidth in the horizontal and vertical directions, and constrain the lower limit of the bandwidth through maximum value calculation to avoid excessive shrinkage of the response region of small targets. The detection terminal can traverse the feature map space pixel by pixel, calculate the Gaussian response value of each target at the current position, and when multiple targets exist at the same position, take the maximum response value as the final output to generate the corresponding level Gaussian space prior map.

[0105] In this embodiment, the detection terminal can complete coordinate mapping by multiplying the normalized target center of the connector pseudo-label with the spatial size of the hierarchical feature map. It then calculates the Gaussian response bandwidth in the horizontal and vertical directions by combining the normalized target width and height with the feature map size and a preset scale adjustment coefficient. At the same time, it constrains the lower limit of the bandwidth through a maximum value function to avoid excessive shrinkage of the response region of small targets. Finally, it traverses the feature map space pixel by pixel to calculate the Gaussian response value of each target and takes the maximum value when multiple targets exist at the same position. This can transform the discrete pseudo-label position information into a continuous and smooth Gaussian distribution prior weight, so that the attention response region is adaptively matched with the actual scale of the target. This effectively avoids the attention failure of small targets on the high-level feature map due to the mapping size being too small, and provides accurate spatial position prior guidance for subsequent cross-domain feature alignment.

[0106] In one embodiment, the expression for the value of the total loss function can be:

[0107] ;

[0108] ;

[0109] ;

[0110] ;

[0111] In the formula, This represents the value of the total loss function. This indicates that the mixed images have supervised loss. This indicates that there is a monitoring loss in the source domain. This indicates unsupervised loss in the target domain. This represents the image-level adversarial alignment loss. This represents instance-level adversarial alignment loss. This represents the foreground feature alignment loss. Represents global features of the source domain. Represents global features of the target domain. This represents the discriminative mapping performed by the image-level domain discriminator on the input features. This represents the binary cross-entropy loss function. This represents the source domain foreground feature vector extracted from the source domain feature map. This represents the target domain foreground feature vector extracted from the target domain feature map. This represents the cosine similarity function.

[0112] For example, the detection terminal can calculate the detection branch loss and the domain alignment branch loss separately, and aggregate them according to preset weights to obtain the total loss function value. The detection terminal can calculate the mixed image supervised loss based on the mixed label and mixed image detection prediction results, and calculate the source domain supervised loss based on the connector labels of the source domain assembly image and the source domain detection prediction results.

[0113] Furthermore, the detection terminal can calculate the unsupervised loss of the target domain based on the pseudo-labels of connectors and the detection prediction results of the strongly enhanced target domain. The detection terminal can input the global features of the source domain and the global features of the target domain into the image-level domain discriminator, respectively, and calculate the image-level adversarial alignment loss using binary cross-entropy. Similarly, it can extract the instance features of the source and target domains and input them into the instance-level domain discriminator to calculate the instance-level adversarial alignment loss. The detection terminal can calculate the foreground feature alignment loss based on the cosine similarity of the foreground feature vectors of the source and target domains. The detection terminal can scale each alignment loss according to preset weight coefficients and then add it to the image-level adversarial alignment loss, the instance-level adversarial alignment loss, and the foreground feature alignment loss to obtain the total loss function value, which is used to drive the iterative update of network parameters.

[0114] In this embodiment, the detection terminal can calculate the mixed image supervised loss, source domain supervised loss, target domain unsupervised loss, image-level adversarial alignment loss, instance-level adversarial alignment loss, and foreground feature alignment loss respectively. After scaling each alignment loss according to preset weights, the total loss function value is obtained by weighted summation with the detection branch loss to drive the iterative update of network parameters. This can enable multi-branch detection supervision signals and multi-level domain alignment constraints to be trained collaboratively under the same optimization objective, so that the main detection task and the domain adaptation task promote each other rather than conflict. This effectively reduces the feature distribution differences between the source domain and the target domain at the three levels of global scene, local instance, and foreground fine granularity while ensuring the source domain detection capability, thereby improving the convergence stability of the cross-domain detection model and the target domain detection accuracy.

[0115] In one embodiment, S2 may include:

[0116] S21. Construct student and teacher networks.

[0117] Optionally, the student network and the teacher network adopt the same network structure, and the parameters of the teacher network can be the parameters of the source domain pre-trained model that have already been initialized.

[0118] For example, the detection terminal can build student and teacher networks with completely identical structures. Both student and teacher networks adopt the YOLOv5L architecture with embedded C3CBAM feature enhancement structure, including three components: backbone feature extraction, multi-scale feature fusion, and detection head output. The detection terminal can load model parameters pre-trained from source domain labeled data and assign these pre-trained model parameters to the teacher network as initial parameters. The student network can be initialized using the same pre-trained parameters. During the training phase, the teacher network does not directly perform backpropagation updates but only synchronizes the iterative results of the student network through a parameter smoothing strategy, ensuring the stability of pseudo-label generation.

[0119] S22. Input the weakly enhanced target domain data into the teacher network, and use the teacher network to perform forward inference on the weakly enhanced target domain data to obtain the initial prediction results corresponding to the weakly enhanced target domain data.

[0120] For example, the detection terminal can input weakly enhanced target domain data into the initialized teacher network. The data is then processed sequentially through convolutional downsampling and feature enhancement of the backbone network, multi-scale feature aggregation of the feature fusion layer, and the detection head decodes and outputs the initial prediction results of all candidate targets. The entire inference process does not trigger gradient calculation and parameter updates.

[0121] Optionally, the initial prediction results corresponding to the weakly enhanced target domain data can be used to characterize the teacher network's preliminary judgment results on connector targets in the weakly enhanced target domain image, including three types of information: the category attributes, spatial location, and confidence level of the candidate targets. This serves as the intermediate basic data for screening and generating reliable connector pseudo-labels.

[0122] S23. Compare the initial prediction result with the preset confidence threshold to obtain the first comparison result. When the confidence level of the first comparison result is greater than the preset confidence threshold, use the initial prediction result as the pseudo-label of the connector in the target domain assembly image.

[0123] For example, the detection terminal can traverse all candidate targets in the initial prediction results, extract the confidence score of each candidate target one by one, and compare the confidence score of each candidate target with the preset confidence threshold to obtain the first comparison result corresponding to each candidate target. When the first comparison result shows that the confidence score of the candidate target is greater than the preset confidence threshold, the candidate target is determined as a valid prediction and retained; when the first comparison result shows that the confidence score of the candidate target is less than or equal to the preset confidence threshold, the candidate target is discarded as a low-confidence, unstable prediction result.

[0124] Furthermore, the detection terminal can perform non-maximum suppression on all retained valid predictions to eliminate overlapping and redundant detection boxes, and use the detection results that are finally retained after confidence screening and non-maximum suppression as the connector pseudo-labels of the target domain assembly image.

[0125] Optionally, the first comparison result can be used to characterize the numerical comparison result between the confidence score of each candidate target in the initial prediction result and the preset confidence threshold.

[0126] In this embodiment, the detection terminal can construct a student network and a teacher network with completely identical structures, both using embedded C3CBAM feature enhancement structures. The teacher network parameters are initialized to the source domain pre-trained model parameters, enabling the teacher network to have reliable connector detection capabilities in the early stages of training. The teacher network then performs forward inference on the weakly enhanced target domain data to obtain initial prediction results. Based on a pre-set confidence threshold, high-confidence candidate targets are selected and retained as connector pseudo-labels for the target domain assembly image. At the same time, the teacher network does not participate in backpropagation but only synchronizes the iterative results of the student network through an exponential moving average strategy. This effectively suppresses noise fluctuations and error accumulation during the pseudo-label generation process, providing high-quality and stable target domain supervision signals for subsequent cross-domain adaptive training.

[0127] In one embodiment, S7 may include:

[0128] S71. During the testing phase, acquire the image of the target domain assembly to be tested of the connector, scale and normalize the image of the target domain assembly to be tested, and obtain the preprocessed image of the target domain assembly to be tested.

[0129] For example, during the testing phase, the detection terminal can acquire an image of the target domain assembly of the connector through an industrial image acquisition link, and scale the image proportionally according to the input size requirements of the target detection model. Edge areas that are insufficient in size after scaling are filled with fixed grayscale values ​​to avoid stretching, deformation, or distortion of the connector's geometric features. The detection terminal can then perform pixel normalization on the scaled image of the target domain assembly, mapping pixel values ​​to a standard distribution range to ensure that the distribution characteristics of the input data remain consistent with those during the training phase, thus obtaining a preprocessed image of the target domain assembly.

[0130] S72. Input the preprocessed image of the target domain assembly to be detected into the trained target detection model for forward inference to obtain the category label, confidence score and bounding box coordinates of each candidate connector in the image of the target domain assembly to be detected. Combine the category label, confidence score and bounding box coordinates of each candidate connector to obtain the initial detection prediction result.

[0131] For example, the detection terminal can input the preprocessed image of the target domain assembly to be detected into the trained target detection model. The preprocessed image of the target domain assembly to be detected is then passed through the convolutional downsampling and feature enhancement of the backbone network, and cross-level feature aggregation of the multi-scale feature fusion layer. Finally, the detection head decodes and outputs the class probability distribution, bounding box coordinate offset, and target confidence score of each candidate target. The detection terminal can decode and restore the bounding box offset to obtain the actual bounding box coordinates of the candidate connector in the image, and select the class with the highest class probability as the class label of the corresponding candidate connector. The class label, confidence score, and bounding box coordinates of each candidate connector are combined one by one to obtain the initial detection prediction result.

[0132] S73. Compare the preset confidence threshold with the initial detection prediction result to obtain the second comparison result. Filter out candidate connectors whose confidence scores are lower than the preset confidence threshold in the second comparison result. Perform non-maximum suppression on the retained candidate connectors to remove redundant detection boxes and obtain the retained detection boxes. Use the retained detection boxes as the small target detection results of the connectors.

[0133] For example, the detection terminal can compare the preset confidence threshold with the initial detection prediction results one by one, filter out low-quality candidate connectors with confidence scores below the threshold, and group the remaining candidate connectors by category, sorting them from highest to lowest confidence score within each group. The detection terminal can then use a non-maximum suppression strategy to traverse the sorted candidate boxes one by one and calculate the intersection-union ratio (IUU) between the current highest confidence box and the remaining candidate boxes.

[0134] Furthermore, the detection terminal can eliminate redundant detection boxes with an intersection-union ratio higher than a preset overlap threshold, and use the retained detection boxes and their corresponding categories and confidence levels as the detection results for the small targets of the connectors.

[0135] Optionally, the second comparison result can be used to characterize the numerical comparison result between the preset confidence threshold and the confidence score of each candidate connector in the initial detection prediction result.

[0136] Optionally, the retained detection boxes can be used to characterize candidate connector detection boxes retained after confidence thresholding and non-maximum suppression. Each retained detection box may include the spatial boundary position, category attribute, and confidence score corresponding to the connector target in the image.

[0137] In this embodiment, the detection terminal can perform forward inference on the trained target detection model after scaling, filling, and pixel normalization of the assembly image of the target domain to be detected during the testing phase. This process yields the category labels, confidence scores, and bounding box coordinates of each candidate connector, which are then combined to form the initial detection prediction result. After filtering out low-confidence candidate targets by setting a pre-set confidence threshold, the remaining candidate connectors are grouped by category and non-maximum suppression is applied to eliminate redundant detection boxes with an intersection-union ratio higher than a preset overlap threshold. This ensures that the distribution of the test input data remains consistent with that of the training phase. Furthermore, the dual mechanisms of confidence screening and redundant box suppression effectively filter out false detections and duplicate detections, outputting accurate and reliable small target detection results for connectors. This meets the actual requirements of real-time monitoring scenarios in the assembly process for detection accuracy and result clarity.

[0138] In a verification embodiment, under the same engine assembly connector dataset and the same evaluation metrics, the unsupervised domain adaptive small target detection method for connectors provided by this invention is compared with existing supervised target detection methods (Mamba You Only Look Once, Mamba YOLO, a single-stage target detection algorithm incorporating the Mamba state space model) and unsupervised domain adaptive target detection methods (Probabilistic Teacher for Domain Adaptive Object Detection, PTDA, Learning Domain Adaptive Object Detection with Probabilistic Teacher; ADDA (Adversarial Discriminative Domain Adaptation, Align and Distill: Unifying and Improving Domain Adaptive Object Detection); and ConfMix (Confidence-based Mixing for Unsupervised Domain Adaptive Object Detection, ConfMix). For Object Detection via Confidence-based Mixing (Dual-Path, Adaptive Cutmix Domain Adaptive Fastener Detection, DPAC: a dual-path domain adaptive fastenersdetection method based on adaptive cutmix strategy), a comparison was made. Evaluation metrics included mAP@0.5, mAP@0.5-0.95, and inference time. The test results are shown in Table 1.

[0139] Table 1 Comparison of Target Domain Connector Detection Accuracy

[0140]

[0141] As shown in Table 1, in the task of detecting small targets in the target domain connectors, the detection accuracy of the model in the target domain is low when using only supervised detection methods. The mAP@50 and mAP@50-95 of Mamba YOLO are 37.9% and 21.8%, respectively, indicating a significant domain difference between the synthetic domain and the real physical domain. Compared with unsupervised domain adaptive methods such as PTDA, ADDA, Confmix, and DPAC, the method of this invention achieves higher detection accuracy, with mAP@50 and mAP@50-95 reaching 82.9% and 50.5%, respectively, both higher than other comparative methods. At the same time, the inference time of the method of this invention is 12.6 ms, which remains at a low level, balancing detection accuracy and inference efficiency. The above results show that this invention can effectively improve the accuracy and stability of small target detection of connectors in assembly scenarios without manual annotation in the target domain training.

[0142] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0143] Based on the same inventive concept, this application also provides a system for implementing the unsupervised adaptive connector small target detection method described above. The solution provided by this system is similar to the implementation described in the above method; therefore, the specific limitations in one or more embodiments of the unsupervised adaptive connector small target detection system provided below can be found in the limitations of the unsupervised adaptive connector small target detection method described above, and will not be repeated here.

[0144] In one exemplary embodiment, such as Figure 7 As shown, an unsupervised domain adaptive connector small target detection system 10 is provided, comprising:

[0145] The data acquisition and enhancement module 11 can be used to acquire source domain assembly images containing connector labels and target domain assembly images without labels, perform strong data enhancement on the source domain assembly images to obtain enhanced source domain data, and perform weak data enhancement and strong data enhancement on the target domain assembly images to obtain weakly enhanced target domain data and strongly enhanced target domain data.

[0146] The pseudo-label generation module 12 can be used to construct a domain adaptive detection framework that includes student networks and teacher networks. It inputs weakly enhanced target domain data into the teacher network in the domain adaptive detection framework for inference and generates pseudo-labels for connectors in the target domain assembly image.

[0147] The adaptive hybrid cropping module 13 can be used to crop and paste the source domain assembly image and the target domain assembly image based on the connector label and connector pseudo label of the source domain assembly image through an adaptive hybrid cropping strategy based on the intersection-union ratio constraint, generating a cross-domain hybrid image and the corresponding hybrid label of the cross-domain hybrid image.

[0148] The feature extraction and detection prediction module 14 can be used to input the enhanced source domain data, strongly enhanced target domain data and cross-domain hybrid image into the student network of the domain adaptive detection framework. The student network performs forward inference on the enhanced source domain data, strongly enhanced target domain data and cross-domain hybrid image through the YOLOv5L backbone network with embedded C3CBAM feature enhancement module to obtain source domain feature map, target domain feature map and hybrid image feature map, as well as source domain detection prediction results corresponding to the source domain feature map, strongly enhanced target domain detection prediction results corresponding to the target domain feature map and hybrid image detection prediction results corresponding to the hybrid image feature map.

[0149] The feature alignment module 15 can be used to generate a multi-scale attention map based on the connector pseudo-label and the target domain feature map, perform foreground region weighted aggregation on the target domain feature map based on the multi-scale attention map to obtain the foreground feature alignment loss, and perform adversarial learning alignment on the global features of the source domain feature map and the global features of the target domain feature map through an image-level domain discriminator to obtain the image-level adversarial alignment loss, and perform adversarial learning alignment on the instance features of the source domain feature map and the instance features of the target domain feature map through an instance-level domain discriminator to obtain the instance-level adversarial alignment loss.

[0150] The joint optimization training module 16 can be used to construct a total loss function based on the connector labels, hybrid labels, connector pseudo labels, image-level adversarial alignment loss, foreground feature alignment loss, instance-level adversarial alignment loss, source domain detection prediction results, strongly enhanced target domain detection prediction results, and hybrid image detection prediction results of the source domain assembly image. The total loss function is optimized to update the parameters of the student network, and the parameters of the teacher network are updated through an exponential moving average strategy to obtain the trained target detection model.

[0151] The test inference module 17 can be used to acquire the image of the target domain assembly of the connector during the testing phase, detect the target domain assembly image through the trained target detection model, and output the small target detection result of the connector.

[0152] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of the unsupervised domain adaptive connector small target detection method as described above.

[0153] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps described in the embodiment of the unsupervised domain adaptive connector small target detection method.

[0154] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0155] The above-described embodiments are merely illustrative of several implementation methods of the embodiments of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the embodiments of this application, and these modifications and improvements all fall within the protection scope of the embodiments of this application.

Claims

1. An unsupervised, domain-adaptive small target detection method for connectors, characterized in that, The method includes: S1. Obtain a source domain assembly image containing connector labels and a target domain assembly image without labels. Perform strong data augmentation on the source domain assembly image to obtain enhanced source domain data. Perform weak data augmentation and strong data augmentation on the target domain assembly image to obtain weakly enhanced target domain data and strongly enhanced target domain data. S2. Construct a domain adaptive detection framework that includes a student network and a teacher network. Input the weakly enhanced target domain data into the teacher network in the domain adaptive detection framework for inference and generate connector pseudo-labels for the target domain assembly image. S3. Based on the connector label and connector pseudo label of the source domain assembly image, the source domain assembly image and the target domain assembly image are cropped and pasted using an adaptive hybrid cropping strategy based on cross-union ratio constraints to generate a cross-domain hybrid image and a hybrid label corresponding to the cross-domain hybrid image. S4. The enhanced source domain data, the strongly enhanced target domain data, and the cross-domain hybrid image are input into the student network in the domain adaptive detection framework. The student network performs forward inference on the enhanced source domain data, the strongly enhanced target domain data, and the cross-domain hybrid image through the YOLOv5L backbone network with embedded C3CBAM feature enhancement module to obtain source domain feature map, target domain feature map, and hybrid image feature map, as well as source domain detection prediction results corresponding to the source domain feature map, strongly enhanced target domain detection prediction results corresponding to the target domain feature map, and hybrid image detection prediction results corresponding to the hybrid image feature map. S5. Based on the connector pseudo-label and the target domain feature map, generate a multi-scale attention map. Based on the multi-scale attention map, perform foreground region weighted aggregation on the target domain feature map to obtain a foreground feature alignment loss. Then, perform adversarial learning alignment on the global features of the source domain feature map and the global features of the target domain feature map through an image-level domain discriminator to obtain an image-level adversarial alignment loss. Finally, perform adversarial learning alignment on the instance features of the source domain feature map and the instance features of the target domain feature map through an instance-level domain discriminator to obtain an instance-level adversarial alignment loss. S6. Based on the connector labels of the source domain assembly image, the hybrid label, the connector pseudo label, the image-level adversarial alignment loss, the foreground feature alignment loss, the instance-level adversarial alignment loss, the source domain detection prediction result, the strongly enhanced target domain detection prediction result, and the hybrid image detection prediction result, a total loss function is constructed. The parameters of the student network are updated by optimizing the total loss function, and the parameters of the teacher network are updated by an exponential moving average strategy to obtain the trained target detection model. S7. During the testing phase, acquire the image of the target domain assembly of the connector to be detected, and use the trained target detection model to detect the image of the target domain assembly, and output the small target detection result of the connector.

2. The method according to claim 1, characterized in that, The step S5, generating a multi-scale attention map based on the connector pseudo-label and the target domain feature map, includes: S51. Based on the feature maps of each level in the target domain feature map, obtain the target box coordinates and size information of the connector pseudo-label, and perform coordinate mapping and Gaussian response calculation on the target box coordinates and size information to obtain the Gaussian space prior map corresponding to each level feature map. S52. Perform convolution processing and activation mapping on each of the layer feature maps to generate feature response maps corresponding to each of the layer feature maps; The expression for the feature response map is: ; In the formula, Indicates the first The feature response map corresponding to the hierarchical feature map, The YOLOv5L backbone network of the student network represents the first... The target domain feature map output by the layer, This indicates mapping the input to... interval Type activation function, Indicates the kernel size as Convolution operations; S53. The Gaussian space prior maps and the corresponding feature response maps are fused and normalized element by element to obtain the initial attention maps of the hierarchical feature maps. The expression for the initial attention map is: ; In the formula, Indicates the first The initial attention map corresponding to the hierarchical feature map, This indicates a normalization operation. Indicates the first The Gaussian space prior map of the hierarchical feature map. This represents element-wise multiplication. This represents the preset characteristic response adjustment coefficient. Indicates the first The feature response map corresponding to the hierarchical feature map; S54. Adjust each of the initial attention maps to a uniform spatial scale and fuse and normalize them to obtain the multi-scale attention map.

3. The method according to claim 2, characterized in that, The Gaussian response value of the Gaussian space prior map is calculated using the following formula: ; ; ; ; ; In the formula, Indicates the first Hierarchical feature maps in spatial location The Gaussian response value at that location, Indicates the first Hierarchical feature map Indicates the first The x-coordinate of any spatial location on the hierarchical feature map. Indicates the first The ordinate of any spatial location on the hierarchical feature map. This indicates the first pseudo-label of the connector. The target center of the connector is mapped to the first... The horizontal axis on the hierarchical feature map The first one in the pseudo-label of the connector The target center of the connector is mapped to the first... The ordinate on the hierarchical feature map, The first one in the pseudo-label of the connector The normalized x-coordinate of the center of each connector target. Indicates the first Spatial width of the hierarchical feature map This indicates the first pseudo-label of the connector. The normalized center ordinate of each connector target. Indicates the first Spatial height of hierarchical feature maps Indicates the first The connector target is in the first The horizontal Gaussian response range of the hierarchical feature map. The first one in the pseudo-label of the connector The normalized width of each connector target. This represents the preset scaling factor. Indicates the first The connector target is in the first The range of Gaussian response in the vertical direction of the hierarchical feature map. The first one in the pseudo-label of the connector The normalized height of each connector target. This function represents the maximum value among the values ​​within the parentheses. Represented by natural constant An exponential function with base 0.

4. The method according to claim 1, characterized in that, The expression for the total loss function is: ; ; ; ; In the formula, The value of the total loss function represents the total loss function. This indicates that the mixed images have supervised loss. This indicates that there is a monitoring loss in the source domain. This indicates unsupervised loss in the target domain. This represents the image-level adversarial alignment loss. This represents the instance-level adversarial alignment loss. This represents the foreground feature alignment loss. Represents global features of the source domain. Represents global features of the target domain. This represents the discriminative mapping performed by the image-level domain discriminator on the input features. This represents the binary cross-entropy loss function. This represents the source domain foreground feature vector extracted from the source domain feature map. This represents the target domain foreground feature vector extracted from the target domain feature map. This represents the cosine similarity function.

5. The method according to claim 1, characterized in that, S2 includes: S21. Construct the student network and the teacher network; wherein the student network and the teacher network adopt the same network structure; S22. Input the weakly enhanced target domain data into the teacher network, and perform forward inference on the weakly enhanced target domain data through the teacher network to obtain the initial prediction result corresponding to the weakly enhanced target domain data. S23. The initial prediction result is compared with a preset confidence threshold to obtain a first comparison result. When the confidence level of the first comparison result is greater than the preset confidence threshold, the initial prediction result is used as the pseudo-label of the connector in the target domain assembly image.

6. The method according to claim 5, characterized in that, The S7 includes: S71. During the testing phase, the image of the target domain assembly to be detected of the connector is acquired, and the image of the target domain assembly to be detected is scaled and normalized to obtain a preprocessed image of the target domain assembly to be detected. S72. Input the preprocessed image of the target domain assembly to be detected into the trained target detection model for forward inference to obtain the category label, confidence score and bounding box coordinates of each candidate connector in the image of the target domain assembly to be detected, and combine the category label, confidence score and bounding box coordinates of each candidate connector to obtain the initial detection prediction result. S73. The preset confidence threshold is compared with the initial detection prediction result to obtain a second comparison result. Candidate connectors whose confidence scores are lower than the preset confidence threshold are filtered out in the second comparison result. Non-maximum suppression is performed on the retained candidate connectors to remove redundant detection boxes and obtain retained detection boxes. The retained detection boxes are used as the small target detection results of the connectors.

7. An unsupervised domain adaptive connector small target detection system, characterized in that, The system includes: The data acquisition and enhancement module is used to acquire a source domain assembly image containing connector labels and a target domain assembly image without labels, perform strong data enhancement on the source domain assembly image to obtain enhanced source domain data, and perform weak data enhancement and strong data enhancement on the target domain assembly image to obtain weakly enhanced target domain data and strongly enhanced target domain data. The pseudo-label generation module is used to construct a domain adaptive detection framework that includes a student network and a teacher network. The weakly enhanced target domain data is input into the teacher network in the domain adaptive detection framework for inference, and pseudo-labels for connectors of the target domain assembly image are generated. An adaptive hybrid cropping module is used to crop and paste the source domain assembly image and the target domain assembly image based on the connector label and the connector pseudo label of the source domain assembly image through an adaptive hybrid cropping strategy based on the intersection-union ratio constraint, thereby generating a cross-domain hybrid image and a hybrid label corresponding to the cross-domain hybrid image. The feature extraction and detection prediction module is used to input the enhanced source domain data, the strongly enhanced target domain data, and the cross-domain hybrid image into the student network in the domain adaptive detection framework. The student network performs forward inference on the enhanced source domain data, the strongly enhanced target domain data, and the cross-domain hybrid image through a YOLOv5L backbone network with embedded C3CBAM feature enhancement module to obtain source domain feature map, target domain feature map, and hybrid image feature map, as well as source domain detection prediction results corresponding to the source domain feature map, strongly enhanced target domain detection prediction results corresponding to the target domain feature map, and hybrid image detection prediction results corresponding to the hybrid image feature map. The feature alignment module is used to generate a multi-scale attention map based on the connector pseudo-label and the target domain feature map, perform foreground region weighted aggregation on the target domain feature map based on the multi-scale attention map to obtain a foreground feature alignment loss, perform adversarial learning alignment on the global features of the source domain feature map and the global features of the target domain feature map through an image-level domain discriminator to obtain an image-level adversarial alignment loss, and perform adversarial learning alignment on the instance features of the source domain feature map and the instance features of the target domain feature map through an instance-level domain discriminator to obtain an instance-level adversarial alignment loss. The joint optimization training module is used to construct a total loss function based on the connector labels, the hybrid labels, the connector pseudo labels, the image-level adversarial alignment loss, the foreground feature alignment loss, the instance-level adversarial alignment loss, the source domain detection prediction results, the strongly enhanced target domain detection prediction results, and the hybrid image detection prediction results of the source domain assembly image. The module then optimizes the total loss function to update the parameters of the student network and updates the parameters of the teacher network using an exponential moving average strategy to obtain the trained target detection model. The test inference module is used to acquire the image of the target domain assembly of the connector during the testing phase, detect the target domain assembly image through the trained target detection model, and output the small target detection result of the connector.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 6.