Target instance segmentation enhancement method and system in mechanical arm assembly operation scene
By introducing the CSSB module into the instance segmentation model, and combining star topology and attention mechanism, the problem of insufficient segmentation accuracy of traditional algorithms in robotic arm assembly scenarios is solved, achieving high-precision segmentation of complex industrial parts, which is suitable for real-time applications of edge computing devices.
Patent Information
- Application Number
- CN202511078105.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-01
- Publication Date
- 2025-11-11
AI Technical Summary
Traditional algorithms are not robust enough in robotic arm assembly scenarios when dealing with lighting changes, occlusion, and segmentation of complex industrial parts, making it difficult to achieve high-precision instance segmentation, especially when dealing with fine edges and occluded scenes, where the segmentation accuracy is insufficient.
The Cross-Stage Star Bottleneck (CSSB) module is designed and integrated into the instance segmentation model by combining star topology and spatial-channel joint attention mechanism. Through multi-scale feature processing and key region focus, the segmentation accuracy is improved.
It improves the segmentation accuracy of industrial parts, especially the edge segmentation accuracy in occluded and high similarity scenarios, reduces computational complexity and the number of parameters, and meets the real-time requirements of robotic arm assembly.
Smart Images

Figure CN120931931A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of instance segmentation technology in computer vision, and in particular to a method and system for enhancing target instance segmentation in a robotic arm assembly operation scenario. Background Technology
[0002] In industrial settings, the environment is complex and lighting conditions are highly variable. Traditional image processing algorithms rely on fixed rules for edge detection and threshold segmentation, which results in poor robustness when faced with changes in lighting, dirt occlusion, reflections, or shadows. For example, reflections on the surface of a container can cause traditional algorithms to misidentify rusted areas, and rainy or foggy weather reduces the accuracy of feature extraction. In strong light conditions, the false detection rate of traditional algorithms can be as high as 35%, making it difficult to meet the requirements for accurate segmentation of industrial parts.
[0003] Industrial parts often have minute damage such as dents and fine cracks. These tiny defects occupy only a few pixels (0.1mm level), and traditional feature descriptors, such as SIFT or SURF, struggle to reliably capture these minute features. Furthermore, traditional methods are also insufficient when processing the fine edges of industrial parts, failing to accurately delineate edge contours.
[0004] In robotic arm assembly scenarios, occlusion between parts is frequent, and some parts have high similarity. Traditional instance segmentation models are prone to false positives and false negatives when dealing with these situations, making it difficult to achieve accurate segmentation of target parts and severely impacting the accuracy and efficiency of robotic arm assembly. In instance segmentation models, traditional modules like the C3k2 in the YOLO series suffer from insufficient feature extraction capabilities, limited receptive fields, and poor information exchange between features of different scales when handling complex industrial scenarios. This results in low segmentation accuracy for industrial parts, especially edge segmentation accuracy (AP75 metric), failing to meet the high-precision visual recognition requirements of industrial production.
[0005] Therefore, in order to address the problems existing in instance segmentation in the above-mentioned industrial scenarios, there is an urgent need for a method that can improve segmentation accuracy, especially when dealing with fine edges and occluded scenes. Summary of the Invention
[0006] In view of the above-mentioned problems in the prior art, the present invention provides a target instance segmentation enhancement method and system for robotic arm assembly operation scenarios, so as to improve the performance of industrial component instance segmentation, especially the accuracy when dealing with fine edges and occlusion scenes.
[0007] The target instance segmentation enhancement method for robotic arm assembly operations provided by this invention specifically includes the following steps:
[0008] Collect an image dataset of industrial components related to robotic arm assembly in industrial scenarios. This dataset should include images of industrial components with fine edges and occlusion to meet the needs of model training and testing in complex scenarios.
[0009] The collected data is preprocessed to improve its quality. This includes data cleaning operations such as deduplication, time sorting, and quality filtering, as well as data augmentation methods such as flipping, rotation, adding Gaussian and ISO noise, and affine transformations to enhance the model's robustness to industrial disturbances and alleviate class imbalance.
[0010] This invention designs a Cross-Stage Star Bottleneck (CSSB) module, which is the core of this invention, and specifically includes the following:
[0011] Star topology design: The input features are divided into two parts. One part is directly transmitted to retain the original feature information, and the other part is transformed through Star Block to obtain enhanced feature representation. Finally, all features are aggregated, and the receptive field is expanded through multi-branch feature interaction to improve the ability to process complex structures and edge details of industrial parts.
[0012] Spatial-channel joint attention mechanism: Introducing a spatial-channel joint attention mechanism into the star topology, by calculating the interaction relationship between query, key, and value matrices and combining it with the channel attention function, enhances the feature weights of component edges and occluded areas, thereby improving the model's attention to key regions.
[0013] The CSSB module is integrated into the existing instance segmentation model, replacing the original C3k2 module, and deployed at the P3, P4, and P5 scale levels of the Feature Pyramid Network (FPN) to form a complete multi-scale feature processing pipeline, thereby improving the segmentation capability of industrial parts of different scales, especially densely arranged parts.
[0014] Industrial part instance segmentation was performed using an instance segmentation model with integrated CSSB module, including the model training and testing process, and the model performance was evaluated.
[0015] Evaluation metrics: Commonly used evaluation metrics in the field of instance segmentation, such as AP50, AP75, and AP50:95, are used to evaluate the model's performance on industrial component instance segmentation tasks, with a focus on improving the model's accuracy when handling fine edges and occluded scenes.
[0016] Model training: Train the model with the integrated CSSB module using the preprocessed dataset, set appropriate training hyperparameters such as batch size, number of training epochs, optimizer, initial learning rate, etc., and continuously adjust the model parameters during training until the model performance tends to stabilize.
[0017] Model testing: The trained model is tested on the test set. The segmentation results output by the model are compared with the labeled real results. The model performance is evaluated according to the evaluation metrics to verify the effect of the CSSB module on improving the segmentation accuracy of industrial parts instances, especially in fine edge and occlusion scenarios.
[0018] The beneficial effects of this invention are as follows:
[0019] The CSSB module designed in this invention realizes multi-branch feature interaction through a star topology, which expands the receptive field and improves the ability to process complex structures and edge details of industrial parts, thereby improving segmentation accuracy, especially edge segmentation accuracy (AP75 index).
[0020] The spatial-channel joint attention mechanism introduced in this invention enhances the feature weights of component edges and occluded regions, improves the model's attention to key regions, and reduces false detections and false negatives in occluded and highly similar component scenarios.
[0021] This invention integrates the CSSB module into the existing instance segmentation model and deploys it across multiple scale levels of FPN to form a multi-scale feature processing pipeline, thereby improving the segmentation capability for industrial parts of different scales, especially densely arranged parts.
[0022] Compared with traditional models, models integrating CSSB modules have reduced computational complexity (FLOPs) and number of parameters, which is beneficial for deployment on edge computing devices and meets the real-time requirements of robotic arm industrial scenarios. Attached Figure Description
[0023] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0024] Figure 1 This is a schematic diagram of the overall structure of the CSSB module in an embodiment of the present invention, with the processes of feature segmentation, star block transformation and feature fusion marked.
[0025] Figure 2 This is a comparison diagram of the segmentation effect of different modules (CSSB vs C3k2) in embodiments of the present invention on obstructed components (such as a scenario where a door handle and a bearing seat overlap).
[0026] Figure 3This is a schematic diagram of the shared convolutional separator segmentation head structure in an embodiment of the present invention, illustrating the specific structure of the Seg SCSS segmentation head. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0028] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0029] like Figure 1 As shown in the figure, this illustrates the overall architecture of the instance segmentation model integrating the Cross-Stage Star Bottleneck (CSSB) module. It comprises three parts: Backbone, Neck, and Head. By combining the CSSB module with multi-scale feature fusion, the model significantly improves the segmentation accuracy of fine edges and occluded parts in industrial scenes while maintaining high computational efficiency. The specific structure is as follows:
[0030] 1. The backbone network consists of multiple stages (Stage 1 to Stage 4). Each stage contains several star blocks and convolutional layers (Conv), which are responsible for extracting features from the input image (size H×W×3).
[0031] Stages 1 through 4 gradually reduce the feature map size through convolution operations (e.g., from H×W to H / 16×W / 16), while increasing the number of channels to achieve multi-scale low-level feature extraction.
[0032] Key modules: Includes the C2PSA module and the SPPF (Spatial Pyramid Pooling-Fast) module, used to enhance feature aggregation capabilities and expand the receptive field.
[0033] 2. Neck network
[0034] A Feature Pyramid Network (FPN) structure is adopted, and the fusion of features at different scales is achieved through Upsample and Concat operations.
[0035] Core improvements: Multiple CSSB modules are deployed in the Neck section, replacing the traditional C3k modules, to handle P3 (small-scale), P4 (medium-scale), and P5 (large-scale) features. The CSSB modules enhance the interaction of multi-scale features and the attention to key regions through a star-shaped topology and a spatial-channel joint attention mechanism.
[0036] Feature transformation: Low-level features (such as Stage4 output) are upsampled and then concatenated with high-level features (such as Stage3 output), and further processed by the CSSB module to form a rich multi-scale feature map.
[0037] 3. Head (Head Network)
[0038] It contains three Seg_SCSS segmentation headers, corresponding to Small, Medium, and Large targets respectively, which are used to output instance segmentation results (mask and category information).
[0039] Each Seg_SCSS segmentation head integrates components such as convolutional layers (Conv), batch normalization (BN), and activation functions (ReLU), achieving accurate segmentation of targets at different scales through shared or non-shared convolutional structures.
[0040] Auxiliary modules: These include GAP (Global Average Pooling) and FC (Fully Connected) layers, used for feature compression and classification decisions.
[0041] like Figure 2 As shown, the performance of the enhanced model based on the CSSB module and the baseline model (without the CSSB module integrated) in the industrial part instance segmentation task is compared. The results include segmentation results from two perspectives. Overall, the two sets of comparison results clearly demonstrate the improvement of the proposed model in industrial part segmentation accuracy (especially edge accuracy and occlusion scenes) from different perspectives, confirming the effectiveness of the CSSB module in complex industrial scenarios. Details are as follows:
[0042] 1. Auxiliary-View Results
[0043] The left side shows the segmentation results of the baseline model, and the right side shows the segmentation results of the present invention model (Ours) with integrated CSSB module.
[0044] The figure shows the segmentation effect of industrial components such as circuit board, bearing, and housing, and marks the confidence level of each component (e.g., circuit board 0.85→0.93, housing 0.90→0.94).
[0045] As can be seen from the comparison, the model of the present invention significantly improves the segmentation confidence of components in the auxiliary view, especially in the more accurate positioning of fine edge regions, which reflects the enhanced ability of the CSSB module to capture subtle features through star-shaped topology and spatial-channel joint attention mechanism.
[0046] 2. Global-View Results
[0047] The left side shows the segmentation result of the baseline model, and the right side shows the segmentation result of the model of this invention.
[0048] The invention highlights the segmentation performance of easily obscured components such as door handles. The baseline model's segmentation confidence for door handles is only 0.77, while the model of this invention improves it to 0.94, significantly reducing the false negative rate.
[0049] This comparison verifies the advantages of the CSSB module in handling occlusion and highly similar components in the global scene. It enhances the attention to key regions through multi-scale feature interaction and attention mechanism, which meets the performance metric of "18.3% reduction in false negative rate in occluded scenes" in the document.
[0050] like Figure 3 As shown, the present invention provides a target instance segmentation enhancement method for robotic arm assembly operations. Specifically, it is an industrial component instance segmentation enhancement method based on the Cross-Stage Star Bottleneck (CSSB) module, which includes the following steps:
[0051] S1: Dataset Preparation and Preprocessing
[0052] Dataset Selection: A dataset specifically designed for industrial assembly scenarios using robotic arms was used. It contains 19 categories of industrial parts (such as screws, nuts, circuit boards, bearing housings, etc.), totaling 9786 instances. The dataset exhibits a long-tail distribution (door handles account for approximately 29%) and covers multi-scale objects ranging from small precision parts to large assemblies. This dataset includes auxiliary views, actuator views, and a global view. Figure 3 This perspective reflects the spatial clustering and size diversity of real industrial assembly layouts.
[0053] Preprocessing steps:
[0054] Data cleaning: Deduplication, time sorting, and quality filtering of raw data to remove fuzzy or invalid samples.
[0055] Data augmentation:
[0056] Geometric transformations: Enhance the model's ability to perceive multi-angle components by randomly flipping (horizontal / vertical) and rotating (-15° to 15°);
[0057] Noise injection: Gaussian noise (mean 0, variance 0.01) and ISO noise (intensity range 100-800) are added to simulate sensor noise and reflection interference caused by changes in lighting in industrial scenarios;
[0058] Affine transformation: applies scaling (0.8-1.2x), translation (±10% pixels), and shearing (±10°) to enhance the geometric invariance of complex shaped parts.
[0059] Dataset partitioning: The dataset was divided into a training set (7829 instances), a validation set (979 instances), and a test set (978 instances) in a ratio of 8:1:1 to ensure that the component category distribution is consistent in each set.
[0060] S2: Design and Integration of the CSSB Module
[0061] 1. Module structure design:
[0062] Core topology: The CSSB module adopts a "feature splitting-star transformation-feature aggregation" architecture. The input features are first divided into two parts: one part is directly transmitted to retain the original feature information, and the other part undergoes multi-branch feature transformation through star blocks. Finally, all features are aggregated through convolutional layers.
[0063] (1) Star Block Design: A star topology is used to achieve multi-scale feature interaction, and its feature transformation can be expressed as:
[0064]
[0065] Where X represents the feature tensor input to the Star Block, and τ is the global transformation function. Let F be a nonlinear mapping function (using LeakyReLU activation), W be a learnable parameter set, and F be a nonlinear mapping function (using LeakyReLU activation). star This is the star-shaped feature extraction function.
[0066] (2) The star-shaped feature extraction function achieves multi-branch interaction through tensor decomposition:
[0067]
[0068] Where X represents the feature tensor input to the Star Block, i = 3 indicates 3 star branches, and P iM is a 3×3 position-sensitive projection matrix. i The branch feature transformation function (composed of 1×1 convolutions and 3×3 depthwise separable convolutions), C i This is the channel attention matrix (generated via the Squeeze-Excitation mechanism). For tensor product operations, F star This is the star-shaped feature extraction function.
[0069] 2. Spatial-channel joint attention mechanism:
[0070] (1) Mechanism formula: Embed a spatial-channel joint attention mechanism in the star-shaped block to enhance the feature weights of component edges and occluded areas:
[0071]
[0072] Where A(X) represents the output of the spatial-channel joint attention mechanism, and X represents the feature tensor input to the Star Block. Let T be the learnable weight matrix, T be the transpose operation, C = 128, σ be the softmax function, and d be the weight matrix. k =C / 8 is the scaling factor, and S(X) is the channel attention function.
[0073] (2) Channel attention function:
[0074] S(X) = σ²(W²σ¹(W¹X) pool ))
[0075] Where S(X) is the channel attention function, X pool This refers to the channel descriptor obtained through global average pooling. Let σ1 be the dimension reduction / incrementing matrix, σ2 be the ReLU activation function, and σ3 be the Sigmoid function.
[0076] Integration with baseline models:
[0077] Using YOLOv11-seg as the baseline model, the CSSB module replaces the C3k2 module in the original model's Feature Pyramid Network (FPN) and is deployed at the P3 (small-scale features), P4 (medium-scale features), and P5 (large-scale features) levels to form a multi-scale feature processing pipeline.
[0078] Module parameter settings: Each CSSB module contains 3 star branches, the depthwise separable convolution kernel size is 3×3, and the number of channels increases with the FPN level (64 for P3, 128 for P4, and 256 for P5).
[0079] S3: Model Training Process
[0080] 1. Training environment and hyperparameters:
[0081] Hardware environment: NVIDIA RTX 3090 GPU (24GB VRAM), Intel i7-9700 CPU, 64GB RAM.
[0082] Software environment: Ubuntu 20.04LTS operating system, PyTorch 1.13.1 deep learning framework, CUDA 11.7 acceleration library.
[0083] Training hyperparameters:
[0084] Optimizer: SGD (Stochastic Gradient Descent), momentum 0.937, weight decay 0.0005;
[0085] Initial learning rate: 0.01, using a cosine annealing strategy (decreasing by 10% every 5 epochs);
[0086] Batch size: 16 (8 samples per GPU);
[0087] Training cycles: 200 epochs, early stop strategy (training is terminated if there is no improvement after 10 consecutive epochs on the validation set AP75).
[0088] 2. Loss Function Design: A multi-task loss function is adopted, including:
[0089] Object detection loss: CIoU loss (bounding box regression) + cross-entropy loss (category classification);
[0090] Instance segmentation loss: Dice loss (mask prediction) + cross-entropy loss (mask classification);
[0091] Attention loss: KL divergence loss (constrains the consistency of attention weight distribution with labeled edge regions).
[0092] 3. Training process:
[0093] Load the preprocessed dataset and use DataLoader to perform batch reading and real-time enhancement.
[0094] Initialize the YOLOv11-seg model of the CSSB module and initialize the weights using the He initialization method;
[0095] In each round of training, forward propagation calculates the loss value, and back propagation updates the model parameters;
[0096] Every 10 epochs, the model performance is evaluated on the validation set (measuring metrics such as AP50, AP75, and AP50:95), and the weights of the best-performing model are saved.
[0097] S4: Model Testing and Performance Verification
[0098] 1. Test metrics:
[0099] Core evaluation indicators:
[0100] AP50: Average accuracy when the IoU threshold is 0.5;
[0101] AP75: Average accuracy with an IoU threshold of 0.75 (focusing on edge segmentation accuracy);
[0102] AP50:95: Average accuracy of IoU threshold from 0.5 to 0.95;
[0103] FLOPs (computational complexity) and number of parameters (metrics for model lightweighting).
[0104] 2. Testing process:
[0105] Load the pre-trained optimal model and set it to evaluation mode (disable Dropout and BatchNorm updates);
[0106] Forward inference is performed on the test set to generate bounding boxes, masks, and class probabilities for each component;
[0107] Compare the values with the marked true values and calculate the various evaluation indicators;
[0108] Visualize and analyze the segmentation results of typical scenarios (such as occluded parts and fine edge parts) and compare them with the baseline model (YOLOv11-seg).
[0109] 3. Performance comparison results:
[0110] Comparison with baseline model:
[0111] AP75 Improvement: From 0.620 to 0.662 (+4.2%);
[0112] Computational efficiency: FLOPs reduced by 22.1%, number of parameters reduced by 37.7%;
[0113] Performance in occlusion scenarios: In scenarios where the door handle and bearing housing overlap, the false negative rate is reduced by 18.3%.
[0114] S5: Deployment of Industrial Applications
[0115] 1. Deployment environment: Edge computing device (NVIDIA Jetson AGX Xavier, 8GB VRAM), adapted to the robotic arm control system (ROS Noetic).
[0116] 2. Real-time performance optimization:
[0117] Using TensorRT for model quantization (INT8 precision), the inference time was reduced from 45ms to 27ms;
[0118] By combining the OpenVINO acceleration library, real-time segmentation processing of 37 frames per second is achieved.
[0119] 3. Application process:
[0120] An industrial camera (1280×720 resolution) captures images of the assembly scene and transmits them to edge devices;
[0121] The model with integrated CSSB module outputs part masks and pose information in real time;
[0122] The robotic arm plans the grasping path based on the segmentation results and completes the assembly operation (such as the alignment and installation of screws and nuts).
[0123] Specifically, the present invention provides a target instance segmentation enhancement system for robotic arm assembly operation scenarios, comprising: a data acquisition module for acquiring an image dataset of industrial components related to robotic arm assembly in industrial scenarios, wherein the dataset contains images of industrial components with fine edges and occlusion.
[0124] A data preprocessing module is used to preprocess the dataset, the preprocessing including data cleaning and data augmentation operations;
[0125] A module design module is used to design CSSB modules, which include a star topology and a space-channel joint attention mechanism.
[0126] The model integration module is used to integrate the CSSB module into the existing instance segmentation model and deploy it on the feature pyramid network to form a multi-scale feature processing flow.
[0127] The segmentation and evaluation module is used for segmenting industrial parts instances using an instance segmentation model that integrates the CSSB module, including the model training and testing process, and evaluating model performance.
[0128] The star topology structure designed by the module includes: dividing the input features into two parts, one part being directly transmitted to retain the original feature information, and the other part being transformed through the Star Block to obtain an enhanced feature representation; finally, all features are aggregated; the formula and channel attention function of the spatial-channel joint attention mechanism are consistent with those described in the above method.
[0129] Furthermore, although the operations of the method of the present invention are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.
[0130] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to the embodiments, those skilled in the art should understand that modifications or equivalent substitutions to the technical solutions of the present invention do not depart from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A target instance segmentation enhancement method for robotic arm assembly operations, comprising the following steps: A dataset of images of industrial components related to robotic arm assembly in industrial settings was collected. The dataset included images of industrial components with fine edges and occlusion. The dataset is preprocessed, including data cleaning and data augmentation operations. Design a CSSB module, which includes a star topology and a space-channel joint attention mechanism; The CSSB module is integrated into the existing instance segmentation model and deployed on the feature pyramid network to form a multi-scale feature processing workflow. Industrial part instance segmentation was performed using an instance segmentation model with integrated CSSB module, including the model training and testing process, and the model performance was evaluated.
2. The method according to claim 1, characterized in that, The data cleaning includes deduplication, time sorting, and quality filtering operations; the data augmentation includes flipping, rotating, adding Gaussian noise, adding ISO noise, and affine transformation operations.
3. The method according to claim 1, characterized in that, The star topology design includes: dividing the input features into two parts, one part being directly transmitted to retain the original feature information, and the other part being transformed through Star Block to obtain an enhanced feature representation, and finally aggregating all features.
4. The method according to claim 3, characterized in that, The Star Block uses a star-shaped topology to achieve multi-scale feature interaction. Its characteristic transformation is expressed as: Where τ is the global transformation function, and X represents the feature tensor input to the Star Block. Let F be a nonlinear mapping function, W be a learnable parameter set, and F be a nonlinear mapping function. star This is the star-shaped feature extraction function. The star-shaped feature extraction function achieves multi-branch interaction through tensor decomposition: Among them, F stat Let X be the star-shaped feature extraction function, where X represents the feature tensor input to the Star Block, i = 3 indicates 3 star branches, and P is the star-shaped feature extraction function. i M is a 3×3 position-sensitive projection matrix. i C is the branch feature transformation function. i For channel attention matrix, This is for tensor product operations.
5. The method according to claim 1, characterized in that, The formula for the spatial-channel joint attention mechanism is: Where A(X) represents the output of the spatial-channel joint attention mechanism, and X represents the feature tensor input to the Star Block. Let T be the learnable weight matrix, T be the transpose operation, C = 128, σ be the softmax function, and d be the weight matrix. k =C / 8 is the scaling factor, and S(X) is the channel attention function. Channel attention function: S(X)=σ2(W2σ1(W1X pool )) Where S(X) is the channel attention function, X pool This refers to the channel descriptor obtained through global average pooling. Let σ1 be the dimension reduction / incrementing matrix, σ2 be the ReLU activation function, and σ3 be the Sigmoid function.
6. The method according to claim 1, characterized in that, The existing instance segmentation model is the YOLOv11-seg model.
7. The method according to claim 1, characterized in that, The loss function used in the model training process is a multi-task loss function, including object detection loss, instance segmentation loss, and attention loss; The target detection loss includes CIoU loss and cross-entropy loss; the instance segmentation loss includes Dice loss and cross-entropy loss. The attention loss is the KL divergence loss.
8. The method according to claim 1, characterized in that, The evaluation metrics used for assessing the performance of the model include AP50, AP75, AP50:95, FLOPs, and the number of parameters.
9. A target instance segmentation and enhancement system for robotic arm assembly operations, characterized in that, include: The data acquisition module is used to acquire image datasets of industrial components related to robotic arm assembly in industrial scenarios. The datasets include images of industrial components with fine edges and occlusion. A data preprocessing module is used to preprocess the dataset, the preprocessing including data cleaning and data augmentation operations; A module design module is used to design CSSB modules, which include a star topology and a space-channel joint attention mechanism. The model integration module is used to integrate the CSSB module into the existing instance segmentation model and deploy it on the feature pyramid network to form a multi-scale feature processing flow. The segmentation and evaluation module is used for segmenting industrial parts instances using an instance segmentation model that integrates the CSSB module, including the model training and testing process, and evaluating model performance.
10. The system according to claim 9, characterized in that, The star topology structure designed by the module includes: dividing the input features into two parts, one part being directly transmitted to retain the original feature information, and the other part being transformed by StarBlock to obtain an enhanced feature representation, and finally aggregating all features; the formula and channel attention function of the spatial-channel joint attention mechanism are consistent with those in claim 5.