Lightweight multi-target instance segmentation method and system for inspection robot
Through the dynamic lightweight network architecture and channel-space global attention mechanism, a lightweight YOLOv11 network is built, which solves the problems of high computing costs and insufficient segmentation accuracy in campus scenarios, and realizes the real-time, high-precision, multi-objective segmentation and obstacle avoidance capabilities of the inspection robot.
Patent Information
- Application Number
- CN202510357065.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
Smart Images

Figure CN120298692A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and deep learning, and relates to a lightweight multi-object instance segmentation method and system for inspection robots. Background Art
[0002] With the accelerating advancement of the construction of smart campuses, campus security management faces increasingly complex challenges. According to statistics, safety accidents caused by crowded people, mixed traffic of vehicles and bicycles on campus account for more than 60%, and most of these accidents are due to the failure to detect and accurately locate targets (such as pedestrians, vehicles, bicycles) in a timely manner. Traditional monitoring systems mostly rely on manual inspections or general target detection technologies, and it is difficult to achieve pixel-level instance segmentation, resulting in the inability to distinguish overlapping targets (such as individuals in a dense crowd, bicycles parked in a mixed manner) in dense scenarios, seriously restricting the efficiency of safety warning and emergency response.
[0003] The emergence of instance segmentation technology provides a new solution for the accurate recognition of multiple targets in campus scenarios. Through pixel-level target contour segmentation, this technology can accurately distinguish the boundaries of different category targets and output category and mask information in real time, thus supporting inspection robots for intelligent obstacle avoidance, path planning and early warning of abnormal events. For example, in the campus main road scenario, the instance segmentation algorithm can identify pedestrians, vehicles and bicycles in real time to avoid traffic congestion and collision risks; in large-scale gathering activities, accurate crowd density analysis and individual positioning help prevent stampede accidents. In addition, combined with the Internet of Things technology, the instance segmentation results can be linked to the campus security system to achieve automated emergency response.
[0004] However, the application of existing instance segmentation models (such as Mask R-CNN, YOLO series) in campus scenarios faces significant bottlenecks. First, the number of parameters of general models is huge (for example, the number of parameters of Mask R-CNN exceeds 44M), and the computational cost is high (the computational volume of YOLOv8 reaches 28.4 GFLOPs), making it difficult to be deployed on resource-constrained embedded inspection robot platforms. Second, general datasets (such as COCO) contain 80 categories of targets, while campus scenarios only need to focus on a few categories such as people, vehicles, bicycles, etc. Redundant category information will lead to waste of model computing resources and a decrease in accuracy. In addition, problems such as changes in campus environmental lighting, differences in target scales (such as pedestrians nearby vs. bicycles in the distance), and dense occlusion further exacerbate the false detection and missed detection rates of the segmentation algorithm.
[0005] In traditional solutions, manual monitoring relies on security personnel to continuously observe the video stream, which has problems such as response delay and high labor costs. Research shows that the target miss detection rate of manual monitoring exceeds 35% in dense scenarios, and it is difficult to achieve 24-hour continuous coverage. On the other hand, methods based on lightweight object detection (such as YOLO Nano) can reduce the computational load, but they can only output bounding box information and cannot meet the pixel-level contour accuracy required for obstacle avoidance navigation. For example, the overlap of the bounding boxes of bicycles and pedestrians may lead to misjudgment of the obstacle avoidance path, while the precise mask provided by instance segmentation can effectively avoid such problems.
[0006] Therefore, there is an urgent need for a lightweight multi-object instance segmentation method for inspection robots, which can significantly compress the model parameter quantity and computational cost while ensuring high-precision segmentation, so as to adapt to the real-time processing requirements of embedded devices and provide reliable technical support for the safety management of smart campuses. Summary of the Invention
[0007] In view of this, the purpose of the present invention is to provide a lightweight multi-object instance segmentation method and system for inspection robots. It provides an efficient solution for dense object segmentation in campus scenarios through a dynamic lightweight network architecture and a channel-spatial global attention mechanism. The dynamic lightweight architecture significantly reduces the model parameter quantity and computational cost through grouped convolution recombination, dynamic weight allocation, and cheap feature transformation, adapting to the low computing power environment of embedded inspection devices. The channel-spatial global attention mechanism (C2CGA) combines cascaded group attention and deformable convolution spatial focusing to screen key features in the channel dimension and locate the target contour in the spatial dimension, thereby improving the segmentation accuracy and complex background robustness in dense occlusion scenarios. Combining the above technologies, an efficient, accurate, and campus-scenario-adapted lightweight instance segmentation model can be constructed to provide real-time and reliable environment perception capabilities for inspection robots.
[0008] To achieve the above purpose, the present invention provides the following technical solutions:
[0009] On the one hand, a lightweight multi-object instance segmentation method for inspection robots is proposed, and the method includes the following steps:
[0010] Obtain a public dataset, perform data cropping and data augmentation on the public dataset, and construct a campus scenario dataset for training;
[0011] Construct an improved lightweight YOLOv11 network, which at least adaptively adjusts weights according to input features through a dynamic convolution module, and collaboratively optimizes feature fusion through cascaded group attention and channel-spatial attention;
[0012] Train the improved lightweight YOLOv11 network according to the constructed campus scenario dataset to obtain its optimal weights;
[0013] Deploy the trained lightweight instance segmentation model to the inspection robot to process the campus scene image data in real time and output the target pixel-level contour and class information.
[0014] Furthermore, the improved lightweight YOLOv11 at least includes a backbone network and a head network. Among them, the backbone network at least includes several EIEStem lightweight feature extraction modules, C3k2_DynamicConv dynamic convolution modules, C3k2_GhostDynamicConv modules, and LAWDS composite lightweight modules, as well as at least one SPPF module and one C2CGA module; in the backbone network, multi-scale feature extraction is performed through the combination of several groups of EIEStem modules, LAWDS modules, C3k2_DynamicConv modules, C3k2_GhostDynamicConv modules, and C2CGA modules. Among them, the EIEStem module reduces the computational complexity through grouped convolution and channel recombination; the C3k2_DynamicConv module adaptively adjusts the weights according to the input features to reduce the number of parameters; the C3k2_GhostDynamicConv combines the cheap linear transformation of Ghost convolution and dynamic convolution to compress the parameters while maintaining the feature expression ability; the LAWDS module retains the high-frequency detail information through adaptive weighted pooling; the extracted highest-scale features are input into the SPPF module and the C2CGA module for multi-scale feature fusion;
[0015] The head network at least includes several Upsample upsampling modules, Concat splicing modules, C3k2_DynamicConv modules, C3k2_GhostDynamicConv modules, LAWDS modules, and Segment_Efficient modules; in the head network, first perform upsampling through the Upsample module, then splice feature maps with different resolutions through the Concat module, and then further process the features using C3k2_DynamicConv, C3k2_GhostDynamicConv, and LAWDS. Finally, the Segment_Efficient layer segments targets at different scales.
[0016] Furthermore, the C3k2_DynamicConv dynamic convolution module in the improved lightweight YOLOv11 network combines a dynamic convolution kernel and a C3 structure for dynamic feature extraction and parameter optimization. Among them, for the input feature map X ∈ R H×W×C , its processing process is as follows:
[0017] First, extract the channel dimension statistical characteristics of the feature map , that is, the mean μ and variance σ2 ; Concatenate the mean and variance into a vector Predict the dynamic weight matrix through a small fully-connected network f FC Predict the dynamic weight matrix K represents the convolution kernel size, R << K 2 If it is a reduction factor, the dynamic weight matrix is expressed as:
[0018] W dyn = f FC (S) = W2·ReLU(W1S + b1) + b2
[0019] In the formula, W1 and W2 represent the weights of the first and second layers respectively, b1 and b2 represent the biases of the first and second layers respectively, and ReLU(·) represents a non-linear activation function;
[0020] Then, combine the base weight with the dynamic weight W dyn to construct the final convolution kernel:
[0021] W final = W base ·W dyn
[0022] Finally, place the obtained final convolution kernel in the second convolutional layer of the C3 module of the YOLOv11 network for multi-branch feature fusion:
[0023] Y dynamic = Conv3(Dyna, ocConv2(Conv1(X)) + X)
[0024] In the formula, Conv1(·) is the first convolutional operation of the C3 module, DynamicConv2(·) is the second dynamic convolutional operation of the C3 module, the dynamic convolution kernel adaptively adjusts the weights according to the input features, and Conv3(·) is the third convolutional operation of the C3 module.
[0025] Furthermore, C3k2_GhostDynamicConv in the improved lightweight YOLOv11 network is based on the lightweight convolution idea and reduces the computational complexity by decomposing the standard convolution into multi-stage operations; its core component GhostModule divides the feature generation into two steps: first, generate some features with ordinary convolution, and then supplement the detailed features through depthwise convolution; the computational process of C3k2_GhostDynamicConv is expressed as:
[0026] Input feature map First, pass through the basic convolutional layer to obtain intermediate features Among them, the number of hidden channels C = C2×e, where C2 is the number of output channels and e∈(0,1) is the expansion factor; subsequently, feature transformation is performed through n cascaded GhostModules. The calculation of each module is divided into a main path and a secondary path, and then the output results of the main path and the secondary path are concatenated. The process is expressed as:
[0027]
[0028] In the formula, m = C / s, s is the splitting ratio, k is the kernel size, Y1 represents the output of the main path, Y2 represents the output of the secondary path, and Y represents the concatenation result of the outputs of the main path and the secondary path;
[0029] In C3k2_GhostDynamicConv, the module type is selected according to the c3k flag: when c3k = True, the C3k structure containing two GhostModules is adopted, and n is used to control the stacking times; otherwise, a single GhostModule is directly used; the parameter g controls the number of groups of grouped convolution, and shortcut determines whether to add a residual connection.
[0030] Furthermore, the EIEStem module in the improved lightweight YOLOv11 network realizes efficient feature encoding through multi-branch feature extraction and fusion. Its processing process is divided into four stages. Let the input feature map be There is:
[0031] Primary convolution, the first 3×3 convolution for spatial compression: input Extract features through Conv1:
[0032]
[0033] Among them, W1 is hidc 3×3 convolution kernels;
[0034] Dual-branch processing, parallel execution of edge enhancement and feature retention: The Sobel branch applies the directional gradient operator:
[0035] Y edge = Soble(Y1; G x , g y )
[0036] Among them, G x = [[-1,0,1],[-2,0,2],[-1,0,1]] and respectively represent the horizontal / vertical Sobel operators;
[0037] The pooling branch adopts special max pooling:
[0038] Y pool= MaxPool(ZeroPad(Y1))
[0039] Maintain the feature resolution through (0,1,0,1) padding and pooling window;
[0040] Feature concatenation and fusion, followed by secondary downsampling along the channel dimension:
[0041] Y2 = Conv(Concat(Y edge ,Y pool ); W2
[0042] where W2 is hidc 3×3 convolutional kernels, and compress the feature map to 1 / 4 of the original size again;
[0043] Channel compression, 1×1 convolution to adjust the final dimension:
[0044]
[0045] where, W3 is ouc 1×1 convolutional kernels, achieving channel compression from C h → C out ; C in ,C h ,C out correspond to the number of input / hidden / output channels respectively, and H / W is the input spatial dimension.
[0046] Furthermore, the lightweight adaptive weight downsampling module LAWDS in the improved lightweight YOLOv11 network realizes adaptive spatial compression by dynamically allocating regional weights. Its principle is divided into two stages: attention weight generation and downsampling fusion. Let the input feature map dimension be where B is the batch size and C is the number of channels, and there is:
[0047] Attention weight generation:
[0048]
[0049] where AvgPool(·) is the average pooling operation; φ is the candidate weight corresponding to the spatial position generated by 1×1 convolution. Among them, the parameters of the 1×1 convolution Softmax normalizes along the last dimension; each element A b,c,i,j,k represents the attention weight of the k-th region at the position (i,j) of the b-th batch, c-th channel;
[0050] Adaptive downsampling:
[0051]
[0052] Generate downsampled features with four times the number of channels through grouped convolution, where the number of groups G = C / group, and reorganize the downsampled features as follows:
[0053]
[0054] The final output is obtained through weighted fusion:
[0055]
[0056] where ⊙ represents element-wise multiplication, and the feature at each position is the weighted sum of the features of its four candidate regions and the corresponding attention weights.
[0057] Furthermore, in the improved lightweight YOLOv11 network, the C2CGA module performs collaborative optimization of cascaded group attention and channel-spatial attention. Specifically, for the input feature map It is divided into G subgroups according to the channel dimension, and each subgroup independently calculates channel and spatial attention. The attention weights output by the previous subgroup are used as the prior information for the next subgroup, and the feature focusing region is gradually refined through cascaded transmission. The process is expressed as:
[0058] Group division: Divide X into {X1, X2, …, X G}, where X g ∈R H×W×C / G ;
[0059] Cascaded processing: The input to the g-th group is X g and the attention weights A g-1 output by the previous group, and the global perception ability is gradually enhanced through iterative optimization;
[0060] Channel-spatial attention collaboration: Perform global average pooling on each group of features X g to generate a channel statistical vector Z g ∈R C / G , and generate channel weights through a lightweight fully connected network
[0061]
[0062] where r is the compression ratio, and σ(·) represents an activation function, usually the Sigmoid function, which restricts the weights to [0,1].
[0063] Introduce deformable convolution to predict the spatial offset K is the convolution kernel size, and the sampling position is dynamically adjusted to focus on the key regions of the target:
[0064]
[0065] where pk is the position of the conventional convolution kernel, and ω g,k is the learnable weight;
[0066] Then, cross-fusion and weight transfer are performed: the channel weight and the spatial weight are multiplied element-wise to generate the subgroup attention map A g , and it is passed to the next group through concatenation:
[0067]
[0068] X g+1 = X g ⊙ A g
[0069] Finally, all subgroup outputs are concatenated into the global feature map X out ∈ R H×W×C .
[0070] Furthermore, in the head network of the improved lightweight YOLOv11 network, the high feature map and the low feature map are aligned and fused. The size consistency is adjusted by upsampling or downsampling. Suppose the low-level feature map is adjusted to the same spatial dimension as the high-level feature map through upsampling:
[0071] X low_upsamp = Upsample(X low , H', W')
[0072] Then, the two are concatenated and fused and used as the input for subsequent feature extraction:
[0073] X fusion = Concat(X low_upsamp , X high )
[0074] The processing methods of the C3k2_DynamicConv module, C3k2_GhostDynamicConv, and LAWDS composite module in the head network are the same as those in the backbone network;
[0075] The Segment_Efficient module includes a detection branch and a segmentation branch. The detection branch predicts the bounding box box and the class cls through multiple layers of convolution; the segmentation branch generates a mask through the prototype mask Proto and the dynamic coefficient. The specific process is as follows:
[0076] Detection branch: The input feature map is extracted by the stem module, and then the regression parameters and class probabilities are output through cv2 and cv3 respectively. The regression parameters are decoded into bounding box coordinates through DFL:
[0077] dbox = DFL(box) × anchors + strides
[0078] Among them, box is the regression parameter, and its shape is B × 4 × reg max × H × W; anchors are the anchor coordinates, strides are the feature map strides, and reg_max represents the number of discretization intervals;
[0079] Segmentation branch:
[0080] Prototype generation: The Proto module converts the input feature map into npr prototype masks:
[0081] P = Conv(ch[0] → npr)
[0082] Among them, ch[0] is the number of channels of the input feature map, npr is the number of prototypes, and the dimension of P is B × npr × H p × W p , representing the basic mask template;
[0083] Coefficient generation: Generate nm dynamic coefficients from the feature maps of each layer through the cv4 module:
[0084] M = Conv(x → c4 → nm)
[0085] Among them, nm is the number of instance masks, is the number of intermediate channels, and the dimension of M is B × nm × H × W, representing the mask combination weight corresponding to each position;
[0086] Mask fusion:
[0087] The final mask of the instance is generated by the linear combination of the prototype mask and the coefficient:
[0088] Mask i = σ(M i · T )
[0089] Among them, M i is the coefficient vector corresponding to the i-th detection box, σ is the Sigmoid function, and P T is the spatial expansion form of the prototype, and the dimension of the final mask is H p × W p , aligned with the resolution of the input image.
[0090] Furthermore, the process of training the improved lightweight YOLOv11 network is as follows:
[0091] Input the preprocessed data into the improved lightweight YOLOv11 network;
[0092] Its backbone network uses the EIEStem module for efficient initial feature extraction, and realizes multi-scale feature abstraction through the C3k2_DynamicConv dynamic convolution module and the C3k2_GhostDynamicConv composite lightweight module, and combines the LAWDS lightweight adaptive weighted downsampling module to retain high-frequency detail information; in the feature fusion stage, the C2CGA channel-spatial global attention module is used to dynamically allocate channel and spatial dimension weights;
[0093] The head network uses the Segment_Efficient lightweight segmentation module, and generates pixel-level mask predictions through depthwise separable convolution and dynamic upsampling, and outputs the class probability, confidence, and contour parameters of the target instance;
[0094] During the training process, a multi-task loss function jointly optimized by the dynamic focal loss and the boundary-aware mask loss is used, which is expressed as:
[0095] L total =λ cls L cls +λ mask L mask +λ boundray L boundary
[0096] In the formula, L cls represents the classification loss, the dynamic focal loss, which solves the problem of class imbalance, and λ cls represents the weight coefficient of the classification loss, which controls the importance of the classification task; L msak represents the mask loss, which measures the overlap degree between the predicted mask and the true mask, and λ msak represents the weight coefficient of the mask loss, which balances the optimization intensity of the mask accuracy; L boundary represents the boundary loss, which penalizes the contour prediction deviation, and λ boundary represents the weight coefficient of the boundary loss, which enhances the contour alignment ability; the network parameters are updated through the Adam optimizer and the gradient descent method, and the accuracy and robustness of the model for multi-object instance segmentation in the campus scene are gradually improved.
[0097] On the other hand, a lightweight multi-object instance segmentation system for inspection robots is also proposed. The system at least includes a robot body, a housing, a four-wheel drive chassis, a support frame, a mobile control component, and a multi-modal sensor component. Among them, the housing is installed outside the robot body, and the four-wheel drive chassis is set at the bottom. The four-wheel drive chassis is equipped with an independent suspension system, and the support frame is rigidly connected to the housing;
[0098] The core of the platform uses a NUC host as the computing processor, runs the ROS robot operating system, and deploys the improved lightweight YOLOv11 network constructed in the aforementioned lightweight multi-object instance segmentation method for inspection robots.
[0099] The platform integrates an Intel RealSense D455 depth camera and a lidar, which are respectively used for high-resolution RGB-D data acquisition and three-dimensional space modeling.
[0100] Through the ROS communication framework, multi-sensor data is input into the segmentation network after being optimized in real time by the C2CGA channel-spatial global attention module, generating pixel-level masks and class probabilities for pedestrians, vehicles, and bicycles.
[0101] The robot supports dual modes of autonomous cruise and remote control. Control commands are transmitted to the brushless DC motor via the CAN bus for centimeter-level positioning accuracy and multi-radius steering control. During the autonomous cruise process, the depth camera continuously collects campus scene images. After being processed in real time by the Jetson platform, the segmentation results are published to the navigation decision-making node through ROS topics to drive the robot to dynamically avoid obstacles. At the same time, the contour and class information of key targets are uploaded to the large screen of the campus security center through a low-latency wireless communication module to support real-time monitoring and emergency dispatching by management personnel.
[0102] The beneficial effects of the present invention are as follows:
[0103] In the instance segmentation network architecture of the present invention, the improved lightweight YOLOv11-seg_multi-object adopts a dynamic convolution module (C3k2_DynamicConv) and a C2CGA channel-spatial global attention module. Through dynamic weight allocation and group-level cascade attention mechanism, the number of model parameters and computational complexity are significantly reduced, enabling the segmentation network to process high-resolution images in real time on an embedded inspection robot platform and ensuring efficient operation in complex campus scenarios.
[0104] Secondly, to improve the segmentation accuracy of dense targets, the present invention introduces a LAWDS lightweight adaptive weighted downsampling module. Through multi-scale pooling fusion and dynamic weight allocation, high-frequency detail information of small targets (such as bicycles and pedestrian limbs) is retained, and the segmentation accuracy of small targets is guaranteed in the COCO subset test.
[0105] In addition, for the complex background interference in the campus scene, the C2CGA module combines the spatial offset prediction of deformable convolution and channel attention screening to accurately locate the target contours (such as vehicle edges and crowd gaps), improving the segmentation robustness in scenarios with light changes and occlusions, which is significantly better than traditional instance segmentation models (such as Mask R-CNN).
[0106] Finally, through the depthwise separable convolution and dynamic upsampling design of the Segment_Efficient lightweight segmentation head, the mask generation speed is improved, and the edge continuity error is reduced, providing a pixel-level reliable input for real-time obstacle avoidance and path planning of the inspection robot.
[0107] Other advantages, objectives, and features of the present invention will be described to some extent in the following specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0108] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:
[0109] Figure 1 It is a schematic structural diagram of the lightweight YOLOv11-seg_multi-object network constructed in the embodiment of the present invention;
[0110] Figure 2 It is a schematic structural diagram of the EIEStem module in the embodiment of the present invention;
[0111] Figure 3 It is a diagram of the DynamicConv expert system in the embodiment of the present invention;
[0112] Figure 4 It is a schematic structural diagram of the GhostModule in the embodiment of the present invention;
[0113] Figure 5 It is a schematic structural diagram of the LAWDS module in the embodiment of the present invention;
[0114] Figure 6 It is a schematic structural diagram of the C2CGA module in the embodiment of the present invention;
[0115] Figure 7 It is an example of the multi-object instance segmentation output in the embodiment of the present invention;
[0116] Figure 8 It is a comparison schematic diagram of the accuracy index mAP and the computational complexity index GFLOPs of the multi-object instance segmentation of the lightweight model YOLOv11-seg_multi-object and the existing YOLOv11 model in the same dataset in the embodiment of the present invention, where Figure 8 (a) is the training result of the original unimproved YOLOv11 model; Figure 8(b) is the training result of the YOLOv11-seg_multi-object model of the present invention. Specific implementation manners
[0117] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0118] Among them, the drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and cannot be understood as a limitation to the present invention; in order to better illustrate the embodiments of the present invention, some components in the drawings will be omitted, enlarged or reduced, and do not represent the dimensions of actual products; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0119] In the drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the drawings are only for illustrative purposes and cannot be understood as a limitation to the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0120] Please refer to Figures 1 to 8 , which is a lightweight multi-object instance segmentation method and system for inspection robots.
[0121] Embodiment 1
[0122] This embodiment provides a detailed process of a lightweight multi-object instance segmentation method for inspection robots. As Figure 1 shown, it specifically includes the following steps:
[0123] S1. Based on the publicly available COCO dataset, construct training data adapted to the campus scenario through data augmentation strategies and cropping the dataset;
[0124] S2. Construct an improved lightweight YOLOv11 network, which at least adaptively adjusts weights according to input features through a dynamic convolution module, and collaboratively optimizes feature fusion by cascading group attention and channel-spatial attention;
[0125] S3. Train the improved lightweight YOLOv11 network to obtain optimal weights;
[0126] S4. Deploy the trained lightweight instance segmentation model to the inspection robot, process campus scene image data in real time, and output target pixel-level contours and category information.
[0127] In step S1 of this embodiment, by randomly cropping the original dataset, a lightweight training subset is generated to reduce the computational overhead of model training. The cropped data subsets (training set, test set) retain all 80 category annotation information of the COCO dataset. In the model training and inference stages, the adaptability of the instance segmentation task for the targets of three categories, namely people, vehicles, and bicycles, adapted to the campus scene, is optimized. The data augmentation strategy adopts the data augmentation method of horizontal flipping and random scaling of the dataset to simulate the perspective of mirror symmetry of the target object and the change of the object's distance, enhancing the model's robustness to left-right direction changes and multi-scale detection ability.
[0128] In step S2 of this embodiment, as Figure 1 shown, improve the latest YOLOv11 model. The structure of the constructed improved lightweight YOLOv11-seg_multi-object network model at least includes a backbone network and a head network. Among them, the backbone network at least includes several EIEStem lightweight feature extraction modules, C3k2_DynamicConv dynamic convolution modules, C3k2_GhostDynamicConv, LAWDS composite lightweight modules, as well as at least one SPPF module and one C2CGA module. In the backbone network, multi-scale feature extraction is performed through the combination of several groups of EIEStem modules, LAWDS modules, C3k2_DynamicConv modules, C3k2_GhostDynamicConv modules, and C2CGA modules. Among them, the EIEStem module reduces the computational amount through grouped convolution and channel recombination; the C3k2_DynamicConv module adaptively adjusts weights according to input features to reduce the number of parameters; the C3k2_GhostDynamicConv combines the cheap linear transformation of Ghost convolution and dynamic convolution to further compress parameters while maintaining the feature expression ability; the LAWDS module retains high-frequency detail information through adaptive weighted pooling to improve the segmentation accuracy of small targets; the extracted highest-scale features are input into the SPPF module and the C2CGA module to enhance the multi-scale feature fusion effect.
[0129] Specifically, as Figure 2 shown, the content of the EIEStem lightweight feature extraction module is as follows:
[0130] The EIEStem module realizes efficient feature encoding through multi-branch feature extraction and fusion. Its processing process is divided into four stages (assuming the input feature map is ):
[0131] Primary convolution: The first layer of 3×3 convolution performs spatial compression
[0132] The input X ∈ R^{N×inc×H×W} extracts features through Conv1 (3×3 convolution, stride = 2, output channels hidc):
[0133]
[0134] where W1 is hidc 3×3 convolution kernels, and the stride s = 2, realizing channel expansion (C in →C h ) and halving the size.
[0135] Two-branch processing: Edge enhancement and feature retention are executed in parallel
[0136] The Sobel branch applies the directional gradient operator:
[0137] Y edge = Soble(Y1; G x , G y )
[0138] G x = [[-1,0,1],[-2,0,2],[-1,0,1]] and represent the horizontal / vertical Sobel operators respectively.
[0139] The pooling branch uses special max pooling:
[0140] Y pool = MaxPool(ZeroPad(Y1))
[0141] Maintains the feature resolution through (0,1,0,1) padding and a 2×2 pooling window (stride 1).
[0142] Feature concatenation and fusion: After concatenation along the channel dimension, secondary downsampling
[0143] Y2 = Conv(Concat(Y edge , Y pool ) ; W2
[0144] Among them, W2 is hidc 3×3 convolutional kernels, and the stride s = 2 compresses the feature map to 1 / 4 of the original size again.
[0145] Channel compression: 1×1 convolution adjusts the final dimension
[0146]
[0147] W3 is ouc 1×1 convolutional kernels, achieving channel compression from C h →C out to C.
[0148] This structure enhances edge features through the Sobel operator, retains detailed features with asymmetric pooling, and finally constructs an efficient feature extraction backbone through hierarchical downsampling and channel adjustment. Among them, C un , C h , C out correspond to the input / hidden / output channel numbers respectively, and H / W is the input spatial dimension.
[0149] As Figure 3 shown, the content of the C3k2_DynamicConv module is as follows:
[0150] The C3k2_DynamicConv dynamic convolution module combines dynamic convolutional kernels with the C3 structure for dynamic feature extraction and parameter optimization. Among them, for the input feature map X ∈ R H×W×C , its processing process is as follows:
[0151] First, extract the statistical characteristics of the channel dimension of the feature map , that is, the mean μ and variance σ 2 ; concatenate the mean and variance into a vector and predict the dynamic weight matrix FC through a small fully connected network f K represents the convolutional kernel size, and R << K 2 is the reduction factor, then the dynamic weight matrix is expressed as:
[0152] W dyn = f FC (S) = W2·ReLU(W1S + b1) + b2
[0153] In the formula, W1 and W2 represent the weights of the first and second layers respectively, b1 and b2 represent the biases of the first and second layers respectively, and ReLU(·) represents a non-linear activation function;
[0154] Then, combine the base weight with the dynamic weight W dyn to construct the final convolutional kernel:
[0155] W final = Wbase ·W dyn
[0156] Finally, place the obtained final convolution kernel in the second convolutional layer of the C3 module of the YOLOv11 network for multi-branch feature fusion:
[0157] Y dynamic = Conv3(DynamicConv2(Conv1(X)) + X)
[0158] In the formula, Conv1(·) is the first convolutional operation of the C3 module, DynamicConv2(·) is the second dynamic convolutional operation of the C3 module, the dynamic convolution kernel adaptively adjusts the weights according to the input features, and Conv3(·) is the third convolutional operation of the C3 module.
[0159] The dynamic convolution kernel adaptively adjusts the weights according to the input features, enhancing the adaptability to multi-scale targets (such as pedestrians and bicycles) in the campus scene.
[0160] Such as Figure 4 As shown, the content of the C3k2_GhostDynamicConv module is as follows:
[0161] C3k2_GhostDynamicConv is based on the lightweight convolution idea and reduces the computational complexity by decomposing the standard convolution into multi-stage operations. Its core component, GhostModule, divides feature generation into two steps: first, generate partial features using ordinary convolution, and then supplement the detailed features through depth convolution. Taking C3k_GhostDynamicConv as an example, its calculation process can be expressed as:
[0162] Input feature map First, pass through the basic convolutional layer to obtain intermediate features Among them, the number of hidden channels C = C2 × e (C2 is the number of output channels, and e ∈ (0, 1) is the expansion factor). Subsequently, perform feature transformation through n cascaded GhostModules. The calculation of each module is divided into:
[0163] Main path: (m = c_ / s, s is the segmentation ratio, default is 2)
[0164] Side path: (k is the convolution kernel size)
[0165] Output concatenation:
[0166] In C3k2_GhostDynamicConv, the module type is selected according to the c3k flag: when c3k = True, the C3k structure containing two GhostModules is adopted (n in the formula controls the stacking times), otherwise a single GhostModule is directly used. The parameter g controls the number of groups for grouped convolution, and shortcut determines whether to add a residual connection. This design reduces the computational amount by about s / (s + 1) while keeping the number of channels c_ unchanged by reducing redundant feature calculations (only calculating the standard convolution of m = c_ / s channels) and reusing depth convolution features.
[0167] As Figure 5 shown, the content of the LAWDS module is as follows:
[0168] The LAWDS module (Lightweight Adaptive Weight Downsampling) realizes adaptive spatial compression by dynamically allocating regional weights. Its principle can be divided into two stages: attention weight generation and downsampling fusion (assuming the input feature map dimension is ), where B is the batch size and C is the number of channels):
[0169] Attention weight generation:
[0170]
[0171] Among them, AvgPool uses a 3×3 kernel for feature smoothing, and φ is a 1×1 convolution (parameter ) to generate 4 candidate weights corresponding to spatial positions, and Softmax normalizes along the last dimension. Each element A b,c,i,j,k represents the attention weight of the kth region at position (i, j) in batch b, channel c.
[0172] Adaptive downsampling:
[0173]
[0174] Generate downsampled features with quadruple channels through a 3×3 grouped convolution with a stride of 2 (the number of groups G = C / group), and reorganize them into:
[0175]
[0176] The final output is obtained through weighted fusion:
[0177]
[0178] Among them, ⊙ represents element-wise multiplication, and the feature at each position is the weighted sum of the features of its four candidate regions and the corresponding attention weights. This method dynamically adjusts the downsampling strategy according to the importance of local regions, reducing information loss while maintaining lightweight.
[0179] As Figure 6 shown, the content of the C2CGA module is as follows:
[0180] The C2CGA (Channel-Spatial Cascaded Group Attention) module is a feature enhancement module improved from the Cascaded Group Attention mechanism. Its core design combines the group cascading strategy and channel-spatial attention for collaborative optimization. The principle is as follows:
[0181] Cascaded Group Attention mechanism:
[0182] The input feature map X ∈ R H×W×C is divided into G subgroups along the channel dimension. In this embodiment, G = 4, and each group independently calculates channel and spatial attention. The attention weights output by the previous group are used as prior information for the next group, and the feature focusing area is gradually refined through cascaded transmission. The specific process is as follows:
[0183] Group division: Divide X into {X1, X2, …, X G}, where X g ∈ R H×w×C / G .
[0184] Cascaded processing: The input of the g-th group is X g and the attention weights A g-1 output by the previous group, and the global perception ability is gradually enhanced through iterative optimization.
[0185] Channel-Spatial Attention Collaboration:
[0186] Channel attention: Perform global average pooling (GAP) on each group of features X g to generate a channel statistical vector Z g ∈ R C / G , and generate channel weights through a lightweight fully connected network
[0187]
[0188] where r is the compression ratio, and σ(·) represents the activation function, usually the Sigmoid function, which restricts the weights to [0, 1].[[]END]
[0189] Spatial attention: Introduce deformable convolution to predict the spatial offset (K is the convolution kernel size), and dynamically adjust the sampling position to focus on the key target area:
[0190]
[0191] where p k is the position of the conventional convolution kernel, and ω g,k is the learnable weight.
[0192] Cross-fusion and weight transfer:
[0193] Multiply the channel weight and the spatial weight element-wise to generate the subgroup attention map A g , and pass it to the next group through concatenation:
[0194]
[0195] X g+1 = X g ⊙ A g
[0196] Finally, all subgroup outputs are concatenated into the global feature map X out ∈ R H×W×C .
[0197] The head network includes at least several Upsample upsampling modules, Concat concatenation modules, C3k2_DynamicConv modules, C3k2_GhostDynamicConv modules, LAWDS modules, and Segment_Efficient modules; first, perform upsampling through the Upsample module, then concatenate feature maps of different resolutions through the Concat module, then further process the features using C3k2_DynamicConv, C3k2_GhostDynamicConv, and LAWDS, and finally segment targets of different scales by the Segment_Efficient layer. The main module processes and principles are as follows:
[0198] The function of the Contact module is to enhance the model's expressive ability by cross-layer feature connection and fusion, retain low-level and high-level feature information, and thus better capture complex patterns.
[0199] Align and fuse the high feature map X high ∈ R H×W×C and the low feature map X low ∈ R H×W×C by adjusting the size consistency through upsampling or downsampling, such as using transposed convolution or pooling operations, to ensure they are aligned in the spatial dimension. Assume that the low-level feature map is adjusted to the same spatial dimension as the high-level feature map through upsampling:
[0200] X low_upsamp = Upsample(X low,H',W')
[0201] Then, the two are spliced and fused and used as the input for subsequent feature extraction;
[0202] X fusion = Concat(X low_upsamp ,X high )
[0203] The processing methods of the C3k2_DynamicConv module, C3k2_GhostDynamicConv module, and LAWDS composite module in the head network are the same as those in the backbone network.
[0204] The content of the Segment_Efficient module in the head network is as follows:
[0205] The Segment_Efficient module inherits from Detect_Efficient and extends the instance segmentation function on the basis of object detection. Its core principle is divided into two parts: the detection branch and the segmentation branch. The detection branch follows the structure of Detect_Efficient and predicts the bounding box (box) and class (cls) through multiple layers of convolution; the segmentation branch generates masks through the prototype mask (Proto) and dynamic coefficients. The specific principle is as follows:
[0206] Detection branch: The input feature map extracts features through the stem module (two 3x3 grouped convolutions), and then outputs the regression parameters and class probabilities through cv2 and cv3 respectively. The regression parameters are decoded into bounding box coordinates through DFL (Distribution Focal Loss):
[0207] dbox = DFL(box) × anchors + strides
[0208] where box is the regression parameter (shape: B×4×reg max ×H×W), anchors are the anchor coordinates, strides are the feature map strides, and reg_max = 16 represents the number of discretization intervals.
[0209] Segmentation branch: Prototype generation: The Proto module converts the input feature map (number of channels: ch[0]) into npr prototype masks:
[0210] P = Conv(ch[0] → npr)
[0211] where npr = 256 is the number of prototypes, and the dimension of P is B×npr×H p ×W p , representing the basic mask template.
[0212] Coefficient generation: Generate nm dynamic coefficients from the feature maps of each layer through the cv4 module (two 3x3 convolutions + 1x1 convolution):
[0213] M = Conv(x -> c4 -> nm)
[0214] where nm = 32 is the number of instance masks, is the number of intermediate channels, and the dimension of M is B × nm × H × W, representing the mask combination weights corresponding to each position.
[0215] Mask fusion:
[0216] The final mask of the instance is generated by the linear combination of the prototype mask and the coefficient:
[0217] Mask i = σ(M i · P T )
[0218] where M i is the coefficient vector corresponding to the i-th detection box (through RoI alignment or coordinate indexing), σ is the Sigmoid function, and P T is the spatially expanded form of the prototype. The dimension of the final mask is H p × W p , aligned with the resolution of the input image.
[0219] In the formula, npr: the number of prototypes, controlling the representational ability of mask generation; nm: the number of instance masks, usually consistent with the maximum number of detected instances; reg max : the discretization granularity of bounding box regression; ch[0]: the number of channels of the input feature map, determining the complexity of prototype generation; strides: the feature map downsampling rate, used to decode the box coordinates to the original image scale.
[0220] This module realizes efficient instance segmentation by decoupling the detection and segmentation branches and using prototype dynamic weighting.
[0221] In step S3 of this embodiment, the process of training the constructed lightweight YOLOv11-seg_multi-object network is as follows:
[0222] The preprocessed data is input into the improved lightweight YOLOv11 network. Its backbone network uses the EIEStem module for efficient initial feature extraction. The C3k2_DynamicConv dynamic convolution module and the C3k2_GhostDynamicConv composite lightweight module are used to achieve multi-scale feature abstraction, and the LAWDS lightweight adaptive weighted downsampling module is combined to retain high-frequency detail information. In the feature fusion stage, the C2CGA channel-spatial global attention module dynamically allocates channel and spatial dimension weights, enhances the cross-level feature fusion ability, and suppresses the interference of complex campus backgrounds. The head network uses the Segment_Efficient lightweight segmentation module to generate pixel-level mask predictions through depthwise separable convolution and dynamic upsampling, and outputs the class probability, confidence, and contour parameters of the target instance. During the training process, a multi-task loss function jointly optimized by the dynamic focal loss (Dynamic FocalLoss) and the boundary-aware mask loss (Boundary-Aware Mask Loss) is used, which is expressed as:
[0223] L total = λ cls L cls + λ mask L mask + λ boundray L boundary
[0224] In the formula, L cls represents the classification loss, the dynamic focal loss, which solves the class imbalance problem. λ cls represents the weight coefficient of the classification loss (such as 1.0), which controls the importance of the classification task; L msak represents the mask loss, which measures the overlap between the predicted mask and the true mask. λ msak represents the weight coefficient of the mask loss, which balances the optimization intensity of the mask accuracy; L boundary represents the boundary loss, which penalizes the contour prediction deviation. λ boundary represents the weight coefficient of the boundary loss, which enhances the contour alignment ability; The network parameters are updated by the Adam optimizer and the gradient descent method, gradually improving the accuracy and robustness of the model for multi-object instance segmentation in campus scenes.
[0225] During the training process of the deep learning model, if the loss function value (Loss) continues to decrease and finally converges to a stable interval, and at the same time the mean average precision (mAP) gradually increases and reaches a steady state synchronously, it indicates that the model has fully learned the feature laws in the representation space of the current dataset, and its optimization process tends to be saturated. At this time, the parameter update of the model no longer significantly affects the performance indicators, and it can be determined that the convergence state has been reached, marking the completion of the training stage.
[0226] The weight of the lightweight model YOLOv11-seg_multi-object is obtained by training according to the original dataset cropped in step S1 and the improved network structure in step S2. This weight is deployed on the campus patrol robot platform to perform multi-object instance segmentation tasks in real time, outputting pixel-level masks and category information of pedestrians, vehicles, and bicycles, and linking with the path planning system to achieve dynamic obstacle avoidance and scene semantic perception. In this embodiment, the training environment is Python-3.10.16 torch-2.6.0+cu124 CUDA:0 (NVIDIA GeForce RTX 4070 Laptop GPU, 8188MiB), the number of training epochs is set to 300, the batch_size is set to 16, and finally the trained model is obtained. After the training of YOLOv11-seg_multi-object is completed, compared with YOLOv11 under the same training environment and parameter settings, the parameters are reduced by more than 19%, the GFLOPs are reduced by more than 13%, and the average detection accuracy mAP_0.5 is increased by 0.1%.
[0227] In step S4 of this embodiment, the trained lightweight model YOLOv11-seg_multi-object is used to segment the image data of pedestrians, vehicles, and bicycles collected by the camera on the TX2 computing processor of the patrol robot, and output the target pixel-level contours and category information, such as Figure 7 is an image segmentation example under this embodiment. Figure 8 is a comparison schematic diagram of the accuracy index mAP and the computational volume index GFLOPs of multi-object instance segmentation of the lightweight model YOLOv11-seg_multi-object in the embodiment of the present invention and the existing YOLOv11 model under the same dataset. Among them, Figure 8 (a) is the training result of the original unimproved YOLOv11 model; Figure 8 (b) is the training result of the YOLOv11-seg_multi-object lightweight model of the present invention. Compared with the two, YOLOv11-seg_multi-object has a better segmentation effect after training. In addition, the lightweight indicators of YOLOv11-seg_multi-object and YOLOv11 are compared as shown in Table 1 below:
[0228] Table 1
[0229]
[0230] As can be seen from Table 1, YOLOv11-seg_multi-object has better lightweight performance and higher accuracy.
[0231] Embodiment 2
[0232] This embodiment presents a campus patrol robot system for implementing the aforementioned lightweight multi-object instance segmentation method, which at least includes a robot body, a housing, a four-wheel drive chassis, a support frame, a motion control component, and a multi-modal sensor assembly. The robot body is protected by a high-strength composite material housing. The four-wheel drive chassis is equipped with an independent suspension system to adapt to the complex terrain of the campus. The support frame is rigidly connected to the housing to ensure the stability of the sensor module. The core of the platform uses a NUC host as the computing processor, runs the ROS robot operating system, and has a lightweight YOLOv11-seg_multi-object instance segmentation model built in. To achieve accurate perception of the campus environment, the platform integrates an Intel RealSense D455 depth camera and a lidar, which are respectively used for high-resolution RGB-D data acquisition and three-dimensional space modeling. Through the ROS communication framework, multi-sensor data is input into the segmentation network after being optimized in real time by the C2CGA channel - spatial global attention module, generating pixel-level masks and class probabilities for pedestrians, vehicles, and bicycles.
[0233] The robot supports two modes: autonomous cruise and remote control. Control commands are transmitted to the brushless DC motor via the CAN bus, achieving centimeter-level positioning accuracy and multi-radius steering control. During the autonomous cruise process, the depth camera continuously captures campus scene images at a frame rate of 30FPS. After being processed in real time by the Jetson platform, the segmentation results are published to the navigation decision-making node through ROS topics, driving the robot to dynamically avoid obstacles (such as pedestrian clusters and illegally parked bicycles). At the same time, the contour and class information of key targets are uploaded to the large screen of the campus security center through a low-latency wireless communication module, supporting real-time monitoring and emergency dispatching by management personnel.
[0234] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not restrictive. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A lightweight multi-object instance segmentation method for inspection robots, characterized in that: The method includes the following steps: Obtain a public dataset, perform data cropping and data augmentation on the public dataset, and construct a campus scene dataset for training; Construct an improved lightweight YOLOv11 network, which at least adaptively adjusts weights according to input features through a dynamic convolution module, and synergistically optimizes feature fusion through cascaded group attention and channel-spatial attention; Train the improved lightweight YOLOv11 network according to the constructed campus scene dataset to obtain its optimal weights; Deploy the trained lightweight instance segmentation model to an inspection robot to process campus scene image data in real time and output target pixel-level contours and class information.
2. The lightweight multi-object instance segmentation method for inspection robots according to claim 1, characterized in that: The improved lightweight YOLOv11 at least includes a backbone network and a head network. Among them, the backbone network at least includes several EIEStem lightweight feature extraction modules, C3k2_DynamicConv dynamic convolution modules, C3k2_GhostDynamicConv modules, and LAWDS composite lightweight modules, as well as at least one SPPF module and one C2CGA module; in the backbone network, multi-scale feature extraction is performed through a combination of several groups of EIEStem modules, LAWDS modules, C3k2_DynamicConv modules, C3k2_GhostDynamicConv modules, and C2CGA modules. Among them, the EIEStem module reduces the computational complexity through grouped convolution and channel recombination; the C3k2_DynamicConv module adaptively adjusts weights according to input features to reduce the number of parameters; the C3k2_GhostDynamicConv combines the cheap linear transformation of Ghost convolution and dynamic convolution to compress parameters while maintaining the feature expression ability; the LAWDS module retains high-frequency detail information through adaptive weighted pooling; the extracted highest-scale features are input into the SPPF module and the C2CGA module for multi-scale feature fusion; The head network at least includes several Upsample upsampling modules, Concat splicing modules, C3k2_DynamicConv modules, C3k2_GhostDynamicConv modules, LAWDS modules, and Segment_Efficient modules; in the head network, first perform upsampling through the Upsample module, then splice feature maps of different resolutions through the Concat module, and then further process the features using C3k2_DynamicConv, C3k2_GhostDynamicConv, and LAWDS. Finally, the Segment_Efficient layer segments targets of different scales.
3. The lightweight multi-object instance segmentation method for inspection robots according to claim 2, wherein: The C3k2_DynamicConv dynamic convolution module in the improved lightweight YOLOv11 network combines dynamic convolution kernels with the C3 structure for dynamic feature extraction and parameter optimization. Among them, for the input feature map X ∈ R H×W×C , its processing process is as follows: First, extract the statistical characteristics of the channel dimension of the feature map , namely the mean μ and variance σ 2 ; Concatenate the mean and variance into a vector Predict the dynamic weight matrix through a small fully connected network f FC K represents the convolutional kernel size, R << K If R is the downsampling factor, the dynamic weight matrix is expressed as: 2 If R is the downsampling factor, the dynamic weight matrix is expressed as: W dyn = f FC (S)= W2·ReLU(W1S + b1)+ b2 Where W1 and W2 represent the weights of the first and second layers respectively, b1 and b2 represent the biases of the first and second layers respectively, and ReLU(·) represents a non-linear activation function; Then, the base weight is combined with the dynamic weight W dyn to construct the final convolution kernel: W final = W base ·W dyn Finally, place the obtained final convolution kernel in the second convolutional layer of the C3 module of the YOLOv11 network for multi-branch feature fusion: Y dynamic = Conv3(DynamicConv2(Conv1(X)) + X) Wherein, Conv1(·) is the first convolutional operation of the C3 module, DynamicConv2(·) is the second dynamic convolutional operation of the C3 module, and the dynamic convolutional kernel adaptively adjusts the weights according to the input features. Conv3(·) is the third convolutional operation of the C3 module.
4. A lightweight multi-object instance segmentation method for inspection robots according to claim 2, characterized in that: Based on the lightweight convolution idea, C3k2_GhostDynamicConv in the improved lightweight YOLOv11 network reduces the computational complexity by decomposing the standard convolution into multi-stage operations. Its core component, GhostModule, generates features in two steps: first, it generates partial features using ordinary convolution, and then supplements the detailed features through depthwise convolution. The computational process of C3k2_GhostDynamicConv is expressed as: Input feature map First, it passes through the basic convolutional layer to obtain intermediate features Among them, the number of hidden channels C = C2×e, where C2 is the number of output channels and e∈(0,1) is the expansion factor; subsequently, feature transformation is performed through n cascaded GhostModules. The calculation of each module is divided into a main path and a secondary path, and then the output results of the main path and the secondary path are concatenated. The process is expressed as: Wherein, m = C / s, s is the segmentation ratio, k is the convolutional kernel size, Y1 represents the output of the main path, Y2 represents the output of the secondary path, and Y represents the concatenation result of the outputs of the main path and the secondary path. In C3k2_GhostDynamicConv, the module type is selected according to the c3k flag: when c3k = True, the C3k structure containing two GhostModules is adopted, and n is used to control the stacking times; otherwise, a single GhostModule is directly used. The parameter g controls the number of groups of grouped convolution, and shortcut determines whether to add a residual connection.
5. The lightweight multi-object instance segmentation method for inspection robots according to claim 2, wherein: The EIEStem module in the improved lightweight YOLOv11 network achieves efficient feature encoding through multi-branch feature extraction and fusion. Its processing process is divided into four stages. Given the input feature map as There is: Primary convolution, spatial compression is performed by the first 3×3 convolution: input Features are extracted by Conv1: Among them, W1 is hidc 3×3 convolutional kernels. Dual-branch processing, performing edge enhancement and feature retention in parallel: The Sobel branch applies the directional gradient operator. Y edge = Soble(Y1; G x , G y ) Among them, G x = [[-1, 0, 1], [-2, 0, 2], [-1, 0, 1]] and respectively represent the horizontal / vertical Sobel operators; The pooling branch adopts special max pooling. Y pool = MaxPool(ZeroPad(Y1)) Maintain the feature resolution through (0,1,0,1) padding and the pooling window. Feature concatenation and fusion, followed by secondary downsampling along the channel dimension. Y2 = Conv(Concat(Y edge , Y pool )); W2 Among them, W2 is hidc 3×3 convolutional kernels, and the feature map is compressed to 1 / 4 of the original size. Channel compression, using 1×1 convolution to adjust the final dimension. Among them, W3 is ouc 1×1 convolutional kernels, realizing the channel compression from C h →C out ; C in , C h , C out correspond to the input / hidden / output channel numbers respectively, and H / W is the input spatial dimension.
6. A lightweight multi-object instance segmentation method for inspection robots according to claim 2, characterized in that: The lightweight adaptive weight downsampling module LAWDS in the improved lightweight YOLOv11 network achieves adaptive spatial compression by dynamically allocating regional weights. Its principle is divided into two stages: attention weight generation and downsampling fusion. Given that the input feature map dimension is where B is the batch size and C is the number of channels, we have: Attention weight generation. where AvgPool(·) is the average pooling operation; φ is the candidate weight corresponding to the spatial position generated by a 1×1 convolution, and the parameters of the 1×1 convolution Softmax normalizes along the last dimension; each element A b,c,i,j,k represents the attention weight of the k-th region at the position (i, j) in the b-th batch and c-th channel Adaptive downsampling. Generate downsampled features with quadruple channels through grouped convolution, and the number of groups G = C / group. The downsampled features are reorganized as: The final output is obtained through weighted fusion. Where ⊙ represents element-wise multiplication, and the feature at each position is the weighted sum of the features of its four candidate regions and the corresponding attention weights.
7. A lightweight multi-object instance segmentation method for inspection robots according to claim 2, characterized in that: The C2CGA module in the improved lightweight YOLOv11 network performs collaborative optimization of cascaded group attention and channel-spatial attention. Among them, for the input feature map is divided into G subgroups according to the channel dimension. Each subgroup independently calculates channel and spatial attention, and the attention weights output by the previous subgroup are used as prior information for the next subgroup. The feature focus area is gradually refined through cascaded transmission. The process is expressed as: Grouping: Divide X into {X1, X2, …, X G}, where X g ∈R H×W×C / G ; Cascaded processing: The input of the g-th group is X g and the attention weights A output by the previous group g-1 , gradually enhancing the global perception ability through iterative optimization; Channel-spatial attention collaboration: For each group of features X g perform global average pooling to generate a channel statistical vector Z g ∈R C / G and generate channel weights through a lightweight fully connected network Among them r is the compression ratio, and σ(·) represents the activation function, usually the Sigmoid function, which restricts the weights to [0, 1]. Introduce deformable convolution to predict spatial offsets K is the convolution kernel size, and the sampling positions are dynamically adjusted to focus on the key regions of the target: where p k is the position of the conventional convolution kernel, and ω g,k is the learnable weight; Then perform cross - fusion and weight transfer: Multiply the channel weights by the spatial weights element - by - element to generate the subgroup attention map A g , and pass it to the next group through concatenation: X g+1 = X g ⊙A g Finally, all subgroup outputs are concatenated into the global feature map X out ∈R H×W×C .
8. A lightweight multi-object instance segmentation method for inspection robots according to claim 2, characterized in that: In the head network of the improved lightweight YOLOv11 network, the high feature map and the low feature map are aligned and fused. The size consistency is adjusted by upsampling or downsampling. Suppose the low-level feature map is adjusted to the same spatial dimension as the high-level feature map by upsampling: X low_upsamp = Upsample(X low , H', W') Then, the two are concatenated and fused and used as the input for subsequent feature extraction. X fusion = Concat(X low_upsamp , X high ) The processing methods of the C3k2_DynamicConv module, C3k2_GhostDynamicConv, and LAWDS composite module in the head network are the same as those in the backbone network. The Segment_Efficient module includes a detection branch and a segmentation branch. The detection branch predicts the bounding box box and the class cls through multiple layers of convolution. The segmentation branch generates a mask through the prototype mask Proto and the dynamic coefficient. The specific process is as follows: Detection branch: The input feature map is extracted by the stem module, and then the regression parameters and class probabilities are output through cv2 and cv3 respectively. The regression parameters are decoded into bounding box coordinates through DFL. dbox = DFL(box) × anchors + strides where box is the regression parameter, with a shape of B×4×re max ×H×W; anchors are the anchor coordinates, strides are the feature map strides, and reg_max represents the number of discretization intervals; Segmentation branch: Prototype generation: The Proto module converts the input feature map into npr prototype masks: P = Conv(ch[0] → npr) where ch[0] is the number of channels of the input feature map, npr is the number of prototypes, and the dimension of P is B×npr×H p ×W p , representing the basic mask template; Coefficient generation: nm dynamic coefficients are generated from the feature maps of each layer through the cv4 module: M = Conv(x → c4 → nm) where nm is the number of instance masks, is the number of intermediate channels, and the dimension of M is B×nm×H×W, representing the mask combination weights corresponding to each position; Mask fusion: The final mask of the instance is generated by the linear combination of the prototype mask and the coefficient: Mask i = σ(M i ·P T ) Among them, M i is the coefficient vector corresponding to the i-th detection box, σ is the Sigmoid function, and P T is the spatial expansion form of the prototype, and the final mask dimension is H p ×W p , which is aligned with the input image resolution.
9. The lightweight multi-target instance segmentation method for inspection robots according to claim 1, characterized in that: The process of training the improved lightweight YOLOv11 network is as follows: Input the preprocessed data into the improved lightweight YOLOv11 network; Its backbone network uses the EIEStem module for efficient initial feature extraction, realizes multi-scale feature abstraction through the C3k2_DynamicConv dynamic convolution module and the C3k2_GhostDynamicConv composite lightweight module, and combines the LAWDS lightweight adaptive weighted downsampling module to retain high-frequency detail information; in the feature fusion stage, the C2CGA channel-spatial global attention module is used to dynamically allocate channel and spatial dimension weights; The head network uses the Segment_Efficient lightweight segmentation module, generates pixel-level mask predictions through depthwise separable convolution and dynamic upsampling, and outputs the class probability, confidence, and contour parameters of the target instance; During the training process, a multi-task loss function jointly optimized by the dynamic focal loss and the boundary-aware mask loss is used, which is expressed as: L total = λ cls L cls + λ mask L mask + λ boundray L boundary where L cls represents the classification loss, the dynamic focal loss, which solves the problem of class imbalance, and λ cls represents the weight coefficient of the classification loss, controlling the importance of the classification task; L msak represents the mask loss, measuring the overlap between the predicted mask and the ground truth mask, and λ msak represents the weight coefficient of the mask loss, balancing the optimization intensity of the mask accuracy; L boundary represents the boundary loss, penalizes the contour prediction deviation, λ boundary represents the weight coefficient of the boundary loss, enhances the contour alignment ability; updates the network parameters through the Adam optimizer and the gradient descent method, and gradually improves the accuracy and robustness of the model for multi-object instance segmentation in campus scenes.
10. A lightweight multi-object instance segmentation system for inspection robots, characterized in that: The system at least includes a robot body, a housing, a four-wheel drive chassis, a support frame, a mobile control component, and a multi-modal sensor component. Among them, the housing is installed outside the robot body, the four-wheel drive chassis is set at the bottom, the four-wheel drive chassis is equipped with an independent suspension system, and the support frame is rigidly connected to the housing; The platform core uses a NUC host as the operation processor, runs the ROS robot operating system, and deploys the improved lightweight YOLOv11 network constructed in the lightweight multi-object instance segmentation method for inspection robots described in any one of the preceding claims 1-9; The platform integrates an Intel RealSense D455 depth camera and a lidar, which are respectively used for high-resolution RGB-D data acquisition and three-dimensional space modeling; Through the ROS communication framework, the multi-sensor data is input into the segmentation network after being optimized in real time by the C2CGA channel-spatial global attention module, and pixel-level masks and class probabilities of pedestrians, vehicles, and bicycles are generated; The robot supports dual modes of autonomous cruise and remote control. The control instructions are transmitted to the brushless DC motor via the CAN bus for centimeter-level positioning accuracy and multi-radius steering control; during the autonomous cruise process, the depth camera continuously collects campus scene images, which are processed in real time by the Jetson platform. The segmentation results are published to the navigation decision-making node through the ROS topic to drive the robot to dynamically avoid obstacles; at the same time, the contour and class information of the key targets are uploaded to the campus security center large screen through the low-latency wireless communication module to support real-time monitoring and emergency dispatch by the management personnel.
Citation Information
Cited By
Acoustic emission source positioning method based on lightweight convolution and attention mechanism
CN120801528A
Model training method, firework detection method, device, equipment, medium and product
CN120953742A
Target detection online learning dynamic sample selection method and system, computer equipment and storage medium
CN120997492A
Re-parameterization unmanned aerial vehicle target detection method based on multi-core fusion and omnidirectional connection
CN121121561A
Lightweight lithium mineral microscopic image real-time detection and instance segmentation method
CN121437530A