A single-object localization method based on a lightweight improved Unet semantic segmentation network

By making lightweight improvements on the Unet semantic segmentation network, combining the Ghost module and Focal Loss loss function, the problem of unstable object positioning in complex lighting environments is solved, and efficient target recognition and positioning under variable lighting conditions is achieved.

CN117237620BActive Publication Date: 2025-06-27HOHAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310559841.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-18
Publication Date
2025-06-27
Estimated Expiration
2043-05-18

AI Technical Summary

Technical Problem

In a complex and changeable lighting environment, the real-time positioning and tracking of objects in industrial production lines is unstable, affecting the grasping and docking stability of the robotic arm.

Method used

A single-objective positioning method based on lightweight improvement of Unet semantic segmentation network is adopted. An industrial camera collects images in a variable lighting environment, performs semantic annotation and training, and uses the Ghost module, Ghost bottleneck and SE attention mechanism, combined with Focal Loss loss function to achieve real-time recognition and positioning of the target.

Benefits of technology

In complex lighting environments, the impact of target recognition is significantly reduced, the recognition accuracy and prediction speed are improved, the target is quickly and accurately positioned, and subsequent real-time tracking and motion control are supported.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117237620B_ABST
    Figure CN117237620B_ABST
Patent Text Reader

Abstract

The present invention discloses a single-object localization method based on a lightweight improved Unet semantic segmentation network, including: performing semantic annotation on the target pixel region in the collected image; building a lightweight improved Unet semantic segmentation network; using the improved Unet semantic segmentation network to train the collected image; using the training results of the images with better pose positions in the training set to extract templates; using the trained weight file to process the input image in real time through the Unet semantic segmentation network; rotating and scaling the template image until it traverses and matches the target in the predicted result image, so as to realize the localization of the target position according to the matching parameters. The present invention introduces the Ghost module, Ghost bottleneck and SE attention mechanism to perform lightweight improvement on it, so as to maintain the original recognition accuracy of Unet and increase the prediction speed by 18%. The present invention can quickly and accurately locate and recognize the pose and position of the target, laying a technical foundation for subsequent motion control of real-time tracking of the target.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention discloses a single-object localization method based on a lightweight improved Unet semantic segmentation network, which relates to the fields of machine vision and industrial production lines. Background Art

[0002] With the further development of society and science and technology and the continuous improvement of industrial automation level, in the process of industrial production, operations such as grasping, docking, and assembly need to be completed. Such production tasks have different requirements, and the positions of the recognition targets and environmental factors are usually variable. Currently, in industrial automated production lines, for problems such as low light, the common method is to install light sources with higher light intensity at fixed positions to provide a stable lighting working environment. In addition to the active light sources of the vision detection system in the production line, there are also interference light sources such as sunlight, workshop lights, and lights from nearby workstations. The robotic arm itself or other parts may also block some light, making the lighting environment complex and variable. Although the active light source with strong light can avoid the phenomena of low light and uneven illumination to a certain extent, it may also cause the phenomenon of strong specular reflection, making it difficult to identify the edge contour of the target, thus affecting the stability and robustness of the robotic arm's actions such as grasping. The same is true in fields such as medical and aviation, where it is difficult to provide a stable and easily recognizable environment for the target. Summary of the Invention

[0003] The purpose of the present invention is to address the problem of unstable real-time positioning and tracking of objects under complex lighting conditions, and to provide a single-object localization method based on a lightweight improved Unet semantic segmentation network, which can significantly reduce the influence of the lighting environment on target recognition during processes such as grasping, docking, and following in industrial production lines.

[0004] The present utility model is realized through the following technical solutions:

[0005] A single-object localization method based on a lightweight improved Unet semantic segmentation network includes the following steps:

[0006] Step 1: Use an industrial camera fixed at the end of the robotic arm to collect a large number of images of a single target (metal part) in a complex and variable lighting environment. After the collection is completed, perform semantic annotation on the pixel area of the target in the image to obtain a json file with image semantic information, and then convert the json file into a visual eight-bit color image as the dataset.

[0007] Step 2: Build a lightweight improved Unet semantic segmentation network. The lightweight improved Unet semantic segmentation network includes a backbone feature extraction network (encoder), an enhanced feature extraction network (decoder), and a prediction network. The modules that make up each of the above networks are 3×3 ordinary convolution, Ghost module, Ghost bottleneck, upsampling, feature layer splicing, SE attention mechanism module, and 1×1 ordinary convolution. In the backbone feature extraction network, the Ghost module replaces the 3×3 convolution in the traditional Unet, and the Ghost bottleneck with a stride of 2 replaces the max pooling layer in the traditional Unet. In the enhanced feature extraction network, a module is composed of one layer of transposed convolution and feature splicing, an SE attention mechanism module, and two 3×3 convolutional layers, and four modules are stacked in sequence.

[0008] Step 3: Use the Unet semantic segmentation network in Step 2 to train the dataset described in Step 1 with a loss function improved based on Focal Loss. Finally, after the training is completed, obtain the.h5 weight file corresponding to each Epoch, and finally select the weight file with a smaller loss.

[0009] Step 4: When the z-axis of the camera in the training set coincides with the z-axis of the spatial pose of the target object, use the training result of the pose position image to extract the template; obtain the target pixel area in the training result through edge feature extraction, and segment the target pixel area as the template. The template is used for subsequent template matching and positioning of the target during image prediction.

[0010] Step 5: Use the trained weight file in Step 3 to process the image input by the camera in real time through the Unet semantic segmentation network, and predict the semantic category of all pixel points on the input image. By setting the color corresponding to each category, the semantic segmentation network displays the pixel areas of each semantics as different color areas, and finally obtains the semantic segmentation prediction result of the real-time image.

[0011] Step 6: Rotate and scale the template in Step 4 until it traverses and matches the target in the prediction result image obtained in Step 5. The matching degree is expressed as the average value of the cosine values of the included angles between all boundary points of the template and the direction vectors of the boundary points of the area to be matched in the prediction result image of Step 5. The value range of the matching degree is (0-1). When the matching degree reaches the standard of 0.75, it is regarded as a successful match, so as to realize the positioning of the target position according to the matching parameters.

[0012] The prediction result image is to predict the image area of the recognition target in the real-time image and display it with a specific color.

[0013] After the above steps, the present invention can identify targets in environments such as bright, dark, and uneven illumination. The improved Unet semantic segmentation network can quickly predict the pixel regions of targets in complex and variable illumination environments in real time and display them in a set color, while the template matching method is used to determine the position of the target in the image on this basis.

[0014] Further, in the backbone feature extraction network in step 2, the Ghost module replaces the 3×3 convolution in the traditional Unet, and the Ghost bottleneck with a stride of 2 replaces the max pooling layer in the traditional Unet. The Ghost module first reduces the channels of the input feature map using a 1×1 convolution to obtain condensed features, and then for each channel of the reduced feature layer, a depthwise separable convolution is used to obtain a feature map similar to the condensed features. Finally, the condensed features obtained by the 1×1 convolution and the feature maps obtained by the depthwise separable convolution are stacked to obtain the output feature layer; the Ghost bottleneck with a stride of 2 is used to compress the width and height of the feature layer. After using a Ghost module for feature extraction in the backbone part, a depthwise separable convolution with a stride of 2 is used to compress the height and width of the input feature layer, and then another Ghost module is used to complete the feature extraction; a layer-by-layer convolution is also performed at the residual edge to compress the height and width of the feature layer, and finally added to the feature extraction result; each feature extraction module consists of a Ghost module and a Ghost bottleneck. There are four feature extraction modules in the backbone feature extraction network and they are stacked in sequence, resulting in a total of five preliminary effective feature layers.

[0015] Further, in the enhanced feature extraction network in step 2, it upsamples the feature layer obtained by the backbone feature extraction network to restore it to the original image size, and finally obtains the mask image of the segmentation result; a module consists of a layer of transposed convolution and feature splicing, an SE attention mechanism module, and two 3×3 convolutional layers. After being processed by four upsampling feature fusion modules, a 1×1 convolution is connected to perform dimensionality reduction to obtain the required number of channels.

[0016] Further, in step 3, the loss function improved based on Focal Loss uses Focal Loss to replace the cross-entropy loss to alleviate the problem of imbalance between positive and negative samples, and introduces a weight factor α∈[0,1]. The weight factor for positive samples is α, and the weight factor for negative samples is 1-α. The loss function is:

[0017]

[0018] A tempering factor is further introduced to focus on difficult-to-separate samples:

[0019]

[0020] In formula (4), γ is a hyperparameter. When p ic tends to 1, it means that the sample is easy to distinguish. When p ic tends to 0, it means that the sample is difficult to distinguish. (1 - p ic ) γ will cause the loss value of the easy-to-distinguish sample to increase by a smaller margin.

[0021] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0022] The present invention adopts a semantic segmentation network structure based on Unet, introduces a Ghost module, a Ghost bottleneck, and an SE attention mechanism to perform lightweight improvement on it, and realizes maintaining the original recognition accuracy of Unet and increasing the prediction speed by 18%.

[0023] The present invention uses a template matching method to quickly and accurately locate the posture and position of the recognition target, laying a technical foundation for subsequent motion control of real-time tracking of the target. Description of the Drawings

[0024] Figure 1 is a lightweight Ghost model structure.

[0025] Figure 2 is a lightweight Ghost bottleneck structure. Detailed Embodiments

[0026] The following further describes the present utility model in detail in conjunction with the drawings and specific embodiments:

[0027] The present invention provides a single-target positioning method based on a lightweight improved Unet semantic segmentation network, including the following steps:

[0028] Step 1: Collect images of a metal target. A total of 2000 images are taken under various environments and angles. After the collection is completed, manually annotate the semantic regions of the target pixels in the image in the labelme software to obtain a json file with image semantic information, and then convert the json file into a visual eight-bit color image.

[0029] Step 2: Build a lightweight improved Unet semantic segmentation network. The lightweight improved Unet semantic segmentation network includes a backbone feature extraction network (encoder), an enhanced feature extraction network (decoder), and a prediction network. In the backbone feature extraction network, the 3×3 convolution in the traditional Unet is replaced by a Ghost module, and the Ghost bottleneck with a stride of 2 replaces the maximum pooling layer in the traditional Unet, so as to reduce the computational amount of the network and improve the solving speed of the network.

[0030] AsFigure 1 The described Ghost module first reduces the channels of the input feature map using a 1×1 convolution to obtain concentrated features. Then, for each channel of the reduced feature layer, a depthwise separable convolution is used to obtain a feature map similar to the concentrated features. Finally, the concentrated features obtained by the 1×1 convolution and the feature map obtained by the depthwise separable convolution are stacked to obtain the output feature layer. As Figure 2 shown, the Ghost bottleneck with a stride of 2 is used to compress the width and height of the feature layer. After using a Ghost module for feature extraction in the main part, a depthwise separable convolution with a stride of 2 is used to compress the height and width of the input feature layer. Then, another Ghost module is used to complete the feature extraction. A layer-by-layer convolution is also performed on the residual side to compress the height and width of the feature layer, and finally, it is added to the feature extraction result.

[0031] Each feature extraction module consists of a Ghost module and a Ghost bottleneck. There are four such feature extraction modules in the main feature extraction network and they are stacked in sequence, resulting in a total of five preliminary effective feature layers. In the enhanced feature extraction network, a module is composed of a layer of deconvolution and feature concatenation, an SE attention mechanism module, and two 3×3 convolutional layers. After each deconvolution, it is concatenated with one of the preliminary effective feature layers corresponding in size in the main feature extraction network. After being processed by four upsampling feature fusion modules, a 1×1 convolution is used for dimensionality reduction to obtain the required number of channels. In the prediction network part, the last effective feature layer in the enhanced feature extraction network is used to classify each feature point, which is equivalent to classifying each pixel point.

[0032] Step 3: Use the improved Unet semantic segmentation network to train the dataset described in Step 1. The dataset is divided into a training set and a validation set. The training set is used as the input of the improved Unet semantic segmentation network, and a loss function improved based on FocalLoss is used. Finally, after the training is completed, the.h5 weight file corresponding to each Epoch is obtained, and the weight file with a smaller loss is finally selected. Then, the validation set is predicted through the improved Unet semantic segmentation network, and the prediction accuracy of the network is judged by the MIou value.

[0033] For the described loss function, the traditional Unet network uses Cross Entropy Loss (cross-entropy loss). For the case of binary classification, its formula is as follows:

[0034]

[0035] In Equation (1), y i represents the label of sample i, where the positive class is 1 and the negative class is 0, and pi It represents the probability that sample i is predicted as the positive class. For the case of multi-classification, its formula is:

[0036]

[0037] In formula (2), N represents the number of samples, M represents the number of classes excluding the background, and y ic takes 1 when the true label of sample i is equal to c, otherwise takes 0, and p i represents the probability that the predicted sample i belongs to class c.

[0038] The present invention uses Focal Loss instead of Cross Entropy Loss. First, a weight factor α∈[0,1] is introduced. The weight factor for positive samples is α, and the weight factor for negative samples is 1-α. Then the loss function can be written as:

[0039]

[0040] On the basis of formula (3), a modulating factor is further introduced to focus on difficult-to-separate samples:

[0041]

[0042] In formula (4), γ is a hyperparameter. When p ic tends to 1, it means that the sample is easy to distinguish. When p ic tends to 0, it means that the sample is difficult to distinguish. Then (1 - p ic ) γ will make the increase amplitude of the loss value of easy-to-distinguish samples decrease, and the increase amplitude of the loss value of difficult-to-distinguish samples increase, thereby prompting the network to pay more attention to difficult-to-distinguish samples.

[0043] For the MIou value evaluation index, a confusion matrix is introduced to show the accuracy of semantic segmentation, and the predicted class and the true class of each pixel in the validation set image are compared. There are a total of k classes excluding the background. For a certain class, this class i is a positive example (Positive), and the remaining classes j are negative examples (Negative). The four basic elements are: P ii represents TP (True Positive, true positive example), and the predicted value of i is consistent with the true value of i; P ji represents FP (False Positive, false positive example), and the predicted value of j is consistent with the true value of i; P ij represents FN (False Negative, false negative example), and the predicted value of i is consistent with the true value of j; P jjIt represents TN (True Negative, true negative example), and the predicted value of j is consistent with the true value of j. The average of the ratios of the intersection to the union of the predicted results and the true values for each class in the model is calculated, and the resulting average Intersection over Union (MIou) can be obtained:

[0044]

[0045] The MIou value of the improved Unet network in the present invention is 88.25, and the real-time image prediction frame rate is 15.97. While the MIou value of the traditional Unet network is 87.81, and the real-time image prediction frame rate is 13.44.

[0046] Step 4: Use the training results of images with good pose positions in the training set to extract templates. The target pixel region in the training results is obtained through edge feature extraction, and the target pixel region is segmented out as the template;

[0047] Step 5: Use the trained weight file to process the input image in real time through the Unet semantic segmentation network, predict the semantic classes of all pixel points on the input image, and display the pixel point regions of different classes in different color regions;

[0048] Step 6: Rotate and scale the template image until it traverses and matches the target in the prediction result image. The template scaling step size is 0.5 pixel units, the template scaling size is 0.5 - 2 times, the template rotation step size is 0.3°, and the template rotation angle is -90° to 90°. The matching degree is expressed as the average value of the cosine values of the angles between the direction vectors of all boundary points of the template and the boundary points of the image region to be matched. There are n edge points on the template, corresponding to the ROI region to be matched on the image. Calculate the cosine values of the angles between the direction vectors of the boundary points of the template and the boundary points of the image region to be matched, which are a1, a2, a3... a n , and the matching value is expressed as:

[0049]

[0050] The closer the value of P is to 1, the higher its matching degree. When the matching degree reaches the standard of 0.6, it is regarded as a successful matching position, thereby realizing the positioning of the target position and pose according to the matching parameters (the horizontal and vertical coordinates of the template center in the image, the horizontal and vertical scaling ratios of the template, and the template rotation angle).

[0051] In the traditional assembly line, there is a high dependence on a stable light source, and under different lighting conditions, for the traditional image processing and target recognition algorithms, workers need to manually adjust parameters such as the camera exposure. The present invention can overcome the influence of lighting on the recognition effect, and can identify the target position in a variety of lighting environments without adjusting the camera parameters, thereby reducing the dependence on a stable and strong light source in the production line and improving the operation efficiency of the assembly line.

[0052] The parts not described in the present invention are the same as or implemented by using the prior art.

Claims

1. A single-object localization method based on a lightweight improved Unet semantic segmentation network, characterized in that It includes the following steps: Step 1: Conduct a large number of image acquisitions on the target. After the acquisition is completed, perform semantic annotation on the pixel area of the target in the image to obtain a json file with image semantic information, and then convert the json file into a visual eight-bit color image. Pixel points with the same semantics in the eight-bit color image have the same color and are distinguished from pixels with other semantics; Step 2: Build a lightweight improved Unet semantic segmentation network, including a backbone feature extraction network, an enhanced feature extraction network, and a prediction network; Backbone feature extraction network: The Ghost module replaces the 3×3 convolution in the traditional Unet, and the Ghost bottleneck with a stride of 2 replaces the max pooling layer in the traditional Unet. Specifically: The Ghost module first uses a 1×1 convolution on the input feature map to reduce the number of channels to obtain a concentrated feature. Then, for each channel of the reduced feature layer, a depthwise separable convolution is used to obtain a feature map similar to the concentrated feature. Finally, the concentrated feature obtained by the 1×1 convolution and the feature map obtained by the depthwise separable convolution are stacked to obtain the output feature layer; The Ghost bottleneck with a stride of 2 is used to compress the width and height of the feature layer. After using a Ghost module for feature extraction in the backbone part, a depthwise separable convolution with a stride of 2 is used to compress the height and width of the input feature layer, and then another Ghost module is used to complete the feature extraction; At the residual edge, a layer-by-layer convolution is also performed to compress the height and width of the feature layer, and finally added to the feature extraction result; Each feature extraction module consists of a Ghost module and a Ghost bottleneck. There are four feature extraction modules in the backbone feature extraction network and they are stacked in sequence, resulting in a total of five preliminary effective feature layers; Enhanced feature extraction network: Use the feature layer obtained by the backbone feature extraction network for upsampling to restore it to the original image size, and finally obtain the mask image of the segmentation result; A module is composed of a layer of transposed convolution and feature splicing, a SE attention mechanism module, and two 3×3 convolutional layers. After being processed by the upsampling feature fusion module four times, a 1×1 convolution is connected to perform dimensionality reduction processing to obtain the required number of channels; Step 3: Use the improved Unet semantic segmentation network to train the dataset in Step 1 with a loss function improved based on FocalLoss. Finally, after the training is completed, obtain the.h5 weight file corresponding to each Epoch, and finally select the weight file with a smaller loss; Step 4: Use the training results of images with better pose positions in the training set to extract templates. The target pixel area in the training results is obtained through edge feature extraction, and the target pixel area is segmented out as a template; Step 5: Use the trained weight file to process the input image in real time through the Unet semantic segmentation network, predict the semantic category of all pixel points on the input image, and display them through different color areas; Step 6: Rotate and scale the template image until it traverses and matches the target in the predicted result image. The matching degree is represented by the average value of the cosine values of the included angles between all boundary points of the template and the direction vectors of the boundary points of the image area to be matched. When the matching degree reaches the standard, it is regarded as the successful matching position, so as to realize the positioning of the target position according to the matching parameters.

2. The single-object localization method based on the lightweight improved Unet semantic segmentation network according to claim 1, wherein The loss function improved based on Focal Loss in Step 3 uses Focal Loss to replace the cross-entropy loss to alleviate the problem of imbalance between positive and negative samples, and introduces a weight factor α ∈ [0, 1]. The weight factor for positive samples is α, and the weight factor for negative samples is 1 - α. The loss function is: Furthermore, a tempering factor is introduced to focus on difficult-to-separate samples: In the above formula: N represents the number of samples, i represents the sample, and α i represents the weight factor of sample i, M represents the number of categories excluding the background, c represents the category, and y ic represents the true label of sample i, taking 1 if the true label of sample i is equal to c, and 0 otherwise; p ic represents the probability that the predicted sample i belongs to category c, γ is a hyperparameter, when p ic tends to 1, it means that the sample is easy to distinguish, when p ic tends to 0, it means that the sample is difficult to distinguish, (1 - p ic ) γ will make the increase in the loss value of easy-to-distinguish samples decrease significantly.

Citation Information

Patent Citations

  • Remote sensing image lightweight semantic segmentation method based on edge decoupling

    CN113159051A

  • Semantic segmentation-based unstructured field road scene recognition method and device

    CN114155481A