A segmentation network system
The split network system solves the positioning error problem caused by light intensity in weld detection, and realizes accurate positioning and real-time detection of welding joints, which is suitable for low-cost equipment in industrial scenarios.
Patent Information
- Application Number
- CN202110282314.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-16
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-03-16
AI Technical Summary
In the existing weld detection methods, due to the excessive ray illumination intensity, the position position error of the weld joint is large, and accurate positioning cannot be performed, and the fitting trajectory deviation is large.
The segmented network system is adopted, including network segmentation method and welding joint positioning method. Through data set preprocessing, real-time classification and segmentation model construction, loss function design, evaluation index setting, learning rate optimization and OpenVINO accelerated inference, the accurate positioning of the weld data set is achieved.
It improves the accuracy and real-timeness of solder joint positioning, is suitable for low-cost embedded equipment, meets the real-time detection needs of industrial scenarios, and has strong local image changes and light interference resistance.
Smart Images

Figure CN113159278B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network cutting, and in particular to a segmentation network system. Background Art
[0002] Currently, the widely used method for weld detection is the radiographic testing method (RT). This method uses the penetrating radiation emitted by an (X, Y) ray source to penetrate the weld and expose the film. The image of the weld is then displayed on the processed radiographic film. The position of the solder joint is located by calculating the inflection points of the weld, and the motion trajectory of the end of the welding device is generated by fitting the calculated solder joint positions, and finally the welding task is completed.
[0003] The above method has low cost and high speed, and this solution is widely used in engineering applications. However, in actual detection, there will be a positioning error problem. Due to the excessive light intensity of the ray, local reflection appears on the metal sheet, directly resulting in the inability to locate the solder joint position, and the deviation of the fitted trajectory is large, making accurate positioning impossible. Summary of the Invention
[0004] Based on the technical problem that the above method proposed in the background art has low cost and high speed and is widely used in engineering applications, but in actual detection, there will be a positioning error problem. Due to the excessive light intensity of the ray, local reflection appears on the metal sheet, directly resulting in the inability to locate the solder joint position, and the deviation of the fitted trajectory is large, making accurate positioning impossible, the present invention proposes a segmentation network system.
[0005] A segmentation network system proposed by the present invention includes a network segmentation method and a solder joint positioning method. The solder joint positioning method includes the following specific processes: input a sample with a dimension of 640*480*1, perform augmentation such as light intensity, random cropping, and rotation on the sample. Augmentation operations are not required during testing. Downsampling is performed to obtain a generalized feature of 80*60*256. One branch passes through a global feature extractor to obtain the depth information of 80*60*256, and the other branch passes through a layer of convolution to obtain the detailed feature of 80*60*256. The detailed feature is fused through the channel selection of the depth information, and then upsampling is performed to return the fused feature. Upsampling is performed, and through a segmentation detection module and a feature point detection module, the detection result, that is, the weld data set, is obtained. The network segmentation method includes the following steps:
[0006] Step 1: Dataset preprocessing:
[0007] Collect the weld dataset, which is collected inside the factory area. By establishing a database for weld images in different scenarios, and then through manual annotation, the specific positions of welds and solder joints with labels are screened out. The data is split into a training set and a test set according to an 8:2 ratio. The training set is augmented with changes such as light intensity, random cropping, and random rotation;
[0008] Step 2: Build a real-time classification and segmentation model, specifically:
[0009] The entire network consists of 5 parts: a downsampling module, a global feature extractor, a feature fusion module, an upsampling module, and a key point detection module;
[0010] Step 3: The loss function of the real-time classification and segmentation model;
[0011] Step 4: Design the evaluation index of the model;
[0012] Step 5: Generate Gaussian key points;
[0013] Step 6: Learning rate and optimizer;
[0014] Step 7: Accelerated inference with OpenVino;
[0015] STEP8: Test results: The model detection results are wire weld detection and solder joint detection.
[0016] Preferably, the downsampling module: This module consists of 3 convolutional layers. The first layer is a common convolutional layer, which reduces the scale of the input image. The latter two layers use depthwise separable convolutions to improve computational efficiency. The stride is 2, the convolutional kernel is 3x3, followed by a BN layer and a ReLU activation function, and the output is at a 1 / 8 scale.
[0017] Preferably, the global feature extractor: Different from the traditional 2-branch structure, this network uses the same downsampling output as the input of the 2-branch structure. The deep semantic feature extraction uses 3 bottleneck residual blocks proposed by MobileNet-v2 to construct the global feature extractor. Among them, the depthwise separable convolution in the bottleneck residual block is beneficial to reducing the number of parameters and computational amount of the global feature extractor. The global feature extractor also includes a pyramid pooling module, namely PPM, for extracting context features at different scales.
[0018] Preferably, the feature fusion module: The feature fusion module is used to fuse the output features of deep semantic information features and shallow position information. The deep semantic information is used to screen the shallow position information feature layer through global pooling, and finally the deep semantic and shallow position information are fused;
[0019] To make the two input branches have the same size, the depth branch is upsampled, and a 1x1 convolution operation is performed during the feature addition stage to adjust the channels to be the same for superposition. The output result is non-linearly transformed using an activation function.
[0020] Preferably, the segmentation upsampling module: This module contains 2 depthwise separable convolutions and 1 ordinary convolution with a convolution kernel of 1x1 to improve network performance. The output result is passed through a softmax operation for calculating the cost function during training. During inference, argmax is used instead of softmax to obtain the segmentation result.
[0021] Preferably, the key point module: A key point branch is added for the solder joint positioning application in the industrial scenario. The branch is led out from the feature fusion result, and the features of the key points are integrated through two layers of convolution. At the same time, the output channels are adjusted. During the training process, dropout following the two convolutional layers promotes the generalization ability of the model, and finally, the sigmoid output is used to obtain the score of the key point position.
[0022] Preferably, in step 3, in mathematical optimization and decision theory, a loss function or cost function is a function that maps the value of an event or one or more variables to a real number, which intuitively represents certain "costs" associated with the event. An optimization problem attempts to minimize the loss function. An objective function can be a loss function or its negative (in a specific domain, differently called a reward function, a profit function, a utility function, a fitness function, etc.). In this case, it is to be maximized. During the training process of the detection model, the output result needs to approximate the real result. The approximation process requires using a loss function (Loss Function) as the objective function. By optimizing the objective function, the predicted value result continuously approaches the real value. The classification branch uses cross-entropy loss as the objective function, and the network predicts the output q i and the sample label p i , the upsampling segmentation decoding branch predicts the output y'. The focal loss and cross-entropy loss functions are designed as the objective functions for the segmentation and classification branches, and the two objective functions are optimized to make the model prediction result continuously approach the real label.
[0023] Overall cost function (objective function) of the model:
[0024] Loss = L cls +F seg
[0025] The loss function of the classification branch is the cross-entropy loss function, which evaluates the difference between the probability distribution obtained from the current training and the real distribution. Its formula is:
[0026]
[0027] where p i is the sample label, and q i is the predicted output.
[0028] The loss function of the segmentation branch is the focal loss. The proportion of background pixels in the segmentation is much larger than the pixel area of the target. To solve the problem of serious imbalance in the proportion of positive and negative samples in detection, the focal loss is designed as the loss function of this branch. This loss function can reduce the weight of a large number of simple negative samples in training. Its formula is
[0029] The loss function of the classification branch is the cross-entropy loss function, which evaluates the difference between the probability distribution obtained by the current training and the true distribution. Its formula is:
[0030]
[0031] where y' is the predicted value, and α and γ are hyperparameters.
[0032] Preferably, in the fourth step, for the model obtained after optimizing and training the objective function, it is necessary to evaluate the detection performance of the model. Only when the detection performance meets the specified requirements does the model have the detection ability for marking. Therefore, evaluation indicators for the model are designed for the two output branches respectively; in the key point branch, first, the predicted results and true labels of the output are statistically analyzed, and the cross-entropy is designed as the cost function of this branch:
[0033]
[0034] p (the ideal result, i.e., the correct label vector) and q (the output result of the neural network, i.e., the result vector after softmax transformation);
[0035] In the segmentation branch, the evaluation index uses mIoU (mean intersection over union). The mask pixels of each class of the segmentation result are statistically analyzed. Its formula is as follows,
[0036]
[0037] where p ij represents the number of samples with the true value of i and the predicted value of j, p ii represents the number of samples with the true value of i and the predicted value of i (true positives), p ij represents the number of samples with the true value of i and the predicted value of j (false positives), p jiIt represents the number (false negative) where the true value is j and the predicted value is i. k + 1 is the number of categories. When the mIoU is close to 1, the predicted value is closer to the true value.
[0038] Preferably, in step five, before the supervised learning of the picture and the label, the Gaussian transformation needs to be performed on the coordinates of the positioning points. Since the model performs weld segmentation and solder joint positioning synchronously, if the coordinate points are directly regressed, the spatial information will be lost, and the two values of x and y cannot converge simultaneously.
[0039]
[0040] Among them, σ takes 30, 20, 10, which decreases sequentially during the training process, gradually reducing the size of the Gaussian points and improving the training convergence speed; x, y are the pixel template coordinates; x0, y0 are the origin coordinates of the template center, defaulting to (0, 0); norm normalizes the output Gaussian heatmap to unify the center value of each solder joint, and finally obtains the Gaussian heatmap heatmap of the solder joint and the heatmap of the weld line.
[0041] In step six, the learning rate is designed with a warm-up scheme of "different values in different stages: rising -> stable -> falling". First, since the convergence direction in the initial training stage of model learning is unstable and oscillates severely, a low learning rate needs to be adopted to ensure good convergence of the model. However, if the learning rate is too low, the learning process will be very slow. Therefore, the low learning rate is gradually increased to a higher learning rate. Then, the model learns at a higher learning rate to continuously reduce the loss function. Finally, when the loss reaches the minimum stage, oscillations will occur. Therefore, a high learning rate is not suitable as it will cause the gradient of the weights to oscillate back and forth and it is difficult to reach the global minimum. So, the learning rate needs to be reduced.
[0042] Preferably, in step seven, when the model test result meets the accuracy requirement and is deployed on an industrial control computer and other devices, the model needs to perform optimization operations such as pruning and quantization to meet the real-time requirements of industrial applications. Since GPU acceleration platforms such as OpenVX, OpenCL, and CUDA require high costs, and OpenVINO-based CPU hardware acceleration is widely used in industry, with advantages such as cost savings and resource reuse, and at the same time integrating tools such as OpenCV, OpenVX, and OpenCL to improve the running performance of the convolutional network. The OpenVINO Toolkit mainly includes two core components: the Model Optimizer and the Inference Engine.
[0043] The inference engine of this tool includes synchronous and asynchronous inferences. During synchronous inference, the calls between functions are implemented sequentially, and no return is made until the operation is completed, and the thread is blocked. Asynchronous inference is divided into two steps: start and wait. To provide multi-device detection, it is necessary to add the continuation of multi-threaded asynchronous operation and the execution of parallel computing.
[0044] The beneficial effects of the present invention are as follows:
[0045] 1. For this segmentation network system, the convolutional neural network (CNN) has strong resistance to local image changes and light interference. With the acceleration support of CNN models by multiple manufacturers, the CNN model has been gradually applied in industry. The demand for applying deep learning algorithms in industrial scenarios is becoming stronger and stronger. We propose a real-time semantic segmentation solution to preprocess weld seams. Since the deep learning solution is applied in industrial scenarios, there are additional requirements for real-time detection algorithms: (1) Algorithm real-time performance. Since semantic segmentation is part of the preprocessing of the visual perception system, the results are often used as the input end of subsequent perception or fusion modules; (2) The algorithm must occupy as little memory as possible to allow deployment on low-cost embedded devices while ensuring that the detection accuracy is within the operating range.
[0046] 2. For this segmentation network system, it adopts a method of gradually increasing the low learning rate to a higher learning rate. Then, the model learns at a higher learning rate to continuously reduce the loss function. Finally, when the loss reaches the minimum stage, oscillations will occur. Therefore, a high learning rate is not suitable as it will cause the gradient of the weights to oscillate back and forth, making it difficult to reach the global minimum. So, it is necessary to reduce the learning rate to improve the stability of the model.
[0047] 3. For this segmentation network system, a key point branch is added for the solder joint positioning application in industrial scenarios. The branch is led out from the feature fusion result, and the features of the key points are integrated through two layers of convolution, while adjusting the output channels. During the training process, dropout follows immediately after the two convolutional layers to promote the generalization ability of the model. Finally, it is output through sigmoid to obtain the score of the key point position, and the position of the key point can be accurately obtained.
[0048] The parts not involved in this device are the same as those in the prior art or can be implemented using the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is a schematic structural diagram of the entire network of a segmentation network system proposed by the present invention;
[0050] Figure 2 It is a schematic structural diagram of the global feature extractor of a segmentation network system proposed by the present invention;
[0051] Figure 3Schematic diagram of the case structure of the feature fusion module of a segmentation network system proposed by the present invention;
[0052] Figure 4 Schematic diagram of the structure of OpenVino accelerated inference of a segmentation network system proposed by the present invention;
[0053] Figure 5 Flowchart of multi-device detection of a segmentation network system proposed by the present invention. Detailed implementation manners
[0054] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0055] In the description of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc. are based on the orientation or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.
[0056] Refer to Figures 1-5 , a segmentation network system includes a network segmentation method and a solder joint positioning method. The solder joint positioning method includes the following specific processes: input a sample with a dimension of 640*480*1, perform augmentation such as light intensity, random cropping, and rotation on the sample, and no augmentation operation is required during testing. Perform downsampling to obtain a generalized feature of 80*60*256. One branch passes through a global feature extractor to obtain depth information of 80*60*256, and the other branch passes through a layer of convolution to obtain detail features of 80*60*256. The detail features are fused through channel selection of the depth information, and then through upsampling, the fused features are returned. Perform upsampling, pass through a segmentation detection module and a feature point detection module to obtain the detection result, that is, the weld dataset. The network segmentation method includes the following steps:
[0057] Step 1: Dataset preprocessing:
[0058] Collect the weld dataset, which is collected inside the factory area. Establish a database for weld images in different scenarios, and then through manual annotation, screen to obtain the specific positions of the welds and solder joints with labels. Split the data according to the ratio of 8:2 to obtain a training set and a test set. The training set performs data augmentation such as light intensity, random cropping, and random rotation;
[0059] Step 2: Build a real-time classification and segmentation model, specifically:
[0060] The entire network consists of five parts: a downsampling module, a global feature extractor, a feature fusion module, an upsampling module, and a key point detection module;
[0061] Step 3: The loss function of the real-time classification and segmentation model;
[0062] Step 4: Design the evaluation metrics of the model;
[0063] Step 5: Gaussian key point generation;
[0064] Step 6: Learning rate and optimizer;
[0065] Step 7: OpenVino accelerated inference;
[0066] STEP8: Test results: The model detection results are wire bonding detection and solder joint detection.
[0067] In the present invention, the downsampling module: This module consists of 3 convolutional layers. The first layer is a common convolutional layer that reduces the scale of the input image. The latter two layers use depthwise separable convolutions to improve computational efficiency. The stride is 2, the convolutional kernel is 3x3, followed by a BN layer and a ReLU activation function, and the output is at 1 / 8 scale.
[0068] In the present invention, the global feature extractor: Different from the traditional two-branch structure, this network uses the same downsampling output as the input of the two-branch structure. The deep semantic feature extraction uses 3 bottleneck residual blocks proposed by MobileNet-v2 to construct the global feature extractor. Among them, the depthwise separable convolution in the bottleneck residual block is beneficial to reducing the number of parameters and computational amount of the global feature extractor. The global feature extractor also includes a pyramid pooling module, namely PPM, for extracting context features of different scales.
[0069] In the present invention, the feature fusion module is used to fuse the output features of deep semantic information features and shallow position information. The deep semantic information is used to screen the shallow position information feature layer through global pooling, and finally the deep semantic and shallow position information are fused;
[0070] In order to make the sizes of the two input branches consistent, the deep branch is upsampled, and at the same time, a 1x1 convolution operation is performed during the feature addition stage to adjust the channels to be the same for superposition, and the output result is non-linearly transformed using an activation function.
[0071] In the present invention, the segmentation upsampling module: This module contains 2 depthwise separable convolutions and 1 ordinary convolution with a convolution kernel of 1x1 to improve the network performance. The output result is passed through the softmax operation for calculating the cost function during training. During inference, argmax is used instead of softmax to obtain the segmentation result.
[0072] In the present invention, the key point module: A key point branch is added for the solder joint positioning application in the industrial scenario. The branch is led out from the feature fusion result, and the features of the key points are integrated through two layers of convolution while adjusting the output channels. During the training process, dropout following the two convolutional layers promotes the generalization ability of the model, and finally, the sigmoid output is used to obtain the score of the key point position.
[0073] In the third step of the present invention, in mathematical optimization and decision theory, a loss function or cost function is a function that maps the value of an event or one or more variables to a real number, which intuitively represents certain "costs" associated with the event. An optimization problem attempts to minimize the loss function. An objective function can be a loss function or its negative (in a specific domain, differently called a reward function, a profit function, a utility function, a fitness function, etc.), in which case it is to be maximized. During the training process of the detection model, the output result needs to be approximated to the real result. The approximation process requires using the loss function as the objective function. By optimizing the objective function, the predicted value result continuously approaches the real value. The classification branch uses the cross-entropy loss as the objective function, and the network predicts the output q i and the sample label p i , the upsampling segmentation decoding branch predicts the output y'. The focal loss and cross-entropy loss functions are designed as the objective functions for the segmentation and classification branches, and the two objective functions are optimized to make the model prediction result continuously approach the real label.
[0074] Overall cost function (objective function) of the model:
[0075] Loss = L cls + F seg
[0076] The loss function of the classification branch is the cross-entropy loss function, which evaluates the difference between the probability distribution obtained from the current training and the real distribution. Its formula is:
[0077]
[0078] where p i is the sample label and q i is the predicted output.
[0079] The segmentation branch loss function is the focal loss. The proportion of segmented background pixels is much larger than the pixel area of the target. To solve the problem of serious imbalance in the ratio of positive and negative samples in detection, the focal loss is designed as the loss function for this branch. This loss function can reduce the weight of a large number of simple negative samples in training, and its formula is
[0080] The classification branch loss function is the cross-entropy loss function, which evaluates the difference between the probability distribution obtained from the current training and the true distribution. Its formula is:
[0081]
[0082] where y' is the predicted value, and α and γ are hyperparameters.
[0083] In the present invention, in step four, for the model obtained after optimizing and training the objective function, it is necessary to evaluate the detection performance of the model. Only when the detection performance meets the specified requirements does the model have the detection ability for marking. Therefore, evaluation indicators for the model are designed for the two output branches respectively; in the key point branch, first, the predicted results and the true labels of the output are statistically analyzed, and the cross-entropy is designed as the cost function for this branch:
[0084]
[0085] p (the ideal result, i.e., the correct label vector) and q (the output result of the neural network, i.e., the result vector after softmax transformation);
[0086] In the segmentation branch, the evaluation index uses the mIoU (mean intersection over union). The mask pixels of each class of the segmentation result are statistically analyzed, and its formula is as follows,
[0087]
[0088] where, p ij represents the number of times the true value is i and is predicted as j, p ii represents the number of times the true value is i and the predicted value is i (true positive), p ij represents the number of times the true value is i and the predicted value is j (false positive), p ji represents the number of times the true value is j and the predicted value is i (false negative), k + 1 is the number of categories. When the mIoU is close to 1, the predicted value is closer to the true value.
[0089] In the present invention, in step five, before the supervised learning of the pictures and labels, it is necessary to perform Gaussian transformation on the positioning point coordinates. Since the model performs weld seam segmentation and solder joint positioning synchronously, if the coordinate points are directly regressed, the spatial information will be lost, and the two values of x and y cannot converge simultaneously;
[0090]
[0091] Among them, σ takes 30, 20, 10, which decreases sequentially during the training process, gradually reducing the size of the Gaussian points and improving the training convergence speed; x and y are the pixel template coordinates; x0 and y0 are the origin coordinates of the template center, defaulting to (0, 0); norm normalizes the output Gaussian heat map, unifying the center value of each solder joint, and finally obtaining the Gaussian heat map heatmap of the solder joint and the heat map of the weld line;
[0092] In step six, the learning rate is designed with a warm-up scheme of "different values in different stages: rising -> stable -> falling". First, since the convergence direction in the initial training stage of model learning is unstable and there is severe oscillation, a low learning rate needs to be adopted to ensure good convergence of the model. However, if the learning rate is too low, the learning process will be very slow. Therefore, the low learning rate is gradually increased to a higher learning rate. Then, the model learns at a higher learning rate to continuously reduce the loss function. Finally, when the loss reaches the minimum stage, there will be oscillation. Therefore, a high learning rate is not suitable as it will cause the gradient of the weights to oscillate back and forth and it is difficult to reach the global minimum. So, the learning rate needs to be reduced.
[0093] In the present invention, in step seven, when the model test result meets the accuracy requirement and is deployed on an industrial control computer and other devices, the model needs to be optimized such as pruned and quantized to meet the real-time requirements of industrial applications. Since GPU acceleration platforms such as OpenVX, OpenCL, and CUDA require high costs, while the acceleration based on CPU hardware of OpenVINO has been widely used in industry, with advantages such as cost savings and resource reuse, and at the same time integrating tools such as OpenCV, OpenVX, and OpenCL to improve the running performance of the convolutional network. The OpenVINO ToolKit mainly includes two core groups, the ModelOptimizer and the Inference Engine;
[0094] The inference engine of this tool includes synchronous and asynchronous inferences. During synchronous inference, the calls between functions are implemented sequentially and do not return until the operation is completed, and the thread is blocked; asynchronous inference is divided into two steps: start and wait, providing the need for multi-device detection to add the continuation of asynchronous running multi-threads and the execution of parallel computing.
[0095] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered within the protection scope of the present invention.
Claims
1. A segmentation network system, including network segmentation and solder joint positioning, characterized in that, The solder joint positioning includes: inputting a sample, performing operations such as light intensity adjustment, random cropping, and rotation augmentation on the sample (augmentation operations are not required during testing), downsampling to obtain generalized features. One branch passes through a global feature extractor to obtain depth information, and the other branch obtains detailed features after passing through a layer of convolution. The detailed features are fused through channel selection of the depth information, then upsampled to return the fused features. After upsampling, it passes through a segmentation detection module and a feature point detection module to obtain the detection result, which is the weld dataset. The network segmentation includes: Dataset preprocessing: Collect the weld dataset, which is collected inside the factory area. By establishing a database for weld images in different scenarios, and then through manual annotation, the specific positions labeled as welds and solder joints are screened. The data is split proportionally to obtain a training set and a test set. The training set is augmented with operations such as light intensity adjustment, random cropping, and random rotation; Building a real-time classification and segmentation model, specifically: The entire network consists of 5 parts: a downsampling module, a global feature extractor, a feature fusion module, an upsampling module, and a key point detection module; The loss function of the real-time classification and segmentation model; Designing the evaluation metrics of the model; Generating Gaussian key points; Learning rate and optimizer; Accelerating inference with OpenVino; Test results: The model test results are wire welding detection and solder joint detection; The downsampling module includes 3 convolutional layers. The first layer is a common convolutional layer for reducing the scale of the input image, and the last two layers use depthwise separable convolutions to improve computational efficiency; The global feature extractor uses the same downsampling output as the input of the 2-branch structure. The deep semantic feature extraction uses 3 bottleneck residual blocks proposed by MobileNet-v2 to construct the global feature extractor. Among them, the depthwise separable convolution in the bottleneck residual block is beneficial to reducing the number of parameters and computational amount of the global feature extractor. The global feature extractor also includes a pyramid pooling module for extracting context features of different scales; The feature fusion module is used to fuse the output features of deep semantic information features and shallow position information. The deep semantic information is used to screen the shallow position information feature layer through global pooling, and finally the deep semantic and shallow position information are fused. To make the two input branches have the same size, the depth branch is upsampled, and at the same time, a 1x1 convolution operation is performed during the feature addition stage to adjust the channels to be the same for superposition. The output result is non-linearly transformed using an activation function; The upsampling module contains 2 depthwise separable convolutions and 1 common convolution to improve the network performance. The output result passes through a softmax operation for calculating the cost function during training. During inference, argmax is used instead of softmax to obtain the segmentation result; The key-point detection module adds a key-point branch for the solder joint positioning application in the industrial scenario. The branch is derived from the feature fusion result, and the features of the key points are integrated through two layers of convolutional layers. At the same time, the output channels are adjusted. During the training process, dropout follows the two convolutional layers to promote the generalization ability of the model. Finally, the score of the key-point position is obtained through sigmoid output.
2. The segmentation network system according to claim 1, characterized in that In the loss function of the real-time classification and segmentation model, in mathematical optimization and decision theory, a loss function or cost function is a function that maps the value of an event or one or more variables to a real number, which intuitively represents some "cost" associated with that event. An optimization problem attempts to minimize the loss function. An objective function can be a loss function or its negative, in which case it is to be maximized. During the training process of the detection model, the output result needs to be approximated to the true result. The approximation process requires using the loss function as the objective function. By optimizing the objective function, the predicted value result can continuously approximate the true value. The classification branch uses cross-entropy loss as the objective function, and the network predicts the output q i and the sample label p i , the upsampling segmentation decoding branch predicts the output y'. The focal loss and cross-entropy loss functions are designed as the objective functions for the segmentation and classification branches, and the two objective functions are optimized to make the model prediction results continuously approximate the true labels; Overall cost function of the model: The classification branch loss function is the cross-entropy loss function, which evaluates the difference between the probability distribution obtained from the current training and the true distribution. Its formula is as follows: where p i is the sample label, and q i is the predicted output; The segmentation branch loss function is the focal loss. The proportion of segmented background pixels is much larger than the pixel area of the target. To solve the problem of serious imbalance in the ratio of positive and negative samples in detection, the focal loss is designed as the loss function for this branch. This loss function can reduce the weight of a large number of simple negative samples in training, and its formula is: where y' is the predicted value, and are hyperparameters.
3. The segmentation network system according to claim 1, characterized in that, Among the evaluation metrics of the designed model, for the model obtained after optimizing and training the objective function, it is necessary to evaluate the detection performance of the model. Only when the detection performance meets the specified requirements can the model have the detection ability for marking. Therefore, the evaluation metrics of the designed model are respectively designed for the two output branches. In the key point branch, first, the predicted results and the true labels of the output are statistically analyzed, and the cross-entropy is designed as the cost function of this branch: p represents the ideal result, that is, the correct label vector, and q represents the output result of the neural network, that is, the result vector after conversion; In the segmentation branch, the evaluation metric is the mean intersection over union (mIoU). , the mask pixels of each class in the segmentation result are counted, and the formula is as follows: where represents the number of pixels with the ground truth value and predicted as . represents the number of pixels with the ground truth value and the predicted value . represents the number of pixels with the ground truth value and the predicted value . represents the number of pixels with the ground truth value and the predicted value . is the number of classes. When is close to 1, the predicted value is closer to the ground truth value.
4. A segmentation network system according to claim 1, wherein, In the generation of Gaussian key points, before supervised learning of the image and label, it is necessary to perform Gaussian transformation on the coordinates of the positioning points. Since the model performs weld seam segmentation and solder joint positioning synchronously, if direct regression of coordinate points is used, spatial information will be lost, and the two values of x and y cannot converge simultaneously; Among them, Take 30, 20, 10, which decrease sequentially during the training process, gradually reducing the size of the Gaussian points to improve the training convergence speed; x, y are the coordinates of the pixel template; x0, y0 are the origin coordinates of the template center, defaulting to (0, 0); Normalize the output Gaussian heat map to unify the central value of each solder joint, and finally obtain the Gaussian heat map heatmap of the solder joint and the heat map of the weld line; In the learning rate and optimizer, the learning rate is designed with a warm-up scheme of "different values in different stages: rising -> stable -> falling". First, since the convergence direction in the initial training stage of model learning is unstable and there are severe oscillations, a low learning rate needs to be adopted to ensure good convergence of the model. However, if the learning rate is too low, the learning process will be very slow. Therefore, the learning rate is gradually increased from a low value to a high value. Then, the model learns at a higher learning rate to continuously reduce the loss function. Finally, when the loss reaches the minimum stage, oscillations will occur. Therefore, a high learning rate is not suitable as it will cause the gradient of the weights to oscillate back and forth and it is difficult to reach the global minimum. So, the learning rate needs to be reduced.
5. A segmentation network system according to claim 1, wherein In the OpenVino accelerated inference, when the model test results meet the accuracy requirements and are then deployed on industrial control computers and other devices, the model needs to be pruned and quantized to meet the real-time requirements of industrial applications. Since the GPU acceleration platforms OpenVX, OpenCL, and CUDA require high costs, while the acceleration based on CPU hardware by OpenVINO has been widely used in industry, with the advantages of cost savings and resource reuse. At the same time, it integrates OpenCV, OpenVX, and OpenCL tools to improve the running performance of the convolutional network. The OpenVino toolkit mainly includes two core groups: the model optimizer and the inference engine. The inference engine of this tool includes synchronous and asynchronous inferences. During synchronous inference, the calls between functions are implemented sequentially and do not return until the operation is completed, blocking the thread. Asynchronous inference is divided into two steps: start and wait, providing the need for multi-device detection to add the continuation of asynchronous multi-threaded operation and the execution of parallel computing.
Citation Information
Patent Citations
Welding robot welding seam recognition method based on deep learning
CN110135513A
Barrel Nut With Helical Wire Insert
US20200408241A1