A workpiece positioning method and system based on machine vision

By improving the SSD algorithm, combined with feature enhancement module, hollow convolution and improved Focal loss loss function, the machine vision system's problem of low accuracy and insufficient robustness when positioning the workpiece, and the precise positioning and high flexibility positioning capabilities of the center point of the workpiece are achieved.

CN119784763BActive Publication Date: 2025-06-10SHENYANG LI DENGWEI AUTO PARTS CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510281191.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-10
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

The existing machine vision system has problems of low accuracy and insufficient robustness when positioning workpieces, which affects the quality of subsequent grabbing, sorting and assembly operations of the robot.

Method used

A workpiece positioning method based on an improved SSD algorithm is proposed. By acquiring the workpiece images acquired by the camera, the preprocessed image is input into the trained improved SSD detection model for positioning, and the three-dimensional coordinates of the center point of the workpiece are obtained. The improved SSD algorithm improves the feature expression ability and robustness of the model in dark light scenarios by introducing feature enhancement modules, hollow convolutions and improved Focal loss loss function.

Benefits of technology

It realizes accurate positioning of the center point of the workpiece, is suitable for positioning detection of different types of workpieces, and improves the flexibility and positioning accuracy of the machine vision system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119784763B_ABST
    Figure CN119784763B_ABST
Patent Text Reader

Abstract

The present invention discloses a workpiece positioning method and system based on machine vision, which relates to the technical field of machine vision. The technical key points of the present invention include: acquiring an image containing a workpiece collected by a camera; preprocessing the image; inputting the preprocessed image into a trained detection model based on an improved SSD algorithm to position the workpiece and obtain the pixel coordinates of the center point of the workpiece; converting the pixel coordinates into the three-dimensional coordinates of the workpiece according to the calibration result; wherein, the improvements of the improved SSD algorithm include: after extracting features of six different scales by the feature extraction module in the original SSD algorithm, introducing a feature enhancement module to further enhance the features of different scales; using dilated convolution to replace the convolution in the original SSD algorithm for downsampling; and using the improved Focal loss function as the classification loss function. The present invention can more accurately locate the center point of the workpiece and is suitable for the positioning detection of different types of workpieces.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine vision, and particularly to a workpiece positioning method and system based on machine vision. Background Art

[0002] The progress of the automation industry has continuously promoted the development of the application of machine vision technology. Machine vision is one of the important components of industrial production. A machine vision system is mainly divided into system software design and system hardware design. The software part mainly refers to image processing software and image processing algorithms; the hardware part mainly consists of a camera, a light source device, a lens, an auxiliary execution device (such as a robotic arm), etc.

[0003] The classic machine vision process is to obtain image information by shooting a target object with an industrial camera under specific light source conditions, then process the image information using image processing software, and transmit corresponding control signals through a connected control device for control operations. The whole process is analogous to a human identifying a target through the eyes, transmitting the target information to the brain, and then controlling the limbs to grab the target object through the brain. In the machine vision structure, an object needs to be photographed by a camera under the illumination of a light source, the collected image is transmitted to an acquisition card, and then sent to a PC for image processing to obtain image information; corresponding control instructions are issued through a control system (such as a PLC), and then the actuator is controlled to grab the corresponding object, which is the most basic machine vision system structure. Among them, the image processing algorithm is the core of the machine vision process, and the result of image processing, that is, the accuracy of target positioning, will directly affect subsequent operations of the robot such as grasping, sorting, and assembling. Therefore, researching a workpiece positioning algorithm with high accuracy and strong robustness is of great significance for improving the flexibility of the entire machine vision system. Summary of the Invention

[0004] In view of the above problems, the present invention proposes a workpiece positioning method and system based on machine vision in an attempt to solve or alleviate one or more of the above problems.

[0005] According to one aspect of the present invention, a workpiece positioning method based on machine vision is proposed, and the method includes:

[0006] Obtaining an image containing a workpiece collected by a camera;

[0007] Preprocessing the image;

[0008] Inputting the preprocessed image into a trained detection model based on an improved SSD algorithm to position the workpiece, and obtaining the pixel coordinates of the center point of the workpiece;

[0009] Converting the pixel coordinates into the three-dimensional coordinates of the workpiece according to the calibration result.

[0010] Further, the preprocessing includes: converting the collected image into a grayscale image; filtering the image using a filter; enhancing the image; and extracting edges from the image.

[0011] Further, the improvements of the improved SSD algorithm include: after extracting features of six different scales by the feature extraction module in the original SSD algorithm, introducing a feature enhancement module to further enhance the features of different scales; using dilated convolution to replace the convolution in the original SSD algorithm for downsampling; and using the improved Focal loss function as the classification loss function.

[0012] Further, the introducing of the feature enhancement module to further enhance the features of different scales includes:

[0013] After increasing the scale of the high-level feature map using the attention mechanism through upsampling operation, adding it element-wise to the low-level feature map using the attention mechanism; applying the attention mechanism again to the fused feature map to obtain an attention feature map; where the attention mechanism uses the CBAM convolutional attention module; then, passing through convolutional layers with different dilation rates on the attention feature map after feature fusion, and introducing an image-level feature of global pooling; splicing the feature maps obtained through convolutional layers with different dilation rates and the global pooling layer together, and connecting a 1*1 convolution to adjust the number of channels to obtain the final enhanced feature.

[0014] Further, the CBAM convolutional attention module includes a channel attention sub-module and a spatial attention sub-module, and adjusting the weights of the channel attention sub-module and the spatial attention sub-module includes channel weight adjustment and spatial weight adjustment.

[0015] Further, the channel weight adjustment includes: performing spatial max-pooling and average pooling respectively on each channel of the feature map F to obtain two corresponding feature maps; arranging the generated max-pooling values and average pooling values in channel order to obtain two result vectors; sending the two result vectors into two-layer fully connected layers for dimensionality reduction and dimensionality increase operations; adding the output feature results and generating a channel adjustment vector through an activation function, that is, the ratio of each channel; the generation process of the channel adjustment vector is expressed as:

[0016]

[0017] In the formula, σ represents the sigmoid activation function; avgpool and maxpool represent average pooling and max-pooling operations; and are the weights of the multi-layer perceptron MLP; 、 They respectively represent two feature maps obtained by performing spatial max pooling and average pooling on each channel of the feature map F.

[0018] Furthermore, the spatial weight adjustment includes: performing max pooling and average pooling at the same position of each channel in the feature map to obtain two corresponding feature maps, then merging the two pooling results to obtain a feature map of (W, H, 2), using a convolution kernel of 7×7 on it, and activating the convolved feature map through an activation function to obtain a position adjustment vector; the generation process of the position adjustment vector is expressed as:

[0019]

[0020] In the formula, is the convolution operation with a convolution kernel of ; They respectively represent two feature maps obtained by performing average pooling and max pooling at the same position of each channel in the feature map

[0021] Furthermore, the improved Focal loss function is expressed as:

[0022]

[0023] In the formula, represents the cross-entropy loss; is a learnable function form.

[0024] Furthermore, after obtaining the three-dimensional coordinates of the workpiece, control the movement of the robotic arm based on the three-dimensional coordinates of the workpiece so that the tool end on the robotic arm aligns with the center point of the workpiece.

[0025] According to another aspect of the present invention, a workpiece positioning system based on machine vision is proposed, and the system includes:

[0026] An image acquisition module configured to acquire an image containing the workpiece collected by the camera;

[0027] A preprocessing module configured to preprocess the image;

[0028] A workpiece positioning module configured to input the preprocessed image into a trained detection model based on the improved SSD algorithm to position the workpiece and obtain the pixel coordinates of the center point of the workpiece; convert the pixel coordinates into the three-dimensional coordinates of the workpiece according to the calibration result.

[0029] The beneficial technical effects of the present invention are:

[0030] ​The present invention provides a workpiece positioning method and system based on machine vision. First, an image containing the workpiece collected by a camera is obtained; the image is preprocessed; the preprocessed image is input into a trained detection model based on the improved SSD algorithm to locate the workpiece, and the pixel coordinates of the center point of the workpiece are obtained; according to the calibration result, the pixel coordinates are converted into the three-dimensional coordinates of the workpiece. The improvements of the improved SSD algorithm include: after extracting features of six different scales by the feature extraction module in the original SSD algorithm, a feature enhancement module is introduced, and the CBAM module with adjusted weights is used to suppress irrelevant information to further enhance features of different scales, improving the feature expression ability and robustness of the model in low-light scenarios; dilated convolution is used to replace the convolution in the original SSD algorithm for downsampling to capture more context information and enrich the feature information; the improved Focal loss function is used as the classification loss function, enabling the sample importance to be adjusted autonomously according to the difficulty and category of the samples, making the model pay more attention to a small number of positive samples and difficult-to-separate samples. The present invention can more accurately locate the center point of the workpiece and is suitable for the positioning detection of different types of workpieces. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present invention will become readily understood. In the drawings, several embodiments of the present invention are shown by way of example and not limitation, wherein:

[0032] Figure 1 is a flowchart of a workpiece positioning method based on machine vision according to an embodiment of the present invention.

[0033] Figure 2 is a framework diagram of the original SSD network.

[0034] Figure 3 is a first process example diagram of feature enhancement in an embodiment of the present invention.

[0035] Figure 4 is a second process example diagram of feature enhancement in an embodiment of the present invention.

[0036] Figure 5 is a schematic diagram of the channel weight adjustment process in an embodiment of the present invention.

[0037] Figure 6 is a schematic diagram of the spatial weight adjustment process in an embodiment of the present invention.

[0038] Figure 7 is a comparison diagram of the prediction of the pixel trajectory of the center point of the workpiece by different algorithms in an embodiment of the present invention.

[0039] Figure 8It is a comparison chart of the average positioning errors of different algorithms in the embodiments of the present invention.

[0040] Figure 9 It is a schematic structural diagram of a workpiece positioning system based on machine vision according to the embodiments of the present invention. Specific embodiments

[0041] The principles and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are given only to enable those skilled in the art to better understand and then implement the present invention, and do not limit the scope of the present invention in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to be able to fully convey the scope of the present disclosure to those skilled in the art.

[0042] Those skilled in the art know that the embodiments of the present invention can be implemented as a system, device, equipment, method or computer program product. Therefore, the present disclosure can be specifically implemented in the following forms, namely: completely hardware, completely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software. In this article, it should be understood that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.

[0043] The present invention proposes a workpiece positioning method and system based on machine vision. The machine vision system is used to realize the fixed-point detection of workpieces by an industrial robotic arm. Its hardware structure includes: an image acquisition module for the workpiece composed of an industrial camera, a lens and an external light source; the human-computer interaction software includes image preprocessing of the images acquired by the acquisition module and workpiece positioning. By establishing communication between the upper computer and the industrial robotic arm, the interaction between the upper computer and the lower computer of the robotic arm is realized, and finally the positioning of the workpiece can be realized through the monitoring of the lower computer by the upper computer.

[0044] An embodiment of the present invention proposes a workpiece positioning method based on machine vision, as Figure 1 shown, the method includes:

[0045] S1. Obtain an image containing the workpiece collected by the camera;

[0046] S2. Preprocess the image;

[0047] S3. Input the preprocessed image into a trained detection model based on the improved SSD algorithm to position the workpiece, and obtain the pixel coordinates of the center point of the workpiece;

[0048] S4. Convert the pixel coordinates into the three-dimensional coordinates of the workpiece according to the calibration result;

[0049] S5. Control the movement of the robotic arm based on the three-dimensional coordinates of the workpiece, so that the end of the tool on the robotic arm is aligned with the center point of the workpiece.

[0050] The method starts from S1. In S1, an image containing the workpiece collected by the camera is obtained.

[0051] According to an embodiment of the present invention, during the workpiece positioning process, first adjust the posture of the camera to keep the camera plane parallel to the workpiece plane and make the workpiece appear in the middle of the camera's field of view, and set the initial pose of the robot; collect the workpiece image and send it to the upper computer.

[0052] Then execute S2. In S2, preprocess the image.

[0053] According to an embodiment of the present invention, the main steps of image preprocessing are as follows: First, convert the collected image into a grayscale image to reduce the color elements in the image and retain the image information features; filter the image with an appropriate filter to reduce the irrelevant elements in the image and emphasize the target features at the same time; perform enhancement processing on the image, select a suitable enhancement operator according to different application scenarios, and selectively highlight the features of the entire image or local areas to improve the visual effect of the image; extract the edges of the workpiece. In the industrial field, mean filtering, median filtering, and Gaussian filtering are three relatively commonly used filtering methods. Image enhancement can be divided into two methods according to the different enhancement spaces: spatial domain method and frequency domain method. In the spatial domain method, various linear or non-linear operations are directly performed on the image itself to improve the pixel values in the image. The main methods include gray-scale transformation, histogram correction, image smoothing, and image sharpening; the frequency domain method regards the image as a two-dimensional signal and amplifies the signal through two-dimensional Fourier transform. The main algorithms include high-pass filtering, low-pass filtering, homomorphic filtering, etc. The purpose of edge extraction is to divide the image into independent regions with different features, and then select the region of interest according to the system requirements.

[0054] In an embodiment of the present invention, image filtering processing is performed on the basis of the grayscale image. Different mask radii are selected for each filtering processing method for experimental demonstration. By comparing the different mask radii in different filtering methods, the median filtering with a mask radius of 5*5 is finally selected as the image filtering processing operator; in order to make the elements in each part of the image display more clearly, the histogram equalization operator is selected to enhance the image. By changing the image structure and making the image histogram evenly distributed, a region of interest with clearer contours can be obtained; the edges of the workpiece are extracted by threshold segmentation, and appropriate pixel values are selected to separate the foreground and background, laying a foundation for subsequent accurate positioning of the workpiece.

[0055] Then execute S3. In S3, input the preprocessed image into a trained detection model based on the improved SSD algorithm to locate the workpiece and obtain the pixel coordinates of the center point of the workpiece.

[0056] According to the embodiments of the present invention, the algorithm framework and implementation details of SSD object detection enable it to have excellent performance in object detection tasks. Considering the characteristics of few workpiece detection features and high positioning accuracy requirements, in-depth research on workpiece positioning technology is carried out based on the SSD object detection algorithm, and the following detection model based on the improved SSD algorithm is designed.

[0057] The SSD algorithm realizes the accurate positioning and classification of objects at one time by converting the detection object into the idea of regular expression. The SSD algorithm is an improvement of the VGG-16 network. The SSD network mainly consists of three parts, and its network framework is as Figure 2 shown. First, Conv1 to Conv5 in VGG-16 are retained as the basic skeleton, and the following two fully connected layers FC6 and FC7 are replaced by two ordinary convolutional layers Conv6 and Conv7 to extract the features of the entire image; subsequently, four convolutional layers Conv8_2 to Conv11_2 with different scales are added to extract features of different scales; finally, the features from six feature layers with different scales are sent to the classification and positioning layer, and non-maximum suppression is applied to screen out the optimal results.

[0058] In the original SSD network, different feature layers are independent of each other and lack feature complementarity. The deeper the feature layer, the lower the resolution and the lower the workpiece positioning accuracy. Workpieces belong to larger objects, and are mainly detected and positioned by the four feature extraction layers Conv8_2, Conv9_2, Conv10_2, and Conv11_2. In order to improve the positioning accuracy of workpieces, the following improvements are introduced based on the original SSD algorithm.

[0059] 1) After extracting features of six different scales, a feature enhancement module is introduced to further enhance the extracted features.

[0060] Specifically, first, the high-level feature map using the attention mechanism is upsampled to increase the scale and then element-wise added to the low-level feature map using the attention mechanism, so as to retain the semantic information of the high-level feature map and the detail information of the low-level feature map; the attention mechanism is used again on the fused feature map to perform feature mapping to obtain the attention feature map; the attention mechanism adopts the convolutional block attention module - CBAM. Figure 3 The feature enhancement process of this part is shown taking the latter three feature maps as an example. D-FEM in the figure is the following process.

[0061] The D-FEM process is as Figure 4As shown, then, on the attention feature map after feature fusion, convolutional layers with different dilation rates (the dilation rates can be 3, 6, 9) are passed through respectively to obtain receptive fields of different scales. In addition, an image-level feature of global pooling is introduced to compensate for the convolution kernel degradation problem caused by large dilation rates; the feature maps obtained through convolutional layers with different dilation rates and the global pooling layer are spliced together, and then a 1*1 convolution is connected to adjust the number of channels to obtain the final enhanced feature.

[0062] The convolutional attention module has 2 consecutive sub-modules: the channel attention sub-module and the spatial attention sub-module. The channel attention module selects which channel features are the most meaningful and contributive; the spatial attention sub-module selects which regions of the feature map the model should focus on and ignores irrelevant regions. Further, the weights of the channel attention sub-module and the spatial attention sub-module are adjusted to focus on useful information with high weights, ignore irrelevant information with low weights, and continuously adjust the weights to obtain more important information in different situations to ensure the usefulness of the feature map in terms of space and channels and make the localization more accurate.

[0063] Channel weight adjustment is used to focus on the target channel information and adjust the weight of the proportion relationship of each channel of the image. Specifically, as Figure 5 shown, spatial max pooling and average pooling are respectively performed on each channel of the feature map F to obtain two corresponding feature maps; the generated max pooling values and average pooling values are arranged in channel order to obtain two result vectors; the two result vectors are respectively sent to a two-layer fully connected layer for dimensionality reduction and dimensionality increase operations; the output feature results are added together, and passed through an activation function to generate a channel adjustment vector, that is, the proportion of each channel; the generation process of the channel adjustment vector is expressed as:

[0064]

[0065] In the formula, σ represents the sigmoid activation function; avgpool and maxpool represent average pooling and max pooling operations; and are the weights of the multi-layer perceptron MLP; 、 respectively represent the two feature maps obtained by performing spatial max pooling and average pooling on each channel of the feature map F.

[0066] Spatial weight adjustment is used to focus on the target position information and adjust the weight of the proportion relationship of each pixel point of the image. Specifically, as Figure 6 shown, on the feature map Max pooling and average pooling are respectively performed at the same position of each channel to obtain two corresponding feature maps, and then the results of the two poolings are merged to obtain a feature map with a width of W, a height of H, and a channel number of 2. A 7×7 convolution kernel is used on it, and the convolutional feature map is activated through an activation function to obtain a position adjustment vector. The generation process of the position adjustment vector is expressed as:

[0067]

[0068] In the formula, represents the convolution operation with a convolution kernel of ; respectively represent two feature maps obtained by performing average pooling and max pooling at the same position of each channel in the feature map ;

[0069] 2) In the SSD algorithm, the receptive field of the convolutional layer added after the backbone extraction network changes little, and the extracted features are not sufficient. Therefore, dilated convolution is used to replace the original convolution for downsampling. For the convolution and the dilated convolution The calculation method of the receptive field is represented by the following formula.

[0070]

[0071] In the formula, represents the receptive field of this convolutional layer, represents the receptive field of the previous layer; represents the stride of the convolution, k represents the size of the convolution kernel; r represents the dilation coefficient. Taking a 3×3 convolution kernel with a stride of 1 as an example, it can be calculated that the receptive field during convolution is 3. When the dilation coefficient r is 1, the dilated convolution is equivalent to convolution, and the receptive field is still 3; when the dilation coefficient is 2, the receptive field when using dilated convolution is 5. It can be seen that using dilated convolution can expand the receptive field and capture more context information.

[0072] 3) Use the improved Focal loss function as the classification loss function.

[0073] In SSD, the loss function is mainly divided into two categories: classification loss and regression loss. The classification loss usually uses cross-entropy loss, and the regression loss uses smooth L1 loss or MSE loss.

[0074] In the embodiments of the present invention, the classification loss with a large influence degree is improved according to the Focal Loss function. The formula of the Focal Loss function is:

[0075]

[0076] Wherein, represents the cross-entropy loss, that is, the probability that the model predicts the positive class; is the class weight, which is used to adjust the ratio of positive and negative samples; is the adjustment factor; is a hyperparameter used to control the attention degree of easy and difficult samples. The FocalLoss loss function can focus the training process on difficult and important samples, alleviating the training problems caused by sample class imbalance.

[0077] However, the above loss function form is fixed and lacks flexibility. It cannot adaptively adjust the shape of the loss function according to the data distribution and task characteristics. The hyperparameters and need to be manually tuned; moreover, for different datasets and tasks, the optimal hyperparameter values may be different and need to be determined through a large number of experiments. At the same time, it does not consider the difficulty differences between different samples, resulting in poor performance when dealing with difficult samples.

[0078] Therefore, the embodiment of the present invention improves the Focal Loss loss function, and the improved Focal loss loss function is expressed as:

[0079]

[0080] Wherein, represents the cross-entropy loss between the predicted probability output by the model and the true label; is a learnable function form that can be adjusted and optimized according to task requirements and data characteristics.

[0081] The improved Focal loss loss function is more flexible and adaptive, and can automatically adjust the loss function parameters according to different data distributions and task requirements, thereby improving the performance and generalization ability of the model.

[0082] After training the detection model based on the improved SSD algorithm, input the preprocessed image to be measured into the trained detection model to locate the workpiece, and obtain the pixel coordinates of the workpiece center point.

[0083] Then execute S4. In S4, convert the pixel coordinates into the three-dimensional coordinates of the workpiece according to the calibration result.

[0084] According to the embodiment of the present invention, the two-dimensional pixel value of the workpiece center point can be converted into the camera coordinate value through the existing camera calibration method. Then, according to the conversion relationship between the camera coordinate system and the world coordinate system, the camera coordinate value is converted into the three-dimensional coordinate of the workpiece.

[0085] Further, it further includes S5. In S5, the movement of the robotic arm is controlled based on the three-dimensional coordinates of the workpiece, so that the tool end on the robotic arm is aligned with the center point of the workpiece.

[0086] According to the embodiment of the present invention, the workpiece position information is transmitted to the robot control cabinet to control the movement of the robotic arm, so that the tool end sleeve is aligned with the center point of the workpiece; after the robot completes the corresponding action, it returns to the initial posture. The workpiece position information needs to be the coordinate position in the tool coordinate system. In the machine vision system, the camera is generally fixed at the tool end of the robot. If the camera coordinate value needs to be converted into the three-dimensional coordinate value in the tool coordinate system, the relationship between the camera and the end tool of the robot also needs to be established. The conversion relationship between the camera and the end tool can be obtained through the existing hand-eye calibration method, so as to realize the transformation of the camera coordinate value to the tool coordinate value. Further, the tool coordinate value can be converted into the coordinate value of the robot according to the transformation matrix obtained by the tool coordinate system calibration.

[0087] Further, the technical effect of the present invention is verified through experiments.

[0088] To verify the effectiveness and robustness of the method proposed by the present invention in workpiece positioning, keep the workpiece stationary, control the camera to move in the X-axis direction, so that the workpiece image is located in the middle of the camera field of view and take this as the base point, and transform six different positions on both the left and right sides of the base point, and collect the corresponding number of workpiece images. Due to the movement of the camera, a movement trajectory will also be generated for the center point of the workpiece. The method proposed by the present invention is compared with the original SSD algorithm, VGG-19-SSD algorithm, and ResNet-50-FPN-SSD algorithm. The prediction results of the center point trajectory are as Figure 7 shown. From Figure 7 it can be seen that the center point trajectory located by the method proposed by the present invention is generally closer to the real trajectory than the other three algorithms. Figure 8 The average positioning errors of different algorithms are shown. From Figure 8 it can be seen that the method proposed by the present invention can keep the average positioning error within 0.6 mm.

[0089] Another embodiment of the present invention proposes a workpiece positioning system based on machine vision. As Figure 9 shown, the system includes:

[0090] An image acquisition module 910 configured to acquire an image containing a workpiece collected by a camera;

[0091] A preprocessing module 920 configured to preprocess the image;

[0092] The workpiece positioning module 930 is configured to input the preprocessed image into the trained detection model based on the improved SSD algorithm to position the workpiece and obtain the pixel coordinates of the center point of the workpiece; and convert the pixel coordinates into the three-dimensional coordinates of the workpiece according to the calibration result.

[0093] The functions of the workpiece positioning system based on machine vision in the embodiments of the present invention can be illustrated by the foregoing workpiece positioning method based on machine vision. Therefore, for the parts not described in detail in this embodiment, reference can be made to the above method embodiments, which will not be elaborated herein.

[0094] It should be noted that although several units, modules or sub-modules are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present invention, the features and functions of two or more modules described above can be embodied in one module. Conversely, the features and functions of one module described above can be further divided and embodied by multiple modules.

[0095] In addition, although the operations of the method of the present invention are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.

[0096] Although the spirit and principles of the present invention have been described with reference to several specific embodiments, it should be understood that the present invention is not limited to the specific embodiments disclosed, and the division of each aspect does not mean that the features in these aspects cannot be combined for benefit. This division is only for the convenience of expression. The present invention aims to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.

Claims

1. A workpiece positioning method based on machine vision, characterized in that: include: Acquire an image containing a workpiece captured by a camera; Preprocessing the image; The preprocessed image is input into the trained detection model based on the improved SSD algorithm to locate the workpiece and obtain the pixel coordinates of the center point of the workpiece; The improvements of the improved SSD algorithm include: after extracting six features of different scales based on the feature extraction module in the original SSD algorithm, introducing a feature enhancement module to further enhance the features of different scales; using a dilated convolution to replace the convolution in the original SSD algorithm for downsampling; using an improved Focal loss loss function as a classification loss function; wherein the introduction of the feature enhancement module to further enhance the features of different scales includes: increasing the scale of the high-level feature map using the attention mechanism through an upsampling operation, and then adding it element by element with the low-level feature map using the attention mechanism; using the attention mechanism to perform feature mapping on the fused feature map again to obtain an attention feature map; wherein the attention mechanism uses a CBAM convolution attention module; then, the attention feature map after feature fusion is passed through convolution layers with different dilation rates respectively, and a global pooling image-level feature is introduced; the feature maps obtained through the convolution layers with different dilation rates and the global pooling layer are spliced ​​together, and the 1*1 convolution is connected to adjust the number of channels to obtain the final enhanced features; the improved Focal loss loss function is expressed as: ; In the formula, represents the cross entropy loss; and It is a learnable function form; The pixel coordinates are converted into three-dimensional coordinates of the workpiece according to the calibration result.

2. A workpiece positioning method based on machine vision according to claim 1, characterized in that: The preprocessing includes: converting the collected image into a grayscale image; filtering the image using a filter; performing enhancement processing on the image; and extracting the edge of the image.

3. The workpiece positioning method based on machine vision according to claim 1, characterized in that: The CBAM convolutional attention module includes a channel attention submodule and a spatial attention submodule, and the weights of the channel attention submodule and the spatial attention submodule are adjusted, including channel weight adjustment and spatial weight adjustment.

4. A workpiece positioning method based on machine vision according to claim 3, characterized in that: The channel weight adjustment includes: performing spatial maximum pooling and mean pooling on each channel of the feature map F to obtain two corresponding feature maps; arranging the generated maximum pooling values ​​and mean pooling values ​​in channel order to obtain two result vectors; sending the two result vectors to two fully connected layers for dimensionality reduction and dimensionality increase operations; adding the output feature results and generating a channel adjustment vector, that is, the ratio of each channel, through an activation function; the generation process of the channel adjustment vector is expressed as: ; In the formula, Represents the sigmoid activation function; and Represents mean pooling and maximum pooling operations; and Multilayer Perceptron The weight of They respectively represent the two feature maps obtained after performing spatial maximum pooling and mean pooling on each channel of the feature map F.

5. A workpiece positioning method based on machine vision according to claim 4, characterized in that: The spatial weight adjustment includes: The same position of each channel is subjected to maximum pooling and mean pooling respectively to obtain the corresponding two feature maps, and then the two pooling results are merged to obtain the feature map of (W, H, 2), and a 7*7 convolution kernel is used for it. The convolution feature map is activated by the activation function to obtain the position adjustment vector; the generation process of the position adjustment vector is expressed as: ; In the formula, The convolution kernel is Convolution operation; ; Respectively represented in the feature map Two feature maps are obtained by performing mean pooling and maximum pooling at the same position of each channel.

6. A workpiece positioning method based on machine vision according to any one of claims 1 to 5, characterized in that: After acquiring the three-dimensional coordinates of the workpiece, the movement of the robot arm is controlled based on the three-dimensional coordinates of the workpiece so that the end of the tool on the robot arm is aligned with the center point of the workpiece.

7. A workpiece positioning system based on machine vision, characterized in that: include: An image acquisition module configured to acquire an image including a workpiece acquired by a camera; A preprocessing module, configured to preprocess the image; The workpiece positioning module is configured to input the preprocessed image into the trained detection model based on the improved SSD algorithm to locate the workpiece and obtain the pixel coordinates of the center point of the workpiece. The improvements of the improved SSD algorithm include: after the feature extraction module in the original SSD algorithm extracts six different scales of features, a feature enhancement module is introduced to further enhance the features of different scales; the convolution in the original SSD algorithm is replaced by the dilated convolution for downsampling; the improved Focal loss loss function as a classification loss function; wherein the introduction of the feature enhancement module to further enhance the features of different scales includes: increasing the scale of the high-level feature map using the attention mechanism through upsampling operation, and then adding it element by element with the low-level feature map using the attention mechanism; using the attention mechanism to perform feature mapping on the fused feature map again to obtain the attention feature map; wherein the attention mechanism adopts the CBAM convolution attention module; then, the attention feature map after feature fusion is passed through convolution layers with different void rates respectively, and a global pooling image-level feature is introduced; the feature maps obtained through the convolution layers with different void rates and the global pooling layer are spliced ​​together, and the 1*1 convolution is connected to adjust the number of channels to obtain the final enhanced features; the improved Focal loss loss function is expressed as: ; In the formula, represents the cross entropy loss; and It is a learnable function form; according to the calibration result, the pixel coordinates are converted into the three-dimensional coordinates of the workpiece.

Citation Information

Patent Citations

  • Method for detecting and identifying floating objects on water based on improved SSD (Solid State Disk) algorithm

    CN114782772A

  • Workpiece surface defect detection method based on improved SSD algorithm

    CN116071331A

  • Pavement disease rapid detection method and system based on YOLOV7 algorithm

    CN117058459A