A Deep Learning-Based Automatic Positioning and Decoding Method for Linear Encoders

Through the coarse positioning and fine positioning decoding model of the code channel based on deep learning, the accuracy and real-time problems of the grating scale under the changes in the lighting environment are solved, and high-precision automatic positioning and decoding of the grating scale is realized, reducing the dependence on high-precision optical devices.

CN116109702BActive Publication Date: 2025-07-18GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310121340.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-13
Publication Date
2025-07-18
Estimated Expiration
2043-02-13

AI Technical Summary

Technical Problem

The existing code channel positioning technology restricts the accuracy and real-time nature of grating scale positioning and decoding, especially when the measurement accuracy is affected when the lighting environment changes, and the calibration process is complex and the information loss is large.

Method used

The code channel coarse positioning decoding model and the code channel fine positioning decoding model are used based on deep learning to obtain the coarse positioning information and fine positioning information of the raster scale image respectively. Combined with the code element width, it is interpreted through the code channel coarse positioning network and the code channel fine positioning network to achieve accurate automatic positioning and decoding of the raster scale.

Benefits of technology

It realizes high-precision grating scale positioning and decoding when the lighting environment changes, meets the real-time measurement requirements, and reduces the dependence on optical amplification devices and high-precision photodetectors, improving the applicability and accuracy of grating scales.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116109702B_ABST
    Figure CN116109702B_ABST
Patent Text Reader

Abstract

The present invention provides a deep learning-based automatic positioning and decoding method for grating scales, comprising the following steps: S1: constructing a coarse positioning and decoding model for code tracks and a fine positioning and decoding model for code tracks; S2: obtaining a binary image of the code track and the abscissa of the center of the code element from the input grating scale image, and obtaining a binary image for the positioning of the end code element from the grating scale image; S3: interpreting the coarse position information of the grating scale image according to the binary image of the code track and the abscissa of the center of the code element, and interpreting the fine position information of the grating scale image according to the binary image for the positioning of the end code element; S4: obtaining the absolute position information based on the code element width, the coarse position information and the fine position information of the grating scale image as the automatic positioning and decoding result. The present invention provides a deep learning-based automatic positioning and decoding method for grating scales, which solves the problem that the existing code track positioning technology restricts the accuracy of grating scale positioning and decoding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of grating scales, and more specifically, to an automatic positioning and decoding method for grating scales based on deep learning. Background Art

[0002] A grating scale, also known as a grating scale displacement sensor (grating scale sensor), is a high-precision position sensor and a core measuring device for micro-nano ultra-precision machining equipment. Generally, there are four levels of precision, and the highest can reach ±1μm. According to different measurement methods, it is mainly divided into absolute grating scales and incremental grating scales. Currently, absolute grating scales have a very wide application prospect in numerical control machine tools. An absolute grating scale consists of a series of absolute codes engraved on the scale. It does not require homing. The position information comes from the grating code disk, and the position value can be obtained immediately after power-on. Moreover, its maximum scanning speed is not affected by the maximum input frequency of the electronic device. Therefore, it can achieve both high speed and high resolution simultaneously.

[0003] An absolute grating scale generally has an absolute code track and an incremental code track. Among them, the incremental code track is used to subdivide the code elements to improve the precision. However, since the resolution of the absolute code track is much lower than that of the incremental code track, an optical amplification device and a very high-precision photodetector are required, resulting in a relatively high cost.

[0004] In addition, the measurement accuracy of an absolute grating scale highly depends on the code track positioning. As the illumination environment changes, the code track image changes accordingly. To ensure the measurement accuracy, the existing code track positioning technology needs to adjust the algorithm on-site to complete the calibration of the grating scale. Moreover, before calibration, the image needs to be digitized, and there is also a certain loss of information during the conversion process. This restricts the real-time performance, applicability, and accuracy of the grating scale positioning and decoding. Summary of the Invention

[0005] The present invention provides an automatic positioning and decoding method for grating scales based on deep learning to overcome the technical defect that the existing code track positioning technology restricts the accuracy of grating scale positioning and decoding.

[0006] To solve the above technical problems, the technical solution of the present invention is as follows:

[0007] An automatic positioning and decoding method for grating scales based on deep learning, comprising the following steps:

[0008] S1: Construct a code track rough positioning and decoding model and a code track fine positioning and decoding model based on deep learning.

[0009] The code track rough positioning and decoding model includes a code track rough positioning network and a code track rough decoding network.

[0010] The code track fine positioning and decoding model includes a code track fine positioning network and a code track fine decoding network.

[0011] S2: Use the coarse code track positioning network to obtain a binary image of the code track and the abscissa of the center of the code element from the input grating scale image,

[0012] Use the fine code track positioning network to obtain a binary image for positioning the end code element from the grating scale image;

[0013] S3: Use the coarse code track decoding network to decode the coarse position information of the grating scale image according to the binary image of the code track and the abscissa of the center of the code element,

[0014] Use the fine code track decoding network to decode the fine position information of the grating scale image according to the binary image for positioning the end code element;

[0015] S4: Obtain the absolute position information based on the code element width, coarse position information, and fine position information of the grating scale image as the automatic positioning decoding result.

[0016] In the above solution, the coarse position information and the fine position information are respectively obtained from the input grating scale image through the coarse code track positioning and decoding model and the fine code track positioning and decoding model based on deep learning, and combined with the code element width of the grating scale image to obtain accurate absolute position information, realizing accurate automatic positioning and decoding of the grating scale. At the same time, the end-to-end model based on deep learning is relatively lightweight and has a fast running speed, which can meet the real-time requirement during measurement.

[0017] Preferably, the coarse code track positioning network includes a first segmentation module and a positioning center module,

[0018] Use the first segmentation module to obtain a binary image of the code track from the grating scale image,

[0019] Use the positioning center module to obtain the abscissa of the center of the code element according to the binary image of the code track;

[0020] The first segmentation module includes a first downsampling module, a second downsampling module, a third downsampling module, a first convolution module, a first upsampling module, a second upsampling module, a third upsampling module, a first convolutional layer, and a first normalization layer connected in sequence; among them, one branch output end of the first downsampling module is connected to the branch input end of the third upsampling module, one branch output end of the second downsampling module is connected to the branch input end of the second upsampling module, and one branch output end of the third downsampling module is connected to the branch input end of the first upsampling module;

[0021] The positioning center module includes a first convolutional pooling module, a second convolutional module, a fourth downsampling module, a fifth downsampling module, a sixth downsampling module, and a first fully connected linear layer connected in sequence; among them, the output of the convolutional pooling module also includes local response normalization processing before inputting into the second convolutional module.

[0022] Preferably, the coarse code track decoding network extracts the center point information of the code element according to the binary image of the code track and the abscissa of the center of the code element, and obtains the coarse position information from the pre-constructed coding position database by means of look-up table method according to the center point information of the code element.

[0023] Preferably, the fine code track positioning network includes a super-resolution module and a second segmentation module;

[0024] The super-resolution module is used to perform feature extraction, non-linear mapping and image reconstruction on the grating scale image, so as to obtain a clear high-resolution image;

[0025] The second segmentation module is used to obtain the binary image of the end code element positioning from the high-resolution image;

[0026] The super-resolution module includes a plurality of convolutional layers;

[0027] The second segmentation module includes a seventh downsampling module, an eighth downsampling module, a ninth downsampling module, a tenth downsampling module, a third convolutional module, a fourth upsampling module, a fifth upsampling module, a sixth upsampling module, a seventh upsampling module, a second convolutional layer and a second normalization layer which are connected in sequence; wherein, one branch output end of the seventh downsampling module is connected to the branch input end of the seventh upsampling module, one branch output end of the eighth downsampling module is connected to the branch input end of the sixth upsampling module, one branch output end of the ninth downsampling module is connected to the branch input end of the fifth upsampling module, and one branch output end of the tenth downsampling module is connected to the branch input end of the fourth upsampling module.

[0028] Preferably, the fine code track decoding network includes an eleventh downsampling module, a twelfth downsampling module, a first depth downsampling module, a second depth downsampling module, a third depth downsampling module and a second fully connected linear layer which are connected in sequence.

[0029] Preferably, the fine code track positioning and decoding model further includes a feedback compensation network;

[0030] The feedback compensation network is used to determine whether the binary image of the end code element positioning is a boundary code element image. When it is determined to be a boundary code element image, 0 is used as the output of the fine code track decoding network. When it is determined not to be a boundary code element image, the actual output of the fine code track decoding network is taken as the standard;

[0031] The feedback compensation network includes a second convolutional pooling module, a third convolutional pooling module, a fourth convolutional pooling module, a fifth convolutional pooling module, a third fully connected layer and a classification module which are connected in sequence.

[0032] Preferably, the loss function of the coarse code track positioning network is as follows:

[0033]

[0034]

[0035]

[0036] Among them, is the loss function of the first segmentation module, l log is the loss function of the positioning center module, μ is and l log 's weight ratio, W is the image width, H is the image height, is the pixel value of each pixel of the image after being segmented by the first segmentation module, is the corresponding pixel value of the original binary map label, p(θ(x)) is the softmax loss function, and θ(x) is the label value of the abscissa pixel point of the code track center.

[0037] Preferably, the loss function of the code track fine positioning decoding model is as follows:

[0038]

[0039]

[0040]

[0041]

[0042]

[0043]

[0044]

[0045]

[0046] Among them, is the loss function of the code track fine positioning network, l acc is the loss function of the code track fine decoding network, θ is and l acc 's weight ratio, is the loss function of the super-resolution module, l elog is the loss function of the second segmentation module, is the mean square error loss function, is the edge loss function, ε is and 's weight ratio, W is the width of the image, H is the height of the image, is the pixel of the super-resolved image, is the pixel of the original image, E x,y is the result of using Canny edge detection, p(θ(x)) is the softmax loss function, θ(x) is the label value of the abscissa pixel point of the code track center, w is the weight value of the code track boundary pixel point, w c is the weight value for balancing the class ratio, d1 is the distance from the symbol pixel point to the nearest pixel point, d2 is the distance from the symbol pixel point to the second nearest pixel point, w0 and σ are different constant values, is the position information label, is the output position information, x is the abscissa of the pixel point, and y is the ordinate of the pixel point.

[0047] Preferably, the loss function l of the feedback compensation network com is as follows:

[0048] l com = -(y1logp(x1)+(1 - y1)log(1 - p(x1))

[0049] where x1 is the output data, y1 is the true label, and p(x1) is the probability function.

[0050] Preferably, the total loss function l is set total as follows:

[0051]

[0052] where is the loss function of the first segmentation module, l loh is the loss function of the positioning center module, l acc is the loss function of the code track fine decoding network, l com is the loss function of the feedback compensation network, is the loss function of the code track fine positioning network, ρ is and the weight value between, σ is l log and the weight value between, τ is l acc and the weight value between, is the weight value between l com and between.

[0053] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:

[0054] The present invention provides a deep learning-based automatic positioning and decoding method for grating scales. By using a deep learning-based coarse positioning and decoding model for code tracks and a fine positioning and decoding model for code tracks, the coarse position information and the fine position information are respectively obtained from the input grating scale image, and combined with the code element width of the grating scale image to obtain accurate absolute position information, realizing accurate automatic positioning and decoding of the grating scale. At the same time, the end-to-end model based on deep learning is relatively lightweight and has a fast running speed, which can meet the real-time requirements during measurement. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 It is a flowchart of the implementation steps of the technical solution of the present invention;

[0056] Figure 2 It is a schematic diagram of the overall framework of the present invention;

[0057] Figure 3 It is a schematic diagram of the grating scale image of the present invention;

[0058] Figure 4 It is a schematic diagram of the network structure of the coarse decoding network for code tracks in the present invention;

[0059] Figure 5 It is a schematic diagram of the table stored in the encoding position database in the present invention;

[0060] Figure 6 It is a schematic diagram of the information marking of the grating scale image in the present invention;

[0061] Figure 7 It is a schematic diagram of the network structure of the fine positioning and decoding model for code tracks in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0062] The drawings are only for illustrative purposes and should not be construed as a limitation of this patent;

[0063] To better illustrate this embodiment, some components in the drawings are omitted, enlarged or reduced, and do not represent the size of the actual product;

[0064] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0065] The technical solution of the present invention will be further described below with reference to the drawings and embodiments.

[0066] Embodiment 1

[0067] As Figures 1-3 shown, a deep learning-based automatic positioning and decoding method for grating scales includes the following steps:

[0068] S1: Construct a deep learning-based coarse positioning and decoding model for code tracks and a fine positioning and decoding model for code tracks,

[0069] The rough track positioning and decoding model includes a rough track positioning network and a rough track decoding network,

[0070] The fine track positioning and decoding model includes a fine track positioning network and a fine track decoding network;

[0071] S2: Use the rough track positioning network to obtain a track binary image and the abscissa of the center of the code element from the input grating scale image,

[0072] Use the fine track positioning network to obtain a binary image for positioning the end code element from the grating scale image;

[0073] S3: Use the rough track decoding network to interpret the rough position information of the grating scale image according to the track binary image and the abscissa of the center of the code element,

[0074] Use the fine track decoding network to interpret the fine position information of the grating scale image according to the binary image for positioning the end code element;

[0075] S4: Obtain the absolute position information based on the code element width, rough position information, and fine position information of the grating scale image as the automatic positioning and decoding result.

[0076] In the specific implementation process, the rough position information and the fine position information are respectively obtained from the input grating scale image through the rough track positioning and decoding model and the fine track positioning and decoding model based on deep learning, and combined with the code element width of the grating scale image to obtain accurate absolute position information, realizing accurate automatic positioning and decoding of the grating scale. At the same time, the end-to-end model based on deep learning is relatively lightweight and has a fast running speed, which can meet the real-time performance during measurement.

[0077] Embodiment 2

[0078] A method for automatic positioning and decoding of a grating scale based on deep learning includes the following steps:

[0079] S1: Construct a rough track positioning and decoding model and a fine track positioning and decoding model based on deep learning,

[0080] The rough track positioning and decoding model includes a rough track positioning network and a rough track decoding network,

[0081] The fine track positioning and decoding model includes a fine track positioning network and a fine track decoding network;

[0082] S2: Use the rough track positioning network to obtain a track binary image and the abscissa of the center of the code element from the input grating scale image,

[0083] Use the fine track positioning network to obtain a binary image for positioning the end code element from the grating scale image;

[0084] S3: Using the coarse decoding network of the code track to interpret the coarse position information of the grating scale image according to the binary image of the code track and the abscissa of the code element center,

[0085] Using the fine decoding network of the code track to interpret the fine position information of the grating scale image according to the binary image of the tail-end code element positioning;

[0086] S4: Obtaining the absolute position information as the automatic positioning decoding result according to the code element width, the coarse position information and the fine position information of the grating scale image.

[0087] More specifically, as Figure 4 shown, the coarse positioning network of the code track includes a first segmentation module and a positioning center module,

[0088] Using the first segmentation module to obtain the binary image of the code track from the grating scale image,

[0089] Using the positioning center module to obtain the abscissa of the code element center according to the binary image of the code track;

[0090] The first segmentation module includes a first downsampling module, a second downsampling module, a third downsampling module, a first convolution module, a first upsampling module, a second upsampling module, a third upsampling module, a first convolution layer and a first normalization layer connected in sequence; wherein, a branch output end of the first downsampling module is connected to a branch input end of the third upsampling module, a branch output end of the second downsampling module is connected to a branch input end of the second upsampling module, and a branch output end of the third downsampling module is connected to a branch input end of the first upsampling module;

[0091] The first downsampling module includes two convolutional layers and one pooling layer. Among them, the first convolutional layer has a convolutional kernel of 3×3 and a depth of 64. After convolving the input image and performing regular padding on the boundaries, a feature map with a depth of 64 is obtained. The second convolutional layer has a convolutional kernel of 1×1 and a depth of 64, and further extracts the feature map. Then, the pooling layer performs regular pooling on the feature map, and the feature map will be reduced to half of the input image and used as the input of the second downsampling module. The second downsampling module and the third downsampling module are similar to the first downsampling module, both including two convolutional layers and one pooling layer. The difference is that the depth of the convolutional kernel of the second downsampling module is 128, and the convolutional kernel of the third downsampling module is 256. And due to the pooling layer, the output feature map of the second downsampling module will be reduced to one-fourth of the input image, and the output feature map of the third downsampling module will be reduced to one-eighth of the input image. The first convolutional module takes a feature map with a depth of 256 and a length and width that are one-eighth of the input image as the input, performs one convolution with a convolutional kernel of 1×1 and a depth of 512, and outputs a feature map with a depth of 512 and a length and width that are one-eighth of the input image, which is used as the input of the first upsampling module. The first upsampling module includes one transposed convolutional layer and two convolutional layers. Among them, the transposed convolutional layer can perform regular expansion on each element of the feature map. The uniform expansion adopted in this embodiment will output an expansion that is twice the original size. Then, it passes through a convolutional layer with a convolutional kernel of 3×3 and a depth of 256, and performs regular padding on the boundaries. Then, it passes through a convolutional layer with a convolutional kernel of 1×1 and a depth of 256 in the first upsampling module to obtain a feature map with a length and width that are one-fourth of the input image, and performs element-wise addition with the output feature map of the third downsampling module without passing through the pooling layer to obtain a feature map after feature fusion, that is, the output of the first upsampling module. The second upsampling module and the third upsampling module also include one transposed convolutional layer and two convolutional layers. The difference is that the depths of the convolutional kernels of their convolutional layers are 128 and 64 respectively, and they perform feature fusion with the intermediate outputs of the second downsampling module and the first downsampling module respectively. Finally, the third upsampling module will output a feature map with a depth of 64 and a length and width consistent with the input image. Then, it passes through a first convolutional layer with a convolutional kernel of 1×1 and a depth of 2 to obtain a probability map with a depth of 2. The softmax function of the first normalization layer is used to judge the two-layer probability map, and the point with the largest probability is retained to obtain the binarized image of the code track.

[0092] When the code track image is processed by the first segmentation module, a feature probability map O1 with two dimensions of 0 and 1 will be obtained. The next classification is required.

[0093] O2 = softmax(O1)

[0094] O2 is the binarized image obtained after classification, which only has two elements, 0 and 1.

[0095] The positioning center module includes a first convolutional pooling module, a second convolutional module, a fourth downsampling module, a fifth downsampling module, a sixth downsampling module, and a first fully-connected linear layer connected in sequence; wherein, before the output of the convolutional pooling module is input into the second convolutional module, local response normalization processing is also performed;

[0096] The code track binary image output by the first segmentation module is used as the input of the first convolutional pooling module. The first convolutional pooling module includes a convolutional layer with a convolution kernel of 3×3 and a depth of 64, and a pooling layer of 2×2. The output is a feature map with a length and width that are half of the binary image and a depth of 64.

[0097]

[0098]

[0099] Among them, is the image after pooling and convolution. In this embodiment, N is 64, n is 5, k is 0, α is 1*e -4 , and β is 0.75.

[0100] The feature map output by the first convolutional pooling module is subjected to local response normalization to amplify the features of each symbol, but the depth still remains 64 layers, and it is used as the input of the second convolutional module. The second convolutional module includes a convolutional layer with a convolution kernel of 1×1 and a depth of 64, and the output is a feature map with a length and width that are half of the binary image and a depth of 64. The subsequent fourth downsampling module, fifth downsampling module, and sixth downsampling module are similar to the first downsampling module, and all include two convolutional layers and a pooling layer, mainly for feature extraction. The difference is that the convolutional depths of the fourth downsampling module, fifth downsampling module, and sixth downsampling module are 128, 256, and 512 respectively, and the lengths and widths of the output feature maps are one-fourth, one-eighth, and one-sixteenth respectively. That is, the output of the sixth downsampling module is a feature map with a depth of 512 and a length and width that are one-sixteenth of the binary image, and it is used as the input of the first fully-connected linear layer. The first fully-connected linear layer consists of three linear layers, which are 1024, 512, and 64 respectively, linearly connecting the feature parameters output by the sixth downsampling module, and finally outputting the parameter of the abscissa of the symbol center.

[0101] More specifically, the code track rough decoding network extracts symbol center point information based on the code track binary image and the abscissa of the symbol center, and obtains the rough position information from the pre-constructed coding position database through the look-up table method according to the symbol center point information.

[0102] In the specific implementation process, the coding position database is constructed by sorting all the center point information of code elements calculated according to the M-sequence generation function of the grating scale, and the center point information of code elements is converted into hexadecimal information. The coarse position information is set according to the sorting. After directly reading binary information such as [1.0, 0.0, 1.0, 1.0, 1.0, 1.0, 0.0, 1.0, 0.0, 0.0, 0.0, 0.0, 1.0, 0.0, 1.0, 1.0] from the code track image and then converting the binary information into hexadecimal information, it can be docked with the coding position database. As Figure 5 shown, the coding position database includes binary information, hexadecimal information and coding position information. As long as the hexadecimal information is input into the coding position database, the coarse position information of the code track can be obtained by using the look-up table method. This is relatively fast, but there is an additional step of converting binary to hexadecimal. Binary information can also be directly input into the database, but it will reduce the speed to a certain extent.

[0103] More specifically, the loss function of the code track rough positioning network is as follows:

[0104]

[0105]

[0106]

[0107] Among them, is the loss function of the first segmentation module, l log is the loss function of the positioning center module, μ is the weight ratio of and l log , W is the image width, H is the image height, is the pixel value of each pixel of the image after being segmented by the first segmentation module, is the corresponding pixel value of the original binary map label, p(θ(x)) is the softmax loss function, and θ: x → {1,...., K} is the label value of the abscissa pixel point of the code track center.

[0108] In the specific implementation process, the information of the absolute grating scale image can be divided into coarse information and fine information. The coarse information is the coding information of the absolute grating scale, and the fine information is the data information of less than one code element at the boundary of the camera range. As Figure 6As shown. Since the grating size track information is obtained from the M-sequence coding, the information of each track is only related to the adjacent tracks and cannot be decoded intuitively, which brings inconvenience to obtaining the coarse information. Therefore, the coarse information is obtained by a phased method using the track coarse positioning decoding model, including three stages: rough segmentation of the track image, obtaining the center point information of each code element, and looking up the table in the database, so as to complete the acquisition of the coarse information. Among them, the rough segmentation of the track image is proposed based on the characteristic that there are large boundaries between code elements. After the segmentation is completed, the data characteristic of the image is obvious binary information, either 0 or 1, which is the premise of the feasibility of the track coarse decoding network. Obtaining the center point information of each code element is also based on the characteristics between the code elements of the grating scale image. Although the boundary is relatively blurred, the width of each code element only differs by a few pixels. Even rough segmentation can accurately obtain binary information from the center point of each code element. The center point information is as Figure 6 shown. When the binary information of the track is obtained, the binary information can be decoded through the pre-established coding position database to obtain the coding information of the absolute grating scale, that is, the coarse information.

[0109] Embodiment 3

[0110] A grating scale automatic positioning and decoding method based on deep learning includes the following steps:

[0111] S1: Construct a track coarse positioning decoding model and a track fine positioning decoding model based on deep learning,

[0112] The track coarse positioning decoding model includes a track coarse positioning network and a track coarse decoding network,

[0113] The track fine positioning decoding model includes a track fine positioning network and a track fine decoding network;

[0114] S2: Use the track coarse positioning network to obtain the track binary image and the horizontal coordinate of the center of the code element from the input grating scale image,

[0115] Use the track fine positioning network to obtain the end code element positioning binary image from the grating scale image;

[0116] S3: Use the track coarse decoding network to interpret the coarse position information of the grating scale image according to the track binary image and the horizontal coordinate of the center of the code element,

[0117] Use the track fine decoding network to interpret the fine position information of the grating scale image according to the end code element positioning binary image;

[0118] S4: Obtain the absolute position information as the automatic positioning and decoding result according to the code element width, coarse position information and fine position information of the grating scale image.

[0119] More specifically, such asFigure 7 As shown, the code track fine positioning network includes a super-resolution module and a second segmentation module;

[0120] Using the super-resolution module to perform feature extraction, non-linear mapping, and image reconstruction on the grating scale image, so as to obtain a clear high-resolution image;

[0121] Using the second segmentation module to obtain the end element positioning binary image from the high-resolution image;

[0122] The super-resolution module includes multiple convolutional layers;

[0123] In the super-resolution module, the input code track image passes through a feature extraction module composed of two convolutional layers to obtain a feature map with a depth of 64 and the same length and width as the input image. Among them, the convolutional kernel of the first convolutional layer of the feature extraction module is 9×9 and the depth is 64. In order to keep the size unchanged, the boundary will be supplemented. The convolutional kernel of the second convolutional layer is 1×1 and the depth is 64, and the size remains unchanged. The output feature map is used as the input of the non-linear mapping module. The non-linear module is composed of two convolutional layers with a convolutional kernel of 1×1 and a depth of 32. Therefore, there is no need to supplement the boundary. Then, through the image reconstruction module, a high-resolution image is obtained. The image reconstruction module is composed of a convolutional layer with a convolutional kernel of 5×5 and a depth of 3;

[0124] Feature extraction: The low-resolution image is blurred through binomial interpolation, and image features are extracted from it. The channel is 3, the convolutional kernel size is 9*9, and the number of convolutional kernels is 64;

[0125] F1(Y) = max(0, w1 * Y + B1)

[0126] w1 = c * 9 * 9 * 64

[0127] B1 = 64

[0128] Among them, F1(Y) is the feature map after preliminary feature extraction, w1 is the weight parameter, and B1 is the output channel.

[0129] Non-linear mapping: Map the low-resolution picture features to high resolution, the convolutional kernel size is 1*1, and the number of convolutional kernels is 32

[0130] F2(Y) = max(0, w2 * F1(Y) + B2)

[0131] w2 = 32 * 1 * 1 * 16

[0132] B2 = 32

[0133] Among them, F2(Y) is the feature map after non-linear mapping, w2 is the weight parameter, and B2 is the output channel. Image reconstruction: restore details to obtain a clear high-resolution image, and the convolution kernel size is 5*5;

[0134] F(Y) = w3 * F2(Y) + B3

[0135] w3 = 32 * 5 * 5 * c

[0136] B3 = c

[0137] Among them, F(Y) is the reconstructed high-resolution image, w3 is the weight parameter, and B3 is the output channel.

[0138] The second segmentation module includes a seventh downsampling module, an eighth downsampling module, a ninth downsampling module, a tenth downsampling module, a third convolution module, a fourth upsampling module, a fifth upsampling module, a sixth upsampling module, a seventh upsampling module, a second convolution layer, and a second normalization layer connected in sequence; among them, one branch output end of the seventh downsampling module is connected to the branch input end of the seventh upsampling module, one branch output end of the eighth downsampling module is connected to the branch input end of the sixth upsampling module, one branch output end of the ninth downsampling module is connected to the branch input end of the fifth upsampling module, and one branch output end of the tenth downsampling module is connected to the branch input end of the fourth upsampling module.

[0139] The second segmentation module is similar to the first segmentation module, both consisting of three modules: a downsampling module, an upsampling module, and a feature fusion module. The seventh downsampling module takes the high-resolution code track image as input, includes two convolutional layers and a pooling layer, and outputs a feature map with a depth of 64 and a length and width that are half of the high-resolution code track image. The eighth, ninth, and tenth downsampling modules are similar to the seventh downsampling module, except that the depths of the convolutional kernels are 128, 256, and 512 respectively, and due to the pooling layer, the size of the feature map will gradually decrease. The tenth downsampling module will output a feature map with a depth of 512 and a length and width that are one-sixteenth of the high-resolution code track image, which serves as the input to the third convolutional module. The third convolutional module consists of a convolutional layer with a 5×5 convolutional kernel and a depth of 1024, and pads the boundaries. The output is a feature map with a depth of 1024 and a length and width that are one-sixteenth of the high-resolution code track image, which serves as the input to the fourth upsampling module. The fourth upsampling module includes a transposed convolutional layer and two convolutional layers, and outputs a feature map with a depth of 512 and a length and width that are one-eighth of the high-resolution code track image. This feature map is added to the output point of the second convolutional layer of the tenth downsampling module for feature fusion, serving as the input to the fifth upsampling module. Similarly, the output feature map depths of the fifth, sixth, and seventh upsampling modules will gradually decrease, while the lengths and widths will gradually increase, and they are respectively fused with the tenth, ninth, and eighth downsampling modules. That is, the output of the seventh upsampling module is a feature map with a depth of 64 and a length and width that are the same as the high-resolution image. After passing through the second convolutional layer with a convolutional kernel size of 1×1 and a depth of 2, a probability map with a depth of 2 will be obtained. After passing through the softmax loss function of the second normalization layer, a binary image for end code element localization will be obtained.

[0140] More specifically, the code track fine decoding network includes an eleventh downsampling module, a twelfth downsampling module, a first depth downsampling module, a second depth downsampling module, a third depth downsampling module, and a second fully connected linear layer connected in sequence.

[0141] The eleventh downsampling module includes two convolutional layers with a 3×3 convolutional kernel and a depth of 64, and the pooling layer is a 2×2 regular reduction. The twelfth downsampling module is similar to the eleventh downsampling module, except that the depth is 128, and the output is a feature map with a depth of 128 and a length and width that are one-fourth of the end symbol positioning binary image, which is used as the input of the first depth downsampling module. The first depth downsampling module, the second depth downsampling module, and the third depth downsampling module all include four convolutional layers and one pooling layer. The first three convolutional layers of the first depth downsampling module have a 3×3 convolutional kernel and a depth of 256, the fourth convolutional layer has a 1×1 convolutional kernel and a depth of 256, and the pooling layer is a 2×2 regular reduction. The second depth downsampling module and the third depth downsampling module are similar to the first depth downsampling module, except that the depths are 512 and 1024 respectively. That is, the output of the third depth downsampling module is a feature map with a depth of 1024 and a length and width that are one-thirty-second of the end symbol positioning binary image, which is used as the input of the second fully connected linear layer. The second fully connected linear layer consists of five linear layers, namely 4096, 1024, 256, 64, and 16, which linearly connect the feature parameters output by the third depth downsampling module, and finally output the parameters of the preliminary fine position information.

[0142] The code track fine decoding network uses the Mean Squared Error (MSE) acc , in order to link the image-data. Specifically, the probability and the position information are linked, that is, the coverage area size of the last symbol is converted into a probability. The more the area is covered, the greater the probability, and the smaller the area is covered, the smaller the probability.

[0143] d = p * m

[0144] where d is the position information of the fine decoding, p is the probability of the corresponding code track image, and m is the length of one symbol.

[0145] More specifically, the loss function of the code track fine positioning decoding model is as follows:

[0146]

[0147]

[0148]

[0149]

[0150]

[0151]

[0152]

[0153]

[0154] Among them, is the loss function of the code track fine positioning network, l acc is the loss function of the code track fine decoding network, θ is and l acc 's weight ratio, is the loss function of the super-resolution module, l elog is the loss function of the second segmentation module, is the mean square error loss function, is the edge loss function, ε is and 's weight ratio, W is the width of the image, H is the height of the image, is the pixel of the image after super-resolution, is the pixel of the original image, E x,y is the result of using Canny edge detection, p(θ(x)) is the softmax loss function, θ: x → {1,...., K} is the label value of the abscissa pixel point of the code track center, w: C ∈ R is the weight value of the code track boundary pixel point, the purpose is to give higher weights to the pixels close to the boundary points in the image, w c : C ∈ R is the weight value for balancing the class ratio, d1: C ∈ R is the distance from the code element pixel point to the nearest pixel point, d2: C ∈ R is the distance from the code element pixel point to the second nearest pixel point, w0 and σ are different constant values, in this embodiment, w0 = 5, σ = 2, is the position information label, is the output position information, x is the abscissa of the pixel point, y is the ordinate of the pixel point.

[0155] In the specific implementation process, the end symbol information of the grating scale image is segmented with high precision by the code track fine positioning network, so that the subsequent code track fine decoding network can accurately interpret it. In the process of fine positioning and segmentation, there are two difficulties. The first is that the boundary between black and white stripes is blurred, and the second is how to segment the same-color stripes. The second point is more difficult than the first point because there is no boundary information, and many traditional methods only consider the first one. To solve these two difficulties, this embodiment incorporates the super-resolution function, improves the resolution of the original image to the sub-pixel level, expands the blurred stripe boundary, and enables more accurate segmentation, thus initially solving the first problem. To solve the second problem, not only the boundary but also the information of the entire code track needs to be considered. Although the same-color stripes cannot provide information at the end boundary, there is rich information provided on the entire code track, such as the start symbol information or adjacent symbol information, which can all be utilized. Moreover, in the process of segmenting the boundary of different-color stripes, the overall information can also provide useful information. After the end symbol is segmented, it will be input into the code track fine decoding network. To enable the code track fine decoding network to complete the regression prediction of the image-data, this embodiment links the probability with the image to obtain the probability of the image as the output of the data, and compresses the information of one symbol between 0 and 1.

[0156] More specifically, the code track fine positioning and decoding model further includes a feedback compensation network;

[0157] The feedback compensation network is used to determine whether the end symbol positioning binary image classification is a boundary symbol image. When it is determined to be a boundary symbol image, 0 is used as the output of the code track fine decoding network. When it is determined not to be a boundary symbol image, the actual output of the code track fine decoding network is used as the standard;

[0158] The feedback compensation network includes a second convolutional pooling module, a third convolutional pooling module, a fourth convolutional pooling module, a fifth convolutional pooling module, a third fully connected layer, and a classification module connected in sequence.

[0159] The binary image with the end symbol located is used as the input of the second convolutional pooling module. The second convolutional pooling module includes a convolutional layer with a convolutional kernel size of 3×3 and a depth of 64, and a 2×2 uniform regular pooling layer. Padding is performed on the boundaries, and the output is a feature map with a depth of 64 and a length and width that are half of the binary image, which is used as the input of the third convolutional pooling module. The third convolutional pooling module, the fourth convolutional pooling module, and the fifth convolutional pooling module all include a convolutional layer and a pooling layer. The differences are that the convolutional kernels are 5×5, 7×7, and 9×9 respectively, and the depths are 128, 512, and 1024 respectively. That is, the output of the fifth convolutional pooling module is a feature map with a depth of 1024, which is used as the input of the third fully connected layer. The third fully connected layer consists of three linear layers, namely 1024, 128, and 32, which linearly connect the feature parameters output by the fifth convolutional pooling module, and finally output 0 and 1 parameters to enter the classification module for discrimination.

[0160] More specifically, the loss function l of the feedback compensation network com is as follows:

[0161] l com = -(y1 log p(x1) + (1 - y1) log(1 - p(x1))

[0162] where x1 is the output data, y1 is the true label, and p(x1) is the probability function.

[0163] In the specific implementation process, when the end symbol of the code track is at or near a complete symbol, due to the edge effect, the information provided to the network is insufficient, and misjudgment is likely to occur. Therefore, in this embodiment, a feedback compensation network is added to the code track fine positioning network. It mainly classifies the code track information where the end symbol of the code track is at or near a complete symbol as one category, and the rest are regarded as another category. It is a classification network. This avoids misjudgment due to the edge effect and greatly improves the measurement accuracy of the absolute grating scale.

[0164] More specifically, the total loss function l is set as follows: total as follows:

[0165]

[0166] where is the loss function of the first segmentation module, l log is the loss function of the positioning center module, l acc is the loss function of the code track fine decoding network, l com is the loss function of the feedback compensation network, is the loss function of the code track fine positioning network, ρ is and the weight value between, σ is l log and The weight value between them, τ is l acc and The weight value between them, is l com and The weight value between them.

[0167] In the specific implementation process, this embodiment adopts the method of full-network self-training. After pre-training each module in each network, the modules are uniformly trained, and the entire network is fine-tuned, including but not limited to freezing several layers of the network. Through fine-tuning, the most suitable loss weight parameter ratio for each module in the full network is found.

[0168] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. An automatic positioning and decoding method for grating scales based on deep learning, characterized in that, Including the following steps: S1: Construct a coarse code track positioning and decoding model and a fine code track positioning and decoding model based on deep learning. The coarse code track positioning and decoding model includes a coarse code track positioning network and a coarse code track decoding network. The coarse code track positioning network includes a first segmentation module and a positioning center module. Use the first segmentation module to obtain a binary image of the code track from the grating scale image. Use the positioning center module to obtain the abscissa of the code element center according to the binary image of the code track. The coarse code track decoding network extracts the information of the code element center point according to the binary image of the code track and the abscissa of the code element center, and obtains the coarse position information from the pre-constructed coding position database by means of a look-up table method according to the information of the code element center point. The fine code track positioning and decoding model includes a fine code track positioning network and a fine code track decoding network. The fine code track positioning network includes a super-resolution module and a second segmentation module. Use the super-resolution module to perform feature extraction, non-linear mapping and image reconstruction on the grating scale image, so as to obtain a clear high-resolution image. Use the second segmentation module to obtain a binary image for positioning the end code element from the high-resolution image. S2: Use the coarse code track positioning network to obtain a binary image of the code track and the abscissa of the code element center from the input grating scale image. Use the fine code track positioning network to obtain a binary image for positioning the end code element from the grating scale image. S3: Use the coarse code track decoding network to interpret the coarse position information of the grating scale image according to the binary image of the code track and the abscissa of the code element center. Use the fine code track decoding network to interpret the fine position information of the grating scale image according to the binary image for positioning the end code element. S4: Obtain the absolute position information as the automatic positioning and decoding result according to the code element width, coarse position information and fine position information of the grating scale image.

2. The automatic positioning and decoding method of a grating scale based on deep learning according to claim 1, characterized in that, The first segmentation module includes a first downsampling module, a second downsampling module, a third downsampling module, a first convolutional module, a first upsampling module, a second upsampling module, a third upsampling module, a first convolutional layer and a first normalization layer connected in sequence; wherein, a branch output end of the first downsampling module is connected to a branch input end of the third upsampling module, a branch output end of the second downsampling module is connected to a branch input end of the second upsampling module, and a branch output end of the third downsampling module is connected to a branch input end of the first upsampling module. The positioning center module includes a first convolutional pooling module, a second convolutional module, a fourth downsampling module, a fifth downsampling module, a sixth downsampling module and a first fully connected linear layer connected in sequence; wherein, the output of the convolutional pooling module also includes local response normalization processing before inputting into the second convolutional module.

3. A grating scale automatic positioning and decoding method based on deep learning according to claim 1, characterized in that, The super-resolution module includes a plurality of convolutional layers. The second segmentation module includes a seventh downsampling module, an eighth downsampling module, a ninth downsampling module, a tenth downsampling module, a third convolution module, a fourth upsampling module, a fifth upsampling module, a sixth upsampling module, a seventh upsampling module, a second convolutional layer, and a second normalization layer, which are connected in sequence; among them, a branch output end of the seventh downsampling module is connected to a branch input end of the seventh upsampling module, a branch output end of the eighth downsampling module is connected to a branch input end of the sixth upsampling module, a branch output end of the ninth downsampling module is connected to a branch input end of the fifth upsampling module, and a branch output end of the tenth downsampling module is connected to a branch input end of the fourth upsampling module.

4. A grating scale automatic positioning and decoding method based on deep learning according to claim 1, characterized in that The code track fine decoding network includes an eleventh downsampling module, a twelfth downsampling module, a first depth downsampling module, a second depth downsampling module, a third depth downsampling module, and a second fully connected linear layer, which are connected in sequence.

5. A grating scale automatic positioning and decoding method based on deep learning according to claim 1, characterized in that, The code track fine positioning decoding model further includes a feedback compensation network; The feedback compensation network is used to determine whether the binary image classification of the tail end symbol positioning is a boundary symbol image. When it is determined to be a boundary symbol image, 0 is used as the output of the code track fine decoding network. When it is determined not to be a boundary symbol image, the actual output of the code track fine decoding network is used as the standard; The feedback compensation network includes a second convolutional pooling module, a third convolutional pooling module, a fourth convolutional pooling module, a fifth convolutional pooling module, a third fully connected layer, and a classification module, which are connected in sequence.

6. The automatic positioning and decoding method of a grating scale based on deep learning according to claim 2, characterized in that, The loss function of the code channel rough positioning network is as follows: Among them, is the loss function of the first segmentation module, is the loss function of the positioning center module, is and 's weight ratio, is the image width, is the image height, is the pixel value of each pixel of the image after being segmented by the first segmentation module, is the corresponding pixel value of the original binary map label, is the softmax loss function, (x) is the label value of the abscissa pixel point of the code track center, x is the abscissa of the pixel point, and y is the ordinate of the pixel point.

7. A grating scale automatic positioning and decoding method based on deep learning according to claim 3, characterized in that The loss function of the code channel fine positioning decoding model is as follows: Among them, is the loss function of the code track fine positioning network, is the loss function of the code track fine decoding network, is and 's weight ratio, is the loss function of the super-resolution module, is the loss function of the second segmentation module, is the mean square error loss function, is the edge loss function, is and 's weight ratio, is the width of the image, is the height of the image, is the pixel of the image after super-resolution, is the pixel of the original image, is the result of using Canny edge detection, is the softmax loss function, is the label value of the abscissa pixel point of the code track center, is the weight value of the code track boundary pixel point, is the weight value for balancing the class ratio, is the distance from the code element pixel point to the nearest pixel point, is the distance from the code element pixel point to the second nearest pixel point, and are different constant values, is the position information label, is the output position information, x is the abscissa of the pixel point, and y is the ordinate of the pixel point.

8. A grating scale automatic positioning and decoding method based on deep learning according to claim 5, characterized in that The loss function of the feedback compensation network is as follows: Among them, is the output data, is the true label, is the probability function.

9. A grating scale automatic positioning and decoding method based on deep learning according to claim 5, characterized in that, Set the total loss function as follows: Among them, is the loss function of the first segmentation module, is the loss function of the positioning center module, is the loss function of the code track fine decoding network, is the loss function of the feedback compensation network, is the loss function of the code track fine positioning network, is and the weight value between them, is and the weight value between them, is and the weight value between them, is and the weight value between them.

Citation Information

Patent Citations

  • Absolute position encoding and decoding method for absolute position displacement sensor gratings

    CN109341545A

  • Pseudo-random code channel grating ruler and reading method therefor

    WO2019192196A1