Method and apparatus for key point detection of reflective two-dimensional codes based on deep neural network

The feature map of the QR code image through the deep neural network model solves the detection problem of QR code key point detection under light and angle changes, improves the detection accuracy and success rate, and reduces the cost.

CN114565776BActive Publication Date: 2025-07-08XIAMEN STAR SMART TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210084942.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-25
Publication Date
2025-07-08
Estimated Expiration
2042-01-25

AI Technical Summary

Technical Problem

The existing QR code key point detection algorithm has poor detection success rate and stability under light and angle changes, and is relatively expensive.

Method used

The deep neural network model is used to feature map the QR code image. Through the construction of the training data set and the model training, more key point coordinates are output to improve detection accuracy, including a combination of two-dimensional convolutional layer, batch normalization layer, linear rectifier function layer, average pooling layer and fully connected layer.

Benefits of technology

It improves the anti-interference ability to light changes and angle changes, increases the accuracy and success rate of key point detection, and reduces equipment costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114565776B_ABST
    Figure CN114565776B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for detecting key points of a reflective two-dimensional code based on a deep neural network. By obtaining a large number of two-dimensional code images annotated with the coordinates of a set number of key points, and then calculating the Euler angles of each two-dimensional code image respectively; performing normalization and standardization processing on the two-dimensional code images; constructing a training data set based on the two-dimensional code images annotated with the coordinates of a set number of key points, Euler angles and the processed two-dimensional code images; constructing a deep neural network model, and then using the training data set to train the deep neural network model to obtain a key point coordinate detection model; inputting the processed image to be detected into the model to obtain the key point coordinate positions. The method and device for detecting key points of a reflective two-dimensional code based on a deep neural network provided by the present invention calculate a feature map for the entire image through the deep neural network, so as to realize two-dimensional code target detection with strong anti-lighting change and angle change capabilities and low cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of two-dimensional code key point detection, and particularly to a method and device for detecting key points of a reflective two-dimensional code based on a deep neural network. Background Art

[0002] Currently, the general steps of two-dimensional code parsing technology are as follows:

[0003] (1) Two-dimensional code target detection;

[0004] (2) Two-dimensional code key point detection;

[0005] (3) Two-dimensional code image calibration;

[0006] (4) Pixel map digital matrix conversion.

[0007] Among them, the two-dimensional code key point detection technology is usually based on the Scan Line Algorithm. This algorithm assumes that there is basically no interference from light changes and angle changes in the two-dimensional code. However, in actual scenarios, the two-dimensional code images captured by the camera are often affected by sunlight, light, and shooting angles, and are often contaminated by strong light interference and angle changes, seriously affecting the detection success rate and stability of the scan line algorithm. To solve this problem, the existing technology filters reflections by adding special optical lenses or barcode readers to the camera, but this will increase the equipment cost at the same time.

[0008] Currently, there is still a lack of a two-dimensional code target detection algorithm with strong resistance to light changes and angle changes and low cost. Summary of the Invention

[0009] The technical problem to be solved by the present invention is to provide a method and device for detecting key points of a reflective two-dimensional code based on a deep neural network, which calculates a feature map for the entire image through the deep neural network, so as to realize two-dimensional code target detection with strong resistance to light changes and angle changes and low cost.

[0010] In a first aspect, the present invention provides a method for detecting key points of a reflective two-dimensional code based on a deep neural network, including: a training data set construction process, a deep neural network model construction and training process, and a two-dimensional code key point detection process;

[0011] The training data set construction process includes: obtaining a large number of two-dimensional code images marked with the coordinates of a set number of key points, and then calculating the Euler angles of each two-dimensional code image respectively; normalizing and standardizing the two-dimensional code images; constructing a training data set based on the two-dimensional code images marked with the coordinates of a set number of key points, Euler angles, and the processed two-dimensional code images;

[0012] The process of constructing and training the deep neural network model includes: constructing a deep neural network model, where the number of output vectors of the deep neural network model is twice the number of key points of each QR code image in the training dataset; then using the training dataset to train the deep neural network model to obtain a key point coordinate detection model.

[0013] The QR code key point detection process includes: obtaining the area of the QR code to be detected in the image to be detected, then performing normalization and standardization processing and inputting it into the key point coordinate detection model to obtain the key point coordinate positions of the QR code to be detected.

[0014] Further, in the process of constructing the training dataset, the specific method of marking the key point positions on the QR code image is as follows: Let the positioning pattern in the upper left corner of the QR code be the first positioning pattern, the positioning pattern in the upper right corner be the second positioning pattern, and the positioning pattern in the lower left corner be the third positioning pattern. Take the four vertices of the outermost square in each of the three positioning patterns, the vertex in the lower right corner of the QR code image, and the intersection points obtained by extending the inner side of the second positioning pattern and the upper side of the third positioning pattern respectively, for a total of 14 key points.

[0015] Further, the specific method for calculating the Euler angles is as follows: Based on the two-dimensional image coordinates and three-dimensional coordinates in the world coordinate system of all key points in each QR code image, calculate the Euler angles of the QR code image relative to the camera as the origin based on the PNP method.

[0016] Further, the deep neural network model includes a two-dimensional convolutional layer, a batch normalization layer, a rectified linear unit layer, an average pooling layer, a fully connected layer, and an activation function layer.

[0017] In a second aspect, the present invention provides a device for detecting key points of a reflective QR code based on a deep neural network, including: a training dataset module, a deep neural network model module, and a QR code key point detection module;

[0018] The training dataset module is used to obtain a large number of QR code images marked with the coordinates of a set number of key points, and then calculate the Euler angles of each QR code image respectively; perform normalization and standardization processing on the QR code images; construct a training dataset based on the QR code images marked with the coordinates of a set number of key points, Euler angles, and the processed QR code images;

[0019] The deep neural network model module is used to construct a deep neural network model, where the number of output vectors of the deep neural network model is twice the number of key points of each QR code image in the training dataset; then use the training dataset to train the deep neural network model to obtain a key point coordinate detection model;

[0020] The QR code key point detection module is used to obtain the QR code area to be detected in the image to be detected, and then perform normalization and standardization processing before inputting it into the key point coordinate detection model to obtain the key point coordinate positions of the QR code to be detected.

[0021] Further, in the training dataset module, the specific positions of the key points marked on the QR code image are as follows: Let the positioning pattern in the upper left corner of the QR code be the first positioning pattern, the positioning pattern in the upper right corner be the second positioning pattern, and the positioning pattern in the lower left corner be the third positioning pattern. Take the four vertices of the outermost square in each of the three positioning patterns, the vertex in the lower right corner of the QR code image, and the intersection points where the inner side of the second positioning pattern and the upper side of the third positioning pattern extend respectively, for a total of 14 key points.

[0022] Further, in the training dataset module, the specific calculation method of the Euler angle is as follows: Based on the two-dimensional image coordinates and three-dimensional coordinates in the world coordinate system of all key points in each QR code image, calculate the Euler angle of the QR code image relative to the camera as the origin based on the PNP method.

[0023] Further, the deep neural network model includes a two-dimensional convolutional layer, a batch normalization layer, a rectified linear unit layer, an average pooling layer, a fully connected layer, and an activation function layer.

[0024] The technical solutions provided in the embodiments of the present invention have the following technical effects or advantages:

[0025] 1. Using a deep neural network model to detect the key points of the QR code, transforming the shape detection using the positioning pattern by the traditional algorithm into feature vector detection, calculating the feature map of the entire image based on the convolutional shared weights, and using both local information and global information. Compared with the traditional algorithm, the anti-interference ability against local reflective pollution and angle transformation is improved.

[0026] 2. Compared with the traditional algorithm that can only output 3 points, the number of output control points is increased to 14, improving the positioning accuracy, and thus greatly improving the success rate of subsequent processing.

[0027] The above description is only an overview of the technical solutions of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the following specifically gives the specific embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The present invention will be further described below with reference to the accompanying drawings in conjunction with embodiments.

[0029] Figure 1 It is the flowchart of the method in Embodiment 1 of the present invention;

[0030] Figure 2 It is a schematic structural diagram of the deep neural network in the first embodiment of the present invention;

[0031] Figure 3 It is a schematic diagram of the detection result of the key points of the reflective two-dimensional code in the first embodiment of the present invention;

[0032] Figure 4 It is a schematic structural diagram of the device in the second embodiment of the present invention. Specific implementation manners

[0033] In the embodiments of the present application, by providing a method and a device for detecting key points of a reflective two-dimensional code based on a deep neural network, a feature map is calculated for the entire image through the deep neural network, so as to achieve strong anti-lighting change and angle change capabilities and low-cost two-dimensional code target detection.

[0034] The overall idea of the technical solution in the embodiments of the present application is as follows:

[0035] A two-dimensional code key point detection algorithm that is lightweight, fast, and has strong stability against reflective interference and angle changes is proposed, and the steps are as follows:

[0036] 1. Train a deep neural network model using a large amount of two-dimensional code data

[0037] 1.1 Perform data augmentation processing on the input training data;

[0038] 1.2 Calculate the Euler angle of the two-dimensional code relative to the camera in three-dimensional space according to the labeled two-dimensional code key point data;

[0039] 1.3 Normalize and standardize the two-dimensional code image data;

[0040] 1.4 Construct two-dimensional code label data based on the two-dimensional code key point coordinates and Euler angles;

[0041] 1.5 Use the two-dimensional code image data and label data to train the deep neural network model.

[0042] 2. Two-dimensional code target key point detection

[0043] 2.1 SSD two-dimensional code target detection, (obtain the position of the two-dimensional code in the entire image);

[0044] 2.2 Crop the two-dimensional code ROI (Region of Interest) area in the entire image to obtain two-dimensional code picture data Im;

[0045] 2.3 Normalize and standardize the two-dimensional code picture data Im;

[0046] 2.4 Use the trained neural network model to predict the positions of key points on the QR code image data Im.

[0047] Example 1

[0048] This example provides a method for detecting key points of a reflective QR code based on a deep neural network. As Figure 1 shown, it includes;

[0049] The training dataset construction process, the deep neural network model construction and training process, and the QR code key point detection process;

[0050] The training dataset construction process includes: obtaining a large number of QR code images annotated with the coordinates of a set number of key points, and then calculating the Euler angles of each QR code image respectively; performing normalization and standardization processing on the QR code images; constructing a training dataset based on the QR code images annotated with the coordinates of a set number of key points, Euler angles, and the processed QR code images;

[0051] The deep neural network model construction and training process includes: constructing a deep neural network model, where the number of output vectors of the deep neural network model is twice the number of key points of each QR code image in the training dataset; then using the training dataset to train the deep neural network model to obtain a key point coordinate detection model.

[0052] The QR code key point detection process includes: obtaining the QR code area to be detected in the image to be detected, and then performing normalization and standardization processing and inputting it into the key point coordinate detection model to obtain the key point coordinate positions of the QR code to be detected.

[0053] A specific implementation of this example:

[0054] In the training dataset construction process, the specific positions of the key points marked on the QR code image are as follows: Let the positioning pattern in the upper left corner of the QR code be the first positioning pattern, the positioning pattern in the upper right corner be the second positioning pattern, and the positioning pattern in the lower left corner be the third positioning pattern. Take the four vertices of the outermost square in each of the three positioning patterns, the vertex in the lower right corner of the QR code image, and the intersection points where the inner side of the second positioning pattern and the upper side of the third positioning pattern extend respectively, for a total of 14 key points, as Figure 3 shown.

[0055] To expand the training data, data augmentation processing is performed on the training data. A random number generator can be used to randomly generate attributes such as image rotation angle, perspective angle, translation amount, etc., calculate the corresponding transformation matrix, and apply this transformation matrix to the QR code image.

[0056] The specific method for calculating the Euler angles is as follows: based on the two-dimensional image coordinates of all key points in each QR code image and the three-dimensional coordinates in the world coordinate system, the Euler angles of the QR code image relative to the camera as the origin are calculated using the PNP (Perspective-n-Point, a general algorithm) method.

[0057] During the construction and training process of the deep neural network model, the deep neural network model includes a two-dimensional convolutional layer (Conv2d layer), a batch normalization layer (Batch Norm Layer), a rectified linear unit layer (ReLu layer), an average pooling layer (average pooling layer), a fully connected layer (fully connect payer), and an activation function layer (sigmoid layer).

[0058] The deep neural network model is designed with a 4-layer neural network, and the input and output of each layer are as follows:

[0059] Layer1: Input a normalized and standardized 112×112 QR code image, and output 64 feature maps of 56×56 (feature map);

[0060] Layer2: Input the 64 feature maps of Layer1 and output 64 feature maps of 28x28;

[0061] Layer3: Input the 64 feature maps of Layer2 and output 16 feature maps of 14×14;

[0062] Layer4:

[0063] (1) Perform 14×14 average pooling on the 16 feature maps output by Layer3 and apply flattening (Flattern) to obtain a feature vector of length 16;

[0064] (2) Apply convolution to the 16 feature maps of 14×14 feature maps output by Layer3 to obtain 32 feature maps of 7×7, perform 7×7 average pooling and flattening to obtain a feature vector of length 32;

[0065] (3) Apply convolution with kernel = 7 to the 32 feature vectors of 7×7 and flatten to obtain a feature vector of length 128;

[0066] (4) Concatenate the above feature vectors of 16 + 32 + 128 to obtain a feature vector of length 176;

[0067] (5) Apply a fully connected layer and a Sigmoid layer to the feature vector of length 176 to obtain a vector with a final length of 28, which is the coordinate positions of the 14 key points of the QR code.

[0068] During the training process, use the deep neural network model to perform forward inference on the training data to obtain the QR code key point and Euler angle prediction results, calculate the mean square error (L2 Loss) between the prediction results and the training data, and perform backpropagation (BP, BackProPagation) on the L2 Loss to optimize the weights of the deep neural network model; repeat the above steps 1×10 4 times to obtain a trained deep neural network model, that is, a key point coordinate detection model.

[0069] The QR code key point detection process specifically includes:

[0070] 1. Use the SSD object detection model to obtain the position of the QR code in a large image, represented by a rectangle (x, y, h, w), where x and y are the coordinates of the center point of the matrix, and h and w are the length and width of the rectangle respectively;

[0071] 2. According to the rectangle (x, y, h, w) obtained in the previous step, intercept the QR code image data Im in the large image;

[0072] 3. Normalize and standardize the QR code image data Im obtained in the previous step;

[0073] The normalization method is:

[0074]

[0075] The standardization method is:

[0076]

[0077] where I ij is the pixel value at the QR code image position (i, j), min(I ij ) is the minimum pixel value among all pixels in the image, max(I ij ) is the maximum pixel value among all pixels in the image, μ is the average value of all pixel values of the entire QR code image, σ is the pixel standard deviation, and N is the number of pixels.

[0078] 4. Input the normalized and standardized image data into the trained deep neural network model, that is, the key point coordinate detection model, as Figure 2 shown, predict the key point positions on the QR code image data Im, and the obtained key point results are as Figure 3 shown.

[0079] Based on the same inventive concept, this application also provides an apparatus corresponding to the method in Embodiment 1. For details, see Embodiment 2.

[0080] Embodiment 2

[0081] In this embodiment, an apparatus for detecting key points of a reflective two-dimensional code based on a deep neural network is provided. As Figure 4 shown, it includes: a training data set module, a deep neural network model module, and a two-dimensional code key point detection module;

[0082] The training data set module is used to obtain a large number of two-dimensional code pictures marked with the coordinates of a set number of key points, and then calculate the Euler angles of each two-dimensional code picture respectively; perform normalization and standardization processing on the two-dimensional code pictures; construct a training data set based on the two-dimensional code pictures marked with the coordinates of a set number of key points, Euler angles, and the processed two-dimensional code pictures;

[0083] The deep neural network model module is used to construct a deep neural network model, and the number of output vectors of the deep neural network model is twice the number of key points of each two-dimensional code picture in the training data set; then use the training data set to train the deep neural network model to obtain a key point coordinate detection model;

[0084] The two-dimensional code key point detection module is used to obtain the area of the two-dimensional code to be detected in the picture to be detected, and then after normalization and standardization processing, input it into the key point coordinate detection model to obtain the key point coordinate positions of the two-dimensional code to be detected.

[0085] Preferably, in the training data set module, the specific position of marking key points on the two-dimensional code picture is as follows: Let the positioning pattern in the upper left corner of the two-dimensional code be the first positioning pattern, the positioning pattern in the upper right corner be the second positioning pattern, and the positioning pattern in the lower left corner be the third positioning pattern. Take the four vertices of the outermost square in each of the three positioning patterns, the vertex in the lower right corner of the two-dimensional code image, and the intersection points of the inner side of the second positioning pattern and the upper side of the third positioning pattern extending respectively, for a total of 14 key points.

[0086] Specifically, in the training data set module, the calculation method of the Euler angle is specifically as follows: Based on the two-dimensional image coordinates and three-dimensional coordinates in the world coordinate system of all key points in each two-dimensional code picture, calculate the Euler angle of the two-dimensional code picture relative to the camera as the origin based on the PNP method.

[0087] Preferably, the deep neural network model includes a two-dimensional convolutional layer, a batch normalization layer, a rectified linear unit layer, an average pooling layer, a fully connected layer, and an activation function layer.

[0088] Since the device introduced in the second embodiment of the present invention is the device adopted for implementing the method of the first embodiment of the present invention, based on the method introduced in the first embodiment of the present invention, those skilled in the art can understand the specific structure and variations of the device, so it will not be elaborated here. Any device adopted for the method of the first embodiment of the present invention falls within the scope of protection of the present invention.

[0089] The embodiment of the present application uses a deep neural network model to detect the key points of the two-dimensional code, converts the shape detection using the positioning pattern by the traditional algorithm into feature vector detection, calculates the feature map for the entire image based on the convolutional shared weights, and utilizes both local information and global information. Compared with the traditional algorithm, it improves the anti-interference ability against local reflective pollution and angle transformation; compared with the traditional algorithm that can only output 3 points, the number of output control points is increased to 14, improving the positioning accuracy, thereby greatly improving the success rate of subsequent processing.

[0090] Although the specific implementation manners of the present invention have been described above, those skilled in the art should understand that the specific embodiments we described are illustrative rather than used to limit the scope of the present invention. Equivalent modifications and variations made by those skilled in the art in accordance with the spirit of the present invention should be covered by the scope of protection of the claims of the present invention.

Claims

1. A method for detecting key points of a reflective two-dimensional code based on a deep neural network, characterized in that, Including: The training dataset construction process, the deep neural network model construction and training process, and the QR code key point detection process; The training dataset construction process includes: obtaining a large number of QR code images labeled with the coordinates of a set number of key points, and then respectively calculating the Euler angles of each QR code image relative to the camera in three-dimensional space according to the labeled QR code key point data; normalizing and standardizing the QR code images; constructing a training dataset based on the QR code images labeled with a set number of key point coordinates, Euler angles, and the processed QR code images; The deep neural network model construction and training process includes: constructing a deep neural network model, where the number of output vectors of the deep neural network model is twice the number of key points of each QR code image in the training dataset, and every two vectors correspond to a key point coordinate; then using the training dataset to train the deep neural network model to obtain a key point coordinate detection model; The QR code key point detection process includes: obtaining the QR code area to be detected in the image to be detected, and then inputting it into the key point coordinate detection model after normalization and standardization to obtain output vectors, so as to obtain the key point coordinate positions of the QR code to be detected.

2. The method according to claim 1, wherein: In the training dataset construction process, the specific positions of the key points marked on the QR code image are as follows: Let the positioning pattern in the upper left corner of the QR code be the first positioning pattern, the positioning pattern in the upper right corner be the second positioning pattern, and the positioning pattern in the lower left corner be the third positioning pattern. Take the four vertices of the outermost square in each of the three positioning patterns, the vertex in the lower right corner of the QR code image, and the intersection points of the inner side of the second positioning pattern and the extended upper side of the third positioning pattern, for a total of 14 key points.

3. The method according to claim 1, wherein: The specific calculation method of the Euler angle is to calculate the Euler angle of the QR code image relative to the camera as the origin based on the two-dimensional image coordinates and three-dimensional coordinates in the world coordinate system of all key points in each QR code image, using the PNP method.

4. The method according to claim 1, characterized in that: The deep neural network model includes a two-dimensional convolutional layer, a batch normalization layer, a rectified linear unit layer, an average pooling layer, a fully connected layer, and an activation function layer.

5. An apparatus for detecting key points of a reflective two-dimensional code based on a deep neural network, characterized in that, Including: A training dataset module, a deep neural network model module, and a QR code key point detection module; The training dataset module is used to obtain a large number of QR code images labeled with the coordinates of a set number of key points, and then respectively calculate the Euler angles of each QR code image relative to the camera in three-dimensional space according to the labeled QR code key point data; normalizing and standardizing the QR code images; constructing a training dataset based on the QR code images labeled with a set number of key point coordinates, Euler angles, and the processed QR code images; The deep neural network model module is used to construct a deep neural network model, where the number of output vectors of the deep neural network model is twice the number of key points of each QR code image in the training dataset, and every two vectors correspond to a key point coordinate; then using the training dataset to train the deep neural network model to obtain a key point coordinate detection model; The QR code key point detection module is used to obtain the QR code area to be detected in the image to be detected, and then after normalization and standardization processing, it is input into the key point coordinate detection model to obtain the key point coordinate positions of the QR code to be detected.

6. The device according to claim 5, characterized in that: In the training data set module, the specific key point positions marked on the QR code image are as follows: Let the positioning pattern in the upper left corner of the QR code be the first positioning pattern, the positioning pattern in the upper right corner be the second positioning pattern, and the positioning pattern in the lower left corner be the third positioning pattern. Four vertices of the outermost square are taken from each of the three positioning patterns, the vertex in the lower right corner of the QR code image, and the intersection points of the inner side of the second positioning pattern and the extended upper side of the third positioning pattern, totaling 14 key points.

7. The device according to claim 5, characterized in that: In the training data set module, the specific calculation method of the Euler angle is as follows: Based on the two-dimensional image coordinates and three-dimensional coordinates in the world coordinate system of all key points in each QR code image, the Euler angle of the QR code image relative to the camera as the origin is calculated based on the PNP method.

8. The device according to claim 5, characterized in that: The deep neural network model includes a two-dimensional convolutional layer, a batch normalization layer, a rectified linear unit layer, an average pooling layer, a fully connected layer, and an activation function layer.

Citation Information

Patent Citations

  • Method and system for identifying two-dimensional code in picture

    CN110321750A

  • Bill detection method and device based on deep learning, equipment and storage medium

    CN110807455A