Two-dimensional code recognition method and device, equipment and storage medium

By aligning correction and multi-channel image data processing of the scanned video frames, and dot matrix prediction combined with the QR code recognition model, the problem of low recognition success rate of QR code in fuzzy situations is solved, and higher recognition accuracy is achieved.

CN120218097APending Publication Date: 2025-06-27SF TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311831255.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

During the QR code recognition process, the QR code recognition success rate is low in the case of blur caused by the movement of the QR code or the jitter of the user terminal.

Method used

By obtaining the first video frame containing the QR code from the scanned video, aligning and correcting the relevant video frame set based on the frame, multi-channel image data is generated, and a QR code recognition model is called to predict the QR code dot matrix to improve the recognition effect.

Benefits of technology

By fusing the information in the multi-frame QR code image in the video stream, the information loss caused by blurring between different frames is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218097A_ABST
    Figure CN120218097A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a two-dimensional code recognition method and device, equipment and a storage medium. The method comprises the steps that a first video frame containing a two-dimensional code is acquired from a video obtained through scanning; based on the two-dimensional code in the first video frame, performing alignment correction on each video frame in a video frame set related to the first video frame to obtain an image feature map sequence, the image feature map sequence corresponding to the first video frame and each video frame in the video frame set; generating multi-channel image data according to the image feature map sequence; calling a two-dimensional code recognition model to carry out two-dimensional code dot matrix prediction on the multi-channel image data to obtain a two-dimensional code dot matrix prediction result; and identifying the two-dimensional code dot matrix prediction result to obtain a two-dimensional code identification result. Information in multiple frames of images in the video stream is fused, information loss caused by blurring between different frames is bridged, and the recognition effect of the two-dimensional code is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of image processing, and in particular, to a method, device, equipment and storage medium for QR code recognition. Background Art

[0002] A QR code image is an image that records information through black and white grids arranged in an orderly two-dimensional manner. Users identify the QR code image by scanning it to decode and obtain the information recorded by the QR code.

[0003] Generally, the recognition success rate for clear QR code images is relatively high. However, during the scanning process, due to the movement of the QR code or the movement of the user's terminal, etc., resulting in a blurred QR code obtained by scanning, the QR code recognition success rate is relatively low. Summary of the Invention

[0004] Therefore, the embodiments of the present application provide a method, device, equipment and storage medium for QR code recognition, which fuse the information in multiple frames of images in a video stream to bridge the information loss caused by blurring between different frames, so as to improve the recognition effect of the QR code.

[0005] In order to achieve the above object, the embodiments of the present application provide the following technical solutions:

[0006] According to the first aspect of the embodiments of the present application, a method for QR code recognition is provided, and the method includes:

[0007] Obtain a first video frame containing a QR code from the scanned video;

[0008] Based on the QR code in the first video frame, perform alignment correction on each video frame in the video frame set related to the first video frame to obtain a sequence of image feature maps, and the sequence of image feature maps corresponds to the first video frame and each video frame in the video frame set;

[0009] Generate multi-channel image data according to the sequence of image feature maps;

[0010] Call a QR code recognition model to perform QR code dot matrix prediction on the multi-channel image data to obtain a QR code dot matrix prediction result;

[0011] Identify the QR code dot matrix prediction result to obtain a QR code recognition result.

[0012] Optionally, the video frame set includes a first set number of video frames before the first video frame in the video, and a second set number of video frames after it;

[0013] The performing alignment correction on each video frame in the video frame set related to the first video frame to obtain a sequence of image feature maps includes:

[0014] In the first video frame, determine a two-dimensional code target frame corresponding to the two-dimensional code, crop the first video frame according to the two-dimensional code target frame to obtain a target image, and expand the two-dimensional code target frame to obtain a two-dimensional code expansion frame;

[0015] Crop each video frame in the video frame set related to the first video frame according to the two-dimensional code expansion frame to obtain a two-dimensional code expansion map sequence;

[0016] Align the two-dimensional codes in the two-dimensional code expansion map sequence with the two-dimensional code in the target image to obtain an image feature map sequence; the image feature map sequence includes the target image and the two-dimensional code expansion map sequence after being corrected.

[0017] Optionally, the aligning the two-dimensional codes in the two-dimensional code expansion map sequence with the two-dimensional code in the target image to obtain the image feature map sequence corresponding to the video frame set includes:

[0018] Extract two-dimensional code feature points from the target image and the two-dimensional code expansion map sequence;

[0019] Based on the two-dimensional code feature points in the target image and the two-dimensional code expansion map sequence, align the two-dimensional codes in each image in the two-dimensional code expansion map sequence with the two-dimensional code in the target image to obtain an image feature map sequence.

[0020] Optionally, the calling the two-dimensional code recognition model to perform two-dimensional code dot matrix prediction on the multi-channel image data to obtain a two-dimensional code dot matrix prediction result includes:

[0021] For any multi-channel image data, perform at least one convolutional downsampling on the multi-channel image data to obtain a to-be-classified feature map with the multi-channel image data reduced to a first set size;

[0022] Classify the to-be-classified feature map to obtain the several two-dimensional code versions and the confidence of each two-dimensional code version;

[0023] Perform at least one deconvolutional upsampling on the to-be-classified feature map to obtain a reconstructed feature map with a second set size;

[0024] Perform convolutional processing on the reconstructed feature map of each two-dimensional code version to obtain a two-dimensional code dot matrix prediction result.

[0025] Optionally, the method further includes:

[0026] Train an initial model with two-dimensional code image samples to obtain the two-dimensional code recognition model;

[0027] The training process includes:

[0028] Input the QR code image sample into the initial model to obtain the predicted result of the output;

[0029] Calculate the target loss, where the target loss includes a version loss and a pixel loss. The version loss represents the difference between several QR code versions in the predicted result and the actual version of the QR code image sample, and the pixel loss represents the difference between the dot matrices of several QR code versions in the predicted result and the pixel dot matrix of the QR code image sample;

[0030] Update the parameters of the initial model according to the target loss to obtain a new initial model, and then execute again the step of inputting the QR code image sample into the initial model to obtain the predicted result of the output until the stop condition is met.

[0031] Optionally, the calling of the QR code recognition model to perform QR code dot matrix prediction on the multi-channel image data to obtain a QR code dot matrix prediction result includes:

[0032] Call the QR code recognition model to perform QR code dot matrix prediction on the multi-channel image data to obtain the confidence levels corresponding to the predicted dot matrices of several QR code versions;

[0033] According to the confidence levels, screen out the target dot matrix from the predicted dot matrices of several QR code versions to obtain the QR code dot matrix prediction result.

[0034] Optionally, the recognizing the QR code dot matrix prediction result to obtain a QR code recognition result includes:

[0035] Perform binarization processing on the QR code dot matrix prediction result according to a set threshold to obtain a QR code binarization result;

[0036] Perform decoding processing on the QR code binarization result to obtain a QR code recognition result.

[0037] According to the second aspect of the embodiments of the present application, there is provided a QR code recognition device, and the device includes:

[0038] A video frame acquisition module, configured to acquire a first video frame containing a QR code from the scanned video;

[0039] A correction module, configured to perform alignment correction on each video frame in the video frame set related to the first video frame based on the QR code in the first video frame to obtain a sequence of image feature maps, and the sequence of image feature maps corresponds to the first video frame and each video frame in the video frame set;

[0040] A multi-channel image data module, configured to generate multi-channel image data according to the sequence of image feature maps;

[0041] A model recognition module, configured to call a two-dimensional code recognition model to perform two-dimensional code dot matrix prediction on the multi-channel image data, so as to obtain a two-dimensional code dot matrix prediction result;

[0042] An identification result screening module, configured to identify the two-dimensional code dot matrix prediction result to obtain a two-dimensional code recognition result.

[0043] Optionally, the video frame set includes a first set number of video frames before the first video frame in the video and a second set number of video frames after the first video frame;

[0044] The multi-channel image data module is specifically configured to:

[0045] In the first video frame, determine a two-dimensional code target frame corresponding to the two-dimensional code, crop the first video frame according to the two-dimensional code target frame to obtain a target image, and expand the two-dimensional code target frame to obtain a two-dimensional code expansion frame;

[0046] Crop each video frame in the video frame set related to the first video frame according to the two-dimensional code expansion frame to obtain a two-dimensional code expansion map sequence;

[0047] Align the two-dimensional codes in the two-dimensional code expansion map sequence with the two-dimensional code in the target image to obtain an image feature map sequence; the image feature map sequence includes the target image and the two-dimensional code expansion map sequence after being corrected.

[0048] According to a third aspect of the embodiments of the present application, there is provided an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor runs the computer program, it is configured to implement the method described in the first aspect above.

[0049] According to a fourth aspect of the embodiments of the present application, there is provided a computer-readable storage medium, on which computer-readable instructions are stored, and the computer-readable instructions can be executed by a processor to implement the method described in the first aspect above.

[0050] In summary, the embodiments of the present application provide a QR code recognition method, device, equipment, and storage medium. The method includes obtaining a first video frame containing a QR code from a scanned video; based on the QR code in the first video frame, aligning and correcting each video frame in a video frame set related to the first video frame to obtain a sequence of image feature maps, where the sequence of image feature maps corresponds to the first video frame and each video frame in the video frame set; generating multi-channel image data according to the sequence of image feature maps; calling a QR code recognition model to perform QR code dot matrix prediction on the multi-channel image data to obtain a QR code dot matrix prediction result; and recognizing the QR code dot matrix prediction result to obtain a QR code recognition result. It can be seen that by adopting this implementation manner, the QR code information in multiple frames of QR code images in the video stream is fused, the information loss caused by blurring between different frames is bridged, and the QR code recognition effect is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary, and those of ordinary skill in the art can also obtain other implementation drawings according to the provided drawings without creative efforts.

[0052] The structures, ratios, sizes, etc. illustrated in this specification are only used to cooperate with the content disclosed in the specification for those who are familiar with this technology to understand and read, and are not used to limit the limited conditions for the implementation of the present invention. Therefore, they do not have technical substance. Any modification of the structure, change of the ratio relationship, or adjustment of the size should still fall within the scope covered by the technical content disclosed in the present invention without affecting the effects that the present invention can produce and the purposes that can be achieved.

[0053] Figures 1a - 1d It is a schematic diagram of the scenario of QR codes in the prior art;

[0054] Figure 2 It is a schematic diagram of the method flow of a QR code recognition method provided by an embodiment of the present application;

[0055] Figure 3 It is a schematic diagram of the scenario of the QR code target box b0 and the extended box b1 provided by an embodiment of the present application;

[0056] Figure 4 It is a schematic diagram of the scenario of the QR code extended graph sequence C provided by an embodiment of the present application;

[0057] Figure 5 It is a schematic diagram of the scenario of the registered and aligned image feature map B provided by an embodiment of the present application;

[0058] Figure 6 Schematic diagram of the training process of the QR code recognition model modified based on the U-net model provided by the embodiment of the present application;

[0059] Figure 7 Block diagram of the QR code recognition device provided by the embodiment of the present application;

[0060] Figure 8 Schematic diagram of the structure of an electronic device provided by the embodiment of the present application;

[0061] Figure 9 Schematic diagram of a computer-readable storage medium provided by the embodiment of the present application. Detailed implementation manners

[0062] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0063] The following introduces the technical scenarios of the embodiments of the present application.

[0064] A QR code image is an image that records information through an orderly two-dimensional arrangement of black and white grids. There are various encoding methods, and common types include qr code, data matrix, etc. The process of a user calling an application to scan a QR code can be a process in which the application obtains a video of the QR code image to decode the QR code image.

[0065] The qr code has strong anti-interference ability and fast recognition speed. Its symbol specifications range from version 1 (21×21 modules) to version 40 (177×177 modules). Fig. 1(a) shows a 21×21-module qr code (21 black and white grids horizontally and vertically), and Fig. 1(b) shows a 29×29-module qr code (29 black and white grids horizontally and vertically). The relationship between the version and the number of modules is: QRwidth=(Version - 1)*4 + 21; higher versions can store longer character lengths and are mostly used in daily life.

[0066] The data matrix QR code can be made in a very small size and can be designed as a square or rectangle to adapt to different-shaped mounting surfaces, and is mostly used in industrial production. Fig. 1(c) shows an 18×18-module data matrix (18 black and white grids horizontally and vertically), and Fig. 1(d) shows a 12×26-module data matrix (26 horizontally and 12 vertically black and white grids).

[0067] For the QR code composed of black and white grids, the embodiment of the present application proposes a QR code detection and recognition solution based on multi-frame fusion, which can effectively improve the success rate of QR code recognition in cases such as motion blur caused by the movement of the QR code or camera shake, out-of-focus blur caused by the camera not being focused, or low-resolution blur caused by long-distance shooting.

[0068] Figure 2 The following shows the QR code recognition method provided by the embodiment of the present application, and the method includes the following steps:

[0069] Step 201: Obtain the first video frame containing the QR code from the scanned video;

[0070] Step 202: Based on the QR code in the first video frame, perform alignment and correction on each video frame in the video frame set related to the first video frame to obtain a sequence of image feature maps, and the sequence of image feature maps corresponds to the first video frame and each video frame in the video frame set.

[0071] Among them, the alignment and correction is a process of aligning the QR codes in each video frame in the video frame set.

[0072] The video frame set related to the first video frame is a set of video frames adjacent to the first video frame.

[0073] Step 203: Generate multi-channel image data according to the sequence of image feature maps.

[0074] Among them, the multi-channel image data is a three-dimensional matrix synthesized by stacking each image in the sequence of image feature maps in the channel direction.

[0075] Step 204: Invoke the QR code recognition model to perform QR code dot matrix prediction on the multi-channel image data to obtain the QR code dot matrix prediction result.

[0076] Among them, the QR code dot matrix prediction result is the QR code dot matrix predicted according to the multi-channel image data.

[0077] Step 205: Recognize the QR code dot matrix prediction result to obtain the QR code recognition result.

[0078] In a possible implementation manner, the video frame set includes the first set number of video frames before the first video frame in the video, and the second set number of video frames after it. The first set number and the second set number can be the same or different.

[0079] In a possible implementation manner, in step 202, performing alignment and correction on each video frame in the video frame set related to the first video frame to obtain a sequence of image feature maps includes:

[0080] In the first video frame, determine the QR code target box corresponding to the QR code, crop the first video frame according to the QR code target box to obtain a target image, and expand the QR code target box to obtain a QR code expansion box; crop each video frame in the video frame set related to the first video frame according to the QR code expansion box to obtain a QR code expansion map sequence; correct the QR codes in the QR code expansion map sequence to be aligned with the QR code in the target image to obtain an image feature map sequence; the image feature map sequence includes the target image and the corrected QR code expansion map sequence.

[0081] In a possible implementation manner, the alignment correction adopts a method including feature point (such as SIFT feature point) registration.

[0082] In a possible implementation manner, correcting the QR codes in the QR code expansion map sequence to be aligned with the QR code in the target image to obtain an image feature map sequence corresponding to the video frame set includes:

[0083] Extract QR code feature points from the target image and the QR code expansion map sequence; based on the QR code feature points in the target image and the QR code expansion map sequence, align the QR codes in each image in the QR code expansion map sequence with the QR code in the target image to obtain an image feature map sequence.

[0084] In the QR code recognition method provided by the embodiments of the present application, by performing cropping and expansion of the same specification on each video frame in the video frame set related to the first video frame, the pertinence of QR code recognition in subsequent QR code detection is improved, and the efficiency and accuracy of QR code recognition are improved.

[0085] In a possible implementation manner, in step 203, generating multi-channel image data according to the image feature map sequence includes:

[0086] Perform grayscale processing on several images in the image feature map sequence respectively to obtain several two-dimensional grayscale images; stack the several two-dimensional grayscale images together to generate multi-channel image data.

[0087] In a possible implementation manner, after performing grayscale processing on several images in the image feature map sequence, the method further includes: performing enhancement processing on the several grayscale images after grayscale processing. The enhancement processing includes, but is not limited to, adding illumination, contrast, saturation, hue, Gaussian blur, motion blur, noise, etc.

[0088] In the QR code recognition method provided by the embodiments of the present application, performing grayscale processing on the image feature map sequence can stack single-channel data, thereby improving the processing efficiency of the QR code image. More image information can be obtained through enhancement processing, enhancing the robustness and invariance.

[0089] In a possible implementation manner, in step 204, a two-dimensional code recognition model is called to perform two-dimensional code dot matrix prediction on multi-channel image data, and a two-dimensional code dot matrix prediction result is obtained, including:

[0090] For any multi-channel image data, perform at least one convolutional downsampling on the multi-channel image data to obtain a feature map to be classified with the multi-channel image data reduced to a first set size; classify the feature map to be classified to obtain several two-dimensional code versions and the confidence of each two-dimensional code version; perform at least one deconvolutional upsampling on the feature map to be classified to obtain a reconstructed feature map with a second set size; perform convolutional processing on the reconstructed feature map of each two-dimensional code version to obtain a two-dimensional code dot matrix prediction result.

[0091] Based on the optimized U-net model provided in the embodiments of the present application, convolutional processing, image reconstruction, and feature fusion are performed on multi-channel image data, improving the image processing efficiency and contributing to feature classification and judgment of multi-channel image data.

[0092] In a possible implementation manner, in step 204, a two-dimensional code recognition model is called to perform two-dimensional code dot matrix prediction on multi-channel image data, and the confidence corresponding to the predicted dot matrix of several two-dimensional code versions is obtained; according to the confidence, the target dot matrix is screened out from the predicted dot matrices of several two-dimensional code versions to obtain a two-dimensional code dot matrix prediction result.

[0093] In the two-dimensional code recognition method provided in the embodiments of the present application, screening is performed based on the confidence, making the two-dimensional code prediction result more accurate.

[0094] In a possible implementation manner, in step 205, binarization processing is performed on the two-dimensional code dot matrix prediction result according to a set threshold to obtain a two-dimensional code binarization result; decoding processing is performed on the two-dimensional code binarization result to obtain a two-dimensional code recognition result.

[0095] In a possible implementation manner, the method further includes: training an initial model with two-dimensional code image samples to obtain a two-dimensional code recognition model.

[0096] In a possible implementation manner, the training process includes:

[0097] Input the QR code image sample into the initial model to obtain the predicted output result. Further, calculate the target loss, which includes the version loss and the pixel loss. The version loss represents the difference between several QR code versions in the predicted result and the actual version of the QR code image sample, and the pixel loss represents the difference between the dot matrices of several QR code versions in the predicted result and the pixel dot matrix of the QR code image sample. Update the parameters of the initial model according to the target loss to obtain a new initial model, and then execute again the step of inputting the QR code image sample into the initial model to obtain the predicted output result until the stop condition is met. The stop condition includes but is not limited to the target loss being less than the set threshold and the number of iterations being greater than the set iteration threshold.

[0098] Through iterative training and optimization of the QR code recognition model, the recognition of the QR code recognition model becomes more accurate.

[0099] The following describes in detail the QR code recognition method provided by the embodiments of the present application with reference to the accompanying drawings.

[0100] The QR code recognition method provided by the embodiments of the present application mainly includes three stages: QR code detection, multi-frame QR code correction and alignment, and construction, training and prediction of the QR code recognition model.

[0101] The first stage: Obtain the first video frame containing the QR code from the scanned video.

[0102] In a possible implementation manner, obtain the video frame containing the QR code from the scanned video through a QR code detection algorithm, denoted as the first video frame. Train a QR code detection algorithm D on a public QR code detection data set, and the yolo detection algorithm (such as yolov8) can be used.

[0103] Specifically, extract the video frames in the existing QR code short video data set S as image frames, and use the above QR code detection algorithm D for QR code detection; the QR code detection results can also be corrected manually to obtain the corrected data set; further optimize the QR code detection algorithm D on the corrected data set to obtain the optimized QR code detection algorithm D'.

[0104] The second stage: Based on the QR code in the first video frame, perform alignment and correction on each video frame in the video frame set related to the first video frame to obtain a sequence of image feature maps, and then generate multi-channel image data according to the sequence of image feature maps.

[0105] The sequence of image feature maps corresponds to the first video frame and each video frame in the video frame set.

[0106] Specifically, first, determine the QR code target box corresponding to the QR code in the first video frame, crop the first video frame according to the QR code target box to obtain a target image, and expand the QR code target box to obtain a QR code expansion box; then crop each video frame in the video frame set related to the first video frame according to the QR code expansion box to obtain a QR code expansion map sequence. Further, extract QR code feature points from the target image and the QR code expansion map sequence; based on the QR code feature points in the target image and the QR code expansion map sequence, align the QR code in each image in the QR code expansion map sequence with the QR code in the target image to obtain an image feature map sequence; the image feature map sequence includes the target image and the corrected QR code expansion map sequence. Further, perform grayscale processing on several images in the image feature map sequence respectively to obtain several two-dimensional grayscale images; stack the several two-dimensional grayscale images together to generate multi-channel image data.

[0107] An example is given for the method provided in the embodiment of the present application above, including the following steps:

[0108] Step 1: Randomly extract a first video frame I containing a QR code from the scanned video i , and use the QR code detection algorithm D’ to detect the QR code to obtain a series of QR code target boxes.

[0109] Step 2: Select a QR code target box b0 from the series of QR code target boxes, and based on the QR code target box b0, expand it to a QR code expansion box b1 with the length and width of b0 in all directions. Figure 3 The schematic diagram of the QR code target box b0 and the QR code expansion box b1 is shown.

[0110] Step 3: Obtain the first set number of video frames before the first video frame and the second set number of video frames after it in the video as the video frame set. That is, obtain the k images before and after the first video frame I i , and get the video frame collection I = {I i-k ,,..., I i-1 , I i+1 ,,..., I i+k}.

[0111] Step 4: Crop the video frame collection I = {I i-k ,,..., I i-1 , I i+1 ,,..., I i+k} according to the size of the QR code expansion box b1 to obtain a 2k QR code expansion map sequence C = {C i-k ,,..., C i-1 , C i+1 ,,..., C i+k}. Figure 4Shows the QR code extended map sequence C.

[0112] Step 5: In video frame I i Intercept the b0 area to obtain the target image B i ; Use the method of registration based on feature points (such as SIFT feature points) to correct and align the QR codes in the 2k QR code extended map sequence C to the QR code in the target image B i ; If the registration fails, directly use B i instead; Finally, obtain the image feature map sequence B = {B i-k, ,..., B i-1 , B i , B i+1 ,..., B i+k}. Figure 5 Shows the image feature map sequence B.

[0113] Step 6: Convert all the images in the image feature map sequence B into grayscale images.

[0114] Step 7: Augment the grayscale images in B using the same set of parameters. The augmentation methods include adding illumination, contrast, saturation, hue, Gaussian blur, motion blur, noise, etc.

[0115] Step 8: Stack the images in the augmented B in the channel direction to obtain a multi-channel image data B*, whose number of channels is 2k + 1. One grayscale image corresponds to a two-dimensional matrix with size h*w, and one RGB color image corresponds to a three-dimensional matrix with size h*w*3. Multiple grayscale images correspond to n two-dimensional matrices, i.e., n*[h*w]. Stack them in the channel direction to synthesize a three-dimensional matrix with size h*w*n.

[0116] Step 9: The stacked image B can be represented by a 0, 1 dot matrix, and the obtained dot matrix L and the image B form an image label pair. For example, for a 21×21 module QR code image, it is represented by a 21×21 0, 1 dot matrix, where 0 represents a black square and 1 represents a white square; while for a 29×29 module QR code, it is represented by a dot matrix with size 29×29.

[0117] The third stage: Application of the QR code recognition model. Call the QR code recognition model to predict the QR code dot matrix for the multi-channel image data to obtain the QR code dot matrix prediction result; and then further identify the QR code dot matrix prediction result to obtain the QR code recognition result.

[0118] Specifically, for any multi-channel image data, at least one convolutional downsampling is performed on the multi-channel image data to obtain a feature map to be classified with the multi-channel image data reduced to a first set size; further, the feature map to be classified is classified to obtain a number of QR code versions and the confidence level of each QR code version; further, at least one deconvolutional upsampling is performed on the feature map to be classified to obtain a reconstructed feature map of a second set size; further, convolutional processing is performed on the reconstructed feature map of each QR code version to obtain a QR code dot matrix prediction result.

[0119] Further, the one with the largest confidence level is selected from the QR code dot matrix prediction results as the target dot matrix; the target dot matrix is binarized according to a set threshold to obtain a QR code binarization result; and the QR code binarization result is decoded to obtain a QR code recognition result.

[0120] The main architecture of the QR code recognition model in the embodiments of the present application is based on the U-net model. The U-net model is divided into a backbone feature extraction network (also called an encoder) and an enhanced feature extraction network (also called a decoder). The backbone feature extraction network performs convolutional downsampling, and the size of the feature map is reduced until it becomes 1*1. The enhanced feature extraction network starts from this 1*1 feature map and performs deconvolutional upsampling until it is reconstructed to 1 / 2 of the original image size. During the decoding process, the encoding layers of the corresponding sizes are concatenated for feature fusion. Since the encoding and decoding processes resemble a "U", it is named U-net. The U-net model is mainly used in tasks such as image segmentation and super-resolution.

[0121] In the application of the QR code recognition model, a multi-channel image data after multi-frame QR code correction and alignment is input, and the output QR code dot matrix prediction result is a number of QR code versions and the corresponding predicted dot matrix. Further, the version number with the largest confidence level is selected from the predicted dot matrices of different QR code versions, and the predicted dot matrix corresponding to the version number with the largest confidence level is binarized using a set threshold (such as 0.5) to obtain a QR code binarization result. Then, the QR code recognition result is decoded using a decoding recognition algorithm (such as ZBar).

[0122] The fourth stage: the construction and training of the QR code recognition model.

[0123] Figure 6 It shows the construction and training process of the QR code recognition model modified based on the U-net model provided by the embodiments of the present application.

[0124] Input multi-channel image data into the backbone feature extraction network. For any multi-channel image data, perform convolutional downsampling, and the size of the feature map gradually shrinks until it becomes 1*1. The QR code recognition model provided by the embodiments of the present application adds a classification layer to the 1*1 feature map at the bottom of the conventional U-net model to classify the feature map to be classified to obtain several QR code versions and the confidence of each QR code version; output the QR code version, and the number of versions is determined by the types of QR codes to be supported. In Figure 6 the example of

[0125] it supports 40 kinds of qr code - symbol specifications range from version 1 (21×21 modules) to version 40 (177×177 modules). Further, perform deconvolutional upsampling starting from the 1*1 feature map to be classified through the enhanced feature extraction network. The QR code recognition model provided by the embodiments of the present application removes the last stage of the decoding process of the conventional U-net model, and the feature map to be classified is only reconstructed to 1 / 4 of the original image size. Multiple SPP (Spatial Pyramid Pooling) layers are also added to pool the feature map to: 177×177 size (corresponding to the size of version 40), 173×173 size (corresponding to the size of version 39),..., 25×25 size (corresponding to the size of version 2), 21×21 size (corresponding to the size of version 1). Then add a 1×1 convolution to each pooled feature map, with the input channel number being 1, to obtain the prediction results of QR code dot matrices of different sizes and output them.

[0126] In a possible implementation manner, an initial model is also trained with QR code image samples to obtain a QR code recognition model; the training process includes:

[0127] Step 1: Input the QR code image samples into the initial model to obtain training results;

[0128] Step 2: Calculate the target loss, where the target loss includes a version loss and a pixel loss. The version loss represents the difference between several QR code versions in the prediction results and the actual versions of the QR code image samples, and the pixel loss represents the difference between the dot matrices of several QR code versions in the prediction results and the pixel dot matrices of the QR code image samples;

[0129] Step 3: Update the parameters of the initial model according to the target loss to obtain a new initial model, and then execute the step of inputting the QR code image samples into the initial model to obtain the predicted output results again until the stop condition is met.

[0130] The supervised process is to optimize the model with labeled data and directly predict unlabeled data based on the trained model. For each model, a QR code version classification vector and prediction dot matrices corresponding to multiple different QR code versions are output. Using the QR code dot matrix label L obtained in the second stage, with its QR code version number being v, the output QR code version is supervised using the QR code version number v, and the version loss of the QR code version is calculated using the cross-entropy loss function; the dot matrix L of the QR code is used to supervise the dot matrix of the size corresponding to version v in the output dot matrices of different sizes of the QR code, and the pixel loss of the dot matrix is calculated using the pixel-based cross-entropy loss function. The version loss of the QR code version and the pixel loss of the dot matrix are added together, and the QR code recognition model is optimized using the principle of gradient descent.

[0131] In summary, the embodiments of the present application provide a QR code recognition method, device, equipment, and storage medium. By obtaining a first video frame containing a QR code from the scanned video; based on the QR code in the first video frame, aligning and correcting each video frame in the video frame set related to the first video frame to obtain a sequence of image feature maps, where the sequence of image feature maps corresponds to the first video frame and each video frame in the video frame set; generating multi-channel image data according to the sequence of image feature maps; calling a QR code recognition model to perform QR code dot matrix prediction on the multi-channel image data to obtain a QR code dot matrix prediction result; and recognizing the QR code dot matrix prediction result to obtain a QR code recognition result. It can be seen that with this implementation method, the QR code information in multiple frames of QR code images in the video stream is fused, bridging the information loss caused by blurring between different frames and improving the QR code recognition effect.

[0132] Based on the same technical concept, the embodiments of the present application also provide a QR code recognition device, as Figure 7 shown. The device includes:

[0133] A video frame acquisition module 701, configured to obtain a first video frame containing a QR code from the scanned video;

[0134] A correction module 702, configured to, based on the QR code in the first video frame, align and correct each video frame in the video frame set related to the first video frame to obtain a sequence of image feature maps, where the sequence of image feature maps corresponds to the first video frame and each video frame in the video frame set;

[0135] A multi-channel image data module 703, configured to generate multi-channel image data according to the sequence of image feature maps;

[0136] A model recognition module 704, configured to call a QR code recognition model to perform QR code dot matrix prediction on the multi-channel image data to obtain a QR code dot matrix prediction result;

[0137] The recognition result screening module 705 is used to recognize the two-dimensional code dot matrix prediction result to obtain the two-dimensional code recognition result.

[0138] In a possible implementation manner, the video frame set includes a first set number of video frames before the first video frame in the video and a second set number of video frames after it.

[0139] In a possible implementation manner, the multi-channel image data module 703 is specifically used for:

[0140] In the first video frame, determine the two-dimensional code target box corresponding to the two-dimensional code, crop the first video frame according to the two-dimensional code target box to obtain the target image, and expand the two-dimensional code target box to obtain the two-dimensional code expansion box; crop each video frame in the video frame set related to the first video frame according to the two-dimensional code expansion box to obtain a two-dimensional code expansion map sequence; align the two-dimensional codes in the two-dimensional code expansion map sequence with the two-dimensional code in the target image to obtain an image feature map sequence; the image feature map sequence includes the target image and the two-dimensional code expansion map sequence after being corrected.

[0141] In a possible implementation manner, the correction module 702 is specifically used to extract two-dimensional code feature points from the target image and the two-dimensional code expansion map sequence; based on the two-dimensional code feature points in the target image and the two-dimensional code expansion map sequence, align the two-dimensional codes in each image in the two-dimensional code expansion map sequence with the two-dimensional code in the target image to obtain an image feature map sequence.

[0142] In a possible implementation manner, the model recognition module 704 is specifically used for any multi-channel image data, perform at least one convolutional downsampling on the multi-channel image data to obtain a to-be-classified feature map with the multi-channel image data reduced to a first set size; classify the to-be-classified feature map to obtain several two-dimensional code versions and the confidence of each two-dimensional code version; perform at least one deconvolutional upsampling on the to-be-classified feature map to obtain a reconstructed feature map with a second set size; perform convolutional processing on the reconstructed feature map of each two-dimensional code version to obtain the two-dimensional code dot matrix prediction result.

[0143] In a possible implementation manner, the model recognition module 704 is specifically used to call a two-dimensional code recognition model to perform two-dimensional code dot matrix prediction on multi-channel image data, obtain the confidence corresponding to the predicted dot matrices of several two-dimensional code versions respectively; according to the confidence, screen out the target dot matrix from the predicted dot matrices of several two-dimensional code versions to obtain the two-dimensional code dot matrix prediction result.

[0144] In a possible implementation manner, the recognition result screening module 705 is specifically used to perform binarization processing on the two-dimensional code dot matrix prediction result according to a set threshold to obtain a two-dimensional code binarization result; perform decoding processing on the two-dimensional code binarization result to obtain the two-dimensional code recognition result.

[0145] In a possible implementation, the system further includes:

[0146] A training module, configured to train an initial model with a two-dimensional code image sample to obtain a two-dimensional code recognition model; the training process includes: inputting the two-dimensional code image sample into the initial model to obtain an output prediction result; calculating a target loss, where the target loss includes a version loss and a pixel loss, the version loss represents the difference between several two-dimensional code versions in the prediction result and the actual version of the two-dimensional code image sample, and the pixel loss represents the difference between the dot matrixes of several two-dimensional code versions in the prediction result and the pixel dot matrix of the two-dimensional code image sample; updating the parameters of the initial model according to the target loss to obtain a new initial model, and repeatedly performing the step of inputting the two-dimensional code image sample into the initial model to obtain an output prediction result until a stop condition is satisfied.

[0147] The embodiment of the present application also provides an electronic device corresponding to the method provided in the foregoing embodiment. Please refer to Figure 8 , which shows a schematic diagram of an electronic device provided in some embodiments of the present application. The electronic device 20 may include: a processor 200, a memory 201, a bus 202, and a communication interface 203. The processor 200, the communication interface 203, and the memory 201 are connected through the bus 202; a computer program that can run on the processor 200 is stored in the memory 201, and when the processor 200 runs the computer program, it executes the method provided in any of the foregoing embodiments of the present application.

[0148] Among them, the memory 201 may include a high-speed random access memory (RAM: Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. The communication connection between the system network element and at least one other network element is realized through at least one physical port 203 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.

[0149] The bus 202 may be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. Among them, the memory 201 is used to store a program. After receiving an execution instruction, the processor 200 executes the program, and the method disclosed in any of the foregoing embodiments of the present application can be applied to the processor 200 or implemented by the processor 200.

[0150] The processor 200 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 200 or the instructions in the form of software. The above-mentioned processor 200 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 201, and the processor 200 reads the information in the memory 201 and combines its hardware to complete the steps of the above method.

[0151] The electronic device provided by the embodiment of the present application and the method provided by the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, run or implemented by it.

[0152] The embodiment of the present application also provides a computer-readable storage medium corresponding to the method provided by the foregoing embodiment. Please refer to Figure 9 , which shows that the computer-readable storage medium is an optical disc 30, on which a computer program (i.e., a program product) is stored. When the computer program is run by the processor, it will execute the method provided by any of the foregoing embodiments.

[0153] It should be noted that examples of computer-readable storage media may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical and magnetic storage media, which will not be elaborated here one by one.

[0154] The computer-readable storage medium provided by the above embodiment of the present application and the method provided by the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, run or implemented by the application program stored in it.

[0155] It should be noted that:

[0156] The algorithms and displays provided herein are not inherently related to any particular computer, virtual apparatus, or other device. Various general-purpose apparatuses may also be used in conjunction with the teachings presented herein. The structure required to construct such apparatuses will be apparent from the above description. Additionally, the present application is not directed to any particular programming language. It should be understood that the content of the present application described herein can be implemented using various programming languages, and the descriptions made above with respect to specific languages are for the purpose of disclosing the best mode of the present application.

[0157] In the specification provided herein, numerous specific details are set forth. However, it can be understood that embodiments of the present application may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.

[0158] Similarly, it should be understood that in order to streamline the present application and assist in understanding one or more of the various inventive aspects, in the above description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that the claimed present application requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspects lie in less than all of the features of the single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim standing on its own as a separate embodiment of the present application.

[0159] Those skilled in the art will appreciate that the modules in the devices in the embodiments can be adaptively changed and disposed in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be adopted for all the features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) can be replaced by an alternative feature providing the same, equivalent, or similar purpose.

[0160] In addition, those skilled in the art can understand that although some embodiments herein include certain features included in other embodiments rather than other features, the combination of features of different embodiments means that it is within the scope of this application and forms different embodiments. For example, in the following claims, any one of the claimed embodiments can be used in any combination.

[0161] Each component embodiment of the present application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components in the virtual machine creation device according to the embodiments of the present application. The present application can also be implemented as a device or device program (such as a computer program and a computer program product) for executing part or all of the methods described herein. Such a program for implementing the present application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0162] It should be noted that the above embodiments illustrate the present application rather than limit the present application, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In a unit claim listing several devices, several of these devices can be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names.

[0163] The above is only the preferred specific implementation mode of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A method for QR code recognition, characterized in that, The method includes: Obtaining a first video frame containing a QR code from the scanned video; Based on the QR code in the first video frame, performing alignment correction on each video frame in the video frame set related to the first video frame to obtain a sequence of image feature maps, where the sequence of image feature maps corresponds to the first video frame and each video frame in the video frame set; Generating multi-channel image data according to the sequence of image feature maps; Invoking a QR code recognition model to perform QR code dot matrix prediction on the multi-channel image data to obtain a QR code dot matrix prediction result; Identifying the QR code dot matrix prediction result to obtain a QR code recognition result.

2. The method according to claim 1, characterized in that The video frame set includes a first set number of video frames before the first video frame in the video and a second set number of video frames after it; The performing alignment correction on each video frame in the video frame set related to the first video frame to obtain a sequence of image feature maps includes: In the first video frame, determining a QR code target box corresponding to the QR code, cropping the first video frame according to the QR code target box to obtain a target image, and expanding the QR code target box to obtain a QR code expansion box; Cropping each video frame in the video frame set related to the first video frame according to the QR code expansion box to obtain a sequence of QR code expansion maps; Aligning the QR codes in the sequence of QR code expansion maps with the QR code in the target image to obtain a sequence of image feature maps; the sequence of image feature maps includes the target image and the sequence of QR code expansion maps after correction.

3. The method according to claim 2, wherein The aligning the QR codes in the sequence of QR code expansion maps with the QR code in the target image to obtain the sequence of image feature maps corresponding to the video frame set includes: Extracting QR code feature points from the target image and the sequence of QR code expansion maps; Based on the QR code feature points in the target image and the sequence of QR code expansion maps, aligning the QR code in each image in the sequence of QR code expansion maps with the QR code in the target image to obtain a sequence of image feature maps.

4. The method according to claim 1, wherein The invoking a QR code recognition model to perform QR code dot matrix prediction on the multi-channel image data to obtain a QR code dot matrix prediction result includes: For any multi-channel image data, performing at least one convolution downsampling on the multi-channel image data to obtain a feature map to be classified with the multi-channel image data reduced to a first set size; Classifying the feature map to be classified to obtain a number of QR code versions and the confidence of each QR code version; Performing at least one deconvolution upsampling on the feature map to be classified to obtain a reconstructed feature map with a second set size; Performing convolution processing on the reconstructed feature map of each QR code version to obtain a QR code dot matrix prediction result.

5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Training an initial model with QR code image samples to obtain the QR code recognition model; The training process includes: Inputting the QR code image samples into the initial model to obtain an output prediction result; Calculate the target loss, where the target loss includes a version loss and a pixel loss. The version loss characterizes the difference between several QR code versions in the prediction result and the actual version of the QR code image sample, and the pixel loss characterizes the difference between the dot matrices of several QR code versions in the prediction result and the pixel dot matrix of the QR code image sample; Update the parameters of the initial model according to the target loss to obtain a new initial model, and perform again the step of inputting the QR code image sample into the initial model to obtain the prediction result of the output until the stop condition is satisfied.

6. The method according to claim 1, wherein The calling of the QR code recognition model to perform QR code dot matrix prediction on the multi-channel image data to obtain a QR code dot matrix prediction result includes: Calling the QR code recognition model to perform QR code dot matrix prediction on the multi-channel image data to obtain the confidence levels corresponding to the predicted dot matrices of several QR code versions; According to the confidence levels, screen out the target dot matrices from the predicted dot matrices of the several QR code versions to obtain a QR code dot matrix prediction result.

7. The method according to claim 1, characterized in that The recognizing of the QR code dot matrix prediction result to obtain a QR code recognition result includes: Performing binarization processing on the QR code dot matrix prediction result according to a set threshold to obtain a QR code binarization result; Performing decoding processing on the QR code binarization result to obtain a QR code recognition result.

8. A QR code recognition device, characterized in that, The device includes: A video frame acquisition module, configured to acquire a first video frame containing a QR code from the scanned video; A correction module, configured to perform alignment correction on each video frame in the video frame set related to the first video frame based on the QR code in the first video frame to obtain a sequence of image feature maps, where the sequence of image feature maps corresponds to the first video frame and each video frame in the video frame set; A multi-channel image data module, configured to generate multi-channel image data according to the sequence of image feature maps; A model recognition module, configured to call a QR code recognition model to perform QR code dot matrix prediction on the multi-channel image data to obtain a QR code dot matrix prediction result; A recognition result screening module, configured to recognize the QR code dot matrix prediction result to obtain a QR code recognition result.

9. An electronic device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that when the processor runs the computer program, it is executed to implement the method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, A computer-readable instruction is stored thereon, and the computer-readable instruction can be executed by a processor to implement the method according to any one of claims 1-7.