Two-dimensional code attached video identification and analysis method based on object perception
Through real-time video acquisition and feature analysis on mobile terminals, the problem of QR code cheating in the audit system is solved, and fast and accurate QR code carrier identification is achieved, which is suitable for scenarios such as product anti-counterfeiting and document verification.
Patent Information
- Application Number
- CN202511248760.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-09-03
AI Technical Summary
In existing sales audit systems, auditors may directly scan pre-prepared QR code images instead of the actual product, resulting in data distortion. Existing anti-cheating methods have low accuracy, are time-consuming, or are unable to cope with advanced cheating.
A QR code attachment video recognition and analysis method based on object perception is adopted. Through real-time video acquisition from mobile terminals, the spatiotemporal and background features of the QR code area are extracted. Combined with lightweight depth estimation and image quality analysis, it can quickly distinguish whether the QR code is attached to a three-dimensional object or a two-dimensional plane.
It can quickly and accurately distinguish whether the carrier attached to the QR code is a physical object or a picture, eliminate the deception of pre-stored pictures or videos, meet the real-time response needs of mobile terminals, reduce data transmission volume and computing overhead, and is suitable for scenarios such as product anti-counterfeiting inspection and document authenticity verification.
Smart Images

Figure CN120726549A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the intersection of computer vision, mobile applications and anti-counterfeiting technology, and in particular to a method for identifying and analyzing QR code attachment videos based on object perception. Background Art
[0002] Existing sales audit systems typically require auditors to complete the audit by scanning QR codes on products. However, to save time, some auditors may directly scan pre-prepared QR code images instead of the actual product. This cheating behavior can distort backend data and affect decision-making.
[0003] Currently, there are some anti-cheating methods on the market, such as: 1. Require users to upload multiple photos from different angles: This method still leaves the possibility that users will use pre-stored photos in their albums, and the cost of manual review is high.
[0004] 2. Using a single image to infer 3D information: Due to the limited information in a single image, the accuracy is low and it cannot cope with advanced cheating (such as high-quality printed images).
[0005] 3. Video-based anti-counterfeiting methods: Some methods use videos, but they usually take a long time (more than 10 seconds) to process, making it difficult to meet the mobile terminal's requirements for real-time response (within 5 seconds), and allowing album uploads has a vulnerability to pre-recorded video fraud.
[0006] Therefore, a method is needed that can quickly and accurately distinguish whether the carrier attached to the QR code is a physical object or a picture, and prevent cheating from the source (i.e., forcing the use of real-time video). Summary of the Invention
[0007] The purpose of the present invention is to solve the shortcomings of the prior art and to propose a method for identifying and analyzing QR code attachment videos based on object perception.
[0008] To achieve the above object, the present invention adopts the following technical solutions: A method for identifying and analyzing a QR code attached video based on object perception includes a mobile terminal equipped with a QR code scanning and verification module; and further includes the following steps: S1: Real-time video acquisition and preprocessing; S11: Startup and constraints of the QR code scanning verification module; Start the QR code scanning and verification module on the mobile terminal, set the permissions of the QR code scanning and verification module, limit the use of only the rear camera and block the album access interface, and the user uses the QR code scanning and verification module to aim at the object with the QR code to be identified and take a picture; S12: short video recording; After the user triggers the shooting, short videos of a certain length are recorded continuously; S13: extraction and decoding of key frames; From the recorded video, samples are taken at fixed intervals to extract valid frames with successful QR code recognition. Each frame is then preprocessed, and the QR code area is located and decoded using a fast QR code recognition library. The bounding box coordinates, center point coordinates, and area of the QR code in each frame are recorded. The frame image is then converted into a grayscale image. S2: Feature extraction; S21: Extraction of spatiotemporal features of QR code regions; S211: Position / posture change analysis; S212: Surface deformation consistency analysis; S213: Determine the movement mode; S214: eigenvalue output; S22: Extraction of background features around the QR code; S221: Background motion consistency analysis; S222: Obtain background depth clues: Use a lightweight depth estimation model to output a depth map and obtain the depth gradient of the QR code boundary; S223: Background stability analysis; S23: Extract image quality and anti-spoofing features; Select the QR code area in the middle valid frame for analysis, and analyze the high-frequency noise energy, moiré effect, and edge sharpness of the QR code area; S3: Feature fusion and classification; S31: extract feature vector; S32: Feature vector input classification model; S4: Result output and anti-cheating judgment; Receive the output result of the classification model and determine whether it is a three-dimensional physical carrier or a two-dimensional plane carrier.
[0009] Preferably, step S211 is to calculate motion features based on the bounding box coordinates, center point coordinates and area of the QR code in adjacent frames. The motion features include: the displacement vector and displacement size of the bounding box center point of the entire frame sequence, the average displacement, the maximum displacement, the displacement standard deviation, the bounding box area change rate, the average area change rate, and the maximum area change rate.
[0010] Preferably, step S212 is to calculate the cosine similarities between all non-zero motion vectors, and then calculate the average value of the cosine similarities of all inter-frame motion vectors in the entire sequence as a motion consistency indicator; S212a: Obtain a non-zero motion vector; Obtain the displacement vectors of the center points of the bounding boxes of all adjacent QR codes to obtain a motion vector set. Remove motion vectors with a modulus less than 10^(-6) to obtain valid motion vectors, and combine the valid vectors into a non-zero motion vector set. S212b: Calculate the cosine similarity of any two different vectors in the set of non-zero motion vectors; Filter out valid vector pairs and calculate the cosine similarity of all valid vector pairs; S212c: obtaining motion consistency index; The average of the cosine similarities of all vector pairs is calculated to obtain the motion consistency index.
[0011] Preferably, it also includes: S213a: Determine whether it is a 2D motion flag; Determine whether the QR code's motion is two-dimensional plane motion: If the following conditions are met: average displacement ≤ 15 pixels, displacement standard deviation < 10 pixels, average area change rate ≤ 5%, and motion consistency > 0.7, then it is determined to be two-dimensional plane motion. If not, then it is determined to be no, not two-dimensional plane motion, and the determination result is converted to 1.0 or 0.0 using a Boolean value; S213b: Determine whether it is a rigid plane mark; Determine whether the QR code is attached to a rigid plane: If the following conditions are met: average area change rate ≤ 3% and maximum area change rate ≤ 10% and displacement standard deviation < 8 pixels, then it is determined to be attached to a rigid plane. If not, then it is determined to be no, not attached to a rigid plane, and the determination result is converted to 1.0 or 0.0 through a Boolean value.
[0012] Preferably, in step S221, a background area is selected outside the QR code bounding box, the optical flow field of the background area and the QR code area is calculated using the Farneback optical flow method, and the difference between the average optical flow vector of the background area and the average optical flow vector of the QR code area is compared; S221a: Calculate the optical flow field of the background area and the QR code area; First, the QR code area and background area are obtained: grayscale images of a group of adjacent frames of the middle valid frame are extracted, and the rectangular area is cropped as the QR code area according to the coordinates of the QR code's bounding box; a background area is selected outside the QR code's bounding box; Next, feature points are extracted: first, feature points of the QR code area and background area of a valid frame are extracted, and then feature points of the QR code area and background area of another frame are extracted using feature point matching; Finally, the optical flow vectors of the feature points in the QR code area and the background area are calculated respectively; The optical flow vector of the QR code area is obtained based on the difference between the feature points matching the QR code area, and the optical flow vector of the background area is obtained based on the difference between the feature points matching the background area; S221b: Calculate the average optical flow vector of all feature points in the QR code area and the background area respectively; S221c: calculating motion difference; The motion difference is obtained by calculating the modulus of the difference between the average optical flow vector of the QR code area and the average optical flow vector of the background area.
[0013] Preferably, step S222 includes: Step 1: Preprocessing of valid frame images; Obtain the grayscale image of the valid frame, and generate an image that meets the input requirements of the lightweight MiDaS model through resizing and normalization; Step 2: Depth map generation; The pre-processed valid frame image is input into the lightweight MiDaS model, and the MiDaS model infers and outputs a depth map; Step 3: Extract the boundary of the QR code; According to the obtained bounding box, the four sides of the QR code bounding box are discretized and sampled by linear interpolation method to obtain the pixel point set of the QR code boundary; Step 4: Depth gradient calculation; The Sobel operator is used to calculate the gradient of the depth map, and the gradient modulus of the extracted pixel points is calculated to obtain the boundary gradient set; Step 5: Get the depth gradient at the boundary of the QR code: The maximum value in the boundary gradient set is obtained as the depth gradient at the boundary of the QR code.
[0014] Preferably, step S223 includes: obtaining the grayscale value of each pixel in the background area of the previous frame and the current frame, calling the corrcoef function according to the grayscale value to directly calculate the correlation coefficient, and obtaining the correlation coefficient of the background area.
[0015] Preferably, the feature vector is a 12-dimensional vector, specifically composed of: 6-dimensional features from S21: Displacement normalization value: the average displacement normalization process is mapped to the range of [0, 1]; Average area change rate, directly taken value; Motion consistency index, directly taking values; 2D motion flag: 1.0 or 0.0; Rigid plane flag: 1.0 or 0.0; Normalized displacement standard deviation value: mapped to the range of [0, 1] by normalization of displacement standard deviation; 3D features from S22: Motion difference normalization value: Motion difference normalization processing is mapped in the range of [0, 1]; Depth gradient normalization value: The depth gradient normalization processing at the boundary of the QR code is mapped in the range of [0, 1]; Background stability, directly obtain the correlation coefficient of the background area; 3D features from S23: High-frequency energy: The high-frequency noise energy in the QR code area is directly measured; Moiré intensity: directly obtain the intensity of the moiré effect in the QR code area; Edge sharpness: The average edge gradient of the QR code area is directly taken.
[0016] Preferably, the classification model is a pre-trained lightweight CNN classifier with three layers of convolution, an input feature dimension of 128, a ReLU activation function, and an output of the probability of a three-dimensional physical carrier and a two-dimensional plane carrier; Preferably, step S4 is specifically: If it is determined to be a "three-dimensional physical carrier", the verification is successful, confirming that the QR code is attached to a real object; If it is determined to be a "two-dimensional flat carrier", the verification fails and is suspected of cheating by using a pre-stored image or screen display. A pop-up window will appear on the user interface, prompting "The QR code may be from an image or screen. Please aim at the real object to rescan and verify." If the number of valid frames is insufficient or the QR code cannot be extracted, the verification fails and the user interface prompts "QR code not detected, please aim at the physical object and rescan for verification"; if the number of valid frames N is ≥ 5, the verification is considered to have failed.
[0017] Compared with the prior art, the beneficial effects of the present invention are: 1. Anti-cheating at the source: forcing real-time video shooting and prohibiting access to local albums, fundamentally eliminating the possibility of cheating using pre-stored QR code pictures or videos.
[0018] 2. High efficiency and real-time performance: The lightweight algorithm design, combined with efficient frame sampling and feature extraction strategies, can complete the entire process from video capture to authenticity determination within 5 seconds on a mobile device or in the cloud, meeting the rapid response requirements of mobile applications.
[0019] 3. High identification accuracy: Multi-dimensional fusion analysis is performed by comprehensively utilizing the spatiotemporal dynamic characteristics of the QR code region itself (position / posture changes, surface deformation consistency), the correlation characteristics between the QR code and the surrounding background (motion consistency, depth continuity), and image anti-spoofing features. This significantly improves the accuracy of distinguishing between three-dimensional physical carriers and two-dimensional flat carriers.
[0020] 4. Low resource consumption: Only 2-3 seconds of video is required (rather than the 10-30 seconds required by traditional solutions), and the processing is optimized, significantly reducing data transmission (traffic consumption) and cloud computing overhead. Furthermore, the shorter video length and processing time make it more robust in weak network environments.
[0021] 5. Wide applicability: This method can be applied to various scenarios requiring authenticity verification of physical objects, such as product anti-counterfeiting inspection, document authenticity verification, and industrial parts traceability. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 This is a flowchart of a method for identifying and analyzing QR code attachment videos based on object perception according to the present invention; Figure 2 This is an overall flow chart of a QR code attachment video recognition and analysis method based on object perception of the present invention. DETAILED DESCRIPTION
[0023] In order to provide a further understanding of the purpose, structure, features, and functions of the present invention, the present invention is described in detail below with reference to the embodiments.
[0024] like Figure 1 and Figure 2 A method for identifying and analyzing a QR code attached video based on object perception includes a mobile terminal equipped with a QR code scanning and verification module, including the following steps: S1: Real-time video acquisition and preprocessing; S11: Startup and constraints of the QR code scanning verification module; Launch the QR code scanning and verification module on the mobile terminal, set permissions for the module, restrict access to the rear camera only, and block access to the photo album. Users are required to use the QR code scanning and verification module to capture an object with the QR code to be identified. This approach, which restricts access to the rear camera and blocks access to the photo album, effectively addresses the vulnerability of existing technologies that allow pre-stored images / videos to be uploaded to the photo album. In sales audit scenarios, this prevents auditors from using pre-saved QR code images to circumvent inspections, ensuring data is derived from actual scans of physical objects and blocking fraudulent methods at the source. The rear camera offers higher resolution and stability on mobile devices, making it more suitable for capturing clear QR code details. The permission-blocking design requires no additional user input, improving ease of use, reducing user learning curves, and increasing efficiency in audit and verification scenarios.
[0025] S12: short video recording; After the user triggers the capture, the app controls the camera to continuously record short videos of a fixed duration (preferably 2-3 seconds). This process requires the user to keep the phone relatively stable (allowing for slight shaking caused by normal hand-held operation) and the camera to maintain a moderate distance from the QR code (ensuring that the QR code is clearly legible and occupies a certain proportion of the screen). This ensures the accuracy of subsequent frame decoding and feature analysis, and reduces invalid verifications caused by poor shooting conditions (such as hand shaking or excessive distance).
[0026] S13: extraction and decoding of key frames; The recorded video is sampled at a fixed interval, such as 0.2s, to extract valid frames where the QR code is successfully recognized (N must be ≥ 5; if this is less than 5, the verification is considered a failure). If no valid QR code frames can be extracted, the verification is considered a failure. Each valid frame is then preprocessed (such as denoising and brightness / contrast adjustment) to eliminate environmental interference and ensure the accuracy of information such as the QR code's bounding box and corner coordinates. A fast QR code recognition library is then used to locate and decode the QR code region, recording the bounding box coordinates, center point coordinates, and area of the QR code in each frame. The frame images are then converted to grayscale for subsequent processing. Invalid frames with blurry or unrecognized QR codes are filtered out to reduce the computational effort of subsequent feature extraction.
[0027] Specifically, load the video and extract all frames, obtaining the image of each frame. For each raw image, use a lightweight QR code detection library (such as pyzbar's decode function) to quickly detect the presence of a QR code. Filter out valid frames containing QR codes (at least five frames) and record the timestamps. Perform image preprocessing on the valid frames and use a fast QR code recognition library (such as zxing) to detect the QR code. Obtain the bounding box ((x, y, width, height)), where x and y are the coordinates of the upper-left corner of the QR code's rectangular area. Width and height are the width and height of the QR code's rectangular area. Obtain the center coordinates using the bounding box calculation (x + w / 2, y + h / 2). Calculate the bounding box area (width × height). Record the bounding box coordinates, center coordinates, and area of the QR code in each frame.
[0028] S2: Feature extraction; S21: Extraction of spatiotemporal features of QR code regions; S211: Position / posture change analysis; Calculate the movement trajectory of the QR code's bounding box center point coordinates between adjacent frames, the change in bounding box area, and the motion vectors (i.e., the displacement vector of the center point) between adjacent frames. Analyze the displacement magnitude and area change rate. Determine whether the QR code exhibits two-dimensional or three-dimensional motion.
[0029] Based on the bounding box coordinates, center point coordinates, and area of the QR code in adjacent frames (t-1 and t), the motion features are calculated: Let the center coordinates of the valid frame t be , the area of the bounding box ; The center coordinates of the valid frame t-1 , the area of the bounding box ; 1) Obtain the displacement vector of the center point of the QR code bounding box of adjacent frames and calculate the displacement size; According to the displacement vector of the center point of the QR code bounding box of the adjacent frame: , calculate the displacement magnitude (Euclidean norm) ;in is the displacement vector of the center point of the QR code bounding box of the adjacent frame, is a vector Traverse all adjacent valid frames and obtain the displacement vector and displacement size of the center point of the bounding box of the entire frame sequence; 2) Obtain the area change rate of the bounding box of adjacent frames; According to the formula, the rate of change of the bounding box area is: , traverse all adjacent frames to obtain the bounding box area change rate.
[0030] 3) Calculate the average displacement of the entire frame sequence; Get the displacement of the center point of the bounding box of the entire frame sequence and calculate the average value; Average displacement of the entire frame sequence: ; Where N is the number of valid frames, is the displacement of the center point of the bounding box.
[0031] 4) Count the maximum displacement of the entire frame sequence; Compare the displacements of the center points of the bounding boxes of the entire frame sequence and obtain the maximum value;
[0032] 5) Calculate the displacement standard deviation of the entire frame sequence; Obtain the displacement of the center point of the bounding box of the entire frame sequence and calculate the standard deviation; Standard deviation of displacement:
[0033] 6) Calculate the average area change rate of the entire frame sequence; Obtain the area change rate of the entire frame sequence and calculate the average value; Average area change rate:
[0034] 7) Calculate the maximum area change rate of the entire frame sequence; Compare the area change rates of the entire frame sequence and obtain the maximum value; Maximum area change rate:
[0035] S212: Surface deformation consistency analysis; The cosine similarity between all non-zero motion vectors (i.e., the cosine value of the angle between the vectors) is calculated, and then the average cosine similarity of the motion vectors between all frames in the entire sequence is calculated as the motion consistency indicator.
[0036] S212a: Obtain a non-zero motion vector; Get the displacement vectors of the center points of the QR code bounding boxes of all adjacent frames and get the motion vector set Then remove vectors with a modulus of approximately zero (e.g. ), to avoid interference with consistency calculation, obtain valid motion vectors, and group the valid vectors into a non-zero motion vector set { , ,..., }; S212b: Calculate the cosine similarity of any two different vectors in the set of non-zero motion vectors; Use a double loop to traverse the non-zero motion vector set and filter out all those that meet i ≠ j The effective vector pair ( , ); Calculate the cosine similarity of all valid vector pairs; The calculation formula is:
[0037] S212c: obtaining motion consistency index; The average of the cosine similarities of all vector pairs is calculated to obtain the motion consistency index.
[0038] Get the cosine values of all vector pairs to form a motion vector cosine similarity set { , ,..., }; according to Calculate and obtain motion consistency index; Where m is the number of valid vector pairs, is the kth motion vector cosine similarity in the motion vector cosine similarity set; Valid vector pairs , where n is the number of vectors in the set of non-zero motion vectors.
[0039] S213: Determine the movement mode; S213a: Determine whether it is a 2D motion flag; Determine whether the QR code's motion is two-dimensional plane motion: If the following conditions are met: average displacement ≤ 15 pixels, displacement standard deviation < 10 pixels, average area change rate ≤ 5%, and motion consistency > 0.7, then it is determined to be two-dimensional plane motion. If not, then it is determined to be negative, indicating that it is not two-dimensional plane motion.
[0040] The judgment result is converted to 1.0 or 0.0 through a Boolean value. When the judgment result is a two-dimensional plane motion, it is converted to 1.0, otherwise, it is converted to 0.0.
[0041] S213b: Determine whether it is a rigid plane mark; Determine whether the QR code is attached to a rigid plane: If the following conditions are met: average area change rate ≤ 3% and maximum area change rate ≤ 10% and displacement standard deviation < 8 pixels, then it is determined to be attached to a rigid plane (rigid plane flag is true). If not, then it is determined to be no and not attached to a rigid plane.
[0042] Convert the judgment result to 1.0 or 0.0 through Boolean value. If the judgment result is attached to the rigid plane, it is converted to 1.0, otherwise, it is converted to 0.0.
[0043] S214: eigenvalue output; The characteristic values are summarized and output, including: average displacement, displacement standard deviation, average area change rate, maximum area change rate, motion consistency index, 2D motion mark, and rigid plane mark.
[0044] S22: Extraction of background features around the QR code; S221: Background motion consistency analysis; A background region is selected outside the QR code bounding box. The Farneback optical flow method is used to calculate the optical flow field of the background region and the QR code region. The difference (i.e., motion difference) between the average optical flow vector of the background region and the average optical flow vector of the QR code region is compared.
[0045] In three-dimensional objects, the background and the QR code area typically exhibit coherent or correlated motion patterns (with minimal motion variability). In two-dimensional images, however, the QR code area may exhibit independent motion relative to the background (with significant motion variability). Using the feature vector of this motion pattern ensures accurate judgment, making it suitable for complex scene analysis.
[0046] S221a: Calculate the optical flow field of the background area and the QR code area; First, obtain the QR code area and background area: extract the grayscale images of a set of adjacent frames (t and t-1) of the middle valid frame, and crop the rectangular area as the QR code area based on the coordinates of the QR code's bounding box. Select a background area outside the QR code bounding box (such as the right or bottom area, which is similar in area to the QR code area and does not overlap, and is 0.5 to 1 times the area of the QR code area) and crop it as the background area.
[0047] Next, feature points are extracted: feature points are extracted from the QR code area and the background area of the valid frame t-1 respectively, and a feature point set of the QR code area and a feature point set of the background area of the valid frame t-1 are obtained; Extract feature points of the two-dimensional code region and the background region of the effective frame t by using feature point matching, and obtain a feature point set of the two-dimensional code region and a feature point set of the background region of the effective frame t; The feature point set of the QR code region of the valid frame t-1 is denoted as {p j} (j = 1 to M, M is the number of feature points in the QR code area) and the background area feature point set {p k} (k = 1 to N, N is the number of feature points in the background area); The feature point set of the QR code region of the valid frame t is denoted as {p j '} (j = 1 to M, M is the number of feature points in the QR code area) and the background area feature point set {p k '} (k = 1 to N, N is the number of feature points in the background area); Finally, the optical flow vectors of the feature points in the QR code area and the background area are calculated respectively; Let the optical flow vector of the feature point in the QR code area be recorded as , the optical flow vector of the feature points in the background area is recorded as ; Obtained based on the difference between the feature points matching the QR code area , obtained based on the difference between the feature points matching the background area ; Right now ; .
[0048] Based on the analysis of feature points and optical flow fields, the calculation is smaller and faster, without relying on high-performance hardware. The optical flow field calculation uses local feature point analysis rather than full-screen calculation, further reducing computing power consumption and achieving real-time response.
[0049] S221b: Calculate the average optical flow vector of all feature points in the QR code area and the background area respectively; Average optical flow vector of the QR code area: ; Average optical flow vector of background area: ; in: Indicates the QR code area The optical flow vector of feature points, Represents the background area The optical flow vector of feature points, Indicates the number of feature points in the QR code area. Indicates the number of feature points in the background area.
[0050] S221c: calculating motion difference; The motion difference is obtained by calculating the modulus of the difference between the average optical flow vector of the QR code area and the average optical flow vector of the background area.
[0051] according to: Get the motion difference.
[0052] By incorporating background motion comparison and background motion consistency analysis through regional correlation verification, we can reduce the limitations of single features and improve the algorithm's adaptability to complex scenarios. For example, some high-quality printed images (2D carriers) may simulate 3D motion patterns, but the motion difference between the QR code area and the background area is greater, while in 3D objects, the motion of the two areas is more coherent, further eliminating advanced cheating behaviors.
[0053] S222: Obtain background depth clues: Use a lightweight depth estimation model to output a depth map and obtain the depth gradient of the QR code boundary; The following steps are involved: Step 1: Preprocessing of valid frame images; Obtain the grayscale image of the valid frame (intermediate valid frame can be selected), and generate an image that meets the input requirements of the lightweight MiDaS model through resizing and standardization. The standardization process includes pixel value normalization and format conversion.
[0054] Step 2: Depth map generation; The pre-processed valid frame image is input into the lightweight MiDaS model, and the MiDaS model infers and outputs the depth map D(x,y), where (x,y) represents the image pixel coordinates. .
[0055] Step 3: Extract the boundary of the QR code; According to the obtained bounding box, the four sides of the QR code bounding box are discretized and sampled by linear interpolation method to obtain the pixel point set Γ of the QR code boundary; Step 4: Depth gradient calculation; The Sobel operator is used to calculate the gradient of the depth map D (x, y) to obtain the horizontal gradient map Gx (x, y) and the vertical gradient map Gy (x, y). The coordinates of all pixels on the boundary of the QR code are extracted, and the gradient modulus of the extracted pixels is calculated. , we get the boundary gradient set {‖∇D(x,y)‖|(x,y)∈Γ}; Step 5: Get the depth gradient at the boundary of the QR code: Obtain the maximum value in the boundary gradient set as the depth gradient at the boundary of the QR code; ; Depth map analysis based on the lightweight MiDaS model captures the depth relationship between the QR code and the background. While the depth of the two is continuous in 3D objects (depth gradient ≤ 0.3 at the boundary), 2D objects exhibit abrupt depth changes (depth gradient > 0.3 at the boundary), significantly improving the ability to detect advanced cheating techniques, such as high-quality composite images. The lightweight MiDaS model quickly outputs a coarse depth map on mobile devices, eliminating the need for cloud-based computing. Furthermore, analyzing the depth gradient at the QR code boundary, rather than the entire image, further reduces computational effort and ensures real-time algorithm performance.
[0056] S223: Background stability analysis; The Pearson correlation coefficient is used to calculate the linear correlation of the grayscale values of the background areas of two adjacent frames, that is, the correlation coefficient of the background area: ,in and The grayscale values of the background area of the previous and current frames are obtained. The cropped background areas are obtained for the previous frame (valid frame t-1) and the current frame (valid frame t). The grayscale values of each pixel in the background areas of the previous and current frames are obtained. The corrcoef function is called based on the grayscale values to directly calculate the correlation coefficient for the background area. This enables fast calculations, improves the algorithm's real-time responsiveness, reduces waiting time, and minimizes time consumption, ensuring that the total analysis time is less than 5 seconds.
[0057] S23: Extract image quality and anti-spoofing features; Print / Screen Characteristic Detection: This function selects the QR code region within the middle valid frame for analysis. This analysis measures high-frequency noise energy (calculated using Fourier transform), moiré effects (calculated using the Laplace operator to calculate variance), and edge sharpness (calculated using the Sobel operator to calculate average gradient). It also measures color channel statistics (calculated using the variance of each channel in the HSV space) and texture complexity (derived from the standard deviation of the grayscale image).
[0058] Selecting the middle valid frame avoids the quality fluctuation of the edge valid frames and ensures the reliability of the basic detection data. The middle valid frame is in the stable stage of video recording, providing accurate and interference-free basic data for print / screen feature detection, avoiding detection errors caused by frame quality fluctuations, ensuring feature representativeness, and avoiding accidental errors in a single frame.
[0059] Supplementing the multi-dimensional identification basis can quickly identify cheating behaviors such as using a mobile phone to shoot the QR code on the screen, make up for the deficiencies of spatiotemporal and background features, deal with new cheating methods, respond to the upgrade of printing synthetic QR code technology, and simulate three-dimensional motion or depth feature shooting scenarios.
[0060] S3: Feature fusion and classification; S31: extract feature vector; The three types of features extracted in step S2 are combined into a 12-dimensional feature vector, which is specifically composed as follows: 6-dimensional features from S21 (QR code area spatiotemporal features): Displacement normalization value: the average displacement Normalized: Preset the expected maximum displacement value, and the average displacement Divide the expected maximum displacement by the preset value to obtain the normalized displacement value, and convert the average displacement The mapping is in the range [0, 1]; Average area change rate: (direct value); Motion consistency index: (direct value); 2D Motion Logo: (Boolean value converted to 1.0 or 0.0); Rigid flat signs: (Boolean value converted to 1.0 or 0.0); Normalized value of displacement standard deviation: Normalized: preset maximum standard deviation expected value, obtained by displacement standard deviation Divide by the preset maximum standard deviation expected value to map the displacement standard deviation to the range of [0, 1].
[0061] 3D features from S22 (QR code surrounding background features): Motion difference normalization value: optical flow difference between background area and QR code area ( ) is obtained by normalization processing: the maximum difference expected value is preset, the optical flow difference between the background area and the QR code area is divided by the preset maximum difference expected value, and the optical flow difference is mapped in the range of [0, 1].
[0062] Depth gradient normalization value: the depth gradient at the boundary of the QR code ( ) is normalized: the maximum expected gradient value is preset, and the depth gradient at the boundary of the QR code ( ) is divided by the preset maximum gradient expectation value to map the depth gradient to the range of [0, 1].
[0063] Background stability: The correlation coefficient of the background area ( , directly get the value); 3D features from S23 (image quality and anti-spoofing features): High-frequency energy: high-frequency noise energy in the QR code area (direct value); Moiré intensity: Moiré effect intensity in the QR code area (direct value); Edge sharpness: the average edge gradient of the QR code area (direct value); After multi-feature fusion, the algorithm can conduct a comprehensive analysis from three dimensions: motion pattern, depth relationship, and physical properties, significantly improving the accuracy of distinguishing three-dimensional objects from two-dimensional carriers.
[0064] Normalization is a simple process, requiring only simple division. The logic is straightforward, requiring minimal computational overhead and requiring minimal computing power. It also minimizes the time required for feature processing. It ensures that all feature vectors fall within the range [0, 1], preventing differences in feature magnitude from dominating model judgments. This ensures that each feature contributes equally to model judgments, accelerating model training convergence. This ensures more balanced gradients across features and more stable weight updates. This significantly reduces model training time, facilitates rapid deployment of lightweight models, and improves model generalization. It also provides more stable predictions for feature data from new scenarios. For example, the distribution of feature values in product inspection and document verification scenarios may differ. Normalization ensures that the model maintains consistent recognition logic across these scenarios, preventing accuracy degradation due to these differences. It also mitigates outlier interference, preventing biased model judgments caused by individual extreme data points. This makes it particularly suitable for complex scenarios such as product inspection and document verification.
[0065] S32: Feature vector input classification model; This feature vector is input into the classification model, a pre-trained lightweight CNN classifier (three-layer convolution, 128-dimensional input features, ReLU activation function). The output is the probability of a 3D physical carrier and a 2D flat carrier. The three-layer convolutional structure avoids the high computational power of deep networks. The 128-dimensional feature vector reduces data processing overhead, and the ReLU activation function is fast, enabling classification on mobile devices within 5 seconds while ensuring accurate classification (after initial feature optimization, the classification accuracy is close to that of deep networks).
[0066] Pre-training is performed by collecting samples from product anti-counterfeiting inspections and document verification, including 3D physical carrier samples (photographing QR codes on real products and physical documents, recording 2-3 second videos, and extracting 12-dimensional feature vectors) and 2D flat carrier samples (printing QR code images, displaying QR codes on screen, and also extracting 12-dimensional feature vectors). The total number of samples is no less than 200, and each group of samples is annotated with the true category label. A lightweight CNN classifier model is pre-trained using these samples to obtain a pre-trained model. Training the classification model is prior art in this field and does not constitute the inventive solution of this application, so it is not detailed here.
[0067] The feature vector is input into the model, and forward propagation calculations are performed. The input layer passes the 12-dimensional vector into the first hidden layer, calculating the 24-dimensional hidden layer output. The first hidden layer output is passed into the second hidden layer, calculating the 12-dimensional hidden layer output, which is then passed into the output layer. Using the Softmax() function, a two-dimensional probability vector [P1, P2] is calculated, where P1 is the probability of a three-dimensional physical carrier and P2 is the probability of a two-dimensional planar carrier; P1 + P2 = 1. The classification logic is clear, and the results are intuitive and reliable, making it easy to determine the carrier type based on probability, without complex interpretation.
[0068] The classification model of the present invention is lightweight and can be quickly loaded and calculated on the mobile terminal, avoiding occupying too much memory and computing power, and the model training and deployment are easy.
[0069] S4: Result output and anti-cheating judgment; Receiving the output of the classification model and determining whether it is a three-dimensional physical carrier or a two-dimensional plane carrier; According to the probability vector, if P1>P2, it is determined to be a three-dimensional physical carrier; otherwise, it is determined to be a two-dimensional plane carrier.
[0070] If it is determined to be a "three-dimensional physical carrier", the verification is passed, confirming that the QR code is attached to a real object.
[0071] If it is determined to be a "two-dimensional flat carrier", the verification fails and it is determined that there is suspicion of cheating by using pre-stored pictures / screen displays. A pop-up window will be displayed on the user interface, prompting "The QR code is detected to be from a picture or screen. Please aim at the real object to rescan and verify."
[0072] Furthermore, if the number of valid frames is insufficient or the QR code cannot be extracted, the verification fails, and a prompt "QR code not detected, please aim at the physical object to rescan and verify" is displayed through the user interface.
[0073] Through user interface prompts, clear result feedback is provided and correct operations are guided, avoiding user confusion caused by "verification failure but unknown reasons" and improving user experience.
[0074] The present invention overcomes the shortcomings of existing QR code inspection technology, such as the inability to identify the physical properties of the carrier, slow response speed, and susceptibility to deception by pre-stored pictures / videos. By forcing the mobile terminal to shoot a short video (usually 2-3 seconds) containing the target QR code in real time, and quickly analyzing the dynamic visual features of the QR code area and its surrounding background in the video frame sequence (total time ≤ 5 seconds), the present invention determines whether the QR code is attached to the surface of a three-dimensional physical object or only exists on a two-dimensional plane medium (picture or screen), thereby preventing cheating using pre-stored QR code pictures / videos from the source.
[0075] The present invention has been described with reference to the above embodiments. However, the above embodiments are merely exemplary embodiments of the present invention. It should be noted that the disclosed embodiments do not limit the scope of the present invention. On the contrary, modifications and improvements that do not depart from the spirit and scope of the present invention are intended to be protected by the present invention.
Claims
1. A method for identifying and analyzing QR code attachment videos based on object perception, characterized by: It includes a mobile terminal, on which a QR code scanning and verification module is installed; The following steps are also included: S1: Real-time video acquisition and preprocessing; S11: Startup and constraints of the QR code scanning verification module; Start the QR code scanning and verification module on the mobile terminal, set the permissions of the QR code scanning and verification module, limit the use of only the rear camera and block the album access interface, and the user uses the QR code scanning and verification module to aim at the object with the QR code to be identified and take a picture; S12: short video recording; After the user triggers the shooting, short videos of a certain length are recorded continuously; S13: extraction and decoding of key frames; From the recorded video, samples are taken at fixed intervals to extract valid frames with successful QR code recognition. Each frame is then preprocessed, and the QR code area is located and decoded using a fast QR code recognition library. The bounding box coordinates, center point coordinates, and area of the QR code in each frame are recorded. The frame image is then converted into a grayscale image. S2: Feature extraction; S21: Extraction of spatiotemporal features of QR code regions; S211: Position / posture change analysis; S212: Surface deformation consistency analysis; S213: Determine the movement mode; S214: eigenvalue output; S22: Extraction of background features around the QR code; S221: Background motion consistency analysis; S222: Obtain background depth clues: Use a lightweight depth estimation model to output a depth map and obtain the depth gradient of the QR code boundary; S223: Background stability analysis; S23: Extract image quality and anti-spoofing features; Select the QR code area in the middle valid frame for analysis, and analyze the high-frequency noise energy, moiré effect, and edge sharpness of the QR code area; S3: Feature fusion and classification; S31: extract feature vector; S32: Feature vector input classification model; S4: Result output and anti-cheating judgment; Receive the output result of the classification model and determine whether it is a three-dimensional physical carrier or a two-dimensional plane carrier.
2. The object-aware QR code attachment video recognition and analysis method according to claim 1, wherein: Step S211 is to calculate motion features based on the bounding box coordinates, center point coordinates and area of the QR code in adjacent frames. The motion features include: the displacement vector and displacement size of the bounding box center point of the entire frame sequence, the average displacement, the maximum displacement, the displacement standard deviation, the bounding box area change rate, the average area change rate, and the maximum area change rate.
3. The object-aware QR code attachment video recognition and analysis method according to claim 2, wherein: Step S212 is to calculate the cosine similarity between all non-zero motion vectors, and then calculate the average value of the cosine similarity of all inter-frame motion vectors in the entire sequence as a motion consistency indicator; S212a: Obtain a non-zero motion vector; Get the displacement vectors of the center points of the QR code bounding boxes of all adjacent frames, get the motion vector set, and remove the modulus The motion vector of the spherical axis is obtained to obtain a valid motion vector, and the valid vectors are combined into a non-zero motion vector set; S212b: Calculate the cosine similarity of any two different vectors in the set of non-zero motion vectors; Filter out valid vector pairs and calculate the cosine similarity of all valid vector pairs; S212c: obtaining motion consistency index; The average of the cosine similarities of all valid vector pairs is calculated to obtain the motion consistency index.
4. The object-aware QR code attachment video recognition and analysis method according to claim 3, wherein: Also includes: S213a: Determine whether it is a 2D motion flag; Determine whether the QR code's motion is two-dimensional plane motion: If the following conditions are met: average displacement ≤ 15 pixels, displacement standard deviation < 10 pixels, average area change rate ≤ 5%, and motion consistency > 0.7, then it is determined to be two-dimensional plane motion. If not, then it is determined to be no, not two-dimensional plane motion, and the determination result is converted to 1.0 or 0.0 using a Boolean value; S213b: Determine whether it is a rigid plane mark; Determine whether the QR code is attached to a rigid plane: If the following conditions are met: average area change rate ≤ 3% and maximum area change rate ≤ 10% and displacement standard deviation < 8 pixels, then it is determined to be attached to a rigid plane. If not, then it is determined to be no, not attached to a rigid plane, and the determination result is converted to 1.0 or 0.0 through a Boolean value.
5. The object-aware QR code attachment video recognition and analysis method according to claim 1, wherein: Step S221: Select a background area outside the QR code bounding box, calculate the optical flow field of the background area and the QR code area using the Farneback optical flow method, and compare the difference between the average optical flow vector of the background area and the average optical flow vector of the QR code area; S221a: Calculate the optical flow field of the background area and the QR code area; First, the QR code area and background area are obtained: grayscale images of a group of adjacent frames of the middle valid frame are extracted, and the rectangular area is cropped as the QR code area according to the coordinates of the QR code's bounding box; a background area is selected outside the QR code's bounding box; Next, feature points are extracted: first, feature points of the QR code area and background area of a valid frame are extracted, and then feature points of the QR code area and background area of another frame are extracted using feature point matching; Finally, the optical flow vectors of the feature points in the QR code area and the background area are calculated respectively; The optical flow vector of the QR code area is obtained based on the difference between the feature points matching the QR code area, and the optical flow vector of the background area is obtained based on the difference between the feature points matching the background area; S221b: Calculate the average optical flow vector of all feature points in the QR code area and the background area respectively; S221c: calculating motion difference; The motion difference is obtained by calculating the modulus of the difference between the average optical flow vector of the QR code area and the average optical flow vector of the background area.
6. The object-aware QR code attachment video recognition and analysis method according to claim 1, wherein: Step S222 includes: Step 1: Preprocessing of valid frame images; Obtain the grayscale image of the valid frame, and generate an image that meets the input requirements of the lightweight MiDaS model through resizing and normalization; Step 2: Depth map generation; The pre-processed valid frame image is input into the lightweight MiDaS model, and the MiDaS model infers and outputs a depth map; Step 3: QR code boundary extraction; According to the obtained bounding box, the four sides of the QR code bounding box are discretized and sampled by linear interpolation method to obtain the pixel point set of the QR code boundary; Step 4: Depth gradient calculation; The Sobel operator is used to calculate the gradient of the depth map, and the gradient modulus of the extracted pixel points is calculated to obtain the boundary gradient set; Step 5: Get the depth gradient at the boundary of the QR code: The maximum value in the boundary gradient set is obtained as the depth gradient at the boundary of the QR code.
7. The object-aware QR code attachment video recognition and analysis method according to claim 1, wherein: Step S223 includes: obtaining the grayscale value of each pixel in the background area of the previous frame and the current frame, calling the corrcoef function to directly calculate the correlation coefficient according to the grayscale value, and obtaining the correlation coefficient of the background area.
8. The object-aware QR code attachment video recognition and analysis method according to claim 4, wherein: Eigenvector It is a 12-dimensional vector, specifically composed of: 6-dimensional features from S21: Displacement normalization value: the average displacement normalization process is mapped to the range of [0, 1]; Average area change rate, directly taken value; Motion consistency index, directly taking values; 2D motion flag: 1.0 or 0.0; Rigid plane flag: 1.0 or 0.0; Normalized displacement standard deviation value: mapped to the range of [0, 1] by normalization of displacement standard deviation; 3D features from S22: Motion difference normalization value: Motion difference normalization processing is mapped in the range of [0, 1]; Depth gradient normalization value: The depth gradient normalization processing at the boundary of the QR code is mapped in the range of [0, 1]; Background stability, directly obtain the correlation coefficient of the background area; 3D features from S23: High-frequency energy: The high-frequency noise energy in the QR code area is directly measured; Moiré intensity: directly obtain the intensity of the moiré effect in the QR code area; Edge sharpness: The average edge gradient of the QR code area is directly taken.
9. The object-aware QR code attachment video recognition and analysis method according to claim 1, wherein: The classification model is a pre-trained lightweight CNN classifier with three layers of convolution, an input feature dimension of 128, a ReLU activation function, and outputs the probability of a three-dimensional physical carrier and a two-dimensional plane carrier.
10. The object-aware QR code attachment video recognition and analysis method according to claim 1, wherein: Step S4 is specifically as follows: If it is determined to be a "three-dimensional physical carrier", the verification is successful, confirming that the QR code is attached to a real object; If it is determined to be a "2D flat carrier", verification fails and is suspected of cheating by using a pre-stored image or screen display. A pop-up window will appear on the user interface, prompting "The QR code may be from an image or screen. Please scan it again with the real object for verification." If the number of valid frames is insufficient or the QR code cannot be extracted, the verification fails and the user interface prompts "QR code not detected, please aim at the physical object and rescan for verification"; if the number of valid frames N is ≥ 5, the verification is considered to have failed.
Citation Information
Patent Citations
Anti-cheating sign-in method and device, computer system and readable storage medium
CN110163314A
Method and device for realizing authenticity query through one-step code scanning
CN114298257A
Anti-counterfeiting traceability method and device based on structural deformation and image watermark
CN115601216A
Authenticity determination system and authenticity determination device
JP2025021519A