Two-dimensional code intelligent scanning method and system based on multi-mode perception

By combining feature extraction from visual, depth, and radar data with evidence theory through multimodal perception technology, the problem of QR code recognition in complex scenarios has been solved, achieving efficient and reliable scanning results, which are suitable for industrial and logistics fields.

CN121009909AActive Publication Date: 2025-11-25GUANGZHOU XUNBAO ELECTRONICS TECH CO LTD

Patent Information

Application Number
CN202510977625.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-11-25
Estimated Expiration
2045-07-16

AI Technical Summary

Technical Problem

Traditional QR code scanning technology is inefficient in complex scenarios such as strong reflections, obstructions, and low light, making it difficult to meet the needs of practical applications.

Method used

A multimodal perception method is adopted to extract texture, curvature and normal vector features by acquiring visual images, depth images, infrared reflectance intensity maps and radar point cloud data, and use the Dempster Shafer evidence theory fusion machine to calculate credibility and dynamically switch scanning modes.

Benefits of technology

It improves the environmental adaptability and recognition reliability of QR code scanning, reduces manual intervention, and significantly improves scanning efficiency and accuracy, making it suitable for high-precision scenarios such as industrial production and logistics warehousing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121009909A_ABST
    Figure CN121009909A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent two-dimensional code scanning method and system based on multi-modal perception, relates to the technical field of two-dimensional code scanning, and solves the technical problem of low recognition efficiency in a complex scene in the prior art. The method comprises the following steps: collecting a visual image, a depth image, an infrared reflection intensity graph and radar point cloud data of a to-be-scanned object to obtain multi-dimensional data; performing feature extraction on the multi-dimensional data to obtain multi-dimensional features; the multi-dimensional features are input into a Dempster Shafer evidence theory fusion device, and credibility is output; and deciding the scanning mode of the to-be-scanned object based on the credibility. The method can be applied to industrial production, logistics storage and other processes with high scanning precision requirements.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of two-dimensional code scanning, and specifically relates to a two-dimensional code intelligent scanning method and system based on multi-modal perception. BACKGROUND

[0002] Nowadays, two-dimensional code scanning technology has been widely applied to many fields such as industrial production, logistics and warehousing, mobile payment and the like. However, most of the traditional two-dimensional code scanning methods are based on single-modal perception, such as relying only on visual images. In the face of complex scenes such as strong light reflection, object occlusion and low light, significant technical defects are exposed. Specifically, when there is strong light reflection on the surface of a two-dimensional code, single-modal visual acquisition will cause the image to be overexposed due to light reflection, the gray scale features of black and white modules are blurred, and thus the decoding algorithm fails; in the case that a two-dimensional code is partially occluded, single-modal perception cannot obtain complete image information, which easily causes incomplete feature extraction and leads to a significant decrease in recognition accuracy; and in a low-light environment, the contrast of visual images is reduced, and texture features become unobvious, which also makes scanning and recognition difficult, and it is difficult to meet the actual application requirements. SUMMARY

[0003] The application provides a two-dimensional code intelligent scanning method and system based on multi-modal perception, which solves the technical problem of low recognition efficiency of existing two-dimensional code scanning technology in complex scenes.

[0004] To achieve the above-mentioned purpose, the application adopts the following technical solutions: In a first aspect, a two-dimensional code intelligent scanning method based on multi-modal perception is provided, comprising: acquiring a visual image, a depth image, an infrared reflection intensity map and radar point cloud data of a to-be-scanned object to obtain multi-dimensional data; performing feature extraction on the multi-dimensional data to obtain multi-dimensional features, wherein the multi-dimensional features include texture features extracted based on the visual image, curvature features calculated based on the infrared reflection intensity map and the depth image, and normal vector features calculated based on the radar point cloud data; inputting the multi-dimensional features into a Dempster Shafer evidence theory fusion device to output a credibility; deciding a scanning mode of the to-be-scanned object based on the credibility, wherein the scanning mode includes a direct scanning mode, a light reflection suppression mode, an occlusion compensation mode and a low light enhancement mode.

[0005] It should be noted that the Dempster Shafer evidence theory fusion unit of this application is a computational framework for realizing multimodal feature confidence fusion. Its core is to convert the texture features from visual images, the curvature features from infrared and depth images, and the normal vector features from radar point cloud data into the confidence of the QR code decodeable state by constructing an identification framework, defining a basic probability allocation model and evidence combination rules, based on Dempster Shafer evidence theory.

[0006] Based on the above technical solutions, this application provides a multimodal perception-based intelligent QR code scanning method. By collecting visual images, depth images, infrared reflectance intensity maps, and radar point cloud data of the object to be scanned, a multidimensional data fusion system is constructed. This effectively solves the recognition failure problem of traditional single-modal scanning in complex scenarios such as strong reflection, occlusion, and low light. This method extracts multidimensional features such as texture, curvature, and normal vectors, and uses Dempster-Shafer evidence theory for credibility fusion. This allows for accurate determination of the QR code's decodeability, enabling the adoption of different scanning modes for intelligent QR code recognition and scanning. This combination of multimodal perception and intelligent decision-making mechanisms not only improves the environmental adaptability and recognition credibility of QR code scanning but also reduces manual intervention through automated mode switching, significantly improving scanning efficiency and reliability. This provides technical support for scenarios with high scanning accuracy requirements, such as industrial production and logistics warehousing.

[0007] In conjunction with the first aspect above, in one possible implementation, the method for obtaining the texture features includes: The visual image is divided into blocks to obtain multiple image blocks; Extract the local binary pattern (LBP) feature vector for each image patch; Calculate the statistical features within each image patch, including the contrast, entropy, and energy of the image patch's gray-level co-occurrence matrix; By combining LBP feature vectors with statistical features, texture features of visual images can be obtained.

[0008] In conjunction with the first aspect above, in one possible implementation, the method for obtaining the curvature feature includes: Calculation of infrared reflectance intensity diagram I ir gradient magnitude and the eigenvalues ​​λ1 and λ2 of the Hessian matrix; Based on gradient magnitude The curvature feature K is calculated using the Hessian matrix eigenvalues ​​λ1 and λ2 and the depth map D(x,y) according to the curvature feature calculation formula.

[0009] In conjunction with the first aspect above, in one possible implementation, the curvature feature calculation formula satisfies: Among them, S pix represents the physical size of a pixel in mm, and μ represents the depth weighting coefficient, with a value range of [0.2, 0.4].

[0010] In conjunction with the first aspect above, in one possible implementation, the method for obtaining the normal vector features includes: Build a KD Tree index for each point in the radar point cloud data; Based on point p i The set of neighboring points N(p) within the spatial location search radius r i ); Calculate the neighborhood point set N(p) i The covariance matrix of the covariance matrix is ​​used, and the eigenvector corresponding to the smallest eigenvalue in the covariance matrix is ​​taken as the point p. i The normal vector; By statistically analyzing the histogram of the normal vector direction for each point, the characteristics of the normal vector are obtained.

[0011] In conjunction with the first aspect above, in one possible implementation, the input of multidimensional features into the Dempster-Shafer evidence theory fusion processor includes: Construct a recognition framework Θ={m1="decodeable",m2="requires enhancement",m3="undecodeable"}; where the recognition framework is used to determine the category to which the QR code scanning result of the object to be scanned belongs; Calculate texture features F respectively t Curvature feature F k , Normal vector feature F n The basic probability m i (A); where the basic probability is represented in feature F i∈{t,k,n} Under the given conditions, determine the confidence level of the QR code scanning result of the object to be scanned to belong to category A∈Θ in the recognition frame; Based on the evidence combination rules of Dempster-Shafer evidence theory, multi-dimensional features are fused to obtain the fused basic probability m. total (A); where, the fusion basic probability represents the comprehensive confidence level of determining whether the QR code scanning result of the object to be scanned belongs to category A after integrating the basic probabilities of multidimensional features; The credibility level Bel=m is calculated according to the formula. total (m1)+λ×m total (m2); where λ represents the weighting coefficient.

[0012] In conjunction with the first aspect above, in one possible implementation, the basic probability m i The formula for (A) is: ;in, This indicates that feature i is the preset baseline feature value of category A, and σ represents the scaling parameter; The fusion basic probability m total The formula for (A) is: ;in, Represents the empty set. The symbol represents the intersection.

[0013] In conjunction with the first aspect above, in one possible implementation, the decision-making process based on the scanning mode of the object to be scanned based on confidence level includes: If the confidence level is greater than T1, then execute the direct scan mode; If T2 ≤ confidence level ≤ T1, then pattern matching is performed based on multidimensional features, including: When the curvature feature of the target region exists K max >K th At that time, the reflection suppression mode is activated; K max K represents the maximum curvature eigenvalue. th Indicates the preset curvature threshold; When the peak ratio of the normal vector histogram of the target region is <H th At that time, the occlusion compensation mode is executed; the normal vector histogram is obtained based on the statistical analysis of the normal vectors of all points, H th This indicates the preset peak-to-peak ratio threshold; When the texture features of the target region exist F t_max <F t_th At that time, low-light enhancement mode is executed; F t_max F represents the maximum texture feature value. t_th Indicates the preset texture feature threshold; If the credibility is less than T2, then the manual verification mode is triggered; The target area refers to the QR code scanning area, which is obtained by edge detection or image recognition of the image through a visual sensor; T1 represents a preset first threshold, and T2 represents a preset second threshold.

[0014] In conjunction with the first aspect above, in one possible implementation, the automatic execution operation of the reflection suppression mode includes: turning off the light of the visual sensor, increasing the infrared emission power to a preset value, and deflecting the scanning angle by a preset angle along the direction of the curvature feature change rate. The automatic execution operation of the occlusion compensation mode includes: controlling the radar sensor to perform point cloud scanning within a preset pitch angle range; The automatic execution of the low-light enhancement mode includes: activating the light of the vision sensor and activating infrared auxiliary illumination.

[0015] Secondly, a QR code intelligent scanning system based on multimodal perception is provided, comprising: The data acquisition module is used to acquire visual images, depth images, infrared reflectance intensity maps, and radar point cloud data of the object to be scanned, thereby obtaining multi-dimensional data. The feature extraction module is used to extract features from multidimensional data to obtain multidimensional features, which include texture features extracted based on visual images, curvature features calculated based on infrared reflectance intensity maps and depth images, and normal vector features calculated based on radar point cloud data. The fusion decision module is used to input multi-dimensional features into the Dempster Shafer evidence theory fusion engine and output the confidence level. The decision on the scanning mode of the object to be scanned is based on confidence level; the scanning mode includes direct scanning mode, reflection suppression mode, occlusion compensation mode and low light enhancement mode.

[0016] Thirdly, this application provides a QR code intelligent scanning device based on multimodal perception, including: a communication unit and a processing unit; The communication unit is used to receive visual images, depth images, infrared reflectance intensity maps, and radar point cloud data of the object to be scanned in order to obtain multi-dimensional data and send the decision results of the scanning mode. The processing unit is used to extract features from multidimensional data to obtain multidimensional features including texture features, curvature features and normal vector features; Input multidimensional features into the Dempster-Shafer evidence theory fusion engine to output credibility; The decision on the scanning mode of the object to be scanned is based on confidence level. The scanning modes include direct scanning mode, reflection suppression mode, occlusion compensation mode, and low light enhancement mode.

[0017] Fourthly, this application provides a QR code intelligent scanning device based on multimodal perception, comprising: a processor and a storage medium; the storage medium includes instructions, and the processor is configured to execute the instructions to implement the method described in the first aspect and any possible implementation thereof. This QR code intelligent scanning device can be an electronic device or a chip within an electronic device.

[0018] Fifthly, this application provides a computer-readable storage medium storing instructions that, when executed on a multimodal sensing-based QR code smart scanning device, cause the multimodal sensing-based QR code smart scanning device to perform the methods described in the first aspect and any possible implementation thereof.

[0019] In a sixth aspect, this application provides a computer program product containing instructions that, when the computer program product is run on a multimodal sensing-based QR code smart scanning device, causes the multimodal sensing-based QR code smart scanning device to perform the methods described in the first aspect and any possible implementation thereof.

[0020] This application provides a QR code intelligent scanning method and system based on multimodal perception, enabling adaptive and accurate decision-making for QR code recognition in complex environments. At the feature extraction level, a geometric feature quantization method fusing infrared reflection intensity and depth data transforms gradient abrupt changes and depth discontinuities caused by surface reflection or occlusion into calculable feature parameters, making the algorithm more sensitive to environmental interference. The statistical analysis method for normal vector features transforms the three-dimensional spatial direction distribution into histogram peak ratios, providing a quantitative basis for occlusion judgment. At the decision fusion level, a credibility calculation framework based on evidence theory effectively handles the conflict and complementarity of multi-source data by constructing a mathematical mapping model of multi-feature confidence. This allows the system to dynamically switch scanning modes based on the comprehensive confidence of multi-dimensional features such as texture, curvature, and normal vectors, avoiding the risk of misjudgment based on a single feature and achieving adaptive optimization of the scanning strategy. Ultimately, this improves the reliability and environmental adaptability of QR code scanning in complex scenarios such as strong reflection, low light, and partial occlusion.

[0021] It should be understood that the descriptions of technical features, technical solutions, beneficial effects, or similar language in this application do not imply that all features and advantages can be achieved in any single embodiment. Rather, it is understood that the description of a feature or beneficial effect means that a specific technical feature, technical solution, or beneficial effect is included in at least one embodiment. Therefore, the descriptions of technical features, technical solutions, or beneficial effects in this specification do not necessarily refer to the same embodiment. Furthermore, the technical features, technical solutions, and beneficial effects described in this embodiment can be combined in any suitable manner. Those skilled in the art will understand that embodiments can be implemented without one or more specific technical features, technical solutions, or beneficial effects of a particular embodiment. In other embodiments, additional technical features and beneficial effects may be identified in specific embodiments that do not embody all embodiments. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1A system architecture diagram of a QR code intelligent scanning method based on multimodal perception provided in this application embodiment; Figure 2 A flowchart illustrating a QR code intelligent scanning method based on multimodal perception provided in this application embodiment; Figure 3 A flowchart illustrating another QR code intelligent scanning method based on multimodal perception provided in this application embodiment; Figure 4 A flowchart illustrating another QR code intelligent scanning method based on multimodal perception provided in this application embodiment; Figure 5 A schematic diagram of a QR code intelligent scanning method based on multimodal perception provided in an embodiment of this application; Figure 6 This is a schematic diagram of the hardware structure of a QR code intelligent scanning method based on multimodal perception, provided in an embodiment of this application. Detailed Implementation

[0024] In the description of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. The "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Furthermore, "at least one" means one or more, and "multiple" means two or more. The terms "first," "second," etc., do not limit the quantity or order of execution, and "first," "second," etc., do not necessarily imply differences.

[0025] It should be noted that, in this application, the terms "exemplary" or "for example" are used to indicate that something is being described as an example, illustration, or illustration. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0026] The QR code intelligent scanning method based on multimodal perception provided in this application embodiment can be applied to, for example... Figure 1 In the QR code intelligent scanning system shown, based on multimodal perception, such as Figure 1 As shown, the communication system includes a data acquisition module, a feature extraction module, and a fusion decision module.

[0027] The data acquisition module is used to acquire visual images, depth images, infrared reflectance intensity maps, and radar point cloud data of the object to be scanned, thereby obtaining multi-dimensional data. The feature extraction module is used to extract features from multidimensional data to obtain multidimensional features, which include texture features extracted based on visual images, curvature features calculated based on infrared reflectance intensity maps and depth images, and normal vector features calculated based on radar point cloud data. The fusion decision module is used to input multi-dimensional features into the Dempster Shafer evidence theory fusion engine and output the confidence level. The decision on the scanning mode of the object to be scanned is based on confidence level; the scanning mode includes direct scanning mode, reflection suppression mode, occlusion compensation mode and low light enhancement mode.

[0028] To address the technical problem of QR code recognition failure in complex scenarios such as strong reflection, occlusion, and low light in existing technologies, this application provides a QR code intelligent scanning method based on multimodal perception, which includes: The system acquires visual images, depth images, infrared reflectance intensity maps, and point cloud data of the object to be scanned to obtain multidimensional data. Feature extraction is performed on multidimensional data to obtain multidimensional features, which include texture features extracted from visual images, curvature features calculated from infrared reflectance intensity maps and depth images, and normal vector features calculated from point cloud data. Input the multidimensional features into the Dempster-Shafer evidence theory fusion engine and output the credibility. The decision on the scanning mode of the object to be scanned is based on confidence level; the scanning modes include direct scanning mode, reflection suppression mode, occlusion compensation mode, and low light enhancement mode.

[0029] Based on this, the method of the present invention can improve the environmental adaptability and recognition reliability of QR code scanning through multimodal data fusion and intelligent decision-making mechanism.

[0030] like Figure 2 As shown in the embodiments of this application, the intelligent QR code scanning method based on multimodal perception includes: S1. Collect visual images, depth images, infrared reflectance intensity maps, and point cloud data of the object to be scanned to obtain multidimensional data.

[0031] Among them, visual images are obtained through visual sensors (such as cameras), depth images are obtained through structured light sensors or binocular stereo vision cameras, infrared reflection intensity maps are collected through infrared light sources in conjunction with infrared cameras, and point cloud data are obtained through LiDAR or millimeter-wave radar scanning.

[0032] In some implementations, the arrangement of the sensors can be coaxial integrated design, array distribution, or modular combination to ensure spatiotemporal synchronous acquisition of multi-source data. For example, cameras, infrared sensors, and lidar can be integrated on the same plane of the scanning device, and synchronous sampling can be achieved through hardware triggering.

[0033] S2. Extract features from the multidimensional data to obtain multidimensional features.

[0034] Among them, the multidimensional features include texture features extracted from visual images, curvature features calculated from infrared reflectance intensity maps and depth images, and normal vector features calculated from point cloud data.

[0035] In some implementations, texture features can be obtained through traditional methods such as Local Binary Pattern (LBP), Histogram of Oriented Gradients (HOG), or Gabor filters. Alternatively, convolutional neural networks (CNNs) can be used to extract end-to-end features from visual images, or statistics such as contrast and entropy can be calculated using the Gray-Level Co-occurrence Matrix (GLCM). Curvature features can be calculated through differential geometry methods (such as gradient-based surface fitting), or by using morphological operators to extract curvature changes at image edges. Normal vector features can be calculated through point cloud registration algorithms, octree indexing, or RANSAC (Random Sample Consensus) algorithms.

[0036] It should be noted that, in order to achieve multi-dimensional representation of the micro-texture, macro-geometry and spatial orientation of the QR code surface, feature extraction is performed on the acquired multi-dimensional data so that different modal data can effectively reflect the impact of environmental interference (such as reflection and occlusion) on the QR code features.

[0037] For example, in low-light scenarios, texture features extracted from infrared reflectance intensity maps can be used to replace visual images, avoiding feature loss due to insufficient light.

[0038] S3. Input the multidimensional features into the Dempster Shafer evidence theory fusion engine and output the credibility.

[0039] Among them, the Dempster Shafer Evidence Theory Fusion Machine is a multi-source information decision-making model built on evidence theory. By defining the identification framework of "decodeable", "needs to be enhanced" and "undecodeable", it transforms the confidence of each modality feature into basic probability assignment (BPA), and then uses evidence combination rules to process the conflict and complementarity of multi-source data, and finally outputs the comprehensive credibility.

[0040] In some implementations, improved Dempster combination rules (such as Yager rules and DSmT rules) can be used to handle highly conflicting evidence, or the feature weight coefficients can be adaptively adjusted through genetic algorithms or particle swarm optimization algorithms.

[0041] It should be noted that this fusion engine does not rely on prior probabilities and can improve the reliability of decisions through evidence accumulation, making it suitable for uncertain reasoning with multimodal data.

[0042] S4. Make a decision on the scanning mode of the object to be scanned based on credibility.

[0043] Among them, credibility reflects the overall confidence that the QR code can be correctly decoded in the current environment, and the scanning strategy is automatically switched based on a preset threshold.

[0044] In some implementations, the scanning mode can be to dynamically adjust sensor parameters (such as camera exposure time and infrared light source power) based on confidence level, or to trigger multi-sensor collaborative scanning (such as multi-angle lidar scanning and multi-spectral image fusion) or to activate hardware compensation mechanisms (such as a robotic arm adjusting the scanning angle).

[0045] It should be noted that the decision-making logic of the scanning mode needs to be decoupled from the hardware execution module in order to support custom policy configuration in different scenarios.

[0046] For example, in an industrial production line scenario, when the confidence level is in the T2T1 range and the curvature characteristics are abnormal, a reflection suppression mode can be automatically triggered. This mode eliminates reflection interference from the metal surface by turning off the visible light source, enhancing infrared illumination, and adjusting the scanning angle.

[0047] Based on the above technical solutions, the QR code intelligent scanning method based on multimodal perception provided in this application solves the technical bottleneck of traditional single-modal scanning in complex environments by constructing a processing system of "multi-source data acquisition, cross-modal feature extraction, evidence theory fusion, and adaptive mode decision-making". This method utilizes the complementary features of multi-dimensional data such as vision, depth, infrared, and point cloud, combined with the uncertainty reasoning ability of evidence theory, to achieve accurate judgment of QR code status and intelligent adaptation of scanning strategies. It effectively solves problems such as image distortion caused by strong reflection, feature loss caused by partial occlusion, and imaging blurring caused by low light, providing a universal technical solution for high-precision scanning needs in fields such as intelligent manufacturing and smart logistics.

[0048] In one possible implementation of this application embodiment, the above-mentioned S1 can be implemented through the following steps, which are described in detail below: S101. Acquire visual images of the object to be scanned using a vision sensor.

[0049] The visual sensor can be a CMOS (Complementary Metal-Oxide-Semiconductor) or CCD (Charge-Coupled Device) camera, which focuses on the area to be scanned through an optical lens and acquires RGB or grayscale images at a preset frame rate (e.g., 30fps) and resolution (e.g., 1920×1080). The camera's exposure time, gain, and other parameters can be automatically adjusted according to the ambient light intensity to avoid overexposure or underexposure.

[0050] In some implementations, timestamp alignment between visual images and other sensor data can be achieved through hardware triggering mechanisms (such as general purpose input / output interface signals GPIO signals) or software synchronization protocols (such as Network Time Protocol NTP) to ensure the spatiotemporal consistency of multimodal data.

[0051] It should be noted that the field of view of the visual sensor needs to cover the maximum spatial range where the QR code may appear, and the lens needs to have an anti-glare coating to reduce the impact of strong reflections on image quality.

[0052] For example, in logistics and warehousing scenarios, industrial-grade area array cameras can be used, and the frame rate can be dynamically adjusted in conjunction with the conveyor belt speed to ensure clear imaging of QR codes in motion.

[0053] S102. Use an infrared sensor and a depth sensor to acquire infrared reflection intensity map and depth image.

[0054] The infrared sensor consists of an infrared light source (such as an 850nm infrared LED) and an infrared camera. It generates an infrared reflection intensity map by emitting infrared light and receiving reflected signals, reflecting the reflection characteristics of an object's surface to infrared light. This is suitable for scenes with strong reflection or low light. The depth sensor can use a Time-of-Flight (ToF) camera or a structured light sensor. It calculates the three-dimensional coordinates of the object's surface by measuring the time of flight of light or pattern distortion, generating a depth image D(x,y), where (x,y) represents the spatial coordinates in millimeters.

[0055] In some implementations, infrared sensors and depth sensors can be integrated into the same module to improve data synchronization by sharing an optical path; alternatively, they can be deployed separately and data registration can be achieved through spatial calibration algorithms (such as hand-eye calibration).

[0056] It should be noted that the power of the infrared light source needs to be adjusted according to the needs of the scene. For example, the power should be reduced when scanning a metal surface to avoid infrared reflection saturation, and the power should be increased in a dark environment to enhance the signal strength.

[0057] For example, in a medical cold chain scenario, a ToF depth sensor combined with a narrowband infrared light source can penetrate condensed water mist to collect the infrared reflectance image and depth data of the QR code on the vaccine packaging box.

[0058] S103. Collect point cloud data using radar sensors and establish a KD Tree index.

[0059] The radar sensor can be a lidar system, which generates point cloud data containing three-dimensional coordinates (x, y, z) and reflection intensity by emitting laser or electromagnetic waves and receiving the echoes. For each point cloud point p... i A KD Tree index is built based on its spatial location for subsequent neighborhood point search and normal vector calculation.

[0060] In some implementations, the lidar can be either mechanically rotating (such as a 16-line lidar) or solid-state, and the point cloud density can be adjusted according to the scanning range and accuracy requirements (such as 10 points / square centimeter).

[0061] It should be noted that point cloud data acquisition needs to be synchronized with visual and infrared data in time. The sampling delay between sensors can be compensated by hardware clock (such as GPS clock) or software interpolation (such as linear interpolation).

[0062] For example, in an industrial robot sorting scenario, solid-state LiDAR is used to collect point cloud data of the package surface in real time. KD Tree indexing can accelerate point cloud registration and obstacle detection, and provide spatial location basis for occlusion compensation mode.

[0063] Based on the above technical solution, a multi-sensor collaborative acquisition and data synchronization mechanism was implemented to accurately acquire multi-dimensional data such as visual, infrared, depth, and point cloud data, providing a diverse information foundation for subsequent feature extraction and fusion decision-making. This data acquisition scheme not only adapts to complex environments such as strong reflections, low light, and partial occlusion, but also improves the spatiotemporal consistency of multimodal data through KD Tree indexing and hardware synchronization technology, laying a data foundation for intelligent QR code scanning.

[0064] In one possible implementation of the embodiments of this application, combined with Figure 2 ,like Figure 3 As shown, the above S2 can be implemented through the following S201, S202 and S203, which are explained in detail below: S201. Perform block processing on the visual image and extract texture features.

[0065] Specifically, the visual image is first divided into multiple 8×8 pixel image blocks, and local texture features are focused through block processing.

[0066] Then, a Local Binary Pattern (LBP) feature vector is calculated for each image patch: by comparing the gray values ​​of the center pixel with those of its 8 neighboring pixels, an 8-bit binary code is generated, and its histogram is used as the LBP feature.

[0067] Next, the gray-level co-occurrence matrix (GLCM) statistical features of the image patch are calculated, including contrast (reflecting texture sharpness), entropy (reflecting texture randomness), and energy (reflecting texture uniformity).

[0068] Finally, the LBP feature vector is combined with the GLCM statistical features to form the final texture feature vector.

[0069] In some implementations, the block size can be adaptively adjusted; for example, high-resolution images can use 16×16 pixel blocks to balance computational efficiency and feature accuracy.

[0070] S202. Calculate curvature features based on infrared reflection intensity map and depth image.

[0071] The formula for calculating curvature characteristics is as follows: ; In the formula: S pix The pixel physical size (unit: mm) is determined by the sensor parameters.

[0072] λ1 and λ2 are infrared reflectance intensity diagrams I ir The eigenvalues ​​of the Hessian matrix are used to characterize the second-order gradient changes of the image.

[0073] This represents the gradient magnitude of the infrared reflectance intensity map, reflecting the rate of change in pixel intensity.

[0074] D(x,y) is the depth map. and The second-order partial derivative of depth represents the trend of depth variation.

[0075] μ is the depth weighting coefficient, with a value range of [0.2, 0.4]. In highly reflective scenes, μ can be increased to enhance the influence of depth features.

[0076] In some implementations, the Hessian matrix can be preprocessed using Gaussian convolution kernels to reduce noise interference with gradient calculation.

[0077] It should be noted that the curvature feature K is used to characterize the local curvature of an object's surface. When there are reflective points on the QR code surface, a sudden change in infrared reflection intensity will lead to... As |λ1|+|λ2| increase, the value of K exceeds the threshold.

[0078] For example, if there are water droplets on the surface of the QR code causing reflection, the gradient amplitude of the infrared reflectance intensity map will increase significantly, and the calculated K value may reach 0.7. If a preset threshold K is set... th If the value is 0.5, the reflection suppression mode will be triggered in subsequent processes.

[0079] S203. Calculate normal vector features based on radar point cloud data. Specifically, this includes: (1) Establish a KDTree index for each point in the radar point cloud data to quickly search for neighboring points.

[0080] (2) Taking point p i Centered on the point cloud, search for a neighborhood point set N(p) within a radius r (0.510 mm, adjusted according to the point cloud density). i ).

[0081] (3) Calculate the neighborhood point set N(p) i The covariance matrix of ) is used to take the eigenvector corresponding to the smallest eigenvalue as the point p. i The normal vector represents the surface orientation at that point.

[0082] (4) Divide the 360° direction into 16 angular regions, and statistically analyze the normal vector direction histogram of all points to obtain the normal vector characteristics.

[0083] In some implementations, the size of the neighborhood point set must satisfy |N(p i )|≥5, to ensure the statistical significance of the covariance matrix calculation.

[0084] It should be noted that the peak ratio of the normal vector histogram is the ratio of the number of intervals with the highest count in the normal vector direction histogram to the number of intervals with the second highest count, i.e., peak bin count / second peak bin count, where bin represents the angular region, which can be used to determine the surface continuity: when the QR code is occluded, the normal vector direction of the occluded region becomes discrete, which will lead to a decrease in the peak ratio.

[0085] Therefore, normal vector features are used to reflect the three-dimensional orientation distribution and continuity of an object's surface. By statistically analyzing the peak ratio of the normal vector direction histogram, it is possible to quantitatively determine whether there are occlusions or other factors on the QR code surface that cause the normal vector direction to be discrete, thus providing a three-dimensional spatial feature basis for scanning mode decision-making.

[0086] For example, the surface normal vectors of a complete QR code have a consistent direction, and the histogram shows a single peak with a peak-to-peak ratio of 3.0. If it is obscured by the edge of a cardboard box, the peak-to-peak ratio of the normal vector histogram may drop to 1.5. If a preset threshold H is set... th If the value is 2.0, the occlusion compensation mode will be triggered in subsequent processes.

[0087] Based on the above technical solution, a multi-dimensional quantitative representation of the QR code surface state is achieved by fusing texture, curvature, and normal vector features. Texture features are used to capture the grayscale distribution pattern of the visual image, curvature features are used to quantify the abrupt changes in infrared and depth caused by reflection, and normal vector features are used to statistically analyze the orientation distribution of the three-dimensional surface. These features provide multi-dimensional input for subsequent credibility calculations of Dempster-Shafer evidence theory, enabling the system to accurately determine the decodeability of QR codes in complex scenarios such as strong reflection, occlusion, and low light, laying the foundation for intelligent decision-making in scanning modes.

[0088] In one possible implementation of the embodiments of this application, combined with Figure 2 ,like Figure 3 As shown, the above S3 can be implemented through the following S301, S302 and S303, which are explained in detail below: S301. Define the recognition framework.

[0089] The recognition framework Θ is defined as {"decodeable", "requires enhancement", "undecodeable"}, used to classify the QR code scanning results. "Decoding available" (m1): Indicates that the QR code can be directly decoded under the current feature conditions; "Needs Enhancement" (m2): This indicates that the scanning mode needs to be adjusted (such as reflection suppression, low light enhancement, etc.) to improve the decoding success rate; "Undecodeable" (m3): This indicates that the current conditions cannot be decoded and manual intervention is required.

[0090] In some implementations, categories can be expanded according to scenario requirements, such as adding sub-states like "blurred and needs to be focused" and "partially occluded" to improve decision-making granularity.

[0091] It should be noted that the category of the recognition framework must be strictly mapped to the subsequent scanning mode. For example, "needs enhancement" corresponds to specific modes such as reflection suppression and occlusion compensation, to ensure that the confidence calculation results can directly guide the operation.

[0092] For example, in an industrial production line scenario, if there is no obvious interference on the surface of the QR code, the recognition frame determines it as m1, and the system performs direct scanning; if there is slight reflection, it is determined as m2, and the reflection suppression mode is triggered.

[0093] S302: Calculate the basic probability distribution of each feature.

[0094] Calculate the basic probability m for texture feature Ft, curvature feature Fk, and normal vector feature Fn respectively. i (A)

[0095] The basic probability calculation formula is as follows: ; In the formula: F i The currently extracted feature values ​​(i=t,k,n correspond to texture, curvature, and normal vector, respectively). This indicates that feature i is a preset baseline feature value for category A, which is determined through training with historical samples (for example, the baseline texture feature value for category m1 corresponds to the LBP feature distribution that is manually judged as a clear QR code). σ is a scaling parameter used to control the smoothness of the probability distribution. It is usually set to 0.1-1.0 and can be optimized through cross-validation.

[0096] In some implementations, the baseline eigenvalue Supports dynamic updates: When the system scans a new scene, if the features of consecutively successfully decoded samples differ from the original baseline by more than 10%, it will automatically update. To improve adaptability.

[0097] It should be noted that the basic probability m i (A) represents the support of a single feature for a class, with a value range of [0,1], and the sum of the probabilities of the same feature for all classes is 1. For example, if the texture feature has a high matching degree with the baseline feature m1, then m t (m1) approaches 1, m t (m2) and m t (m3) approaches 0.

[0098] For example, suppose the texture feature F of a QR code t Baseline characteristics of category m1 The Euclidean distance is 0.2, σ = 0.5, then: ;like and If all are greater than 0.5, then m t (m1)≈0.67 indicates that the confidence level of the texture features supporting “decoding” is 67%.

[0099] S303. Based on the Dempster combination rule, the basic probability is fused to obtain the confidence level Bel.

[0100] The specific steps include: 1. Evidence Combination: Calculate the basic probability m of fusion using the combination rules of Dempster-Shafer evidence theory. total (A): ; The numerator is the sum of the products of all sub-features whose basic probabilities intersect at point A, and the denominator is the normalization factor (excluding cases of conflicting evidence).

[0101] 2. Credibility Calculation: Calculate the final credibility based on the fusion results. Where λ is the weighting coefficient (0≤λ≤1), used to adjust the impact of the "needs enhancement" category on credibility. In industrial scenarios, λ is often taken as 0.3-0.5 to balance efficiency and accuracy.

[0102] In some implementations, when the degree of conflict of evidence (in the denominator) If the value exceeds 0.7, it is considered that the multiple features are significantly contradictory, and data needs to be collected again or the feature extraction module needs to be checked to avoid misjudgment.

[0103] It should be noted that Dempster's combination rules can effectively handle the complementarity and conflict of multi-source data: for example, texture features support m1(m t =0.6), but the curvature feature supports m2 (m k When m = 0.7), the fused m total (m1) will combine the support of both to avoid a single feature dominating the decision.

[0104] For example, suppose: Texture features: m t (m1)=0.6, m t (m2)=0.3, m t (m3)=0.1; Curvature characteristics: m k (m1)=0.2, m k (m2)=0.7, m k (m3)=0.1; Normal vector characteristics: m n (m1)=0.5, m n (m2)=0.3, m n (m3)=0.2.

[0105] During fusion, the sum of products with intersection m1 is calculated: m t (m1)×m k (m1)×m n (m1) = 0.06; Calculate the sum of products whose intersection is m²: m t (m2)×m k (m2)×m n (m2)=0.063; Calculate the sum of products whose intersection is m³: m t (m3)×m k (m3)×m n (m3)=0.002; Apart from the three types of intersections of the same category mentioned above, all other combinations are conflicting cases, such as m1×m2×m3, m1×m2×m2, etc., which require calculating the product of each item and summing them to obtain the value of the denominator. Then normalize according to the formula, assuming that the normalized result is m. total (m1)≈0.42, m total (m2)≈0.51; Note: The example here is a hypothetical result. In actual applications, all conflicting cases need to be multiplied and summed before calculation.

[0106] If λ=0.4, then the confidence level Bel=0.42+0.4×0.51=0.622, indicating that after integrating multi-dimensional features such as texture, curvature, and normal vector, the overall support for the "decodeable" state of the QR code is about 62%. Its essence is a quantitative index obtained by fusing the confidence level of multi-source features through Dempster-Shafer evidence theory. Therefore, the higher the value of Bel∈[0,1], the closer the QR code is to the "decodeable" state.

[0107] Based on the above technical solution, a three-layer recognition framework, a basic probability calculation model, and Dempster combination rules were constructed to achieve confidence fusion of texture, curvature, and normal vector features. This mechanism not only solves the problem of misjudgment of single-modal features in complex scenes (such as when strong light causes texture features to fail, curvature features can compensate), but also dynamically adjusts the influence of the "needs enhancement" category through λ weights, so that the confidence calculation results can accurately reflect the decodeability status of the QR code.

[0108] In one possible implementation of the embodiments of this application, combined with Figure 2 ,like Figure 3 As shown, the decision-making conditions for S4 above specifically include: (1) When the confidence level Bel>T1, it is determined that the QR code can be directly decoded and the default scanning parameters are executed.

[0109] Where T1 is the high confidence threshold, typically set to 0.7-0.9, with the specific value depending on the scenario's requirements for scanning reliability. For example, 0.85 might be used in industrial scenarios, while 0.7 might be used in consumer scenarios. The execution logic of the direct scanning mode includes: 1. The vision sensor acquires images using standard exposure parameters, eliminating the need for additional lighting; 2. Trigger the QR code decoding algorithm to directly parse the QR code information in the visual image.

[0110] It should be noted that the direct scanning mode does not completely disregard environmental interference; rather, it determines that interference does not affect decoding based on a comprehensive assessment of multi-dimensional features. For example, if slight reflection exists, but the confidence level of the fusion of curvature features and other features is still >T1, then there is no need to activate the reflection suppression mode.

[0111] For example, in a warehouse management scenario, a QR code is affixed to a flat surface of a cardboard box. The collected visual image has clear texture, continuous depth map, and consistent normal vector distribution. The calculated value is Bel=0.88>T1=0.8. The system directly triggers decoding and completes the scan within 10ms.

[0112] (2) When the confidence level is between T2 and T1, the system executes the corresponding enhancement mode based on the multidimensional feature matching results.

[0113] T2 is a low confidence threshold (typically 0.3-0.5), at which point feature analysis is needed to locate the type of interference, specifically including: 1. Reflection suppression mode.

[0114] Triggering condition: Curvature feature K max >K th ; Among them, K th The curvature threshold (e.g., 0.5) is set when the maximum curvature feature value of the target region exceeds K. th When this occurs, it indicates a sudden change in infrared reflection intensity caused by reflection.

[0115] Perform the following operation: 1. Turn off the vision sensor light to avoid visible light reflection; 2. Increase the infrared emission power to a preset value (e.g., 80% of the maximum power) and use infrared light to penetrate the reflective layer; 3. The scanning angle is deflected by a preset angle along the direction of the greatest rate of change of curvature feature, for example, with an angle increment of 2°, gradually increasing from 5°. The confidence level is recalculated each time it is deflected, until the confidence level is greater than T1.

[0116] 2. Occlusion compensation mode.

[0117] Triggering condition: Peak ratio of normal vector histogram <H th ; Among them, H th The peak-to-peak ratio threshold (e.g., 2.0) is calculated as: Peak-to-peak bin count / Number of peak bins. When this value is reached... <H th When the normal vector distribution is discrete, it indicates that occlusion exists.

[0118] Operation: Control the radar sensor to scan back and forth within a preset pitch angle, such as [10°, +10°], and reconstruct the outline of the occluded area using multi-angle point cloud data.

[0119] 3. Low-light enhancement mode Triggering condition: Texture feature F t_max <F t_th ; Among them, F t_this the texture feature threshold (e.g., 0.6). When the maximum texture feature value F t_max is lower than F t_th , it indicates that the image contrast is insufficient (caused by low light).

[0120] Perform the operation: activate the LED fill light of the vision sensor and start the infrared auxiliary lighting.

[0121] In some implementations, if multiple enhancement mode trigger conditions are satisfied simultaneously, the system executes according to the preset priority: specular reflection suppression > occlusion compensation > low light enhancement.

[0122] It should be noted that the positioning of the target area can be achieved through vision edge detection (such as the Canny operator) or deep learning recognition (such as the YOLO model) to ensure that the feature analysis is only targeted at the area where the QR code is located.

[0123] Exemplarily, when scanning the QR code on the glass, if the curvature feature K max = 0.7 > K th = 0.5, then execute the specular reflection suppression mode: turn off the vision light, increase the infrared power to 80%, deflect the scanning angle 10° to the right, and then recalculate the multi-dimensional features. If the curvature feature K max drops to 0.4 and the confidence level increases to 0.75, then direct scanning can be triggered.

[0124] When the confidence level Bel < T2, the system determines that the QR code cannot be automatically decoded and starts the manual intervention mechanism.

[0125] Among them, the execution logic of the manual verification mode includes: Output the visual superposition results of the visual image, depth map, and infrared reflection intensity map, and mark the suspicious areas (such as specular reflection points, occlusion boundaries) in red; Generate a feature analysis report, including data such as texture feature vectors, curvature heat maps, and normal vector histograms; Prompt the operator to intervene through audible and visual alarms or the host computer interface. For example, display "The QR code cannot be decoded. Please check for surface damage or occlusion".

[0126] In some implementations, the manual verification mode can integrate the physical button function of the barcode scanner. After the operator confirms the QR code status, they can force a secondary scan (such as rescan after adjusting the angle) or manually enter the QR code information.

[0127] It should be noted that the manual verification mode is a fault tolerance mechanism of the system, applicable to extreme scenarios such as serious QR code damage, multi-angle occlusion, or sensor failure, where the fusion confidence level of multi-dimensional features is no longer sufficient to support automatic decision-making.

[0128] Exemplarily, if the QR code is covered by oil stains and the visual image texture is blurred (F t_max = 0.3 < F t_th = 0.6), the peak ratio of the normal vector histogram = 1.2 < H th = 2.0, it is calculated that Bel = 0.25 < T2 = 0.3. The system pops up a prompt box and highlights the oil stain area. After the operator wipes it, the credibility is recalculated.

[0129] It should be noted that the preset thresholds of T1 and T2 can be optimized through offline training (such as collecting a large number of industrial scenario samples) or online adaptation (such as adjusting according to historical decoding success rates) to ensure that the decision boundary meets the actual application requirements.

[0130] Based on the above technical solutions, through the three-layer decision-making mechanism driven by credibility, the multi-modal perception data and the scanning strategy can be directly associated through a mathematical model, avoiding reliance on empirical rules, and enabling the system to have the ability of autonomous learning and environmental adaptation. This mechanism can accurately judge the decodable state of the QR code, realize the adaptive optimization of the scanning strategy, and can improve the recognition reliability and environmental adaptability in complex scenarios such as strong reflection, low light, and partial occlusion.<关于上述公式中的部分数据是去除量纲取其数值计算,公式是由采集的大量数据经过软件模拟得到最接近真实情况的一个公式;公式中的预设参数和预设阈值由本领域的技术人员根据实际情况设定或者通过大量数据模拟获得。

[0131] Some of the data in the above formula are calculated by removing the dimension and taking their numerical values. The formula is obtained by software simulation of a large amount of collected data to get a formula closest to the actual situation; the preset parameters and preset thresholds in the formula are set by those skilled in the art according to the actual situation or obtained through simulation of a large amount of data. [[ID=关于上述主要从设备实现的角度对本申请实施例的方案进行了介绍。可以理解的是,各个设备,例如,基于多模态感知的二维码智能扫描装置为了实现上述功能,其包含了执行各个功能相应的硬件结构和软件模块中的至少一个。本领域技术人员应该很容易意识到,结合本文中所公开的实施例描述的各示例的单元及算法步骤,本申请能够以硬件或硬件和计算机软件的结合形式来实现。某个功能究竟以硬件还是计算机软件驱动硬件的方式来执行,取决于技术方案的特定应用和设计约束条件。专业技术人员可以对每个特定的应用来使用不同方法来实现所描述的功能,但是这种实现不应认为超出本申请的范围。

[0132] The above mainly introduces the solutions of the embodiments of the present application from the perspective of device implementation. It can be understood that for each device, for example, the QR code intelligent scanning device based on multi-modal perception includes at least one of the corresponding hardware structures and software modules for executing each function in order to implement the above functions. Those skilled in the art should easily realize that, combining the units and algorithm steps of the examples described in the embodiments disclosed in this article, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0133] It should be noted that there seems to be some incorrect or inconsistent formatting in the original text, especially in the part where the Chinese text is mixed in the English translation section. I have tried my best to translate according to the requirements while keeping the original structure. If possible, it is recommended to check and correct the original text for a more accurate translation.This application embodiment can divide the multimodal perception-based QR code intelligent scanning device into functional units according to the above method example. For example, each function can be divided into a separate functional unit, or two or more functions can be integrated into one processing unit. The integrated unit can be implemented in hardware or as a software functional unit. It should be noted that the unit division in this application embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.

[0134] When using integrated units, Figure 5 A possible structural schematic diagram of the QR code intelligent scanning device based on multimodal perception (referred to as QR code intelligent scanning device 50 based on multimodal perception) involved in the above embodiments is shown. The QR code intelligent scanning device 50 based on multimodal perception includes a processing unit 501 and a communication unit 502, and may also include a storage unit 503. Figure 5 The structural diagram shown can be used to illustrate the structure of the QR code intelligent scanning device based on multimodal perception involved in the above embodiments.

[0135] when Figure 5 The schematic diagram shown illustrates the structure of the QR code intelligent scanning device based on multimodal perception involved in the above embodiments. The processing unit 501 is used to control and manage the operation of the QR code intelligent scanning device based on multimodal perception. Specifically, it includes extracting features from the collected multidimensional data to obtain multidimensional features including texture features, curvature features, and normal vector features. The multidimensional features are input into the DempsterShafer evidence theory fusion unit to output confidence. Based on the confidence, a decision is made on the scanning mode of the object to be scanned. The scanning mode includes direct scanning mode, reflection suppression mode, occlusion compensation mode, and low light enhancement mode.

[0136] The communication unit 502 is used for communication between the QR code intelligent scanning device based on multimodal perception and other devices. Specifically, it receives visual images, depth images, infrared reflection intensity maps and radar point cloud data of the object to be scanned to obtain multidimensional data and sends the decision results of the scanning mode.

[0137] The storage unit 503 is used to store the program code and data of the QR code intelligent scanning device based on multimodal perception, including storing the algorithm program for realizing multimodal data acquisition, feature extraction, evidence theory fusion and scanning mode decision, as well as preset benchmark feature values, threshold parameters and other data.

[0138] In one possible implementation, the processing unit 501 can also be used to schedule the workflow of the feature extraction module and the fusion decision module, such as controlling the timing of feature extraction and coordinating the order of multi-dimensional feature fusion.

[0139] In one possible implementation, the communication unit 502 is also used to establish a data transmission link with external sensors (such as vision sensors, depth sensors, infrared sensors, and radar sensors) to acquire multidimensional perception data in real time. The processing unit 501 is also used to preprocess the received multidimensional data, such as denoising and format conversion, to ensure that the data meets the input requirements of the feature extraction module.

[0140] The processing unit 501 can be a processor or a controller, and the communication unit 502 can be a communication interface, transceiver, transceiver circuit, transceiver device, etc. The term "communication interface" is a general term and may include one or more interfaces. The storage unit 503 can be a memory. When the multimodal sensing-based QR code intelligent scanning device 50 is a chip, the processing unit 501 can be a processor or a controller, and the communication unit 502 can be an input interface and / or an output interface, pins, or circuits, etc. The storage unit 503 can be a storage unit within the chip (e.g., a register, cache, etc.) or a storage unit located outside the chip (e.g., read-only memory (ROM), random access memory (RAM, etc.).

[0141] The communication unit can also be called a transceiver unit. The antenna and control circuit with transceiver functions in the multimodal sensing-based QR code intelligent scanning device 50 can be considered as the communication unit 502 of the multimodal sensing-based QR code intelligent scanning device 50, and the processor with processing functions can be considered as the processing unit 501 of the multimodal sensing-based QR code intelligent scanning device 50. Optionally, the device in the communication unit 502 used to implement the receiving function can be considered as the communication unit, which is used to execute the receiving steps in the embodiments of this application. The communication unit can be a receiver, a receiver circuit, etc. The device in the communication unit 502 used to implement the transmitting function can be considered as the transmitting unit, which is used to execute the transmitting steps in the embodiments of this application. The transmitting unit can be a transmitter, a transmitter, a transmitting circuit, etc.

[0142] Figure 5If the integrated units in the process are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, in essence, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. Storage media for storing computer software products include various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.

[0143] Figure 5 The units in the process can also be called modules; for example, a processing unit can be called a processing module.

[0144] This application also provides a hardware structure diagram of a QR code intelligent scanning method and device based on multimodal perception (referred to as multimodal perception-based QR code intelligent scanning device 60), see [link to diagram]. Figure 6 The QR code smart scanning device 60 based on multimodal perception includes a processor 601, and optionally, a memory 602 connected to the processor 601.

[0145] In the first possible implementation, see Figure 6 The QR code intelligent scanning device 60 based on multimodal perception also includes a transceiver 603. The processor 601, memory 602, and transceiver 603 are connected via a bus. The transceiver 603 is used to communicate with other devices or communication networks. Optionally, the transceiver 603 may include a transmitter and a receiver. The device in the transceiver 603 that implements the receiving function can be considered as a receiver, which is used to perform the receiving steps in the embodiments of this application. The device in the transceiver 603 that implements the transmitting function can be considered as a transmitter, which is used to perform the transmitting steps in the embodiments of this application.

[0146] Based on the first possible implementation method Figure 6 The structural diagram shown can be used to illustrate the structure of the QR code intelligent scanning device based on multimodal perception involved in the above embodiments.

[0147] in, Figure 6 This can also be illustrated by the system chip in a QR code smart scanning device based on multimodal perception. In this case, the actions performed by the aforementioned QR code smart scanning device based on multimodal perception can be implemented by this system chip. The specific actions performed can be found above and will not be repeated here.

[0148] In implementation, each step of the method provided in this embodiment can be completed by integrated logic circuits in the processor or by instructions in software form. The steps of the method disclosed in the embodiments of this application can be directly manifested as being executed by a hardware processor, or being executed by a combination of hardware and software modules in the processor.

[0149] The processor in this application may include, but is not limited to, at least one of the following: a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a microcontroller unit (MCU), or an artificial intelligence processor, etc., which are various computing devices that run software. Each computing device may include one or more cores for executing software instructions to perform calculations or processing. The processor may be a separate semiconductor chip or integrated with other circuits into a single semiconductor chip. For example, it may be integrated with other circuits (such as encoding / decoding circuits, hardware acceleration circuits, or various bus and interface circuits) to form a SoC (System-on-a-Chip), or it may be integrated as a built-in processor within an ASIC. The ASIC with the integrated processor may be packaged separately or together with other circuits. In addition to the cores for executing software instructions to perform calculations or processing, the processor may further include necessary hardware accelerators, such as field-programmable gate arrays (FPGAs), PLDs (programmable logic devices), or logic circuits that implement dedicated logic operations.

[0150] The memory in the embodiments of this application may include at least one of the following types: read-only memory (ROM) or other types of static storage devices capable of storing static information and instructions; random access memory (RAM) or other types of dynamic storage devices capable of storing information and instructions; or electrically erasable programmable read-only memory (EEPROM). In some scenarios, the memory may also be a compact disc read-only memory (CDROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer, but is not limited thereto.

[0151] This application also provides a computer-readable storage medium including instructions that, when run on a computer, cause the computer to perform any of the methods described above.

[0152] This application also provides a computer program product containing instructions that, when run on a computer, cause the computer to perform any of the methods described above.

[0153] This application also provides a chip including a processor and an interface circuit. The interface circuit is coupled to the processor. The processor is used to run computer programs or instructions to implement the above-described method. The interface circuit is used to communicate with other modules outside the chip.

[0154] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software programs, implementation can be, in whole or in part, in the form of a computer program product. This computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device containing one or more servers, data centers, etc., that can be integrated with the medium. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state disks (SSDs)).

[0155] Although this application has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings, disclosure, and appended claims, will understand and implement other variations of the disclosed embodiments in carrying out the claimed application. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple instances. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce good results.

[0156] Although this application has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made thereto without departing from the spirit and scope of this application. Accordingly, this specification and drawings are merely illustrative descriptions of the application as defined by the appended claims, and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from the spirit and scope of this application. Thus, if such modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is also intended to include such modifications and variations.

Claims

1. A QR code intelligent scanning method based on multimodal perception, characterized in that, include: The system collects visual images, depth images, infrared reflectance intensity maps, and radar point cloud data of the object to be scanned to obtain multidimensional data. Multidimensional features are obtained by extracting features from multidimensional data. These multidimensional features include texture features extracted from visual images, curvature features calculated from infrared reflectance intensity maps and depth images, and normal vector features calculated from radar point cloud data. Multidimensional features are input into a Dempster-Shafer evidence theory fusion unit, which outputs confidence scores. The fusion unit is used to perform fusion calculations on multidimensional features based on Dempster-Shafer evidence theory to obtain confidence scores. The decision on the scanning mode of the object to be scanned is based on confidence level; the scanning mode includes direct scanning mode, reflection suppression mode, occlusion compensation mode and low light enhancement mode.

2. The QR code intelligent scanning method based on multimodal perception according to claim 1, characterized in that, The methods for obtaining the curvature features include: Calculation of infrared reflectance intensity diagram I ir gradient magnitude and the eigenvalues ​​λ1 and λ2 of the Hessian matrix; Based on gradient magnitude The curvature feature K is calculated using the Hessian matrix eigenvalues ​​λ1 and λ2 and the depth map D(x,y) according to the curvature feature calculation formula.

3. The QR code intelligent scanning method based on multimodal perception according to claim 2, characterized in that, The curvature characteristic calculation formula satisfies: Among them, S pix represents the physical size of a pixel, and μ represents the depth weighting coefficient.

4. The QR code intelligent scanning method based on multimodal perception according to claim 1, characterized in that, The methods for obtaining the texture features include: The visual image is divided into blocks to obtain multiple image blocks; Extract the local binary pattern (LBP) feature vector for each image patch; Calculate the statistical features within each image patch, including the contrast, entropy, and energy of the image patch's gray-level co-occurrence matrix; By combining LBP feature vectors with statistical features, texture features of visual images can be obtained.

5. The QR code intelligent scanning method based on multimodal perception according to claim 1, characterized in that, The methods for obtaining the normal vector features include: Build a KD Tree index for each point in the radar point cloud data; Based on point p i The set of neighboring points N(p) within the spatial location search radius r i ); Calculate the neighborhood point set N(p) i The covariance matrix of the covariance matrix is ​​used, and the eigenvector corresponding to the smallest eigenvalue in the covariance matrix is ​​taken as the point p. i The normal vector; By statistically analyzing the histogram of the normal vector direction for each point, the characteristics of the normal vector are obtained.

6. The intelligent QR code scanning method based on multimodal perception according to claim 1, characterized in that, The input of multidimensional features into the Dempster-Shafer evidence theory fusion processor includes: Construct a recognition framework Θ={m1="decodeable",m2="requires enhancement",m3="undecodeable"}; where the recognition framework is used to determine the category to which the QR code scanning result of the object to be scanned belongs; Calculate texture features F respectively t Curvature feature F k , Normal vector feature F n The basic probability m i (A); where the basic probability is represented in feature F i∈{t,k,n} Under the given conditions, determine the confidence level of the QR code scanning result of the object to be scanned to belong to category A∈Θ in the recognition frame; Based on the evidence combination rules of Dempster-Shafer evidence theory, multi-dimensional features are fused to obtain the fused basic probability m. total (A); where, the fusion basic probability represents the comprehensive confidence level of determining whether the QR code scanning result of the object to be scanned belongs to category A after integrating the basic probabilities of multidimensional features; The credibility level Bel=m is calculated according to the formula. total (m1)+λ×m total (m2); where λ represents the weighting coefficient.

7. The QR code intelligent scanning method based on multimodal perception according to claim 6, characterized in that, The basic probability m i The formula for (A) is: ;in, This indicates that feature i is the preset baseline feature value of category A, and σ represents the scaling parameter; The fusion basic probability m total The formula for (A) is: ;in, Represents the empty set. The symbol represents the intersection.

8. The intelligent QR code scanning method based on multimodal perception according to claim 1, characterized in that, The decision-making process for the scanning mode of the object to be scanned based on credibility includes: If the confidence level is greater than T1, then execute the direct scan mode; If T2 ≤ confidence level ≤ T1, then pattern matching is performed based on multidimensional features, including: When the curvature feature of the target region exists K max >K th At that time, the reflection suppression mode is activated; K max K represents the maximum curvature eigenvalue. th Indicates the preset curvature threshold; When the peak ratio of the normal vector histogram of the target region is <H th At that time, the occlusion compensation mode is executed; the normal vector histogram is obtained based on the statistical analysis of the normal vectors of all points, H th This indicates the preset peak-to-peak ratio threshold; When the texture features of the target region exist F t_max <F t_th At that time, low-light enhancement mode is executed; F t_max F represents the maximum texture feature value. t_th Indicates the preset texture feature threshold; If the credibility is less than T2, then the manual verification mode is triggered; The target area refers to the QR code scanning area, which is obtained by edge detection or image recognition of the image through a visual sensor; T1 represents a preset first threshold, and T2 represents a preset second threshold.

9. The intelligent QR code scanning method based on multimodal perception according to claim 1, characterized in that, The automatic execution operation of the reflection suppression mode includes: turning off the light of the vision sensor, increasing the infrared emission power to a preset value, and deflecting the scanning angle by a preset angle along the direction of the curvature feature change rate. The automatic execution operation of the occlusion compensation mode includes: controlling the radar sensor to perform point cloud scanning within a preset pitch angle range; The automatic execution of the low-light enhancement mode includes: activating the light of the vision sensor and activating infrared auxiliary illumination.

10. A QR code intelligent scanning system based on multimodal perception, characterized in that, include: The data acquisition module is used to acquire visual images, depth images, infrared reflectance intensity maps, and radar point cloud data of the object to be scanned, thereby obtaining multi-dimensional data. The feature extraction module is used to extract features from multidimensional data to obtain multidimensional features, which include texture features extracted from visual images, curvature features calculated from infrared reflectance intensity maps and depth images, and normal vector features calculated from radar point cloud data; the fusion unit is used to perform fusion calculations on the multidimensional features based on Dempster-Shafer evidence theory. The fusion decision module is used to input multi-dimensional features into the Dempster Shafer evidence theory fusion engine and output the confidence level. The decision on the scanning mode of the object to be scanned is based on confidence level; the scanning mode includes direct scanning mode, reflection suppression mode, occlusion compensation mode and low light enhancement mode.

Citation Information

Patent Citations

  • Code recognition method and mobile terminal

    CN107992780A

  • Robot positioning method and system based on binocular vision and laser scanning fusion

    CN118050734A

  • Image processing method and device, computing equipment and storage medium

    CN118762081A

  • Intelligent code reading equipment based on AI

    CN119830932A

  • Intelligent spectrum tuning code scanning auxiliary lighting method, system and equipment and medium

    CN120264524A

Cited By

  • Positioning method and device based on two-dimensional code, equipment, medium and product

    CN122133688A