An image barcode multi-code recognition method, device, system and storage medium
By introducing dynamic region growth and parallel decoding methods into barcode recognition technology, the recognition problems in complex backgrounds, intensive arrangements and mixed-type scenarios are solved, and efficient and accurate multi-code recognition effect is achieved.
Patent Information
- Application Number
- CN202510380376.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-28
AI Technical Summary
Existing barcode recognition technology does not perform well in complex backgrounds, intensive arrangements or mixed-type scenarios, especially in dense barcode separation, mixed-type compatibility, real-time and complex background interference.
The multi-code recognition method of image barcode based on dynamic region growth and parallel decoding is adopted. Through technical means such as multi-modal image preprocessing, multi-task detection network, dynamic region growth algorithm, geometric correction and GPU parallel decoding pipeline, precise separation of dense barcodes, adaptive recognition and real-time processing of mixed types are achieved.
It significantly improves the accuracy and efficiency of multi-code recognition in complex scenarios, and solves the problems of difficult separation of dense barcodes, poor compatibility of mixed types, insufficient real-time performance and complex background interference.
Smart Images

Figure CN119903863B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of barcode recognition, and specifically, to an image barcode multi-code recognition method, device, system, and storage medium. Background Art
[0002] As a key means of automatic information collection, barcode recognition technology has been widely applied in fields such as logistics, retail, and industrial inspection. Existing methods are mainly divided into two categories: traditional image processing-based methods and deep learning-based methods. Traditional methods (such as open-source libraries like ZBar and ZXing) locate and decode barcodes through steps such as image preprocessing, edge detection, and projection analysis. However, they rely on manually designed features and have insufficient robustness in complex backgrounds, densely arranged, or uneven illumination scenarios. In particular, the separation and recognition effects of overlapping barcodes and mixed types (coexistence of one-dimensional and two-dimensional barcodes) are poor. In recent years, deep learning-based object detection models (such as YOLO and Faster R-CNN) have significantly improved barcode localization accuracy, but there are still the following problems:
[0003] Difficult separation of dense barcodes: Existing detection models are prone to missing detections or false detections for closely arranged or partially occluded barcodes, and it is difficult to achieve precise segmentation relying on post-processing algorithms (such as non-maximum suppression);
[0004] Poor compatibility for mixed types: Most solutions need to pre-specify the barcode type and cannot adaptively recognize mixed scenarios of one-dimensional barcodes, two-dimensional barcodes, and color barcodes;
[0005] Insufficient real-time performance: The serial decoding process results in high processing delays for multiple barcodes and is difficult to meet the real-time requirements of scenarios such as industrial sorting and mobile payment;
[0006] Interference from complex backgrounds: Barcodes are easily confused with similar texture backgrounds (such as packaging patterns and metal reflections), leading to an increased decoding failure rate. Summary of the Invention
[0007] The present invention proposes an image barcode multi-code recognition method, device, system, and storage medium based on dynamic region growth and parallel decoding, which solves the problems of difficult separation of dense barcodes, poor compatibility for mixed types, and insufficient real-time performance in related technologies.
[0008] The technical solution of the present invention is as follows:
[0009] An image barcode multi-code recognition method, characterized by including the following steps:
[0010] S100. Perform multi-modal image preprocessing on the input image to be recognized. The multi-modal image preprocessing includes performing dynamic illumination compensation on the image to be recognized and performing multi-scale feature fusion on the image to be recognized, and an enhanced image is generated after the multi-modal image preprocessing;
[0011] S200. Input the enhanced image into a multi-task detection network, which simultaneously outputs the bounding box, segmentation mask, and type probability of the barcode.
[0012] S300. Using the center point of the bounding box as a seed point, control the growth direction by combining the black-and-white alternating frequency of the barcode texture, and separate densely arranged or partially overlapping barcode regions through a dynamic region growing algorithm.
[0013] S400. Geometrically correct the separated barcode regions and select the corresponding decoding strategy according to the barcode type judgment result.
[0014] S500. Based on the GPU parallel decoding pipeline, simultaneously perform binarization, modular parsing, and error correction decoding on multiple barcode regions.
[0015] S600. Verify the logical and spatial consistency of the decoding results through a multi-modal verification mechanism and output structured recognition information.
[0016] Further, in step S100,
[0017] The dynamic light compensation includes decomposing the original image into a reflection component and a light component based on the Retinex theory, adaptively enhancing the light component, generating multi-angle polarization simulation images for the reflective area, and selecting the image with the highest local contrast as the preprocessing result.
[0018] The multi-scale feature fusion includes constructing a Gaussian pyramid for the image after dynamic light compensation, generating multi-scale downsampled images, extracting the high-frequency edge features and low-frequency texture features of each scale image, and weighted fusing the multi-scale features through a channel attention mechanism to generate an enhanced output image.
[0019] Further, in step S200, the structure of the multi-task detection network is as follows:
[0020] The backbone network uses an improved MobileNetV3 with an efficient channel attention module embedded in MobileNetV3.
[0021] The output head includes a bounding box regression branch, a class probability branch, and a segmentation mask branch. Among them, the bounding box regression branch outputs the bounding box coordinates and confidence of the barcode, and uses an improved intersection over union loss function to optimize the positioning accuracy. The class probability branch outputs the barcode type probability and uses a focal loss function to alleviate the class imbalance problem. The segmentation mask branch outputs the pixel-level mask of the barcode and uses a similarity-based Dice loss function to improve the segmentation fitting degree.
[0022] The total loss function of the multi-task detection network is the weighted sum of the above three losses, and its calculation formula is
[0023] ,
[0024] wherein, , , are preset weight coefficients, is the improved intersection over union loss, is the class imbalance optimization loss, is the similarity loss of the segmentation mask.
[0025] Furthermore, in step S300, the specific steps of the dynamic region growing algorithm are as follows:
[0026] S310. Estimate the barcode direction angle through Hough transform , and extend the region boundary along the direction;
[0027] S320. Calculate the black and white pixel alternation frequency within the local window f . If f < f th then stop growing;
[0028] S330. For the overlapping regions, preferentially separate the barcodes containing the positioning marks.
[0029] Furthermore, in step S400, the processing method of the geometric correction for the one-dimensional code or two-dimensional code is as follows:
[0030] For the one-dimensional code, detect the tilt angle through Hough transform and rotate it to the horizontal direction;
[0031] For the two-dimensional code, extract the corner points of the positioning marks to calculate the homography matrix H and perform perspective transformation.
[0032] Furthermore, in step S500, the running steps of the GPU parallel decoding pipeline are as follows:
[0033] S510. Allocate the barcode region to multiple CUDA thread blocks and perform binarization and module parsing in parallel;
[0034] S520. Dynamically allocate computing resources according to the barcode complexity score . In the calculation formula of the barcode complexity score, is the blur weight coefficient, is the error correction level weight coefficient, is the blur score, is the error correction level quantization value;
[0035] S530. Preprocess the blurred area by preferentially invoking the super-resolution reconstruction model.
[0036] Further, in step S600, the multimodal verification mechanism includes:
[0037] Logical verification to verify whether the check bits and encoding formats conform to the standards;
[0038] Spatial consistency verification to perform majority voting fusion on the multi-view recognition results.
[0039] An image barcode multi-code recognition device includes:
[0040] A multimodal image preprocessing module that performs multimodal image preprocessing on the input image to be recognized. The multimodal image preprocessing includes performing dynamic light compensation on the image to be recognized and performing multi-scale feature fusion on the image to be recognized. After the multimodal image preprocessing, an enhanced image is generated;
[0041] A multi-task detection module that inputs the enhanced image into a multi-task detection network. The multi-task detection network simultaneously outputs the bounding box, segmentation mask, and type probability of the barcode;
[0042] A dynamic segmentation module that uses the center point of the bounding box as a seed point, combines the black-and-white alternating frequency of the barcode texture to control the growth direction, and separates densely arranged or partially overlapping barcode areas through a dynamic region growing algorithm;
[0043] A barcode correction module that performs geometric correction on the separated barcode areas and selects a corresponding decoding strategy according to the barcode type judgment result;
[0044] A GPU parallel decoding module that simultaneously performs binarization, modular parsing, and error correction decoding on multiple barcode areas based on a GPU parallel decoding pipeline;
[0045] A structured output module that performs logical and spatial consistency verification on the decoding results through a multimodal verification mechanism and outputs structured recognition information. The structured recognition information includes JSON data of barcode type, coordinates, and association relationships
[0046] The structure of the multi-task detection network is:
[0047] The backbone network uses an improved MobileNetV3 with an efficient channel attention module embedded in MobileNetV3;
[0048] The output head includes a bounding box regression branch, a class probability branch, and a segmentation mask branch. Among them, the bounding box regression branch outputs the bounding box coordinates and confidence of the barcode, and uses an improved intersection over union loss function to optimize the positioning accuracy. The class probability branch outputs the barcode type probability and uses a focal loss function to alleviate the class imbalance problem. The segmentation mask branch outputs the pixel-level mask of the barcode and uses a similarity-based Dice loss function to improve the segmentation fitting degree;
[0049] The total loss function of the multi-task detection network is the weighted sum of the above three losses, and its calculation formula is
[0050] ,
[0051] Among them, , , are preset weight coefficients, is the improved intersection over union loss, is the class imbalance optimization loss, is the similarity loss of the segmentation mask.
[0052] An image barcode multi-code recognition system, the system includes:
[0053] One or more memories for storing instructions; and
[0054] One or more processors for calling and running the instructions from the memory to execute the image barcode multi-code recognition method as described above.
[0055] A computer-readable storage medium, the computer-readable storage medium includes:
[0056] A program, when the program is run by a processor, the image barcode multi-code recognition method as described above is executed.
[0057] The working principle and beneficial effects of the present invention are:
[0058] The present invention realizes the precise separation of dense barcodes through a multi-task detection network and a dynamic region growing algorithm, combines a GPU-accelerated parallel decoding pipeline to improve the processing efficiency, and designs a multi-modal image preprocessing and structured output mechanism, effectively solving the core pain points in the prior art such as difficult separation of dense arrangements, insufficient support for mixed types, poor real-time performance, and complex background interference, and significantly improving the multi-code recognition accuracy and efficiency in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] The present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0060] Figure 1It is the flowchart of the method in Embodiment 1;
[0061] Figure 2 It is the flowchart of the multi-modal image preprocessing in Embodiment 1;
[0062] Figure 3 It is the block diagram of the multi-task detection network structure in Embodiment 1;
[0063] Figure 4 It is the flowchart of the dynamic region growing in Embodiment 1;
[0064] Figure 5 It is the flowchart of the multi-modal verification mechanism in Embodiment 1;
[0065] Figure 6 It is the block diagram of the device structure in Embodiment 2. Specific implementation manners
[0066] Next, in combination with the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the scope of protection of the present invention.
[0067] Embodiment 1
[0068] As Figures 1 to 5 shown, this embodiment proposes an image barcode multi-code recognition method, including the following steps:
[0069] S100. Perform multi-modal image preprocessing on the input image to be recognized. The multi-modal image preprocessing includes performing dynamic illumination compensation on the image to be recognized and performing multi-scale feature fusion on the image to be recognized. After the multi-modal image preprocessing, an enhanced image is generated.
[0070] In step S100, the illumination distribution of the image is adjusted by dynamic illumination compensation to reduce the interference of low-illumination or reflective areas; multi-scale feature fusion extracts edge and texture features of different resolutions to enhance the saliency of the barcode area. The multi-modal image preprocessing can improve the image quality, ensure the robustness of the subsequent detection and segmentation modules under complex illumination (such as dim environment, metal reflection), and avoid feature loss or false detection caused by uneven illumination.
[0071] S200. Input the enhanced image into a multi-task detection network, and the multi-task detection network simultaneously outputs the bounding box, segmentation mask, and type probability of the barcode.
[0072] Step S200 uses a multi-task detection network to output the bounding box (localization) for generating barcodes synchronously, the segmentation mask (pixel-level region), and the type probability (classification), achieving end-to-end detection and preliminary classification. It avoids the error accumulation in the step-by-step processing of traditional methods (such as detecting first and then segmenting), improves the processing efficiency, and reduces the computational redundancy through multi-task feature sharing.
[0073] S300: Using the center point of the bounding box as the seed point, combined with the black-and-white alternating frequency of the barcode texture to control the growth direction, the densely arranged or partially overlapping barcode regions are separated by the dynamic region growing algorithm.
[0074] Step S300 uses the dynamic region growing algorithm to start from the center of the bounding box, combined with the black-and-white alternating frequency to control the growth direction, and adaptively expands the boundary of the barcode region. The dynamic region growing algorithm can accurately separate densely packed or partially overlapping barcodes (such as in the scenario of stacked packages), avoiding the problem of region adhesion caused by traditional connected component analysis or fixed threshold segmentation.
[0075] S400: Geometric correction is performed on the separated barcode regions, and the corresponding decoding strategy is selected according to the barcode type judgment result.
[0076] Step S400 performs geometric correction (such as rotation, perspective transformation) on the inclined and distorted barcodes, and calls the corresponding decoding strategy according to the type (one-dimensional code / two-dimensional code). It can solve the problem of decoding failure of deformed barcodes (such as inclined two-dimensional codes), improve the compatibility of cross-type barcodes, and avoid decoding errors caused by misjudgment of types.
[0077] S500: Based on the GPU parallel decoding pipeline, binarization, modular parsing, and error correction decoding are performed on multiple barcode regions simultaneously.
[0078] Step S500 uses the parallel computing power of the GPU to perform binarization, modular parsing, and error correction decoding on multiple barcode regions simultaneously. It significantly reduces the processing delay in the multi-barcode scenario, meets the real-time requirements of the multi-code scenario, and the decoding efficiency is improved several times compared with the CPU serial decoding.
[0079] S600: The decoding results are verified for logical and spatial consistency through a multi-modal verification mechanism, and structured recognition information is output.
[0080] Step S600 filters out the incorrect decoding results through logical verification (such as check digit verification) and spatial consistency verification (such as multi-view result fusion), and organizes them into structured data (such as JSON). It ensures the reliability of the output results (such as avoiding misreading and garbled codes), and the structured data can be directly connected to the business system, improving the practicality of the technical solution.
[0081] In this embodiment, in step S100,
[0082] Dynamic illumination compensation includes decomposing the original image into a reflection component and an illumination component based on the Retinex theory, adaptively enhancing the illumination component, generating multi-angle polarization simulation images for the specular reflection area, and selecting the image with the highest local contrast as the preprocessing result;
[0083] Specifically, dynamic illumination compensation decomposes the original image into a reflection component and an illumination component , satisfying: , estimating the illumination component through Gaussian filtering , calculating the reflection component as , and performing contrast stretching on R . Dynamic illumination compensation can solve the problem of insufficient contrast caused by low illumination and specular reflection;
[0084] Multi-scale feature fusion includes constructing a Gaussian pyramid for the image after dynamic illumination compensation, generating multi-scale downsampled images, extracting high-frequency edge features and low-frequency texture features of each scale image, and performing weighted fusion on the multi-scale features through a channel attention mechanism to generate an enhanced output image.
[0085] Specifically, generating a Gaussian{ G 0 , G 1 , G 2} for the image after illumination compensation, where G k is the downsampled image of the k th layer.
[0086] High-frequency edge feature extraction: Extracting the edges of each layer through the Laplacian operator ;
[0087] Low-frequency texture feature extraction: Extracting the textures of each layer through mean filtering ;
[0088] Channel attention weighted fusion of multi-scale features , where is dynamically calculated by the attention network, represents channel concatenation. Multi-scale fusion can enhance the cross-scale feature expression of barcode edges and textures, and improve the robustness of subsequent detection.
[0089] In this embodiment, in step S200, the structure of the multi-task detection network is:
[0090] The backbone network uses an improved MobileNetV3, and an Efficient Channel Attention (ECA-Net) module is embedded in the penultimate layer of MobileNetV3. Through the channel weights: , where is the Sigmoid function, W is the learnable parameter, F c is the input feature map. The ECA module can enhance the feature response of small targets (such as small barcodes).
[0091] The output head contains a bounding box regression branch, a class probability branch, and a segmentation mask branch. Among them, the bounding box regression branch outputs the bounding box coordinates of the barcode ( x , y , w , h ) and the confidence p . The improved Intersection over Union loss function (CIoU loss) is used to optimize the localization accuracy. The class probability branch outputs the barcode type probability P type . The focal loss function is used to alleviate the class imbalance problem. The segmentation mask branch outputs the pixel-level mask of the barcode M . The similarity-based Dice loss function is used to improve the segmentation fitting degree.
[0092] The total loss function of the multi-task detection network is the weighted sum of the above three losses, and its calculation formula is
[0093] ,
[0094] where , , are the preset weight coefficients, is the improved Intersection over Union loss, is the class imbalance optimization loss, is the similarity loss of the segmentation mask.
[0095] The calculation formula of the Intersection over Union loss is: , where IoU is the Intersection over Union of the predicted box and the ground truth box, is the Euclidean distance between the centers of the predicted box and the ground truth box, c is the diagonal length of the smallest enclosing box, v is the aspect ratio penalty term, is the weight coefficient, .
[0096] The calculation formula of the class imbalance optimization loss is: , where is the predicted probability of the model for the true class; is the class weight, which solves class imbalance (such as the uneven number of one-dimensional codes and two-dimensional codes); is the focusing parameter (usually ≥0), which reduces the weight of easily classifiable samples and enables the model to focus on difficult samples.
[0097] The similarity loss calculation formula for the segmentation mask is: , where is the segmentation mask predicted by the model, is the true segmentation mask.
[0098] The multi-task loss collaboratively optimizes the localization, classification, and segmentation accuracy, reducing error accumulation.
[0099] In this embodiment, in step S300, the specific steps of the dynamic region growing algorithm are as follows:
[0100] S310. Estimate the barcode orientation angle through Hough transform , and extend the region boundary along the direction;
[0101] Specifically, taking the center of the BBox output by the detection network ( x c , y c ) as the seed point, perform Hough transform on the neighborhood of the seed point to detect the barcode orientation angle , The linear parametric equation is: , where x is the column coordinate (horizontal position) of a certain pixel point in the image, y is the row coordinate (vertical position) of a certain pixel point in the image.
[0102] S320. Calculate the black-and-white pixel alternation frequency within the local window f , if f < f th then stop growing.
[0103] Calculate the black-and-white pixel alternation frequency within the local window f The calculation formula is: . Where N is the total number of pixels within the local window, that is, the number of all pixels within the window; i is the position index of the current pixel in the window pixel sequence, and the value range is ; I is the pixel value sequence arranged in spatial order within the local window (usually the binarized value, 0 represents black, and 1 represents white); is the iThe binary value of a pixel is used to calculate the difference between adjacent pixels.
[0104] S330. For the overlapping area, preferentially separate the barcode containing the positioning mark.
[0105] Specifically, preferentially separate the barcode area containing the positioning mark (such as the corner points of a QR code).
[0106] The dynamic region growing algorithm controls the growth by combining direction and texture frequency, avoiding over-segmentation or under-segmentation of dense barcodes in traditional methods; in the overlapping scenario, preferentially separate through the positioning mark to improve the recognition priority of key barcodes.
[0107] In this embodiment, in step S400, the processing method for geometric correction of one-dimensional barcodes or two-dimensional barcodes is as follows:
[0108] For one-dimensional barcodes, detect the tilt angle through the Hough transform and rotate it to the horizontal direction.
[0109] Specifically, detect the tilt angle through the Hough transform , and the rotation matrix aligns with the horizontal direction. By rotating and correcting one-dimensional barcodes, the sampling deviation caused by tilt can be solved.
[0110] For two-dimensional barcodes, extract the corner points of the positioning mark to calculate the homography matrix H and perform perspective transformation.
[0111] Specifically, extract the corner points of the positioning mark , calculate the homography matrix H, and minimize the reprojection error. The calculation formula of the homography matrix H is: , where is the homography matrix that minimizes the objective function through optimization and solution H , that is, find the optimal H such that the sum of the squares of the reprojection errors of all corresponding points is minimized; is the standard coordinate of the i th corner point in the corrected target coordinate system (such as the theoretical corner point coordinates of the QR code positioning mark); is the actual coordinate of the i th corner point detected in the original image (the tilted or distorted position caused by perspective deformation); H is the homography matrix (a 3×3 projection transformation matrix) used to map the corner points in the original image to the standard coordinate system , realizing geometric correction. By performing perspective transformation on the two-dimensional barcode to restore the planar deformation, the decoding success rate can be improved.
[0112] In this embodiment, in step S500, the running steps of the GPU parallel decoding pipeline are as follows:
[0113] S510. Allocate the barcode area to multiple CUDA thread blocks and perform binarization and module parsing in parallel.
[0114] Specifically, for binarization, local thresholding is used for one-dimensional barcodes, and the OTSU global threshold is used for two-dimensional barcodes. Module parsing means parsing the data matrix of the two-dimensional barcode according to the mask pattern. Parallel decoding reduces latency and meets the requirements of industrial real-time performance.
[0115] S520. According to the barcode complexity score Dynamically allocate computing resources, and perform super-resolution reconstruction on the blurred area with priority.
[0116] In the barcode complexity score calculation formula:
[0117] is the blurriness weight coefficient, and the value range is 0 ≤ α ≤ 1, which is used to adjust the influence weight of the blurriness degree of the barcode area on the overall complexity score.
[0118] is the error correction level weight coefficient, and the value range is 0 ≤ β ≤ 1, which is used to adjust the influence weight of the barcode error correction level on the overall complexity score, and satisfies α + β = 1.
[0119] is the blurriness score, and the value range is [0, 1], which is obtained by normalizing the gradient variance of the barcode area image. The larger the value, the blurrier the image.
[0120] is the error correction level quantization value, and the value range is [0, 1], which is mapped to a numerical value according to the error correction ability level of the barcode (for example, the error correction levels L / M / Q / H of the QR code correspond to 0.2 / 0.5 / 0.8 / 1.0 respectively).
[0121] S530. Preprocess the blurred area by preferentially calling the super-resolution reconstruction model.
[0122] Step S520 and step S530 achieve dynamic scheduling to optimize resource utilization and improve the decoding priority of high-difficulty barcodes.
[0123] In this embodiment, in step S600, the multi-modal verification mechanism includes:
[0124] Logical verification, verifying whether the check digit and the encoding format meet the standards.
[0125] Specifically, logical verification includes one-dimensional barcode check digit verification (such as the modulo 10 verification of EAN-13) and two-dimensional barcode format character and version information verification. Logical verification can filter out data with format errors.
[0126] Spatial consistency check, and perform majority voting fusion on the multi-view recognition results.
[0127] Fusing multi-view information through spatial consistency can improve the reliability of the results.
[0128] Embodiment 2
[0129] As Figure 6 shown, this embodiment proposes an image barcode multi-code recognition device, including:
[0130] A multi-modal image preprocessing module that performs multi-modal image preprocessing on the input image to be recognized. The multi-modal image preprocessing includes performing dynamic illumination compensation on the image to be recognized and performing multi-scale feature fusion on the image to be recognized. After the multi-modal image preprocessing, an enhanced image is generated.
[0131] The multi-modal image preprocessing module can solve the problem of image quality degradation caused by complex illumination conditions (such as dim environments and metal reflections); enhance the cross-scale expression of barcode edges and textures, and provide more robust input data for subsequent modules.
[0132] A multi-task detection module that inputs the enhanced image into a multi-task detection network, and the multi-task detection network simultaneously outputs the bounding box, segmentation mask, and type probability of the barcode.
[0133] The multi-task detection module can achieve end-to-end multi-task processing to reduce computational redundancy and avoid error accumulation in traditional step-by-step detection; and enhance the feature response of small targets (such as small barcodes) through an efficient channel attention module (ECA-Net) to improve the detection accuracy in dense scenes.
[0134] A dynamic segmentation module that uses the center point of the bounding box as a seed point, combines the black-and-white alternating frequency of the barcode texture to control the growth direction, and separates densely arranged or partially overlapping barcode regions through a dynamic region growing algorithm.
[0135] The dynamic segmentation module can accurately separate closely arranged barcodes (such as in a package stacking scenario) that are difficult to handle by traditional methods; control the growth direction through texture frequency to avoid over-segmentation or under-segmentation problems.
[0136] A barcode correction module that performs geometric correction on the separated barcode regions and selects a corresponding decoding strategy based on the barcode type judgment result.
[0137] The barcode correction module can solve the problem of decoding failure for deformed barcodes (such as tilted and distorted); and support the adaptive processing of mixed-type barcodes (such as the coexistence of one-dimensional and two-dimensional barcodes).
[0138] The GPU parallel decoding module simultaneously performs binarization, modular parsing, and error correction decoding on multiple barcode regions based on the GPU parallel decoding pipeline.
[0139] The GPU parallel decoding module can significantly reduce the processing latency of multi-barcode scenarios (with several times higher efficiency); and dynamic resource allocation optimizes the computational load, improving the decoding success rate of high-difficulty barcodes.
[0140] The structured output module verifies the logical and spatial consistency of the decoding results through a multi-modal verification mechanism and outputs structured recognition information, which includes JSON data on barcode types, coordinates, and association relationships.
[0141] The structured output module can ensure the accuracy and reliability of the output data (such as avoiding misreading and garbled codes); and its structured data can be directly connected to the business system, improving the practicality of the technical solution.
[0142] The structure of the multi-task detection network is as follows:
[0143] The backbone network uses an improved MobileNetV3 with an efficient channel attention module embedded in MobileNetV3;
[0144] The output head includes a bounding box regression branch, a class probability branch, and a segmentation mask branch. Among them, the bounding box regression branch outputs the bounding box coordinates and confidence of the barcode, and uses an improved intersection over union loss function to optimize the positioning accuracy. The class probability branch outputs the barcode type probability and uses a focal loss function to alleviate the class imbalance problem. The segmentation mask branch outputs the pixel-level mask of the barcode and uses a similarity-based Dice loss function to improve the segmentation fitting degree;
[0145] The total loss function of the multi-task detection network is the weighted sum of the above three losses, and its calculation formula is
[0146] ,
[0147] where, , , are preset weight coefficients, is the improved intersection over union loss, is the class imbalance optimization loss, is the similarity loss of the segmentation mask.
[0148] Example 3
[0149] This example proposes an image barcode multi-code recognition system, which includes:
[0150] One or more memories for storing instructions; and
[0151] One or more processors, configured to call and run instructions from a memory and execute the image barcode multi-code recognition method as described above.
[0152] Embodiment 4
[0153] This embodiment provides a computer-readable storage medium, which includes:
[0154] A program, which when run by a processor, executes the image barcode multi-code recognition method as described above.
[0155] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for multi-code recognition of image barcodes, characterized in that: The following steps are involved: S100, performing multimodal image preprocessing on the input image to be recognized, wherein the multimodal image preprocessing includes performing dynamic illumination compensation on the image to be recognized and performing multi-scale feature fusion on the image to be recognized, and generating an enhanced image after the multimodal image preprocessing; S200, inputting the enhanced image into a multi-task detection network, wherein the multi-task detection network simultaneously outputs a bounding box, a segmentation mask, and a type probability of the barcode; S300, using the center point of the bounding box as a seed point, combining the black and white alternating frequency of the barcode texture to control the growth direction, and separating the densely arranged or partially overlapping barcode regions through a dynamic region growing algorithm; S400, performing geometric correction on the separated barcode area, and selecting a corresponding decoding strategy according to the barcode type judgment result; S500, based on GPU parallel decoding pipeline, performs binarization, modular analysis and error correction decoding on multiple barcode areas simultaneously; S600, verifying the logic and spatial consistency of the decoding result through a multi-modal verification mechanism, and outputting structured identification information; In step S200, the structure of the multi-task detection network is: The backbone network uses an improved MobileNetV3, in which an efficient channel attention module is embedded; The output head includes a bounding box regression branch, a category probability branch, and a segmentation mask branch. The bounding box regression branch outputs the bounding box coordinates and confidence of the barcode, and uses an improved intersection-over-union loss function to optimize the positioning accuracy. The category probability branch outputs the barcode type probability, and uses a focal loss function to alleviate the category imbalance problem. The segmentation mask branch outputs the pixel-level mask of the barcode, and uses a similarity-based Dice loss function to improve the segmentation fit. The total loss function of the multi-task detection network is the weighted sum of the above three losses, and its calculation formula is: , in, , , is the preset weight coefficient, For the improved intersection-over-union loss, Optimize loss for class imbalance, is the similarity loss of the segmentation mask.
2. The method for multi-code recognition of image barcodes according to claim 1, characterized in that: In step S100, The dynamic illumination compensation includes decomposing the original image into a reflection component and an illumination component based on the Retinex theory, adaptively enhancing the illumination component, generating a multi-angle polarization simulation image for the reflective area, and selecting the image with the highest local contrast as the preprocessing result; The multi-scale feature fusion includes constructing a Gaussian pyramid for the image after dynamic illumination compensation, generating a multi-scale down-sampled image, extracting high-frequency edge features and low-frequency texture features of each scale image, performing weighted fusion of the multi-scale features through a channel attention mechanism, and generating an enhanced output image.
3. The method for multi-code recognition of image barcodes according to claim 1, characterized in that: In step S300, the specific steps of the dynamic region growing algorithm are as follows: S310, estimating the barcode direction angle by Hough transform ,along Direction expansion area boundaries; S320, calculating the alternation frequency of black and white pixels in the local window f ,like f < f th Then growth stops; S330. For the overlapping area, the barcode containing the positioning mark is separated preferentially.
4. The method for multi-code recognition of image barcodes according to claim 1, characterized in that: In step S400, the geometric correction processing method for the one-dimensional code or the two-dimensional code is: For one-dimensional codes, the tilt angle is detected by Hough transform and rotated to the horizontal direction; For the QR code, the corner points of the positioning mark are extracted to calculate the homography matrix H and perform perspective transformation.
5. The method for multi-code recognition of image barcodes according to claim 1, characterized in that: In step S500, the operation steps of the GPU parallel decoding pipeline are: S510, allocating the barcode area to multiple CUDA thread blocks, and performing binarization and module parsing in parallel; S520, score based on barcode complexity Dynamically allocate computing resources, in the barcode complexity score calculation formula, is the fuzziness weight coefficient, is the error correction level weight coefficient, Score the fuzziness, is the quantitative value of the error correction level; S530, preferentially calling a super-resolution reconstruction model to pre-process the blurred area.
6. The method for multi-code recognition of image barcodes according to claim 1, characterized in that: In step S600, the multimodal verification mechanism includes: Logical check to verify whether the check bit and encoding format meet the standards; Spatial consistency check, and majority voting fusion of multi-view recognition results.
7. An image barcode multi-code recognition device, characterized in that: include: A multimodal image preprocessing module performs multimodal image preprocessing on the input image to be identified, wherein the multimodal image preprocessing includes dynamic illumination compensation for the image to be identified and multi-scale feature fusion for the image to be identified, and an enhanced image is generated after the multimodal image preprocessing; A multi-task detection module, which inputs the enhanced image into a multi-task detection network, which simultaneously outputs the barcode's bounding box, segmentation mask, and type probability; A dynamic segmentation module, based on the center point of the bounding box as a seed point, controls the growth direction in combination with the black and white alternating frequency of the barcode texture, and separates densely arranged or partially overlapping barcode regions through a dynamic region growing algorithm; The barcode correction module performs geometric correction on the separated barcode area and selects the corresponding decoding strategy according to the barcode type judgment result; GPU parallel decoding module, based on GPU parallel decoding pipeline, performs binarization, modular analysis and error correction decoding on multiple barcode areas at the same time; The structured output module verifies the logic and spatial consistency of the decoding results through a multi-modal verification mechanism and outputs structured identification information, which includes JSON data of barcode type, coordinates and association relationships; The structure of the multi-task detection network is: The backbone network uses an improved MobileNetV3, in which an efficient channel attention module is embedded; The output head includes a bounding box regression branch, a category probability branch, and a segmentation mask branch. The bounding box regression branch outputs the bounding box coordinates and confidence of the barcode, and uses an improved intersection-over-union loss function to optimize the positioning accuracy. The category probability branch outputs the barcode type probability, and uses a focal loss function to alleviate the category imbalance problem. The segmentation mask branch outputs the pixel-level mask of the barcode, and uses a similarity-based Dice loss function to improve the segmentation fit. The total loss function of the multi-task detection network is the weighted sum of the above three losses, and its calculation formula is: , in, , , is the preset weight coefficient, For the improved intersection-over-union loss, Optimize loss for class imbalance, is the similarity loss of the segmentation mask.
8. An image barcode multi-code recognition system, characterized in that: The system comprises: one or more memories for storing instructions; and One or more processors, configured to call and execute the instructions from the memory to perform the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer readable storage medium comprises: The program, when the program is executed by a processor, the method according to any one of claims 1 to 6 is executed.
Citation Information
Patent Citations
Method for rapidly positioning and recognizing one-dimensional barcode of outer commodity package
CN104463066A
Multi-task detection method for surface abnormal region pixel-level segmentation
CN112669274A