Industrial whole disk multi-target intelligent code scanning method and system based on deep learning

Through the YOLOv8-SSD model of deep learning multi-exposure simulation, saliency modulation layer and residual feature fusion path, combined with the improved Soft-NMS algorithm and image enhancement processing, the problem of low recognition accuracy of QR codes on reflective or transparent film surfaces is solved, and automatic recognition of multi-target QR codes under complex lighting conditions is achieved.

CN120633694APending Publication Date: 2025-09-12无锡宇宁科技集团股份有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510701986.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

In scenarios such as industrial manufacturing, warehousing and logistics, and electronic packaging, the accuracy and reliability of QR code recognition are limited by reflective interference. Especially under reflective or transparent film conditions, traditional methods have difficulty in effectively identifying multi-target QR codes.

Method used

A deep learning-based industrial whole-plate multi-target intelligent scanning method is adopted. Through multi-exposure simulation processing, the YOLOv8-SSD lightweight neural network model with saliency modulation layer and residual feature fusion path, the improved Soft-NMS algorithm and local image enhancement processing, reflective interference is suppressed and the QR code recognition accuracy is improved.

Benefits of technology

It significantly improves the recognition accuracy and edge positioning robustness of QR codes on reflective or transparent film surfaces, and is suitable for automatic recognition of multi-target QR codes under complex lighting conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633694A_ABST
    Figure CN120633694A_ABST
Patent Text Reader

Abstract

The invention discloses an industrial whole disk multi-target intelligent code scanning method and system based on deep learning, and the method comprises the steps: obtaining an original image containing a plurality of two-dimensional codes; performing multi-exposure simulation processing on the original image to generate a plurality of enhanced images with different brightness levels; inputting the enhanced image into a YOLOv8-SSD lightweight neural network model fused with a significance modulation layer and a residual feature fusion path, and extracting a two-dimensional code candidate box; the saliency modulation layer adjusts the channel attention weight according to the local brightness gradient and the standard deviation in the image so as to suppress reflection interference; applying a Soft-NMS algorithm with a reflective area adaptive weight strategy to a candidate box output by the model, and reserving an effective detection result of a shielded or reflective area; performing contrast enhancement and texture detail restoration processing on the two-dimensional code image in each detection frame; and the processed two-dimensional code image is sent to a decoding module. The technical scheme of the invention aims to improve the accuracy and reliability of two-dimensional code recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of two-dimensional code recognition technology, and in particular to a method and system for industrial whole-plate multi-target intelligent code scanning based on deep learning. Background Art

[0002] In modern industrial manufacturing, warehousing and logistics, and electronic packaging, QR codes, as the primary carrier of information identification, are widely used for automated tracking and identification of materials such as pallets, devices, and components. However, as QR codes become increasingly complex in real-world environments, their recognition accuracy faces numerous challenges, particularly in the presence of surface reflections, which severely limits the stability and reliability of conventional recognition methods. Summary of the Invention

[0003] The purpose of the embodiments of the present application is to propose an industrial whole-plate multi-target intelligent scanning method and system based on deep learning to solve the technical problems of poor accuracy and reliability of QR code recognition.

[0004] To solve the above technical problems, the present application provides a method for intelligently scanning industrial whole-plate multi-target codes based on deep learning, which adopts the following technical solutions:

[0005] A method for intelligently scanning industrial whole-plate multi-target codes based on deep learning, comprising the following steps:

[0006] Acquire an original image containing a plurality of two-dimensional codes, wherein the two-dimensional codes are attached to an object having a reflective or transparent film surface;

[0007] Perform multi-exposure simulation processing on the original image to generate multiple enhanced images with different brightness levels;

[0008] Inputting the enhanced image into a YOLOv8-SSD lightweight neural network model that integrates a saliency modulation layer and a residual feature fusion path to extract QR code candidate frames;

[0009] The saliency modulation layer adjusts the channel attention weight according to the local brightness gradient and standard deviation in the image to suppress reflection interference;

[0010] Apply the Soft-NMS algorithm with an adaptive weight strategy for reflective areas to the candidate boxes output by the model to retain valid detection results in occluded or reflective areas;

[0011] Perform contrast enhancement and texture detail restoration on the QR code image within each detection frame;

[0012] The processed QR code image is fed into the decoding module to extract the QR code content.

[0013] In a possible implementation, the step of performing multi-exposure simulation processing on the original image to generate a plurality of enhanced images with different brightness levels includes:

[0014] The brightness enhancement method based on Retinex theory performs multiple exposure simulation processing on the original image;

[0015] Construct a multi-branch convolutional network to extract features of images with different exposures;

[0016] The weighted fusion mechanism is used to fuse the feature maps of each branch to form an enhanced input feature map.

[0017] In one possible implementation, in the step of extracting a QR code candidate frame from a YOLOv8-SSD lightweight neural network model that integrates a saliency modulation layer and a residual feature fusion path into the enhanced image input, the saliency modulation layer includes:

[0018] Estimate the reflective interference area in the image by using the local brightness gradient entropy and pixel standard deviation;

[0019] A reflective mask map is constructed and inhibitory attention adjustment is performed on the feature map channel weights based on the mask map.

[0020] In a possible implementation, after the step of extracting a QR code candidate frame by inputting the enhanced image into a YOLOv8-SSD lightweight neural network model that is fused with a saliency modulation layer and a residual feature fusion path, the method further includes:

[0021] The unenhanced original image features are adjusted by 1×1 convolution and then channel-wise fused with the deep feature maps in the backbone network to retain the edge information of the QR code.

[0022] In one possible implementation, the Soft-NMS algorithm includes:

[0023] Calculate the reflection probability of the candidate frame area based on its average brightness and edge direction distribution;

[0024] The confidence attenuation function between candidate frames is dynamically adjusted according to the reflection probability to achieve fault tolerance for weakly reflective QR codes.

[0025] In a possible implementation, the step of performing contrast enhancement and texture detail restoration processing on the QR code image within each detection frame includes:

[0026] Perform local histogram equalization to improve the contrast of the QR code area;

[0027] Edge-preserving filtering is used to restore the texture details of the QR code and enhance decoding stability.

[0028] In a possible implementation, in the step of performing multiple exposure simulation processing on the original image using the brightness enhancement method based on Retinex theory, the multiple exposure simulation processing simulates the visual performance of highlights and dark areas in the RGB image domain by adjusting logarithmic brightness transformation parameters.

[0029] In order to solve the above technical problems, the embodiment of the present application also provides an industrial whole-plate multi-target intelligent code scanning system based on deep learning, which adopts the following technical solutions:

[0030] An industrial whole-plate multi-target intelligent scanning system based on deep learning, including:

[0031] An acquisition module is used to acquire an original image containing a plurality of two-dimensional codes, wherein the two-dimensional codes are attached to an object having a reflective or transparent film surface;

[0032] A first processing module is used to perform multi-exposure simulation processing on the original image to generate multiple enhanced images with different brightness levels;

[0033] A first extraction module is configured to input the enhanced image into a YOLOv8-SSD lightweight neural network model fused with a saliency modulation layer and a residual feature fusion path to extract a QR code candidate frame;

[0034] An adjustment module, wherein the saliency modulation layer adjusts the channel attention weight according to the local brightness gradient and standard deviation in the image to suppress reflection interference;

[0035] The application module is used to apply the Soft-NMS algorithm with the adaptive weight strategy of the reflective area to the candidate boxes output by the model, retaining the valid detection results of the occluded or reflective areas;

[0036] The second processing module is used to perform contrast enhancement and texture detail restoration processing on the QR code image in each detection frame;

[0037] The second extraction module is used to send the processed two-dimensional code image to the decoding module to extract the content of the two-dimensional code.

[0038] In order to solve the above technical problems, the embodiment of the present application further provides a computer device, which adopts the following technical solution:

[0039] A computer device includes a memory and a processor, wherein the memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, it implements the steps of the industrial whole-plate multi-target intelligent scanning method based on deep learning as described above.

[0040] In order to solve the above technical problems, the embodiment of the present application further provides a computer-readable storage medium, which adopts the following technical solution:

[0041] A computer-readable storage medium having computer-readable instructions stored thereon, which, when executed by a processor, implement the steps of the industrial whole-disk multi-target intelligent code scanning method based on deep learning as described above.

[0042] Compared with the prior art, the embodiments of the present application have the following beneficial effects:

[0043] The deep learning-based industrial whole-plate multi-target intelligent scanning method disclosed in this application forms an integrated recognition link by introducing multi-exposure image simulation, saliency modulation attention mechanism, residual feature fusion path, improved Soft-NMS algorithm, and image local enhancement and decoding process in the reflective interference scenario. It can significantly improve the overall recognition accuracy and edge positioning robustness of QR codes under typical interference conditions such as transparent film, reflective material or local highlight on the QR code surface. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the solutions in this application, a brief introduction will be given below to the drawings required for use in the description of the embodiments of this application. Obviously, the drawings described below are some embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0045] Figure 1 This is a flowchart of an embodiment of the industrial whole-plate multi-target intelligent code scanning method based on deep learning according to the present application;

[0046] Figure 2 This is a structural diagram of an embodiment of an industrial whole-plate multi-target intelligent code scanning system based on deep learning according to the present application;

[0047] Figure 3 It is a structural diagram of an embodiment of a computer device according to the present application. DETAILED DESCRIPTION

[0048] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0049] In the following typical scenarios, the accuracy of QR code recognition is significantly reduced: the QR code is attached to a transparent plastic film, an antistatic bag or a PET label, and its surface produces mirror reflections or highlight artifacts; the QR code is located on a polished or coated metal surface, and strong reflected light causes partial overexposure of the image; the QR code is artificially laminated or covered with tape, resulting in complex reflection phenomena such as fine stripe interference and bubble artifacts; in the low-light environment of the factory, the lighting compensation is insufficient, and the contrast of the local reflective area is higher than that of the QR code body, causing confusion in edge recognition.

[0050] In these situations, traditional QR code recognition methods (including handheld scanners, fixed industrial camera systems, and YOLO-based target detection models) face the following problems: The edge information of the QR code is suppressed or disrupted by reflective areas, resulting in positioning failure or offset; reflective areas generate bright saturated noise, which the model interprets as background or other non-QR code targets, resulting in missed or false detections; the recognition model cannot adapt to local exposure differences, and image enhancement processing has poor versatility; standard post-processing algorithms (such as NMS) incorrectly suppress the recognition results of obscured QR codes, affecting recall rates. Although some studies have attempted to alleviate the problem of reflective interference by increasing front-end light sources, using HDR cameras, and using image enhancement algorithms, they are subject to practical limitations such as high cost, complex deployment, and the inability of edge devices to support them. This makes it difficult to promote in industrial scenarios that require full-disk multi-target recognition, rapid deployment, and edge-side operation.

[0051] In view of this, reference Figure 1 , shows a flow chart of an embodiment of the industrial whole disk multi-target intelligent code scanning method based on deep learning according to the present application. The industrial whole disk multi-target intelligent code scanning method based on deep learning includes the following steps:

[0052] Step S101 : acquiring an original image containing a plurality of two-dimensional codes, wherein the two-dimensional codes are attached to an object having a reflective or transparent film surface.

[0053] In this embodiment, the electronic device on which the deep learning-based industrial whole-disk multi-target intelligent code scanning method is running can send or receive data via a wired connection or a wireless connection. It should be noted that the above-mentioned wireless connection method may include but is not limited to 3G / 4G / 5G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wideband) connection, and other wireless connection methods currently known or to be developed in the future.

[0054] In this embodiment, a raw image containing multiple QR code targets is captured as the basic data source for subsequent recognition algorithm processing. QR codes are typically affixed to the surfaces of industrial materials, such as device packaging boxes, electronic pallets, pharmaceutical packages, and logistics containers. However, in actual applications, these surfaces are often covered with transparent plastic film, anti-static bags, PET protective film, or other smooth coated materials. These materials have significant specular reflection properties, which can cause highlights, exposure saturation, and edge bleeding in the captured image. This significantly affects the integrity and contrast of the QR code in the image, especially in low-light or side-light conditions. Therefore, it is necessary to capture images containing these interfering features at the source.

[0055] Step S102 : performing multi-exposure simulation processing on the original image to generate a plurality of enhanced images with different brightness levels.

[0056] In this embodiment, software enhancement is mainly used to generate image versions with multiple simulated exposure levels based on a single-frame image, avoiding hardware requirements such as high dynamic range (HDR) for the acquisition equipment. The specific processing method is: based on the Retinex theory or its variant method, separation and mapping are performed between the illumination component and the reflection component of the image, and by adjusting the brightness transformation parameters in the logarithmic domain of the image, multiple versions of the image are generated. These images emphasize the contrast details of dark areas, bright areas or neutral areas respectively, thereby simulating the real multi-exposure sampling process; finally, a set of image sequences is formed, covering the brightness distribution from high exposure to low exposure, to make up for the missing QR code information in the original image due to local strong reflection or overexposure, and to avoid ignoring the QR code edge or background boundary due to insufficient exposure. No need for multi-frame shooting, high computational efficiency, suitable for deployment in edge-side industrial recognition terminals.

[0057] Step S103: Input the enhanced image into a YOLOv8-SSD lightweight neural network model that integrates a saliency modulation layer and a residual feature fusion path to extract a QR code candidate frame.

[0058] In this embodiment, a lightweight neural network with a fusion structure is used to uniformly process multiple enhanced images and output candidate frames for the positions of QR codes in the images. The target detection framework combining the YOLOv8 and SSD structures is used as the basic backbone, which can efficiently handle multi-target recognition tasks. To enhance the robustness of the model under reflective interference, two key structures are embedded in the model: a saliency modulation layer and a residual feature fusion path. The saliency modulation layer is used to identify areas where reflective disturbances may exist based on the local brightness gradient, standard deviation, and edge change characteristics of the image, and generates a saliency weight map based on this. By fusing it with the intermediate features extracted by the backbone network, attention adjustment is applied in the channel dimension, which reduces the channel response of the model in reflective areas and enhances the channel response in QR code areas. The residual feature fusion path preserves key structural information such as QR code edges, positioning frames, and corners by channel-wise splicing the shallow edge feature map of the network with the deep semantic feature map, effectively compensating for edge blur or information loss that may occur during the image enhancement process. After end-to-end training and optimization, the entire network model can maintain sensitivity and stability to the QR code structure under complex lighting backgrounds and output multiple preliminary located QR code candidate frames.

[0059] In step S104, the saliency modulation layer adjusts the channel attention weight according to the local brightness gradient and standard deviation in the image to suppress reflection interference.

[0060] In this embodiment, a window sliding calculation is first performed on the image in the network's intermediate layer or input stage to extract the local brightness gradient change value and brightness standard deviation. Regions of high reflectivity within the image, such as blocks with local highlights, high contrast, and low texture, are identified by setting a threshold. These regions are then constructed into a saliency weight map and mapped to the neural network's feature channels. By weighting the response strength of different channels, the activation of the QR code-related texture channel is enhanced while suppressing the influence of the reflective noise-related channel. This attention mechanism differs from conventional CBAM and SE channel attention in that its activation factor is derived from the image's physical reflective structure, resulting in stronger semantic consistency and correspondence with reflective features. This significantly improves the model's ability to discriminate QR code regions in complex backgrounds.

[0061] Step S105: Apply the Soft-NMS algorithm with the reflective area adaptive weight strategy to the candidate boxes output by the model to retain the valid detection results of the occluded or reflective areas.

[0062] In this embodiment, the QR code candidate frames in the model output results are further optimized, especially when the candidate frames overlap, are partially occluded, or light spot reflection areas frequently appear, and an improved Soft-NMS (non-maximum suppression) algorithm is used to dynamically retain and adjust the confidence of the detection frames. Unlike the standard NMS algorithm that only suppresses confidence based on IoU, the present invention introduces an adaptive adjustment strategy for reflective areas based on Soft-NMS, that is, a reflection probability estimation is performed on the image area of ​​each candidate frame. The estimation is based on factors such as the regional pixel brightness mean, brightness variance, and edge direction gradient distribution, and comprehensively evaluates whether there is strong reflection interference in the area; if an overlapping candidate frame is located in a high-reflection area, the confidence decay rate of the frame is slowed down by setting a nonlinear confidence decay curve, thereby retaining the QR code target in the edge area that may be interfered with and occluded, improving the overall recall rate and reducing the situation where valid targets are mistakenly killed due to reflection.

[0063] Step S106 : performing contrast enhancement and texture detail restoration processing on the two-dimensional code image within each detection frame.

[0064] In this embodiment, the purpose is to improve the local image quality of each QR code target detected by the model, providing a more stable and clear QR code image input for the subsequent decoding module. The contrast enhancement part uses image enhancement technologies such as local histogram equalization and adaptive gamma correction to make the black and white contrast in the QR code image more obvious, making it easier to identify its structural boundaries and coding dot matrix; the texture detail restoration part uses edge-preserving filtering, guided filtering or nonlinear sharpening methods to enhance the boundary lines and corner information of the QR code image while avoiding amplifying image noise. This processing process is designed for the QR code structure rather than general image enhancement, and places special emphasis on protecting the integrity of the coding structure and the clarity of the identification points. This enhancement process can significantly improve the decoding success rate of QR code images in the case of blurring, local saturation or distortion caused by reflective interference, and has strong adaptability and robustness.

[0065] Step S107: sending the processed QR code image to a decoding module to extract the QR code content.

[0066] In this embodiment, the QR code image after detection, cropping and image enhancement is sent to the QR code decoding module. The decoding module can be a QR code decoding engine in the traditional ZBar or OpenCV library, or an end-to-end image decoder based on deep learning training. The decoding process identifies the QR code format (such as QR Code, DataMatrix, etc.) and extracts the encoded information such as serial number, batch number, task number, etc.; for the QR code sample with improved image quality, the decoding module can complete the analysis in a shorter time, reduce the bit error rate, and can further connect with the upper system (such as MES) to complete the traceability or sorting task, and finally realize the end-to-end multi-target QR code automatic recognition process suitable for complex scenes with reflective interference.

[0067] This application introduces multi-exposure image simulation, saliency modulation attention mechanism, residual feature fusion path, improved Soft-NMS algorithm, and image local enhancement and decoding process to form an integrated recognition chain in the reflective interference scenario. It can significantly improve the overall recognition accuracy and edge positioning robustness of QR codes under typical interference conditions such as transparent film, reflective material or local highlight on the QR code surface.

[0068] In some optional implementations of this embodiment, the step of performing multi-exposure simulation processing on the original image to generate multiple enhanced images with different brightness levels includes:

[0069] The brightness enhancement method based on Retinex theory performs multiple exposure simulation processing on the original image;

[0070] Construct a multi-branch convolutional network to extract features of images with different exposures;

[0071] The weighted fusion mechanism is used to fuse the feature maps of each branch to form an enhanced input feature map.

[0072] In this embodiment, a brightness enhancement method based on Retinex theory is first used to perform multiple exposure simulations on the original image. This method uses a human eye perception model to separate the illumination component and the reflection component in the image. Through logarithmic domain transformation and global or local illumination estimation functions, the brightness response of the image under three typical exposure conditions, high exposure, low exposure, and standard exposure, is simulated. Multiple sets of image copies with significantly different brightness distributions are output, further addressing the problem of missing image information in the QR code area due to local overexposure or shadow obscuration. Subsequently, a multi-branch convolutional neural network is constructed to extract features from each set of exposure image inputs. Each branch network shares part of the backbone structure but independently extracts texture details, edge information, and structural brightness maps at the corresponding exposure level. This setting enables the system to learn the potential representation features of the QR code area in a real environment under different illumination simulation conditions. Finally, a weighted fusion mechanism is used to fuse channel-level or spatial-level features of all branch feature maps. The fusion method can adopt attention mechanism weighting, convolution compression after feature map splicing, or average pooling superposition, so that the effectively supplemented QR code area features in multiple simulation images are unified and integrated, ultimately forming an enhanced high-quality input feature map for use by subsequent network modules.

[0073] This application combines the brightness enhancement strategy of Retinex theory with the multi-branch convolutional network feature extraction structure to simulate the appearance of QR codes under various exposure environments, allowing the network to perceive dark details and highlight occlusion information simultaneously during training and inference, further enhancing the model's ability to extract the saliency of QR codes in scenes with drastic brightness changes. The weighted fusion strategy ensures the coordinated use of multi-branch feature information, improving the overall perception tolerance while ensuring the lightweight of the model, and is particularly suitable for on-site application needs with severe reflective interference or severe image quality compression.

[0074] In some optional implementations of this embodiment, in the step of extracting a QR code candidate frame from the YOLOv8-SSD lightweight neural network model that integrates the enhanced image input with the saliency modulation layer and the residual feature fusion path, the saliency modulation layer includes:

[0075] Estimate the reflective interference area in the image by using the local brightness gradient entropy and pixel standard deviation;

[0076] A reflective mask map is constructed and inhibitory attention adjustment is performed on the feature map channel weights based on the mask map.

[0077] In this embodiment, in the step of extracting QR code candidate frames from a YOLOv8-SSD lightweight neural network model that integrates the enhanced image input with the saliency modulation layer and the residual feature fusion path, the saliency modulation layer is used to dynamically identify areas in the image that may produce reflective interference and suppress their feature responses. Its internal structure includes the following processing mechanism: First, by setting a local sliding window in the spatial dimension of the image or feature map, the brightness gradient entropy and the standard deviation of the pixel value are calculated for each window area, where the brightness gradient entropy is used to measure the complexity of the grayscale change of the image in the area, and the standard deviation reflects the discreteness of the brightness value. If a region has both a high brightness gradient and a high standard deviation, it usually indicates that there is a drastic light change in the region, which is very likely to be a mirror reflection, highlight overflow or light transmission interference area, so it can be determined as a potential reflective interference area; then, based on all areas marked as having high reflection probability, a reflection mask map (Saliency Reflection Mask) is generated. The mask image can be multiplied or weighted with the intermediate feature map channel by channel to achieve explicit feature suppression of the reflective interference area while retaining the feature expression ability of the non-interference area. This allows the neural network to pay more attention to the texture, edge and structural features of the QR code body in the subsequent processing process, avoiding the misidentification of reflective highlights as targets or background boundaries, and improving detection accuracy and positioning stability.

[0078] This application introduces a saliency modulation layer and constructs a reflective mask map based on brightness gradient entropy and pixel standard deviation, thereby realizing active recognition of highly reflective areas and channel-level attention suppression by the neural network, effectively reducing the risk of misleading target detection output by reflective noise, and enhancing the model's ability to focus on the structural features of QR codes under complex backgrounds. Compared with conventional CBAM or SE modules, its modulation is based on the real illumination disturbance structure of the image, and has stronger semantic consistency and actual suppression effect, significantly improving the target positioning stability of the model in mirror interference scenarios.

[0079] In some optional implementations of this embodiment, after extracting the QR code candidate frame by inputting the enhanced image into the YOLOv8-SSD lightweight neural network model that fuses the saliency modulation layer and the residual feature fusion path, the following steps are further included:

[0080] The unenhanced original image features are adjusted by 1×1 convolution and then channel-wise fused with the deep feature maps in the backbone network to retain the edge information of the QR code.

[0081] In this embodiment, an auxiliary feature fusion branch is introduced into the network structure to introduce the unenhanced original image feature map into the deep semantic feature stream of the backbone network. The specific method is to first perform 1×1 convolution processing on the shallow feature map of the original image (such as the first or second layer convolution output) to match the number of channels, feature size and data distribution, so that it can be structurally consistent with the deeper semantic feature map in the backbone network. Then, the transformed original image features are concatenated (concat) with the deep feature map in the channel dimension, so that the network retains the original structural information such as the QR code edge, geometric framework, dense texture, etc. while learning abstract semantics. This strategy is particularly suitable for situations where the QR code edge becomes blurred and the boundary is discontinuous due to enhancement processing or reflection interference. By supplementing the original low-level image information, the positioning accuracy and boundary consistency of the candidate box can be improved, and the model's anti-interference ability to adverse image factors such as QR code occlusion, degradation, and local defocus can be enhanced.

[0082] This application introduces the unenhanced original image features into the deep semantic expression through the residual feature fusion path, which can effectively retain the edge lines, contour information and positioning framework of the QR code, compensate for the local information loss caused by enhancement, compression or interference, and improve the accuracy and continuity of the boundaries while ensuring detection accuracy, and enhance the model's ability to recognize QR codes with blurred edges and partially occluded. At the same time, the residual fusion path is an independent branch structure with low computational complexity and flexible deployment, which is suitable for efficient operation on mobile terminals and edge devices.

[0083] In some optional implementations of this embodiment, the Soft-NMS algorithm includes:

[0084] Calculate the reflection probability of the candidate frame area based on its average brightness and edge direction distribution;

[0085] The confidence attenuation function between candidate frames is dynamically adjusted according to the reflection probability to achieve fault tolerance for weakly reflective QR codes.

[0086] In this embodiment, first, in the multiple candidate frames output by the model, the average brightness value and edge direction distribution characteristics of the image area corresponding to each candidate frame are extracted, wherein the average brightness can be obtained through grayscale statistics, reflecting the overall light intensity of the area, and the edge direction distribution is extracted through edge detection operators such as Sobel or Canny to obtain the histogram of the edge gradient direction, which is used to evaluate the concentration degree and direction change trend of the texture direction in the area. If a candidate frame area presents high brightness and the edge direction distribution is messy or has a strong directional jump, it usually indicates that there is mirror reflection, high spot disturbance or occlusion noise in the area. Based on this, a reflection probability model can be constructed, which takes the above brightness features and edge direction statistics as input and outputs the reflection interference probability score of each candidate frame; after obtaining the reflection of all candidate frames, After the light probability, the present invention further introduces the probability value into the Soft-NMS confidence suppression function. Specifically, on the basis of the original IoU (overlap) attenuation confidence, the confidence attenuation amplitude of the candidate box in the high reflection probability area is appropriately relaxed, that is, a reflection adjustment factor is added to the Gaussian or linear suppression function of Soft-NMS, so that this type of candidate box still has a certain retention weight when the confidence overlap conflicts, thereby increasing the recall probability of weakly reflective or edge-blurred QR code targets and reducing the false deletion under the conventional NMS mechanism. While ensuring the overall detection accuracy, this improved mechanism significantly enhances the model's fault tolerance and edge perception flexibility in reflection interference scenarios, and is particularly suitable for industrial recognition tasks where QR codes are densely arranged, covered with films or severely interfered with by light sources.

[0087] This application introduces a reflection probability estimation mechanism and dynamically adjusts the confidence attenuation strategy between candidate frames based on the average brightness of the image area and the edge direction distribution. This allows the low-confidence QR code candidate frames originally caused by reflection interference to be retained during the NMS process with fault tolerance, avoiding the false positives of conventional NMS algorithms when screening overlapping targets. This mechanism has a higher recall rate and tolerance for QR code targets such as weak reflections, semi-occlusions, and blurred edges, and is particularly suitable for stable detection in scenarios with densely arranged multiple QR codes.

[0088] In some optional implementations of this embodiment, the step of performing contrast enhancement and texture detail restoration processing on the QR code image within each detection frame includes:

[0089] Perform local histogram equalization to improve the contrast of the QR code area;

[0090] Edge-preserving filtering is used to restore the texture details of the QR code and enhance decoding stability.

[0091] In this embodiment, a local histogram equalization operation is first performed to enhance the contrast of the two-dimensional code image area. This operation is based on the grayscale distribution characteristics of the image. By dividing the two-dimensional code area into local windows, the grayscale value distribution in each window is stretched by histogram, so that the image originally concentrated in the grayscale range due to reflection, uneven exposure or compression noise can redistribute its light and dark contrast, enhance the black and white structure boundaries of the two-dimensional code, and have a significant extraction effect on the two-dimensional code with sparse structure, thin lines or strong background interference. Subsequently, the edge-preserving filter is used to perform texture detail restoration processing on the enhanced two-dimensional code image. This filtering method is different from ordinary high-pass sharpening or mean smoothing filtering. Its core is to construct an edge guide map of the image (for example, through a gradient field or brightness change model), perform denoising and smoothing processing on the non-edge area while retaining or enhancing the texture response of the edge area, for example, using a guided filter (Guided Filter), bilateral filtering or L0 smoothing, etc., can maintain the clarity and edge sharpness of key decoding structures such as positioning graphics, functional graphics, and data cells of the QR code image while removing light spot reflection and compression artifacts, thereby effectively improving the recognizability and stability of the QR code image after it is sent to the decoding module. This processing flow does not require changing the model architecture and has the advantages of strong versatility, low computational cost, and good compensation for local occlusion and structural missing of the QR code.

[0092] This application can significantly improve the local contrast and structural clarity of the QR code image by performing local histogram equalization and edge-preserving filtering before decoding the QR code image, enhance the visibility of key patterns such as the QR code edge lines, dot matrix and positioning graphics, effectively reduce the decoding failure rate caused by underexposure, reflective spots or compression artifacts, further improve the fault tolerance of the decoding module for complex image input, and ensure the stable operation of the system's end-to-end recognition success rate in industrial deployment.

[0093] In some optional implementations of this embodiment, in the step of performing multiple exposure simulation processing on the original image by the above-mentioned brightness enhancement method based on Retinex theory, the multiple exposure simulation processing simulates the visual performance of highlights and dark areas in the RGB image domain by adjusting the logarithmic brightness transformation parameters.

[0094] In this embodiment, without relying on an image acquisition device to provide real multi-exposure images, the Retinex image enhancement theory is used to separate and remap the luminance components of the original image in the RGB image domain. By adjusting control parameters in the logarithmic luminance transformation function, such as the logarithmic basis, luminance stretch coefficient, or illumination restoration factor, image versions corresponding to three typical lighting scenarios, high exposure, medium exposure, and low exposure, are simulated. When simulating high-exposure images, the logarithmic mapping function is set to a form with a relatively flat response curve, thereby enhancing the luminance details of dark areas in the original image and simulating the visual effect under high lighting conditions. When simulating low-exposure images, the response curve is set to compress the highlight area to reduce the overexposure effect of local strong reflective areas, thereby more realistically restoring the grayscale distribution of the QR code background in the image. All simulation processes are completed in the RGB image domain, without the need for conversion to perceptual spaces such as HSV or Lab. This has the advantages of simple implementation, strong color consistency, and intact preservation of structural textures.

[0095] This application realizes multiple exposure simulations by adjusting the logarithmic brightness mapping function in the RGB image domain. Without relying on HDR shooting equipment, it simulates the visibility differences of QR codes under high light, low light and normal light. The simulated image has the advantages of rich brightness levels and complete structural details, which can effectively enhance the network's ability to model target textures under complex lighting conditions, while avoiding color distortion and computational redundancy, so that the model has consistent and stable perception effects under different brightness conditions, effectively expanding the environmental adaptability of the system.

[0096] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.

[0097] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0098] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware via computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes in the above-described method embodiments. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0099] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0100] Further references Figure 2 , as a response to the above Figure 1 The present application provides an embodiment of an industrial multi-target intelligent code scanning system based on deep learning. Figure 2 Corresponding to the method embodiment shown, the system can be specifically applied to various electronic devices.

[0101] like Figure 2 As shown, the deep learning-based industrial multi-target intelligent scanning system 200 described in this embodiment includes: an acquisition module 201, a first processing module 202, a first extraction module 203, an adjustment module 204, an application module 205, a second processing module 206 and a second extraction module 207. Among them:

[0102] An acquisition module 201 is configured to acquire an original image containing a plurality of two-dimensional codes, wherein the two-dimensional codes are attached to an object having a reflective or transparent film surface;

[0103] The first processing module 202 is configured to perform multi-exposure simulation processing on the original image to generate multiple enhanced images with different brightness levels;

[0104] A first extraction module 203 is configured to input the enhanced image into a YOLOv8-SSD lightweight neural network model fused with a saliency modulation layer and a residual feature fusion path, to extract a QR code candidate frame;

[0105] Adjustment module 204, the saliency modulation layer adjusts the channel attention weight according to the local brightness gradient and standard deviation in the image to suppress reflection interference;

[0106] Application module 205, for applying the Soft-NMS algorithm with a reflective area adaptive weight strategy to the candidate boxes output by the model, retaining valid detection results of occluded or reflective areas;

[0107] The second processing module 206 is used to perform contrast enhancement and texture detail restoration processing on the QR code image within each detection frame;

[0108] The second extraction module 207 is used to send the processed two-dimensional code image to the decoding module to extract the content of the two-dimensional code.

[0109] The industrial whole-plate multi-target intelligent code scanning system based on deep learning provided by the embodiment of the present invention can realize all the processes of the industrial whole-plate multi-target intelligent code scanning method based on deep learning in the above-mentioned embodiment. The functions of each module in the system and the technical effects achieved are respectively the same as the functions and technical effects achieved by the industrial whole-plate multi-target intelligent code scanning method based on deep learning in the above-mentioned embodiment, and will not be repeated here.

[0110] To solve the above technical problems, the present application also provides a computer device. Figure 3 , Figure 3 This is a basic structural block diagram of the computer device in this embodiment.

[0111] The computer device 3 includes a memory 31, a processor 32, and a network interface 33 that are interconnected through a system bus. It should be noted that the figure only shows a computer device 3 with components 31-33, but it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Among them, those skilled in the art can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0112] The computer device may be a desktop computer, notebook computer, PDA, cloud server, etc. The computer device may interact with the user via a keyboard, mouse, remote control, touchpad, or voice control device.

[0113] The memory 31 includes at least one type of readable storage medium, including flash memory, a hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic storage, a magnetic disk, an optical disk, etc. In some embodiments, the memory 31 may be an internal storage unit of the computer device 3, such as the hard disk or internal memory of the computer device 3. In other embodiments, the memory 31 may also be an external storage device of the computer device 3, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash memory card, etc. equipped on the computer device 3. Of course, the memory 31 may also include both the internal storage unit of the computer device 3 and its external storage device. In this embodiment, the memory 31 is generally used to store the operating system and various application software installed on the computer device 3, such as computer-readable instructions for the industrial full-disk multi-target intelligent code scanning method based on deep learning. In addition, the memory 31 can also be used to temporarily store various types of data that have been output or are to be output.

[0114] In some embodiments, the processor 32 can be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 32 is generally used to control the overall operation of the computer device 3. In this embodiment, the processor 32 is used to execute computer-readable instructions or process data stored in the memory 31, such as the computer-readable instructions for executing the deep learning-based industrial full-disk multi-target intelligent code scanning method.

[0115] The network interface 33 may include a wireless network interface or a wired network interface. The network interface 33 is generally used to establish a communication connection between the computer device 3 and other electronic devices.

[0116] The computer device provided in this application, by introducing multi-exposure image simulation, saliency modulation attention mechanism, residual feature fusion path, improved Soft-NMS algorithm and image local enhancement and decoding process to form an integrated recognition chain in the reflective interference scenario, can significantly improve the overall recognition accuracy and edge positioning robustness of the QR code under typical interference conditions such as transparent film, reflective material or local highlight on the QR code surface.

[0117] The present application also provides another embodiment, namely, providing a computer-readable storage medium, which stores computer-readable instructions, and the computer-readable instructions can be executed by at least one processor to enable the at least one processor to perform the steps of the above-mentioned industrial whole-plate multi-target intelligent scanning method based on deep learning.

[0118] The computer-readable storage medium provided in this application, by introducing multi-exposure image simulation, saliency modulation attention mechanism, residual feature fusion path, improved Soft-NMS algorithm and image local enhancement and decoding process to form an integrated recognition link in the reflective interference scenario, can significantly improve the overall recognition accuracy and edge positioning robustness of the QR code under typical interference conditions such as transparent film, reflective material or local highlight on the QR code surface.

[0119] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0120] The above are only preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present application should be included in the scope of protection of the present application.

Claims

1. A method for intelligent multi-target scanning of industrial whole-plate based on deep learning, characterized by: The steps include: Acquire an original image containing a plurality of two-dimensional codes, wherein the two-dimensional codes are attached to an object having a reflective or transparent film surface; Perform multi-exposure simulation processing on the original image to generate multiple enhanced images with different brightness levels; Inputting the enhanced image into a YOLOv8-SSD lightweight neural network model that integrates a saliency modulation layer and a residual feature fusion path to extract QR code candidate frames; The saliency modulation layer adjusts the channel attention weight according to the local brightness gradient and standard deviation in the image to suppress reflection interference; Apply the Soft-NMS algorithm with an adaptive weight strategy for reflective areas to the candidate boxes output by the model to retain valid detection results in occluded or reflective areas; Perform contrast enhancement and texture detail restoration on the QR code image within each detection frame; The processed QR code image is fed into the decoding module to extract the QR code content.

2. The industrial whole-plate multi-target intelligent code scanning method based on deep learning according to claim 1 is characterized in that: The step of performing multi-exposure simulation processing on the original image to generate multiple enhanced images with different brightness levels includes: The brightness enhancement method based on Retinex theory performs multiple exposure simulation processing on the original image; Construct a multi-branch convolutional network to extract features of images with different exposures; The weighted fusion mechanism is used to fuse the feature maps of each branch to form an enhanced input feature map.

3. The industrial whole-plate multi-target intelligent code scanning method based on deep learning according to claim 1 is characterized in that: In the step of extracting a QR code candidate frame by inputting the enhanced image into a YOLOv8-SSD lightweight neural network model that is fused with a saliency modulation layer and a residual feature fusion path, the saliency modulation layer includes: Estimate the reflective interference area in the image by using the local brightness gradient entropy and pixel standard deviation; A reflective mask map is constructed and inhibitory attention adjustment is performed on the feature map channel weights based on the mask map.

4. The industrial whole-plate multi-target intelligent code scanning method based on deep learning according to claim 3 is characterized in that: After the step of inputting the enhanced image into the YOLOv8-SSD lightweight neural network model fused with the saliency modulation layer and the residual feature fusion path, and extracting the QR code candidate frame, the method further includes: The unenhanced original image features are adjusted by 1×1 convolution and then channel-wise fused with the deep feature maps in the backbone network to retain the edge information of the QR code.

5. The industrial whole-plate multi-target intelligent code scanning method based on deep learning according to claim 1 is characterized in that: The Soft-NMS algorithm includes: Calculate the reflection probability of the candidate frame area based on its average brightness and edge direction distribution; The confidence attenuation function between candidate frames is dynamically adjusted according to the reflection probability to achieve fault tolerance for weakly reflective QR codes.

6. The industrial whole-plate multi-target intelligent code scanning method based on deep learning according to claim 1 is characterized in that: The step of performing contrast enhancement and texture detail restoration processing on the two-dimensional code image within each detection frame includes: Perform local histogram equalization to improve the contrast of the QR code area; Edge-preserving filtering is used to restore the texture details of the QR code and enhance decoding stability.

7. The method for industrial multi-target intelligent scanning based on deep learning according to claim 2 is characterized in that: In the step of performing multiple exposure simulation processing on the original image by the brightness enhancement method based on Retinex theory, the multiple exposure simulation processing simulates the visual performance of highlights and dark areas in the RGB image domain by adjusting logarithmic brightness transformation parameters.

8. An industrial multi-target intelligent scanning system based on deep learning, characterized by: include: An acquisition module is used to acquire an original image containing a plurality of two-dimensional codes, wherein the two-dimensional codes are attached to an object having a reflective or transparent film surface; A first processing module is used to perform multi-exposure simulation processing on the original image to generate multiple enhanced images with different brightness levels; A first extraction module is configured to input the enhanced image into a YOLOv8-SSD lightweight neural network model fused with a saliency modulation layer and a residual feature fusion path to extract a QR code candidate frame; An adjustment module, wherein the saliency modulation layer adjusts the channel attention weight according to the local brightness gradient and standard deviation in the image to suppress reflection interference; The application module is used to apply the Soft-NMS algorithm with the adaptive weight strategy of the reflective area to the candidate boxes output by the model, retaining the valid detection results of the occluded or reflective areas; The second processing module is used to perform contrast enhancement and texture detail restoration processing on the QR code image in each detection frame; The second extraction module is used to send the processed two-dimensional code image to the decoding module to extract the content of the two-dimensional code.

9. A computer device, characterized in that: It includes a memory and a processor, wherein the memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, it implements the steps of the industrial whole-plate multi-target intelligent scanning method based on deep learning as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the industrial whole-disk multi-target intelligent code scanning method based on deep learning as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Winter jujube defect identification method and system based on deep learning

    CN122368643A