Intelligent visual defect detection method and system for ton barrel production line and storage medium

CN122550504APending Publication Date: 2026-08-11SHANGHAI SAKURA PLASTIC PROD CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-15
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0006]本发明提供了一种吨桶生产线的智能视觉缺陷检测方法、系统及存储介质,以解决现有基于单一光学模态的视觉检测技术,在面对吨桶这类多材质嵌套、结构复杂且存在视觉遮挡的工业容器时,所存在的特征串扰、检测盲区以及全局三维重建算力成本高昂、无法满足高速在线检测需求的核心技术问题

Benefits of technology

[0056](1) This invention acquires multi-modal orthogonal images such as transient temperature field, polarization-preserving reflection and depolarization transmission, and extracts the topological mask matrix representing the rigid skeleton from them. A series of image decoupling and feature extraction operations are performed to achieve pixel-level physical separation of the multi-material features of the ton barrel. This eliminates the crosstalk between the optical features of the metal and plastic regions from the source, so that subsequent defect identification can be based on pure material features, which greatly improves the accuracy and robustness of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122550504A_ABST
    Figure CN122550504A_ABST
Patent Text Reader

Abstract

This invention discloses an intelligent visual defect detection method, system, and storage medium for ton barrel production lines, relating to the field of industrial automation quality inspection technology. The method includes: acquiring multimodal orthogonal image data of ton barrels to be inspected on the production line; extracting features from transient temperature field images in the multimodal orthogonal image data to generate a topological mask matrix; performing spatial registration and logical operations on the mask matrix to extract external rigid feature maps, external flexible feature maps, and transmitted light intensity distribution feature maps; segmenting the external flexible feature map into multiple local coordinate sub-panes based on the topological mask matrix, and calculating the relative curvature gradient of the optical texture to obtain local differential morphology features; and determining multi-dimensional defect judgment results based on the external rigid feature map, external flexible feature map, transmitted light intensity distribution feature map, and local differential morphology features to trigger corresponding operations, thereby achieving comprehensive, accurate, and efficient detection of all types of defects across the entire ton barrel area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial automation quality inspection technology, and in particular to an intelligent visual defect detection method, system and storage medium for ton barrel production lines. Background Technology

[0002] Existing technologies face profound contradictions when dealing with high-speed online visual inspection of interlocking container bins (IBCs). The core issue lies in the "structural mismatch" between the single-modal (e.g., visible light) surface optical data processing paradigm and the IBC's multi-material (metal frame and plastic liner) and complex nested topological structure. Specifically, this manifests as follows:

[0003] 1. Material crosstalk and interference: The ton container is made of highly reflective metal and diffuse reflective plastic nested together. Traditional single visible light imaging is easily affected by ambient light interference. Under dirty working conditions, it is easy to cause material misjudgment and feature crosstalk, resulting in instability of the detection system.

[0004] 2. Blind spots exist in the inspection: The outside of the ton container is fixed with a large metal nameplate, and traditional external light source imaging cannot obtain image data of the plastic inner liner behind it, forming a fatal blind spot in quality inspection, which may lead to the failure to detect internal defects.

[0005] 3. Low efficiency of 3D inspection: To detect the deformation (bulging / shrinkage) of the plastic inner liner, traditional methods use 3D laser point cloud scanning for global surface reconstruction. This method has high hardware costs, large data processing latency (on the order of seconds), and huge computing power overhead, which cannot meet the millisecond-level online real-time full inspection requirements of the production line. Under high concurrency, it is prone to system response delays. Summary of the Invention

[0006] This invention provides an intelligent visual defect detection method, system, and storage medium for ton barrel production lines, in order to solve the core technical problems of existing single-optical-modality visual inspection technologies when dealing with industrial containers such as ton barrels, which have multiple nested materials, complex structures, and visual occlusions. These problems include feature crosstalk, detection blind spots, and high computational costs for global 3D reconstruction, which cannot meet the requirements of high-speed online inspection.

[0007] Firstly, an intelligent visual defect detection method for ton drum production lines is provided, including...

[0008] Acquire multimodal orthogonal image data of the ton barrels to be tested on the production line under multidimensional physical field excitation, the multimodal orthogonal image data including transient temperature field image, polarization-preserving reflection image and depolarization transmission image;

[0009] Feature extraction is performed on the transient temperature field image to generate a topological mask matrix representing the rigid skeleton of the target to be detected;

[0010] The topological mask matrix is ​​spatially registered and logically operated with the polarization-preserving reflection image and the depolarization-transmission image respectively, and the external rigid feature map and the external flexible feature map are decoupled and extracted. Based on the occlusion region coordinates of the topological mask matrix, the transmitted light intensity distribution feature map is extracted from the depolarization-transmission image.

[0011] Based on the topological mask matrix, the external flexible feature map is divided into multiple local coordinate sub-panes, and the relative curvature gradient of the optical texture is calculated in each local coordinate sub-pane to obtain local differential morphology features.

[0012] The external rigidity feature map, the external flexible feature map, the transmitted light intensity distribution feature map, and the local differential morphology feature are input into a pre-trained multi-task visual detection model to output multi-dimensional defect judgment results.

[0013] Based on the defect determination result, the corresponding operation on the ton to be tested is triggered according to the preset rules.

[0014] Optionally, the step of extracting features from the transient temperature field image to generate a topological mask matrix representing the rigid skeleton of the target to be detected includes:

[0015] Calculate the global thermal gradient histogram of the transient temperature field image;

[0016] Based on the global thermal gradient histogram, an adaptive threshold algorithm is used to determine the target segmentation threshold;

[0017] The transient temperature field image is binarized using the target segmentation threshold to generate a two-dimensional matrix composed of a first logical value and a second logical value. The two-dimensional matrix is ​​used as the topological mask matrix, wherein the first logical value represents a rigid skeleton region with high thermal conductivity.

[0018] Optionally, the step of spatially registering and logically operating the topological mask matrix with the polarization-preserving reflectance image and the depolarization-depolarizing transmission image, respectively, to decouple and extract the external rigid feature map and the external flexible feature map, includes:

[0019] Obtain a pre-calibrated affine transformation matrix, and use the affine transformation matrix to map the polarization-preserving reflection image and the depolarization-transmission image to the pixel coordinate system where the topological mask matrix is ​​located, thus completing spatial registration;

[0020] The external rigidity feature map is extracted by performing a Hadamard product operation on the topological mask matrix and the registered polarization-preserving reflection image.

[0021] A bitwise NOT operation is performed on the topological mask matrix to obtain an inverted mask matrix. The inverted mask matrix is ​​then subjected to a Hadamard product operation with the registered depolarized transmission image to extract the external flexible feature map.

[0022] Optionally, extracting the transmitted light intensity distribution feature map from the depolarized transmission image based on the occlusion region coordinates of the topological mask matrix includes:

[0023] Perform a connected component labeling algorithm on the topology mask matrix to extract target connected components with an area greater than a preset area threshold;

[0024] Obtain the set of bounding box coordinates of the target connected component, and use the set of bounding box coordinates as the coordinates of the occluded region;

[0025] Based on the coordinates of the occlusion region, the corresponding pixel sub-matrix is ​​extracted from the registered depolarized transmission image and used as the transmitted light intensity distribution feature map.

[0026] Optionally, based on the topological mask matrix, the external flexible feature map is segmented into multiple local coordinate sub-panes, and the relative curvature gradient of the optical texture is calculated within each local coordinate sub-pane to obtain local differential topography features, including:

[0027] Extract the set of edge contour points representing the rigid skeleton from the topological mask matrix;

[0028] Using the set of edge contour points as a closed boundary, the external flexible feature map is divided into multiple independent local coordinate sub-panes;

[0029] Within each of the local coordinate sub-panes, calculate the Euclidean distance from the pixel to the nearest set of edge contour points;

[0030] Calculate the grayscale gradient magnitude of the pixel in the external flexible feature map;

[0031] The ratio of the grayscale gradient magnitude to the Euclidean distance is used as the relative curvature gradient. When the relative curvature gradient exceeds a preset deformation tolerance threshold, the corresponding pixel region is extracted as the local differential morphology feature.

[0032] Optionally, the step of inputting the external rigidity feature map, the external flexible feature map, the transmitted light intensity distribution feature map, and the local differential morphology feature into a pre-trained multi-task visual detection model to output multi-dimensional defect judgment results includes:

[0033] The external rigidity feature map, the external flexible feature map, the transmitted light intensity distribution feature map, and the local difference morphology feature are respectively converted into high-dimensional input tensors with a preset dimension.

[0034] Each of the high-dimensional input tensors is input into the corresponding independent feature extraction backbone branch in the multi-task visual detection model, and deep semantic feature tensors are extracted through multi-layer two-dimensional convolution operations.

[0035] Using the deep semantic feature tensor corresponding to the external rigid feature map as the query vector and the deep semantic feature tensor corresponding to the external flexible feature map as the key vector and value vector, perform cross-modal cross attention matrix multiplication to generate a rigid-flexible topological space attention weight matrix.

[0036] The rigid-flexible topological space attention weight matrix and each of the deep semantic feature tensors are concatenated through channels to generate a global fusion feature tensor.

[0037] The global fusion feature tensor is distributed to multiple independent task prediction heads, and the multi-dimensional defect determination results, including rigid defect bounding boxes, flexible defect bounding boxes, blind zone anomaly classification probabilities, and local form and position tolerance regression values, are output in parallel.

[0038] Secondly, an intelligent visual defect detection system for a ton-of-ton (TOT) production line is provided. This system is used to execute the intelligent visual defect detection method for a TOT production line according to any embodiment of the present invention, including:

[0039] The acquisition module is used to acquire multimodal orthogonal image data of the ton barrels to be tested on the production line under multidimensional physical field excitation. The multimodal orthogonal image data includes transient temperature field images, polarization-preserving reflection images, and depolarization-depolarizing transmission images.

[0040] The first feature extraction module is used to extract features from the transient temperature field image and generate a topological mask matrix that characterizes the rigid skeleton of the target to be detected.

[0041] The second feature extraction module is used to perform spatial registration and logical operations between the topological mask matrix and the polarization-preserving reflection image and the depolarization-transmission image, respectively, to decouple and extract the external rigid feature map and the external flexible feature map, and to extract the transmitted light intensity distribution feature map from the depolarization-transmission image based on the occlusion region coordinates of the topological mask matrix.

[0042] The segmentation module is used to segment the external flexible feature map into multiple local coordinate sub-panes based on the topological mask matrix, and to calculate the relative curvature gradient of the optical texture in each local coordinate sub-pane to obtain local differential morphology features.

[0043] The output module is used to input the external rigid feature map, the external flexible feature map, the transmitted light intensity distribution feature map, and the local differential morphology feature into the pre-trained multi-task visual detection model, and output multi-dimensional defect judgment results.

[0044] The execution module is used to trigger corresponding operations on the ton barrel to be inspected according to preset rules based on the defect determination result.

[0045] Optionally, the first feature extraction module is specifically used for:

[0046] Calculate the global thermal gradient histogram of the transient temperature field image;

[0047] Based on the global thermal gradient histogram, an adaptive threshold algorithm is used to determine the target segmentation threshold;

[0048] The transient temperature field image is binarized using the target segmentation threshold to generate a two-dimensional matrix composed of a first logical value and a second logical value. The two-dimensional matrix is ​​used as the topological mask matrix, wherein the first logical value represents a rigid skeleton region with high thermal conductivity.

[0049] Optionally, the segmentation module is specifically used for:

[0050] Extract the set of edge contour points representing the rigid skeleton from the topological mask matrix;

[0051] Using the set of edge contour points as a closed boundary, the external flexible feature map is divided into multiple independent local coordinate sub-panes;

[0052] Within each of the local coordinate sub-panes, calculate the Euclidean distance from the pixel to the nearest set of edge contour points;

[0053] Calculate the grayscale gradient magnitude of the pixel in the external flexible feature map;

[0054] The ratio of the grayscale gradient magnitude to the Euclidean distance is used as the relative curvature gradient. When the relative curvature gradient exceeds a preset deformation tolerance threshold, the corresponding pixel region is extracted as the local differential morphology feature.

[0055] Thirdly, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions, the computer instructions being used to cause a processor to execute and implement the intelligent visual defect detection method for a ton barrel production line according to any embodiment of the present invention.

[0056] (1) This invention acquires multi-modal orthogonal images such as transient temperature field, polarization-preserving reflection and depolarization transmission, and extracts the topological mask matrix representing the rigid skeleton from them. A series of image decoupling and feature extraction operations are performed to achieve pixel-level physical separation of the multi-material features of the ton barrel. This eliminates the crosstalk between the optical features of the metal and plastic regions from the source, so that subsequent defect identification can be based on pure material features, which greatly improves the accuracy and robustness of detection.

[0057] (2) This invention utilizes a topological mask to precisely locate the area obscured by the metal nameplate and extracts the transmitted light intensity distribution characteristics of this area based on depolarized transmission images. This eliminates the detection blind spots caused by structural component obstruction at the data level. By analyzing the uniformity of light intensity penetrating the plastic liner, it can effectively detect hidden defects such as cracks and foreign objects in the area covered by the metal nameplate, achieving "see-through" inspection of the ton container and overcoming the inherent limitations of traditional external imaging techniques.

[0058] (3) This invention uses the edges of the rigid skeleton in the topological mask matrix as natural boundaries to divide the flexible feature map into multiple local coordinate sub-panes, and calculates the relative curvature gradient of the optical texture in situ within each sub-pane to characterize deformation. The complex global 3D point cloud reconstruction and registration problem is reduced to local gradient calculation in a 2D image plane, which significantly reduces the algorithm complexity. It achieves millisecond-level, high-precision online detection of bulging, shrinkage and other geometric tolerances in plastics with extremely low hardware cost, completely avoiding the computational disaster of traditional 3D scanning methods.

[0059] (4) This invention uses a multi-task visual detection model with independent branches and cross-modal attention mechanism to process and infer the decoupled multi-dimensional heterogeneous feature maps in parallel. This enables the model to output multi-dimensional judgment results such as rigid defects, flexible defects, blind zone anomalies and deformation tolerances at one time, and integrates multiple detection processes that originally needed to be called serially into a single efficient forward propagation. This reduces the system processing latency from hundreds of milliseconds to tens of milliseconds, significantly improving data throughput and system real-time performance, perfectly meeting the high-speed cycle requirements of industrial production lines.

[0060] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0061] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0062] Figure 1 A flowchart of an intelligent visual defect detection method for a ton barrel production line is provided in Embodiment 1 of the present invention;

[0063] Figure 2 This is a schematic diagram of the framework of an intelligent visual defect detection system for a ton barrel production line provided in Embodiment 3 of the present invention;

[0064] Figure 3 A schematic diagram of an electronic device that can be used to implement embodiments of the present invention is shown. Detailed Implementation

[0065] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0066] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0067] Application Overview:

[0068] The inherent characteristics of existing technologies at the principle level have gradually revealed deep-seated contradictions when addressing the new challenges of high-speed online visual inspection of ton containers (such as IBC containers). Fundamentally, this stems from a structural mismatch between the single-modal surface optical data processing paradigm and the three-dimensional spatial characteristics of multi-material, complex topological targets. On one hand, ton containers consist of a highly reflective metal outer frame nested within a diffusely reflective, semi-transparent plastic liner. Existing image processing systems, when faced with mixed materials, rely on external industrial control computers for frequent sensor scheduling and image semantic segmentation model switching. This semantic segmentation based on a single visible light image is highly susceptible to ambient light interference, and under extremely dirty conditions, it is prone to material misjudgment and feature crosstalk, leading to image processing pipeline failure. On the other hand, a large area of ​​rigid metal nameplate is fixed to the outside of the ton container, and traditional external light source imaging cannot acquire image data of the plastic liner behind this metal plate, creating a fatal blind spot in quality inspection data. Furthermore, because plastic liners are prone to irregular bulging or shrinking after blow molding, traditional 3D laser point cloud scanning requires global coordinate registration and surface reconstruction of the entire giant ton container. This not only results in high hardware costs but also enormous computational overhead for point cloud data, with data processing latency typically exceeding seconds. This is completely unacceptable for the millisecond-level online real-time full inspection requirements of production lines. This crude approach of "global 3D reconstruction and single-modal recognition" not only wastes significant computing resources but also makes the system highly susceptible to response delays or even service interruptions under high-concurrency image stream input.

[0069] This invention, by introducing multimodal orthogonal image fusion and topological mask-based image decoupling technology, achieves physical-level separation of the multi-material features of the ton container (i.e., the metal frame and the high-molecular-weight, high-density polyethylene inner container) at the pixel level, breaking the feature crosstalk bottleneck of traditional single optical detection. It acquires depolarized transmission images through multi-dimensional physical field excitation and uses a topological mask to accurately locate the metal nameplate occlusion area, directionally extracting the transmitted light intensity distribution feature map, thus eliminating the detection blind spot caused by occlusion of the ton container structural components from the data source. Based on the extracted rigid topological mask, it constructs natural local coordinate sub-panes, reducing the global 3D point cloud reconstruction to in-situ differential curvature gradient calculation within a 2D image plane, accurately focusing on local deformation areas, and completely eliminating the computational disaster of global 3D data processing for large-size flexible bodies. The decoupled multi-dimensional heterogeneous image features are input in parallel into a multi-task visual detection model, significantly improving the feature extraction efficiency and classification throughput of complex defects in the ton container. It not only reduces image processing latency to the millisecond level and supports high-speed production line operation, but also significantly reduces GPU computing overhead through image logic operations and topology guidance. At the same time, the multimodal data fusion architecture provides a foundation for comprehensive quality control of complex containers, and promotes the leap of visual inspection of ton barrels from "surface visible light recognition" to "multi-dimensional data perspective and intelligent perception".

[0070] Example 1:

[0071] Figure 1 This is a flowchart illustrating an intelligent visual defect detection method for a ton-of-ton (TOT) production line according to Embodiment 1 of the present invention. This embodiment is applicable to situations where intelligent visual defect detection is performed on TOTs on a production line. This method can be executed by an intelligent visual defect detection system for the TOT production line. Figure 1 As shown, the method includes:

[0072] S110. Acquire multimodal orthogonal image data of the target to be inspected on the production line under multidimensional physical field excitation, wherein the multimodal orthogonal image data includes transient temperature field image, polarization-preserving reflection image and depolarization-depolarizing transmission image.

[0073] S120. Extract features from the transient temperature field image to generate a topological mask matrix that characterizes the rigid skeleton of the target to be detected.

[0074] S130. Spatial registration and logical operations are performed between the topological mask matrix and the polarization-preserving reflection image and the depolarization-transmission image, respectively, to decouple and extract the external rigid feature map and the external flexible feature map. Based on the occlusion region coordinates of the topological mask matrix, the transmitted light intensity distribution feature map is extracted from the depolarization-transmission image.

[0075] S140. Using the topological mask matrix as a reference, the external flexible feature map is divided into multiple local coordinate sub-panes, and the relative curvature gradient of the optical texture is calculated in each local coordinate sub-pane to obtain local differential morphology features.

[0076] S150. Input the external rigidity feature map, the external flexible feature map, the transmitted light intensity distribution feature map, and the local differential morphology feature into the pre-trained multi-task visual detection model, and output multi-dimensional defect judgment results.

[0077] S160. Based on the defect determination result, trigger the corresponding operation on the ton barrel to be tested according to the preset rules.

[0078] In this embodiment, multimodal orthogonal image fusion and topological mask-based image decoupling technology are used to achieve physical feature separation of the metal frame and plastic liner of the ton container at the pixel level, fundamentally breaking the bottleneck of crosstalk between multiple material features in traditional single optical inspection. By using topological masks to locate the occlusion area and extract the transmitted light intensity features, the detection blind spot caused by the metal nameplate is eliminated from the data source, realizing perspective detection of hidden defects. Based on the rigid topological mask, a local coordinate sub-pane is constructed, reducing the complex problem of global three-dimensional point cloud reconstruction to in-situ differential curvature gradient calculation of the two-dimensional image plane, achieving millisecond-level accurate detection of flexible liner deformation with extremely low computing power cost. Finally, the decoupled multidimensional heterogeneous features are input in parallel into a multi-task visual inspection model for joint inference, and multidimensional defect judgment results are output at one time, greatly compressing the system processing delay and significantly improving the detection throughput and real-time performance, perfectly meeting the high-speed requirements of industrial production lines.

[0079] Example 2:

[0080] In step S110, multimodal orthogonal image data of the ton to be tested under multidimensional physical field excitation is acquired. The multimodal orthogonal image data includes transient temperature field image, polarization-preserving reflection image and depolarization transmission image.

[0081] The ton container to be inspected refers to an industrial container composed of nested combinations of different materials, specifically a ton container with an external galvanized steel pipe frame and metal nameplate, and an internal blow-molded polyethylene liner. Multidimensional physical field excitation refers to the composite physical energy applied to the ton container to be inspected within an extremely short time window, including transient thermal pulses, high-frequency polarized light, and internal penetrating diffuse light. Multimodal orthogonal image data refers to a collection of multidimensional digital matrices captured by different types of image sensors, which are independent of each other in terms of physical properties (such as infrared band, specific polarization state, and visible light band). Transient temperature field image refers to a two-dimensional matrix captured by an infrared thermal imaging sensor, where each element represents the quantified temperature value of the corresponding spatial location on the surface of the target object instantaneously after the thermal pulse. Polarization-preserving reflection image refers to a two-dimensional matrix captured by an image sensor equipped with a polarization filter in the same direction as the excitation light source. Since the reflection from the metal surface usually maintains the polarization state of the original incident light, this matrix mainly records the surface optical details of the metal frame. Depolarized transmission images are two-dimensional matrices captured by image sensors equipped with orthogonal polarization filters. Since the diffuse reflection and transmission of the plastic liner destroy the polarization state of light, this matrix mainly records the surface texture of the plastic liner and the distribution of transmitted light intensity inside.

[0082] Specifically, within a very short time window, three composite physical energies are applied to the ton container to be tested. The ton container has an external metal frame and nameplate, and an internal high-molecular-weight polyethylene blow-molded liner, which can be understood as a plastic liner. The three composite physical energies can be transient thermal pulses, high-frequency polarized light, and internal penetrating diffused light. The transient thermal pulse is generated by an external transient thermal excitation source to excite a transient temperature field on the object's surface; for example, the external transient thermal excitation source can be a high-power infrared radiator. The high-frequency polarized light is generated by an external polarization optical excitation source to acquire the polarization optical response of the object's surface; for example, the external polarization optical excitation source can be an LED light source with a specific polarization state. The internal penetrating diffused light is generated by a penetrating diffused excitation source protruding into the ton container to illuminate the plastic liner from within.

[0083] Three types of image sensors were used simultaneously to capture responses sensitive to the aforementioned physical fields. These included an infrared thermal imaging sensor, an industrial camera equipped with a co-polarizing filter, and an industrial camera equipped with a cross-polarizing filter. The industrial camera with the co-polarizing filter had its filter polarization direction aligned with the polarization direction of the excitation light source. The industrial camera with the cross-polarizing filter had its filter polarization direction perpendicular to the polarization direction of the excitation light source. The infrared thermal imaging sensor captured the temperature distribution of the object's surface after a thermal pulse, generating a transient temperature field image. This image is a two-dimensional matrix, where each pixel value represents the quantized temperature at the corresponding point. The industrial camera equipped with the co-polarizing filter captured a polarization-preserving reflection image. Since reflection from the metal surface preserves the polarization state of light, this image primarily records the surface optical details of the metal frame. The industrial camera equipped with the cross-polarizing filter captured a depolarized transmission image. Since diffuse reflection and transmission from the plastic liner disrupt the polarization state of the (depolarized) light, this image primarily records the surface texture of the plastic liner and the internal transmitted light intensity distribution.

[0084] To ensure strict temporal alignment of multimodal images and prevent misalignment due to barrel movement, synchronization is achieved through a dedicated hardware network: when the barrel enters the inspection station, a photoelectric sensor sends a trigger signal. A programmable logic controller (PLC) generates a global master trigger pulse and sends it to a delay controller, such as an FPGA. Based on the different physical response time constants of each excitation source, the delay controller sequentially sends sub-trigger pulses with precise microsecond delays to the external transient thermal excitation source, the internal penetrating diffuse excitation source, and the external polarized optical excitation source, enabling them to reach peak output simultaneously. Within the time window of overlapping outputs from all excitation sources, the delay controller simultaneously sends synchronous exposure pulses to the infrared thermal imager, the polarization-preserving camera, and the depolarization-depolarizing camera, triggering their global shutters to open within the same microsecond, thereby acquiring three images that are strictly aligned in the temporal domain.

[0085] Image sensors transmit raw image data streams to edge computing nodes (such as industrial computers) via high-speed interfaces (e.g., GigE). The system asynchronously writes the data stream into a circular memory pool constructed from page-locked memory blocks pre-allocated in host memory via a direct memory access controller, bypassing the CPU to reduce copy overhead. The main processor's (CPU) image preprocessing thread reads the two-dimensional matrix data of the raw transient temperature field image, polarization-preserving reflectance image, and depolarization-depolarizing transmission image from the circular memory pool and feeds it into the subsequent processing pipeline.

[0086] In this embodiment, through a sophisticated hardware system (excitation source, sensor, synchronization controller) and underlying data transmission mechanism (DMA, ring memory pool), three orthogonal images reflecting different physical properties (thermal, polarization reflection, depolarization transmission) of the ton barrel are actively excited and captured at strictly synchronized time points, providing raw data for subsequent decoupling analysis and defect identification.

[0087] In step S120, the step of extracting features from the transient temperature field image to generate a topological mask matrix characterizing the rigid skeleton of the target to be detected includes:

[0088] Calculate the global thermal gradient histogram of the transient temperature field image;

[0089] Based on the global thermal gradient histogram, an adaptive threshold algorithm is used to determine the target segmentation threshold;

[0090] The transient temperature field image is binarized using the target segmentation threshold to generate a two-dimensional matrix composed of a first logical value and a second logical value. The two-dimensional matrix is ​​used as the topological mask matrix, wherein the first logical value represents a rigid skeleton region with high thermal conductivity.

[0091] The global thermal gradient histogram (GJH) can be defined as a statistical distribution map describing the drastic temperature changes (gradient magnitude) of all pixels in the entire transient temperature field image. It reflects the distribution of the number of "edges" (high-temperature difference areas, such as metal-plastic boundaries) and "flat areas" (low-temperature difference areas) in the image. An adaptive thresholding algorithm can be a type of image processing algorithm that automatically calculates the optimal segmentation threshold based on the current image content, such as the Otsu method. It eliminates the need for manually setting a fixed threshold and can adapt to changes in overall image brightness caused by different ambient temperatures and different batches of workpieces. The target segmentation threshold can be a specific temperature gradient value calculated by the adaptive thresholding algorithm. This threshold is the critical point for distinguishing between "suspected metal areas" (high gradient) and "non-metal areas" (low gradient).

[0092] Binarization refers to converting a grayscale image (transient temperature field image) into an image containing only two values ​​based on a threshold. All pixels greater than or equal to the threshold are grouped into one category, and those less than the threshold into another. In the topological mask matrix, pixels with a first logical value of "1" represent rigid skeleton regions with high thermal conductivity, i.e., the metal frame of the ton container. Pixels with a second logical value of "0" represent all other parts of the non-rigid skeleton, i.e., the flexible inner liner of the ton container or background space, which is a part of the inspection station other than the ton container itself, such as the conveyor belt, support structure, or the surrounding environment. A rigid skeleton refers to physical components in the target object that have high thermal conductivity, a fixed structure, and are not easily deformed, such as the external metal pipe mesh and metal nameplate of the ton container. A topological mask matrix is ​​a binary two-dimensional array that matches the resolution of the original image. Elements with a logical value of 1 correspond to the spatial position of the rigid skeleton, while elements with a logical value of 0 correspond to the spatial position of the flexible inner shell or background. It is used as a spatial index or logical operator in subsequent image processing.

[0093] Specifically, due to temperature fluctuations in industrial environments and initial temperature differences between different batches of ton containers, using a fixed temperature threshold cannot accurately separate metal from plastic. This invention utilizes a global thermal gradient histogram combined with Otsu's method to dynamically find the optimal segmentation point, ensuring the robustness of the mask. First, the transient temperature field image matrix is ​​obtained, denoted as T(x,y), where x∈[0,W-1], y∈[0,H-1], and W and H are the width and height of the image, respectively. The gradient image is calculated by performing a two-dimensional convolution operation on the temperature field image T(x,y) using image gradient operators (such as the Sobel operator and the Scharr operator) to obtain the temperature change components of each pixel in the x and y directions. and The gradient magnitude, i.e., the temperature gradient magnitude G(x,y) at each location (x,y) in the image, is typically calculated using the following formula: Or approximately This gradient value represents the degree of temperature change at that point. The edges of the metal skeleton exhibit high gradient values ​​due to the large temperature difference with the plastic; while the interior of the plastic area has a uniform temperature and low gradient values. The gradient magnitude distribution of all pixels in the entire gradient image is statistically analyzed. The gradient magnitude range (e.g., from 0 to the maximum value) is divided into several intervals (bars), and the number of pixels whose gradient magnitude falls within each interval is counted. The resulting "global thermal gradient histogram" has the horizontal axis representing the gradient magnitude interval and the vertical axis representing the pixel frequency of the corresponding interval. This histogram visually displays the overall proportion of strong and weak edges in the image.

[0094] Otsu's method is an automatic method for determining the image binarization threshold. Its goal is to find a threshold that maximizes the inter-class variance between the foreground (assuming metal, high gradient) and background (assuming plastic, low gradient) segments obtained by this threshold. A larger inter-class variance indicates a greater difference between the foreground and background, resulting in better segmentation. The specific process of determining the target segmentation threshold using an adaptive thresholding algorithm can be as follows: Given a possible threshold t, iterate through all possible values ​​of the gradient magnitude, for example, from the minimum to the maximum gradient. For each threshold t, divide the gradient image pixels into two classes: C0 (gradient < t) and C1 (gradient ≥ t). Calculate the proportion of pixels in each class relative to the entire image. , and their average gradient , Calculate the inter-class variance at the current threshold t:

[0095] ;

[0096] After traversing all t, find the class variance that is equal to the variance between classes. The largest threshold .Should That is, it is determined as the target segmentation threshold. This threshold is best able to distinguish between high-gradient regions (metal edges) and low-gradient regions in an image.

[0097] Iterate through each pixel coordinate (x, y) in the transient temperature field image T(x, y). Read the temperature value T(x, y) at that coordinate, and compare T(x, y) with the target segmentation threshold determined in the previous step. Compare them. If T(x,y)≥ If the pixel is considered to belong to a rigid skeleton region (metal) with high thermal conductivity, then a first logical value "1" is assigned to the corresponding position (x,y) in the topological mask matrix M. If T(x,y) < If a pixel is not found to be part of the plastic liner or background area, then the second logical value "0" is assigned to M(x,y). After traversing and assigning values ​​to all pixels in the entire image, the resulting two-dimensional matrix M containing only 0s and 1s is the topological mask matrix M(x,y).

[0098]

[0099] The matrix acts like a precise map, pinpointing the location of all the metal components in the original image.

[0100] In this embodiment, the physical difference in thermal conductivity between the metal frame and the plastic liner is utilized, rather than relying on easily disturbed optical appearance. By performing adaptive threshold segmentation on transient temperature field images, the metal frame structure of the ton container can be stably and accurately identified from complex backgrounds. The generated topological mask matrix acts like a precise "skeleton map," providing an unreliable benchmark for subsequent decoupling of multimodal images into pure rigid and flexible features. This fundamentally overcomes the feature crosstalk problem caused by material reflection and dirt in traditional visual inspection.

[0101] In step S130, the step of spatially registering and logically operating the topological mask matrix with the polarization-preserving reflection image and the depolarization-depolarizing transmission image, respectively, to decouple and extract the external rigid feature map and the external flexible feature map, includes:

[0102] Obtain a pre-calibrated affine transformation matrix, and use the affine transformation matrix to map the polarization-preserving reflection image and the depolarization-transmission image to the pixel coordinate system where the topological mask matrix is ​​located, thus completing spatial registration;

[0103] The external rigidity feature map is extracted by performing a Hadamard product operation on the topological mask matrix and the registered polarization-preserving reflection image.

[0104] A bitwise NOT operation is performed on the topological mask matrix to obtain an inverted mask matrix. The inverted mask matrix is ​​then subjected to a Hadamard product operation with the registered depolarized transmission image to extract the external flexible feature map.

[0105] The pre-calibrated affine transformation matrix can refer to a 3×3 mathematical matrix pre-calculated and stored through multi-sensor joint calibration. It defines the geometric mapping relationship from the visible light camera pixel coordinate system to the infrared thermal imager pixel coordinate system (i.e., the coordinate system where the topological mask matrix resides), including translation, rotation, scaling, and shearing transformations. The pixel coordinate system can refer to a digital image coordinate system with the top-left corner of the image as the origin (0,0), the positive x-axis (column) to the right, and the positive y-axis (row) downwards. Each integer coordinate pair (x, y) corresponds to one pixel. Spatial registration refers to the process of using the affine transformation matrix to perform geometric transformations on the image, such as bilinear interpolation, to eliminate parallax caused by differences in sensor installation position, viewing angle, and lens distortion, ensuring a one-to-one correspondence between pixels in different modalities of the image in spatial location. The Hadamard product operation refers to a matrix operation, specifically the element-wise multiplication of two matrices of the same dimension. For example, in the Hadamard product C of matrix A and matrix B, elements... Used as a logic filter, it uses 0 values ​​to mask irrelevant regions. Bitwise NOT operation refers to a logical operation on a binary matrix, inverting the logical value of each element. That is, "1" becomes "0", and "0" becomes "1". An inverted mask matrix refers to the new binary matrix obtained by performing a bitwise NOT operation on the topological mask matrix. In this matrix, the original metal regions (values ​​of 1) become 0, and the original plastic / background regions (values ​​of 0) become 1. An external rigid feature map refers to a two-dimensional image matrix that retains only the optical information of the metal frame surface after completely removing the plastic background and ambient light scattering interference at the pixel level. An external flexible feature map refers to a two-dimensional image matrix that exposes the optical information of the plastic inner liner surface after removing pixels obscured by the metal frame at the pixel level.

[0106] Specifically, the affine transformation matrix, pre-calculated using the checkerboard calibration method, is obtained. During system deployment, a one-time multimodal joint calibration is performed: a specially designed target (such as a chromium film cross-shaped target deposited on an alumina ceramic substrate) is used to generate high-contrast feature points under both infrared (thermal radiation) and visible light (reflection), with these feature points being physically coplanar and centered. The target is placed at the detection station, and the system synchronously triggers the excitation source to acquire infrared calibration images. and visible cursor fixed image Subpixel-level corner extraction algorithms are applied to both images, such as Shi-Tomasi corner detection + subpixel thinning, resulting in two sets of coordinates: an infrared feature point set. and visible light feature point set Using the infrared image coordinate system as a reference, the least squares method (such as singular value decomposition (SVD) is employed to solve the problem from... arrive homography matrix This matrix is ​​used as an affine transformation matrix and is stored on the system hard drive.

[0107] Load the precalibrated affine transformation matrix from storage. For real-time acquired polarization-preserving reflection images and depolarized transmission images ,application Since perspective transformation is performed, and the transformed pixel coordinates are usually non-integer, algorithms such as bilinear interpolation are needed to calculate the pixel values ​​of the target coordinates to generate a registered image with a resolution completely consistent with the topological mask matrix M. and At this point, the three images (M, , Achieve pixel-level spatial alignment.

[0108] The registered polarization-preserving reflectance image Perform the Hadamard product operation with the topological mask matrix M(x,y):

[0109] = ⊙M(x,y);

[0110] Here, ⊙ represents the Hadamard product (element-wise multiplication). Iterating through each pixel coordinate (x, y), if M(x, y) = 1 (metallic region), then... = Preserve the optical reflection information at that point. If M(x,y)=0 (plastic / background), then =0, masking the information at that point. The resulting matrix This is the external rigid feature map, which is an image with pixel values ​​only at the metal skeleton and all other areas are completely black, completely eliminating the interference of light scattered from the plastic background.

[0111] Plastic features are extracted by inverting the mask. A bitwise NOT operation is performed on the topological mask matrix M(x,y):

[0112] ;

[0113] Obtain the inversion mask matrix The original metal region is represented by 0, and the original plastic / background region by 1. The registered depolarized transmission image is then used. With the antiphase mask matrix Perform the Hadamard product operation:

[0114] = ⊙ ;

[0115] Iterate through each pixel coordinate (x, y), if =1 (plastic area), then = Preserve the transmission or texture information of that point. If =0 (metallic region), then =0, masking the information at that point (i.e., removing the portion obscured by the metal frame). The resulting matrix... This is the external flexible feature map, which is an image with pixel values ​​only in the plastic inner liner area and a completely black area in the metal frame, achieving pure extraction of features of flexible components.

[0116] In this embodiment, spatial registration ensures pixel alignment of multimodal images, and a topological mask matrix representing the metal skeleton is used as a precise "logic sieve." Through Hadamard product operations, the hybrid image can be efficiently and losslessly decoupled into pure external rigid feature maps and external flexible feature maps, fundamentally achieving physical separation of multi-material features at the pixel level. This completely eliminates crosstalk between optical features of heterogeneous materials, providing substrate-clean input data for subsequent high-precision defect identification.

[0117] In step S140, extracting the transmitted light intensity distribution feature map from the depolarized transmission image based on the occlusion region coordinates of the topological mask matrix includes:

[0118] Perform a connected component labeling algorithm on the topology mask matrix to extract target connected components with an area greater than a preset area threshold;

[0119] Obtain the set of bounding box coordinates of the target connected component, and use the set of bounding box coordinates as the coordinates of the occluded region;

[0120] Based on the coordinates of the occlusion area, the corresponding pixel sub-matrix is ​​extracted from the registered depolarized transmission image and used as the transmitted light intensity distribution feature map.

[0121] Calculate the two-dimensional Laplace energy sum of the transmitted light intensity distribution feature map. When the two-dimensional Laplace energy sum is greater than a preset uniformity threshold, generate a blind zone anomaly marker.

[0122] The connected component labeling algorithm can refer to an image processing algorithm used to identify all interconnected regions in a binary image that have the same pixel value (such as "1") and assign a unique label to each independent region. The document specifies the use of eight-connected component labeling, meaning a pixel is connected to its neighbors in the top, bottom, left, right, and four diagonal directions. The preset area threshold can refer to a pre-defined number of pixels used to filter out large, "meaningful" metal areas. Its purpose is to filter out scattered noise points or small metal structures (such as welds), retaining only large, rigid components that cause significant visual occlusion, such as metal nameplates. The target connected component can refer to the connected regions in the topological mask matrix whose total pixel area is greater than the preset area threshold after connected component labeling and area filtering. These regions correspond to large metal components that may cause occlusion blind spots and require special attention. The bounding box coordinate set can refer to the smallest bounding rectangle parameter describing the position of a target connected component in the image. The registered depolarized transmission image can refer to the depolarized transmission image that has been spatially aligned with the topological mask matrix. It records the light intensity distribution information transmitted from inside the ton container and passing through the plastic liner. A pixel submatrix can refer to a smaller, continuous rectangular image patch extracted from a large image matrix according to a specified range of rows and columns.

[0123] The occlusion region coordinates refer to the bounding box position parameters of the connected domain representing a large-area opaque rigid component (such as a metal nameplate) in the topological mask matrix, such as the pixel coordinates and width and height of the top-left corner. The transmitted light intensity distribution feature map refers to the pixel submatrix corresponding to the coordinate range of the occlusion region in the depolarized transmission image. This submatrix records the light intensity attenuation information after the internal light source penetrates the plastic wall. The two-dimensional Laplacian energy sum can refer to the sum of the absolute values ​​of all pixel values ​​in the resulting matrix after convolving the transmitted light intensity distribution feature map with the discrete Laplacian operator (a second-order differential operator). This value quantifies the total intensity of high-frequency changes (abrupt changes) in the image; a small value indicates uniform light intensity, while a large value indicates abrupt changes in light intensity (such as a drastic change in transmittance due to breakage). The preset uniformity threshold can refer to a pre-set scalar value used as a benchmark for judging whether the transmitted light intensity is uniform. If the calculated Laplacian energy sum exceeds this threshold, the transmitted light intensity distribution is considered non-uniform, potentially indicating a defect. A blind zone anomaly flag can be a binary logic signal or label (usually a Boolean value of True / False or 1 / 0). This flag is generated when the system determines that an anomaly exists within the blind zone, and serves as one of the key inputs for subsequent multi-task model determination of the "blind zone breakage" defect.

[0124] Specifically, given a binary topological mask matrix M(x,y), the algorithm scans the entire matrix. When a pixel with a value of "1" is found, it examines its eight neighboring pixels (up, down, left, right, top-left, top-right, bottom-left, bottom-right). All regions connected by pixels with "1" values ​​are grouped into connected components and assigned a unique label (e.g., 1, 2, 3...). Simultaneously or after labeling, the algorithm calculates the statistical properties of each connected component, most importantly its area (the total number of "1" pixels within that area). It iterates through all found connected components and calculates their areas. Compared with the preset area threshold Compare them. Only keep those that satisfy the criteria. The connected components are labeled as target connected components. These regions typically correspond to the large metal nameplates on the ton drum.

[0125] Input the coordinates of the occluded region (bounding box set) and the registered depolarized transmission image In memory, the image It is a two-dimensional array, based on the coordinate range x min To x max (Column range) and y min to y max (Row range), from The corresponding rectangular data block is copied from the original. This data block is the transmitted light intensity distribution characteristic map, denoted as... It is a small, independent image matrix, the size of which is equal to the width and height of the bounding box, and its content is entirely derived from the transmitted light intensity information of the area obscured by the metal nameplate.

[0126] Typically, a 3x3 discrete Laplacian convolution kernel is used, for example:

[0127] ;

[0128] Combine the Laplacian kernel with the feature map Perform convolution operations to obtain the Laplace response map. The absolute value of each pixel in the Laplacian response map reflects the drastic change in light intensity at that point. Summing the absolute values ​​of all pixels in the Laplacian response map L(x,y) yields the two-dimensional Laplacian energy. .

[0129] ;

[0130] Calculated With the preset uniformity threshold Compare. If This indicates an abnormal high-frequency abrupt change (non-uniformity) in the transmitted light intensity within the blocked area, potentially corresponding to a crack, foreign object, or severe thickness unevenness in the plastic wall. In this case, a blind zone anomaly flag is generated, for example, by setting the state variable to True or 1. If... This indicates that the transmitted light intensity distribution is relatively uniform, and the blind area may be normal. In this case, no abnormality flag is generated or the flag is False / 0.

[0131] In this embodiment, the area obscured by a large metal component (such as a nameplate) is precisely located, and the transmitted light intensity distribution of this area is extracted directionally from the depolarized transmission image. By analyzing the uniformity of this light intensity distribution—that is, calculating its two-dimensional Laplace energy—it is possible to effectively detect hidden defects such as cracks, foreign objects, or uneven thickness in the plastic liner within the obscured blind area. This achieves effective "see-through" detection of "visual blind spots" that traditional external imaging techniques cannot reach, fundamentally eliminating quality inspection blind spots.

[0132] In step S150, the step of segmenting the external flexible feature map into multiple local coordinate sub-panes based on the topological mask matrix, and calculating the relative curvature gradient of the optical texture within each local coordinate sub-pane to obtain local differential topography features includes:

[0133] Extract the set of edge contour points representing the rigid skeleton from the topological mask matrix;

[0134] Using the set of edge contour points as a closed boundary, the external flexible feature map is divided into multiple independent local coordinate sub-panes;

[0135] Within each of the local coordinate sub-panes, calculate the Euclidean distance from the pixel to the nearest set of edge contour points;

[0136] Calculate the grayscale gradient magnitude of the pixel in the external flexible feature map;

[0137] The ratio of the grayscale gradient magnitude to the Euclidean distance is used as the relative curvature gradient. When the relative curvature gradient exceeds a preset deformation tolerance threshold, the corresponding pixel region is extracted as the local differential morphology feature.

[0138] In this context, the reference refers to the set of absolute coordinates in the image processing coordinate system, serving as a positional reference and scale measure. Specifically, in this invention, it refers to the pixel boundaries representing the metal frame within the topological mask matrix. Local coordinate sub-panes refer to multiple independent polygonal image matrix regions naturally divided by the crisscrossing grid boundaries in the rigid topological mask, each region possessing an independent local two-dimensional coordinate system. Optical texture refers to the grayscale variation patterns or structured light stripes formed in the external flexible feature map due to surface undulations of the plastic liner or uneven transmission of internal light sources. Relative curvature gradient refers to the ratio of the grayscale change rate (gradient magnitude) of the optical texture in the two-dimensional image matrix to its physical distance to the nearest rigid reference boundary, used to equivalently characterize the degree of deformation in three-dimensional space in two-dimensional space. Local differential morphological features refer to the quantitative data set characterizing abnormal bulging, shrinkage, collapse, and other dimensional tolerance exceedance phenomena occurring within specific sub-panes of the ton's plastic liner.

[0139] Specifically, a binary topological mask matrix M(x,y) is input. A slight Gaussian blur is applied to M(x,y) using the Canny edge detection operator to smooth noise. The gradients of the image in the x and y directions are calculated using operators such as Sobel. and The gradient magnitude and direction are obtained. Along the gradient direction, pixels with the largest gradient magnitude are retained to refine the edges. Two thresholds are set: high and low. Pixels with gradient magnitudes greater than the high threshold are marked as strong edges, pixels with gradient magnitudes lower than the low threshold are suppressed, and pixels in between are only retained if they are connected to strong edges. The coordinates of all pixels marked as edges are calculated. These points are collected to form edge contour point set B. This point set accurately describes the contour of the metal skeleton in the image.

[0140] Identify individual closed contours from point set B. Each closed contour corresponds to the boundary of a "cell" in the metal mesh. Create a corresponding binary mask for each closed contour, filling the interior (plastic region) with 1s and the exterior with 0s. Assign each contour mask to the external flexible feature map. Perform a logical AND operation or directly use a mask for pixel indexing. The result of each operation is a local coordinate subpane. It contains image information of the plastic liner portion surrounded by the metal mesh. The outer flexible feature map is naturally segmented into K sub-images, each corresponding to a plastic region within a grid cell, which can be processed independently. For each pixel within a sub-pane, the image is processed using distance transformation or nearest neighbor search. Calculate its point set Each boundary point The Euclidean distance. From all these distances, take the minimum value as the distance from that pixel to the nearest metallic boundary. :

[0141] ;

[0142] In practical programming, to improve efficiency, distance transformation algorithms (such as the Felzenszwalb algorithm) are often used to calculate the distance from all pixels within a sub-pane to the nearest zero-value pixel (boundary) in one step. Here, the zero-value pixel is the boundary point. The Sobel operator is then used for the external flexible feature map. At pixel Perform a two-dimensional convolution on the values ​​of the convolution kernel and its neighborhood. Define a Sobel kernel in the horizontal direction. Vertical Sobel kernel Calculate the convolution response and horizontal gradient. Vertical gradient Where * represents the convolution operation. Calculate the gradient magnitude:

[0143] ;

[0144] Alternatively, for computational efficiency, an approximate formula can be used. . A larger value indicates a more drastic change in the optical texture (such as structured light stripes or natural texture) of the plastic surface near that point, potentially suggesting surface undulations (deformation). For sub-pane W... k Each pixel within Use the results from the first two steps to perform the calculation:

[0145] ;

[0146] Where ϵ is a very small positive number (e.g., 1e−6) to prevent the denominator from being zero. Iterate through all pixels within the sub-pane and determine whether their relative curvature gradient exceeds a preset deformation tolerance threshold, i.e., check the condition. Check if the condition is true. Record all pixels that satisfy the above condition and count them. Normally, overall deformation is not determined based on a single pixel exceeding the limit. An area percentage threshold (e.g., 15%) is introduced for secondary determination. The ratio of the area of ​​the exceeding pixel to the total area of ​​the sub-pane is calculated. If the area ratio exceeds a preset threshold, it is considered that significant local deformation has occurred in the plastic area within that sub-pane. In this case, the coordinate set of all pixels exceeding the threshold (or the identifier of the entire sub-pane) is extracted and output as a local differential topography feature. This feature indicates the specific location and extent of the "excessive bulging" or "excessive shrinkage collapse."

[0147] In this embodiment, the complex problem of 3D surface reconstruction and point cloud registration is cleverly reduced to the calculation of local gradient and distance transformation of a 2D image matrix, thus reducing the algorithm's time complexity from O(N^2) to O(N^2). 3 The computational cost has plummeted to O(N). This not only avoids the computational disaster caused by global scanning of large-size flexible inner liner, but also uses the topological mask of the metal frame as an in-situ reference to eliminate measurement errors caused by the overall displacement or rotation of the ton on the conveyor belt, achieving online full inspection of form and position tolerances with extremely low hardware cost and ultra-high frame rate.

[0148] In step S160, the input of the external rigidity feature map, the external flexible feature map, the transmitted light intensity distribution feature map, and the local differential topography feature into the pre-trained multi-task visual detection model, and the output of multi-dimensional defect judgment results, includes:

[0149] The external rigidity feature map, the external flexible feature map, the transmitted light intensity distribution feature map, and the local difference morphology feature are respectively converted into high-dimensional input tensors with a preset dimension.

[0150] Each of the high-dimensional input tensors is input into the corresponding independent feature extraction backbone branch in the multi-task visual detection model, and deep semantic feature tensors are extracted through multi-layer two-dimensional convolution operations.

[0151] Using the deep semantic feature tensor corresponding to the external rigid feature map as the query vector and the deep semantic feature tensor corresponding to the external flexible feature map as the key vector and value vector, perform cross-modal cross attention matrix multiplication to generate a rigid-flexible topological space attention weight matrix.

[0152] The rigid-flexible topological space attention weight matrix and each of the deep semantic feature tensors are concatenated through channels to generate a global fusion feature tensor.

[0153] The global fusion feature tensor is distributed to multiple independent task prediction heads, and the multi-dimensional defect determination results, including rigid defect bounding boxes, flexible defect bounding boxes, blind zone anomaly classification probabilities, and local form and position tolerance regression values, are output in parallel.

[0154] Among them, the multi-task visual inspection model refers to an artificial intelligence algorithm model based on a deep convolutional neural network (CNN) or visual transformer (ViT) architecture. This model includes a shared low-level feature extraction network and multiple independent task-specific prediction heads, capable of simultaneously processing heterogeneous input tensors and performing classification, object detection, and regression tasks. The multi-dimensional defect determination results refer to the structured data output by the model inference, including the bounding box coordinates and confidence scores of surface defects on the metal frame (such as rust and scratches), the bounding box coordinates and confidence scores of surface defects on the plastic liner (such as bubbles and black spots), the classification labels of hidden defects in blind spots, and the regression values ​​of the geometric tolerances of each sub-pane.

[0155] Tensor quantization refers to the process of converting a two-dimensional image pixel matrix into a multi-dimensional array (Tensor) that supports automatic differentiation and matrix multiplication operations by deep learning frameworks (such as PyTorch or TensorFlow). A multi-stream independent feature extraction backbone network refers to a network topology structure at the front end of a neural network, where multiple convolutional layers without shared weights are set up to extract low-level texture and edge features from different modalities of the image. Cross-modal cross-attention mechanism is a matrix operation method based on a variant of self-attention, which achieves weighted fusion of heterogeneous features by calculating the spatial correlation weights of one modality's features (such as metal topology) to another modality's features (such as plastic texture). A task prediction head is a lightweight network structure connected to the end of the backbone network, consisting of a small number of convolutional or fully connected layers, specifically designed to map high-dimensional feature tensors to a task-specific output format, such as coordinate values ​​or probability distributions. A joint loss function is an objective function that combines mathematical penalty terms from multiple different optimization objectives into a single scalar value through linear weighting, used to guide the entire multi-task network in a unified gradient descent process. Focus loss is a variant of cross-entropy loss that reduces the weight of easily classified samples by introducing a modulation factor, thus allowing the model to focus more on difficult-to-classify samples (such as small defects) during training. Generalized intersection-union loss is a bounding box regression evaluation metric that considers not only the overlap area between the predicted and ground truth boxes, but also the proportion of the non-overlapping region to the minimum closure region.

[0156] Specifically, algorithms such as bilinear interpolation are used to uniformly scale all feature maps to a fixed spatial resolution, such as H×W=512×512 mentioned in the document. This ensures the consistency of the network input size. A channel dimension is added to each scaled image matrix. For example, a single-channel grayscale image is converted into a tensor of shape (1, 512, 512); if needed, the channels can be duplicated to make it a 3-channel (3, 512, 512) tensor to fit the pre-trained backbone network. Multiple samples are stacked to form a batch processing dimension, ultimately resulting in four input tensors, for example: ; ...and so on. B represents the set of real numbers, B is the batch size, and C is the number of channels.

[0157] Four parallel feature extraction backbone branches are constructed, for example, using the convolutional block structure of a Residual Network (ResNet). Each branch has its own independent weight parameters. , Forward propagation process: ... Input branch ,Will Input branch And so on. Within each branch, the tensor undergoes a series of stacks of "convolutional layers, batch normalization layers, activation functions (such as ReLU), and pooling layers." As the network deepens, the spatial dimensions (height and width) of the feature maps gradually decrease (downsampling), while the number of channels gradually increases, and the semantic information becomes increasingly abstract. The end of each branch outputs a deep semantic feature tensor, for example: , Where h, w, and c are the downsampled height, width, and number of channels. These X tensors contain high-level features for defect identification.

[0158] rigidity and flexible features Using three different learnable weight matrices W respectively Q W K W V Perform a linear transformation to generate the query vector Q, the key vector K, and the value vector V:

[0159] , , ;

[0160] Calculate the dot product of the transposes of Q and K, then divide by a scaling factor. , The dimension of the key vector is then normalized using the Softmax function to obtain the attention weights.

[0161] ;

[0162] Multiplying the attention weights obtained in the previous step by the value vector V yields a weighted output, namely the attention weight matrix for rigid-flexible topological space. :

[0163] ;

[0164] By using an attention weight matrix, the model learns "which features of the plastic region should receive more attention based on a certain feature pattern of the metal region", thereby modeling the topological relationship between the two.

[0165] Based on the attention matrix Across and four deep semantic feature tensors , , , Assuming each tensor has a spatial dimension of h×w, and the number of channels is respectively , , , , Connect all these tensors end-to-end along the channel dimension: By concatenating along the channel axis, the global fused feature tensor is obtained. Its shape is h×w×( + + + + It centralizes the semantic information of all modalities as well as the interaction information between rigid and flexible modalities.

[0166] exist Then connect four independent task prediction heads, and Simultaneously, the data is fed into four heads, each composed of different network layers: The rigid defect detection head typically includes a classification branch (outputting defect category confidence) and a bounding box regression branch (outputting bounding box coordinate offsets), employing a detection head design similar to Faster R-CNN or YOLO. Output: (Category probability) (Bounding box coordinates). Flexible defect detection head: Similar in structure to the rigid head, but with independent parameters, specifically designed for plastic defects. Output: , Blind zone anomaly classification head: Typically contains a global average pooling (GAP) layer to compress spatial features into a vector, followed by a fully connected layer and a sigmoid activation function. Output: A scalar. ∈[0,1] represents the probability of an anomaly in the blind zone. Form and position tolerance regression head: may contain a fully connected layer, outputting one or more continuous numerical values. This represents the degree of deformation of each local sub-pane. In one forward propagation, the model generates all four outputs in parallel, forming a multi-dimensional defect determination result.

[0167] This embodiment constructs a multi-stream independent feature extraction and cross-modal cross-attention mechanism, enabling the model to autonomously learn the physical occlusion and topological nesting relationships between the metal frame and the plastic liner in a high-dimensional tensor space, completely avoiding feature annihilation caused by direct stitching of heterogeneous images. The introduction of a joint loss function allows the model to simultaneously complete three major computer vision tasks—object detection, image classification, and numerical regression—in a single forward inference process. This reduces the inference latency of the four independent models that previously required sequential calls from 800 milliseconds to less than 45 milliseconds, significantly improving the throughput of image data processing and perfectly meeting the industrial-grade requirements for efficient processing of massive image data.

[0168] In step S170, based on the defect determination result, a corresponding operation is triggered on the ton barrel to be inspected according to a preset rule.

[0169] The preset rules refer to the business logic judgment conditions composed of conditional statements pre-configured in the Manufacturing Execution System (MES) or Programmable Logic Controller (PLC). Its core function is to take the defect judgment results (such as defect type, level, and location) output by the multi-task visual inspection model as input, and through logical judgments, such as "if...then...", automatically trigger the corresponding control instruction sequence to drive subsequent operations such as audible and visual alarms, visual annotation, and sorting by the actuator, thereby achieving an automated closed loop of inspection and execution.

[0170] Specifically, the multi-dimensional judgment results output by the multi-task model are encapsulated into a structured data packet according to a preset format (such as JSON or XML). This data packet includes at least: product ID, timestamp, defect list (each defect includes type, coordinates, confidence level, and grade), blind zone status, deformation parameters, and the global result of this inspection (OK / NG). The data packet is published to a specific topic subscribed to by the MES system via industrial Ethernet using standard protocols (such as MQTT, OPC UA) or uploaded directly by calling the RESTful API interface provided by the MES. The entire process is typically completed within 100 milliseconds, achieving "real-time" processing. After receiving the data, the MES system immediately sends it to the rule engine. The engine compares the data with preset rules. For example, the rule might be: "IF global result is 'NG' AND any defect grade is 'critical' THEN trigger action: [alarm-emergency, sorting-scrap]". After rule matching, the MES generates a specific control command sequence. The MES outputs digital signals to the field alarm via the PLC, triggering audible and visual effects. For example, a tri-color light illuminates red and flashes, and a buzzer sounds. The MES pushes defect data to the SCADA / HMI system in real time. The SCADA system, in its rendered production line animation, highlights the defect location with a red border on the corresponding bin graphic and displays a defect details card. The MES or PLC calculates the sorting timing based on the product ID and defect location. When a defective product arrives at the sorting station via the conveyor belt, the PLC precisely controls the pneumatic pusher to push it into the scrap channel; or guides a robot to grab and remove it.

[0171] All data (raw images, feature maps, arbitration logs, etc.) is packaged and sent to the enterprise's data lake or dedicated storage server. Large files such as images may be stored directly in binary format, with their metadata (index) stored in the database. All data is ultimately archived in a relational database or data lake, indexed using key fields such as product serial number, timestamp, and production line number to form a complete "data history." When the market reports a problem with a batch of tonnes, the product batch number is entered into the traceability system. The system can immediately retrieve the complete inspection reports for all products in that batch, including the defect information identified at the time. Furthermore, the system can trace back to view the original multimodal images of the product during production to confirm whether the defect was missed or misjudged, and review the arbitration logs to understand the decision-making process. This achieves "one item, one file, lifelong traceability."

[0172] In this embodiment, based on the defect determination result, corresponding operations are triggered on the ton drum to be inspected according to preset rules. The virtual determination of the AI ​​model is transformed into physical control and information action of the production line in real time. Through the flexible configuration of preset rules in MES, automatic, accurate, and rapid sorting of non-conforming products is achieved (e.g., by pneumatic push rods or robots), and audible and visual alarms and defect visualization on the SCADA interface are triggered simultaneously, greatly improving the real-time performance of production response and the accuracy of human intervention. At the same time, the entire chain of data (from original images, intermediate features to determination results and arbitration logs) is forcibly structured and stored and associated with production batches. This not only provides an irrefutable data foundation for the full life cycle quality traceability of "one item, one file", but also allows any market feedback issues to be traced back to the original evidence and decision-making logic at the moment of production.

[0173] Optionally, pre-training a multi-task visual detection model is a systematic process, the core of which lies in using a joint loss function and a weight update strategy based on gradient norm dynamic balancing. The specific steps are as follows:

[0174] 1. Constructing the training dataset: Collect a large number of multimodal orthogonal images (transient temperature field, polarization-preserving reflection, and depolarization-transmission images) of ton barrels with known defects. Automated logical decoupling is performed using a topological mask matrix to generate a massive number of external rigid feature maps, external flexible feature maps, and transmitted light intensity distribution feature map samples. Local differential morphological feature samples are then generated through two-dimensional relative curvature gradient calculation. These samples are concatenated along the channel dimension into a high-dimensional heterogeneous training tensor, and a multi-dimensional ground truth label vector containing bounding box coordinates, defect categories, blind zone labels, and geometric tolerance values ​​is appended to it.

[0175] 2. Define the joint loss function: total loss It is the weighted sum of the losses from the four tasks:

[0176] The detection task loss is mainly for rigid / flexible defect bounding boxes: it includes focal loss to address the imbalance between positive and negative samples; and generalized intersection-union loss to optimize the position regression of the boxes.

[0177] The classification task loss mainly targets the probability of classifying abnormalities in the blind zone: a binary cross-entropy loss is adopted.

[0178] The regression task loss mainly targets local geometric tolerance values: smoothed L1 loss is used, which is more robust to outliers.

[0179] 3. Perform dynamic balancing training: In each training iteration, forward propagation calculates the independent loss value for each task. Calculate the gradient vector of the loss for each task relative to the weights of the shared feature extraction backbone network. Calculate the L2 norm of each gradient vector. and its global average Dynamically update the weighting coefficients of each loss term. The formula is:

[0180] ;

[0181] Where α is the momentum update rate and ϵ is a local constant. This mechanism automatically reduces the weights of tasks with excessively large gradient norms and increases the weights of tasks with excessively small gradient norms, thus balancing the contributions of each task to the network weight updates and effectively preventing the "negative transfer" phenomenon in multi-task learning. Using the updated... Calculate the weighted total loss And perform backpropagation and parameter updates.

[0182] Through the above process, the model can be jointly optimized to complete multiple tasks such as defect detection, classification and regression with high accuracy while sharing the extraction of underlying features.

[0183] Example 3:

[0184] Figure 2 This is a schematic diagram of the framework of an intelligent visual defect detection system for a ton-barrel production line provided in Embodiment 3 of the present invention. Figure 2 As shown, the system includes:

[0185] The acquisition module 210 is used to acquire multimodal orthogonal image data of the ton barrel to be tested under multidimensional physical field excitation. The multimodal orthogonal image data includes transient temperature field image, polarization-preserving reflection image and depolarization transmission image.

[0186] The first feature extraction module 220 is used to extract features from the transient temperature field image and generate a topological mask matrix that characterizes the rigid skeleton of the target to be detected.

[0187] The second feature extraction module 230 is used to perform spatial registration and logical operations on the topological mask matrix with the polarization-preserving reflection image and the depolarization-transmission image, respectively, to decouple and extract the external rigid feature map and the external flexible feature map, and to extract the transmitted light intensity distribution feature map from the depolarization-transmission image based on the occlusion region coordinates of the topological mask matrix.

[0188] The segmentation module 240 is used to segment the external flexible feature map into multiple local coordinate sub-panes based on the topological mask matrix, and to calculate the relative curvature gradient of the optical texture in each local coordinate sub-pane to obtain local differential morphology features.

[0189] The output module 250 is used to input the external rigid feature map, the external flexible feature map, the transmitted light intensity distribution feature map, and the local differential morphology feature into a pre-trained multi-task visual detection model, and output multi-dimensional defect judgment results.

[0190] The execution module 260 is used to trigger corresponding operations on the ton barrel to be inspected according to the defect determination result and preset rules.

[0191] Optionally, the first feature extraction module 220 is specifically used for:

[0192] Calculate the global thermal gradient histogram of the transient temperature field image;

[0193] Based on the global thermal gradient histogram, an adaptive threshold algorithm is used to determine the target segmentation threshold;

[0194] The transient temperature field image is binarized using the target segmentation threshold to generate a two-dimensional matrix composed of a first logical value and a second logical value. The two-dimensional matrix is ​​used as the topological mask matrix, wherein the first logical value represents a rigid skeleton region with high thermal conductivity.

[0195] Optionally, the second feature extraction module 230 is specifically used for:

[0196] Obtain a pre-calibrated affine transformation matrix, and use the affine transformation matrix to map the polarization-preserving reflection image and the depolarization-transmission image to the pixel coordinate system where the topological mask matrix is ​​located, thus completing spatial registration;

[0197] The external rigidity feature map is extracted by performing a Hadamard product operation on the topological mask matrix and the registered polarization-preserving reflection image.

[0198] A bitwise NOT operation is performed on the topological mask matrix to obtain an inverted mask matrix. The inverted mask matrix is ​​then subjected to a Hadamard product operation with the registered depolarized transmission image to extract the external flexible feature map.

[0199] Optionally, the second feature extraction module 230 is further used for:

[0200] Perform a connected component labeling algorithm on the topology mask matrix to extract target connected components with an area greater than a preset area threshold;

[0201] Obtain the set of bounding box coordinates of the target connected component, and use the set of bounding box coordinates as the coordinates of the occluded region;

[0202] Based on the coordinates of the occlusion region, the corresponding pixel sub-matrix is ​​extracted from the registered depolarized transmission image and used as the transmitted light intensity distribution feature map.

[0203] Optionally, the segmentation module 240 is specifically used for:

[0204] Extract the set of edge contour points representing the rigid skeleton from the topological mask matrix;

[0205] Using the set of edge contour points as a closed boundary, the external flexible feature map is divided into multiple independent local coordinate sub-panes;

[0206] Within each of the local coordinate sub-panes, calculate the Euclidean distance from the pixel to the nearest set of edge contour points;

[0207] Calculate the grayscale gradient magnitude of the pixel in the external flexible feature map;

[0208] The ratio of the grayscale gradient magnitude to the Euclidean distance is used as the relative curvature gradient. When the relative curvature gradient exceeds a preset deformation tolerance threshold, the corresponding pixel region is extracted as the local differential morphology feature.

[0209] Optional, output module 250, specifically used for:

[0210] The external rigidity feature map, the external flexible feature map, the transmitted light intensity distribution feature map, and the local difference morphology feature are respectively converted into high-dimensional input tensors with a preset dimension.

[0211] Each of the high-dimensional input tensors is input into the corresponding independent feature extraction backbone branch in the multi-task visual detection model, and deep semantic feature tensors are extracted through multi-layer two-dimensional convolution operations.

[0212] Using the deep semantic feature tensor corresponding to the external rigid feature map as the query vector and the deep semantic feature tensor corresponding to the external flexible feature map as the key vector and value vector, perform cross-modal cross attention matrix multiplication to generate a rigid-flexible topological space attention weight matrix.

[0213] The rigid-flexible topological space attention weight matrix and each of the deep semantic feature tensors are concatenated through channels to generate a global fusion feature tensor.

[0214] The global fusion feature tensor is distributed to multiple independent task prediction heads, and the multi-dimensional defect determination results, including rigid defect bounding boxes, flexible defect bounding boxes, blind zone anomaly classification probabilities, and local form and position tolerance regression values, are output in parallel.

[0215] The intelligent visual defect detection system for ton drum production lines provided in this embodiment of the invention can execute the intelligent visual defect detection method for ton drum production lines provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.

[0216] Example 4:

[0217] Figure 3A schematic diagram of an electronic device that can be used to implement embodiments of the present invention is shown. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0218] like Figure 3 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor 11, and the computer program is executed by the at least one processor 11 to enable the at least one processor 11 to perform the method provided by the present invention.

[0219] The processor 11 can perform various appropriate actions and processes based on a computer program stored in the read-only memory (ROM) 12 or a computer program loaded from the storage unit 18 into the random access memory (RAM) 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0220] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0221] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as the intelligent visual defect detection method in a tonne drum production line.

[0222] In some embodiments, the intelligent visual defect detection method for a tonne drum production line can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the intelligent visual defect detection method for a tonne drum production line described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the intelligent visual defect detection method for a tonne drum production line by any other suitable means (e.g., by means of firmware).

[0223] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard parts (ASSPs), systems-on-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0224] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0225] In the context of this invention, a computer-readable storage medium stores computer instructions that, when executed by a processor, implement the intelligent visual defect detection method for a ton-barrel production line provided by this invention. The computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. The computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, the computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0226] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD)) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0227] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0228] A computing system can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is established by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system. It addresses the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.

[0229] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0230] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for intelligent visual defect detection in a ton-barrel production line, characterized in that, include: Acquire multimodal orthogonal image data of the ton barrels to be tested on the production line under multidimensional physical field excitation, the multimodal orthogonal image data including transient temperature field image, polarization-preserving reflection image and depolarization transmission image; Feature extraction is performed on the transient temperature field image to generate a topological mask matrix representing the rigid skeleton of the target to be detected; The topological mask matrix is ​​spatially registered and logically operated with the polarization-preserving reflection image and the depolarization-transmission image respectively, and the external rigid feature map and the external flexible feature map are decoupled and extracted. Based on the occlusion region coordinates of the topological mask matrix, the transmitted light intensity distribution feature map is extracted from the depolarization-transmission image. Based on the topological mask matrix, the external flexible feature map is divided into multiple local coordinate sub-panes, and the relative curvature gradient of the optical texture is calculated in each local coordinate sub-pane to obtain local differential morphology features. The external rigidity feature map, the external flexible feature map, the transmitted light intensity distribution feature map, and the local differential morphology feature are input into a pre-trained multi-task visual detection model to output multi-dimensional defect judgment results. Based on the defect determination result, the corresponding operation on the ton to be tested is triggered according to the preset rules.

2. The method according to claim 1, characterized in that, The step of extracting features from the transient temperature field image to generate a topological mask matrix representing the rigid skeleton of the target to be detected includes: Calculate the global thermal gradient histogram of the transient temperature field image; Based on the global thermal gradient histogram, an adaptive threshold algorithm is used to determine the target segmentation threshold; The transient temperature field image is binarized using the target segmentation threshold to generate a two-dimensional matrix composed of a first logical value and a second logical value. The two-dimensional matrix is ​​used as the topological mask matrix, wherein the first logical value represents a rigid skeleton region with high thermal conductivity.

3. The method according to claim 1, characterized in that, The step of spatially registering and logically operating the topological mask matrix with the polarization-preserving reflectance image and the depolarization-depolarizing transmission image, respectively, to decouple and extract the external rigid feature map and the external flexible feature map, includes: Obtain a pre-calibrated affine transformation matrix, and use the affine transformation matrix to map the polarization-preserving reflection image and the depolarization-transmission image to the pixel coordinate system where the topological mask matrix is ​​located, thus completing spatial registration; The external rigidity feature map is extracted by performing a Hadamard product operation on the topological mask matrix and the registered polarization-preserving reflection image. A bitwise NOT operation is performed on the topological mask matrix to obtain an inverted mask matrix. The inverted mask matrix is ​​then subjected to a Hadamard product operation with the registered depolarized transmission image to extract the external flexible feature map.

4. The method according to claim 1, characterized in that, The extraction of transmitted light intensity distribution feature map from the depolarized transmission image based on the occlusion region coordinates of the topological mask matrix includes: Perform a connected component labeling algorithm on the topology mask matrix to extract target connected components with an area greater than a preset area threshold; Obtain the set of bounding box coordinates of the target connected component, and use the set of bounding box coordinates as the coordinates of the occluded region; Based on the coordinates of the occlusion region, the corresponding pixel sub-matrix is ​​extracted from the registered depolarized transmission image and used as the transmitted light intensity distribution feature map.

5. The method according to claim 1, characterized in that, Based on the topological mask matrix, the external flexible feature map is segmented into multiple local coordinate sub-panes, and the relative curvature gradient of the optical texture is calculated within each local coordinate sub-pane to obtain local differential topography features, including: Extract the set of edge contour points representing the rigid skeleton from the topological mask matrix; Using the set of edge contour points as a closed boundary, the external flexible feature map is divided into multiple independent local coordinate sub-panes; Within each of the local coordinate sub-panes, calculate the Euclidean distance from the pixel to the nearest set of edge contour points; Calculate the grayscale gradient magnitude of the pixel in the external flexible feature map; The ratio of the grayscale gradient magnitude to the Euclidean distance is used as the relative curvature gradient. When the relative curvature gradient exceeds a preset deformation tolerance threshold, the corresponding pixel region is extracted as the local differential morphology feature.

6. The method according to claim 1, characterized in that, The external rigidity feature map, the external flexible feature map, the transmitted light intensity distribution feature map, and the local differential topography feature map are input into a pre-trained multi-task visual detection model to output multi-dimensional defect judgment results, including: The external rigidity feature map, the external flexible feature map, the transmitted light intensity distribution feature map, and the local difference morphology feature are respectively converted into high-dimensional input tensors with a preset dimension. Each of the high-dimensional input tensors is input into the corresponding independent feature extraction backbone branch in the multi-task visual detection model, and deep semantic feature tensors are extracted through multi-layer two-dimensional convolution operations. Using the deep semantic feature tensor corresponding to the external rigid feature map as the query vector and the deep semantic feature tensor corresponding to the external flexible feature map as the key vector and value vector, perform cross-modal cross attention matrix multiplication to generate a rigid-flexible topological space attention weight matrix. The rigid-flexible topological space attention weight matrix and each of the deep semantic feature tensors are concatenated through channels to generate a global fusion feature tensor. The global fusion feature tensor is distributed to multiple independent task prediction heads, and the multi-dimensional defect determination results, including rigid defect bounding boxes, flexible defect bounding boxes, blind zone anomaly classification probabilities, and local form and position tolerance regression values, are output in parallel.

7. An intelligent visual defect detection system for a ton-barrel production line, characterized in that, The system is used to execute the intelligent visual defect detection method for the ton barrel production line according to any one of claims 1-6, including: The acquisition module is used to acquire multimodal orthogonal image data of the ton barrels to be tested on the production line under multidimensional physical field excitation. The multimodal orthogonal image data includes transient temperature field images, polarization-preserving reflection images, and depolarization-depolarizing transmission images. The first feature extraction module is used to extract features from the transient temperature field image and generate a topological mask matrix that characterizes the rigid skeleton of the target to be detected. The second feature extraction module is used to perform spatial registration and logical operations between the topological mask matrix and the polarization-preserving reflection image and the depolarization-transmission image, respectively, to decouple and extract the external rigid feature map and the external flexible feature map, and to extract the transmitted light intensity distribution feature map from the depolarization-transmission image based on the occlusion region coordinates of the topological mask matrix. The segmentation module is used to segment the external flexible feature map into multiple local coordinate sub-panes based on the topological mask matrix, and to calculate the relative curvature gradient of the optical texture in each local coordinate sub-pane to obtain local differential morphology features. The output module is used to input the external rigid feature map, the external flexible feature map, the transmitted light intensity distribution feature map, and the local differential morphology feature into the pre-trained multi-task visual detection model, and output multi-dimensional defect judgment results. The execution module is used to trigger corresponding operations on the ton barrel to be inspected according to preset rules based on the defect determination result.

8. The system according to claim 1, characterized in that, The first feature extraction module is specifically used for: Calculate the global thermal gradient histogram of the transient temperature field image; Based on the global thermal gradient histogram, an adaptive threshold algorithm is used to determine the target segmentation threshold; The transient temperature field image is binarized using the target segmentation threshold to generate a two-dimensional matrix composed of a first logical value and a second logical value. The two-dimensional matrix is ​​used as the topological mask matrix, wherein the first logical value represents a rigid skeleton region with high thermal conductivity.

9. The system according to claim 1, characterized in that, The segmentation module is specifically used for: Extract the set of edge contour points representing the rigid skeleton from the topological mask matrix; Using the set of edge contour points as a closed boundary, the external flexible feature map is divided into multiple independent local coordinate sub-panes; Within each of the local coordinate sub-panes, calculate the Euclidean distance from the pixel to the nearest set of edge contour points; Calculate the grayscale gradient magnitude of the pixel in the external flexible feature map; The ratio of the grayscale gradient magnitude to the Euclidean distance is used as the relative curvature gradient. When the relative curvature gradient exceeds a preset deformation tolerance threshold, the corresponding pixel region is extracted as the local differential morphology feature.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the intelligent visual defect detection method for the ton barrel production line according to any one of claims 1-6.