A visual inspection system for foreign objects in a can
Patent Information
- Application Number
- CN202610688808.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-19
- Publication Date
- 2026-08-18
AI Technical Summary
[0003]针对上述情况,为克服现有技术的缺陷,本发明提供了一种罐内异物视觉检测系统,针对镀锡铁罐的强镜面反射、铝塑复合罐的纹理干扰、塑料罐的颜色变异等材质特性差异,导致异物特征难以提取的问题;本方案集成环形LED明场光源、侧向暗场光源及偏振光源,根据开盖罐体材质(如金属罐、玻璃罐、塑料罐)动态切换照明模式,并使用旋转定位模块驱动开盖罐体绕中心轴匀速旋转,配合高速相机阵列从顶部俯视和侧向斜视角度同步采集图像,实现对罐壁、罐底、液面等全方位覆盖;针对传统基于CNN或Transformer的检测模型在处理高分辨率图像时,往往因计算资源限制而采用下采样策略,导致细节信息丢失,无法有效识别微小异常的问题,本方案采用将异常检测转化为视觉标记预测任务,使用预训练的DINO模型提取多层级视觉token(视觉令牌),通过多方向扫描函数(MDS)将二维特征图转换为四个方向的一维token序列,采用轻量化的Mamba自回归模型对所有序列进行统一建模,有效捕获罐内空间结构的全局依赖与局部细节,同时引入动态异常阈值机制,结合正常样本的统计特性自动调整判别边界,不仅显著缓解了源域(通用预训练数据)与目标域(奶粉罐实际场景)之间的分布差异,还实现了“一次训练、多场景适用”的泛化能力,该方法在大幅降低计算负担的同时,有效避免了因图像降采样导致的微小异物漏检的问题
[0034] (1) In response to the problem that the differences in material properties, such as strong specular reflection of tin-plated iron cans, texture interference of aluminum-plastic composite cans, and color variation of plastic cans, make it difficult to extract foreign object features, this solution integrates ring LED bright field light source, side dark field light source and polarized light source. The lighting mode is dynamically switched according to the material of the can (such as metal can, glass can, plastic can), and the rotating positioning module is used to drive the can to rotate at a constant speed around the central axis. With the help of a high-speed camera array, images are collected simultaneously from the top view and the side oblique view angle to achieve all-round coverage of the can wall, can bottom, liquid surface, etc.
Smart Images

Figure CN122591688A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial automation inspection technology, specifically to a visual inspection system for foreign objects inside a tank. Background Technology
[0002] In the filling and production of highly sensitive foods such as milk powder cans, foreign object contamination (such as metal fragments, plastic particles, hair, fibers, insect remains, etc.) is a quality risk that runs through the entire chain of "people, machines, materials, methods, and environment." Its sources include raw material introduction, equipment wear and tear, packaging material defects, and human operational negligence. Differences in material characteristics, such as the strong specular reflection of tin-plated iron cans, the texture interference of aluminum-plastic composite cans, and color variations of plastic cans, coupled with the fact that the deep cavity structure of the can easily form visual blind spots, make it difficult to extract foreign object features, further increasing the risk of missed detection. Existing visual detection models based on CNN or Transformer are often forced to downsample due to computational resource constraints, and their deployment costs are high and production line switching is cumbersome, resulting in serious loss of texture and edge features of micron-sized foreign objects such as hair and fibers, leading to missed detection or misjudgment. Summary of the Invention
[0003] To address the aforementioned issues and overcome the shortcomings of existing technologies, this invention provides a visual detection system for foreign objects inside cans. It addresses the problem of difficulty in extracting foreign object features due to differences in material properties, such as the strong specular reflection of tin-plated iron cans, the texture interference of aluminum-plastic composite cans, and color variations in plastic cans. This solution integrates a ring-shaped LED bright-field light source, a side-mounted dark-field light source, and a polarized light source. The illumination mode dynamically switches according to the material of the can (e.g., metal, glass, plastic). A rotation positioning module drives the can to rotate uniformly around its central axis. A high-speed camera array simultaneously acquires images from top-down and side-mounted angles, achieving comprehensive coverage of the can walls, bottom, and liquid surface. Furthermore, traditional CNN or Transformer-based detection models often employ downsampling strategies when processing high-resolution images due to computational resource limitations, resulting in the loss of detailed information and hindering effective detection. To address the problem of effectively identifying minute anomalies, this solution transforms anomaly detection into a visual tag prediction task. A pre-trained DINO model is used to extract multi-level visual tokens. A multi-directional scan function (MDS) is then used to convert the two-dimensional feature map into a one-dimensional token sequence in four directions. A lightweight Mamba autoregressive model is employed to uniformly model all sequences, effectively capturing the global dependencies and local details of the can's internal spatial structure. Simultaneously, a dynamic anomaly threshold mechanism is introduced, automatically adjusting the discrimination boundary based on the statistical characteristics of normal samples. This not only significantly alleviates the distribution differences between the source domain (general pre-training data) and the target domain (the actual scenario of a milk powder can), but also achieves the generalization capability of "one-time training, applicable to multiple scenarios." This method significantly reduces the computational burden while effectively avoiding the problem of missing minute foreign objects due to image downsampling.
[0004] The present invention provides a visual detection system for foreign objects inside a tank, the system comprising a rotation positioning module, a multi-source illumination module, a multimodal imaging module, an image processing decision module, and a visualization interface module;
[0005] The rotary positioning module uses a conveyor belt to transport the open cans to be inspected one by one to the inspection station and accurately position them. It uses a servo motor and a clamping mechanism to fix the open cans to be inspected, so that the open cans rotate around their central axis.
[0006] The multi-source lighting module integrates a ring-shaped LED bright field light source, a side dark field light source, and a polarized light source, and switches the lighting mode according to the material of the open can.
[0007] The multimodal imaging module uses a high-speed industrial camera array to simultaneously acquire images from top-down and side-oblique angles, and combines them with the light source of the multi-source illumination module to obtain bright-field images, dark-field images, and polarized images.
[0008] The image processing decision module constructs a lightweight anomaly detection model, analyzes and processes the acquired bright-field images, dark-field images, and polarization images to obtain anomaly detection results and generate decision instructions.
[0009] The visualization interface module provides a visualized image of the inside of the tank. The foreign object area is displayed as a semi-transparent color heat map overlaid, and a bounding box is automatically drawn for each detected foreign object, with the foreign object type, size and confidence level labeled.
[0010] Furthermore, the image processing decision module constructs a lightweight anomaly detection model to analyze and process the acquired bright-field images, dark-field images, and polarization images, specifically including the following steps:
[0011] Step S1: Image preprocessing. High dynamic range (HDR) fusion is performed on the acquired bright field image, dark field image and polarization image. The uniform illumination correction algorithm and image registration technology are used to correct and spatially align the image to obtain the HDR composite image.
[0012] Step S2: Visual token sequence extraction. Use the pre-trained DINO model to extract features from the input HDR composite image. Extract multi-level tokens from the 4th, 8th and 12th layers of the pre-trained DINO model. Each layer has M feature channels, forming a multi-level visual token sequence.
[0013] Step S3: Multi-directional token sequence generation. MDS is used to convert the 2D visual token sequence into a 1D sequence. Four token sequences are generated for each level of visual token sequence in four directions. The token sequences are then adjusted through residual connections to obtain the initial token sequence.
[0014] Step S4: Regression modeling and anomaly detection. The Mamba model is used for autoregressive modeling to predict and score anomalies in the token sequence, resulting in predicted token sequences and anomaly scores.
[0015] Step S5: Multi-level information fusion decision-making, transforming the predicted token sequence into anomaly graphs, and fusing anomaly graphs at different levels to generate anomaly location maps, indicating the location and size of foreign objects;
[0016] Step S6: Anomaly Decision Making. Extract multi-dimensional features of the anomaly region from the anomaly localization map, use a classification model to identify the type of foreign object, and perform multi-level decision making to obtain the detection results.
[0017] Step S7: Optimize feedback, set the iteration threshold, repeat steps S2 to S6 until the iteration threshold is reached, generate the trained lightweight anomaly detection model, and feed the real-time detection results back to the lightweight anomaly detection model for continuous optimization.
[0018] Furthermore, in step S4, regression modeling and anomaly detection specifically include the following steps:
[0019] Step S41: Mamba Autoregressive Model. Mamba is used as an autoregressive model to replace the traditional Transformer. Global information is captured with linear complexity through State Space Model (SSM). An independent Mamba model is trained for each token sequence to capture information from different directions.
[0020] Step S42: Anomaly prediction. Set a prediction step size L for each token sequence, that is, exclude the most recent L tokens from prediction. Add [BOS] and [EOS] tags to each token sequence. Generate the predicted token sequence through the Mamba model.
[0021] Step S43: Anomaly score calculation. Calculate the Euclidean distance between the predicted token sequence and the initial token sequence to obtain the anomaly score. The formula used is as follows: ;
[0022] In the formula, The index of the token sequence. For the initial token sequence, To predict the token sequence, The square of the Euclidean distance. Indicates abnormal rating;
[0023] Furthermore, in step S5, the multi-level information fusion decision-making specifically includes the following steps:
[0024] Step S51: Multi-source information fusion, average the predicted token sequence in all directions, and restore the averaged predicted token sequence to a 2D image through inverse MDS to generate anomaly maps for each level.
[0025] Step S52: Multi-level anomaly map fusion. Anomaly maps of different levels are weighted and fused. Local detail features are extracted from low-level anomaly maps, and global context features are extracted from high-level anomaly maps. The fusion ratio is dynamically adjusted by weighting coefficients to obtain a fused anomaly map.
[0026] Step S53: Adaptive thresholding. Based on the statistical characteristics of the anomaly map, the anomaly threshold is dynamically calculated. Combining the complexity inside the tank and the background, the threshold is automatically adjusted to generate an anomaly location map, indicating the location and size of the foreign object. The formula for calculating the anomaly threshold is as follows: ;
[0027] In the formula, This is the abnormal threshold. Indicates a normal sample index. and This represents the mean and standard deviation of the anomaly confidence level for the current normal samples. Sensitivity modulator;
[0028] Furthermore, in step S6, the anomaly decision-making specifically includes the following steps:
[0029] Step S61: Foreign object feature extraction. Extract multi-dimensional features from the abnormal region of the abnormal location map. The multi-dimensional features include shape, texture, motion trajectory and spectral characteristics. Use a pre-trained classification model to identify the type of foreign object and obtain the classification result.
[0030] Step S62: Anomaly Confidence. The reasonableness confidence is calculated based on multi-dimensional features. The anomaly score and reasonableness confidence are then weighted and fused to obtain the final anomaly confidence. The formula for reasonableness confidence is as follows: ;
[0031] In the formula, , and These are the weighting coefficients. Indicates semantic consistency score, This indicates the score for motion stability. Indicates the score for texture anomalies. The calculated confidence level of reasonableness;
[0032] Step S63: Comprehensive decision-making. Based on the anomaly threshold and combined with the anomaly confidence, classification results and foreign object location, a detection result is generated. A multi-level decision-making mechanism is used to make a decision and generate a decision instruction.
[0033] The beneficial effects achieved by the present invention using the above solution are as follows:
[0034] (1) In response to the problem that the differences in material properties, such as strong specular reflection of tin-plated iron cans, texture interference of aluminum-plastic composite cans, and color variation of plastic cans, make it difficult to extract foreign object features, this solution integrates ring LED bright field light source, side dark field light source and polarized light source. The lighting mode is dynamically switched according to the material of the can (such as metal can, glass can, plastic can), and the rotating positioning module is used to drive the can to rotate at a constant speed around the central axis. With the help of a high-speed camera array, images are collected simultaneously from the top view and the side oblique view angle to achieve all-round coverage of the can wall, can bottom, liquid surface, etc.
[0035] (2) In view of the problem that traditional detection models based on CNN or Transformer often adopt downsampling strategies due to computational resource limitations when processing high-resolution images, resulting in the loss of detailed information and the inability to effectively identify small anomalies, this solution transforms anomaly detection into a visual label prediction task. It uses a pre-trained DINO model to extract multi-level visual tokens, and converts the two-dimensional feature map into a one-dimensional token sequence in four directions through a multi-directional scan function (MDS). It uses a lightweight Mamba autoregressive model to uniformly model all sequences, effectively capturing the global dependence and local details of the spatial structure inside the tank. At the same time, a dynamic anomaly threshold mechanism is introduced, which automatically adjusts the discrimination boundary in combination with the statistical characteristics of normal samples. This not only significantly alleviates the distribution difference between the source domain (general pre-training data) and the target domain (actual scene of the open tank), but also achieves the generalization ability of "one training, multiple scenarios". This method significantly reduces the computational burden while effectively avoiding the missed detection of small foreign objects caused by image downsampling, and significantly improves the robustness of the system to irregular and low-contrast foreign objects. Attached Figure Description
[0036] Figure 1 This is a schematic diagram of a visual detection system for foreign objects inside a can, as proposed in this invention.
[0037] Figure 2 This is a flowchart illustrating the image processing decision module.
[0038] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof. Detailed Implementation
[0039] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0040] Example 1, see Figure 1 The present invention provides a visual detection system for foreign objects inside a tank, which includes a rotation positioning module, a multi-source illumination module, a multimodal imaging module, an image processing decision module, and a visualization interface module;
[0041] The rotary positioning module uses a conveyor belt to transport the open cans to be inspected one by one to the inspection station and accurately position them. It uses a servo motor and a clamping mechanism to fix the open cans to be inspected, so that the open cans rotate around their central axis.
[0042] The multi-source lighting module integrates a ring-shaped LED bright field light source, a side dark field light source, and a polarized light source. It switches lighting modes according to the material of the open can to suppress specular reflection and enhance diffuse reflection characteristics from foreign objects. The lighting modes corresponding to different can materials are as follows:
[0043] For metal tanks (stainless steel / aluminum alloy), polarized light sources should be used preferentially to eliminate mirror reflections on the tank walls;
[0044] Highly reflective plastic containers: Use side-lit, dark-field light sources first, highlighting the edges of foreign objects in the bright areas;
[0045] For glass / transparent tanks, use ring-shaped bright field LEDs to provide uniform lighting and clearly show internal impurities;
[0046] Matte, non-reflective can body: Direct use of ring-shaped bright field, sufficient brightness and clear features;
[0047] The multimodal imaging module uses a high-speed industrial camera array to simultaneously acquire images from top-down and side-oblique angles. Combined with the light source of the multi-source illumination module, it obtains bright-field images, dark-field images, and polarized images. The image resolution is set to 1024×1024.
[0048] The image processing decision module constructs a lightweight anomaly detection model, analyzes and processes the acquired bright-field images, dark-field images, and polarization images to obtain anomaly detection results and generate decision instructions.
[0049] The visualization interface module provides a visualized image of the inside of the tank. The foreign object area is displayed as a semi-transparent color heat map overlaid, and a bounding box is automatically drawn for each detected foreign object, with the foreign object type, size and confidence level labeled.
[0050] Example 2: Based on the above examples, in the multimodal imaging module, the camera for acquiring bright-field images is positioned above the open tank to capture images of the tank wall and specular reflections, thus obtaining bright-field images; the camera for acquiring dark-field images is positioned diagonally below the bottom of the open tank, in conjunction with a lateral dark-field light source, to capture the diffuse reflection characteristics of the tank bottom and small foreign objects; the device for acquiring polarized images includes a polarizer placed in front of the light source and an analyzer placed in front of the camera lens, with cross-polarization configuration capable of filtering out more than 90% of liquid surface glare.
[0051] By performing the aforementioned operations, this solution addresses the problem of difficulty in extracting foreign object features due to differences in material properties such as strong specular reflection in tin-plated iron cans, texture interference in aluminum-plastic composite cans, and color variations in plastic cans. It integrates a ring-shaped LED bright field light source, a side dark field light source, and a polarized light source. The illumination mode is dynamically switched according to the material of the can (such as metal cans, glass cans, and plastic cans). A rotation positioning module drives the can to rotate uniformly around the central axis. In conjunction with a high-speed camera array, images are simultaneously acquired from top-down and side-oblique angles, achieving comprehensive coverage of the can wall, bottom, and liquid surface.
[0052] Example 3, see Figure 2 This embodiment is based on the above embodiment. The image processing decision module constructs a lightweight anomaly detection model to analyze and process the acquired bright-field images, dark-field images, and polarization images. Specifically, it includes the following steps:
[0053] Step S1: Image preprocessing. High dynamic range (HDR) fusion is performed on the acquired bright field image, dark field image and polarization image. The uniform illumination correction algorithm and image registration technology are used to correct and spatially align the image to obtain the HDR composite image.
[0054] Step S2: Visual token sequence extraction. Use the pre-trained DINO model to extract features from the input HDR composite image. Extract multi-level tokens from the 4th, 8th and 12th layers of the pre-trained DINO model. Each layer has M feature channels, forming a multi-level visual token sequence.
[0055] Step S3: Multi-directional token sequence generation. MDS is used to convert the 2D visual token sequence into a 1D sequence. Four token sequences are generated for each level of visual token sequence in four directions. The token sequences are then adjusted through residual connections to obtain the initial token sequence.
[0056] Step S4: Regression modeling and anomaly detection. The Mamba model is used for autoregressive modeling to predict and score anomalies in the token sequence, resulting in predicted token sequences and anomaly scores.
[0057] Step S5: Multi-level information fusion decision-making, transforming the predicted token sequence into anomaly graphs, and fusing anomaly graphs at different levels to generate anomaly location maps, indicating the location and size of foreign objects;
[0058] Step S6: Anomaly Decision Making. Extract multi-dimensional features of the anomaly region from the anomaly localization map, use a classification model to identify the type of foreign object, and perform multi-level decision making to obtain the detection results.
[0059] Step S7: Optimize feedback, set the iteration threshold, repeat steps S2 to S6 until the iteration threshold is reached, generate a trained lightweight anomaly detection model, and feed the real-time detection results back to the lightweight anomaly detection model for continuous optimization; for false positives and false negatives, automatically add them to the training dataset used for training, and gradually improve the model's ability to identify new types of foreign objects through an online learning mechanism.
[0060] Example 4, based on the above examples, includes the following steps in step S4: regression modeling and anomaly detection.
[0061] Step S41: Mamba Autoregressive Model. Mamba is used as an autoregressive model to replace the traditional Transformer. Global information is captured with linear complexity through SSM. An independent Mamba model is trained for each token sequence to capture information from different directions. The SSM is a selective state-space model, which is the core of the Mamba algorithm and is based on the discretization of continuous-time SSM.
[0062] Step S42: Anomaly prediction. Set a prediction step size L for each token sequence, that is, exclude the most recent L tokens from prediction. Add [BOS] and [EOS] tags to each token sequence. Generate the predicted token sequence through the Mamba model.
[0063] Step S43: Anomaly score calculation. Calculate the Euclidean distance between the predicted token sequence and the initial token sequence to obtain the anomaly score. The formula used is as follows: ;
[0064] In the formula, The index of the token sequence. For the initial token sequence, To predict the token sequence, The square of the Euclidean distance. This indicates an abnormal score.
[0065] Example 5, based on the above examples, includes the following steps in step S5, where multi-level information fusion decision-making:
[0066] Step S51: Multi-source information fusion, average the predicted token sequence in all directions, and restore the averaged predicted token sequence to a 2D image through inverse MDS to generate anomaly maps for each level.
[0067] Step S52: Multi-level anomaly map fusion. The anomaly maps of the 4th, 8th and 12th layers are weighted and fused. Local detail features are extracted from the low-level anomaly map and global context features are extracted from the high-level anomaly map. The fusion ratio is dynamically adjusted by the weighting coefficient to obtain the fused anomaly map.
[0068] Step S53: Adaptive thresholding. Based on the statistical characteristics of the anomaly map, the anomaly threshold is dynamically calculated. Combining the complexity inside the tank and the background, the threshold is automatically adjusted to generate an anomaly location map, indicating the location and size of the foreign object. The formula for calculating the anomaly threshold is as follows: ;
[0069] In the formula, This is the abnormal threshold. Indicates a normal sample index. and This represents the mean and standard deviation of the anomaly confidence level for the current normal samples. It is a sensitivity modulator.
[0070] Example 6, based on the above examples, includes the following steps in step S6 for anomaly decision-making:
[0071] Step S61: Foreign object feature extraction. Extract multi-dimensional features from the abnormal region of the abnormal location map. The multi-dimensional features include shape, texture, motion trajectory and spectral characteristics. Use a pre-trained classification model to identify the type of foreign object and obtain the classification result.
[0072] Step S62: Anomaly Confidence. The reasonableness confidence is calculated based on multi-dimensional features. The anomaly score and reasonableness confidence are then weighted and fused to obtain the final anomaly confidence. The formula for reasonableness confidence is as follows: ;
[0073] In the formula, , and These are the weighting coefficients. Indicates semantic consistency score, This indicates the score for motion stability. Indicates the score for texture anomalies. The calculated confidence level of reasonableness;
[0074] Step S63: Comprehensive decision-making. Based on the anomaly threshold and combined with the anomaly confidence level, classification results, and foreign object location, a detection result is generated. A multi-level decision-making mechanism is used to make a decision and generate decision instructions. The decision instructions include manual re-inspection for low confidence, marking for medium confidence, and direct rejection for high confidence.
[0075] By performing the aforementioned operations, this solution addresses the problem that traditional CNN- or Transformer-based detection models often employ downsampling strategies due to computational resource limitations when processing high-resolution images, leading to the loss of detailed information and an inability to effectively identify minute anomalies. Instead, it transforms anomaly detection into a visual tag prediction task. A pre-trained DINO model is used to extract multi-level visual tokens, and a multi-directional scan function (MDS) is used to convert the two-dimensional feature map into a one-dimensional token sequence in four directions. A lightweight Mamba autoregressive model is then used to uniformly model all sequences, effectively capturing the global dependencies and local details of the internal spatial structure of the tank. Simultaneously, a dynamic anomaly threshold mechanism is introduced, automatically adjusting the discrimination boundary based on the statistical characteristics of normal samples. This not only significantly alleviates the distribution differences between the source domain (general pre-training data) and the target domain (the actual scenario of an open tank), but also achieves the generalization capability of "one-time training, applicable to multiple scenarios." This method significantly reduces the computational burden while effectively avoiding the missed detection of minute foreign objects caused by image downsampling, significantly improving the system's robustness to irregular, low-contrast foreign objects.
[0076] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0077] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
[0078] The present invention and its embodiments have been described above. This description is not restrictive, and the accompanying drawings are only one embodiment of the present invention; the actual structure is not limited thereto. In conclusion, if those skilled in the art are inspired by this description and design similar structures and embodiments without departing from the spirit of the invention, such designs should fall within the protection scope of the present invention.
Claims
1. A visual inspection system for foreign objects inside a tank, characterized in that: The system includes a rotation positioning module, a multi-source illumination module, a multimodal imaging module, an image processing decision module, and a visualization interface module; The rotary positioning module uses a conveyor belt to transport the open cans to be inspected one by one to the inspection station and accurately position them. It uses a servo motor and a clamping mechanism to fix the open cans to be inspected, so that the open cans rotate around their central axis. The multi-source lighting module integrates a ring-shaped LED bright field light source, a side dark field light source, and a polarized light source, and switches the lighting mode according to the material of the open can. The multimodal imaging module uses a high-speed industrial camera array to simultaneously acquire images from top-down and side-oblique angles, and combines them with the light source of the multi-source illumination module to obtain bright-field images, dark-field images, and polarized images. The image processing decision module constructs a lightweight anomaly detection model, analyzes and processes the acquired bright-field images, dark-field images, and polarization images to obtain anomaly detection results and generate decision instructions. The visualization interface module provides a visualized image of the inside of the tank. The foreign object area is displayed as a semi-transparent color heat map overlaid, and a bounding box is automatically drawn for each detected foreign object, with the foreign object type, size and confidence level labeled.
2. The visual detection system for foreign objects inside a tank according to claim 1, characterized in that: The image processing decision module constructs a lightweight anomaly detection model to analyze and process the acquired bright-field images, dark-field images, and polarization images, specifically including the following steps: Step S1: Image preprocessing, high dynamic range fusion of the acquired bright field image, dark field image and polarization image, and correction and spatial alignment of the image using uniform illumination correction algorithm and image registration technology to obtain HDR composite image; Step S2: Visual token sequence extraction. Use the pre-trained DINO model to extract features from the input HDR composite image. Extract multi-level tokens from the 4th, 8th and 12th layers of the pre-trained DINO model. Each layer has M feature channels, forming a multi-level visual token sequence. Step S3: Multi-directional token sequence generation. MDS is used to convert the 2D visual token sequence into a 1D sequence. Four token sequences are generated for each level of visual token sequence in four directions. The token sequences are then adjusted through residual connections to obtain the initial token sequence. Step S4: Regression modeling and anomaly detection. The Mamba model is used for autoregressive modeling to predict and score anomalies in the token sequence, resulting in predicted token sequences and anomaly scores. Step S5: Multi-level information fusion decision-making, transforming the predicted token sequence into anomaly graphs, and fusing anomaly graphs at different levels to generate anomaly location maps, indicating the location and size of foreign objects; Step S6: Anomaly Decision Making. Extract multi-dimensional features of the anomaly region from the anomaly localization map, use a classification model to identify the type of foreign object, and perform multi-level decision making to obtain the detection results. Step S7: Optimize feedback, set the iteration threshold, repeat steps S2 to S6 until the iteration threshold is reached, generate the trained lightweight anomaly detection model, and feed the real-time detection results back to the lightweight anomaly detection model for continuous optimization.
3. The visual detection system for foreign objects inside a tank according to claim 2, characterized in that: In step S4, regression modeling and anomaly detection specifically include the following steps: Step S41: Mamba Autoregressive Model. Mamba is used as the autoregressive model. Global information is captured with linear complexity through SSM. An independent Mamba model is trained for each token sequence to capture information from different directions. Step S42: Anomaly prediction. Set a prediction step size L for each token sequence, that is, exclude the most recent L tokens from prediction. Add [BOS] and [EOS] tags to each token sequence. Generate the predicted token sequence through the Mamba model. Step S43: Anomaly score calculation. Calculate the Euclidean distance between the predicted token sequence and the initial token sequence to obtain the anomaly score.
4. The visual detection system for foreign objects inside a tank according to claim 2, characterized in that: In step S5, the multi-level information fusion decision-making specifically includes the following steps: Step S51: Multi-source information fusion, average the predicted token sequence in all directions, and restore the averaged predicted token sequence to a 2D image through inverse MDS to generate anomaly maps for each level. Step S52: Multi-level anomaly map fusion. Anomaly maps of different levels are weighted and fused. Local detail features are extracted from low-level anomaly maps, and global context features are extracted from high-level anomaly maps. The fusion ratio is dynamically adjusted by weighting coefficients to obtain a fused anomaly map. Step S53: Adaptive thresholding. Based on the statistical characteristics of the anomaly map, the anomaly threshold is dynamically calculated. Combining the complexity inside the tank and the background, the threshold is automatically adjusted to generate an anomaly location map, indicating the location and size of the foreign object. The formula for calculating the anomaly threshold is as follows: ; In the formula, This is the abnormal threshold. Indicates a normal sample index. and This represents the mean and standard deviation of the anomaly confidence level for the current normal samples. It is a sensitivity modulator.