An AI vision-based online detection method and system for defects in spunlace filter material

By adjusting the polarization relationship between polarized light and the fiber texture of spunlace filter media, and combining multi-source detection modules and deep learning technology, the problems of low detection efficiency and insufficient accuracy of traditional spunlace filter media are solved, achieving full-coverage accurate identification and stable detection on high-speed continuous production lines.

CN122448848APending Publication Date: 2026-07-24HUAJI ENVIRONMENTAL PROTECTION (WUHAN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAJI ENVIRONMENTAL PROTECTION (WUHAN) CO LTD
Filing Date
2026-04-16
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Traditional spunlace filter media defect detection is inefficient, lacks accuracy and stability, and is difficult to adapt to high-speed continuous production. Existing machine vision inspection methods are easily affected by fiber texture interference and light fluctuations, cannot identify minute defects, and lack multimodal feature fusion mechanisms.

Method used

By adjusting the polarization relationship between polarized light and the fiber texture of spunlace filter media, and combining a multi-source detection module and dynamic light source parameters, deep learning texture stripping and feature enhancement techniques are used to achieve separation of texture and defects and feature fusion. A dynamic deblurring algorithm eliminates motion blur, and a multi-modal feature fusion model is established for comprehensive judgment.

Benefits of technology

It significantly improves the detection rate and stability of minute defects, achieves accurate identification with full coverage from surface to interior, adapts to the detection needs of different materials and batches, and supports real-time detection on high-speed continuous production lines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122448848A_ABST
    Figure CN122448848A_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of online detection of water jet filter material defects, and particularly relates to an online detection method and system for water jet filter material defects based on AI vision, which adjusts the polarization relationship between polarized light and the fiber texture of water jet filter material, uses the difference in optical refractive index between defects and fibers to make the defects form specific optical signals, preliminarily separates the texture and defect characteristics, adjusts the optical parameters of the near-infrared structured light source according to the light transmission characteristics of the filter material, and forms a uniform active light field; the imaging parameters are dynamically adjusted according to the running state of the production line, the spectral data, internal structure related data, material distribution related data and functional index data of the filter material are respectively collected through a multi-source detection module, and through a synchronous control mechanism and a production line control unit linkage, the data of the same detection area collected by each module is realized space-time registration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of online defect detection technology for spunlace filter media, and particularly relates to an online defect detection method and system for spunlace filter media based on AI vision. Background Technology

[0002] Traditional defect detection in spunlace filter media relies primarily on manual visual inspection, which suffers from low efficiency, high labor intensity, and inconsistent judgment standards, making it difficult to meet the online inspection needs of high-speed continuous production. Conventional machine vision inspection methods are easily affected by factors such as the fiber texture of the filter media itself, fluctuations in production line lighting, and motion blur. They can only identify obvious surface defects and have a low detection rate for hidden defects such as micro-pinholes, fiber agglomerations, and internal delamination. Furthermore, they cannot simultaneously acquire material distribution, internal structure, and functional indicators, resulting in a single detection dimension and insufficient accuracy and stability.

[0003] Existing detection technologies that integrate vision and spectroscopy mostly use fixed polarization angles and light source parameters, which cannot be dynamically adjusted according to the filter material, light transmission characteristics, and production line operation status. Multi-source detection modules lack a unified spatiotemporal benchmark, making it difficult to achieve accurate data registration. Deep learning texture stripping and deblurring algorithms have not been lightweighted and dynamically adapted to the characteristics of spunlace filter materials, which can easily lead to the loss of information on minor defects. At the same time, there is a lack of deep fusion and correlation judgment mechanisms for multimodal features of surface, internal, material, and functional aspects. The model has weak generalization ability and cannot adapt to changes in different materials and batches of filter materials, making it difficult to achieve comprehensive and accurate integrated judgment of defects. Summary of the Invention

[0004] In view of the aforementioned problems, and in conjunction with the first aspect of the present invention, embodiments of the present invention provide an online defect detection method for spunlace filter media based on AI vision, the method comprising: By adjusting the polarization relationship between polarized light and the fiber texture of spunlace filter media, the difference in optical refractive index between defects and fibers is used to make defects form specific optical signals, thus initially separating texture and defect features. At the same time, based on the light transmission characteristics of the filter media, the optical parameters of the near-infrared structured light source are adjusted to form a uniform active light field. The imaging parameters are dynamically adjusted according to the production line operation status. The multi-source detection module collects the spectral data, internal structure data, material distribution data and functional index data of the filter material. Through the synchronous control mechanism, it links with the production line control unit to achieve spatiotemporal registration of the data collected by each module in the same detection area. A deep learning-based texture stripping method is adopted, which combines the texture distribution law of filter material to generate an adaptive texture template. The background texture is stripped by dynamically adjusted texture separation logic while retaining the information of minute defects. Then, feature enhancement technology is used to enhance the edge and gray-scale difference of minute defects. Feature spectral information corresponding to filter material and internal defects is selected, and spectral data and spatial image details are fused to generate an enhanced image. A dynamic deblurring algorithm is used to eliminate image blur caused by production line movement, and edge sharpening technology is used to restore the true boundary of defects. Surface defect features, internal structural features, material distribution features, and functional features are input into a multimodal feature fusion model. A mapping relationship between each feature is established through a cross-modal correlation mechanism. Based on the preset multi-dimensional defect judgment rules, a comprehensive judgment of surface defects, internal defects, and functional defects of the filter material is completed. The defect judgment results are uploaded to the production management system. If an unqualified defect is detected, a marking or automatic rejection action is triggered. At the same time, the defect data is correlated with the production process parameters for analysis.

[0005] Furthermore, embodiments of the present invention also provide an online defect detection system for spunlace filter media based on AI vision, characterized in that it includes: A processor; a machine-readable storage medium for storing machine-executable instructions of the processor; wherein the processor is configured to perform the above-described AI vision-based online defect detection method for spunlace filter media by executing the machine-executable instructions.

[0006] In another aspect, embodiments of the present invention also provide a computer program product, the computer program product including machine-executable instructions, the machine-executable instructions being stored in a computer-readable storage medium, a processor of a computer device reading the machine-executable instructions from the computer-readable storage medium, the processor executing the machine-executable instructions, causing the computer device to execute the above-described AI vision-based online defect detection method for spunlace filter media.

[0007] Based on the above, by dynamically adjusting the polarization angle and light source parameters, and with the spatiotemporal precise registration of the multi-source detection module, the detection errors caused by filter material fiber texture interference, production line illumination fluctuations, and motion blur are effectively eliminated. This significantly improves the detection rate and detection stability of hidden defects such as micro-pinholes, fiber agglomeration, and internal delamination, achieving full-coverage accurate identification from surface defects to internal defects. The overall detection accuracy and anti-interference ability are far superior to traditional visual inspection methods.

[0008] This invention employs a lightweight deep learning algorithm tailored to the characteristics of spunlace filter media, combined with a multimodal feature deep fusion and correlation judgment mechanism. While fully preserving information on minute defects, it significantly improves the model's real-time performance and generalization ability. It can quickly adapt to the testing requirements of filter media of different materials and batches, enabling full-process online real-time testing on high-speed continuous production lines. The overall testing efficiency and automation level are significantly improved, providing stable and reliable technical support for the high-quality continuous production of spunlace filter media. Attached Figure Description

[0009] Figure 1 This is a schematic diagram of the execution flow of the online defect detection method for spunlace filter media based on AI vision provided in an embodiment of the present invention.

[0010] Figure 2 This is a schematic diagram of exemplary hardware and software components of the AI ​​vision-based online defect detection system for spunlace filter media provided in an embodiment of the present invention. Detailed Implementation

[0011] The present invention will now be described in detail with reference to the accompanying drawings. Figure 1 This is a flowchart illustrating an online defect detection method for spunlace filter media based on AI vision, according to an embodiment of the present invention. The following is a detailed description of this online defect detection method for spunlace filter media based on AI vision.

[0012] Step S110: Adjust the polarization relationship between polarized light and the fiber texture of the spunlace filter material. Utilize the difference in optical refractive index between defects and fibers to make defects form specific optical signals, initially separating texture and defect features. At the same time, adjust the optical parameters of the near-infrared structured light source according to the light transmission characteristics of the filter material to form a uniform active light field. This is achieved through a dual polarizer adjustment system working in conjunction with a near-infrared structured light source. The dual polarizer adjustment system includes a polarizer, an analyzer, and a stepper motor drive module. The polarizer is fixed at the light source's emission end, and the analyzer is mounted at the front of the camera lens. Both use high extinction ratio polarizers (extinction ratio ≥1000:1). Utilizing the difference in optical refractive index between defects and spunlace filter fibers (fiber refractive index 1.5-1.6, defects such as pinholes and impurities refractive index 1.0-1.3), the polarization relationship between the two is adjusted to make the defects form specific optical signals, thus initially separating the texture from the defects. Transmittance characteristic detection uses a transmittance photoelectric sensor (model E3Z-T61) that emits detection light in the same wavelength band (850nm) as the near-infrared structured light source, which penetrates vertically through the filter material. Combined with thickness data collected by a laser displacement sensor (accuracy ±0.01mm), the actual transmittance is calculated according to "transmittance = transmitted light intensity / incident light intensity × 100%". The near-infrared structured light source uses a dot matrix LED light source. The light spot density (10-30 spots / cm²), power (10-50W), and projection angle (15°-45°) are dynamically adjusted according to the light transmittance data. Zoned adjustment is initiated for areas with differences in light transmittance uniformity to ensure the formation of a uniform active light field and to offset the fluctuations in ambient light in the production line (fluctuation amplitude ≤ ±10%).

[0013] Step S111: Fix the polarizer to the light source emission end, install the analyzer at the front of the camera lens, start the dual polarizer adjustment system, and drive the analyzer to rotate around the optical axis by a stepper motor. The polarizer is fixed to the emission port of the near-infrared structured light source by a metal bracket. The bracket and the light source housing are fastened with bolts to ensure that the coaxiality deviation between the polarizer and the light source optical axis is ≤0.1mm. The analyzer is mounted on a rotatable bracket at the front of the camera lens. The bracket has a gear ring on its inner side, which meshes with the output gear of the stepper motor (model 42HS4013, step angle 1.8°), with a meshing clearance ≤0.02mm. After the dual polarizer adjustment system is started, the central control unit (model PLC S7-1200) sends pulse signals to the stepper motor through the driver, driving the analyzer to rotate around the optical axis with a rotation accuracy of ±0.05°. During the motor rotation, the encoder provides real-time feedback on the actual position of the analyzer, forming a closed-loop position control to avoid cumulative errors. The initial relative angle between the polarizer and the analyzer is set to 0°. After the system starts up and stabilizes (stabilization time ≤3s), the subsequent angle adjustment process begins to ensure the accuracy and stability of polarization adjustment, laying the equipment foundation for the separation of texture and defect features.

[0014] Step S112: Using the pre-calibrated main polarization direction of the filter fiber as a reference, gradually adjust the relative angle between the polarizer and the analyzer, and simultaneously acquire filter images at different angles. The optimal polarization angle is selected through the image grayscale correlation algorithm, so that the polarization reflection signal of the fiber texture is suppressed, and the defects form specific optical signals due to the difference in optical refractive index with the fiber, thus achieving the initial separation of the optical features of the texture and defects. In the pre-calibration stage, texture images of defect-free spunlace filter media are acquired using a line scan camera (model MV-CH080-10GM). The Histogram of Oriented Gradients (HOG) algorithm is used to analyze the main polarization direction of the fibers, and the angle corresponding to the peak value of the histogram is used as the baseline value, with a calibration error ≤0.5°. Starting from this baseline value, the central control unit controls a stepper motor to drive the analyzer to gradually adjust the relative angle between the polarizer and the analyzer, with an adjustment step size of 0.5° and an adjustment range of 0°-90°. For each adjustment step, the camera simultaneously acquires images of the filter media (frame rate 50fps). The image analysis unit uses a grayscale variance maximization algorithm to process the images, calculates the grayscale variance of the images at different angles, and selects the angle corresponding to the maximum variance as the optimal polarization angle. At this point, the polarization reflection signal of the fiber texture is suppressed (reflection intensity reduced by ≥60%), and defects form non-polarization reflection signals due to refractive index differences, increasing the grayscale contrast with the background by ≥3 times, achieving preliminary separation of the optical features of the texture and defects.

[0015] Step S113: Based on the optimal polarization angle, the image analysis unit monitors the comparison parameters between the defect area and the background in real time. If the comparison parameters do not meet the preset standard, the analyzer is triggered to make fine adjustments within the preset range. The signal change data during the adjustment process is recorded, and a mapping relationship between the polarization angle and the defect signal intensity is established to form a dynamic adjustment model to adapt to the changes in the filter material texture during production line operation. After the optimal polarization angle is determined, the image analysis unit processes images synchronously with the camera's frame rate at a frequency of 50fps, monitoring the contrast parameters between the defect area and the background in real time. Preset standards are determined through sample statistics during system initialization, with thresholds set for different fiber materials (PP, PET, viscose) and defect types (pinholes, fiber agglomerations). When the contrast parameter is below the preset threshold for three consecutive frames, and the fluctuation exceeds 5% of the threshold, a fine-tuning command for the analyzer is triggered. The fine-tuning range is set to ±5°, employing a forward-then-reverse traversal strategy with a step size of 0.5°. Each adjustment synchronously acquires images and calculates the contrast parameters, recording the adjustment direction, step size, current angle, and contrast data to form an adjustment process dataset. Based on this dataset, a polynomial fitting algorithm is used to establish a mapping model between the polarization angle and the defect signal intensity. A sliding window algorithm (window size 100 data sets) is used to update the model parameters in real time, adapting to the dynamic changes in filter material texture on the production line (e.g., texture density fluctuation ±10%). Meanwhile, during production line operation, the filter material texture feature parameters are extracted in real time and input into the dynamic adjustment model to predict the optimal target angle and drive the analyzer to adjust. If the contrast does not meet expectations after adjustment, a second fine-tuning step of 0.1° is performed to form a closed-loop adaptation mechanism. If the contrast still does not meet the standard within the fine-tuning range, a first-level feedback optimization of ±5% of the light source power is triggered. If this is ineffective, an alarm signal is sent to prompt manual inspection.

[0016] Step S1131: Using the grayscale contrast between the defect area and the background area as the comparison parameter, the image analysis unit performs noise reduction filtering and grayscale preprocessing on the acquired filter material image, delineates the minimum bounding rectangle of the defect area as the target ROI, selects areas with equal area and no texture abrupt changes around the target ROI as the background ROI, and calculates the contrast using a standardized formula to eliminate the interference of the absolute grayscale value under different light intensities and ensure the consistency of the comparison parameters. Using the grayscale contrast between the defect area and the background area as the core contrast parameter, the image analysis unit first performs Gaussian filtering noise reduction on the acquired filter image (filter kernel size 3×3, standard deviation 0.8), and then converts the color image to a grayscale image using a weighted average grayscale conversion algorithm, with weighting coefficients set at R:G:B=0.299:0.587:0.114. The Canny edge detection algorithm (low threshold 50, high threshold 150) is used to identify the defect area, and the smallest bounding rectangle of the defect area is defined as the target ROI, with an ROI area ≥0.01mm² (corresponding to image pixels ≥3×3). Within a 5mm radius of the target ROI, areas with the same area as the target ROI and determined to have no texture abrupt changes by the texture abrupt change detection algorithm (grayscale gradient threshold ≤10) are selected as the background ROI. The contrast is calculated using the standardized formula: contrast = (average grayscale value of defect ROI - average grayscale value of background ROI) / average grayscale value of background ROI. This formula can eliminate the interference of absolute grayscale values ​​under different illumination intensities.

[0017] Step S1132: For different fiber materials and defect types, during the system initialization phase, multiple sets of defect-background sample images are collected. The minimum contrast value under each scene is obtained through statistical analysis as the initial threshold. During the production line operation, the texture image of defect-free filter material is automatically collected at preset time intervals. The grayscale fluctuation range under normal texture is calculated, and the initial threshold is dynamically corrected. During system initialization, 50 sets of defect-background sample images were collected for each of the following: different fiber materials (PP, PET, viscose, etc.) and different defect types (pinholes, fiber agglomerations, impurities, etc.), resulting in a total of 150 sample datasets. The contrast ratio of each sample set was calculated, and the minimum contrast ratio for each material-defect type combination was determined using the arithmetic mean. This minimum contrast ratio was used as the initial threshold for that scene; for example, the initial threshold for pinhole defects in PP was set to 0.15, and the initial threshold for fiber agglomeration defects in PET was set to 0.20. During production line operation, 20 frames of texture images of defect-free filter material were automatically collected every hour. Standard deviation analysis was used to calculate the grayscale fluctuation range under normal texture (e.g., fluctuation range ≤ 0.05). Based on the fluctuation range, the initial threshold was dynamically corrected within ±10%. For example, when the fluctuation range increased to 0.06, the initial threshold of 0.15 was corrected to 0.165. This dynamic calibration mechanism avoids threshold mismatch caused by batch variations in filter material (e.g., differences in fiber thickness and weaving density).

[0018] Step S1133: The image analysis unit processes the image at a frequency synchronized with the camera's frame rate, calculates the contrast of each defect area in real time and compares it with a preset threshold. If the contrast of the same defect area is lower than the preset threshold for a consecutive preset number of frames and the fluctuation exceeds the preset proportion of the threshold, it is determined that the comparison parameter does not meet the standard and triggers the analyzer fine-tuning command. The image analysis unit uses an FPGA chip (model XC7K325T) to synchronize processing with the camera's frame rate. The camera's frame rate is set to 50fps, and the image analysis unit's processing frame rate is also synchronized to 50fps, with a single frame image processing time ≤20ms. For each defect area in each frame, the contrast is calculated in real-time according to the standardized formula in step S1131 and compared with a preset threshold. To avoid false triggering caused by single-image noise (such as random salt-and-pepper noise), a continuous triggering condition is set: if the contrast of the same defect area is lower than the preset threshold for three consecutive frames, and the contrast fluctuation between the three frames exceeds 5% of the threshold (e.g., when the threshold is 0.15, the contrasts of the three frames are 0.13, 0.12, and 0.11 respectively, with a fluctuation exceeding 5%), then the contrast parameter is deemed not to meet the standard. At this time, the image analysis unit sends a polarizer fine-tuning command to the controller of the dual polarizer adjustment system via industrial Ethernet (Profinet). The command includes parameters such as the fine-tuning start signal, initial angle, and fine-tuning range.

[0019] Step S1134: Determine the preset fine-tuning range of the analyzer based on the pre-calibrated optimal polarization angle. Use a preset traversal adjustment strategy to drive the stepper motor to rotate the analyzer with a fixed step size. Simultaneously acquire the filter material image at the current angle for each adjustment step size, calculate the contrast of the defect area and store it in real time. At the same time, record the adjustment direction, step size, current polarization angle and corresponding contrast value to form a complete adjustment process dataset. Based on the optimal polarization angle pre-calibrated in step S112, the preset fine-tuning range of the analyzer is set to ±5°. This range can cover the polarization response shift caused by changes in filter material texture during production. A forward-then-reverse traversal adjustment strategy is adopted, starting from the optimal angle, first rotating the analyzer forward to +5° in 0.5° steps, then rotating it backward to -5° in 0.5° steps, for a total of 21 adjustment nodes. After each adjustment step, the stepper motor driver feeds back the positioning signal, and the camera immediately acquires the filter material image at the current angle. The image analysis unit calculates the contrast of the defect area in real time and stores it in the local database of the industrial control computer (storage capacity ≥1TB). Simultaneously, the adjustment direction (forward / reverse), step size (0.5°), current polarization angle (accurate to 0.01°), and corresponding contrast value (accurate to 0.001) are recorded to form a complete adjustment process dataset. The dataset is sorted by timestamp, and the data of each adjustment node is stored in association to ensure the integrity and traceability of the data when building the subsequent mapping relationship model.

[0020] Step S1135: Preprocess the stored adjustment process dataset, remove abnormal data, establish a nonlinear mapping relationship model between polarization angle as independent variable and defect contrast as dependent variable using a multinomial fitting algorithm, and update the parameters of the mapping relationship model in real time using a sliding window algorithm to adapt to the dynamic changes of filter material texture in the production line. The stored adjustment process dataset is preprocessed, and outlier data is removed using the 3σ criterion. This involves calculating the mean u and standard deviation σ of the dataset and removing abrupt contrast changes (such as sudden increases / decreases in contrast caused by filter media vibration) that exceed the range [u-3σ, u+3σ]. The outlier removal rate is ≤2%. A nonlinear mapping model is established between the polarization angle (x) as the independent variable and the defect contrast as the dependent variable (y) using a third-order polynomial fitting algorithm. The fitting formula is y=a3x. 3 +a2x 2The coefficients a3, a2, a1, and a0 are solved using the least squares method, with a fitting error ≤5%. To address the dynamic changes in filter material texture during production (e.g., texture interlacing angle fluctuations of ±5°), a sliding window algorithm is used to update model parameters in real time. The window size is set to the 100 most recent valid data sets, and a parameter update is triggered every 10 new data sets, with an update time ≤100ms. This dynamic update mechanism ensures that the mapping model can reflect the latest texture-polarization response characteristics in real time.

[0021] Step S1136: The mapping relationship model is fused with the filter material texture feature parameters to construct a multi-input single-output dynamic adjustment model. The gradient descent algorithm is used to train the model with the optimization objective of maximizing contrast. A regularization term is introduced to suppress overfitting. Cross-validation is used to ensure the generalization ability of the model. The output value of the mapping model established in step S1135 (the predicted contrast value under the current polarization angle) is fused with the filter material texture feature parameters (texture density, interlacing angle, gray-level variance, and texture uniformity) to construct a multi-input single-output dynamic adjustment model. The texture feature parameters are extracted using the gray-level co-occurrence matrix to obtain energy, entropy, and correlation indices. Combined with the texture direction histogram, the main direction and distribution dispersion are statistically analyzed. After quantization and extraction, the parameters are concatenated according to the feature dimensions to form a 128-dimensional input feature vector. The output parameter is the optimal polarization angle adjustment amount (in °). The model adopts a lightweight shallow neural network architecture with 128 neurons in the input layer, 2 hidden layers (192 neurons per layer), and the ReLU activation function. The output layer is a single neuron (linear activation). The loss function is set as loss value = 1 - (predicted contrast / historical maximum contrast) + λ × Σ weight. 2 (λ=0.005), the Adam optimizer was selected (initial learning rate 0.001), batch size 32, number of iterations 1000, and an early stopping mechanism was introduced (stopping if the loss on the validation set does not decrease for 5 consecutive rounds). 5-fold cross-validation was used, and the dataset was divided into training and validation sets in a 7:3 ratio.

[0022] Step S11361: Define the set of input parameters for the dynamic adjustment model, including the output value of the mapping relationship model between polarization angle and defect contrast, and the filter material texture feature parameters. Extract texture feature indicators through the gray-level co-occurrence matrix, and combine the texture direction histogram to statistically analyze texture-related parameters to complete the quantitative extraction of multi-dimensional texture features. Concatenate the mapping relationship model-related parameters and texture feature parameters according to feature dimensions to form a unified multi-dimensional input feature vector. Set the model output parameter as the optimal polarization angle adjustment amount to achieve the goal of building a multi-input, single-output model. The model outputs the mapping relationship between polarization angle and defect contrast (i.e., the predicted contrast value under the current polarization angle, with an accuracy of 0.001) and the core texture feature parameters of the filter material (texture density, interlacing angle, gray-level variance, and texture uniformity). Texture feature parameters are extracted using the following methods: energy (range 0-1), entropy (range 0-8), and correlation (range -1-1) indices are extracted through the gray-level co-occurrence matrix (distance 1, angles 0°, 45°, 90°, 135°); the main direction (range 0°-180°) and distribution dispersion (range 0-1) are statistically analyzed through the texture direction histogram (bins=16), completing the quantization extraction of multi-dimensional texture features. Each feature parameter is normalized to the [0,1] interval. The relevant parameters of the mapping relationship model (including the current polarization angle and the predicted contrast value) and the four texture feature parameters, totaling six dimensions, are concatenated in the order of polarization angle, predicted contrast value, texture density, interlacing angle, gray-level variance, and texture uniformity to form a 6-dimensional input feature vector. The model output parameters are set to the optimal polarization angle adjustment amount (range -1° to 1°, accuracy 0.01°), and the model structure of 6 inputs and 1 output is clearly defined to ensure that the model can accurately predict the polarization angle adjustment amount based on multi-dimensional information.

[0023] Step S11362: Collect the adjustment process dataset during production line operation and the texture feature dataset of different batches of filter material, align them according to timestamp to form sample data pairs, use the standardization method to eliminate the difference in the dimension of the feature vector, remove abnormal samples through the outlier detection algorithm, and then divide the training set and the validation set according to the preset ratio. Data sets of adjustment processes accumulated during production line operation (including polarization angle, defect contrast, and adjustment timestamps) and texture feature datasets of different batches of filter media (including material type, texture parameters, and collection timestamps) were collected. These datasets were aligned by timestamp (1ms accuracy) to form one-to-one corresponding sample data pairs, resulting in 10,000 valid samples. The input feature vectors were processed using Z-Score normalization, where each feature parameter was transformed using "(xu) / σ" to eliminate dimensional differences between feature dimensions, ensuring a mean of 0, a variance of 1, and a normalization error ≤1%. An isolated forest algorithm (outlier ratio set to 0.02) was used to remove outliers. Outliers primarily included contrast spikes and texture parameter anomalies caused by filter media vibration, momentary equipment malfunctions, etc., leaving 9,800 valid samples. The training set (6,860 samples) and validation set (2,940 samples) were randomly divided in a 7:3 ratio using stratified sampling to ensure consistent material type and defect type distributions between the training and validation sets. The dataset is stored in CSV format and organized by sample ID, input feature vector, output adjustment amount, and label to ensure data integrity and traceability.

[0024] Step S11363: Construct a lightweight shallow neural network architecture. The number of neurons in the input layer is matched with the dimension of the fused features. The number of hidden layers is set to a preset number. The number of neurons in each layer is configured according to a preset ratio. The activation function is a non-linear activation function. The output layer is a single neuron and uses a linear activation function to output the polarization angle adjustment amount, taking into account both the model fitting ability and real-time inference efficiency. A lightweight shallow neural network architecture is constructed to adapt to the real-time inference requirements of the production line. The total number of model parameters is ≤500,000, and the inference time is ≤10ms. The number of neurons in the input layer matches the fused 6-dimensional feature vector, with 6 neurons receiving normalized input feature parameters. Two hidden layers are set, with the number of neurons in each layer configured according to "input dimension × 32" (i.e., 6 × 32 = 192 neurons). A fully connected layer structure is adopted, and the ReLU function (f(x) = max(0,x)) is selected as the activation function. This activation function can effectively enhance the nonlinear fitting ability of the model while avoiding the gradient vanishing problem. The output layer is a single neuron, using a linear activation function (f(x) = x), directly outputting the optimal polarization angle adjustment. The output range is limited to -1° to 1°, and it is automatically truncated to the boundary value when it exceeds this range. A batch normalization layer (BN layer) is embedded in the network structure, placed after each fully connected layer and before the activation function. Standardization (mean 0, variance 1) accelerates model convergence and improves training stability. The model is built using the TensorFlow framework, with optimized computation graph structure. Quantization training (INT8 quantization) further reduces the model size, ensuring real-time inference on an industrial control computer (CPU model i5-10400) and meeting the production line's 50fps detection frame rate requirement.

[0025] Step S11364: Construct a loss function with the core optimization objective of maximizing contrast. The loss function includes a correlation term between predicted contrast and maximum possible contrast and an L2 regularization term. The regularization strength is adjusted by the regularization coefficient to suppress model overfitting. A loss function is constructed with maximizing contrast as the core optimization objective. This loss function consists of two parts: the first part is the correlation term between predicted contrast and maximum possible contrast, and the second part is the L2 regularization term. The correlation term is expressed as "1 - (predicted contrast / maximum possible contrast)", where the predicted contrast is calculated from the polarization adjustment amount of the model output (using the standardized formula in step S1131), and the "maximum possible contrast" is the contrast peak value in the historical best sample (preset to 0.8). This part is used to constrain the adjustment amount of the model output to make the contrast as close to the maximum value as possible. The L2 regularization term is expressed as "λ × Σw²", where λ is the regularization coefficient (initial value 0.005), and w is all the weight parameters of the model. By calculating the sum of squares of the weight parameters and multiplying it by λ, overfitting caused by excessively large weight parameters is suppressed. The total loss function is expressed as "Loss = 1 - (Pred Contrast / Max Contrast The regularization coefficient λ can be adaptively adjusted according to the model's training state. The initial value is set empirically, and it is dynamically corrected during the training process based on the loss trends of the training and validation sets.

[0026] Step S11365: Select an adaptive momentum estimation optimizer to implement gradient descent, adopt a dynamic adjustment strategy to adjust the learning rate, configure the preset batch size and number of iterations, introduce an early stopping mechanism, and stop iterating and save the optimal model parameters when the validation set loss has not decreased for a preset number of consecutive rounds. The Adaptive Momentum Estimation (Adam) optimizer is chosen to implement gradient descent. This optimizer combines momentum gradient descent with an adaptive learning rate strategy, enabling fast convergence and reducing the likelihood of getting trapped in local optima. The initial learning rate is set to 0.001, and an exponential decay strategy is used to dynamically adjust the learning rate, decreasing it by 10% every 100 iterations. The learning rate update formula is: lr = 0.001 × 0.9 (epoch / 100) This ensures rapid exploration of the parameter space in the early stages of training and accurate convergence in the later stages. The batch size is configured as 5% of the number of training set samples. For a training set of 6860 samples, the batch size is 32 (6860 × 5% ≈ 343, taking the nearest power of 2, 32), balancing training efficiency and memory usage. The number of iterations is set to 1000 rounds, and an early stopping mechanism is introduced. When the loss on the validation set does not decrease for 5 consecutive rounds (loss change ≤ 0.001), the iteration stops and the current optimal model parameters are saved to avoid overfitting due to overtraining. The momentum parameters β1, β2, and epsilon of the optimizer are set to 0.9, 0.999, and 1e-8. This parameter configuration ensures the stability and accuracy of gradient updates, accelerating the model's convergence to the optimal solution.

[0027] Step S11366: Initialize the model weight parameters, input the training set samples into the model in batches, calculate the polarization angle adjustment through forward propagation, calculate the loss value with regularization term by combining the true contrast, solve the gradient of the loss function with respect to each weight parameter using the backpropagation algorithm, update the weight parameters using the optimizer, and repeat the above process until the preset number of iterations is reached or the early stopping mechanism is triggered. The Xavier initialization method is used to initialize the model weight parameters, ensuring consistent variance between the input and output of each layer and preventing gradient vanishing or exploding in the early stages of training. Training set samples are grouped into batches of 32 and input into the model batch by batch for training. During forward propagation, the input feature vector is passed through the input layer and hidden layers (including BN and ReLU activation) to the output layer, adjusting the output polarization angle. The total loss value (including L2 regularization) is calculated based on the corresponding true contrast ratio. During backpropagation, the chain rule is used to solve for the gradient of the loss function with respect to each weight parameter. The Adam optimizer is used to update the weight parameters based on the gradient direction and momentum β1 and β2 information. Forward and backpropagation of all training set samples is completed in each iteration. During iteration, the loss values ​​and contrast compliance rates of the training and validation sets are recorded in real time. Training stops and the optimal model parameters are saved (in .h5 format) when the preset number of iterations (1000 iterations) is reached or the early stopping mechanism is triggered (no decrease in validation set loss for 5 consecutive iterations). The entire training process is completed on an industrial control computer, with a training time of ≤2 hours, ensuring rapid model deployment and application.

[0028] Step S11367: During the training process, monitor the loss change trend of the training set and validation set in real time, and adaptively adjust the regularization coefficient according to the overfitting or underfitting state of the model, so that the loss of the model on the training set and validation set tends to be stable and the difference is within the preset range. During training, the loss trends of the training and validation sets are monitored in real time every 10 iterations. The model's fit is judged by calculating the loss difference (validation set loss - training set loss). If the training set loss continues to decrease (≥5% decrease every 10 iterations) while the validation set loss increases (≥3% increase every 10 iterations), the model is considered overfitting. In this case, the regularization coefficient λ is gradually increased by 20% each time (e.g., from 0.005 to 0.006) until the validation set loss stabilizes. If both the training and validation set losses remain high (loss ≥0.2 after 50 iterations), the model is considered underfitting. λ is gradually decreased by 20% each time (e.g., from 0.005 to 0.004), while checking the completeness of feature extraction and supplementing texture feature parameters if necessary. This dynamic adjustment mechanism stabilizes the model's loss on the training and validation sets, with a loss difference of ≤0.05. This ensures that the model fully learns the patterns in the training data without overfitting, thus possessing good generalization ability and adapting to the filter material testing needs of different batches and working conditions.

[0029] Step S11368: Using the k-fold cross-validation method, the dataset is divided into multiple non-overlapping subsets. Some subsets are selected in turn as the training set and the remaining subsets are selected as the validation set. The training and validation process is repeated. The average loss value and contrast compliance rate of multiple validations are calculated. The model structure or feature fusion method is adjusted according to the validation results to ensure that the model's generalization ability meets the standard. A 5-fold cross-validation method was used to verify the model's generalization ability. The preprocessed 9800 valid samples were randomly divided into 5 non-overlapping subsets (1960 samples per subset), with consistent material and defect type distributions across subsets. Four subsets (7840 samples) were selected alternately as the training set, and the remaining subset as the validation set. This training and validation process was repeated 5 times, using the same model structure, optimizer parameters, and training procedure each time. The average loss and average contrast compliance rate were calculated across the 5 validations. Evaluation criteria were set as follows: average loss ≤ 0.08, coefficient of variation (standard deviation / mean) of the loss at each fold ≤ 10%, and average contrast compliance rate ≥ 92%. If the coefficient of variation was too high (> 10%), it indicated insufficient generalization ability. In this case, the model structure was adjusted, such as increasing the number of hidden layer neurons to 256, or optimizing the feature fusion method (increasing the dimensionality of texture feature parameters), and training and cross-validation were repeated. This cross-validation process ensures the model's performance is stable across different subsets of data, avoids misjudgments of performance due to data partitioning bias, and ultimately enables the model's generalization ability to meet the testing requirements of different batches of filter media on the production line.

[0030] Step S11369: Evaluate the model performance using an independent test set. If the performance does not meet the preset standard, adjust the model structure, number of neurons, or activation function type, and re-execute the training, regularization adjustment, and cross-validation process to form a closed-loop iteration until the model performance meets the production line detection requirements.

[0031] After training, an independent test set is selected to evaluate the model's performance. The test set contains 2000 samples (20% of the total samples), which were not used in training and cross-validation, and covers different materials, defect types, and production line conditions. Evaluation metrics include contrast prediction error (|predicted contrast - true contrast| / true contrast × 100%) and polarization adjustment accuracy (number of samples meeting the contrast standard after adjustment / total number of samples × 100%). Performance standards are set as follows: contrast prediction error ≤ 8%, and polarization adjustment accuracy ≥ 90%. If the test set evaluation does not meet the standards (e.g., polarization adjustment accuracy ≥ 85%), the model returns to the model structure design stage for optimization. For example, the activation function can be replaced with LeakyReLU (negative slope 0.01), or one hidden layer (128 neurons) can be added. The training, regularization adjustment, and cross-validation process is then re-executed, forming a closed-loop iteration of "design-training-validation-optimization". A maximum of three iterations of optimization are performed.

[0032] Step S1137: During production line operation, the image analysis unit extracts the texture feature parameters of the filter material in real time and inputs them into the dynamic adjustment model. The model predicts the target angle that can make the contrast reach the optimal value based on the current polarization angle and texture features. It drives the stepper motor to adjust the polarizer to the target angle and monitors the contrast after adjustment in real time. If the expected optimal value is not reached, a second fine adjustment is made based on the model prediction error, forming a closed-loop adaptation mechanism of monitoring, prediction, adjustment and verification. During production line operation, the image analysis unit synchronizes with the camera's frame rate (50fps) to extract texture feature parameters of the filter material in real time. The extraction process is as follows: after each frame is acquired, noise reduction and grayscale preprocessing are performed first. Then, four core parameters, including texture density and interlacing angle, are calculated using the grayscale co-occurrence matrix and texture direction histogram, with an extraction time of ≤5ms. The extracted texture feature parameters are concatenated with the current polarization angle and the contrast prediction value output by the mapping model to form a 6-dimensional input feature vector, which is then input into the trained dynamic adjustment model. Based on the current polarization angle and texture features, the model quickly predicts the target angle that will achieve the optimal contrast through forward propagation, with a prediction time of ≤10ms. After receiving the target angle signal, the central control unit drives the stepper motor to adjust the analyzer to the target angle, with an adjustment accuracy of ±0.01° and an adjustment time of ≤50ms. Meanwhile, the image analysis unit monitors the adjusted contrast in real time. If the adjusted contrast is lower than 90% of the predicted value, it performs a second fine-tuning based on the model prediction error (predicted contrast - actual contrast) with a fine-tuning step size of 0.1° until the contrast reaches the expected optimal value (≥90% of the predicted value).

[0033] Step S1138: If the contrast of the defect area still does not reach the preset threshold after the analyzer has been adjusted throughout the preset fine-tuning range, the system will automatically determine that it is an abnormal situation and trigger two levels of feedback; the first level of feedback is to adjust the power of the light source to assist in optimization, and the second level of feedback is to send an alarm signal to the control system to prompt the operator to check the texture of the filter material or the status of the equipment.

[0034] The preset fine-tuning range of the analyzer is ±5°. After the analyzer completes forward and reverse traversal adjustments within this range in 0.5° steps, if the contrast of the defective area is still lower than the preset threshold (e.g., 0.15), the system automatically determines it as an abnormal situation and triggers a two-level feedback mechanism. The first-level feedback is to adjust the light source power to assist in optimization: the central control unit sends a power adjustment command to the near-infrared structured light source controller, with a power adjustment range of ±5% (e.g., if the current power is 30W, adjust it to 28.5W or 31.5W). After adjustment, the image is re-acquired and the contrast is calculated. If the contrast meets the standard (≥ the threshold), the current parameters are maintained; if it still does not meet the standard, the second-level feedback is activated. The secondary feedback sends an alarm signal to the control system (PLC S7-1200). The alarm signal is a digital signal (active high) transmitted via the Profinet bus. After receiving the signal, the control system triggers the audible and visual alarm (model TBJ-100) and displays the alarm information on the human-machine interface (HMI). The prompts include "Abnormal filter material texture, please check fiber distribution" and "Polarizer may be contaminated, please clean it." At the same time, the time of the abnormality, production line speed, and filter material batch information are recorded.

[0035] Step S114: The transmission photoelectric sensor emits detection light in the same wavelength band as the near-infrared structured light source, which penetrates vertically through the filter material and collects the transmitted light signal. Combined with the filter material thickness detection data, the actual transmittance of the filter material is calculated. At the same time, the light reflection distribution on the filter material surface is collected to analyze the difference in transmittance uniformity. A transmittance photoelectric sensor (model E3Z-T61) is installed below the filter media conveying path. The sensor's transmitter and receiver are fixed to opposite sides of the filter media conveying roller, ensuring that the emitted detection light is in the same wavelength band (850nm) as the near-infrared structured light source and penetrates the filter media surface perpendicularly (perpendicularity deviation ≤1°). The transmitter outputs a constant intensity detection light (1000 lux), and the receiver uses a photodiode to receive the transmitted light signal, converting it into a voltage signal (output range 0-10V). The voltage signal is acquired by an ADC module (sampling rate 1kHz, accuracy 12-bit) to calculate the transmitted light intensity. Simultaneously, a laser displacement sensor (model Keyence IL-300) is installed next to the sensor to synchronously acquire filter media thickness data (measurement range 0-5mm, accuracy ±0.01mm). The actual transmittance of the filter media is calculated using the formula "transmittance = (transmitted light intensity / incident light intensity) × (standard thickness / actual thickness) × 100%", with the standard thickness preset to 1mm. In addition, an image of the light reflection distribution on the surface of the filter material (resolution 1920×1080) was acquired by an area array image sensor (model IMX290), and the difference in light transmission uniformity was analyzed by a gray-scale variance algorithm. When the gray-scale variance was ≥0.05, it was judged as poor uniformity.

[0036] Step S115: Based on the transmittance detection results, call the preset parameter mapping model to dynamically adjust the light spot density, power and projection angle of the dot matrix near-infrared light source. For areas with large differences in transmittance uniformity, start the light source partition adjustment mode to adjust the light source unit corresponding to the weak transmittance area individually. The preset parameter mapping model is a BP neural network-based model, which has been trained using historical data. The input is the filter media transmittance (range 0-100%), and the output is the optimal parameters for the light spot density, power, and projection angle of the near-infrared structured light source. Based on the actual transmittance of the filter media calculated in step S114, the central control unit calls this mapping model to dynamically adjust the core parameters of the dot matrix near-infrared light source: when the transmittance is below 30%, the light spot density is adjusted to 25-30 spots / cm², the power to 40-50W, and the projection angle to 30°-45° to enhance surface penetration; when the transmittance is above 70%, the light spot density is adjusted to 10-15 spots / cm², the power to 10-20W, and the projection angle to 15°-20° to avoid reflection interference; when the transmittance is between 30% and 70%, the parameters are adjusted using linear interpolation. For areas with significant differences in light transmittance uniformity (grayscale variance ≥ 0.05), the light source zoning adjustment mode is activated. The light source is divided into 8 independent adjustment zones along the width of the filter material. Each zone corresponds to a set of LED light-emitting units. By detecting the light transmittance data of the zone, the power of the light source units corresponding to the weak light transmittance areas (light transmittance 10% lower than the surrounding areas) is increased individually (by 10%-20%) to achieve preliminary light field balance and light field uniformity ≥ 85%.

[0037] Step S116: The light field distribution on the filter material surface is collected in real time by the light sensor array. The light field uniformity evaluation algorithm is used to judge the current light field uniformity. If the preset requirements are not met, secondary adjustment is performed by independently controlling and adjusting the vertical distance between the light source and the filter material and the focusing state of the lens group. At the same time, the power of the near-infrared light source is dynamically compensated according to the changes in the ambient light intensity of the production line. A two-dimensional light sensor array (model TSL2591, array density 16×16, detection range 0-65535 lux) is installed in the projection path of a near-infrared structured light source. The distance between the sensor array and the filter surface is 80 mm, and the light field distribution image of the filter surface is acquired in real time (sampling rate 10 fps). The light field uniformity evaluation algorithm (coefficient of variation method) is used to determine the current light field uniformity. The coefficient of variation = standard deviation of light field intensity / average value of light field intensity. The preset uniformity requirement is a coefficient of variation ≤ 5%. If the coefficient of variation is higher than 5%, a secondary adjustment of the light source parameters is triggered: the power of each LED light-emitting unit is independently controlled by the light source controller (adjustment accuracy 1%) to compensate for the light intensity in the weakly lit areas; at the same time, the vertical distance between the light source and the filter material is adjusted by a stepper motor (adjustment range 50-100 mm, step size 1 mm), combined with the focusing state of the lens group (focal length adjustment range 20-50 mm), so that the light spot forms a continuous and uniform light coverage on the filter surface. Ambient light sensors (model BH1750) are installed around the production line to collect ambient light intensity in real time (sampling rate 1Hz). When the ambient light fluctuation exceeds 10%, the power of the near-infrared light source is dynamically compensated within ±8%, with a compensation response time ≤100ms, effectively offsetting ambient light interference and ensuring the stability of the light field.

[0038] Step S117: After completing the adjustment of the polarization angle and the light source parameters, acquire the filter material image under the current state, and perform collaborative verification through the defect signal-to-noise ratio. If the signal-to-noise ratio does not reach the optimal value, simultaneously fine-tune the polarization angle and the light source power, repeatedly acquire images and calculate the signal-to-noise ratio, until a collaborative closed loop of polarization adjustment, light source adjustment and signal-to-noise ratio verification is formed.

[0039] After adjusting the polarization angle and light source parameters, the camera acquires images of the filter material at a frame rate of 50fps, for a total of 10 frames for collaborative verification. Image quality is evaluated using the defect signal-to-noise ratio (SNR), calculated as: SNR = (Defect region signal intensity - Average gray value of defect ROI - Average gray value of background ROI) / (Background noise intensity - Standard deviation of background ROI gray value). The preset optimal SNR value is ≥20dB. If the average SNR of the 10 acquired frames does not reach 20dB, a synchronous fine-tuning process is initiated: the two parameters are adjusted synchronously in steps of ±0.5° for the polarization angle and ±3% for the light source power. Five frames are acquired and the average SNR is calculated after each adjustment. A greedy algorithm is used to find the optimal parameter combination, meaning that each adjustment only retains the parameter changes that improve the SNR, until the average SNR of the 10 frames is ≥20dB, forming a collaborative closed loop of polarization adjustment, light source adjustment, and SNR verification. The adjustment cycle of the collaborative closed loop is ≤1s, ensuring that stable image quality can be maintained even if the characteristics of the filter material or environmental conditions change during production line operation, providing high-quality raw data for subsequent image processing and defect detection.

[0040] Step S120: Dynamically adjust imaging parameters according to the production line operation status. Collect spectral data, internal structure data, material distribution data and functional index data of the filter material through the multi-source detection module. Link with the production line control unit through the synchronous control mechanism to achieve spatiotemporal registration of the data collected by each module in the same detection area. This is achieved through production line status sensing, dynamic adjustment of imaging parameters, multi-source data acquisition, and spatiotemporal registration and coordination. A speed sensor (Keyence FS-V21), a tension sensor (HBM U9C), and a near-infrared material identification sensor (Micro-Epsilon NIRscan) are deployed to collect production line speed (accuracy ±0.1m / min), filter material tension (range 0-50N), and material information. This data is transmitted in real-time to the central control unit (PLCS7-1500) via the Profinet industrial bus. Simultaneously, process switching and start / stop signals from the production line PLC are input, constructing an 8-dimensional operational status sensing matrix. Imaging parameters include sampling frequency (50-200fps) and exposure time (10-100us), dynamically output through an associative mapping model. The multi-source detection module includes a hyperspectral camera (400-1700nm), an ultrasonic phased array probe (5MHz), a near-infrared sensor (850nm), and an air permeability sensor array, with a unified acquisition frame rate of 50fps. An NI PXIe-6674 synchronization controller is configured to generate synchronization trigger signals based on the position pulse signals from the production line encoder (1000 lines pulse resolution). The trigger interval is matched to the production line speed (0.01s trigger interval at v=100m / min). A reference point is calibrated using a laser tracker to construct a Cartesian coordinate system. Affine transformation is used to correct field-of-view deviations. Combined with an NTP time synchronization and delay compensation model (linear fitting to correct delays ≤0.5ms), spatiotemporal registration of data in the same detection area is achieved.

[0041] Step S121: Deploy multiple sensors to collect data on production line conveying speed, filter material tension fluctuation, and filter material material information. Transmit the collected data to the central control unit in real time via the industrial bus. At the same time, access the operating status signal of the production line control unit to obtain production process switching and start / stop status information, and construct a production line operating status perception matrix. A speed sensor (Keyence FS-V21) is deployed at the end of the filter media conveyor roller to collect real-time production line conveyor speed (measurement range 0-300m / min, accuracy ±0.1m / min); a tension sensor (HBMU9C) is installed at the tension roller along the conveyor path to monitor filter media tension fluctuations (range 0-50N, resolution 0.01N); and a near-infrared material identification sensor (Micro-Epsilon NIRscan) is deployed at the front end of the detection area to identify material type and batch information by analyzing the filter media spectral response (800-1000nm). Data from these three types of sensors is transmitted in real-time to the central control unit (PLCS7-1500) via a Profinet industrial bus (transmission rate 100Mbps), with a transmission delay ≤10ms. Simultaneously, the system accesses the production line PLC's operating status signals via an EtherCAT interface to obtain key information such as production process switching (e.g., material change, speed adjustment) and start / stop status (running / stopping / standby). Eight types of data, including speed, tension, material, and process status, are fused according to timestamps (accuracy 1ms) to construct an 8×N-dimensional production line operation status perception matrix (N is the number of sampling points). The weighted average method is used to fuse multi-sensor data to reduce noise interference (data fluctuation amplitude ≤2%).

[0042] Step S122: Based on historical production data, establish a correlation mapping model between production line operation status and imaging parameters, clarify the correspondence between imaging parameters and production line speed, filter tension, and material type, embed adaptive adjustment rules and parameter adjustment priorities into the model, and balance imaging quality and real-time performance. Based on production data from the past 12 months (covering 5 materials, 30 batches, and speeds of 50-200 m / min), a backpropagation (BP) neural network correlation mapping model was constructed (8 neurons in the input layer, 2 hidden layers with 64 neurons each, and 4 neurons in the output layer). The correspondence between core imaging parameters and production line status was defined: sampling frequency f = 0.02 × v + 30 (v is the production line speed, in m / min), exposure time t = 500 / v + 5 (in µs), gain coefficient g = 0.01 × T + 1 (T is the tension, in N), and light source intensity I = 2 × M + 10 (M is the material absorption coefficient, 0-5). An adaptive adjustment rule was embedded in the model: when production line speed fluctuations > 10%, the sampling frequency is adjusted first; when tension fluctuations > 5%, the gain is adjusted first; and when materials are changed, the light source intensity and exposure time are updated synchronously. The parameter adjustment priority was set as: sampling frequency > exposure time > light source intensity > gain coefficient, ensuring a balance between imaging quality (defect resolution ≥ 0.01 mm) and real-time performance (single frame processing time ≤ 20 ms). The model weights are optimized using the random forest algorithm.

[0043] Step S123: The central control unit calls the associated mapping model to output target imaging parameters based on the real-time perceived production line operating status, and sends them to the imaging device through the signal transmission module. After receiving the instruction, the imaging device adjusts the core imaging parameters and the relevant parameters of the supporting light source in real time, and records the parameter adjustment log simultaneously, forming a closed-loop response mechanism of status perception, model calculation and parameter adjustment. The central control unit (PLC S7-1500) reads the production line operation status perception matrix data in real time, calls the associated mapping model every 10ms, and outputs the target imaging parameters (sampling frequency, exposure time, gain coefficient, and light source intensity). The parameter commands are sent to the imaging equipment (MV-CH080-10GM line scan camera and HL-850-50 near-infrared light source) via an EtherNet / IP signal transmission module, with a command transmission delay of ≤5ms. After receiving the commands, the imaging equipment adjusts the parameters in real time through its built-in control module: the sampling frequency is adjusted in 0.1fps steps (range 50-200fps), the exposure time in 1µs steps (10-100µs), the gain coefficient in 0.01 steps (1.0-2.0), and the light source intensity in 1% steps (10-50W). A parameter adjustment log is recorded synchronously, including a timestamp (1ms accuracy), current production line status, target parameter value, and actual adjustment value. The log is stored on an industrial SD card (64GB capacity) and retained for 3 months.

[0044] Step S124: Along the filter media conveying path, deploy multi-source detection modules at corresponding positions in the same detection area. Each module collects spectral data, internal structure-related data, material distribution-related data, and functional index data of the filter media. Unify the acquisition frame rate of each module to ensure consistent data output rhythm. Along the filter media conveying path (detection area length 500mm), multi-source detection modules are deployed at the following locations: a hyperspectral camera (Headwall Hyperspec, band 400-1700nm, resolution 5nm) is installed 50mm from the front end of the detection area to collect spectral data from the filter media surface; an ultrasonic phased array probe (Olympus OmniScan X3, center frequency 5MHz, detection depth 0-5mm) is installed below the conveying roller to collect internal structural data from the filter media conveying surface; a near-infrared sensor (Hamamatsu G12669, wavelength 850nm, measurement distance 50mm) is deployed next to the hyperspectral camera to collect material distribution data; and a miniature air permeability sensor array (Sensirion SDP810, range 0-500sccm / cm², accuracy ±2%) is embedded in the support mechanism 30mm from the rear end of the detection area to collect functional index data. All modules are connected via a synchronization signal control line, with a uniform acquisition frame rate of 50fps (acquiring data once every 20ms), and the data output format is uniformly 16-bit RAW format (image data) and CSV format (sensor data).

[0045] Step S125: Configure a high-precision synchronous controller, establish linkage with the production line control unit through the communication interface, obtain the position pulse signal of the production line encoder as the time and space reference, and output synchronous trigger signal to each multi-source detection module. The trigger interval is matched with the production line speed to ensure that each module starts collecting data synchronously when the filter material runs to the target detection area. A high-precision NI PXIe-6674 synchronous controller (timing accuracy ±1ns) is configured and linked with the production line PLC via an EtherCAT communication interface to acquire the position pulse signals of the production line encoder (Omron E6B2-CWZ6C, pulse resolution 1000 lines) in real time, using this as a time and space reference. The synchronous controller calculates the real-time position of the filter material based on the encoder pulse signals (pulse count × roller circumference / number of pulses), and generates a synchronous trigger signal (TTL level, active high) based on the detection area length (500mm). The trigger interval is calculated using the formula t=L / (v×1000 / 60) (L is the detection area length, v is the production line speed). For example, when v=100m / min, the trigger interval t=0.03s. The synchronous controller sends trigger signals to each multi-source detection module through a multi-channel output interface (8 channels), with a signal transmission delay ≤0.1ms. After receiving the trigger signal, each module synchronously starts data acquisition, ensuring that all modules start acquiring data at the same time when the filter media moves to the center of the target detection area. The acquisition start time difference is ≤1ms, realizing synchronous acquisition in space and avoiding misalignment of the detection area caused by the movement of the filter media.

[0046] Step S126: Using positioning calibration technology, set calibration reference points in the detection area, measure the spatial position coordinates of each detection module and the reference points, construct a unified spatial coordinate system based on the coordinate data, and correct the detection field deviation of each module through coordinate transformation algorithm so that the detection range of all modules accurately covers the same target area. Repeat the calibration process according to the preset cycle to compensate for spatial deviation. Positioning calibration was performed using a laser tracker (Leica AT960-MR, measurement accuracy ±0.5μm / m). Five calibration reference points (2mm diameter reflective marks) were set at the four corners and the center of the detection area. The spatial coordinates (X, Y, Z, accuracy ±1μm) of the lens / probe center of each detection module (hyperspectral camera, ultrasonic probe, etc.) relative to each reference point were measured using the laser tracker. Based on the coordinate data of the five reference points, a unified Cartesian coordinate system (origin as the central reference point) was constructed. To address the detection field of view deviation of each module, a perspective transformation algorithm was used to perform coordinate transformation, establishing a mapping relationship between the module's local coordinate system and the unified coordinate system (the transformation matrix was solved using the least squares method), correcting the field of view deviation (after correction, the field of view overlap rate ≥98%). The preset calibration cycle is every 8 hours or every 1000 meters of filter media produced. The automatic calibration process is started: the filter media is transported to the calibration position, the coordinates of the reference point are remeasured, the transformation matrix is ​​updated, and the spatial deviation caused by equipment vibration and installation displacement is compensated (the deviation after compensation is ≤0.1mm) to ensure that the detection range of all modules accurately covers the same target area.

[0047] Step S127: When the synchronous controller sends the acquisition trigger signal, it records the current encoder pulse count and system time, generates a unique timestamp and position tag, and automatically associates the timestamp and position tag with the data collected by each multi-source detection module. The synchronization controller (NI PXIe-6674) simultaneously sends the acquisition trigger signal and records the pulse count of the current production line encoder (accuracy 1000 lines / revolution) using a counter. Combined with the roller circumference (500mm), it calculates the spatial position of the filter material (position = pulse count × 500mm / 1000). It synchronizes the system time via the NTP protocol (accuracy ±1ms) to generate a unique timestamp (format: YYYYMMDDHHMMSSmmm). The pulse count (spatial position) is combined with the timestamp to generate a unique spatiotemporal tag (format: timestamp-pulse count). Each multi-source detection module (camera, sensor) has a built-in trigger response module that immediately records the spatiotemporal tag upon receiving the trigger signal. After data acquisition is complete (acquisition time ≤ 10ms), the spatiotemporal tag is automatically embedded in the data file header (EXIF information of image files, first column of sensor data CSV). This ensures that each set of spectral data, internal structure data, material distribution data, and functional data is associated with a unique spatiotemporal tag.

[0048] Step S128: The data collected by each module is transmitted using real-time industrial Ethernet. The central control unit has a built-in data transmission delay monitoring module to calculate the data transmission delay in real time. Based on the delay data, a delay compensation model is established to calibrate the time axis of the data from different modules and correct the time difference caused by the transmission delay. Data collected by each module is transmitted via EtherCAT real-time industrial Ethernet at a rate of 1Gbps with a data frame period of ≤1ms. The central control unit (PLC S7-1500) has a built-in data transmission delay monitoring module. This module sends test data packets (1KB in size) to each detection module, records the sending and receiving response times, and calculates the transmission delay from acquisition to reception (delay = reception time - acquisition time - processing time). Based on 100 sets of past delay data, a delay compensation model is established using a linear fitting algorithm (compensation value = 0.98 × delay + 0.02ms) to calibrate the time axis of data from different modules. For example, if the hyperspectral camera has a transmission delay of 12ms and the ultrasonic probe has a delay of 8ms, the timestamp of the ultrasonic probe data is corrected backward by 4ms to ensure that the time difference between multi-source data under the same spatiotemporal label is ≤0.5ms. The calibrated data is stored in an industrial database (MySQL, 1TB storage capacity) categorized by spatiotemporal label.

[0049] Step S129: The central control unit performs consistency verification on the registered multi-source data, evaluates the registration effect by calculating the spatial overlap rate and time synchronization error, and if the preset standard is not met, it triggers the spatial calibration process or adjusts the synchronization trigger signal interval, optimizes the delay compensation model, records the registration parameter adjustment data, and continuously iterates and optimizes the correlation mapping model. The central control unit verifies the consistency of the registered multi-source data, calculating the spatial overlap rate (IOU = intersection of detection regions / union of detection regions) and the time synchronization error (|data1 timestamp - data2 timestamp|). Preset verification standards: spatial overlap rate ≥ 98%, time synchronization error ≤ 1ms. If the spatial overlap rate is below 98%, the spatial calibration process is triggered (automatic calibration in step S126 is executed); if the time synchronization error exceeds 1ms, the synchronization trigger signal interval is adjusted (step size 0.001s) or the delay compensation model is optimized (refitting the delayed data). Data on each registration parameter adjustment (calibration time, correction value, registration effect) is recorded, and the parameters of the association mapping model are updated using an incremental learning algorithm (weights are updated every 1000 data sets), continuously optimizing the model's adaptability to changes in production line conditions. Through this dynamic optimization mechanism, registration accuracy gradually improves over time.

[0050] Step S1210: When the data acquisition of a certain detection module is abnormal or the registration error continues to exceed the standard, the system automatically marks the abnormal module and switches to the redundant acquisition mode, and sends an early warning signal to the production line control unit.

[0051] The system monitors the data acquisition status (data missing rate, signal strength) and registration error of each detection module in real time, and sets anomaly judgment thresholds: data missing rate ≥5%, and 10 consecutive sets of registration errors >1ms. When a module triggers the anomaly threshold, the system automatically marks the module as abnormal (e.g., "hyperspectral camera acquisition abnormality") and switches to redundant acquisition mode: activating a backup detection module (same model equipment, hot backup status), or calling the module's historical registration parameters (average parameters over the past 5 minutes) for emergency calculation. Simultaneously, a warning signal (digital high-level active) is sent to the production line PLC via the Profinet bus. Upon receiving the signal, the PLC triggers an audible and visual alarm (TBJ-100) with a sound level ≥85dB. The alarm information is also displayed on the HMI interface, including the abnormal module name, anomaly type (acquisition abnormality / registration overrun), occurrence time, and suggested troubleshooting measures (e.g., "check the cleanliness of the hyperspectral camera lens" and "calibrate the ultrasonic probe position").

[0052] Step S130: A deep learning-based texture stripping method is adopted, and an adaptive texture template is generated by combining the texture distribution law of the filter material. The background texture is stripped by dynamically adjusted texture separation logic while retaining the information of minor defects. Then, the edge and gray-level difference of minor defects are enhanced by feature enhancement technology. Feature spectral information corresponding to filter material and internal defects is selected, and spectral data and spatial image details are fused to generate an enhanced image. A dynamic deblurring algorithm is used to eliminate the image blur caused by production line movement. The true boundary of the defect is restored by edge sharpening technology. Image enhancement and defect boundary restoration are achieved through a seven-stage collaborative processing approach. First, a lightweight U-Net texture stripping model (≤1 million parameters) is invoked to generate an adaptive template based on the texture patterns of the filter material. Otsu adaptive thresholding and pixel-level subtraction are used to strip the background texture, with a defect retention threshold set to a gray-level difference ≥3, ensuring complete preservation of minute defects <0.1mm. The stripped image undergoes a three-layer wavelet multi-scale decomposition, and high-frequency defect features are weighted and enhanced using a CBAM attention mechanism (weight coefficients are dynamically adjusted based on defect response values), resulting in a defect edge contrast improvement of ≥2 times after reconstruction. A hyperspectral camera acquires data across the 400-1700nm band, which is preprocessed with Savitzky-Golay smoothing and Min-Max normalization. Feature bands are selected using spectral angle matching (threshold ≤0.1rad) and PCA (first 3 principal components) (retaining 30% of the effective bands). An adaptive weighted fusion strategy (higher texture complexity entropy values ​​result in higher spatial weights) is adopted, and pixel-level weighted fusion (fused pixels = α × spectral pixels + (1-α) × spatial pixels) is used to generate enhanced images (signal-to-noise ratio ≥ 30dB). Based on the CGAN dynamic deblurring model, the production line speed label is input to generate an adaptive deblurring kernel (blurring direction 0-180°, scale 1-10 pixels), and the deblurring processing time is ≤ 15ms. Finally, Canny-Laplacian multi-scale edge detection is used, combined with gradient boosting and pixel-level aggregation techniques to repair edge breaks. The sharpening parameters are adaptively adjusted according to the defect size (<0.1mm intensity × 1.5) to restore the true boundary of the defect (accuracy ≤ 0.05mm).

[0053] Step S131: Collect texture images of spunlace filter media of different materials and samples with minor defects, label the background texture area and defect area, construct a lightweight U-Net texture stripping model with encoder and decoder structure, extract multi-scale features of filter media texture through encoder, restore image size through decoder and introduce skip connections to fuse encoder features, optimize model parameters by combining cross-entropy loss function and structural similarity loss, so that the model can learn the texture distribution pattern of filter media of different materials; We collected 1000 defect-free texture images and 800 images with minor defects (pinholes <0.1mm, fiber agglomerations) for each of five types of spunlace filter media: PP, PET, and viscose. Pixel-level annotations were performed using the LabelMe tool to distinguish between background textures and defect areas, generating binary annotation masks (1 for defects, 0 for background). Data augmentation techniques such as random rotation (±10°), horizontal mirroring, grayscale perturbation (±5%), and Gaussian noise addition (variance 0.01) were used to expand the sample to 15000 images, which were then divided into a 7:2:1 ratio for training (10500 images), validation (3000 images), and test (1500 images). A lightweight U-Net model was constructed, using depthwise separable convolutions (number of groups = number of input channels) instead of standard convolutions. The input size was set to 512×512 pixels (adapting to line scan camera resolution). The encoder-decoder core structure (4 encoder layers, 4 decoder layers) was retained, while redundant fully connected layers were removed. The encoder extracts multi-scale texture features using 3×3 and 5×5 convolutional kernels. The lower layers extract edge details, while the higher layers aggregate global patterns through downsampling with a stride of 2. Each layer is followed by batch normalization (momentum 0.99) and the ReLU activation function. The decoder upsamples through 2×2 transposed convolutions, concatenates the encoder's features via skip connections, and outputs a 2-channel feature map (texture + defects) via a final 1×1 convolution. A combined loss function is used, fusing cross-entropy loss (weight 0.7) and SSIM loss (weight 0.3). The Adam optimizer (initial learning rate 0.001, cosine annealing decay) is employed with a batch size of 32 and 200 iterations. Dropout (rate 0.2) and early stopping (patience value 10) are introduced to suppress overfitting.

[0054] Step S1311: Collect standard texture images without defects and defect detection images of spunlace filter media of various materials, divide them into training set, validation set and test set, distinguish background texture area and defect area by pixel-level annotation, generate corresponding binarized annotation mask, expand sample diversity by data augmentation, simulate image differences of actual working conditions on the production line, and form a standardized training dataset. 1000 defect-free standard texture images of each of five types of spunlace filter media (PP, PET, viscose, polyester, and polypropylene) were collected (collection conditions: production line speed 100-200 m / min, light intensity 500-1000 lux). 800 inspection images of each material containing minor defects (pinholes <0.1 mm, fiber agglomeration, impurity embedding) were also collected, covering different production conditions (variable speed operation, slight vibration). Pixel-level annotations were performed using the LabelMe annotation tool to accurately delineate the boundaries between background texture areas and defect areas, generating corresponding binary annotation masks (single-channel image, defect area pixel value 255, background 0). Data augmentation techniques were employed to expand sample diversity: random rotation (-10°~+10°), horizontal / vertical mirror flipping, grayscale perturbation (±5%), Gaussian noise addition (variance 0.01), and random cropping (512×512 pixels), increasing the total sample size from 9000 to 15000 images. The training set (10,500 images), validation set (3,000 images), and test set (1,500 images) were randomly divided in a 7:2:1 ratio, ensuring uniform distribution of samples for each material and defect type (bias ≤3%). All images were normalized (pixel values ​​scaled to [0,1]) and Z-Score normalization was used to eliminate the influence of lighting differences, forming a standardized training dataset. The data format was unified as PNG (images) and JSON (annotation information).

[0055] Step S1312: With the goal of balancing real-time performance and detection accuracy, depthwise separable convolution is used to replace standard convolution to reduce the number of model parameters. The model input size is set to match the resolution of the filter material image acquired by the line scan camera. The encoder-decoder structure is retained and redundant network layers are removed to construct a lightweight U-Net texture stripping model. The model output is a texture and defect separation feature map. To balance real-time performance and detection accuracy, a lightweight U-Net texture stripping model was constructed. Depthwise separable convolutions (number of groups = number of input channels, 3×3 kernels) were used instead of standard convolutions, reducing the number of parameters by 70% (from 5 million to 1.5 million) and improving inference speed to ≤20ms / frame. The model was adapted to the resolution of filter images (4096×2048 pixels) acquired by a line scan camera (MV-CH080-10GM), with an input size of 512×512 pixels (large images were processed using a sliding window), and an output of 2-channel feature maps (channel 1: texture features, channel 2: defect features). The model retains the encoder-decoder core structure of U-Net, with a 4-layer encoder (convolution + batch normalization + ReLU + max pooling) and a 4-layer decoder (transposed convolution + batch normalization + ReLU). Redundant fully connected layers and Dropout layers in traditional U-Net were removed (Dropout was only retained in layers 3 and 4 of the encoder, with a rate of 0.2). The model was quantized using TensorRT to reduce its size further (from 20MB to 5MB), ensuring that the model could run efficiently on an industrial PC (CPU i5-10400, GPU RTX3060) and meet the production line's 50fps detection frame rate requirement.

[0056] Step S1313: Construct an encoder module by using multiple layers of lightweight convolutional units in series. Extract filter material texture features using convolutional kernels with different receptive fields. The bottom convolutional units extract shallow fine-grained texture features. The high-level convolutional units aggregate global information through downsampling and learn the deep semantic texture features corresponding to the material. Each convolutional unit is followed by a batch normalization layer and an activation function layer in sequence. Temporarily store the multi-scale feature maps output by each level of the encoder. A lightweight convolutional unit (CLU) with four cascaded layers is constructed as the encoder module. Each CLU consists of a depthwise separable convolution (3×3 or 5×5), a batch normalization layer, and a ReLU activation function. Layers 1 and 2 are the bottom-level CLUs, using 3×3 kernels (3×3 receptive field), a stride of 1, and padding of 1. They extract shallow, fine-grained features of the filter material texture (such as fiber edges and local interweaving structures), with output feature map sizes of 256×256×64 and 128×128×128, respectively. Layers 3 and 4 are the high-level CLUs, using 5×5 kernels (5×5 receptive field), a stride of 2, and padding of 2. They aggregate global information through downsampling to learn the deep semantic texture features corresponding to the material (such as texture density, overall interweaving patterns, and grayscale distribution characteristics), with output feature map sizes of 64×64×256 and 32×32×512, respectively. Each convolutional unit is followed by a batch normalization layer (momentum 0.99, epsilon = 1e-5) to accelerate model convergence and reduce the risk of overfitting; the ReLU activation function introduces nonlinearity to improve feature extraction capability. Multi-scale feature maps (64 channels, 128 channels, 256 channels, 512 channels) output from each level of the encoder are temporarily stored in a buffer to provide feature input for the skip connections of the decoder.

[0057] Step S1314: Upsample the deep feature map output by the encoder through the transposed convolutional layer to restore the image size. After each upsampling, the decoder features and the multi-scale feature map of the corresponding level of the encoder are concatenated through a skip connection structure to fuse shallow detail features and deep semantic features. The decoder outputs pixel-level texture prediction results through a 1×1 convolutional layer. The decoder consists of 4 layers of transposed convolutional units. Each transposed convolutional unit uses a 2×2 transposed convolutional kernel with a stride of 2 and padding of 1. It upsamples the deep feature map output by the encoder to gradually restore the image size: 32×32×512→64×64×256→128×128×128→256×256×64→512×512×2. After each upsampling stage, a skip connection structure is used to concatenate the current layer feature map of the decoder with the multi-scale feature map of the corresponding layer of the encoder: the first layer of the decoder (64×64×256) is concatenated with the third layer of the encoder (64×64×256), outputting 64×64×512; the second layer of the decoder (128×128×128) is concatenated with the second layer of the encoder (128×128×128), outputting 128×128×256, and so on, to achieve the fusion of shallow detail features and deep semantic features, compensating for the information loss during downsampling. At the end of the decoder, a 1×1 convolutional layer (64 input channels, 2 output channels) maps the high-dimensional features to a 2-channel feature map. Channel 1 is the texture prediction result, and channel 2 is the defect prediction result, completing pixel-level texture prediction. The output feature map size is consistent with the input image (512×512).

[0058] Step S1315: Define a cross-entropy loss function to distinguish between background texture and defect areas, minimize pixel classification error, define a structural similarity loss function to preserve texture structure distribution features, constrain the model to learn the texture rules of real materials, and fuse the two loss functions through weighted coefficients to form the total loss function; The cross-entropy loss function is defined to accurately distinguish between background texture and defect regions. The formula is: L ce =-∑(y i log(p i )+(1-y i log(1-p) i ), where y i For the true labeling mask (0=background, 1=defect), p i To predict defect probabilities for the model, this loss function minimizes pixel classification error, improving the accuracy of separating texture from defects. A structural similarity (SSIM) loss function is defined to preserve the structural distribution features of the filter material texture; the formula is: L ssim =1-(2μ x μ y +C1)(2σ xy +C2) / [(μ x² +μ y² +C1)(σ x² +σ y² +C2)], where μ is the mean, σ is the variance, and σ xy For covariance, C1=0.01² and C2=0.03². This loss function constrains the model to learn the true material texture patterns and avoids excessive texture stripping leading to the loss of defect information. The two loss functions are fused by weighted coefficients, and the total loss function is: L total =0.7×L ce +0.3×L ssim The weighting coefficients are determined through validation set optimization (the classification accuracy is optimal when the cross-entropy weight is 0.7, and the texture structure fidelity is optimal when the SSIM weight is 0.3), ensuring that the model balances classification accuracy and texture structure integrity.

[0059] Step S1316: Select an adaptive optimizer to configure training parameters, set the initial learning rate and learning rate decay mechanism, use mini-batch sample input to perform forward propagation, calculate the prediction error through the total loss function, update the network parameters of the encoder and decoder based on the backpropagation algorithm, so that the model learns the texture distribution pattern of different filter materials; The Adam adaptive optimizer was selected for training parameters, set as follows: β1=0.9, β2=0.999, epsilon=1e-8, with an initial learning rate of 0.001. A cosine annealing learning rate decay mechanism was employed, with the learning rate decreasing with each iteration according to the formula lr=lr0×(1+cos(π×epoch / epochs)) / 2 (lr0=0.001, epochs=200), ensuring rapid convergence in the early stages of training and precise fine-tuning in the later stages. A mini-batch input method was used, with a batch size of 32 (adapting to 6GB of GPU memory). Each iteration input consisted of 32 images of 512×512 pixels and their corresponding labeled masks. During forward propagation, the images were input into the model, and the predicted feature maps were output. The prediction error was calculated using the total loss function. The gradient of the loss function with respect to the encoder and decoder weights was calculated using the backpropagation algorithm (chain rule), with a gradient clipping threshold of 1.0 to avoid gradient explosion. The model parameters are updated once per iteration. During the iteration process, filter material samples of different materials and defect types are continuously input, so that the model can gradually learn the texture distribution rules corresponding to various materials.

[0060] Step S1317: Monitor the changes in texture stripping accuracy and loss value in real time through the validation set. When the loss value is stable and the separation accuracy meets the standard, the model is determined to have converged. A random deactivation layer and an early stopping mechanism are introduced. During training, the model performance was evaluated every 10 iterations using a validation set (3000 images). Monitoring metrics included texture stripping accuracy (number of correctly separated pixels / total number of pixels × 100%), defect retention rate (number of correctly retained defect pixels / number of true defect pixels × 100%), and loss value (L). total The convergence criteria are set as follows: if the loss value fluctuation on the validation set is ≤0.001 for 5 consecutive rounds, and the texture stripping accuracy is ≥95% and the defect retention rate is ≥98%, the model is considered to have converged, and training is stopped. To suppress overfitting, Dropout layers are introduced in layers 3 and 4 of the encoder, with an inactivation rate of 0.2, randomly discarding some neurons to reduce the model's dependence on training samples; at the same time, an early stopping mechanism is introduced, with a patience value of 10 rounds. If the loss value on the validation set does not decrease for 10 consecutive rounds, training is automatically stopped and the current optimal model parameters (the parameters with the smallest loss value) are saved.

[0061] Step S1318: Evaluate the model's texture stripping effect, defect preservation integrity, and inference speed using a test set, verify the model's learning effect on texture distribution patterns, complete lightweight fine-tuning, and solidify the model parameters to obtain a texture stripping model that can be used for online detection.

[0062] The model performance was evaluated using a test set (1500 images, including 5 materials and 3 types of minor defects). Test metrics included texture stripping accuracy, defect preservation integrity, and inference speed. Texture stripping accuracy was calculated using pixel-level IOU (IOU = predicted texture area ∩ true texture area / predicted texture area ∪ true texture area), with a requirement of ≥95%. Defect preservation integrity was calculated using the defect area deviation rate (|predicted defect area - true defect area| / true defect area × 100%), with a requirement of ≤3%. Inference speed was tested on an industrial PC (CPU i5-10400, GPU RTX3060), with a requirement of ≤20ms inference time per frame. Test results showed an average texture stripping accuracy of 96.8%, a defect area deviation rate of 2.1%, and an inference time of 15.3ms, meeting the design requirements. This was also tested for a slightly underfit material (polypropylene).

[0063] Step S132: Real-time acquisition of the surface image of the filter material to be tested, extraction of the texture features of the current filter material through a pre-trained texture feature recognition network, matching of adaptive texture templates from a preset texture template library based on the extracted texture features, and dynamic texture separation logic. First, the background texture and potential defect areas are distinguished by an adaptive threshold segmentation algorithm. Then, the background texture is stripped by template matching and pixel-level subtraction operation. A defect retention threshold is set to fully retain the information of minute defects. The preliminary image after background stripping is output. Real-time images of the filter material surface under test were acquired using a line scan camera (MV-CH080-10GM) (resolution 4096×2048 pixels, frame rate 50fps). Each frame was preprocessed by noise reduction (Gaussian filtering 3×3) and grayscale conversion (weighted average R:G:B=0.299:0.587:0.114) before being input into a pre-trained MobileNetV2 texture feature recognition network (3.4 million parameters, inference time ≤5ms). The core texture features of the current filter material were extracted: texture density (pixel percentage), interlacing angle (hibogram peak angle), grayscale variance (0-255), and texture uniformity (entropy value 0-8), and the features were normalized to the [0,1] interval. Based on the extracted features, a normalized cross-correlation algorithm was used to match suitable texture templates from a preset texture template library (stored according to 5 materials and 20 texture types, containing 1000 templates). The matching threshold was set to 0.85 (matching degree ≥0.85 is considered valid). The dynamic texture separation logic is adopted: first, the background texture and potential defect areas are distinguished by the Otsu adaptive threshold segmentation algorithm (based on real-time image grayscale variance dynamic threshold adjustment); then, the background texture is stripped by template matching and pixel-level subtraction operation (image-template), and a defect retention threshold (grayscale difference ≥ 3) is set to prevent the stripping of potential small defect areas.

[0064] Step S133: Perform wavelet multi-scale decomposition on the preliminary image after background stripping to obtain low-frequency components and multiple high-frequency components. Introduce an attention mechanism module into the high-frequency components to weight and enhance the defect-related features and suppress the high-frequency components corresponding to residual noise. Reconstruct the enhanced high-frequency components and low-frequency components to obtain a small defect feature image with enhanced edges and gray-level differences. The initial image after background removal is subjected to a 3-layer wavelet multi-scale decomposition (using the db4 wavelet basis). The decomposition yields one low-frequency component (containing residual overall texture information) and three high-frequency components (corresponding to minute defect edges, gray-level differences, and other details in the horizontal, vertical, and diagonal directions, respectively). CBAM (Convolutional Block Attention Module) is introduced into the high-frequency component processing. This module first calculates the channel weights of each high-frequency component through a channel attention submodule (global average pooling + fully connected layer) (increasing the weights of defect-related channels by 2-3 times). Then, a spatial attention submodule (max pooling + average pooling + convolution) locates the spatial position of the defect region, weighting and enhancing the defect-related features in the high-frequency components while suppressing the high-frequency components corresponding to residual noise (reducing the weight of noise regions to below 0.3). The enhanced three high-frequency components and the low-frequency components are then reconstructed using inverse wavelet transform to obtain a feature image of minute defects with enhanced edges and gray-level differences.

[0065] Step S134: Synchronously acquire full-band spectral data of the filter material using a hyperspectral camera, perform smoothing and noise reduction and band normalization preprocessing on the spectral data, and use a combination of spectral angle matching algorithm and principal component analysis to screen out characteristic spectral bands corresponding to the filter material and internal defects, remove redundant bands and establish a characteristic spectral band index. A headwall hyperspec camera simultaneously acquired full-band spectral data (400-1700nm, wavelength resolution 5nm, 261 bands) of the filter material, maintaining a stable light source intensity (850nm near-infrared light source, 30W power) during acquisition. Preprocessing of the spectral data included: using a Savitzky-Golay smoothing algorithm (11-point window) to remove noise and reduce spectral fluctuations (variance ≤5% after smoothing); and using a Min-Max normalization algorithm ((spectral value - minimum value) / (maximum value - minimum value)) to standardize the spectral values ​​of each band to the [0,1] interval, eliminating the influence of light intensity differences. Feature bands were selected using a combination of spectral angle matching algorithm and principal component analysis (PCA): The spectral angle matching algorithm calculated the angle between each band and the standard spectrum of the filter material (five pre-stored standard spectra of materials), and bands with an angle ≤0.1 rad were considered material-related feature bands; PCA reduced the dimensionality of the entire band data, and selected the bands corresponding to the top three principal components (cumulative contribution rate ≥90%) as internal defect feature bands. Redundant bands unrelated to the material and defects were removed (removal rate ≥50%), and finally 60-80 feature bands were retained. A feature spectral band index was established (recording the detection target corresponding to each band: material distribution / internal defect type).

[0066] Step S135: Convert the filtered feature spectral band data into corresponding spatial grayscale images. Adopt an adaptive weight fusion strategy, combine the texture complexity and defect distribution of the micro-defect feature images, dynamically adjust the fusion weight of each feature spectral grayscale image and the spatial image, and integrate multiple feature spectral grayscale images and spatial images into an enhanced image with both spatial details and spectral features through a pixel-level weighted fusion algorithm. The selected 60-80 characteristic spectral band data were converted into corresponding spatial grayscale images (512×512 pixels resolution) according to the formula "band grayscale value = spectral reflectance × 255". Each characteristic band corresponds to one grayscale image, reflecting the material or defect distribution characteristics of the filter material under that band (e.g., the 850nm band image highlights the material distribution, and the 1300nm band image highlights internal layering defects). An adaptive weight fusion strategy was adopted to calculate the texture complexity (entropy value 0-8) and defect distribution density (defect pixel ratio 0-5%) of the small defect characteristic images: the spatial image weight was increased for textured complex regions (entropy value > 5) (α=0.7) to avoid interference from spectral information; the characteristic spectral image weight was increased for small defect regions (density > 1%) (α=0.3) to enhance the distinction between material and defects; and the remaining regions were given a balanced weight (α=0.5). By using a pixel-level weighted fusion algorithm, multiple feature spectral grayscale images and spatial images are integrated according to the formula "fusion pixel = α × spatial pixel + (1-α) × mean spectral pixel value" to generate an enhanced image that combines spatial details (defect edges, shape) and spectral features (material composition, internal structure).

[0067] Step S136: Based on motion-blurred samples and corresponding clear micro-defect samples at different operating speeds of the production line, a dynamic generative adversarial deblurring model with a conditional generative adversarial network architecture is constructed. The model input layer includes motion-blurred images and production line speed labels. The current filter material conveying speed is obtained in real time and used as the speed label input to the model. The model automatically generates a deblurring kernel at the corresponding speed and performs real-time deblurring processing on the enhanced image. Based on motion-blurred samples (1000 sets for each speed, blur level 1-10 pixels) at production line speeds of 50-200 m / min and corresponding clear samples of minor defects, a dynamic deblurring model with a conditional generative adversarial network (CGAN) architecture is constructed. The model has two input channels: one channel receives a 256×256 pixel motion-blurred image (3-channel RGB), and the other channel receives quantized production line speed labels (discretely divided into 10 levels, one-hot encoded as a 10-dimensional vector). The generator adopts an encoder-decoder architecture (4-layer encoder + 4-layer decoder), embedding the speed labels into a high-dimensional feature vector (256-dimensional), which is then concatenated with the multi-scale feature channels of the image. The decoder has two output branches: the main branch generates a clear image, and a dedicated branch outputs a 7×7 dynamic deblurring kernel (containing blur direction parameters 0-180° and scale parameters 1-10 pixels). The discriminator is a conditional discrimination structure, receiving the generated image / real image and speed labels as dual inputs, and outputting a dual judgment result of realism and speed matching degree. The model was trained using Wasserstein adversarial loss, L1 reconstruction loss, and MSE velocity matching loss (weighted ratio 0.3:0.5:0.2), achieving a deblurring accuracy of ≥92% after 200 iterations. During online detection, velocity data was collected in real-time via a production line encoder (accuracy 0.1 m / min), converted into velocity labels, and input into the model to generate an adaptive deblurring kernel. This kernel was then used to deblur the enhanced image (time ≤15 ms).

[0068] Step S1361: Collect motion blur images of the spunlace filter media at different production line conveying speeds and clear standard images of the corresponding positions. Assign production line speed as a classification label to each group of blur images. Expand the sample through data augmentation to simulate the blur effect caused by the variable speed operation and shaking of the production line. Divide the training set, validation set and test set according to a preset ratio to complete the dataset standardization preprocessing. Within a production line speed range of 50-200 m / min (step size 10 m / min), motion-blurred images of the spunlace filter media were acquired using a line scan camera (MV-CH080-10GM), with 1000 images acquired per speed. Simultaneously, clear standard images (256×256 pixels resolution) were acquired at the corresponding stationary positions to ensure complete spatial consistency between the blurred and clear images. Each group of blurred images was assigned a corresponding production line speed as a classification label (e.g., 50 m / min was labeled 0, 200 m / min was labeled 9, for a total of 10 levels). Data augmentation techniques were used to expand the sample to 20,000 images: random blur kernel perturbation (blur direction ±10°, scale ±2 pixels), speed interpolation transformation (generating intermediate speed samples such as 55 m / min and 65 m / min), and adaptive grayscale adjustment (±10%) to simulate the complex blurring effects caused by variable speed operation and jitter on the production line. The training set (14,000 images), validation set (4,000 images), and test set (2,000 images) were divided in a 7:2:1 ratio, ensuring uniform distribution of samples across different speed levels and blur levels (deviation ≤3%). All images were normalized (pixel values ​​scaled to [0,1]), and Z-Score normalization was used to eliminate illumination differences, completing the dataset standardization preprocessing. The data format was unified to PNG.

[0069] Step S1362: Construct a dual-main-body model architecture containing a generator and a discriminator, embed the production line speed as a conditional constraint variable into the network framework, set the model to have two input channels to receive motion-blurred images and quantized production line speed labels respectively, and set a deblurring kernel generation branch and a clear image reconstruction branch at the model output end to realize customized deblurring processing based on speed conditions. A Conditional Generative Adversarial Network (CGAN) architecture was constructed, consisting of a generator (G) and a discriminator (D) as the basic components. Production line speed was embedded as a conditional constraint variable within the network framework. The model had two input channels: the first channel was an image input channel, receiving a 256×256×3 motion-blurred image (RGB format); the second channel was a speed label input channel, receiving quantized 10-dimensional one-hot encoded production line speed labels (corresponding to 10 levels from 50-200 m / min). The model output consisted of two branches: a deblurring kernel generation branch outputting a 7×7×1 dynamic deblurring kernel (single-channel floating-point feature map), and a sharp image reconstruction branch outputting a 256×256×3 sharp image (RGB format). The generator and discriminator achieved collaborative optimization through adversarial training. The conditional constraint variable (speed label) was used throughout the training process, ensuring that the generator could generate customized deblurring kernels and sharp images based on different speed labels. This achieved accurate deblurring based on speed conditions, avoiding the poor performance caused by a single deblurring kernel adapting to all speeds.

[0070] Step S1363: The generator backbone is built using an encoder-decoder architecture. The production line speed label is encoded and its features are embedded, and converted into a high-dimensional conditional feature vector. The encoder performs multi-scale feature extraction on the motion-blurred image. In the middle layer of the encoder, the speed conditional feature vector and the image features are fused at the channel level. The decoder has two branches. The main branch is used for image sharpening and reconstruction. The branch outputs a dynamic deblurring kernel that matches the current production line speed. The dynamic deblurring kernel includes motion blur direction and blur scale parameters. The generator employs an encoder-decoder architecture as its backbone. The encoder contains four convolutional layers (3×3 convolutions, stride 2, padding 1), with output feature map sizes of 128×128×64, 64×64×128, 32×32×256, and 16×16×512, respectively. These layers are used for multi-scale feature extraction of motion-blurred images, capturing blur features, texture features, and defect features layer by layer. The input production line speed label (10-dimensional one-hot encoded) is embedded using a fully connected layer, transforming it into a 16×16×512 high-dimensional conditional feature vector, which matches the size of the image feature map output from the last layer of the encoder. In the intermediate layer of the encoder (16×16×512), the speed conditional feature vector is fused with the image features through channel-level concatenation, enabling the network to learn the correlation between different speeds and motion-blurred features. The decoder contains four transposed convolutional layers (2×2 transposed convolution, stride 2, padding 1), divided into two branches: the main branch (output 256×256×3) gradually restores the image size through transposed convolution, achieving image sharpening reconstruction; the dedicated branch (output 7×7×1) outputs a dynamic deblurring kernel that perfectly matches the current production line speed through three layers of 1×1 convolution and the Sigmoid activation function. The kernel parameters include motion blur direction (0-180°) and blur scale (1-10 pixels).

[0071] Step S1364: Design a discriminator using a conditional discriminant structure. Set up dual input interfaces to receive the clear image and the real clear image output by the generator respectively. At the same time, the production line speed label is used as an auxiliary condition input to the discriminator. The discriminator extracts image features through convolution and downsampling. After fusing the image features and speed condition features, it outputs dual discrimination results, which are the image authenticity discrimination result and the deblurring effect and speed condition matching discrimination result, respectively. The generator outputs either a clear image (256×256×3) or a truly clear image (256×256×3), along with the corresponding production line speed label (10-dimensional one-hot encoding). The discriminator has dual input interfaces: the image input interface extracts image features through four convolutional layers (3×3 convolution, stride 2, padding 1), outputting a 1×1×256 feature vector; the speed label input interface embeds the 10-dimensional label into a 1×1×256 feature vector through two fully connected layers. The image feature vector and the speed condition feature vector are fused by element-wise multiplication. The fused vector is then input to two convolutional layers (3×3 convolution, stride 1, padding 0) and one fully connected layer, outputting dual discrimination results: the first result is a binary classification probability (0-1), determining whether the input image is a truly clear image (approaching 1) or a generated clear image (approaching 0); the second result is a binary classification probability (0-1), determining whether the image deblurring effect matches the input speed condition (approaching 1 indicates a match). By optimizing the generator's direction through dual-discrimination constraints, we ensure that the generated clear images are both realistic and highly adapted to speed conditions.

[0072] Step S1365: Construct adversarial loss, image reconstruction loss and velocity conditional matching loss, and form a total loss function by weighted fusion. The adversarial loss is used to reduce the distribution difference between the generated image and the real clear image, the image reconstruction loss is used to constrain the image detail restoration degree, and the velocity conditional matching loss is used to ensure the adaptability of the dynamic deblurring kernel and the input velocity label. Three types of constraint losses are constructed and weighted to form the total loss function. The adversarial loss uses the Wasserstein GAN loss type, which uses a discriminator to determine the distribution difference between the generator's deblurred image and the real sharp image, constraining the generated image to maintain consistency with the real image in texture details and grayscale distribution, thus reducing the distribution gap between the two. The image reconstruction loss uses L1 loss, which calculates the mean pixel-level difference between the generated image and the real image, focusing on constraining the restoration of details such as filter material defect edges and minute textures, avoiding detail loss or over-smoothing during the deblurring process. The speed condition matching loss uses mean squared error loss, which compares the generated dynamic deblurring kernel with the real blurring kernel corresponding to the current speed label, ensuring that key parameters such as the blur direction and blur scale of the deblurring kernel are accurately adapted to the actual operating speed of the production line. The weighting coefficients of the three types of losses are determined by a validation set grid search, with the adversarial loss weighted at 0.3, the image reconstruction loss weighted at 0.5, and the speed condition matching loss weighted at 0.2. These weighted coefficients are then fused to form the total loss function. When the total loss function value converges to below 0.05, the structural similarity index of the generated image is ≥0.92, which meets the detail preservation requirement for defect detection.

[0073] Step S1365: Initialize the network parameters of the generator and discriminator, select the adaptive optimizer to configure the training parameters, input the training set samples into the model in batches, the generator generates a dynamic deblurring kernel based on the blurred image and velocity label and completes the image deblurring, the discriminator performs dual discrimination, and updates the generator and discriminator parameters alternately through the backpropagation algorithm based on the total loss function. The Xavier initialization method was used to initialize the parameters of the fully connected and convolutional layers of the generator, ensuring consistent variance between the input and output of each layer. The convolutional layer parameters of the discriminator were initialized using the He initialization method, with the bias term uniformly initialized to 0.01, providing a stable starting point for model training. The Adam adaptive optimizer was selected to configure the training parameters, with the initial learning rate of the generator set to 0.001 and the discriminator set to 0.0001. The momentum parameters β1=0.9, β2=0.999, and the numerical stability parameter epsilon=1e-8. Training set samples were input into the model in batches of 32 images. After receiving the motion-blurred image and its corresponding velocity label, the generator first generated a suitable dynamic deblurring kernel, then used this kernel to deblur the image, outputting clear image candidates. The discriminator simultaneously received the clear image, the truly clear image, and the corresponding velocity label from the generator, performing dual discrimination: first, determining whether the image is truly clear; and second, determining the fit between the deblurring kernel and the velocity label. Based on the error value calculated by the total loss function, the network parameters of the generator and discriminator are updated alternately through the backpropagation algorithm. A full parameter update is completed in each iteration, and the parameters are continuously optimized during the iteration process to reduce the total loss.

[0074] Step S1366: Monitor the changes in deblurring accuracy, speed matching degree and loss value in real time through the validation set. When the indicators tend to stabilize, the model is determined to have converged. A regularization layer and an early stopping mechanism are introduced. During training, the model performance is monitored in real time every 10 iterations using a validation set. Monitoring metrics include deblurring accuracy, velocity matching, and total loss. Deblurring accuracy is evaluated using the structural similarity index and peak signal-to-noise ratio (PSNR), requiring a structural similarity index ≥ 0.9 and a PSNR ≥ 35 dB. Velocity matching is calculated by comparing the parameters of the generated deblurring kernel and the real blur kernel, requiring a similarity ≥ 90%. The total loss must remain stable below 0.05. A convergence criterion is set: the model is considered converged when the fluctuation range of all three metrics on the validation set is ≤ 3% for five consecutive iterations, with no significant upward or downward trend. To suppress overfitting, a Dropout regularization layer is embedded after the generator's encoder intermediate layer and the discriminator's convolutional layer, with a dropout rate set to 0.2, randomly discarding some neurons to reduce the model's dependence on training samples. An early stopping mechanism is also introduced, with a patience value of 10 iterations. If the total loss on the validation set does not decrease for 10 consecutive iterations and shows an upward trend, training is automatically stopped and the current optimal model parameters are saved to ensure the model has good generalization ability.

[0075] Step S1367: Solidify the trained model and connect it to the detection system. The current filter material conveying speed is collected in real time through the production line encoder, converted into a speed label that the model can recognize and input into the model. The model calls the learned speed-blur kernel mapping relationship to generate a dynamic deblur kernel that matches the current running speed, providing parameters for image deblurring processing. The trained model is quantized and solidified using TensorRT, converting it into an engine format directly usable by industrial inspection systems. Redundant computational layers from the training process are eliminated, reducing the model size to below 5MB and ensuring inference efficiency. The current filter material conveying speed is acquired in real-time via a production line encoder (1000 lines of pulse resolution), with the acquisition frequency synchronized with production line operation and a speed measurement accuracy of ±0.1m / min. The acquired speed values ​​are quantized into 10 levels according to preset intervals, converting them into 10-dimensional one-hot encoded speed labels recognizable by the model. For example, 50-65m / min corresponds to label 0, 66-80m / min corresponds to label 1, and so on. The inspection system synchronously inputs the speed labels and the blurred image to be processed into the solidified model. The model utilizes the speed-blur kernel mapping learned during training to quickly generate a dynamic deblurring kernel that precisely matches the current operating speed. This kernel includes key parameters such as blur direction (0-180°) and blur scale (1-10 pixels), providing customized parameter support for subsequent image deblurring processing. The kernel generation time is ≤5ms.

[0076] Step S1368: Calculate the sharpness index and defect edge integrity index of the deblurred image in real time. If the index does not meet the standard, trigger the model fine-tuning process and update the network parameters based on the current speed and image data.

[0077] Clarity and Defect Edge Integrity Indicators. Clarity is determined by calculating the grayscale variance gradient and information entropy value of the image, requiring a variance gradient ≥ 0.8 and an information entropy value ≥ 6.5. Defect edge integrity is evaluated by statistically analyzing edge continuity rate and the number of break points, requiring an edge continuity rate ≥ 95% and ≤ 2 break points per defect edge. Standards for achieving these standards are set: both standards must be met to be considered satisfactory; failure to meet either standard triggers a model fine-tuning process. After the fine-tuning process begins, the system automatically collects 50 pairs of blurred-clear images at the current speed as small-batch fine-tuning data. A low learning rate of 0.0001 is used, updating only the deblurring kernel generation branch parameters of the generator, iterating 20 times to complete the fine-tuning. During fine-tuning, changes in the indicators are monitored in real time until both indicators meet the standards, at which point updates cease, ensuring the model can adapt to the attenuation of deblurring effect caused by production line speed fluctuations or changes in filter material characteristics.

[0078] Step S137: Use a multi-scale edge detection algorithm to extract defect edge information from the deblurred image, introduce a gradient boosting algorithm to enhance the edge gradient, and use pixel-level gradient aggregation technology to repair edge breakage and blurring problems; set an adaptive adjustment rule for sharpening intensity, adjust the sharpening parameters according to the defect size and edge complexity, and restore the true boundary of the tiny defect.

[0079] A multi-scale edge detection algorithm is used to extract defect edge information from the deblurred image. A five-level Gaussian pyramid is constructed to decompose the image into multiple scales. Convolutional kernels of different sizes (3×3, 5×5, and 7×7) are used for edge detection at each scale. The detection results from each scale are fused to obtain a preliminary edge map, avoiding the omission of minute edges or false detection of coarse edges caused by single-scale detection. A gradient boosting algorithm is introduced to weight and enhance the gradient magnitude in the preliminary edge map. The gradient magnitude of the defect edge region is multiplied by an enhancement coefficient of 1.5-2.0, while the gradient magnitude of the background region is multiplied by a suppression coefficient of 0.3-0.5, highlighting the gradient difference between the defect edge and the background. Pixel-level gradient aggregation technology is used to repair edge breaks and blurred areas. The gradient direction and magnitude of adjacent pixels at the break are calculated, and linear interpolation is used to fill in the broken pixels, restoring edge continuity. The sharpening intensity adaptive adjustment rule is set: the defect pixel area is determined by the defect size detection algorithm. For tiny defects with an area < 0.01 mm², the sharpening intensity is adjusted to 1.8-2.0, and for defects with an area ≥ 0.01 mm², it is adjusted to 1.2-1.5. At the same time, the sharpening parameters are dynamically corrected according to the edge complexity (edge ​​pixel density). If the complexity is high, the sharpening intensity is reduced by 0.2-0.3 to avoid over-sharpening and producing false edges. In the end, the true boundary of tiny defects is accurately restored with a boundary restoration accuracy ≤ 0.05 mm.

[0080] Step S140: Input surface defect features, internal structure features, material distribution features and functional features into the multimodal feature fusion model, establish the mapping relationship between each feature through the cross-modal association mechanism, and complete the comprehensive judgment of surface defects, internal defects and functional defects of filter material according to the preset multi-dimensional defect judgment rules. Comprehensive defect identification of filter media is achieved through multimodal feature fusion and multidimensional rule-based judgment. Four modalities of data—surface defect features, internal structural features, material distribution features, and functional features—are simultaneously input into a multimodal feature fusion model. The model incorporates a cross-modal correlation mechanism, establishing mapping relationships between surface defects and internal structure, and between material distribution and functional indicators through an attention mechanism, thereby uncovering the inherent correlation patterns among different modal features. Based on a pre-defined multidimensional defect judgment rule library, initial morphological defects are first judged based on surface and internal structural features, and then performance defects are verified by combining material and functional features, achieving a hierarchical comprehensive judgment of surface defects, internal defects, and functional defects. This judgment method overcomes the limitations of single-modal detection and can identify hidden defects that cannot be detected by a single feature alone.

[0081] Step S141: Quantify and extract the surface defect features, internal structural features, material distribution features and functional features of the spunlace filter media respectively. Normalize each modal feature using a standardization method to eliminate dimensional differences. Based on the label information after spatiotemporal registration, stitch the four modal features of the same detection area according to dimensions to form a multimodal original feature matrix. Four modal features were quantitatively extracted: surface defect features (defect area, edge gradient, grayscale contrast, etc.), internal structure features (ultrasonic amplitude, interlayer difference, density distribution, etc.), material distribution features (spectral intensity, uniformity, etc.), and functional features (air permeability, filtration efficiency, tensile strength, etc.). Z-Score normalization was used to normalize each modal feature, mapping all feature values ​​to a range with a mean of 0 and a variance of 1, eliminating differences in dimensions and numerical ranges. Based on the unified label information after spatiotemporal registration, the feature vectors of the four modalities in the same detection area were concatenated in a fixed-dimensional order to form a multimodal original feature matrix. This ensured that each row of matrix data corresponded to the same detection position on the filter material, providing a standardized and aligned data foundation for subsequent fusion model input.

[0082] Step S142: Build a Transformer-based cross-modal feature fusion encoder architecture. Set up four independent modal embedding layers to convert modal features of different dimensions into a unified dimensional feature vector. After the embedding layer, connect the position encoding module to add spatial position information. The core layer adopts a cross-modal multi-head attention module to capture the semantic association between modalities through multi-head parallel computing. The serial layer normalization and feedforward neural network perform nonlinear transformation and dimensionality optimization on the fused features and output a unified multimodal fusion feature vector. A Transformer-based cross-modal feature fusion encoder architecture was constructed, featuring four independent modal embedding layers. These layers convert features from four different modalities into a unified-dimensional feature vector, eliminating dimensionality differences. A learnable positional encoding module follows the embedding layers, adding spatial location identifiers to each modal feature and clarifying the spatial relationships between modalities. The core layer employs a cross-modal multi-head attention module, using parallel multi-head computation to capture semantic relationships between surface defects, internal structures, material distribution, and functional features. A serial encoder layer normalizes and incorporates a feedforward neural network, performing nonlinear transformations and dimensionality optimization on the fused features to enhance feature representation capabilities. The final output is a unified multimodal fusion feature vector with fixed dimensions, achieving deep fusion of the four modalities and providing high-dimensional, strongly correlated feature support for defect identification.

[0083] Step S1421: Based on the input specifications of the surface defect characteristics, internal structural characteristics, material distribution characteristics and functional characteristics of the spunlace filter media, define the original input dimensions of each modal feature, and complete the feature vector formatting and organization based on the number of parameters of each modal feature, so that the four types of modal features are input into the model in the form of independent vectors. Based on the actual input specifications of the four modal features, the original input dimensions of each modality are defined: surface defect features are set to 8 dimensions, internal structure features to 10 dimensions, material distribution features to 6 dimensions, and functional features to 4 dimensions. Based on the number of parameters for each modal feature, the original data is formatted and organized, invalid parameters and outliers are removed, missing data is filled in, and each feature is organized into a fixed-length independent one-dimensional vector. The four types of vectors are input into the fusion model in independent channels without pre-mixing, preserving the original independence of each modal feature. This ensures that subsequent embedding layers can perform targeted mapping on each modal feature separately, providing a standardized input format for unified dimension transformation and cross-modal interaction.

[0084] Step S1422: Construct four independent modality embedding layers. Each embedding layer uses a single-layer fully connected linear transformation network to perform dimension mapping calculation on the original input vector of the corresponding modality, and uniformly transform the modality features of the four different dimensions to the preset high-dimensional common feature space, and output standardized modality feature vectors with consistent length. Four independent modality embedding layers are constructed, each corresponding to one class of modality features. All layers employ a single-layer fully connected linear transformation network for dimensionality mapping. For differentiated input vectors of 8, 10, 6, and 4 dimensions, linear weighting is used to uniformly transform the four types of features into a pre-defined 128-dimensional high-dimensional common feature space, outputting standardized modality feature vectors of identical length. The linear transformation weights are simultaneously optimized during model training, adaptively learning the importance weights of each modality feature. This eliminates fusion barriers caused by dimensionality differences while preserving the core feature information of each modality, enabling fair and effective interactive computation of the four types of features within the same dimensional system.

[0085] Step S1423: After the modality embedding layer, build a learnable modality position encoding module, assign a unique modality position identifier to the four types of modality features, generate a position encoding vector that matches the dimension of the standardized modality feature vector, and integrate the position encoding into the standardized modality features by adding the vector element by element, thus injecting spatial position information; A learnable modal position encoding module is built after the modal embedding layer, assigning unique position identifiers (0-3) to four modal features: surface defects, internal structure, material distribution, and functionality. Based on the 128-dimensional standardized feature vector length, dimension-matched learnable position encoding vectors are generated, with encoding parameters dynamically optimized during model training. Position encoding is integrated into the standardized modal features through element-wise vector addition, providing the model with modal sequence position information. This enables the model to identify spatial and sequential associations between different modalities, preventing fusion failure due to changes in modal order and improving the stability and accuracy of cross-modal association modeling.

[0086] Step S1424: Set the number of parallel heads for multi-head attention, use the four types of modal features after embedding and adding position encoding as the basic sequence for attention calculation, adopt a full-modal interactive calculation mechanism, use each type of modal feature as the query vector, and the other three types of modal features as the key vector and value vector, generate query, key, and value feature matrices through linear projection, calculate the inter-modal attention weight distribution head by head, perform weighted summation on the value feature matrix after normalization, and concatenate the output features of all parallel heads to generate cross-modal attention features; Eight parallel attention heads were set up, and four modal features with embedded and positionally encoded features were used as the basic sequence for attention calculation. A full-modal interaction mechanism was adopted, where each modal feature was used as a query vector, and the other three were used as key and value vectors, respectively. Corresponding feature matrices were generated through linear projection. The inter-modal attention weight distribution was calculated head-by-head, and after Softmax normalization, the value features were weighted and summed to extract cross-modal correlation features for each head. The output features of the eight parallel heads were concatenated along the channel dimension to generate cross-modal attention features that integrate global correlation information, comprehensively capturing the potential correlations between the four feature classes, such as the correspondence between internal defects and decreased air permeability, and surface pinholes and abnormal material distribution.

[0087] Step S1425: After the cross-modal multi-head attention module, a normalization structure is connected in series. Combined with the residual connection mechanism, the input features and output features of the attention module are added by residual addition. The feature vector numerical distribution is standardized by layer normalization to unify the feature mean and variance. A normalization structure is concatenated after the cross-modal multi-head attention module, combined with a residual connection mechanism. The input and output features of the attention module are directly added as residuals, preserving the original feature information and avoiding gradient vanishing. Layer normalization standardizes the numerical distribution of feature vectors, unifying the feature mean and variance, and suppressing numerical fluctuations and gradient explosion during training. This structure stabilizes the feature distribution, accelerates model convergence, and improves the robustness of cross-modal feature fusion, ensuring stable fused features even under fluctuating production line data, meeting the stability requirements of industrial online inspection.

[0088] Step S1426: Construct a two-layer fully connected feedforward neural network. The first layer maps the fused features to a high-dimensional space through linear transformation and introduces non-linear fitting ability with the GELU activation function. The second layer restores the features to the target output dimension through linear transformation. A random deactivation layer is embedded in the network to complete the non-linear transformation and dimensionality reduction of the fused features. A two-layer fully connected feedforward neural network is constructed. The first layer maps the 128-dimensional fused features to a 512-dimensional high-dimensional space through linear transformation, and introduces non-linear fitting capability by using the GELU activation function to explore deep correlations between features. The second layer restores the features to the 128-dimensional target output dimension through linear transformation, achieving dimensionality reduction. A random deactivation layer is embedded in the network with a deactivation rate of 0.1, randomly discarding some neuron outputs to suppress model overfitting. Through non-linear transformation and dimensionality optimization, the expressive power and generalization of the fused features are improved, providing high-quality feature input for subsequent defect judgment.

[0089] Step S1427: Combine the cross-modal multi-head attention module, layer normalization structure, residual connection mechanism and feedforward neural network into a single-level encoder unit. Stack multiple encoder units according to the filter material defect detection accuracy requirements. After the output of the multi-level encoder is connected to the global feature aggregation layer, the multi-level fused features are integrated into a single fixed-dimensional feature vector through adaptive pooling and channel compression operations, and a unified multi-modal fused feature vector is output. A single-level encoder unit is constructed by combining cross-modal multi-head attention, layer normalization, residual connections, and a feedforward neural network. Three levels of encoder units are stacked according to the defect detection accuracy requirements, progressively deepening the intermodal feature fusion. The outputs of the multi-level encoders are then fed into a global feature aggregation layer. Through adaptive average pooling and 1×1 convolutional channel compression, the multi-level fused features are integrated into a single 256-dimensional fixed-dimensional feature vector, outputting a unified multimodal fused feature vector. This structure balances deep feature fusion with computational efficiency, effectively extracting cross-modal correlation information while meeting the real-time inference requirements of the production line.

[0090] Step S1428: Solidify the forward propagation calculation process of the encoder architecture, set inference parameters adapted to the industrial production line, optimize matrix calculation efficiency, and ensure that the model can achieve high-precision cross-modal feature fusion while meeting the real-time requirements of online detection.

[0091] The forward propagation computation flow of the solidified encoder architecture is clearly defined, outlining the complete computational logic from feature input, embedding, position encoding, attention calculation, residual normalization, feedforward transformation to feature aggregation. Inference parameters adapted to industrial production lines are set, with a batch size of 1 and FP16 half-precision computation for inference accuracy, optimizing matrix multiplication and tensor operation efficiency. Through model quantization and operator fusion techniques, the single-frame inference time is controlled within 20 milliseconds, ensuring that the model achieves high-precision cross-modal feature fusion while meeting the real-time online detection requirements of 50 frames / second for production lines, and can be stably embedded into industrial inspection systems.

[0092] Step S143: Construct a labeled sample set containing the original feature matrix of multimodal modes and the corresponding defect labels, input the sample set into the fusion model, calculate the fusion feature vector through forward propagation, introduce the cross-modal association loss function to constrain the model to learn the association law between modes, update the model parameters through the backpropagation algorithm, and establish the cross-modal association mapping matrix; A labeled sample set is constructed, containing original multimodal feature matrices and corresponding defect labels. The labels cover defect type, level, and pass / fail status, and are manually labeled based on industry standards and production requirements. The sample set is input into the fusion model, and the fused feature vector is calculated through forward propagation. A cross-modal association loss function is introduced, with the optimization objective being the correlation between intermodal features and the defect judgment error, constraining the model to learn the association patterns of the four types of features. The gradient is calculated using the backpropagation algorithm, updating the parameters of the encoder embedding layer, attention layer, and feedforward layer, progressively optimizing the cross-modal association mapping matrix, enabling the model to have the inference capability of "single modal anomaly → associated multimodal verification".

[0093] Step S144: Review industry standards and production requirements, construct a judgment rule base in three categories: surface defects, internal defects, and functional defects, embed the rule base into the model output layer, set the rule matching priority, form a dual judgment logic of morphological judgment and performance verification, introduce the correlation confidence threshold, and trigger the supplementary judgment process for fuzzy boundary cases. This study analyzes industry standards and production quality requirements for filter media testing, constructing a quantitative judgment rule base categorized into three types: surface defects, internal defects, and functional defects. The size, amplitude, and performance thresholds for each defect type are clearly defined. The rule base is embedded in the model output layer, with rule matching priorities set. First, morphological defects are judged based on surface and internal structural features, then performance defects are verified based on material and functional features, forming a dual judgment logic of morphological judgment and performance verification. An association confidence threshold is introduced. For cases with ambiguous boundaries, such as slight fiber unevenness and minor defects, a supplementary judgment process is triggered when the confidence level exceeds the threshold, improving the accuracy of judging edge cases.

[0094] Step S145: Divide the labeled sample set into training set, validation set and test set according to the proportion, select an adaptive optimizer to configure training parameters, set learning rate and decay strategy, monitor the defect judgment accuracy and association mapping error of the validation set during training, adjust the model structure, introduce regularization layer and early stopping mechanism to avoid overfitting, evaluate model performance through test set, and supplement samples to strengthen training for weak links. The labeled sample set was divided into training, validation, and test sets in a 7:2:1 ratio. The Adam adaptive optimizer was used to configure training parameters, with an initial learning rate of 0.001, dynamically adjusted using a cosine decay strategy. During training, the validation set defect detection accuracy and association mapping error were monitored in real time, and structural parameters such as the number of attention heads and encoder layers were adjusted based on the error. L2 regularization and an early stopping mechanism were introduced into the network; training was stopped when the validation set loss did not decrease for 10 consecutive epochs to avoid overfitting. The overall model performance was evaluated using the test set, and supplementary samples were added to strengthen training for weak links such as minor internal defects and weak functional defects, thereby improving the model's detection accuracy.

[0095] Step S146: During online detection, the model receives the preprocessed multimodal feature matrix in real time, calculates the association weights through the cross-modal attention module to generate a fusion feature vector, calls the rule base for multi-dimensional matching, first judges morphological defects by surface defects and internal structural features, and then verifies performance defects by material distribution and functional features. If there is a modal feature conflict, cross-modal association verification is initiated, and the final judgment result containing defect type, level and association basis is output. During online detection, the model receives the preprocessed multimodal feature matrix in real time, calculates the association weights through a cross-modal attention module, and generates a fused feature vector. It then calls a judgment rule base for multi-dimensional matching, first performing an initial judgment of morphological defects based on surface defects and internal structural features, and then verifying performance based on material distribution and functional features. If modal feature conflicts occur, such as no surface defects but functional indicators exceeding limits, a cross-modal association verification is initiated, using a mapping matrix to find the source of the associated anomaly. The final output includes a comprehensive judgment result containing the defect type, level, and association criteria, providing a clear basis for defect handling.

[0096] Step S147: Establish a closed-loop verification mechanism, regularly compare the results of manual review with the results of model judgment, count misjudgment cases and analyze the reasons. If the misjudgment is due to inaccurate association mapping, adjust the cross-modal attention weights and loss function parameters and retrain the model. If the misjudgment rule threshold is unreasonable, revise the rule base quantification index based on actual detection data, and update the optimized rules and model parameters to the detection system in real time.

[0097] A closed-loop verification mechanism is established, with daily comparisons between manual review results and automatic model judgment results. Missed and misjudged cases are statistically analyzed and categorized to identify the causes. If misjudgments are due to inaccurate cross-modal association mapping, attention weights and loss function parameters are adjusted, and the model is retrained and optimized. If the judgment rule thresholds are unreasonable, the quantitative indicators of the rule base are revised based on actual production data. The optimized model parameters and rule base data are updated to the online detection system in real time, forming a "judgment-verification-analysis-optimization" closed loop to continuously improve the model's adaptability and judgment accuracy, adapting to changes in filter material, process, and operating conditions.

[0098] Step S150: Upload the defect judgment results to the production management system. If an unqualified defect is detected, trigger a marking or automatic rejection action. At the same time, perform correlation analysis between the defect data and the production process parameters.

[0099] The comprehensive defect assessment results are uploaded to the production management system in real time via industrial Ethernet. The system stores data such as defect type, location, level, and correlation basis. If a non-conforming defect is detected, the system immediately triggers an online marking device to print a defect label at the corresponding location on the filter media, or activates an automatic rejection mechanism to separate the non-conforming filter media from the discharge stream. Simultaneously, the defect data is correlated with production process parameters such as production line speed, tension, light source parameters, and polarization angle. Statistical analysis identifies process ranges with high defect incidence, providing data support for production process adjustments, equipment maintenance, and quality control, achieving closed-loop optimization of detection and production.

[0100] Based on the same inventive concept, please refer to Figure 2 This paper shows a schematic block diagram of an AI vision-based online detection system 100 for performing the above-described AI vision-based online detection method for defects in spunlace filter media. The AI ​​vision-based online detection system 100 may include a communication unit 110, a machine-readable storage medium 120, and a processor 130.

[0101] In this embodiment, both the machine-readable storage medium 120 and the processor 130 are located within the AI ​​vision-based online defect detection system for spunlace filter media 100 and are separately configured. However, it should be understood that the machine-readable storage medium 120 may also be independent of the AI ​​vision-based online defect detection system for spunlace filter media 100 and may be accessed by the processor 130 via a bus interface. Alternatively, the machine-readable storage medium 120 may be integrated into the processor 130 and may communicate with external systems via the communication unit 110.

[0102] The processor 130 is the control center of the AI ​​vision-based online defect detection system 100 for spunlace filter media. It connects to various parts of the system via various interfaces and lines. By running or executing software programs and / or modules stored in the machine-readable storage medium 120, and by calling data stored in the machine-readable storage medium 120, it performs various functions and processes data of the AI ​​vision-based online defect detection system 100, thereby providing overall monitoring of the system. Optionally, the processor 130 may include one or more processing cores; for example, the processor 130 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into the processor. The machine-readable storage medium 120 is used to store machine-executable instructions for executing the scheme of this application, and the processor 130 is used to execute the machine-executable instructions stored in the machine-readable storage medium 120 to realize the online detection method for defects in spunlace filter media based on AI vision provided in the aforementioned method embodiments.

[0103] It should be noted that, in order to simplify the description of the present invention and thus help to understand one or more embodiments of the invention, multiple features may sometimes be grouped into one embodiment, drawing or description thereof in the foregoing description of the embodiments of the present invention.

[0104] The embodiments of this application have been described above with reference to the accompanying drawings. Unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other. This application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. An online defect detection method for spunlace filter media based on AI vision, characterized in that: Includes the following steps: By adjusting the polarization relationship between polarized light and the fiber texture of spunlace filter media, the difference in optical refractive index between defects and fibers is used to make defects form specific optical signals, thus initially separating texture and defect features. At the same time, based on the light transmission characteristics of the filter media, the optical parameters of the near-infrared structured light source are adjusted to form a uniform active light field. The imaging parameters are dynamically adjusted according to the production line operation status. The multi-source detection module collects the spectral data, internal structure data, material distribution data and functional index data of the filter material. Through the synchronous control mechanism, it links with the production line control unit to achieve spatiotemporal registration of the data collected by each module in the same detection area. A deep learning-based texture stripping method is adopted, which combines the texture distribution law of filter material to generate an adaptive texture template. The background texture is stripped by dynamically adjusted texture separation logic while retaining the information of minute defects. Then, feature enhancement technology is used to enhance the edge and gray-scale difference of minute defects. Feature spectral information corresponding to filter material and internal defects is selected, and spectral data and spatial image details are fused to generate an enhanced image. A dynamic deblurring algorithm is used to eliminate image blur caused by production line movement, and edge sharpening technology is used to restore the true boundary of defects. Surface defect features, internal structural features, material distribution features, and functional features are input into a multimodal feature fusion model. A mapping relationship between each feature is established through a cross-modal correlation mechanism. Based on the preset multi-dimensional defect judgment rules, a comprehensive judgment of surface defects, internal defects, and functional defects of the filter material is completed. The defect judgment results are uploaded to the production management system. If an unqualified defect is detected, a marking or automatic rejection action is triggered. At the same time, the defect data is correlated with the production process parameters for analysis.

2. The online defect detection method for spunlace filter media based on AI vision according to claim 1, characterized in that: The adjustment of the polarization relationship between the polarized light and the fiber texture of the spunlace filter material utilizes the difference in optical refractive index between defects and fibers to create specific optical signals from the defects, thus initially separating texture and defect features. Simultaneously, based on the light transmission characteristics of the filter material, the optical parameters of the near-infrared structured light source are adjusted to form a uniform active light field, including: Fix the polarizer to the light source emission end, install the analyzer at the front of the camera lens, start the dual polarizer adjustment system, and drive the analyzer to rotate around the optical axis through the stepper motor. Using the pre-calibrated main polarization direction of the filter fiber as a reference, the relative angle between the polarizer and the analyzer is gradually adjusted, and filter images at different angles are acquired simultaneously. The optimal polarization angle is selected by the image grayscale correlation algorithm, so that the polarization reflection signal of the fiber texture is suppressed, and the defects form specific optical signals due to the difference in optical refractive index with the fiber, thus achieving the initial separation of the optical features of texture and defects. Based on the optimal polarization angle, the image analysis unit monitors the comparison parameters between the defect area and the background in real time. If the comparison parameters do not meet the preset standard, the analyzer is triggered to make fine adjustments within the preset range. The signal change data during the adjustment process is recorded, and a mapping relationship between the polarization angle and the defect signal intensity is established to form a dynamic adjustment model to adapt to the changes in filter material texture during production line operation. The transmission photoelectric sensor emits detection light in the same wavelength band as the near-infrared structured light source, which penetrates vertically through the filter material and collects the transmitted light signal. Combined with the filter material thickness detection data, the actual transmittance of the filter material is calculated. At the same time, the light reflection distribution on the filter material surface is collected to analyze the difference in transmittance uniformity. Based on the transmittance test results, the preset parameter mapping model is called to dynamically adjust the light spot density, power and projection angle of the dot matrix near-infrared light source. For areas with large differences in transmittance uniformity, the light source zoning adjustment mode is activated to adjust the light source unit corresponding to the weak transmittance area individually. The light field distribution on the filter material surface is collected in real time by an array of light sensors. The light field uniformity evaluation algorithm is used to judge the current light field uniformity. If the preset requirements are not met, secondary adjustment is performed by independently controlling and adjusting the vertical distance between the light source and the filter material and the focusing state of the lens group. At the same time, the power of the near-infrared light source is dynamically compensated according to the changes in the ambient light intensity of the production line. After adjusting the polarization angle and light source parameters, the filter material image under the current state is acquired, and the signal-to-noise ratio of the defect signal is used for collaborative verification. If the signal-to-noise ratio does not reach the optimal value, the polarization angle and light source power are adjusted synchronously, the image is acquired again and the signal-to-noise ratio is calculated, until a collaborative closed loop of polarization adjustment, light source adjustment and signal-to-noise ratio verification is formed.

3. The online defect detection method for spunlace filter media based on AI vision according to claim 2, characterized in that: The image analysis unit monitors the contrast parameters between the defect area and the background in real time. If the contrast parameters do not meet the preset standard, the analyzer is triggered to perform fine-tuning within a preset range. Signal change data is recorded during the adjustment process, establishing a mapping relationship between the polarization angle and the defect signal intensity, forming a dynamic adjustment model to adapt to changes in filter material texture during production line operation. This includes: Using the grayscale contrast between the defect area and the background area as the comparison parameter, the image analysis unit performs noise reduction filtering and grayscale preprocessing on the acquired filter material image. The minimum bounding rectangle of the defect area is defined as the target ROI. Areas with equal area and no texture abrupt changes around the target ROI are selected as the background ROI. The contrast is calculated using a standardized formula to eliminate the interference of the absolute grayscale value under different light intensities and ensure the consistency of the comparison parameters. For different fiber materials and defect types, multiple defect-background sample images are collected during the system initialization phase. The minimum contrast value under each scene is obtained through statistical analysis as the initial threshold. During the production line operation, texture images of defect-free filter material are automatically collected at preset time intervals, the grayscale fluctuation range under normal texture is calculated, and the initial threshold is dynamically corrected. The image analysis unit processes the image at a frequency synchronized with the camera's frame rate, calculates the contrast of each defect area in real time and compares it with a preset threshold. If the contrast of the same defect area is lower than the preset threshold for a consecutive preset number of frames and the fluctuation exceeds the preset proportion of the threshold, it is determined that the comparison parameter does not meet the standard and triggers the analyzer fine-tuning command. Based on the pre-calibrated optimal polarization angle, the preset fine-tuning range of the analyzer is determined. A preset traversal adjustment strategy is adopted to drive the stepper motor to rotate the analyzer with a fixed step size. Each time the step size is adjusted, the filter material image under the current angle is acquired synchronously, the contrast of the defect area is calculated and stored in real time, and the adjustment direction, step size, current polarization angle and corresponding contrast value are recorded to form a complete adjustment process dataset. The stored adjustment process dataset is preprocessed to remove abnormal data. A nonlinear mapping relationship model between the polarization angle and the defect contrast is established using a multinomial fitting algorithm. The parameters of the mapping relationship model are updated in real time using a sliding window algorithm to adapt to the dynamic changes in the filter material texture in the production line. The mapping relationship model is fused with the filter material texture feature parameters to construct a multi-input single-output dynamic adjustment model. The gradient descent algorithm is used to train the model with the optimization objective of maximizing contrast. A regularization term is introduced to suppress overfitting, and cross-validation is used to ensure the generalization ability of the model. During production line operation, the image analysis unit extracts the texture feature parameters of the filter material in real time and inputs them into the dynamic adjustment model. The model predicts the target angle that can make the contrast reach the optimal value based on the current polarization angle and texture features. It drives the stepper motor to adjust the polarizer to the target angle and monitors the contrast after adjustment in real time. If the expected optimal value is not reached, a second fine adjustment is made based on the model prediction error, forming a closed-loop adaptation mechanism of monitoring, prediction, adjustment and verification. If the contrast of the defect area still does not reach the preset threshold after the detector has been adjusted throughout the preset fine-tuning range, the system will automatically determine that it is an abnormal situation and trigger two levels of feedback. The first level of feedback is to adjust the power of the light source to assist in optimization, and the second level of feedback is to send an alarm signal to the control system to prompt the operator to check the texture of the filter material or the status of the equipment.

4. The online defect detection method for spunlace filter media based on AI vision according to claim 3, characterized in that: The process involves fusing the mapping relationship model with filter material texture feature parameters to construct a multi-input single-output dynamic adjustment model. A gradient descent algorithm is used to train the model with the goal of maximizing contrast. A regularization term is introduced to suppress overfitting, and cross-validation is employed to ensure the model's generalization ability. This includes: The set of input parameters for the dynamic adjustment model is clearly defined, including the output value of the model of the mapping relationship between polarization angle and defect contrast, and the texture feature parameters of the filter material. Texture feature indicators are extracted through gray-level co-occurrence matrix, and texture-related parameters are statistically analyzed by combining texture direction histogram to complete the quantitative extraction of multi-dimensional texture features. The relevant parameters of the mapping relationship model and the texture feature parameters are concatenated according to feature dimensions to form a unified multi-dimensional input feature vector. The model output parameter is set as the optimal polarization angle adjustment amount to achieve the goal of building a multi-input, single-output model. Collect the adjustment process dataset during production line operation and the texture feature dataset of different batches of filter material. Align the data into sample data pairs by timestamp. Use a standardization method to eliminate the difference in the dimension of the feature vector. Use an outlier detection algorithm to remove outlier samples. Then divide the training set and validation set according to a preset ratio. A lightweight shallow neural network architecture is constructed, with the number of neurons in the input layer matching the dimension of the fused features, a preset number of hidden layers, a preset ratio of neurons in each layer, a non-linear activation function, and a single neuron in the output layer using a linear activation function to output the polarization angle adjustment, thus balancing model fitting ability and real-time inference efficiency. A loss function is constructed with the core optimization objective of maximizing contrast. The loss function includes a correlation term between predicted contrast and maximum possible contrast and an L2 regularization term. The regularization strength is adjusted by the regularization coefficient to suppress model overfitting. An adaptive momentum estimation optimizer is selected to implement gradient descent. A dynamic adjustment strategy is used to adjust the learning rate. A preset batch size and number of iterations are configured. An early stopping mechanism is introduced. When the validation set loss does not decrease for a preset number of consecutive rounds, the iteration is stopped and the optimal model parameters are saved. Initialize the model weight parameters, input the training set samples into the model in batches, calculate the polarization angle adjustment through forward propagation, calculate the loss value with regularization term by combining the true contrast, solve the gradient of the loss function with respect to each weight parameter using the back propagation algorithm, update the weight parameters using the optimizer, and repeat the above process until the preset number of iterations is reached or the early stopping mechanism is triggered. During training, the loss trend of the training set and validation set is monitored in real time. Based on the overfitting or underfitting state of the model, the regularization coefficient is adaptively adjusted to make the loss of the model on the training set and validation set tend to be stable and the difference is within a preset range. The k-fold cross-validation method is adopted to divide the dataset into multiple non-overlapping subsets. Some subsets are selected in turn as the training set and the rest as the validation set. The training and validation process is repeated. The average loss value and contrast compliance rate of multiple validations are calculated. The model structure or feature fusion method is adjusted according to the validation results to ensure that the model's generalization ability meets the standard. The model performance is evaluated using an independent test set. If the performance does not meet the preset standard, the model structure, number of neurons, or activation function type are adjusted, and the training, regularization adjustment, and cross-validation process is re-executed to form a closed-loop iteration until the model performance meets the production line detection requirements.

5. The online defect detection method for spunlace filter media based on AI vision according to claim 1, characterized in that: The imaging parameters are dynamically adjusted according to the production line's operating status. A multi-source detection module collects spectral data, internal structure data, material distribution data, and functional index data of the filter material. A synchronous control mechanism links this data with the production line control unit, enabling spatiotemporal registration of data collected by each module from the same detection area. This includes: Multiple sensors are deployed to collect data on production line conveying speed, filter material tension fluctuations, and filter material material information. The collected data is transmitted to the central control unit in real time via an industrial bus. At the same time, the operating status signals of the production line control unit are accessed to obtain production process switching and start / stop status information, and a production line operating status perception matrix is ​​constructed. Based on historical production data, a correlation mapping model between production line operation status and imaging parameters is established to clarify the correspondence between imaging parameters and production line speed, filter tension, and material type. Adaptive adjustment rules and parameter adjustment priorities are embedded in the model to balance imaging quality and real-time performance. Based on the real-time perception of the production line's operating status, the central control unit calls the associated mapping model to output target imaging parameters, which are then sent to the imaging equipment via the signal transmission module. After receiving the instruction, the imaging equipment adjusts the core imaging parameters and related parameters of the supporting light source in real time, and records the parameter adjustment log simultaneously, forming a closed-loop response mechanism of status perception, model calculation, and parameter adjustment. Along the filter media conveying path, multi-source detection modules are deployed at corresponding positions in the same detection area. Each module collects spectral data, internal structure data, material distribution data, and functional index data of the filter media. The acquisition frame rate of each module is unified to ensure consistent data output rhythm. A high-precision synchronous controller is configured to establish linkage with the production line control unit through a communication interface. The position pulse signal of the production line encoder is obtained as a time and space reference. The synchronous controller outputs synchronous trigger signals to each multi-source detection module. The trigger interval is matched with the production line speed to ensure that each module starts collecting data synchronously when the filter material runs to the target detection area. The positioning and calibration technology is adopted. A calibration reference point is set in the detection area, and the spatial position coordinates of each detection module and the reference point are measured. A unified spatial coordinate system is constructed based on the coordinate data. The detection field deviation of each module is corrected by the coordinate transformation algorithm, so that the detection range of all modules accurately covers the same target area. The calibration process is repeated according to a preset cycle to compensate for spatial deviation. When the synchronous controller sends a data acquisition trigger signal, it records the current encoder pulse count and system time, and generates a unique timestamp and location tag. After each multi-source detection module acquires data, it automatically associates the timestamp and location tag. The system uses real-time industrial Ethernet to transmit data collected by each module. The central control unit has a built-in data transmission delay monitoring module that calculates data transmission delay in real time, establishes a delay compensation model based on the delay data, and performs time axis calibration on the data of different modules to correct the time difference caused by transmission delay. The central control unit verifies the consistency of the registered multi-source data, evaluates the registration effect by calculating the spatial overlap rate and time synchronization error, and if the preset standard is not met, it triggers the spatial calibration process or adjusts the synchronization trigger signal interval, optimizes the delay compensation model, records the registration parameter adjustment data, and continuously iterates and optimizes the correlation mapping model. When a certain detection module experiences abnormal data acquisition or registration error that continues to exceed the standard, the system automatically marks the abnormal module and switches to redundant acquisition mode, while simultaneously sending an early warning signal to the production line control unit.

6. The online defect detection method for spunlace filter media based on AI vision according to claim 1, characterized in that: The method employs a deep learning-based texture stripping approach, combining the texture distribution patterns of the filter material to generate a suitable texture template. Background texture is stripped using dynamically adjusted texture separation logic while preserving minor defect information. Feature enhancement techniques are then used to strengthen the edges and grayscale differences of these minor defects. Feature spectral information corresponding to the filter material and internal defects is selected, and spectral data is fused with spatial image details to generate an enhanced image. A dynamic deblurring algorithm is used to eliminate image blurring caused by production line movement, and edge sharpening techniques are employed to restore the true boundaries of defects. This includes: We collected texture images of spunlace filter media of different materials and samples with minor defects, labeled the background texture region and the defect region, and constructed a lightweight U-Net texture stripping model with encoder and decoder structure. The encoder extracts multi-scale features of filter media texture, the decoder restores the image size and introduces skip connections to fuse encoder features. The model parameters are optimized by combining cross-entropy loss function and structural similarity loss, so that the model can learn the texture distribution rules of filter media of different materials. The system acquires real-time images of the filter media surface to be tested, extracts the texture features of the current filter media through a pre-trained texture feature recognition network, matches suitable texture templates from a preset texture template library based on the extracted texture features, and adopts dynamic texture separation logic. First, it distinguishes the background texture from the potential defect area through an adaptive threshold segmentation algorithm, and then it peels off the background texture through template matching and pixel-level subtraction operation. A defect retention threshold is set to fully retain the information of minute defects, and the preliminary image after background removal is output. The initial image after background removal is decomposed by wavelet multi-scale decomposition to obtain low-frequency components and multiple high-frequency components. An attention mechanism module is introduced into the high-frequency components to weight and enhance the defect-related features and suppress the high-frequency components corresponding to residual noise. The enhanced high-frequency components and low-frequency components are reconstructed to obtain a small defect feature image with enhanced edges and gray-level differences. The full-band spectral data of the filter material is acquired synchronously by a hyperspectral camera. The spectral data is smoothed, denoised, and normalized before processing. The characteristic spectral bands corresponding to the filter material and internal defects are screened out by combining the spectral angle matching algorithm with principal component analysis. Redundant bands are removed and a characteristic spectral band index is established. The filtered feature spectral band data are converted into corresponding spatial grayscale images. An adaptive weight fusion strategy is adopted, which combines the texture complexity and defect distribution of the micro-defect feature images to dynamically adjust the fusion weight of each feature spectral grayscale image and the spatial image. Through a pixel-level weighted fusion algorithm, multiple feature spectral grayscale images and spatial images are integrated into an enhanced image that has both spatial details and spectral features. Based on motion-blurred samples and corresponding clear micro-defect samples at different operating speeds of the production line, a dynamic generative adversarial deblurring model with a conditional generative adversarial network architecture is constructed. The model input layer includes motion-blurred images and production line speed labels. The current filter material conveying speed is obtained in real time and used as the speed label input to the model. The model automatically generates a deblurring kernel at the corresponding speed and performs real-time deblurring processing on the enhanced image. A multi-scale edge detection algorithm is used to extract defect edge information from the deblurred image, and a gradient boosting algorithm is introduced to enhance the edge gradient. Pixel-level gradient aggregation technology is used to repair edge breakage and blurring problems. An adaptive sharpening intensity adjustment rule is set to adjust the sharpening parameters according to the defect size and edge complexity to restore the true boundary of tiny defects.

7. The online defect detection method for spunlace filter media based on AI vision according to claim 6, characterized in that: The lightweight U-Net texture stripping model, which includes an encoder and decoder structure, extracts multi-scale features of the filter material texture through the encoder, and the decoder restores the image size and introduces skip connections to fuse the encoder features. The model parameters are optimized by combining cross-entropy loss and structural similarity loss, enabling the model to learn the texture distribution patterns of different filter materials. This includes: Collect defect-free standard texture images and defect-containing detection images of spunlace filter media of various materials, divide them into training set, validation set and test set, distinguish background texture area and defect area by pixel-level annotation, generate corresponding binary annotation mask, expand sample diversity by data augmentation, simulate image differences of actual production line working conditions, and form a standardized training dataset. With the goal of balancing real-time performance and detection accuracy, depthwise separable convolution is used to replace standard convolution to reduce the number of model parameters. The model input size is set to adapt to the resolution of filter material images acquired by the line scan camera. The encoder-decoder structure is retained and redundant network layers are removed to construct a lightweight U-Net texture stripping model. The model output is a texture and defect separation feature map. The encoder module is built by cascading multiple lightweight convolutional units. Convolutional kernels with different receptive fields are used to extract the texture features of the filter material. The bottom convolutional units extract shallow fine-grained texture features, and the high-level convolutional units aggregate global information and learn the deep semantic texture features corresponding to the material through downsampling. Each convolutional unit is followed by a batch normalization layer and an activation function layer to temporarily store the multi-scale feature maps output by each level of the encoder. The deep feature map output by the encoder is upsampled by the transposed convolutional layer to restore the image size. After each upsampling, the decoder features and the multi-scale feature map of the corresponding level of the encoder are concatenated through a skip connection structure to fuse shallow detail features and deep semantic features. The decoder outputs pixel-level texture prediction results through a 1×1 convolutional layer. Define a cross-entropy loss function to distinguish between background texture and defect regions, minimizing pixel classification error; define a structural similarity loss function to preserve texture structure distribution features, constraining the model to learn the texture patterns of real materials; and fuse the two loss functions through weighted coefficients to form a total loss function. An adaptive optimizer was selected to configure training parameters, and an initial learning rate and learning rate decay mechanism were set. Forward propagation was performed using mini-batch sample input. The prediction error was calculated using the total loss function. The network parameters of the encoder and decoder were updated based on the backpropagation algorithm, enabling the model to learn the texture distribution patterns of different filter materials. By monitoring the changes in texture stripping accuracy and loss value in real time using the validation set, the model is determined to converge when the loss value is stable and the separation accuracy meets the standard. Random deactivation layer and early stopping mechanism are introduced. The model's texture stripping effect, defect preservation integrity, and inference speed are evaluated using a test set. The model's learning effect on texture distribution patterns is verified. After lightweight fine-tuning, the model parameters are solidified to obtain a texture stripping model that can be used for online detection.

8. The online defect detection method for spunlace filter media based on AI vision according to claim 6, characterized in that: The constructed dynamic generative adversarial deblurring model employs a conditional generative adversarial network architecture. The model input layer includes a motion-blurred image and production line speed labels. The current filter material conveying speed is acquired in real-time and used as the speed label input to the model. The model automatically generates a deblurring kernel corresponding to the speed, including: Motion blur images of spunlace filter media at different production line conveying speeds and corresponding clear standard images are collected. Production line speed is used as a classification label for each group of blur images. The sample is expanded by data augmentation to simulate the blur effect caused by the variable speed operation and shaking of the production line. The training set, validation set and test set are divided according to a preset ratio to complete the dataset standardization preprocessing. A dual-entity model architecture containing a generator and a discriminator is constructed. The production line speed is embedded into the network framework as a conditional constraint variable. The model has two input channels, which respectively receive motion-blurred images and quantized production line speed labels. The model output is set with a deblurring kernel generation branch and a clear image reconstruction branch to achieve customized deblurring processing based on speed conditions. A generator backbone is built using an encoder-decoder architecture. The production line speed label is encoded and embedded with features, and converted into a high-dimensional conditional feature vector. The encoder performs multi-scale feature extraction on the motion-blurred image. In the middle layer of the encoder, the speed conditional feature vector and the image features are fused at the channel level. The decoder has two branches. The main branch is used for image sharpening and reconstruction. The branch outputs a dynamic deblurring kernel that matches the current production line speed. The dynamic deblurring kernel includes motion blur direction and blur scale parameters. A discriminator is designed using a conditional discriminant structure. It has two input interfaces to receive the clear image and the real clear image output by the generator, respectively. At the same time, the production line speed label is used as an auxiliary condition input to the discriminator. The discriminator extracts image features through convolution and downsampling. After fusing the image features and speed condition features, it outputs dual discrimination results, namely the image authenticity discrimination result and the deblurring effect and speed condition matching discrimination result. Adversarial loss, image reconstruction loss, and velocity conditional matching loss are constructed and weighted to form a total loss function. The adversarial loss is used to reduce the distribution difference between the generated image and the real clear image, the image reconstruction loss is used to constrain the image detail restoration, and the velocity conditional matching loss is used to ensure the adaptability of the dynamic deblurring kernel and the input velocity label. Initialize the network parameters of the generator and discriminator, select the adaptive optimizer to configure the training parameters, input the training set samples into the model in batches, the generator generates a dynamic deblurring kernel based on the blurred image and velocity label and completes the image deblurring, the discriminator performs dual discrimination, and the generator and discriminator parameters are alternately updated through the backpropagation algorithm based on the total loss function. The model converges when the validation set is monitored in real time for changes in deblurring accuracy, speed matching degree and loss value. A regularization layer and an early stopping mechanism are introduced. The trained model is solidified and connected to the detection system. The current filter material conveying speed is collected in real time through the production line encoder, converted into a speed label that the model can recognize, and input into the model. The model calls the learned speed-blur kernel mapping relationship to generate a dynamic deblur kernel that matches the current running speed, providing parameters for image deblurring processing. The system calculates the sharpness index and defect edge integrity index of the deblurred image in real time. If the index does not meet the standard, it triggers the model fine-tuning process and updates the network parameters based on the current speed and image data.

9. The online defect detection method for spunlace filter media based on AI vision according to claim 1, characterized in that: The process involves inputting surface defect features, internal structural features, material distribution features, and functional features into a multimodal feature fusion model. A mapping relationship between these features is established through a cross-modal correlation mechanism. Based on preset multi-dimensional defect judgment rules, a comprehensive judgment of surface defects, internal defects, and functional defects of the filter material is completed, including: The surface defect features, internal structural features, material distribution features and functional features of the spunlace filter media were extracted by quantitative methods. The modal features were normalized by standardization methods to eliminate dimensional differences. Based on the label information after spatiotemporal registration, the four modal features of the same detection area were spliced ​​by dimension to form a multimodal original feature matrix. A cross-modal feature fusion encoder architecture based on Transformer is constructed. Four independent modal embedding layers are set up to convert modal features of different dimensions into a unified dimensional feature vector. After the embedding layer, a position encoding module is connected to add spatial position information. The core layer adopts a cross-modal multi-head attention module to capture the semantic association between modalities through multi-head parallel computing. The concatenated layer normalization and feedforward neural network perform nonlinear transformation and dimensionality optimization on the fused features, and output a unified multimodal fusion feature vector. A labeled sample set containing the original feature matrix of multimodal modes and the corresponding defect labels is constructed. The sample set is input into the fusion model, and the fusion feature vector is calculated through forward propagation. A cross-modal association loss function is introduced to constrain the model to learn the association law between modes. The model parameters are updated through the backpropagation algorithm, and a cross-modal association mapping matrix is ​​established. Industry standards and production requirements were reviewed, and a judgment rule base was constructed in three categories: surface defects, internal defects, and functional defects. The rule base was embedded into the model output layer, and rule matching priority was set to form a dual judgment logic of morphological judgment and performance verification. An association confidence threshold was introduced to trigger a supplementary judgment process for cases with fuzzy boundaries. The labeled sample set is divided into training set, validation set and test set according to the proportion. An adaptive optimizer is selected to configure training parameters, and learning rate and decay strategy are set. During training, the defect judgment accuracy and association mapping error of the validation set are monitored. The model structure is adjusted, and a regularization layer and early stopping mechanism are introduced to avoid overfitting. The model performance is evaluated through the test set, and samples are added to strengthen the training for weak links. During online detection, the model receives the preprocessed multimodal feature matrix in real time, calculates the association weights through the cross-modal attention module to generate a fused feature vector, calls the rule base for multi-dimensional matching, first judges morphological defects by surface defects and internal structural features, and then verifies performance defects by material distribution and functional features. If there is a modal feature conflict, cross-modal association verification is initiated, and the final judgment result containing defect type, level and association basis is output. Establish a closed-loop verification mechanism, regularly compare the results of manual review with the results of model judgment, count misjudgment cases and analyze the reasons. If the misjudgment is due to inaccurate association mapping, adjust the cross-modal attention weights and loss function parameters and retrain the model. If the misjudgment rule threshold is unreasonable, revise the quantitative indicators of the rule base based on actual detection data, and update the optimized rules and model parameters to the detection system in real time.

10. An online defect detection system for spunlace filter media based on AI vision, characterized in that, include: processor; A machine-readable storage medium for storing machine-executable instructions of the processor; The processor is configured to execute the AI ​​vision-based online defect detection method for spunlace filter media according to any one of claims 1 to 9 by executing the machine-executable instructions.