Intelligent identification method for flaws of plush fabric based on multi-modal fusion
By using a multimodal fusion method, combining RGB, infrared, and depth images, a lightweight YOLOv11 model is constructed, which solves the problem of detecting complex defects in plush fabrics. It achieves real-time and stable defect recognition and quality level evaluation, and is suitable for industrial quality inspection scenarios.
Patent Information
- Application Number
- CN202511135875.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2025-11-28
AI Technical Summary
Existing technologies are ineffective in detecting complex defects in plush fabrics, especially in cases of high texture complexity and multi-scale conditions. They are easily affected by light and shadow interference and texture artifacts, and lack deep semantic feature representation and detection generalization ability, thus failing to meet the real-time detection needs of production lines.
A multimodal fusion approach is adopted, combining RGB images, infrared images, and depth images. Features are extracted through a lightweight parallel encoder, and semantic projection is used for feature mapping and fusion. A lightweight YOLOv11 model is constructed, and a multi-scale hybrid attention mechanism is used for defect recognition. Temporal analysis and dynamic compensation mechanisms are set up to achieve real-time detection of defects in plush fabrics.
It enables stable, real-time detection of defects in plush fabrics, improves the comprehensiveness and accuracy of detection, reduces computational complexity and storage scale, adapts to the deployment requirements of edge computing devices, and meets the high standards required for unattended quality inspection.
Smart Images

Figure CN121032950A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of computer systems, in particular to a plush fabric defect intelligent identification method based on multi-modal fusion. BACKGROUND
[0002] Fabric defect detection has become a core link in the textile industry under the development of intelligent manufacturing. The surface structure of plush fabric is complex and has many common defects including jump wool, broken weft, color difference and pile shedding, which makes it very difficult to detect defects. Now there is a technology that uses Gaussian mixture model and gray feature for segmentation and classification, which realizes automatic detection in part. However, this method relies on a single visual mode and is difficult to deal with the high texture complexity and multi-scale defects of plush fabric, and is easily affected by light and shadow interference and texture artifacts, resulting in missed detection and misjudgment. Moreover, this method is complex in space feature calculation and cannot meet the real-time detection needs of the production line, and the deep semantic feature expression lacks and the detection generalization ability is insufficient.
[0003] The application of new version models of YOLO series in industrial quality inspection is gradually rising. Through lightweight and adding multi-scale mixed attention mechanism, the feature extraction and multi-target detection ability of YOLOv11 are optimized, but the existing YOLO model is mainly used for natural scenes, and there are still limitations in recognizing complex defects of plush fabric in this detection scene.
[0004] Therefore, a plush fabric defect intelligent identification method based on multi-modal fusion is proposed. SUMMARY
[0005] The purpose of the present application is to provide a plush fabric defect intelligent identification method based on multi-modal fusion, combining synchronous acquisition and feature alignment, based on lightweight YOLOv11 and multi-scale attention, realizing multi-modal collaborative defect detection under the continuous movement of plush fabric.
[0006] To achieve the above purpose, the present application provides the following technical scheme:
[0007] Acquire RGB images, infrared images and depth images during the continuous movement of the plush fabric, construct a lightweight parallel encoder to extract features from the three kinds of images respectively, and obtain multi-modal local feature information;
[0008] Map the multi-modal local feature information to a unified dimensional space through semantic projection, fuse the features of each mode to generate a joint feature map; perform gradient operation and local morphology analysis on the depth image to convert it into surface relief and pile distribution features, and output a surface feature map;
[0009] Construct a lightweight model, and reduce the load through structure reparameterization, hardware perception compression and dynamic calculation optimization;
[0010] Based on the joint feature map and the surface feature map, multimodal image features are extracted through a multi-scale hybrid attention mechanism to identify and output defect features on plush fabrics. The multi-scale hybrid attention mechanism consists of multi-scale convolutional layers, channel attention layers, and spatial attention layers working together.
[0011] In the continuous conveying and inspection of plush fabrics, defect labels are determined and the defects are evaluated based on the characteristics of the defects, and the quality level evaluation results are output.
[0012] Preferably, the steps of constructing a lightweight parallel encoder to extract features from the three types of images include:
[0013] The lightweight parallel encoder includes RGB image paths, infrared image paths, and depth image paths. The acquired multimodal images of continuous plush fabric are input into the corresponding paths, and modal features are extracted step by step.
[0014] The modal features extracted step by step are output as multimodal local features.
[0015] Preferably, the steps of mapping multimodal local feature information to a unified dimensional space through semantic projection, fusing features from various modalities to generate a joint feature map, and performing gradient calculations and local morphological analysis on the depth image to transform it into surface undulation and velvet distribution features, and outputting a surface feature map include:
[0016] A 1×1 convolutional layer with shared weights is used to uniformly map the features of RGB images, infrared images, and depth images to the same number of channels;
[0017] During the mapping process, batch normalization and activation functions are used to normalize the distribution of each modality feature, and the unified features are then subjected to weighted fusion and splicing operations.
[0018] Gradient calculations are performed on depth images to extract the rate of change of surface undulations. At the same time, local morphology analysis is combined to identify the shape and distribution characteristics of the microstructure of the villi, and finally, a surface feature map is output.
[0019] Preferably, the step of fusing features from various modalities to generate a joint feature map includes:
[0020] The multimodal image is passed to a sliding window segmentation, and the entire image is divided into a 256×256 pixel block sequence with a preset overlap ratio;
[0021] The segmented image block sequence enters the edge gradient enhancement stage, where the Sobel operator is used to calculate the gradient magnitude map, and it is fused with the original image with preset weights to generate feature blocks that enhance texture boundaries.
[0022] Perform data augmentation by randomly applying rotation, cropping, and flipping operations to the feature blocks and outputting a training sample set.
[0023] Preferably, the steps for constructing a lightweight YOLOv11 defect detection model include:
[0024] During the training phase, a multi-branch convolutional structure is constructed, and multi-scale features are extracted in parallel using convolutions of different scales. The multi-branch convolution is transformed into a single equivalent convolution through parameter merging operations.
[0025] Preferably, the steps for load reduction through structural reparameterization, hardware-aware compression, and dynamic computation optimization include:
[0026] The structural reparameterization is used to extract multi-level features from RGB images, infrared images, and depth maps; the convolutional combinations and normalization parameters during training are folded into a single-path convolutional structure.
[0027] The hardware-aware compression further performs cropping, quantization, and channel rearrangement on the model based on structural reparameterization, thereby compressing the parameter scale of multimodal images.
[0028] The dynamic computation optimization adjusts the computation path based on the features of the multimodal image input in each frame, skipping some computations for flawless background areas and retaining flawed areas.
[0029] Preferably, the step of extracting multimodal image features based on the joint feature map and the surface feature map through a multi-scale hybrid attention mechanism to identify and output defect features on the plush fabric, wherein the multi-scale hybrid attention mechanism is a step in which multi-scale convolutional layers, channel attention layers, and spatial attention layers work together includes:
[0030] By combining the surface dynamics of plush fabrics during continuous transport, an optimized feature map is generated by processing multimodal features through a three-layer cascaded structure:
[0031] The multi-scale convolutional layer uses three sets of parallel dilated convolutional kernels of 1×1, 3×3, and 5×5 to extract microscopic fluff shaking features, defect contour features, and fabric deformation features, respectively. The multi-scale convolutional layer is used to extract local and global features of defects of different sizes.
[0032] The channel attention layer dynamically generates channel masks based on the standard deviation of the temperature distribution in the infrared image, suppresses noise channels caused by reflection in the RGB features, and generates weight coefficients for each channel through lightweight fully connected layer operations, learning the contribution of different modalities and channels to defect recognition.
[0033] The spatial attention layer calculates the local structural similarity between the depth map and the RGB image; it then focuses on key defect areas in the image.
[0034] Output the category, size, and spatial location of the defect, and output an optimized feature map that incorporates key information from multiple modalities.
[0035] Preferably, the continuous conveying and detection step of the plush fabric includes:
[0036] Real-time monitoring of input and defect distribution characteristics;
[0037] Utilize historical defect detection feature cache to reduce redundant calculations.
[0038] Preferably, the step of determining and outputting defect labels and evaluating defects based on defect characteristics, and outputting quality grade evaluation results includes:
[0039] Based on the characteristics of the defects, classify and determine the types of defects, and output category labels for damage, color difference, foreign matter, and pilling;
[0040] By combining the size, location, and density of defects, the severity of the defects is assessed, and a quality grade evaluation result is generated and output.
[0041] Preferably, the method further includes setting a timing analysis and dynamic compensation mechanism to correlate and analyze defects in consecutive frames, specifically including the following steps:
[0042] Temporal convolution and sliding window methods are used to analyze the temporal trend of multimodal images of continuously moving plush fabrics and extract multimodal temporal features.
[0043] A temporal correlation model is constructed to dynamically match the multidimensional features of consecutive frames and compare the changes in defects by comparing the spatial location information of consecutive frames with modal consistency features.
[0044] To address the feature loss caused by variations in the conveying speed of plush fabrics, image blurring, and asynchronous modal data, a dynamic compensation method is employed, combining data from previous and subsequent frames for interpolation and correction to generate a continuous temporal feature sequence.
[0045] The continuously time-series feature sequence after dynamic compensation is input into the time-series decision unit. Combined with the spatial location, volume and morphology of the defect, the correlation analysis of the continuous frames is performed to determine whether the defect persists, expands or disappears.
[0046] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0047] 1. This invention acquires RGB, infrared, and depth images of a continuously moving plush fabric and inputs them into a lightweight parallel encoder. It extracts spatial texture features, temperature features, and 3D morphological features from each modality and combines them with a shared semantic projection layer to map the multimodal features to a unified feature space, generating a multimodal fused image for subsequent detection. This solution addresses the problem that single-modal image detection methods cannot fully perceive the fabric surface structure, temperature variations, and depth details.
[0048] 2. This invention further constructs a lightweight YOLOv11 defect detection model based on multimodal fusion images. It employs structural reparameterization technology to convert the multi-branch structure during training into a single-branch equivalent convolution during inference, reducing the model's computational complexity. Simultaneously, by combining hardware-aware model pruning and quantization, it compresses the model's storage size, improves inference efficiency, and adapts to the deployment requirements of edge computing devices.
[0049] 3. In the model feature extraction process, a multi-scale hybrid attention mechanism is adopted. Multi-scale convolutional layers are used to extract features from different receptive fields. Channel attention and spatial attention are combined to allocate feature weights, focusing on extracting features from local areas related to defects. Considering the continuous transport characteristics of plush fabrics, this invention sets up a temporal analysis and dynamic compensation mechanism. Combining sliding windows and temporal convolution, it analyzes feature changes in continuous frames of multimodal data, dynamically compensating for feature loss caused by speed fluctuations, modal asynchrony, and image blurring, maintaining the temporal consistency of the detection results. Attached Figure Description
[0050] Fig. 1 This is a flowchart of the intelligent defect recognition method for plush fabrics based on multimodal fusion according to the present invention.
[0051] Fig. 2 This is a schematic diagram of the intelligent defect recognition structure for plush fabrics based on multimodal fusion according to the present invention.
[0052] Fig. 3 This is a diagram illustrating the interaction process in an embodiment of the intelligent identification of defects in plush fabrics based on multimodal fusion according to the present invention. Detailed Implementation
[0053] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0054] Example 1:
[0055] The following example, using a production line for online defect detection of plush fabrics, will be used to illustrate this embodiment in detail.
[0056] In a specific embodiment, the present invention provides an intelligent detection method for defects in plush fabrics based on multimodal fusion, the process of which is as follows:
[0057] Please see Figs. 1 to 3 The intelligent detection method for defects in plush fabrics proposed in this invention has the following technical solution:
[0058] This method includes multimodal data acquisition, lightweight feature extraction and fusion, lightweight YOLOv11 detection model, multi-scale hybrid attention mechanism, temporal analysis and dynamic compensation mechanism, and feedback control mechanism. Through the above structure, real-time and stable detection of surface defects of continuously conveyed plush fabrics is achieved.
[0059] First, the process involves acquiring multimodal image data of the continuously transmitted plush fabric surface. Specifically, the acquisition includes: first, using an industrial RGB camera to obtain the color texture information of the plush fabric, with a preferred resolution of 1280×720 and a frame rate of 50 frames per second; second, using an infrared camera and a depth camera to acquire temperature distribution data and three-dimensional height information of the plush fabric surface, respectively, to ensure the identification of changes in thermal properties and geometric features caused by material undulations, indentations, wear, etc. To ensure the temporal synchronization of multimodal data, each modal camera is bound to a unified timestamp output through a time synchronization mechanism during acquisition. The method has a built-in motion compensation mechanism that can adjust frame synchronization in real time to prevent modal misalignment caused by transmission speed fluctuations. For example, on a production line where the plush fabric moves at a speed of 30 meters per minute, by configuring an industrial camera group, stable acquisition of multimodal data can be achieved, ensuring complete data input for subsequent feature extraction and detection.
[0060] Preferably, the sampling accuracy of the acquisition device is set between 30 and 60 frames per second, and the sensor response delay is less than 10 milliseconds, so as to ensure real-time matching between the acquisition and the production line speed.
[0061] This invention enables simultaneous perception of both visible and latent defects, fully capturing color, temperature, and morphological information. It covers various defect types such as stains, color differences, weft breaks, lint skipping, dents, and wear, comprehensively characterizing the multidimensional state changes of the plush fabric surface. Through hardware-level time synchronization and spatial registration, it ensures tight alignment of multimodal data in high-speed operating scenarios, avoiding error accumulation caused by speed fluctuations or modal delays, and ensuring high consistency of each frame of data in spatial location and time node. It provides a complete and stable data foundation for subsequent feature extraction, defect localization, and classification decisions, improving the reliability and data integrity of detection in continuous production processes and realizing an efficient and controllable online detection process.
[0062] Next, lightweight feature encoding and unified feature space mapping are performed on the acquired multimodal data. Based on the acquired output, local features of RGB, infrared and depth maps are extracted in parallel through a lightweight convolutional encoder to obtain surface texture, temperature gradient and spatial geometric information, respectively.
[0063] This invention extracts features from different modalities using a lightweight convolutional coding network, combines shared semantic projection to achieve feature alignment and spatial fusion, and introduces a modality-aware dynamic feature compression mechanism. This allows features of different modalities to adaptively adjust their expression density according to the data content during the encoding process, avoiding the accumulation of redundant features and further reducing computational redundancy in the multimodal fusion process. This method supports the direct processing of continuous streaming data on edge devices, and combined with pipelined parallel inference, it reduces storage and caching pressure, ensures low-latency output of detection results, achieves efficient closed-loop data flow processing, and improves real-time intelligent detection capabilities under complex working conditions.
[0064] Furthermore, it can be refined as follows: First, the feature tensors of each modality are mapped to a unified feature dimension through a shared semantic projection layer, and a 1×1 convolutional kernel is used to align the number of channels, so that the output feature tensor has a unified 64-dimensional channel; Second, the features of each modality are fused to generate a multimodal joint feature map, which provides input for the subsequent detection model.
[0065] This invention solves the problems of inconsistent dimensions and semantic distribution differences in multimodal data by uniformly mapping RGB, infrared and depth features through shared semantic projection. It achieves feature-level fusion, retains modal complementary information, takes into account color, thermal characteristics and spatial morphology features, improves the comprehensive expression ability of diverse and complex defects, reduces model redundancy, simplifies the structure and enhances the generalization and adaptability of different plush fabric types and production environments.
[0066] Next, the system determines whether the fused multimodal features meet the model input conditions. First, the system checks the size, number of channels, and data integrity of the feature tensor. If it matches the preset model input specifications, the feature tensor is directly input into the lightweight YOLOv11 model to perform defect identification and classification, outputting the spatial coordinates, category label, and confidence score of the defect. If the feature tensor is found to have size deviations, missing channels, misaligned boundaries, or feature loss due to abnormal data flow, an anomaly detection mechanism is activated to automatically calibrate and correct the abnormal features. If necessary, previous steps are called for local re-acquisition to fill in the data gaps.
[0067] This process ensures the integrity and standardization of the model input data, avoids false detections or missed detections due to input anomalies, ensures the stable operation and inference accuracy of subsequent detection tasks, reduces the system's sensitivity to data anomalies, and enhances the fault tolerance of online detection.
[0068] This invention ensures the quality of data entering the model by dynamically verifying the size, channels, and integrity of multimodal feature tensors, avoiding recognition errors caused by data anomalies, acquisition errors, or synchronization deviations. Through a built-in anomaly detection and correction mechanism, it automatically identifies feature loss, size misalignment, and data distortion issues, supporting local reacquisition and online repair, effectively reducing the risk of detection interruptions caused by environmental factors, equipment aging, or fluctuations in conveyor speed. This design guarantees the long-term stable operation of the system in the complex environment of actual industrial production lines, improving the continuity and reliability of online detection and meeting the high standards of unattended quality inspection.
[0069] When the feature integrity requirement is met, the model enters the detection and inference stage. First, the fused multimodal feature tensor is input into the lightweight YOLOv11 model. The model backbone simplifies the multi-branch convolutional structure in the training stage through structural reparameterization technology, merging the original multi-path convolution, batch normalization and activation function into a single equivalent convolutional kernel, eliminating redundant calculations and inference redundancy, and reducing the complexity of the network forward propagation.
[0070] Next, the model performs hardware-aware pruning based on the computing power environment of the target hardware, analyzes the weight contribution of each channel, prunes low-importance convolution channels and redundant feature branches, and optimizes the model's computation path. After the model is compressed, all convolution kernels and bias parameters are quantized using INT8, mapping floating-point weights to 8-bit integers, and reconstructing the quantization-aware inference graph to minimize memory bandwidth and computation latency.
[0071] The entire process does not affect the original detection framework of the model, maintaining the accuracy of defect identification, while significantly reducing the number of model parameters and computational load. This enables efficient deployment of the model on edge devices such as the built-in computing units of industrial cameras and embedded AI terminals, meeting the real-time inference requirements of continuous streaming detection.
[0072] This invention retains multi-branch convolution, batch normalization, and multi-scale feature extraction structures during the training phase to fully guarantee the model's expressive power and feature learning depth, ensuring efficient learning and differentiation of different types of defects in plush fabrics. During the inference phase, structural reparameterization transforms the complex network into an equivalent single-path convolution structure, significantly reducing the number of computational layers and convolutional operations during inference to avoid redundant computation. Combined with pruning and quantization techniques, it effectively reduces the model parameter size and memory footprint, achieving optimal compression of computational paths to reduce power consumption and resource consumption. This design supports direct deployment of the model on industrial edge devices, compatible with low-computing-power chips and embedded hardware, ensuring real-time online inference in continuous production scenarios to meet the comprehensive requirements of unattended quality inspection systems for rapid response and efficient operation.
[0073] The system then employs a multi-scale hybrid attention mechanism to further optimize feature extraction. Multi-scale convolutional layers extract features from different receptive fields, capturing the scale differences between large areas and minute defects. A channel attention mechanism is used to calculate global pooling weights, highlighting defect-related feature channels. Finally, a spatial attention mechanism generates local saliency maps, enhancing the response to key regions and ensuring the model accurately distinguishes between background texture and real defects. This feature extraction mechanism works in conjunction with the detection head in the backbone network, achieving refined multi-modal feature attention.
[0074] During the detection process, temporal analysis and dynamic compensation steps are performed, and temporal correlation judgment is established by combining the multimodal features of continuous frames. A sliding window mechanism is used to form a feature sequence of the current frame and the two frames before and after it, and the evolution trend of the defect target on the time axis is analyzed. If the defect response appears continuously at the same position in multiple frames, it is judged as a real defect; if it appears only in a single frame and no defect is found in the frames before and after it, it is judged as transient interference, and the output is automatically filtered and the confidence level is reduced.
[0075] Ultimately, by combining dynamic analysis of multimodal information, this invention achieves stable detection of defects in plush fabrics and supports real-time monitoring under continuous production line operation conditions.
[0076] Example 2:
[0077] The following example, using the online grading and testing scenario in the plush fabric supply chain, will be used to illustrate this embodiment in detail.
[0078] In a specific embodiment, the present invention provides an intelligent detection method for defects in plush fabrics based on multimodal fusion, the process of which is as follows:
[0079] Please see Figs. 1 to 3 The present invention proposes a multimodal fusion-based intelligent defect detection method for plush fabrics, the technical solution of which is as follows:
[0080] It includes multimodal data acquisition, parallel lightweight feature encoding, multimodal feature fusion and mapping units, lightweight hierarchical detection models, multi-scale hybrid attention mechanisms, temporal dynamic analysis, level determination, and hierarchical feedback control units. By classifying and grading different types of defects, it not only identifies the existence of defects but also automatically grades them according to their severity, meeting the comprehensive needs of batch quality inspection, sorting, re-inspection, and big data traceability in the supply chain.
[0081] The implementation steps are detailed below:
[0082] First, multimodal data of plush fabrics are collected in real time during the outbound process of the fabric supply chain.
[0083] Specifically, the process includes the following sub-steps: First, a high-resolution RGB industrial camera is used to acquire the color texture features of the plush fabric surface, capturing visible defects such as stains, color differences, and surface damage; second, an infrared thermal imaging sensor is used to acquire a temperature distribution map to help determine hidden thermal anomalies caused by fabric accumulation, uneven filling, or plush fabric indentations; third, a laser structured light depth camera is used to scan the three-dimensional morphology of the plush fabric surface to identify spatial structural defects such as undulations, shedding, and depressions.
[0084] Preferably, multimodal image acquisition can also incorporate a modality-coordinated data cleaning mechanism. This mechanism comprises three stages: modality consistency screening, timestamp alignment verification, and defect sensitivity enhancement filtering. Modality consistency screening automatically removes low-quality images caused by occlusion, lighting interference, or blurring during acquisition by analyzing the spatial registration degree and image integrity of the three modal images. The timestamp alignment verification stage uses hardware synchronization markers to perform frame-level pairing verification of the three modal images, correcting frame errors and missing frames caused by differences in sensor response. Finally, defect sensitivity enhancement filtering combines RGB edge intensity maps, infrared thermal gradient maps, and depth fluctuation maps for joint filtering, highlighting signals in potential defect areas and suppressing background redundancy.
[0085] The addition of cleaning significantly improves the utilization rate and detection stability of the original data, providing a high-quality input foundation for subsequent feature extraction and recognition. It enhances the signal response to genuine defect features and avoids interference from false information in model discrimination. Simultaneously, this mechanism possesses real-time and adaptive capabilities, dynamically adjusting the cleaning intensity based on image quality during acquisition to ensure the stability and consistency of input data during large-scale continuous detection, thereby reducing the risk of false positives and false negatives and improving the overall robustness and controllability of the detection system in complex industrial scenarios. Compared to existing data screening methods that rely on manual thresholds or fixed templates, the data cleaning scheme of this invention enables dynamic adjustment and noise removal based on image content, offering greater flexibility and adaptability, and significantly improving the generalization ability and operational efficiency of the multimodal detection system in complex working conditions.
[0086] Next, feature extraction and spatial mapping are performed on the acquired multimodal images. Based on the acquired output, a lightweight convolutional coding network is used to extract local features from the RGB image, infrared image, and depth image, respectively, and output three sets of intermediate feature tensors.
[0087] Furthermore, the trimodal feature tensors are uniformly mapped to the same feature space through a shared semantic projection layer, outputting a multimodal fusion feature map.
[0088] Preferably, the semantic mapping process also includes introducing a feature alignment mechanism guided by contrastive learning to construct cross-modal positive and negative sample pairs during the training phase, so as to constrain the features of different modalities to maintain high consistency in a unified semantic space and enhance alignment accuracy through contrastive loss constraints.
[0089] This method retains the perceptual advantages of RGB, infrared, and depth maps while effectively suppressing semantic biases between modalities, thus improving the robustness and defect representation capabilities of the fused features. Compared to existing single-feature fusion techniques, the semantically aligned feature map can serve as a highly consistent fusion input, enhancing the stability of subsequent hierarchical models in recognizing complex and multi-morphological defects.
[0090] Next, it determines whether the multimodal fusion features meet the input requirements for graded detection. If the size of the fusion feature tensor is consistent with the model input specifications, it is directly input into the lightweight YOLOv11 graded detection model. The model output includes: defect category, location coordinates, defect confidence, and grade labels: A (minor), B (moderate), and C (severe).
[0091] Conversely, if the feature tensor contains missing data, misaligned dimensions, or anomalous data, the method will automatically invoke anomaly detection and correction to recover missing modalities and perform feature interpolation to ensure data integrity. This mechanism avoids inference interruptions caused by lost modal data, ensuring the continuity of the detection process.
[0092] The approach activates a multi-scale hybrid attention mechanism to further enhance multimodal features. 1×1, 3×3, and 5×5 multi-scale convolutional kernels are used to extract features from different receptive fields to capture large-area defects and minor flaws. A channel attention mechanism is combined to assign weights to each feature channel, highlighting feature responses related to defects; then, a spatial attention mechanism is used to generate a saliency weight map to automatically focus on abnormal areas on the fabric surface.
[0093] This invention extracts large-scale defects and minute flaws through multi-scale convolution, taking into account both local details and global structure to improve the model's ability to perceive defects of different sizes; it combines channel attention to automatically learn the importance of each modality and feature channel, enhances the expression of information related to defects, and suppresses irrelevant features and background interference; and it uses spatial attention to accurately focus on abnormal areas to reduce computational redundancy and improve detection efficiency, achieving high-precision, real-time detection of surface defects on complex plush fabrics.
[0094] Subsequently, temporal dynamic analysis and defect level determination are performed on consecutive frames. A sliding window mechanism is used to dynamically correlate the current frame with the two consecutive frames before and after it to analyze the stability of the defect target over time. If the same type of defect appears continuously at the same location in multiple frames, this method will determine whether to upgrade it to a higher-level defect based on the duration and area changes. For occasional interference in a single frame, comparison with preceding and following frames can be used to filter out false positives.
[0095] If feature loss is caused by fluctuations in fabric conveying speed, sensor jitter, or modal out-of-synchronization, the method automatically calls the dynamic compensation mechanism to restore the defect information of the current frame by interpolating the features of the previous and next frames, thereby realizing time continuity judgment and smooth output of grade.
[0096] Finally, the method implements hierarchical feedback control to achieve intelligent routing based on the defect level of the above output.
[0097] Grade A fabrics pass through automatically without processing; Grade B fabrics are conveyed to the re-inspection area for secondary judgment by manual assistance; Grade C plush fabrics are directly removed by robotic arms or pneumatic sorting devices to prevent unqualified products from flowing into subsequent stages.
[0098] Meanwhile, all test data, grading results, and corresponding images are stored in a quality traceability database to enable full-process data traceability, supporting subsequent big data analysis and supply chain quality optimization.
[0099] This invention enables the storage of detection data and images throughout the entire process, ensuring that the defect information of each piece of plush fabric is searchable and traceable, facilitating subsequent responsibility identification and quality review; it provides a foundation for big data analysis, assists in defect pattern recognition, production line fault early warning and long-term quality trend analysis, and improves the intelligence and controllability of the overall manufacturing process.
[0100] Example 3:
[0101] The following example, using the finished product quality inspection scenario of a plush fabric manufacturing enterprise, will be used to illustrate this embodiment in detail.
[0102] In a specific embodiment, the present invention provides an intelligent detection method for defects in plush fabrics based on multimodal fusion, the process of which is as follows:
[0103] Please see Figs. 1 to 3 The present invention proposes a multimodal fusion-based intelligent defect detection method for plush fabrics, the technical solution of which is as follows:
[0104] This method includes multimodal acquisition, lightweight parallel coding and feature fusion, a lightweight YOLOv11 model specifically for finished product inspection, a multi-scale attention mechanism, and a temporal correlation and compensation unit. Through the collaborative work of the above methods, multimodal detection and screening of surface defects on plush toys leaving the factory can be achieved.
[0105] First, multimodal appearance data collection is performed on the finished plush toys. Specifically, this includes the following operations: each plush toy to be inspected is placed on an automatic rotating platform, and an RGB industrial camera is used to collect the color texture features of the toy surface to capture visible defects such as color difference, stains, and gaps; simultaneously, an infrared thermal imaging device is used to collect surface thermal distribution maps to help detect whether the internal filling of the toy is uniform and to identify potential indentations or missing filling problems; a structured light depth camera is used to collect three-dimensional contour data of the toy surface to identify spatial structural defects such as depressions, protrusions, and misalignments.
[0106] In this system, all modal sensors are synchronously acquired via an industrial camera-triggered controller to avoid temporal misalignment. Preferably, the rotating platform allows the toy to complete a 360-degree multi-angle scan within 10 seconds, ensuring comprehensive detection without blind spots. The infrared sensing sensitivity is set to 0.1℃, and the depth measurement accuracy is preferably ±0.5mm, guaranteeing the comprehensiveness and high precision of the three-modal data.
[0107] Next, lightweight encoding and fusion processing is performed on the multimodal acquired data. Based on the acquired output, local features are extracted from the RGB, infrared, and depth data respectively through a parallel lightweight convolutional coding network, and a multidimensional feature tensor is output.
[0108] Furthermore, a shared semantic projection layer is used to map the three-modal features to a unified feature space to generate a multimodal fusion feature map. This process employs 1×1 convolution combined with BatchNorm for channel compression and feature normalization to reduce computational load and ensure the real-time inference performance of edge computing devices.
[0109] Subsequently, the fused features are input into a lightweight YOLOv11 model specifically designed for finished product inspection to identify and locate defect categories.
[0110] The model output includes: defect location coordinates, defect category label, confidence score, and corresponding defect level classification.
[0111] Preferably, the lightweight model construction process also includes a joint optimization strategy of introducing quantization-aware training and knowledge distillation. During the training phase, low-bit-width quantization simulation is pre-set, and distillation guidance is performed using the output of a high-precision teacher model, enabling the lightweight detection model to be compressed to the INT8 level. The teacher model serves as the learning target, outputting soft labels and intermediate feature representations containing more information to guide the learning of a student model with a lighter structure and lower computational cost.
[0112] This invention achieves more refined and accurate structural compression and accelerated optimization of lightweight models by introducing channel weight contribution analysis and an INT8 quantization-aware inference reconstruction strategy. Compared with traditional coarse pruning methods based on fixed ratios, this approach can dynamically identify and prune low-importance convolutional channels and redundant branches. Simultaneously, after compression, it performs INT8 quantization on all convolutional kernels and bias parameters to construct a quantization-aware inference graph, significantly reducing memory bandwidth and computational latency, and improving the model's execution efficiency on edge devices. This method not only ensures detection accuracy and semantic preservation but also significantly shortens inference time, meeting the practical requirements of low latency and high throughput in real-time quality inspection scenarios.
[0113] By extracting defect features of varying sizes through convolution at different scales, and combining this with a channel attention mechanism to strengthen feature channels that are highly correlated with defects, the model can automatically focus on defect-sensitive areas using a spatial attention mechanism, thus significantly improving the detection rate of small defects.
[0114] Subsequently, time-series correlation analysis and dynamic compensation were performed on the continuous detection data.
[0115] Because plush toys may experience occlusion or image blurring at certain angles during rotation detection, a sliding window time-series analysis of the defect features between consecutive frames is used to determine whether a defect is genuine. For short-term occlusion or missed detection in a single frame, feature interpolation between consecutive frames is used to correct and restore the defect state, resulting in a smooth and consistent detection output.
[0116] This invention effectively distinguishes between occasional interference and real defects through continuous frame dynamic analysis, reducing the false judgment rate of single-frame detection; by associating the persistence and area changes of defects through a sliding window, it supports hierarchical management of defects and improves the response capability to serious defects; combined with time dimension analysis, it adapts to continuous detection scenarios on production lines, achieving stable output and reducing detection anomalies caused by fluctuations in transmission speed, environmental interference, or changes in lighting, ensuring more accurate and reliable defect identification.
[0117] Finally, the comprehensive test results are sent to the sorting department for sorting based on defect type and grade.
[0118] Qualified products are conveyed to the packaging area; products with minor defects enter the re-inspection station; products with serious defects are directly rejected by the robotic arm or returned to the rework area. All inspection data is simultaneously uploaded to the quality inspection database to form a complete factory quality inspection report, supporting subsequent traceability and statistical analysis.
[0119] This invention dynamically corrects detection anomalies, eliminating feature loss caused by speed fluctuations or image quality changes, ensuring long-term stable system operation. It suppresses false detections caused by occasional interference, improving the temporal continuity and spatial stability of defect detection, making it suitable for real-time detection needs in continuous production environments.
[0120] This invention achieves simultaneous perception of both explicit and implicit defects by fusing RGB, infrared, and depth images, significantly improving detection comprehensiveness and accuracy. It employs lightweight parallel encoding and semantic projection to effectively address the issues of inconsistent dimensions and semantic differences in multimodal data, enhancing feature fusion and simplifying the model structure. The constructed lightweight YOLOv11 model combines structural reparameterization and hardware-aware pruning to reduce computational overhead and support edge deployment and real-time inference. A multi-scale hybrid attention mechanism is introduced to enhance the perception of both minute and complex defects. Finally, it outputs defect labels and quality levels, realizing a closed-loop detection process from data acquisition to intelligent recognition. This approach offers comprehensive advantages such as high accuracy, high speed, flexible deployment, and data traceability, making it suitable for high-efficiency industrial quality inspection scenarios.
[0121] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for intelligent identification of defects in plush fabrics based on multimodal fusion, characterized in that, include: RGB, infrared, and depth images of a plush fabric are acquired during continuous movement. A lightweight parallel encoder is constructed to extract features from the three types of images to obtain multimodal local feature information. Multimodal local feature information is mapped to a unified dimensional space through semantic projection, and joint feature maps are generated by fusing features from various modalities. Gradient operations and local morphological analysis are performed on depth images to transform them into surface undulation and villi distribution features, and surface feature maps are output. A lightweight model is constructed, and load reduction is achieved through structural reparameterization, hardware-aware compression, and dynamic computation optimization. Based on the joint feature map and the surface feature map, multimodal image features are extracted through a multi-scale hybrid attention mechanism to identify and output defect features on plush fabrics. The multi-scale hybrid attention mechanism consists of multi-scale convolutional layers, channel attention layers, and spatial attention layers working together. In the continuous conveying and inspection of plush fabrics, defect labels are determined and the defects are evaluated based on the characteristics of the defects, and the quality level evaluation results are output.
2. The intelligent identification method for defects in plush fabrics based on multimodal fusion according to claim 1, characterized in that, The steps of constructing a lightweight parallel encoder to extract features from three types of images include: The lightweight parallel encoder includes RGB image paths, infrared image paths, and depth image paths. The acquired multimodal images of continuous plush fabric are input into the corresponding paths, and modal features are extracted step by step. The modal features extracted step by step are output as multimodal local feature information.
3. The intelligent identification method for defects in plush fabrics based on multimodal fusion according to claim 1, characterized in that, The process involves mapping multimodal local feature information to a unified dimensional space through semantic projection, and fusing features from various modalities to generate a joint feature map. The steps involved in performing gradient calculations and local morphological analysis on depth images to convert them into surface undulation and villi distribution features, and outputting a surface feature map, include: A 1×1 convolutional layer with shared weights is used to uniformly map the features of RGB images, infrared images, and depth images to the same number of channels; During the mapping process, batch normalization and activation functions are used to normalize the distribution of each modality feature, and the unified features are then subjected to weighted fusion and splicing operations. Gradient calculations are performed on depth images to extract the rate of change of surface undulations. At the same time, local morphology analysis is combined to identify the shape and distribution characteristics of the microstructure of the villi, and finally, a surface feature map is output.
4. The intelligent identification method for defects in plush fabrics based on multimodal fusion according to claim 3, characterized in that, The step of fusing features from various modalities to generate a joint feature map includes: The multimodal image is passed to a sliding window segmentation, and the entire image is divided into a 256×256 pixel block sequence with a preset overlap ratio; The segmented image block sequence enters the edge gradient enhancement stage, where the Sobel operator is used to calculate the gradient magnitude map, and it is fused with the original image with preset weights to generate feature blocks that enhance texture boundaries. Perform data augmentation by randomly applying rotation, cropping, and flipping operations to the feature blocks and outputting a training sample set.
5. The intelligent identification method for defects in plush fabrics based on multimodal fusion according to claim 1, characterized in that, The steps for constructing the lightweight YOLOv11 defect detection model include: During the training phase, a multi-branch convolutional structure is constructed, and multi-scale features are extracted in parallel using convolutions of different scales. The multi-branch convolution is transformed into a single equivalent convolution through parameter merging operations.
6. The intelligent identification method for defects in plush fabrics based on multimodal fusion according to claim 1, characterized in that, The steps for load reduction through structural reparameterization, hardware-aware compression, and dynamic computation optimization include: The structural reparameterization is used to extract multi-level features from RGB images, infrared images, and depth maps; the convolutional combinations and normalization parameters during training are folded into a single-path convolutional structure. The hardware-aware compression further performs cropping, quantization, and channel rearrangement on the model based on structural reparameterization, thereby compressing the parameter scale of multimodal images. The dynamic computation optimization adjusts the computation path based on the features of the multimodal image input in each frame, skipping some computations for flawless background areas and retaining flawed areas.
7. The intelligent identification method for defects in plush fabrics based on multimodal fusion according to claim 1, characterized in that, Based on the joint feature map and the surface feature map, the multi-modal image features are extracted through a multi-scale hybrid attention mechanism to identify and output defect features on the plush fabric. The multi-scale hybrid attention mechanism is a step in which multi-scale convolutional layers, channel attention layers, and spatial attention layers work together. Combining surface feature maps from continuous fabric transport, an optimized feature map is generated by processing multimodal features through a three-layer cascaded structure: The multi-scale convolutional layer uses three sets of parallel dilated convolutional kernels of 1×1, 3×3, and 5×5 to extract microscopic fluff shaking features, defect contour features, and fabric deformation features, respectively. The multi-scale convolutional layer is used to extract local and global features of defects of different sizes. The channel attention layer dynamically generates channel masks based on the standard deviation of the temperature distribution in the infrared image, suppresses noise channels caused by reflection in the RGB features, and generates weight coefficients for each channel through lightweight fully connected layer operations, learning the contribution of different modalities and channels to defect recognition. The spatial attention layer calculates the local structural similarity between the depth map and the RGB image; it then focuses on key defect areas in the image. Output the category, size, and spatial location of the defect, and output an optimized feature map that incorporates key information from multiple modalities.
8. The intelligent identification method for defects in plush fabrics based on multimodal fusion according to claim 1, characterized in that, The steps for continuous conveying and detection of the plush fabric include: Real-time monitoring of input and defect distribution characteristics; Invoke the historical defect detection feature cache.
9. The intelligent identification method for defects in plush fabrics based on multimodal fusion according to claim 1, characterized in that, The steps of determining and outputting defect labels and evaluating defects based on defect characteristics, and outputting quality grade evaluation results include: Based on the characteristics of the defects, classify and determine the types of defects, and output category labels for damage, color difference, foreign matter, and pilling; By combining the size, location, and density of defects, the severity of the defects is assessed, and a quality grade evaluation result is generated and output.
10. The intelligent identification method for defects in plush fabrics based on multimodal fusion according to claim 1, characterized in that, The method also includes setting a timing analysis and dynamic compensation mechanism to correlate and analyze defects in consecutive frames. Specific steps include: Temporal convolution and sliding window methods are used to analyze the temporal trend of multimodal images of continuously moving plush fabrics and extract multimodal temporal features. A temporal correlation model is constructed to dynamically match the multidimensional features of consecutive frames and compare the changes in defects by comparing the spatial location information of consecutive frames with modal consistency features. To address the feature loss caused by variations in the conveying speed of plush fabrics, image blurring, and asynchronous modal data, a dynamic compensation method is employed, combining data from previous and subsequent frames for interpolation and correction to generate a continuous temporal feature sequence. The continuously time-series feature sequence after dynamic compensation is input into the time-series decision unit. Combined with the spatial location, volume and morphology of the defect, the correlation analysis of the continuous frames is performed to determine whether the defect persists, expands or disappears.
Citation Information
Cited By
Textile cloth defect real-time detection method and system based on multi-modal feature fusion
CN121527094A
Multi-sensor collaborative fabric defect detection method and device
CN121685537A
Toy plastic coating layer on-line quality detection method based on multi-mode visual perception
CN122156213A
Fabric multi-mode layered asynchronous calibration fusion estimation method and system
CN122265295A
A fabric multi-modal layered asynchronous calibration fusion estimation method and system
CN122265295B