A multi-sensor cooperative fabric defect detection method and device

By employing a multi-sensor collaborative detection method, fabric images are acquired using a dual-line array camera and an infrared camera. Combined with YOLO-FD and DeepLabv3-FD models, the problem of low efficiency and insufficient accuracy in existing textile inspection is solved, enabling efficient and accurate identification and quality assessment of defects in complex fabrics.

CN121685537BActive Publication Date: 2026-05-19WUHAN TEXTILE UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
WUHAN TEXTILE UNIV
Filing Date
2026-02-10
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing textile inspection methods rely on manual inspection, which is inefficient and highly subjective. Traditional machine vision inspection is difficult to adapt to complex and varied fabric textures and defect morphologies. Single imaging systems have limited ability to identify low-contrast defects and lack multimodal collaborative mechanisms, resulting in insufficient inspection stability and accuracy.

Method used

A multi-sensor collaborative detection method is adopted, which acquires fabric images and environmental state information through dual-line array cameras and infrared cameras, performs image registration, stitching, and multimodal fusion, combines YOLO-FD target detection and DeepLabv3-FD semantic segmentation, extracts the geometric and texture features of defects, and introduces environmental state information for dynamic compensation, so as to realize fabric quality assessment and grading.

Benefits of technology

While maintaining detection efficiency, it significantly improves the accuracy and reliability of fabric defect identification. It can stably identify defects with low contrast and weak texture disturbance, overcome high-frequency texture background interference, and achieve accurate positioning and quality assessment of small defects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121685537B_ABST
    Figure CN121685537B_ABST
Patent Text Reader

Abstract

The application provides a multi-sensor cooperative fabric defect detection method and device, and relates to the technical field of image recognition.The method comprises the following steps: collecting multi-modal information of a fabric through a double linear array camera, an infrared camera and a humidity sensor in cooperation, constructing a full-width fusion image, rapidly completing preliminary positioning of defects in a multi-scale feature space, and only acquiring a high-resolution defect image for pixel-level fine segmentation of a defect region positioned, and on the basis of the pixel-level fine segmentation, combining geometric features, texture features and environmental state information to perform dynamic compensation and comprehensive evaluation.The application can improve recognition accuracy while maintaining detection efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image recognition technology, and more specifically to a multi-sensor collaborative method and apparatus for detecting fabric defects. Background Technology

[0002] As the world's largest producer and exporter of textiles, China's textile manufacturing industry plays a crucial role in the national economy. Ensuring textile quality has become a vital link in promoting the healthy development of the industry. However, defects such as holes, broken warp threads, abrasion marks, knots, and dirt are prone to occur during textile production. These defects seriously damage the quality of textiles, and if they are not identified and addressed in a timely manner, they will affect product sales and lead to economic losses. Therefore, the detection of fabric defects has become a core process in textile quality control.

[0003] Currently, the quality monitoring system in the textile industry uses two main inspection methods. The first is manual inspection, which relies entirely on the subjective judgment of workers and is generally inefficient, with a high rate of missed detections. It struggles to comprehensively inspect wide-width fabrics and provides little real-time feedback on quality issues. The second method is traditional machine vision inspection. Algorithms based on manual feature extraction struggle to adapt to the complex and varied textures and defect morphologies of fabrics, resulting in insufficient inspection stability. Single imaging systems have limited ability to identify low-contrast defects, making it difficult to balance efficiency and accuracy. Furthermore, there is a lack of multimodal collaboration mechanisms. Cameras on the production line often operate independently, unable to transmit and utilize captured information, hindering coordination. Therefore, a method is needed to improve recognition accuracy while maintaining inspection efficiency. Summary of the Invention

[0004] This invention provides a method and apparatus for detecting fabric defects using a multi-sensor collaborative approach, which can improve recognition accuracy while maintaining detection efficiency.

[0005] A first aspect of the present invention provides a method for detecting fabric defects using a multi-sensor collaborative approach, the method comprising:

[0006] Local images and infrared images of the fabric are acquired using a dual-line array camera and an infrared camera, and environmental status information corresponding to the detection environment is obtained using a humidity sensor.

[0007] Image registration and stitching are performed on the local image to form a panoramic image; spatial registration and grayscale normalization are performed on the infrared image; and multimodal fusion is performed on the infrared image and the panoramic image to generate a fused image.

[0008] Target detection processing is performed on abnormal regions in the fused image within a multi-scale feature space, and pixel coordinates and range information corresponding to different types of defects are output, thereby forming a preliminary localization result of the defect region.

[0009] Based on the preliminary positioning results, a defect image of the defect area is acquired using an area scan camera;

[0010] Pixel-level semantic segmentation is performed on the defective regions in the defective image to output a segmentation mask corresponding to the defective regions;

[0011] Based on the segmentation mask, the geometric and texture feature parameters of the defects are extracted respectively, and the environmental state information is fused for dynamic compensation to complete the quality assessment and grading of the fabric.

[0012] In a second aspect, the present invention provides a multi-sensor collaborative fabric defect detection device, the device being used to perform a multi-sensor collaborative fabric defect detection method as described in any of the above embodiments, the device comprising an acquisition module, a processing module, and an output module, wherein:

[0013] The acquisition module is used to acquire local images and infrared images of the fabric through a dual-line array camera and an infrared camera, and to acquire environmental state information corresponding to the detection environment through a humidity sensor.

[0014] The processing module is used to perform image registration and stitching processing on the local image to form a panoramic image, perform spatial registration and grayscale normalization processing on the infrared image, and perform multimodal fusion processing on the infrared image and the panoramic image to generate a fused image.

[0015] The processing module is used to perform target detection processing on abnormal regions in the fused image within a multi-scale feature space, and output pixel coordinates and range information corresponding to different types of defects, thereby forming a preliminary localization result of the defect region.

[0016] The processing module is used to acquire a defect image of the defect area using an area scan camera based on the preliminary positioning result.

[0017] The processing module is used to perform pixel-level semantic segmentation processing on the defective region in the defective image, thereby outputting a segmentation mask corresponding to the defective region.

[0018] The output module is used to extract the geometric and texture feature parameters of the defects based on the segmentation mask, and to perform dynamic compensation by fusing the environmental state information, thereby completing the quality assessment and grading of the fabric.

[0019] In a third aspect of the invention, an electronic device is provided, including a processor, a memory, a user interface, and a network interface, wherein the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any of the preceding embodiments.

[0020] In a fourth aspect of the invention, a non-transitory computer-readable storage medium is provided, the computer-readable storage medium storing instructions that, when executed, perform the method as described in any of the preceding claims.

[0021] In summary, one or more technical solutions provided in the embodiments of the present invention have at least the following technical effects or advantages:

[0022] 1. This invention achieves high-speed continuous imaging of wide-width fabrics by using a dual-line array camera and an infrared camera working in parallel, enabling rapid localization of abnormal areas across the entire fabric width. Target detection only undertakes coarse localization to control computational overhead. Furthermore, only the initially located defect areas are triggered by the area array camera to acquire high-resolution defect images and perform pixel-level semantic segmentation. This limits high-precision, computationally intensive processing to a small area, avoiding the efficiency degradation caused by full-width fine analysis. Simultaneously, by incorporating infrared information, transmission and reflection structural information, and environmental state information through multimodal fusion, small defects are no longer judged solely by grayscale contrast but by a comprehensive feature analysis of texture continuity disruption, internal structural anomalies, and environmental compensation. This enables stable identification of low-contrast, weakly textured defects under high-speed online detection conditions, ultimately significantly improving the accuracy and reliability of fabric defect identification without sacrificing detection efficiency.

[0023] 2. By introducing the YOLO-FD target detection model into the multi-scale feature space and combining it with multi-level convolutional feature extraction, attention enhancement, and cross-scale feature fusion, defects of different scales and shapes in the fabric can be stably perceived and distinguished under high-speed detection conditions. This enables reliable preliminary localization of small defects, linear defects, and block defects while maintaining the efficiency of online detection of wide-width fabrics, providing accurate spatial constraints for subsequent fine detection.

[0024] 3. By introducing the DeepLabv3-FD semantic segmentation model, pixel-level semantic segmentation is performed on the defect image after preliminary localization, so that the defect region can obtain clear, continuous and realistic boundary segmentation results at high resolution scale. This effectively overcomes the interference of high-frequency texture background of fabric on defect boundary, improves the accuracy and stability of defect region extraction, and provides an accurate spatial basis for subsequent feature extraction.

[0025] 4. By jointly extracting the geometric and texture features of defects under the constraint of segmentation mask, and further introducing environmental state information for dynamic compensation, the morphological features, texture damage features and environmental influences of defects are modeled in a unified manner, thereby eliminating the interference of environmental changes on the consistency of detection results, realizing objective evaluation and repeatable grading of fabric quality, and improving the reliability and engineering applicability of quality judgment results.

[0026] 5. By constructing a texture consistency benchmark representation and introducing texture consistency violation analysis, the target detection process no longer relies solely on grayscale or geometric saliency, but can actively perceive the subtle disturbances caused by defects to the periodic texture structure of the fabric, thereby significantly improving the detectability of low-contrast, weak-boundary small defects and reducing the masking effect of high-frequency texture background on abnormal areas.

[0027] 6. By introducing a light-transmitting conveyor belt and an auxiliary light source to form transmitted light imaging, and constructing a dual-branch feature extraction structure for reflected light and transmitted light in the target detection stage, the model can simultaneously utilize the fabric surface texture information and internal structure information. Under the constraint of the cross-modal attention mechanism, information complementarity is achieved, thereby significantly enhancing the detection capability for defects such as internal structural anomalies and thickness differences, and improving the detection integrity in complex defect scenarios.

[0028] 7. By introducing a dual-branch coding structure for reflected and transmitted light and a cross-modal attention mechanism in the semantic segmentation stage, the pixel-level judgment process simultaneously integrates evidence of surface texture damage and evidence of internal structural anomalies. This enables a unified and accurate characterization of surface defects and internal defects at the segmentation level, significantly improving the accuracy and robustness of defect boundary segmentation and providing highly reliable segmentation results for subsequent detailed analysis and quality grading. Attached Figure Description

[0029] Figure 1 This is a schematic flowchart of a multi-sensor collaborative fabric defect detection method disclosed in an embodiment of the present invention;

[0030] Figure 2 This is a schematic diagram of a multi-sensor collaborative fabric defect grading and detection platform disclosed in an embodiment of the present invention;

[0031] Figure 3 This is a schematic diagram of a multi-sensor collaborative fabric defect detection device disclosed in an embodiment of the present invention;

[0032] Figure 4 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of the present invention.

[0033] Explanation of reference numerals in the attached drawings: 201, light-transmitting conveyor belt; 202, area array camera; 203, servo motor; 204, adjustable joint; 205, moving device; 206, auxiliary light source; 207, humidity sensor; 208, dual-line array camera; 209, infrared camera; 210, auxiliary heating lamp; 301, acquisition module; 302, processing module; 303, output module; 401, processor; 402, communication bus; 403, user interface; 404, network interface; 405, memory. Detailed Implementation

[0034] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0035] In the description of the embodiments of the present invention, words such as "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "for example" or "for instance" in the embodiments of the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Rather, the use of words such as "for example" or "for instance" is intended to present the relevant concepts in a specific manner.

[0036] In the description of the embodiments of the present invention, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0037] Textile production is prone to various defects such as holes, broken warp threads, abrasion marks, knots, and dirt, which seriously affect product quality and economic benefits. Existing quality inspection methods mainly rely on manual inspection or traditional machine vision inspection. Manual inspection is inefficient, subjective, and difficult to cover wide fabrics. Traditional machine vision inspection is limited by single imaging and manual feature extraction methods, making it difficult to adapt to the actual working conditions such as complex fabric textures, varied defect morphologies, and low-contrast defects. At the same time, it lacks multimodal collaboration and equipment linkage mechanisms, resulting in insufficient detection stability and accuracy. Therefore, there is an urgent need for a detection method that can improve the accuracy and reliability of fabric defect identification through multi-source information collaboration while ensuring detection efficiency.

[0038] This embodiment discloses a multi-sensor collaborative method for fabric defect detection, referring to... Figure 1 This includes the following steps S110-S160:

[0039] S110 acquires local and infrared images of the fabric using a dual-line array camera and an infrared camera, and obtains environmental status information corresponding to the detection environment using a humidity sensor.

[0040] This invention discloses a multi-sensor collaborative fabric defect detection method applied to a server. The server includes, but is not limited to, electronic devices such as mobile phones, tablets, wearable devices, and PCs (Personal Computers), and can also be a backend server running a multi-sensor collaborative fabric defect detection method. The server can be implemented using a standalone server or a server cluster composed of multiple servers.

[0041] Reference Figure 2 This invention discloses a multi-sensor collaborative fabric defect grading and detection platform, including a light-transmitting conveyor belt 201, an area array camera 202, a servo motor 203, an adjustable joint 204, a moving device 205, an auxiliary light source 206, a humidity sensor 207, a dual-line array camera 208, an infrared camera 209, and an auxiliary heating lamp 210. The light-transmitting conveyor belt 201 is used to carry and continuously transport the textile to be inspected, so that the textile maintains a stable and uniform running state within the detection area to meet the imaging conditions for continuous online detection.

[0042] A phase-array camera 202 is positioned above the light-transmitting conveyor belt 201 to acquire real-time high-definition images of the fabric in motion. A servo motor 203 is connected to the high-speed phase-array camera 202 and controls its position adjustment and start / stop status during the inspection process. An adjustable joint 204 and a moving device 205 are coordinated to adjust the spatial orientation and inspection position of the phase-array camera 202 to adapt to the imaging requirements of different inspection areas. An auxiliary light source 206 is positioned below the light-transmitting conveyor belt 201 in the inspection area to provide stable illumination conditions for transmission imaging, thereby enhancing the imaging effect of the internal structural features of the fabric. Humidity sensors 207 are positioned on both sides of the light-transmitting conveyor belt 201 in the inspection area to simultaneously acquire ambient humidity data during fabric inspection and provide environmental status information for subsequent dynamic compensation of image features. The dual-line array camera 208 and the infrared camera 209 are jointly installed on the upper transverse guide rail of the first adjustable gantry. The dual-line array camera 208 is used to perform continuous scanning imaging on the fabric surface along the fabric width direction to obtain a high-resolution visible light image of the fabric surface. The infrared camera 209 is used to acquire an infrared thermal imaging image that corresponds to the visible light image in time and space. The auxiliary heating lamp 210 is set on both sides of the transverse guide rail of the first adjustable gantry to stabilize the infrared imaging conditions.

[0043] The dual-line array camera 208 and the infrared camera 209 are arranged in parallel in space and achieve imaging alignment through synchronous control. This enables them to collaboratively acquire surface texture features and internal thermal features of the fabric, achieving stable capture of fabric surface texture and defects. Through the collaborative control of the above-mentioned multiple devices and the fusion processing of multi-source sensor information, the device can simultaneously complete fabric motion control, multi-view image acquisition, and environmental parameter monitoring during continuous fabric operation. Based on the real-time acquired multimodal information, it performs dynamic correction and compensation, thereby providing a complete and reliable data foundation for fabric defect identification and quality analysis, and achieving accurate and efficient online monitoring of fabric defects.

[0044] While the fabric is in continuous transport, a dual-line array camera 208 mounted on a first adjustable gantry performs synchronous line scanning imaging of the fabric surface along the fabric width direction. This allows the dual-line array camera 208 to continuously acquire high-resolution visible light local images covering a local width range at a preset line frequency during fabric transport. Each visible light local image corresponds one-to-one with the fabric transport position in the time dimension, thus completely recording the continuous changes in the fabric surface texture along the transport direction. Simultaneously with the line scanning imaging performed by the dual-line array camera 208, an infrared camera 209, also mounted on the first adjustable gantry, performs synchronous infrared imaging of the fabric within the same detection area to acquire infrared images that correspond to the visible light local images in time and space. The infrared camera 209... The field of view is adjusted to match the scanning area of ​​the dual-line array camera 208 by adjusting the pitch and yaw angles of the pan-tilt unit. Auxiliary heating lamps 210 are installed on both sides of the horizontal guide rail of the first adjustable gantry to stabilize the infrared radiation conditions on the surface and inside of the fabric, thereby improving the response capability of the infrared image to changes in fabric thickness and internal structural anomalies. While completing the acquisition of visible light local images and infrared images, humidity sensors 207 deployed in the detection areas on both sides of the transmission unit are used to collect humidity parameters in the detection environment in real time. The humidity parameters are synchronously correlated with the visible light local images and infrared images according to the acquisition time, so that the humidity parameters, as environmental state information, establish a mapping relationship with the imaging data at the corresponding time, thereby forming a raw observation data set containing visible light local images, infrared images, and environmental state information.

[0045] S120 performs image registration and stitching processing on the local image to form a panoramic image, performs spatial registration and grayscale normalization processing on the infrared image, and performs multimodal fusion processing on the infrared image and the panoramic image to generate a fused image.

[0046] After acquiring visible light local images continuously captured by the dual-line array camera 208, the visible light local images are first initially sorted according to time sequence based on the imaging parameters of the dual-line array camera 208 and the fabric transmission speed, so that adjacent visible light local images form a continuous relationship in the fabric transmission direction. On this basis, image registration processing is performed on the imaging overlap area between adjacent visible light local images in the width direction. By matching the texture features in the overlap area, the spatial correspondence between adjacent visible light local images is determined, and the visible light local images are geometrically corrected according to the spatial correspondence. After completing the geometric correction, the registered visible light local images are image stitched according to the preset stitching order, so that multiple visible light local images are continuously connected in space, thereby forming a panoramic image covering the full width of the fabric, which is used to completely characterize the overall distribution of the fabric surface texture.

[0047] While constructing the panoramic image, spatial registration processing is performed on the infrared image acquired synchronously with the visible light local image. This establishes a one-to-one correspondence between the infrared image and the panoramic image at the pixel level. The spatial registration processing is based on geometric mapping correction of the installation position relationship and imaging field of view relationship between the infrared camera 209 and the dual-line array camera 208. After completing the spatial registration, grayscale normalization processing is performed on the infrared image. By uniformly mapping the infrared grayscale values, the influence of changes in infrared radiation intensity at different acquisition times or under different environmental conditions on the imaging results is eliminated, making the infrared images comparable in grayscale distribution.

[0048] After completing the panoramic image construction and the spatial registration and grayscale normalization of the infrared image, under the constraint of pixel-level alignment, multimodal fusion processing is performed on the infrared image and the panoramic image. By fusing the visible light information in the panoramic image that represents the continuity of the fabric surface texture with the infrared information in the infrared image that represents the internal structure and thickness difference of the fabric, the fused image contains both the fabric surface texture features and the internal thermal response features, thereby generating a fused image for defect detection and analysis.

[0049] S130 performs target detection processing on abnormal regions in the fused image within a multi-scale feature space, outputting pixel coordinates and range information corresponding to different types of defects, thereby forming preliminary localization results of defect regions.

[0050] After the fused image is fed into the pre-trained YOLO-FD object detection model, the backbone feature extraction network performs multi-level convolutional feature extraction on the fused image, enabling the same fused image to form multi-scale first feature representations at different downsampling ratios. The backbone feature extraction network refers to the main structure of the convolutional network used to progressively extract semantic features from the original image. Multi-level convolutional feature extraction refers to the sequential execution of convolution and non-linear mapping at multiple depth layers, so that shallow features emphasize texture details and deep features emphasize semantic contours. The receptive field scale refers to the area covered by the original image corresponding to a single feature point on the feature layer. The size of the receptive field is important; a larger receptive field is better for expressing the overall outline of blocky defects, while a smaller receptive field is better for expressing the local texture perturbations of small defects. The fused image is first subjected to continuous convolution and downsampling to form several layers of feature maps. Each layer of feature map is used as the first feature representation to participate in the subsequent fusion. This makes small defects more likely to be preserved in high-resolution feature layers, linear defects have both length continuity and texture difference in medium-resolution feature layers, and blocky defects are more likely to form stable semantic aggregation responses in low-resolution feature layers. This provides a homogeneous and complementary feature basis for subsequent cross-scale fusion.

[0051] In the feature extraction process of the backbone feature extraction network, a feature enhancement process combining channel attention and spatial attention is introduced into at least one intermediate feature layer. This allows for the synchronous enhancement of the channel responses and spatial location responses related to defects in the first feature representation. The intermediate feature layer refers to a feature layer that retains certain texture details while possessing a certain semantic abstraction capability. Channel attention involves assigning weights to different semantic channels in the channel dimension to highlight defect-sensitive channels, while spatial attention involves assigning weights to potential defect locations in the spatial dimension to suppress regular texture backgrounds. In practical implementation, a CBAM attention structure can be used to perform channel attention followed by spatial attention on the input feature map. Channel attention obtains channel descriptions through global average pooling and global max pooling, and generates channel weight vectors through a multilayer perceptron with shared parameters. Then, the input feature map is weighted by channel. The formula for calculating channel attention is:

[0052]

[0053] Where F represents the feature map of the intermediate feature layer to be enhanced. This represents the channel description vector obtained by performing global average pooling on F in the spatial dimension. This represents the channel description vector obtained by performing global max pooling on F in the spatial dimension. A multilayer perceptron with shared parameters is used to map channel descriptions to channel weights. The Sigmoid function is used to normalize the weights to the range of 0 to 1. This formula characterizes channel importance through the complementarity of average and extreme value statistics, making channels related to defects respond more strongly after weighting, while suppressing the responses of channels related to the periodic texture of the fabric. Subsequently, spatial attention generates a spatial weight map after channel aggregation to highlight the defect location. The formula for calculating spatial attention is:

[0054]

[0055] in, This represents the feature map after channel attention weighting. This represents the single-channel spatial map obtained by average pooling along the channel dimension. This represents the single-channel spatial graph obtained by max pooling along the channel dimension. This indicates that two spatial maps are stitched together along the channel dimension. This indicates that a 7×7 convolution is used to fuse spatial context and generate spatial weights. The Sigmoid function is used to normalize spatial weights to the range of 0 to 1. This formula remodels the spatial response after channel aggregation, making the defect area form a more concentrated high response in space, thereby achieving joint weighted enhancement of the response intensity of defect-related feature channels and the response intensity of defect-related spatial locations.

[0056] After obtaining the first feature representations at multiple scales, upsampling, downsampling, and feature concatenation are performed on these representations through a neck feature fusion network. This enables cross-scale fusion of first feature representations from different levels within a unified semantic space. The neck feature fusion network refers to the feature integration structure connecting the backbone feature extraction network and the detection head. Upsampling involves enlarging low-resolution feature maps to higher resolutions to supplement higher-level semantics and details. Downsampling involves shrinking high-resolution feature maps to lower resolutions to supplement details and semantics. Feature concatenation refers to combining features from different scales in the channel dimension. After the feature maps are aligned, they are merged into a single fusion feature map. First, the deep semantic features are upsampled and spatially aligned with the shallow detail features, and then stitched together. This allows the small defects and texture perturbations preserved in the shallow layer to obtain the category discrimination support of the deep semantics. At the same time, some shallow features are downsampled and stitched together with the mid-deep features, so that the continuous structure of linear defects maintains consistent expression under a larger receptive field. Through multiple scale alignments and stitchings, small defects, linear defects, and block defects share comparable feature descriptions in a unified semantic space, thereby providing the detection head with fusion feature input that contains both details and semantics.

[0057] Within a unified semantic space, the detection head performs defect category prediction and defect bounding box regression on the fused defect features. It outputs defect category information, defect pixel coordinates of the defect bounding box in the fused image, and the spatial distribution range of the defect in the fabric width and transmission directions. Here, "detection head" refers to the network branch that performs the final prediction of the fused features; "defect category prediction" refers to the probability determination of which defect category a candidate region belongs to; "defect bounding box regression" refers to the continuous value estimation of the position and size of the candidate region's bounding box; "defect pixel coordinates" refers to the pixel-level position parameters of the bounding box in the fused image coordinate system; and "spatial distribution range" refers to mapping the coverage area of ​​the bounding box in the fused image coordinate system to the coverage area in the fabric width and transmission directions to support subsequent directional imaging by the area array camera. For each spatial location, the detection head outputs a category probability vector and a bounding box parameter vector for the fused features. The bounding box parameters can be in the form of a center point and width / height. Or in the form of top left and bottom right corners To enhance training stability, relative offset regression can be used, whereby the offset between the predicted bounding box and the anchor box is defined as:

[0058]

[0059] in, The coordinates of the anchor frame's center and its width and height are preset or generated by the detection head on feature layers at various scales and correspond to the positions of the fused features. This indicates the center coordinates and width and height of the predicted bounding box. This represents the normalized translation offset of the center coordinates relative to the anchor frame. This represents the logarithmic scale offset of width and height relative to the anchor frame. By normalizing the translation offset and logarithmically compressing the width and height changes, the bounding box regression of defects at different scales is numerically more stable and conducive to unified optimization. This yields the defect pixel coordinates of the defect bounding box in the fused image, and, combined with the bounding box coverage area, determines the spatial distribution range of defects in the width and transmission directions, ensuring that the detection output can be directly used as a constraint for subsequent area scan camera positioning and shooting parameter adjustment.

[0060] When generating preliminary localization results based on defect pixel coordinates and spatial distribution range, the bounding box pixel coordinates output by the detection head are bound to the coordinate system of the fused image. This ensures that each defect obtains a unique spatial location identifier at the panoramic scale, and this spatial location identifier is converted into localization information that can drive subsequent device linkage. The preliminary localization result refers to a reusable localization description of the defect area at the panoramic scale. The panoramic scale refers to a unified coordinate scale covering the entire width of the fabric in the panoramic image and its aligned fused image. In specific implementation, confidence filtering and overlap suppression are first performed on multiple candidate bounding boxes of the same defect to obtain stable... Define the bounding box, and then write the center pixel coordinates of the bounding box and the bounding box coverage area as the core fields for defect localization into the localization record. Establish the association between the localization record and the fabric transmission time and corresponding frame sequence, so that subsequent steps can drive the second and third adjustable gantry frames to move the area array camera to the spatial position corresponding to the defect center and adjust the shooting window according to the defect coverage area based on the preliminary localization result. This realizes the continuous spatial transfer from panoramic detection to local fine inspection, ensuring the consistency and coherence of the target detection step with the subsequent defect image acquisition step and semantic segmentation step in terms of spatial position and semantic object.

[0061] The YOLO-FD object detection model uses panoramic images or multimodal fused images obtained by image stitching as a unified input, enabling wide-width fabrics to form a continuous field of view in the same coordinate system. This avoids the positioning breaks caused by segmented scanning and allows the subsequent output defect bounding boxes to directly correspond to the spatial distribution range of the fabric width and transmission direction. After the input enters the backbone, it undergoes progressive feature extraction through multiple Conv and C3K2 structures. Conv is used to establish local representations of basic textures and edges, while C3K2 is used to enhance feature reuse and multi-scale expression under controllable computational requirements. This allows shallow features to retain more high-frequency texture details of the fabric, while mid-to-deep features gradually aggregate the morphological and contextual differences of defects. Deep features form a more stable semantic discrimination criterion, thus providing a distinguishable yet fusionable feature base for small defects, linear defects, and blocky defects. Because the fabric background texture has obvious periodicity and often has a weak contrast relationship with small defects, the Backbone inserts a CBAM attention structure after the second C3K2 module. Channel attention highlights feature channels sensitive to defects, and spatial attention highlights the location of potential defects and suppresses the high response of regular texture main frequency regions. This transforms the response of defects such as broken warp, scratches, knots, and dirt in the feature map from a weak response submerged by texture to a sparse and significant response amplified by attention, thereby improving the detectability in small-scale and low-contrast scenes. The SPPF structure at the bottom of the Backbone is used to expand the effective receptive field and aggregate multi-scale context at a low cost, enabling the model to obtain a larger semantic coverage without significantly reducing resolution. This provides a more stable basis for judging the overall outline of block defects, the continuous extension of linear defects, and the spatial consistency of local texture anomalies. The C2PSA structure adjacent to SPPF is used to further enhance the selective response and feature recalibration of key regions, so that deep semantic features obtain a clearer defect-related expression before entering the Neck, reducing the risk of being pulled by background texture in the subsequent fusion stage.

[0062] Upon entering the Neck, the model employs a cross-scale fusion path centered on Upsample and Concat. Deep semantic features are upsampled and spatially aligned with shallow detail features before being stitched together. This allows the high-resolution texture differences relied upon by small defects to receive deep semantic category discrimination support. Simultaneously, some features are fused again via a downlink path to supplement the detail consistency of the mid-to-low resolution layers, achieving a balance between length continuity and boundary separability for linear defects. Ultimately, multi-level fused features are formed for use by different detection heads. The figure introduces a CA attention structure before two key fusion nodes. CA jointly encodes channel dependence and positional information, enabling the attention to not only answer which channels are more important but also which positions of important channels are more important in the width and transmission directions. This is more effective for fabric defects, which have directional background textures, strong positional correlation, elongated shapes, or sparse distribution. It retains more accurate positioning cues in the fused features of the Neck and reduces the impact of scale changes caused by stitching on positioning. The Head section is configured with multi-scale Detect branches, each corresponding to a fusion feature layer of different resolutions. The small-scale Detect branch targets fine and dot-like defects, the medium-scale Detect branch targets linear defects and medium-sized stains, and the large-scale Detect branch targets blocky holes and large-area scratches. Each Detect branch outputs defect category information and bounding box regression results for its corresponding scale candidate region, and provides the defect pixel coordinates and spatial coverage of the defect bounding box in the fusion image coordinate system, thus forming a preliminary localization result of the defect region. This preliminary localization result can serve as a spatial constraint for subsequent area scan camera directional shooting and DeepLabv3-FD pixel-level segmentation.

[0063] Furthermore, the fabric surface is formed by the regular interweaving of yarns, naturally exhibiting a high-frequency texture structure with significant periodicity and directionality. In this context, minor defects usually do not manifest as strong gray-scale abrupt changes or clear geometric boundaries, but only as local, slight disturbances to the original weaving pattern. For example, the interruption of texture continuity caused by micro-broken yarns, periodic misalignment caused by skipped patterns, texture energy attenuation caused by slight abrasion, or local directional disorder caused by the exposure of a single yarn. These disturbances are extremely weak in amplitude and spatially scattered, and are easily masked by the main texture frequency, making it difficult for minor defects to be effectively distinguished and reliably identified in conventional imaging and detection methods based on significant edges or gray-scale contrast.

[0064] When constructing a texture consistency benchmark representation for a fused image, the fused image is divided into a continuous set of local regions according to the width and transmission directions. Periodic and directional parameters for characterizing the dominant texture structure are extracted in each local region, and this set of parameters forms a texture consistency benchmark representation. The texture consistency benchmark representation is used to characterize the stable form of the dominant texture direction, dominant texture period, and dominant texture energy distribution generated by the yarn arrangement in the normal fabric region. The texture consistency benchmark representation can be composed of a local texture direction field and a local periodic field. The local texture direction field is used to characterize the dominant direction distribution of the texture in space, and the local periodic field is used to characterize the stable distribution of the texture repetition interval in space. The dominant texture structure refers to the texture component that dominates the energy in the frequency domain or spatial domain, which corresponds to the regular repetition pattern of the yarn arrangement in the normal fabric. This allows subsequent analysis to use the dominant texture structure as a normal reference rather than relying solely on gray-level abrupt changes as an abnormal reference.

[0065] When texture consistency violation analysis is performed under the constraint of the texture consistency benchmark representation, texture direction continuity metric, texture energy distribution stability metric, and texture repetition period offset metric are calculated for each local region. These three metrics are then jointly represented as texture violation intensity to reflect the degree of perturbation of the local weaving pattern. Specifically, the texture direction continuity metric characterizes the deviation of the texture direction from the texture consistency benchmark representation within the local region; the texture energy distribution stability metric characterizes the attenuation or anomalous enhancement of the texture energy relative to the benchmark representation; and the texture repetition period offset metric characterizes the misalignment of the texture repetition interval relative to the benchmark representation. To enable the three metrics to be jointly represented at the same scale, they are normalized and weighted fused to generate the texture violation intensity of the local region. The fusion expression for the texture violation intensity is as follows:

[0066]

[0067] Where p represents the center position or pixel index of the local region. This indicates the intensity of texture destruction at position p. This represents the normalized texture orientation deviation metric, which is obtained from the angle difference between the local texture orientation at position p and the main texture orientation in the texture consistency benchmark representation and mapped to a range of 0 to 1. This represents the normalized texture energy difference measure, which is obtained from the relative difference between the local texture energy at position p and the dominant frequency texture energy in the texture consistency benchmark representation, and mapped to the range of 0 to 1. This represents the normalized texture period offset metric, which is determined by position. The relative offset between the local texture period and the main texture period in the texture consistency benchmark representation is obtained and mapped to the range of 0 to 1. , , These represent the fusion weights of the three metrics, with values ​​ranging from 0 to 1 and satisfying the following conditions: This formula weights and superimposes three complementary sources of texture disturbance—direction deviation, energy difference, and period shift—so that the subtle texture damage caused by minor defects can be amplified and expressed in the comprehensive metric. It also allows for consistent response under different fabric texture conditions through weight adjustment, thereby generating a texture damage response map that reflects the degree of disturbance to local weaving patterns.

[0068] When spatially aligning the texture destruction response map with the fused image, the relationship between the texture destruction response map and the fused image originating from the same coordinate system is utilized to align each position of the texture destruction response map. Same location in the fused image Pixel-level correspondences are established, and scale mapping is performed for cases with resolution differences to ensure that the spatial resolution of the texture destruction response map is consistent with that of the fused image, thereby obtaining additional features. These additional features are used as texture destruction cue information synchronized with the fused image in subsequent target detection processing, enabling target detection processing to directly use the texture destruction response map to provide saliency for potential defect locations without changing the pixel coordinate system of the fused image. Spatial alignment refers to ensuring that the same pixel location represents the same fabric spatial location under a unified coordinate system, and scale mapping refers to interpolating or aggregating the texture destruction response map to match the resolution requirements of the fused image.

[0069] When additional features are introduced into the target detection process, they are either used as an extra channel at the same scale as the fused image and concatenated with the fused image, or embedded as attention-guided signals into the intermediate feature layer of the backbone feature extraction network. This allows target detection to simultaneously represent abnormal regions based on gray-level variation features, geometric features, and texture consistency disruption features in a multi-scale feature space. Gray-level variation features refer to local change information represented by the pixel intensity differences and spatial gradients of the fused image. Geometric features refer to the shape information jointly represented by the edge direction, regional connectivity, aspect ratio, and scale-level response of the abnormal region. Texture consistency disruption features refer to the comprehensive perturbation information represented by the directional deviation, energy anomaly, and periodic misalignment represented by the additional features. By establishing a correlation between the additional features and the feature representations in the multi-scale feature space, even small defects with extremely weak gray-level differences and blurred boundaries can form candidate responses that can be used by the detection head based on texture consistency disruption features. This reduces the masking effect of high-frequency texture background on abnormal regions and improves the recall ability of small defects.

[0070] When the detection head performs defect category discrimination and defect bounding box regression processing on abnormal areas, it generates a set of candidate regions based on the multi-scale feature representation after introducing additional features. For each candidate region, it outputs the defect category probability and bounding box parameters. The defect category discrimination result is used to distinguish different defect types such as broken warp, skipped stitches, abrasion marks, knots, and dirt. The defect bounding box regression result is used to provide the defect pixel coordinates and spatial range information of the defect in the fused image coordinate system. The defect pixel coordinates are used to characterize the position parameters of the bounding box in the fused image, and the spatial range information is used to characterize the pixel interval covered by the bounding box and can be mapped to the coverage interval in the fabric width direction and transmission direction, thus forming a preliminary positioning result. The preliminary positioning result binds the defect category information with the defect pixel coordinates and spatial range information, so that the subsequent area array camera directional imaging and pixel-level semantic segmentation can use the same object identifier and the same spatial coordinate constraints.

[0071] S140: Based on the preliminary positioning results, acquire defect images of the defect area using an area array camera.

[0072] After obtaining the preliminary positioning results output by the target detection processing, the preliminary positioning results are used as a unified basis for spatial and temporal constraints to drive the area scan camera 202 to perform directional imaging of the defect area. The preliminary positioning results include the defect pixel coordinates of the defect in the fused image coordinate system and the corresponding spatial range information. The defect pixel coordinates are used to determine the center position of the defect in the width direction, and the spatial range information is used to characterize the coverage area of ​​the defect in the transmission direction. Based on the pre-established mapping relationship between the fused image coordinate system and the physical coordinate system of the detection platform, the defect pixel coordinates are converted into the corresponding physical position information. Combined with the real-time transmission speed of the fabric and the time of defect appearance, the arrival time of the defect on the detection platform is predicted, thereby forming the positioning control parameters used to drive the movement of the area scan camera 202 and trigger acquisition.

[0073] Under the constraints of positioning control parameters, the area scan camera 202, mounted on an adjustable gantry, is moved along the width direction to align its imaging center with the center of the defect area. Simultaneously, the size of the shooting window and the field of view coverage of the area scan camera 202 are adjusted according to spatial range information to ensure that the defect area is completely covered within the imaging field of view of the area scan camera 202. After the area scan camera 202 completes the positioning alignment, the exposure time of the area scan camera 202 is synchronously controlled based on the predicted arrival time of the defect in the transmission direction, so that the area scan camera 202 completes image acquisition the instant the defect area enters the imaging field of view, thereby avoiding imaging offset or omission caused by continuous fabric movement.

[0074] During the defect image acquisition process, based on the size information of the defects in the preliminary localization results, the focal length and shooting height of the array camera 202 are adaptively adjusted to ensure that the defect area occupies sufficient pixel resolution in the acquired defect image, so as to meet the requirements of subsequent pixel-level semantic segmentation for boundary accuracy and texture details. Among them, the defect size information is used to distinguish between small defects and large-scale defects, so that small defects can obtain higher magnification during imaging, while larger defects can take into account both overall shape and local details during imaging, thereby forming a scale-matched defect image.

[0075] S150: Perform pixel-level semantic segmentation processing on the defective region in the defective image, and output the segmentation mask corresponding to the defective region.

[0076] After the defect image is fed into the pre-trained DeepLabv3-FD semantic segmentation model, the model's encoding structure performs multi-level feature encoding processing on the defect image, enabling the same defect image to form multi-scale second feature representations at different downsampling scales. Here, the encoding structure refers to the first half of the network used to extract semantic features and compress spatial resolution step by step, the multi-level feature encoding refers to the feature levels formed at different network depths, and the downsampling scale refers to the resolution ratio of the feature map relative to the original defect image. In this process, the shallow second feature representation focuses on characterizing the local texture anomalies of the defect area, such as fine-grained grayscale and texture changes caused by yarn breakage, exposed fuzz, or slight abrasion. The mid-deep second feature representation gradually integrates the contextual relationship between the defect area and the surrounding fabric texture background. The deep second feature representation forms a stable expression of the overall semantic attributes of the defect, so that the defect area has both distinguishable local features and global semantic support at different scales.

[0077] During feature encoding, a multi-branch dilated convolution structure is set in the encoding structure, enabling the second feature representation to obtain contextual information within different receptive fields without further reducing spatial resolution. Dilated convolution refers to introducing interval sampling inside the convolution kernel to expand the receptive field without increasing the number of parameters. Multi-branch refers to setting multiple sets of convolutional paths with different dilation rates in parallel. The receptive field refers to the size of the original image coverage area corresponding to a single position in the feature map. Convolutional branches with different dilation rates correspond to contextual ranges of different scales, allowing small defects to maintain boundary clarity in branches with small receptive fields, and linear or blocky defects to obtain more complete structural information in branches with large receptive fields. The features output by multiple branches together constitute the second feature representation, thereby improving the model's adaptability to defects with significant scale changes and irregular shapes.

[0078] After obtaining the multi-scale second feature representation, feature fusion is performed on the second feature representation to combine defect features at different scales within a unified semantic space. Feature fusion refers to splicing or weighted integration of feature maps from different scales after spatial alignment. A unified semantic space means that the fused features have consistent category discrimination meaning at the semantic level. A spatial attention mechanism is introduced during feature fusion to assign enhanced weights to the spatial location of defect regions. Spatial attention refers to assigning different weights to different locations in the spatial dimension, making the model pay more attention to potential defect regions and suppress the response of regular fabric texture regions. Spatial attention can be achieved by statistically analyzing the fused features in the channel dimension and generating a spatial weight map, the expression of which is:

[0079]

[0080] Where F represents the second feature representation after fusion. This represents the average pooling result along the channel dimension, used to reflect the average response intensity at each spatial location. This represents the max pooling result along the channel dimension, used to reflect the extreme response at each spatial location. This indicates that splicing is performed at the channel level. Indicates the kernel size as Convolution operations are used to fuse spatial context information. The Sigmoid function is used to map spatial weights to the range of 0 to 1. This formula generates spatial weights by combining the average response and the extreme response, so that defective regions can gain more attention in the fusion features.

[0081] After feature fusion is completed, the fused second feature representation is subjected to progressive upsampling through the decoding structure. This process gradually restores the feature map to a spatial resolution consistent with the defect image. The decoding structure refers to the latter half of the network used to restore the spatial resolution and output pixel-level prediction results. Progressive upsampling refers to gradually increasing the feature map resolution in reverse order of the downsampling path in the encoding stage. Boundary enhancement processing is introduced into the decoding structure to strengthen the feature responses corresponding to the edges of the defect region. This ensures that the defect boundaries remain clear and continuous during the upsampling process. Boundary enhancement assigns higher weights to locations with significant gradient changes in the features to prevent the defect boundaries from being smoothed by background textures or interpolation processes, thereby improving the fit between the pixel-level segmentation results and the true contours of the defects.

[0082] After decoding and obtaining feature representations consistent with the defect image space, the DeepLabv3-FD semantic segmentation model performs category determination processing on each pixel in the defect image at the pixel level, distinguishing pixels belonging to the defect region from pixels belonging to the background region. Pixel-level category determination refers to outputting the probability distribution of the category to which each pixel belongs. Based on the pixel-level determination results, the set of pixels determined to be of the defect category forms a segmentation mask, so that the segmentation mask corresponds one-to-one with the defect region in space. The segmentation mask is used to clearly identify the precise location and boundary range of the defect region in the defect image.

[0083] The DeepLabv3-FD semantic segmentation model takes the defect image constrained by the initial localization results, allowing the model to avoid searching for the target in the full-width background. Instead, it performs pixel-level analysis of the defect region on the high-resolution defect image acquired by the area scan camera, thus outputting a segmentation mask that corresponds one-to-one with the defect region. The segmentation mask is used as a unified spatial constraint for subsequent geometric feature extraction, texture feature extraction, and quality assessment and grading. Due to the periodicity and directionality of fabric textures, defects often exhibit weak contrast local perturbations with blurred boundaries. This model uses a combination of multi-scale context aggregation in the encoding structure and detail restoration and boundary enhancement in the decoding structure to ensure that defect boundaries can be stably separated from the high-frequency texture background.

[0084] In the Encoder section, DCNN serves as the backbone encoding structure to perform multi-level feature encoding on the input defect image, forming a feature representation that includes low-level details and high-level semantics. Atrous Conv is used to expand the receptive field without significantly reducing spatial resolution, enabling the model to utilize a wider range of context to distinguish normal periodic variations in texture from abnormal variations caused by broken veins, skipped patterns, abrasions, knots, and dirt. The parallel multi-branch structure within the Encoder corresponds to ASPP-style multi-scale context aggregation. Each branch models the same feature at different receptive field scales, including a 1×1 convolutional branch to preserve local detail baselines, a 3×3 Deformable Conv branch for adaptive sampling of irregular defect boundaries, allowing the convolutional sampling position to shift with the defect morphology, thus better fitting the contours of linear defects, exposed fuzz, or broken boundaries, a 3×3 Depthwise Sep Conv branch to enhance the expression of texture and boundary details under controllable computational requirements, 3×3 convolutional branches with different hole rates to obtain contextual relationships at different scales, and an Img branch. The pooling branch is used to introduce global semantic priors to suppress missegmentation caused by large-scale texture backgrounds; the outputs of multiple branches are spliced ​​to form a multi-scale second feature representation, so that the defect region has both local texture anomaly features and contextual semantic relationship features.

[0085] After the multi-branch outputs of the Encoder converge, the diagram uses 1×1 convolution for channel compression and semantic reorganization, and introduces Spatial Attention to reweight spatial locations. This causes the model to give higher responses to locations that are more likely to belong to defect areas and lower responses to regular texture areas in the spatial dimension. The role of Spatial Attention is to explicitly inject the spatial saliency of defects into the subsequent decoding process, thereby reducing the interference of high-frequency textures of the fabric on the judgment of defect area boundaries. The high-level semantic features output at this stage contain multi-scale context and spatial saliency while maintaining a certain spatial resolution, and are the main semantic basis for the subsequent generation of segmentation masks.

[0086] In the Decoder part, the model extracts Low-Level Features from the encoding structure and performs channel dimensionality reduction using 1×1 convolutions. This allows the edge and texture details preserved in the low-level features to participate in the fusion with a controllable channel scale. The Low-Level Features are used to compensate for the loss of detail caused by semantic abstraction during the encoding stage, especially for restoring the true boundary direction of small defects. The diagram shows that CBAM is introduced in both the Low-Level Features path and the high-level semantic path. CBAM enhances defect-related features and suppresses background texture responses through joint channel attention and spatial attention, making the fusion in the decoding stage more inclined to preserve defect edges rather than amplify the periodic texture of the fabric. Subsequently, the high-level semantic features are upsampled and fused with the CBAM-enhanced Low-Level Features. Features are concatenated to combine high-level semantic discrimination capabilities with low-level boundary details in a unified semantic space. Then, 3×3 convolution is used to reconstruct the local consistency of the fused features, which enhances the internal connectivity and boundary continuity of the defect region. Finally, upsampling is performed to restore the spatial resolution consistent with the defect image and output pixel-level segmentation results. The pixel-level results are classified to form a segmentation mask, thereby distinguishing the defect region from the background region at the pixel level and maintaining the spatial constraints consistent with the subsequent feature extraction and quality grading steps.

[0087] Furthermore, referring to Figure 2This invention introduces structural improvements to the light-transmitting conveyor belt 201 and the auxiliary light source 206 below, enabling imaging to move beyond the apparent reflection information of the fabric and acquire the internal yarn structure and thickness difference features of the fabric formed by bottom transmitted light. Therefore, based on the enhanced structural information of transmitted light, a dual-branch feature extraction mechanism is introduced. A transmitted light feature channel is added to both the YOLO-FD target detection model and the DeepLabv3-FD semantic segmentation model, allowing the models to learn the texture continuity corresponding to the reflection map and the yarn density distribution corresponding to the transmitted map, respectively. Furthermore, a cross-modal attention mechanism is employed in the neck feature fusion stage to enhance the expression of thickness anomalies in the transmitted light region, preventing them from being masked by high-frequency textures, thereby improving the detectability of extremely low-contrast fine points and internal defects.

[0088] Transmitted light imaging is performed in conjunction with the light-transmitting conveyor belt 201 and the auxiliary light source 206 below. The light-transmitting conveyor belt 201 is used to allow light from below to penetrate the fabric and form a transmitted light field while carrying and continuously transporting the fabric. The auxiliary light source 206 is used to provide a stable and uniform illumination intensity below the detection area, so that the yarn arrangement, local density differences, and thickness changes inside the fabric form observable differences in the transmitted energy distribution in the transmitted light field. Transmitted light imaging refers to imaging and acquiring the light intensity distribution of the transmitted light field after it has been modulated by the fabric. This imaging result can characterize abnormal information related to the internal structure of the fabric, which complements the reflected light imaging, which focuses on surface texture information. This ensures that the fused image has both surface texture continuity information and internal structural difference information.

[0089] When decoupling the reflected and transmitted light information in the fused image using feature channels, the fused image is split into reflected light components and transmitted light components according to the imaging source. The two types of components are then mapped to independent feature channel sets. The reflected light information refers to the pixel response formed by the reflection of the fabric surface acquired by the imaging device above, which mainly reflects the texture continuity changes of the yarn surface texture and surface defects. The transmitted light information refers to the pixel response after the auxiliary light source below penetrates the fabric and is modulated by the internal structure of the fabric, which mainly reflects the attenuation or enhancement effect of the yarn density distribution, pore structure, and thickness changes on the transmitted energy. Feature channel decoupling refers to separating the information of different imaging mechanisms into channel representations that can be modeled by the network at the data representation layer, thereby constructing reflected light feature channels and transmitted light feature channels respectively. The reflected light feature channels are used to characterize the texture continuity of the fabric surface, and the transmitted light feature channels are used to characterize the yarn density distribution and thickness differences inside the fabric.

[0090] When the reflected light feature channel and the transmitted light feature channel are input in parallel into the dual-branch feature extraction structure of the YOLO-FD target detection model, the backbone of the YOLO-FD target detection model consists of a first feature extraction branch and a second feature extraction branch. The first feature extraction branch receives the reflected light feature channel and extracts texture continuity features related to the periodic texture background of the fabric. Texture continuity features are used to characterize the main texture direction of the yarn, the stability of the texture period, and the abnormal patterns corresponding to local perturbations of the surface texture. The second feature extraction branch receives the transmitted light feature channel and extracts the transmitted energy distribution variation features. Transmitted energy distribution variation features are used to characterize the transmitted light attenuation gradient, local enhanced patches of transmitted light, and discontinuous structure of transmitted light energy distribution caused by single yarn unraveling, local loose yarn, micropores, or uneven thickness. The dual-branch feature extraction structure refers to using parallel networks with independent or partially shared parameters to extract features from feature channels from different sources. This allows the two types of information to form their own feature representations that are most conducive to the expression of defects before entering the fusion stage, thereby avoiding the transmission information being masked by the surface high-frequency texture response or the reflected information being misled by the difference in transmitted energy.

[0091] When introducing a cross-modal attention mechanism in the neck feature fusion stage, the cross-modal attention mechanism is used to model the complementarity and consistency between the feature representations corresponding to the reflected light feature channel and the transmitted light feature channel. The neck feature fusion stage refers to the cross-scale feature fusion stage located between the dual-branch backbone feature extraction and the detection head, which is used to combine features from different scales and sources into fused features that can be used for detection head prediction. The cross-modal attention mechanism generates attention weights simultaneously in the spatial location dimension and the channel dimension, enabling the model to identify whether regions with thickness anomalousness in the transmitted light features also have texture continuity perturbations in the reflected light features, and whether regions with texture anomalousness in the reflected light features also have abrupt energy distribution changes in the transmitted light features, thereby providing interpretable correlation evidence for subsequent weighted fusion.

[0092] When modeling the correlation between reflected light feature channels and transmitted light feature channels in terms of spatial location and channel dimension, a cross-modal attention mechanism is used to generate spatial correlation weights and channel correlation weights, respectively. The spatial correlation weights characterize the consistency of the two types of features in defect representation at the same spatial location, while the channel correlation weights characterize the complementarity of the two types of features at the semantic channel level. Correlation modeling can be achieved by constructing query vectors, key vectors, and value vectors from feature mappings to establish an attention map, forming a bidirectional association from transmitted light features to reflected light features and from reflected light features to transmitted light features. This ensures that internal defects with significant thickness differences but weak surface texture changes can still receive higher weights through transmitted light features, while surface defects with significant surface texture perturbations but weak transmitted light changes can still receive higher weights through reflected light features, avoiding missed or false detections caused by single-modal dominance. To ensure that the correlation weights are numerically usable for fusion, the attention weights can be normalized. The attention normalization expression is:

[0093]

[0094] in, This represents the attention weight of the feature at index j at spatial location or channel index i. This represents the query vector generated at index i based on the reflected light or transmitted light characteristics. Let d represent the key vector generated at index j by another modality feature, d represent the dimension of the query vector and the key vector, which is used to scale the inner product magnitude to stabilize the gradient, and N represent the number of key vectors participating in the normalization. This formula exponentializes and normalizes the similarity between the query vector and the key vector, so that the weights reflect the degree of matching of cross-modal features in the spatial location dimension and the channel dimension, thereby forming a relevance description that can be used for weighted fusion.

[0095] When performing weighted fusion of the first and second feature extraction branches based on correlation, the spatial correlation weight and channel correlation weight output by the cross-modal attention mechanism are applied to the outputs of the first and second feature extraction branches, respectively. This allows the two types of features to be adaptively weighted according to their degree of correlation with defects during fusion. Weighted fusion refers to adding or concatenating the two types of features at the corresponding scale level and then recalibrating the weights. This makes the transmission energy distribution variation feature contribute more to the internal structural anomaly region and the texture continuity feature contribute more to the surface texture anomaly region. The dual-branch multi-scale feature refers to the fused feature set that contains both reflected light contribution and transmitted light contribution at multiple scale levels after weighted fusion. This can simultaneously retain the high-resolution detail information required for small defects and the high semantic context information required for blocky defects.

[0096] When mapping bi-branch multi-scale features to a unified multi-scale feature space, the bi-branch multi-scale features at each scale level are used to establish a unified scale system through upsampling, downsampling, and feature alignment operations. Within this unified scale system, a fused feature representation that can be shared by the detection head is formed. The multi-scale feature space refers to the semantic space composed of fused features at different resolution levels, which serves the detection needs of defects at different scales. When the detection head performs defect category discrimination and defect bounding box regression on abnormal regions within the multi-scale feature space, defect category discrimination outputs category information for different types of defects, and defect bounding box regression outputs the defect pixel coordinates and spatial range information of the defect bounding box in the fused image coordinate system. The defect pixel coordinates are used to determine the precise location of the defect in the fused image, and the spatial range information is used to characterize the defect coverage area and can be mapped to the coverage range in the fabric width direction and transmission direction. The defect pixel coordinates and spatial range information output by the detection head together constitute the preliminary localization result, which reflects the spatial positional relationship between abnormal fabric surface texture and abnormal fabric internal structure.

[0097] The channel-level decoupling of reflected and transmitted light information is accomplished based on the difference in imaging sources. This difference refers to the fact that the pixel responses in the defect image originate from the reflected light field formed by the upper imaging and the transmitted light field formed by the lower auxiliary light source penetrating the fabric. The reflected light information mainly reflects the continuity of the fabric surface texture and the texture disturbance of surface defects, while the transmitted light information mainly reflects the attenuation or enhancement effect of the yarn density distribution and thickness difference within the fabric on the transmitted energy. The channel-level decoupling refers to decomposing the defect image into two sets of channel data that can be modeled separately by the network at the data representation layer. This organizes the reflected light information as reflected light feature input and the transmitted light information as transmitted light feature input. The reflected light feature input is used to stably express surface texture attributes such as the main texture direction, texture period, and texture energy of the yarn, while the transmitted light feature input is used to stably express internal structural attributes such as pore structure, density abrupt changes, and thickness anomalies, thereby providing clear feature source boundaries for subsequent dual-branch encoding.

[0098] After the reflected light feature input and the transmitted light feature input are fed into the dual-branch coding structure of the DeepLabv3-FD semantic segmentation model in parallel, the first coding branch performs multi-level feature encoding specifically on the reflected light feature input to extract multi-scale detail features related to the high-frequency periodic texture of the fabric. The multi-scale detail features are used to express the disruption of the surface texture continuity caused by the defect area. The disruption of the surface texture continuity refers to the slight perturbation of the yarn texture in terms of directional continuity, periodic stability, or local texture energy. The second coding branch performs multi-level feature encoding specifically on the transmitted light feature input to extract the transmission energy distribution change features caused by single yarn unraveling, local loose yarn, or thickness abnormalities. The transmission energy distribution change features are used to express the internal structural abnormalities corresponding to the defect area. The internal structural abnormalities refer to the local abrupt change in transmitted light intensity caused by the decrease in yarn density, the increase in porosity, or the change in thickness. The dual-branch coding structure means that the two coding paths are independent or partially independent in terms of parameters, so that the reflected light features and the transmitted light features form their own feature representations that are more conducive to segmentation judgment during the coding stage, thereby avoiding the masking of internal structural abnormalities by high-frequency surface texture or the interference of transmission energy differences on surface texture judgment.

[0099] After introducing the first and second coding branches respectively, the multi-scale receptive field expansion structure enables the reflected light feature input and the transmitted light feature input to form corresponding feature representations at different spatial scales. The multi-scale receptive field expansion structure is used to obtain contextual semantic information of different ranges while maintaining a certain spatial resolution. The receptive field refers to the size of the original image coverage area corresponding to a certain position of the feature map. The multi-scale receptive field expansion structure of the first coding branch enables the edge perturbation of small defects to remain clear in a small receptive field, while maintaining the continuity of linear defects in a medium receptive field and enabling blocky defects to obtain complete contour semantics in a large receptive field. The multi-scale receptive field expansion structure of the second coding branch enables the transmission energy distribution variation feature to simultaneously express the structural context of local density abrupt changes and larger-scale thickness gradual changes, so that internal structural anomalies can be stably expressed at different scales, ensuring that the subsequent fusion stage can perform correlation modeling of the two types of features under scale matching conditions.

[0100] After completing the bi-branch multi-scale feature encoding, a cross-modal attention mechanism is introduced in the feature fusion stage to model the correlation between reflected light features and transmitted light features. This cross-modal attention mechanism characterizes the complementarity and consistency of the two types of features in both spatial location and channel dimensions. The spatial location correlation represents the degree of joint support between reflected light features and transmitted light features for defect boundaries and regions at the same spatial location. The channel correlation represents the complementary contribution of reflected light feature channels and transmitted light feature channels at the semantic level. Correlation modeling is achieved through an attention weight matrix, enabling the model to distinguish between surface defects with only surface texture anomalies and stable transmission structures, and internal defects with significant transmission structure anomalies and weak surface texture changes during fusion. It also adaptively selects the more reliable feature source for each type of defect at its corresponding spatial location. By exponentializing and normalizing the similarity between the query vector and the key vector, the weights reflect the matching degree of the two types of features in both spatial location and channel dimensions, thus forming a correlation description that can be used for weighted fusion.

[0101] When performing weighted fusion of bi-branch features based on correlation, the spatial weights and channel weights generated by the cross-modal attention mechanism are applied to the feature representations corresponding to reflected light features and transmitted light features, respectively. This makes the fusion result more biased towards the expression supported by both types of features near the defect boundary, and more biased towards the expression of only one type of feature in the region where only one type of feature is significant. Weighted fusion refers to performing element-wise weighted summation or weighted concatenation of the two types of features under the condition of same scale alignment, and then recalibrating to form a fused multi-scale feature set. This allows the fused multi-scale features to simultaneously contain information on the disruption of surface texture continuity and information on internal structural anomalies, and the two types of information maintain a distinguishable but synergistic relationship in terms of spatial location and channel semantics.

[0102] After the fused multi-scale features enter the decoding structure, spatial resolution is restored through stepwise upsampling and boundary enhancement processing is introduced. Stepwise upsampling is used to gradually restore low-resolution semantic features to a spatial resolution consistent with the defect image, while boundary enhancement processing is used to strengthen the feature responses corresponding to the edges of the defect region to maintain clear and continuous boundaries. When decoding under cross-modal attention constraints, the boundary responses of the defect region supported by both reflected and transmitted light features are strengthened. Boundary response refers to the edge feature response formed by gradient abrupt changes, texture direction disorder, or transmission energy abrupt changes at the junction of the defect region and the background region. "Supported by both" means that the edge position shows anomalous evidence in both reflected and transmitted light features. By giving higher weight to the boundary positions supported by both in the feature restoration and boundary enhancement process during decoding, the boundaries of small defects remain separable against the background of high-frequency textures, and the boundaries of internal defects are not smoothed or broken due to surface texture interference.

[0103] After decoding, the DeepLabv3-FD semantic segmentation model performs category determination processing on each pixel in the defect image at the pixel level. Pixel-level category determination processing refers to outputting the determination result of whether each pixel belongs to the defect category or the background category. The determination result of combining reflected light features and transmitted light features means that the pixel-level classifier or output layer simultaneously uses evidence from reflected light and transmitted light in the fused features, so that surface defects and internal defects can be consistently assigned pixels under the same determination framework. After distinguishing pixels belonging to defect areas from pixels belonging to non-defect areas, a segmentation mask corresponding to the defect area is generated. The segmentation mask is a binary or multi-class annotation map with the same spatial resolution as the defect image, which is used to accurately identify the spatial range of the defect area at the pixel level.

[0104] S160, based on segmentation mask, extracts the geometric and texture feature parameters of defects respectively, and integrates environmental state information for dynamic compensation, thereby completing the quality assessment and grading of the fabric.

[0105] When a segmentation mask is used as a spatial constraint, a pixel-level one-to-one correspondence is established between the segmentation mask and the defect image. This ensures that the set of pixels marked as defects in the segmentation mask is identified as defect regions in the defect image, and the set of pixels marked as non-defects in the segmentation mask is identified as non-defect regions in the defect image. The spatial constraint means that the segmentation mask limits subsequent operations to be carried out only within the marked region or within the marked region and its neighborhood, thereby avoiding the misinclusion of the periodic background texture of the fabric in the defect analysis. Through this spatial constraint, defect regions and non-defect regions are strictly distinguished in the same defect image coordinate system, providing a unified definition of the region boundary for the subsequent extraction of geometric feature parameters and texture feature parameters.

[0106] When performing feature extraction under spatial constraints, the defect region is used as the sole analysis object. Geometric feature parameters and texture feature parameters are obtained separately. Feature extraction refers to the process of quantitatively describing the defect region. Geometric feature parameters are a set of parameters used to characterize the morphology and boundary structure of the defect, while texture feature parameters are a set of parameters used to characterize the degree of damage caused by the defect region to the continuity of the fabric texture. Geometric feature parameters describe the shape of the defect, while texture feature parameters describe the damage caused by the defect to the texture. The two types of parameters together constitute a complementary representation of the defect, so that quality assessment and grading do not depend on a single scale or a single visual cue.

[0107] Contour analysis and region statistical processing revolve around the boundary pixel set and internal pixel set of the defect region, quantifying the defect geometry into comparable quantitative representations. Contour analysis involves extracting the outer boundary curve of the defect region from the segmentation mask and measuring its shape, while region statistics involve statistically analyzing the area, distribution, and shape of the internal pixel set of the defect region. Contour analysis yields the curvature change and boundary complexity of the defect boundary, while region statistics determine the defect coverage and morphological scale. This allows for the formation of a set of distinguishable quantitative indicators for defects such as holes, linear breaks, and blocky dirt at the geometric level. To form core geometric indicators suitable for grading, area, perimeter, and compactness can be combined to construct a geometric representation vector. Compactness reflects the degree of deviation of the defect boundary from a regular circle; the expression for compactness is:

[0108]

[0109] Where C represents compactness, ranging from 0 to 1, with values ​​closer to 1 indicating a shape closer to a circle; A represents the area of ​​the defect region, determined by the mapping relationship between the number of defect pixels in the segmentation mask and pixel resolution; and P represents the perimeter of the defect region, obtained by the length of the contour pixel chain or the sum of boundary line segments. This formula normalizes the relationship between area and perimeter, allowing the morphological complexity of defects at different scales to be compared at the same scale. Defects with broken or elongated lines usually correspond to a smaller compactness, while blocky defects with smooth boundaries usually correspond to a larger compactness, thus forming a quantitative characterization result of the defect geometry.

[0110] Texture analysis, under spatial constraints, examines the texture information corresponding to defect areas, systematically extracting grayscale statistical distribution features, texture direction consistency features, and local texture roughness features. Grayscale statistical distribution features characterize the concentration trend and dispersion of pixel intensity in defect areas; texture direction consistency features characterize the stability of the main texture direction relative to the background texture direction within defect areas; and local texture roughness features characterize the strength and spatial fluctuation of high-frequency components in defect areas. Comparative analysis simultaneously extracts similar texture features from adjacent non-defect areas and calculates the differences between defect and non-defect areas to form texture feature parameters characterizing the degree of texture continuity disruption. The degree of texture continuity disruption refers to the interruption, misalignment, energy attenuation, or directional disorder caused by defect areas to the original periodic and directional texture structure of the fabric. To simultaneously characterize grayscale differences, direction differences, and roughness differences, a texture disruption index can be constructed as one of the texture feature parameters. The expression for the texture disruption index is:

[0111]

[0112] Where T represents the texture destruction index, which ranges from 0 to 1, with a larger value indicating more significant texture destruction. This represents the normalized result of the difference in gray-scale statistical distribution, calculated from the difference in mean, variance, or histogram distance between defective and non-defective regions and mapped to a range of 0 to 1. The normalized result representing the difference in texture orientation consistency is calculated from the angle difference between the principal direction of the defective region and the principal direction of the non-defective region and mapped to the range of 0 to 1. This represents the normalized result of local texture roughness differences, calculated from the high-frequency energy difference or gradient energy difference between defective and non-defective regions and mapped to the range of 0 to 1. , , These represent the fusion weights for the three types of differences, with values ​​ranging from 0 to 1 and satisfying the following conditions: This formula weights and fuses grayscale differences, directional differences, and roughness differences, so that defects with extremely weak contrast but obvious directional disorder can still obtain a high texture destruction index, and stain-like defects with obvious grayscale differences but small directional changes can also obtain a high texture destruction index, thus forming texture feature parameters that can be compared across materials and lighting conditions.

[0113] The fusion processing of geometric feature parameters, texture feature parameters, and environmental state information incorporates environmental humidity data into the feature interpretation framework. Environmental state information includes at least environmental humidity data, which is synchronously collected by a humidity sensor and associated with the defect image timestamp, ensuring that each defect image corresponds to unique environmental humidity data. When establishing the correlation between environmental state information and imaging features, environmental humidity data is considered as an external disturbance source affecting fabric surface reflectivity, yarn fuzziness, and transmission energy attenuation. Based on this, dynamic compensation is performed on geometric and texture feature parameters. Dynamic compensation refers to correcting feature parameters according to environmental humidity data, ensuring comparable feature values ​​for the same type of defect under different humidity conditions. Dynamic compensation can be achieved by constructing a humidity compensation factor, the expression of which is:

[0114]

[0115] in, This represents the humidity compensation factor, with a value range from... Together with the humidity range, H is determined and remains a positive value, representing the ambient humidity data at the time of defect image acquisition. This represents the baseline humidity data, used to define the zero-bias reference for compensation. This indicates the upper limit of humidity within the statistical period or under preset conditions. This indicates the lower limit of humidity within the statistical period or under preset conditions. This represents the compensation sensitivity coefficient, which is used to adjust the effect of humidity changes on the feature compensation intensity and is determined by calibration samples. This formula normalizes the humidity deviation and linearly maps it to a compensation factor, allowing the contrast attenuation or texture energy drift caused by increased humidity to be corrected at the feature layer. The compensation factor is then applied to the feature parameters to obtain dynamically compensated feature parameters. For example, it can compensate for the texture destruction index and geometric scale parameters respectively. The expression is:

[0116]

[0117] in, This represents the texture destruction index after dynamic compensation. This represents the area parameter after dynamic compensation. The index represents the texture degradation index before compensation, and A represents the area parameter before compensation. This represents the humidity compensation factor; by uniformly scaling the feature parameters, this formula weakens the feature drift caused by changes in environmental humidity, thereby eliminating the impact of environmental changes on the defect feature judgment results.

[0118] The comprehensive defect characterization result is constructed from dynamically compensated geometric feature parameters and dynamically compensated texture feature parameters. This allows the comprehensive defect characterization result to express the morphological scale, boundary complexity, and texture damage intensity of the defect within the same vector space. The comprehensive defect characterization result serves as a unified input for quality assessment and grading. The comprehensive defect characterization result can be formed into a comprehensive feature vector using a normalized concatenation method, or a comprehensive score can be formed using a weighted aggregation method. This ensures that the influence of the same defect in both the geometric and texture dimensions is reflected simultaneously, and the comparability between different types of defects is guaranteed by the normalization and weighting system. The comprehensive score can be constructed using linear weighting, and the expression for the defect score is:

[0119]

[0120] Where S represents the defect score, which ranges from 0 to 1, with a larger value indicating a more severe impact of the defect on quality. This represents the normalized result of the area parameter after dynamic compensation. This represents the normalized result of the perimeter parameter after dynamic compensation. This represents the normalized result of firmness. This represents the normalized result of the texture corruption index after dynamic compensation. , , , These represent the area contribution weight, perimeter contribution weight, compactness contribution weight, and texture contribution weight, respectively, with values ​​ranging from 0 to 1 and satisfying the following conditions: This formula maps defect size, boundary complexity, and texture destruction intensity into comparable defect scores, enabling subsequent grading models to directly use defect scores for grade classification.

[0121] When the comprehensive defect characterization results are matched with preset quality assessment rules or grading judgment models, the comprehensive defect characterization results are used as the judgment basis and the fabric quality grade is output. The quality assessment rules refer to the set of rules that map the comprehensive feature vector or defect score to the quality grade, and the grading judgment model refers to the classification model obtained through training or calibration to map the comprehensive defect characterization results to the quality grade. The matching process aggregates all comprehensive defect characterization results within the same fabric inspection batch, so that the fabric quality assessment reflects both the severity of individual defects and the cumulative impact of the number and distribution of defects on the overall quality. Based on preset thresholds or models, the quality grading judgment result of the fabric is determined, thereby completing the assessment and grading judgment of fabric quality and providing an executable judgment basis for subsequent process adjustments, sorting and disposal, or quality traceability.

[0122] This embodiment also discloses a multi-sensor collaborative fabric defect detection device, referring to... Figure 3 The device includes an acquisition module 301, a processing module 302, and an output module 303. It is used to execute any of the multi-sensor collaborative fabric defect detection methods described above, wherein:

[0123] The acquisition module 301 is used to acquire local images and infrared images of the fabric through a dual-line array camera and an infrared camera, and to acquire environmental state information corresponding to the detection environment through a humidity sensor.

[0124] Processing module 302 is used to perform image registration and stitching processing on local images to form panoramic images, perform spatial registration and grayscale normalization processing on infrared images, and perform multimodal fusion processing on infrared images and panoramic images to generate fused images.

[0125] Processing module 302 is used to perform target detection processing on abnormal regions in the fused image in a multi-scale feature space, and output the pixel coordinates and range information corresponding to different types of defects, thereby forming a preliminary localization result of the defect region;

[0126] Processing module 302 is used to acquire defect images of the defect area using an area array camera based on the preliminary positioning results;

[0127] Processing module 302 is used to perform pixel-level semantic segmentation processing on the defect region in the defect image, thereby outputting a segmentation mask corresponding to the defect region;

[0128] The output module 303 is used to extract the geometric and texture feature parameters of the defects based on the segmentation mask, and to perform dynamic compensation by integrating environmental state information, so as to complete the quality assessment and grading of the fabric.

[0129] It should be noted that the above embodiments of the apparatus are only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0130] This embodiment also discloses an electronic device, as shown in the reference. Figure 4 The electronic device may include: at least one processor 401, at least one communication bus 402, user interface 403, network interface 404, and at least one memory 405.

[0131] The communication bus 402 is used to enable communication between these components.

[0132] The user interface 403 may include a display screen and a camera. Optionally, the user interface 403 may also include a standard wired interface and a wireless interface.

[0133] The network interface 404 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).

[0134] The processor 401 may include one or more processing cores. The processor 401 connects to various parts of the server using various interfaces and lines, and performs various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory 405, and by calling data stored in memory 405. Optionally, the processor 401 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 401 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem may also be implemented as a separate chip without being integrated into the processor 401.

[0135] The memory 405 may include random access memory (RAM) or read-only memory. Optionally, the memory may include a non-transitory computer-readable storage medium. The memory 405 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 405 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 405 may also be at least one storage device located remotely from the aforementioned processor 401. As a computer storage medium, the memory 405 may include an operating system, a network communication module, a user interface 403 module, and an application program for a multi-sensor collaborative fabric defect detection method.

[0136] exist Figure 4In the electronic device shown, the user interface 403 is mainly used to provide an input interface for the user and to obtain the user input data; while the processor 401 can be used to call an application program stored in the memory 405 for a multi-sensor collaborative fabric defect detection method. When executed by one or more processors 401, the electronic device performs one or more methods as described in the above embodiments.

[0137] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, as some steps can be performed in other orders or simultaneously according to the present invention. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0138] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0139] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some service interface; the indirect coupling or communication connection between apparatuses or units may be electrical or other forms.

[0140] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0141] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0142] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory 405 and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned memory 405 includes various media capable of storing program code, such as a USB flash drive, external hard drive, magnetic disk, or optical disk.

[0143] The present invention also discloses a non-transitory computer-readable storage medium storing instructions. When executed by one or more processors 401, these instructions cause an electronic device to perform one or more methods as described in the above embodiments.

[0144] The above are merely exemplary embodiments of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Those skilled in the art will readily conceive of other embodiments of this disclosure upon considering the specification and the disclosure of practical truths. This invention is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure. The specification and embodiments are to be considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.

Claims

1. A method for detecting fabric defects using a multi-sensor collaborative approach, characterized in that, The method includes: Local images and infrared images of the fabric are acquired using a dual-line array camera and an infrared camera, and environmental status information corresponding to the detection environment is obtained using a humidity sensor. Image registration and stitching are performed on the local image to form a panoramic image; spatial registration and grayscale normalization are performed on the infrared image; and multimodal fusion is performed on the infrared image and the panoramic image to generate a fused image. Target detection processing is performed on abnormal regions in the fused image within a multi-scale feature space, and pixel coordinates and range information corresponding to different types of defects are output, thereby forming a preliminary localization result of the defect region. Based on the preliminary positioning results, a defect image of the defect area is acquired using an area scan camera; Pixel-level semantic segmentation is performed on the defective regions in the defective image to output a segmentation mask corresponding to the defective regions; Based on the segmentation mask, the geometric and texture feature parameters of the defects are extracted respectively, and the environmental state information is fused for dynamic compensation to complete the quality assessment and grading of the fabric. Based on the segmentation mask, the geometric and texture feature parameters of the defects are extracted respectively, and the environmental state information is fused for dynamic compensation to complete the quality assessment and grading of the fabric. Specifically, this includes: Using the segmentation mask as a spatial constraint, the defective region and the non-defective region are distinguished in the defective image; Under the constraints of the spatial constraints, feature extraction processing is performed on the defect area, wherein the feature extraction processing includes obtaining geometric feature parameters for characterizing the morphological features of the defect and texture feature parameters for characterizing the degree of fabric texture damage. By performing contour analysis and region statistical processing on the defect area, the quantitative characterization results of the defect area on the geometric morphology of the defect are extracted. Under the constraints of the spatial constraints, texture analysis processing is performed on the image texture information corresponding to the defect area to extract the gray-level statistical distribution features, texture direction consistency features and local texture roughness features in the defect area. By comparing and analyzing the texture features with those of adjacent non-defect areas, texture feature parameters are formed to characterize the degree of damage of defects to the continuity of fabric texture. The geometric feature parameters, the texture feature parameters, and the environmental state information are fused together. The environmental state information includes at least environmental humidity data. By establishing the correlation between the environmental state information and the imaging features, dynamic compensation processing is performed on the geometric feature parameters and the texture feature parameters to eliminate the influence of environmental changes on the defect feature determination results. A comprehensive defect characterization result is constructed based on dynamically compensated geometric and texture feature parameters. The comprehensive defect characterization results are matched with preset quality assessment rules or grading judgment models to complete the assessment and grading judgment of fabric quality.

2. The method for detecting fabric defects using multi-sensor collaboration according to claim 1, characterized in that, The step of performing target detection processing on abnormal regions in the fused image within a multi-scale feature space, and outputting defect pixel coordinates and range information corresponding to different types of defects, thereby forming a preliminary localization result of the defect region, specifically includes: The fused image is input into a pre-trained YOLO-FD object detection model. The YOLO-FD object detection model's backbone feature extraction network performs multi-level convolutional feature extraction on the fused image to obtain multi-scale first feature representations corresponding to different receptive field scales at different feature levels. Each first feature representation is used to characterize the spatial distribution differences of defects, linear defects, and block defects of different scales in the fabric. In the feature extraction process of the backbone feature extraction network, a feature enhancement process combining channel attention and spatial attention is introduced into at least one intermediate feature layer to weight and enhance the feature channel response intensity and spatial location response intensity related to defects. The first feature representation is upsampled, downsampled, and concatenated by a neck feature fusion network to achieve cross-scale fusion of defect features at different scales within a unified semantic space. The detection head performs defect category prediction and defect bounding box regression on the fused defect features in the unified semantic space, thereby outputting defect category information corresponding to each defect, defect pixel coordinates of the defect bounding box in the fused image, and spatial distribution range of defects in the fabric width direction and transmission direction. Based on the defect pixel coordinates and the spatial distribution range, a preliminary localization result is generated to characterize the spatial positional relationship of the defect at a panoramic scale.

3. The method for detecting fabric defects using multi-sensor collaboration according to claim 1, characterized in that, The step of performing pixel-level semantic segmentation processing on the defective regions in the defective image to output a segmentation mask corresponding to the defective regions specifically includes: The defect image is input into the pre-trained DeepLabv3-FD semantic segmentation model. The defect image is processed by multi-level feature encoding through the encoding structure of the DeepLabv3-FD semantic segmentation model to form multi-scale second feature representations at different downsampling scales. Each second feature representation is used to characterize the local texture anomaly features of the defect area and the contextual semantic relationship of the defect area in the fabric texture background. During the feature encoding process of the defective image, by setting a multi-branch dilated convolution structure in the encoding structure, the second feature representation can obtain contextual information within different receptive fields while maintaining spatial resolution. The second feature representation is subjected to feature fusion processing, so that the defect features at different scales are combined in a unified semantic space, and a spatial attention mechanism is introduced in the feature fusion process to assign enhanced weights to the spatial location of the defect region. The fused feature representation is subjected to progressive upsampling through a decoding structure, and boundary enhancement processing is introduced into the decoding structure to strengthen the feature response corresponding to the edge of the defect region. The DeepLabv3-FD semantic segmentation model performs category determination processing on each pixel in the defect image at the pixel level, distinguishing pixels belonging to the defect area from pixels belonging to the background area, and generating a segmentation mask that corresponds one-to-one with the defect area based on the pixel-level determination results.

4. The method for detecting fabric defects using multi-sensor collaboration according to claim 1, characterized in that, The step of performing target detection processing on abnormal regions in the fused image within a multi-scale feature space, and outputting pixel coordinates and range information corresponding to different types of defects to form a preliminary localization result of the defect region, specifically also includes: Based on the periodic and directional characteristics of fabric texture, a texture consistency benchmark representation is constructed for the fused image. The texture consistency benchmark representation is used to characterize the dominant frequency texture structure formed by the yarn arrangement in the normal fabric area. Under the constraints of the texture consistency benchmark representation, texture consistency violation analysis processing is performed on the fused image. By jointly characterizing the texture direction continuity, texture energy distribution stability and texture repetition period offset of the local region, a texture violation response map is generated to reflect the degree of disturbance of the local weaving pattern. The texture destruction response map is spatially aligned with the fused image to obtain additional features; The additional features are introduced into the target detection process, enabling the target detection to simultaneously express the abnormal region based on gray-scale change features, geometric morphology features, and texture consistency destruction features in a multi-scale feature space. The detection head performs defect category discrimination and defect bounding box regression processing on the abnormal area, and outputs the defect pixel coordinates and spatial range information corresponding to the defect, thereby forming the preliminary positioning result.

5. The method for detecting fabric defects using multi-sensor collaboration according to claim 1, characterized in that, The step of performing target detection processing on abnormal regions in the fused image within a multi-scale feature space, and outputting defect pixel coordinates and range information corresponding to different types of defects to form a preliminary localization result of the defect region, specifically also includes: Complete the acquisition of transmitted light imaging based on the light-transmitting conveyor belt and the auxiliary light source below, and form the fused image; The reflected light information and transmitted light information in the fused image are decoupled by feature channels, and a reflected light feature channel for characterizing the continuity of fabric surface texture and a transmitted light feature channel for characterizing the yarn density distribution and thickness difference inside the fabric are constructed respectively. The reflected light feature channel and the transmitted light feature channel are input in parallel into the dual-branch feature extraction structure of the YOLO-FD target detection model, so that the first feature extraction branch extracts the texture continuity features of the fabric periodic texture background for the reflected light feature channel, and the second feature extraction branch extracts the transmission energy distribution variation features for the transmitted light feature channel. A cross-modal attention mechanism is introduced in the neck feature fusion stage to model the correlation between the reflected light feature channel and the transmitted light feature channel in terms of spatial position dimension and channel dimension; Based on the correlation, the first feature extraction branch and the second feature extraction branch are weighted and fused to obtain dual-branch multi-scale features; The dual-branch multi-scale features are mapped to a unified multi-scale feature space, and the detection head performs defect category discrimination and defect bounding box regression processing on the abnormal region in the multi-scale feature space, outputting defect pixel coordinates and spatial range information corresponding to different types of defects, thereby forming the preliminary localization result.

6. The method for detecting fabric defects using multi-sensor collaboration according to claim 1, characterized in that, The step of performing pixel-level semantic segmentation processing on the defective regions in the defective image to output a segmentation mask corresponding to the defective regions further includes: Based on the difference in imaging sources, the reflected light information and transmitted light information in the defect image are decoupled at the channel level to form reflected light feature inputs for characterizing the continuity of fabric surface texture and transmitted light feature inputs for characterizing the yarn density distribution and thickness difference inside the fabric. The reflected light feature input and the transmitted light feature input are input in parallel into the dual-branch encoding structure of the DeepLabv3-FD semantic segmentation model. The first encoding branch extracts multi-scale detail features related to the high-frequency periodic texture of the fabric based on the reflected light feature input, which is used to characterize the damage of the defect area to the continuity of the surface texture. The second encoding branch extracts the transmitted energy distribution change features caused by single yarn unraveling, local loose yarn, or thickness abnormality based on the transmitted light feature input, which is used to characterize the internal structural abnormalities corresponding to the defect area. In the dual-branch coding process, a multi-scale receptive field extension structure is introduced in the first coding branch and the second coding branch respectively, so that the reflected light feature input and the transmitted light feature input form corresponding feature representations at different spatial scales; After completing the bi-branch multi-scale feature encoding, a cross-modal attention mechanism is introduced in the feature fusion stage to model the relationship between the feature representation corresponding to the reflected light feature input and the feature representation corresponding to the transmitted light feature input in the spatial position dimension and channel dimension, and to perform weighted fusion of the bi-branch features based on the relationship. Under the cross-modal attention constraint, the fused multi-scale features are input into the decoding structure. Through step-by-step upsampling and boundary enhancement processing, the boundary response of the defect region supported by both reflected light features and transmitted light features is strengthened during the decoding process. After decoding is completed, the DeepLabv3-FD semantic segmentation model performs category determination processing on each pixel in the defect image at the pixel level. By combining the determination results of the reflected light features and the transmitted light features, pixels belonging to the defect area are distinguished from pixels belonging to the non-defect area, thereby generating a segmentation mask corresponding to the defect area.

7. A multi-sensor collaborative fabric defect detection device, characterized in that, The device is used to perform a multi-sensor collaborative fabric defect detection method as described in any one of claims 1-6, the device comprising an acquisition module, a processing module, and an output module, wherein: The acquisition module is used to acquire local images and infrared images of the fabric through a dual-line array camera and an infrared camera, and to acquire environmental state information corresponding to the detection environment through a humidity sensor. The processing module is used to perform image registration and stitching processing on the local image to form a panoramic image, perform spatial registration and grayscale normalization processing on the infrared image, and perform multimodal fusion processing on the infrared image and the panoramic image to generate a fused image. The processing module is used to perform target detection processing on abnormal regions in the fused image within a multi-scale feature space, and output pixel coordinates and range information corresponding to different types of defects, thereby forming a preliminary localization result of the defect region. The processing module is used to acquire a defect image of the defect area using an area scan camera based on the preliminary positioning result. The processing module is used to perform pixel-level semantic segmentation processing on the defective region in the defective image, thereby outputting a segmentation mask corresponding to the defective region. The output module is used to extract the geometric and texture feature parameters of the defects based on the segmentation mask, and to perform dynamic compensation by fusing the environmental state information, thereby completing the quality assessment and grading of the fabric.

8. An electronic device, characterized in that, The device includes a processor, a communication bus, a user interface, a network interface, and a memory. The memory is used to store instructions. The user interface and the network interface are both used to communicate with other devices. The communication bus is used to enable communication between the components within the electronic device. The processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any one of claims 1-6.

9. A non-transitory computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed, perform the method as described in any one of claims 1-6.