Multi-scene equipment defect detection method, control device and equipment
Through multimodal detection methods, dynamic adjustment of feature channel weights, separation of noise and defect features, and cross-scene feature alignment loss, the problems of feature drift and noise interference in equipment detection in complex environments are solved, and high-precision and low-power equipment defect detection is achieved.
Patent Information
- Application Number
- CN202510551413.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-09-26
AI Technical Summary
Existing equipment defect detection methods face problems such as feature drift caused by environmental interference, high false detection rate, and poor model generalization ability when applied across scenarios. In addition, single-modality detection is susceptible to noise interference, lacks a multi-source data collaborative verification mechanism, and has low detection reliability.
A multimodal detection method is adopted, the feature channel weights are adjusted through the dynamic feature calibration module, the anti-interference decoupling head is used to separate noise and defect features, and the cross-scene feature alignment loss is combined to achieve multi-source data collaborative verification, thereby enhancing the adaptability and accuracy of the model in complex environments.
It improves the accuracy and reliability of equipment defect detection, reduces the false alarm rate, adapts to the switching needs of complex scenarios, and meets the real-time and low power consumption requirements of industrial sites.
Smart Images

Figure CN120707914A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of defect detection, and in particular to a multi-scenario equipment defect detection method, control device and equipment. Background Art
[0002] In the field of intelligent operation and maintenance of industrial equipment, multi-scenario equipment defect detection technology is crucial to ensuring the safe operation of equipment. However, existing detection methods face severe challenges when applied across scenarios: environmental interference such as high dust in mining areas, low light in pipeline corridors, and complex backgrounds in scenic areas cause the visual features of equipment defects to drift significantly. Traditional detection models are difficult to adapt to dynamically changing feature distributions, resulting in a surge in false detection rates and poor model generalization capabilities. Existing technologies often use manual adjustment of model parameters or the addition of scene-specific detection modules to address these issues, which is not only inefficient but also difficult to achieve real-time adaptation. In addition, single-modality detection methods are easily affected by environmental noise and lack a multi-source data collaborative verification mechanism, resulting in a sharp drop in detection reliability under harsh working conditions.
[0003] Therefore, in response to the above pain points, this application proposes an environment-adaptive multimodal defect detection method control device and equipment, aiming to break through the technical bottleneck of cross-scene feature drift. Summary of the Invention
[0004] The present application provides a multi-scenario equipment defect detection method, control device and equipment. The method uses a dynamic feature calibration module to adjust the feature channel weights according to environmental parameters, so that the model focuses on the key features of the scene; the anti-interference decoupling head separates noise and defect features to suppress environmental interference; the cross-scenario feature alignment loss enforces the consistency of similar defect features. The three work together to solve the feature drift problem and improve the detection accuracy in complex scenarios.
[0005] In a first aspect, a multi-scenario device defect detection method is provided, the method comprising:
[0006] S1: Acquire RGB images, infrared thermal images, and environmental parameters of the device to be inspected through a multi-source sensor, wherein the environmental parameters include one or more of light intensity, dust concentration, and humidity;
[0007] S2: The RGB image is input into an improved YOLOv8 model, wherein the improvements include: adding a dynamic feature calibration module at the Backbone end, which dynamically adjusts feature weights based on environmental parameters; deploying an anti-interference decoupling head in front of the detection head, and using a dual-branch structure to extract environmental invariant features and defect discriminant features respectively;
[0008] S3: Calculate cross-scene feature alignment loss to constrain the cosine similarity of feature vectors of the same type of defects under different environment parameters;
[0009] S4: Fuse the abnormal area of the infrared thermal image with the detection results of the improved YOLOv8 model to output defect location and classification information.
[0010] It should be understood that by integrating multi-source data such as RGB images, infrared thermal images, and environmental parameters, a dynamic feature calibration module (DFCM) and an anti-interference decoupling head are embedded in the improved YOLOv8 model, and combined with cross-scene feature alignment loss to form a multimodal collaborative detection framework. The dynamic feature calibration module dynamically adjusts feature weights based on environmental parameters, allowing the model to focus on effective features in the current scene; the anti-interference decoupling head (ADH) reduces the influence of interference factors such as dust reflection by separating environmental noise from defect features; and the cross-scene feature alignment loss constrains the feature consistency of similar defects in different environments and suppresses feature distribution drift. Ultimately, through the infrared thermal image verification mechanism, the defect detection accuracy is improved and the false alarm rate is reduced in complex environments, especially adapting to the needs of complex scene switching, solving the problem of poor environmental adaptability of traditional single-modal detection models.
[0011] In conjunction with the first aspect, in certain implementations of the first aspect, implementation of the dynamic feature calibration module includes:
[0012] Inputting the environmental parameters into a two-layer fully connected network to generate a scene encoding vector;
[0013] Mapping the scene encoding vector to a weight vector of each channel of the feature map through a learnable matrix;
[0014] A soft threshold function is used to perform a sparse processing on the weight vector.
[0015] It should be understood that a learnable matrix is a linear transformation matrix that automatically optimizes parameters using training data. In the technical solution of this application, the learnable matrix maps the scene encoding vector to the calibration weights of each channel of the feature map. Its function is to establish a dynamic association between environmental parameters and the importance of feature channels. Compared with manually designed weight rules, the learnable matrix captures complex nonlinear relationships between environment and features through end-to-end training, making feature calibration more accurate.
[0016] It should be understood that the dynamic feature calibration module uses a fully connected network to map environmental parameters to feature channel weights. It then sparsifies these weights using a soft threshold function, retaining key feature channels sensitive to the current scene (such as edge texture channels in high-dust environments) while suppressing interference from irrelevant channels. This design dynamically associates feature selection with scene parameters. Compared to fixed-weight feature extraction methods, it significantly improves the signal-to-noise ratio of feature maps in scenes with sudden changes in illumination. It also reduces computational redundancy through weight sparsification, balancing detection accuracy and efficiency.
[0017] In combination with the first aspect, in some implementations of the first aspect, the anti-interference decoupling head includes:
[0018] An environmental noise branch, which uses depthwise separable convolution to extract environmental related features;
[0019] A defect detection branch, wherein the defect detection branch uses deformable convolution to extract device defect features;
[0020] The environmental noise branch is connected in parallel with the defect detection branch.
[0021] In combination with the first aspect, in certain implementations of the first aspect, the outputs of the environmental noise branch and the defect detection branch are subjected to a subtraction operation to eliminate environmental interference.
[0022] It should be understood that the anti-interference decoupling head uses a dual-branch structure to extract environmental noise characteristics and defect characteristics respectively, using deformable convolution to enhance adaptability to device deformation and eliminating environmental interference through feature subtraction. The noise branch uses lightweight depthwise separable convolution to effectively identify interference patterns such as dust adhesion and metal reflections while ensuring real-time performance. The deformable convolution of the detection branch can adaptively adapt to changes in defect morphology, improving detection accuracy and reducing detection errors.
[0023] In combination with the first aspect, in some implementations of the first aspect, the cross-scene feature alignment loss is calculated as follows:
[0024] The feature vector group {v1, v2, v3......v n The calculation formula is:
[0025]
[0026] in, is a permutation and combination, and n represents the number of characteristic vectors of the same type of defects under different environmental parameters.
[0027] It should be understood that the cross-scenario feature alignment loss addresses the issue of feature distribution drift at the model training level by forcing feature vectors of the same defect type to maintain high cosine similarity across different environments. Compared to traditional contrastive learning loss, this loss function is designed specifically for aligning defect features across multiple scenarios. In cross-dataset testing, migration detection accuracy from mining sites to pipeline corridors was significantly improved.
[0028] In combination with the first aspect, in certain implementations of the first aspect, step S4 includes: when the visual detection confidence level is in a first interval, starting infrared thermal imaging verification;
[0029] If the temperature difference between the defect area and the surrounding temperature is greater than the first threshold, the confidence level is increased.
[0030] It should be understood that when the confidence of visual inspection results is in doubt, the multimodal verification mechanism introduces infrared thermal imaging temperature difference analysis for cross-validation, and combines temperature differences to identify real defects, effectively solving the problem of traditional pure visual methods misjudging dust accumulation and other problems as equipment defects.
[0031] In conjunction with the first aspect, in some implementations of the first aspect, the improved YOLOv8 model adopts a staged training strategy:
[0032] The first stage is pre-trained on a synthetic dataset, using GAN to generate defect images under different environmental parameters;
[0033] The second stage is fine-tuning on real datasets, adopting a curriculum learning strategy to gradually increase training samples from low to high environment complexity.
[0034] It should be understood that the phased training strategy generates multi-environment synthetic data through GAN, expands the scenario diversity of training samples, and solves the problem of high cost of collecting real defect data; the curriculum learning strategy conducts progressive training according to the complexity of the environment, so that the model gradually adapts to the changes in feature distribution from simple to complex.
[0035] In conjunction with the first aspect, in certain implementations of the first aspect, the implementation optimization of the method on the embedded device side includes:
[0036] Replacing the fully connected network of the dynamic feature calibration module with a grouped sparse structure;
[0037] The dual-branch convolutional layer of the anti-interference decoupling head is quantized to 8 bits, and the 16-bit high-precision calculation of the detection branch is retained.
[0038] It should be understood that the grouped sparse structure and quantization compression reduce the computational load while maintaining detection accuracy. The grouped sparse network reduces the number of parameters in the DFCM module, and 8-bit quantization improves the inference speed of the ADH branch, while the detection branch retains 16-bit computation to ensure accurate extraction of key features. The optimized model significantly reduces power consumption on embedded platforms, meeting the low-power and real-time requirements of industrial detection systems.
[0039] In a second aspect, a control device is provided, which includes a processor and a memory, wherein the processor is coupled to the memory, the memory is used to store computer programs or instructions, and the processor is used to execute the computer programs or instructions in the memory, so that any implementation method described in any one of the first aspects is executed.
[0040] In a third aspect, a device is provided, comprising the control device as described in the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 This is a flowchart for implementing a multi-scenario device defect detection method provided in an embodiment of the present application. DETAILED DESCRIPTION
[0042] The terms used in the following embodiments are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification of this application and the appended claims, the singular expressions "a", "an", "said", "above", "the" and "this" are intended to also include expressions such as "one or more", unless there is a clear contrary indication in the context. It should also be understood that in the following embodiments of the present application, "at least one", "one or more" refer to one, two or more. The term "and / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist; for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship.
[0043] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0044] In industrial intelligent operation and maintenance, multi-scenario equipment defect detection technology faces the core challenge of insufficient adaptability to dynamic environments. Cross-scenario interference such as high dust in mining areas, low illumination in pipeline corridors, and complex backgrounds in scenic areas cause nonlinear drift in the visual features of equipment. Traditional detection models lack a dynamic feature adaptation mechanism, resulting in a cliff-like drop in generalization performance and a surge in false detection rates when feature distribution shifts. Existing technologies rely on manual parameter tuning or the addition of scene-specific modules, making it difficult to achieve real-time adaptive engineering deployment. At the same time, single-visual modality detection systems are susceptible to environmental noise pollution under harsh working conditions. Due to the lack of a collaborative verification mechanism for multi-source sensor data, detection reliability exhibits scene-sensitive fluctuations. Breaking through the technical bottlenecks of robust modeling of dynamic features and cross-modal information fusion has become a key path to achieving accurate identification of equipment defects in complex industrial scenarios.
[0045] The embodiments of the present application provide a multi-scenario equipment defect detection method, control device and equipment, which can effectively overcome the above-mentioned problems.
[0046] The technical solutions of the embodiments of the present application will be described below with reference to the accompanying drawings.
[0047] Figure 1 This is a flowchart of a multi-scenario device defect detection method provided by an embodiment of the present application. In some examples, the method includes:
[0048] S1: Acquire RGB images, infrared thermal images, and environmental parameters of the device to be inspected through a multi-source sensor, wherein the environmental parameters include one or more of light intensity, dust concentration, and humidity;
[0049] S2: The RGB image is input into an improved YOLOv8 model, wherein the improvements include: adding a dynamic feature calibration module at the Backbone end, which dynamically adjusts feature weights based on environmental parameters; deploying an anti-interference decoupling head in front of the detection head, and using a dual-branch structure to extract environmental invariant features and defect discriminant features respectively;
[0050] S3: Calculate cross-scene feature alignment loss to constrain the cosine similarity of feature vectors of the same type of defects under different environment parameters;
[0051] S4: Fuse the abnormal area of the infrared thermal image with the detection results of the improved YOLOv8 model to output defect location and classification information.
[0052] In some examples, implementation of the dynamic feature calibration module includes:
[0053] Inputting the environmental parameters into a two-layer fully connected network to generate a scene encoding vector;
[0054] Mapping the scene encoding vector to a weight vector of each channel of the feature map through a learnable matrix;
[0055] A soft threshold function is used to perform a sparse processing on the weight vector.
[0056] In one possible implementation, the dynamic feature calibration module consists of an environmental parameter encoder and a weight generator. Environmental parameters (lighting, dust, humidity) are encoded into a 128-dimensional vector via a two-layer fully connected network. This vector is then converted into a channel weight vector using a learnable matrix (size 128 × C, where C is the number of feature map channels). Low-weight channels are then filtered using a soft threshold function. Finally, the weight vector is multiplied channel-by-channel with the original feature map to output the calibrated feature map.
[0057] In one possible implementation, a soft threshold function is used to perform sparse processing on the weight vector. The calculation formula is:
[0058]
[0059] Among them, k is 0.5 and τ=0.3 is an empirical value.
[0060] In some examples, the anti-interference decoupling head includes:
[0061] An environmental noise branch, which uses depthwise separable convolution to extract environmental related features;
[0062] A defect detection branch, wherein the defect detection branch uses deformable convolution to extract device defect features;
[0063] The environmental noise branch is connected in parallel with the defect detection branch.
[0064] In one possible implementation, the anti-interference decoupling head adopts a dual-branch parallel structure: the environmental noise branch uses 3×3 depthwise separable convolution to extract light / dust interference features, and the detection branch uses 5×5 deformable convolution to extract deformation-invariant defect features.
[0065] In some examples, the outputs of the environmental noise branch and the defect detection branch are subjected to a subtraction operation to eliminate environmental interference.
[0066] In one possible implementation, the following formula is used to eliminate environmental interference:
[0067] F clean =F detect -α·F noise ,
[0068] Among them, F dectet is the defect feature map output by the detection branch, F noise is the environmental interference feature map extracted by the noise branch, and α is the dynamic adjustment coefficient.
[0069] In some examples, the cross-scene feature alignment loss is calculated as follows:
[0070] The feature vector group {v1, v2, v3......v n The calculation formula is:
[0071]
[0072] in, is a permutation and combination, and n represents the number of characteristic vectors of the same type of defects under different environmental parameters.
[0073] In some examples, step S4 includes:
[0074] When the visual inspection confidence level is in the first interval, infrared thermal imaging verification is started;
[0075] If the temperature difference between the defect area and the surrounding temperature is greater than the first threshold, the confidence level is increased.
[0076] In a possible implementation, the first interval is [0.4, 0.6], and the first threshold is the temperature difference between the defect area and the surrounding area, which is empirically determined to be 3°C.
[0077] In some examples, the improved YOLOv8 model adopts a phased training strategy:
[0078] The first stage is pre-trained on a synthetic dataset, using GAN to generate defect images under different environmental parameters;
[0079] The second stage is fine-tuning on real datasets, adopting a curriculum learning strategy to gradually increase training samples from low to high environment complexity.
[0080] In some examples, the implementation optimization of the method on the embedded device side includes:
[0081] Replacing the fully connected network of the dynamic feature calibration module with a grouped sparse structure;
[0082] The dual-branch convolutional layer of the anti-interference decoupling head is quantized to 8 bits, and the 16-bit high-precision calculation of the detection branch is retained.
[0083] In one possible implementation, the fully connected DFCM network is modified to a grouped sparse structure, with only 40% of neurons connected within each group, reducing computational complexity. The noise branch convolutional layers of the ADH are quantized to 8-bit integers, while the detection branch retains 16-bit floating-point computation. This optimization significantly reduces the model's peak memory usage on embedded platforms, improves inference speed, and reduces power consumption, meeting the 24 / 7 real-time detection requirements of industrial sites.
[0084] An embodiment of the present application provides a control device, which includes a processor and a memory, wherein the processor is coupled to the memory, the memory is used to store computer programs or instructions, and the processor is used to execute the computer program or instructions in the memory, so that the method described in any of the above examples is executed.
[0085] An embodiment of the present application also provides a device, which includes the control device described in the above example.
[0086] The above are only preferred embodiments of the present invention. The protection scope of the present invention is not limited to the above embodiments. Any equivalent modifications or changes made by ordinary technicians in this field based on the contents disclosed in the present invention should be included in the protection scope recorded in the claims.
Claims
1. A multi-scenario device defect detection method, characterized in that: The method comprises: S1: Acquire RGB images, infrared thermal images, and environmental parameters of the device to be inspected through a multi-source sensor, wherein the environmental parameters include one or more of light intensity, dust concentration, and humidity; S2: The RGB image is input into an improved YOLOv8 model, wherein the improvements include: adding a dynamic feature calibration module at the Backbone end, which dynamically adjusts feature weights based on environmental parameters; deploying an anti-interference decoupling head in front of the detection head, and using a dual-branch structure to extract environmental invariant features and defect discriminant features respectively; S3: Calculate cross-scene feature alignment loss to constrain the cosine similarity of feature vectors of the same type of defects under different environment parameters; S4: Fuse the abnormal area of the infrared thermal image with the detection results of the improved YOLOv8 model to output defect location and classification information.
2. The method according to claim 1, characterized in that The implementation of the dynamic feature calibration module includes: Inputting the environmental parameters into a two-layer fully connected network to generate a scene encoding vector; Mapping the scene encoding vector to a weight vector of each channel of the feature map through a learnable matrix; A soft threshold function is used to perform a sparse processing on the weight vector.
3. The method according to claim 1, characterized in that The anti-interference decoupling head includes: An environmental noise branch, which uses depthwise separable convolution to extract environmental related features; A defect detection branch, wherein the defect detection branch uses deformable convolution to extract device defect features; The environmental noise branch is connected in parallel with the defect detection branch.
4. The method according to claim 3, characterized in that The outputs of the environmental noise branch and the defect detection branch are subjected to a subtraction operation to eliminate environmental interference.
5. The method according to claim 1, wherein The calculation method of the cross-scene feature alignment loss is: The feature vector group {v1, v2, v3......v n The calculation formula is: in, is a permutation and combination, and n represents the number of characteristic vectors of the same type of defects under different environmental parameters.
6. The method according to claim 1, characterized in that The step S4 comprises: When the visual inspection confidence level is in the first interval, infrared thermal imaging verification is started; If the temperature difference between the defect area and the surrounding temperature is greater than the first threshold, the confidence level is increased.
7. The method according to claim 1, characterized in that The improved YOLOv8 model adopts a phased training strategy: The first stage is pre-trained on a synthetic dataset, using GAN to generate defect images under different environmental parameters; The second stage is fine-tuning on real datasets, adopting a curriculum learning strategy to gradually increase training samples from low to high environment complexity.
8. The method according to claim 1, characterized in that The implementation optimization of the method on the embedded device side includes: Replacing the fully connected network of the dynamic feature calibration module with a grouped sparse structure; The dual-branch convolutional layer of the anti-interference decoupling head is quantized to 8 bits, and the 16-bit high-precision calculation of the detection branch is retained.
9. A control device, characterized in that: The method comprises a processor and a memory, wherein the processor is coupled to the memory, the memory is used to store computer programs or instructions, and the processor is used to execute the computer programs or instructions in the memory, so that the method according to any one of claims 1 to 8 is performed.
10. A device, characterized in that The apparatus comprises a control device as claimed in claim 9.
Citation Information
Cited By
Transformer abnormal state intelligent identification system
CN121612888A
An intelligent transformer abnormal state identification system
CN121612888B