Cross-scene belt defect identification method based on adaptive typical sample learning technology

By adopting adaptive typical sample learning method in belt defect recognition technology, combining prototype adaptive modules and learnable adaptive modules, the computing resource efficiency and sample learning problems in cross-scene belt defect recognition are solved, and efficient and accurate belt defect recognition is achieved.

CN119942061AActive Publication Date: 2025-05-06UNIV OF SCI & TECH BEIJING
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202411841865.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-05-06
Estimated Expiration
2044-12-13

AI Technical Summary

Technical Problem

Existing belt defect identification technology has shortcomings in adapting to new scenarios, computing resource efficiency, and learning with few samples, making it difficult to effectively identify belt defects across scenarios.

Method used

A cross-scene belt defect recognition method based on adaptive typical sample learning technology is adopted. By building a supporting image dataset and query image dataset, combining the prototype adaptive module and the learnable adaptive module, feature enhancement and model optimization are achieved, computing resource consumption is reduced, and model universality and migration capabilities are improved.

Benefits of technology

It significantly improves the versatility and migration capabilities of the belt defect recognition model, reduces the amount of fine-tuning calculations, enhances the belt defect feature capture capability in new scenarios, and achieves efficient and accurate cross-scene belt defect recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942061A_ABST
    Figure CN119942061A_ABST
Patent Text Reader

Abstract

The invention provides a cross-scene belt defect identification method and device based on an adaptive typical sample learning technology, and relates to the technical field of defect detection. The method comprises the following steps: constructing a support image data set and a query image data set; inputting the support image data set and the query image data set into a defect identification basic model to obtain a first support image feature and a first query image feature; inputting the first support image feature and the first query image feature into a prototype adaptive module for feature enhancement; constructing a loss function according to the support image data set, the query image data set and the defect detection image data set; optimizing the prototype adaptive model to obtain an optimized prototype adaptive model; and according to the to-be-identified belt image data set, carrying out belt defect identification based on the defect identification basic model and the optimized prototype adaptive model. The cross-scene belt defect identification method is based on typical sample learning, and only needs a small amount of annotations, and is efficient and accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of defect detection technology, and in particular to a cross-scenario belt defect recognition method and device based on adaptive typical sample learning technology. Background Art

[0002] In industrial production, the belt conveyor system is an important part of material transportation, and its operating status directly affects the efficiency and safety of the entire production line. After long-term, high-intensity operation, the belt may have various types of defects, such as wear, tear, foreign body scratches, etc. If these problems are not discovered and handled in time, they will not only affect production efficiency, but may also cause serious safety accidents or equipment damage. Therefore, belt defect detection is of great significance in industrial production.

[0003] Traditional detection methods mainly rely on manual visual inspection, which is not only time-consuming and laborious, but also prone to missed detection or false detection due to human negligence. In recent years, with the development of computer vision and deep learning technology, image-based automatic defect detection methods have gradually become a research hotspot. These methods identify defects on belts by training deep neural networks, which have high accuracy and efficiency, but usually require a large amount of labeled data to train the model. In actual industrial environments, it is extremely costly to obtain a large amount of high-quality and diverse labeled data. The belt materials and operating conditions in different industrial scenarios vary greatly, and the belt image background and defect types vary greatly in different scenarios. The complex types of defects and uneven distribution make it difficult to effectively identify them in traditional target detection algorithms.

[0004] While the various existing belt defect recognition basic models often use different deep learning algorithms (such as convolutional neural networks or Transformers), there is currently no universal migration method to apply these basic models to different scenarios. The structures of these basic models are usually designed for specific scenarios. Once the environment or data distribution changes, the entire model structure needs to be retrained or modified, which increases the time and computing cost of model migration and reduces the flexibility and scalability of the system.

[0005] Existing technologies rely on full fine-tuning to adjust the belt defect recognition model; this method has the problem of excessive consumption of computing resources, making it difficult to efficiently apply the model in industrial scenarios. Since full fine-tuning requires re-optimization of all parameters of the model, and the scale of model parameters in industrial scenarios is large and hardware resources are limited, full fine-tuning not only has high computational overhead, but may also cause overfitting problems when there is insufficient data in the new scenario, further reducing the stability of the model.

[0006] Existing belt defect recognition technology has problems such as model collapse, feature forgetting, and decreased detection accuracy under the condition of a small amount of labeled data, which makes it difficult to deploy the model in new scenarios with a small number of typical samples. Existing algorithms rely heavily on large-scale, high-quality labeled data, and the significant feature of belt damage industrial scene defect data is the low sample size, so it is difficult to obtain data that meets the needs. In addition, a small amount of labeled data can easily cause category imbalance problems, making it difficult for the algorithm to effectively learn the characteristics of new defect categories. It may also forget the knowledge of existing categories, limiting the generalization ability of the model and the actual application effect.

[0007] In the existing technology, there is a lack of an efficient and accurate cross-scene belt defect recognition method based on typical sample learning and requiring only a small amount of annotation. Summary of the invention

[0008] In order to solve the technical problems of the existing technology in adapting to new scenarios, computing resource efficiency and few-sample learning, the embodiment of the present invention provides a cross-scenario belt defect recognition method and device based on adaptive typical sample learning technology. The technical solution is as follows:

[0009] On the one hand, a cross-scenario belt defect recognition method based on an adaptive typical sample learning technology is provided, the method is implemented by a cross-scenario belt defect recognition device, and the method includes:

[0010] Use cameras to capture images of belts in various industrial scenarios, build supporting image datasets, and query image datasets;

[0011] Inputting the supporting image data set and the query image data set into a defect recognition basic model for feature extraction to obtain a first supporting image feature and a first query image feature;

[0012] Inputting the first supporting image feature and the first query image feature into a prototype adaptive module for feature enhancement to obtain a third supporting image feature and a third query image feature;

[0013] The step of inputting the first supporting image feature and the first query image feature into the prototype adaptive module for feature enhancement to obtain the third supporting image feature and the third query image feature includes:

[0014] Inputting the first supporting image feature and the first query image feature into a prototype enhancement module for feature enhancement to obtain a second supporting image feature and a second query image feature;

[0015] Inputting the second supporting image feature and the second query image feature into a learnable adaptive module for adaptive adjustment to obtain a third supporting image feature and a third query image feature;

[0016] Inputting the third supporting image feature and the third query image feature into a target detection module for target detection to obtain a detection defect image dataset; constructing a loss function according to the supporting image dataset, the query image dataset and the detection defect image dataset;

[0017] According to the loss function, the prototype adaptive model is optimized to obtain an optimized prototype adaptive model;

[0018] Acquire a belt image data set to be identified; perform belt defect identification based on the belt image data set to be identified and the defect identification basic model and the optimized prototype adaptive model.

[0019] On the other hand, a cross-scenario belt defect recognition device based on adaptive typical sample learning technology is provided, and the device is applied to a cross-scenario belt defect recognition method based on adaptive typical sample learning technology, and the device includes:

[0020] The dataset construction module is used to capture images of belts in various industrial scenes through cameras, build supporting image datasets, and query image datasets;

[0021] A feature extraction module, used for inputting the supporting image data set and the query image data set into the defect recognition basic model for feature extraction, so as to obtain a first supporting image feature and a first query image feature;

[0022] A feature enhancement module, configured to input the first supporting image feature and the first query image feature into a prototype adaptive module for feature enhancement to obtain a third supporting image feature and a third query image feature;

[0023] Wherein, the feature enhancement module is further used for:

[0024] Inputting the first supporting image feature and the first query image feature into a prototype enhancement module for feature enhancement to obtain a second supporting image feature and a second query image feature;

[0025] Inputting the second supporting image feature and the second query image feature into a learnable adaptive module for adaptive adjustment to obtain a third supporting image feature and a third query image feature;

[0026] A loss function construction module is used to input the third supporting image feature and the third query image feature into a target detection module for target detection to obtain a detection defect image dataset; and to construct a loss function according to the supporting image dataset, the query image dataset and the detection defect image dataset;

[0027] A model optimization module, used to optimize the prototype adaptive model according to the loss function to obtain an optimized prototype adaptive model;

[0028] The belt defect recognition module is used to obtain a belt image data set to be recognized; and perform belt defect recognition based on the belt image data set to be recognized and based on the defect recognition basic model and the optimized prototype adaptive model.

[0029] On the other hand, a cross-scene belt defect identification device is provided, which includes: a processor; a memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, any one of the above-mentioned cross-scene belt defect identification methods based on adaptive typical sample learning technology is implemented.

[0030] On the other hand, a computer-readable storage medium is provided, in which at least one instruction is stored. The at least one instruction is loaded and executed by a processor to implement any one of the above-mentioned cross-scenario belt defect identification methods based on adaptive typical sample learning technology.

[0031] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0032] The present invention proposes a cross-scenario belt defect recognition method based on adaptive typical sample learning technology. Aiming at the problem that the existing belt defect recognition basic model cannot be directly adapted to the new scene due to structural differences, a plug-and-play modular fine-tuning method is proposed, which can flexibly adapt to belt defect recognition algorithms of various structures, greatly improving the versatility and migration ability of the model; in order to solve the problem of excessive consumption of computing resources when the belt defect recognition model is fully fine-tuned, a learnable adaptive module is designed, which effectively reduces the amount of fine-tuning calculation by combining the down-sampling projection and up-sampling projection of the features, and accurately captures the belt defect features in the new scene; in view of the limitation of the belt defect recognition algorithm to be trained under a small amount of labeled data, a prototype enhancement module is proposed to construct a defect category prototype library, and realize efficient matching of feature extraction and category information through a similarity mechanism. While learning the new defect category information, the feature information of the original category is fully retained, and the detection performance in the typical sample scene is significantly improved. The present invention is an efficient and accurate cross-scenario belt defect recognition method based on typical sample learning that only requires a small amount of annotation. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0034] Figure 1 This is a flow chart of a cross-scenario belt defect recognition method based on an adaptive typical sample learning technology provided by an embodiment of the present invention;

[0035] Figure 2 It is a schematic diagram of the structure of a prototype enhancement module provided by an embodiment of the present invention;

[0036] Figure 3 is a schematic diagram of a learnable adaptive module structure provided by an embodiment of the present invention;

[0037] Figure 4 is a schematic diagram of the structure of a target detection module provided by an embodiment of the present invention;

[0038] Figure 5 It is a block diagram of a cross-scenario belt defect recognition device based on an adaptive typical sample learning technology provided by an embodiment of the present invention;

[0039] Figure 6 It is a structural schematic diagram of a cross-scenario belt defect recognition device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0040] The technical solution of the present invention is described below in conjunction with the accompanying drawings.

[0041] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "example" is intended to present the concept in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.

[0042] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same. "of", "corresponding, relevant" and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same.

[0043] In the embodiments of the present invention, sometimes the subscripts such as W1 It may be written in non-subscript form such as W1. When the difference is not emphasized, the meaning is the same.

[0044] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.

[0045] The embodiment of the present invention provides a cross-scenario belt defect recognition method based on adaptive typical sample learning technology, which can be implemented by a cross-scenario belt defect recognition device, which can be a terminal or a server. Figure 1 The flowchart of the cross-scenario belt defect recognition method based on the adaptive typical sample learning technology is shown. The processing flow of the method may include the following steps:

[0046] S1. Use cameras to capture images of belts in various industrial scenes, build support image datasets and query image datasets.

[0047] Optionally, a camera is used to capture images of belts in various industrial scenes, and a supporting image dataset and a query image dataset are constructed, including:

[0048] The camera is used to shoot belts in various industrial scenes to obtain a belt image dataset containing normal states and various defect states;

[0049] Label the belt image dataset to obtain a typical defect sample dataset;

[0050] The typical defect sample data set is divided according to a preset ratio to obtain a supporting image data set and a query image data set.

[0051] In a feasible implementation, in a belt working scenario, a 2048-resolution linear array camera (with an image acquisition resolution of 2048×1) is used to photograph the running belt, and 2000 consecutive frames are taken at a certain frequency to form a 2048×2000 high-resolution two-dimensional belt tear image. Defective and non-defective images are screened to construct a data set.

[0052] The defects in the above images are divided into C different defect categories according to different shape, texture features and severity. The belt images are labeled using the labelimg tool. The labeling format uses the YOLO target detection labeling format (c t , x t , y t , h t , w t), where subscript t represents the true value of the label, c represents the defect category number in the range of (0, C), (x, y) represents the center coordinates of the target detection box, (h, w) represents the height and width of the target detection box, x±w / 2 should be in the range of 0~2048, y±h / 2 should be in the range of 0~2000, and the output is stored in txt format.

[0053] The typical defect sample dataset includes a support image dataset and a query image dataset. The support image dataset and the query image dataset are composed of images and target detection labels, respectively. The support dataset is used to provide feature information of the target category to help the model identify similar samples during reasoning, and the samples in the support set extract features through the model to build the feature representation of the category and provide a comparative "template" for the model; the query dataset is the image set that the model needs to detect, which is used to evaluate the model's recognition ability for typical belt defect samples. The distribution of targets in the query set does not overlap with the samples in the support set, in order to enhance the generalization ability of the model. Finally, the data is flipped, translated, and other image enhancements and data normalization are performed before being sent to the model for training.

[0054] S2. Input the supporting image dataset and the query image dataset into the defect recognition basic model for feature extraction to obtain a first supporting image feature and a first query image feature.

[0055] In a feasible implementation, the present invention fine-tunes the existing belt defect recognition basic model by inserting modules. This structural design makes the present invention flexible and applicable to the application of belt defect recognition models with different structures to new scenarios. The input data is processed layer by layer by the trained backbone network basic model to extract the first support image feature and the first query image feature, which are denoted as F s_n and F q_n .

[0056] Among them, the basic model of defect recognition is a pre-trained deep neural network.

[0057] In a feasible implementation, the defect recognition basic model has been fully trained on a large-scale data set and has good feature extraction capabilities. The model is used as the basic model for the subsequent adaptive belt defect typical sample learning stage.

[0058] S3, inputting the first supporting image feature and the first query image feature into the prototype adaptive module for feature enhancement to obtain a third supporting image feature and a third query image feature;

[0059] The first supporting image feature and the first query image feature are input into the prototype adaptive module for feature enhancement to obtain the third supporting image feature and the third query image feature, including:

[0060] Inputting the first supporting image feature and the first query image feature into the prototype enhancement module for feature enhancement to obtain the second supporting image feature and the second query image feature;

[0061] The second supporting image feature and the second query image feature are input into a learnable adaptive module for adaptive adjustment to obtain a third supporting image feature and a third query image feature.

[0062] In a feasible implementation, the prototype adaptive module proposed in the present invention is composed of a prototype enhancement module and a learnable adaptive module. The former aims to register the defect category prototypes in the support image and enhance the model's feature expression capabilities for the support image and the query image. The latter aims to enable the model to effectively adapt to defect feature extraction in new scenarios through training.

[0063] Optionally, inputting the first supporting image feature and the first query image feature into a prototype enhancement module for feature enhancement to obtain a second supporting image feature and a second query image feature includes:

[0064] Based on the preset target detection frame, the defect local details are cropped according to the supporting image data set to obtain the defect local image data set;

[0065] Performing element-by-element multiplication according to the first supporting image feature and the defect local image data set to obtain a first category prototype;

[0066] Based on multiple category prototypes in the defect category prototype warehouse, similarity measurement is performed according to the first category prototype to obtain the second category prototype;

[0067] Perform weighted calculation based on the first category prototype and the second category prototype to obtain the third category prototype;

[0068] Based on the third category prototype, feature enhancement is performed according to the first supporting image feature and the first query image feature to obtain the second supporting image feature and the second query image feature.

[0069] In a feasible implementation, the prototype enhancement module is a key component of the prototype adaptive module. The module structure is as follows: Figure 2 As shown in the figure, the main function is to dynamically enhance the category features of supporting samples so that they can more effectively characterize the target category and improve the performance of identifying typical samples of belt defects.

[0070] The input of the prototype enhancement module is the first support image feature F s_n and the first query image feature F q_n , by calculating the support image feature F s_n And support image defect cropping I s_cThe first category prototype P of the target class is obtained by multiplying the defect local detail image (cropped from the support image according to the target detection frame annotated in the support image) t .

[0071] Calculate P t and each category prototype in the existing defect category prototype warehouse (P 1 , P 2 , ..., P n ) and select the second category prototype P corresponding to the i-th category with the highest score i It is used to enhance image features in the future. t and P i The weighted sum of the third category prototypes is used as the new P i Through continuous training, the category prototypes in the defect category prototype warehouse are becoming more and more accurate. In this way, the defect category prototype warehouse is dynamically updated to support the model's generalization ability for multiple categories in new scenarios, while retaining the original defect category feature information extraction ability and updating and registering new defect category information.

[0072] Optionally, based on the third category prototype, feature enhancement is performed according to the first supporting image feature and the first query image feature to obtain the second supporting image feature and the second query image feature, including:

[0073] Perform matrix multiplication according to the first supporting image feature and the third category prototype to obtain a supporting image similarity matrix; perform calculation according to the first supporting image feature and the supporting image similarity matrix to obtain an enhanced supporting image feature;

[0074] Perform matrix multiplication based on the first query image feature and the third category prototype to obtain a query image similarity matrix; perform calculation based on the first query image feature and the query image similarity matrix to obtain an enhanced query image feature;

[0075] Adding the first supporting image feature and the enhanced supporting image feature element by element to obtain a fused supporting image feature;

[0076] Adding the first query image feature and the enhanced query image feature element by element to obtain a fused query image feature;

[0077] The fused supporting image features and the fused query image features are input into the defect recognition basic model for high-dimensional feature extraction to obtain the second supporting image features and the second query image features.

[0078] In a feasible implementation, based on the third category prototype obtained in the above steps, the module measures the first supporting image feature F by similarity. s_n and the first query image feature Fq_n , respectively with the category prototype P i Perform association calculations to generate prototype enhancement features corresponding to defect categories.

[0079] Image feature F s_n Or the image feature F q_n With the third category prototype P i Perform matrix multiplication to obtain a similarity matrix, and use the similarity matrix as a weight to perform dot multiplication with the image features to obtain image features enhanced by the category prototype. This process guides feature adjustment by supporting the defect category prototype representation of the sample, thereby improving the correlation between the query feature and the supporting feature.

[0080] The image features after the third category prototype enhancement are added to the original image features for feature fusion to generate support image features or query image features containing category prototype information. This layer-by-layer feature enhancement and fusion strategy fully exploits the defect characteristics of the support image and query image, and improves the accuracy and robustness of feature representation. The fused features continue to be input into the subsequent network layers to further extract higher-level semantic information.

[0081] Optionally, inputting the second supporting image feature and the second query image feature into a learnable adaptive module for adaptive adjustment to obtain a third supporting image feature and a third query image feature includes:

[0082] According to the second supporting image features and the second query image features, a dimensionality reduction mapping is performed through a downsampling projection layer to obtain a reduced dimensionality supporting image feature and a reduced dimensionality query image feature;

[0083] According to the reduced dimension support image features and the reduced dimension query image features, the ReLU activation function is used to process them to obtain the activated support image features and the activated query image features;

[0084] According to the activated supporting image features and the activated query image features, high-dimensional mapping is performed through an upsampling projection layer to obtain a third supporting image feature and a third query image feature.

[0085] In a feasible implementation, the learnable adaptive module is another important component of the prototype adaptive module, which aims to further optimize the expression ability of the supporting image features and the query image features, so that they can be adaptively adjusted in a variety of scenarios to enhance the defect recognition and detection effect. Figure 3 The input of the module is the second support image features and the second query image features containing the category prototype information. These features have been processed by the prototype enhancement module and have preliminary category relevance.

[0086] The feature input passes through the downsampling projection layer, and the high-dimensional features are mapped to a more compact feature space through dimensionality reduction operations, reducing the interference of redundant information and extracting more essential feature representations. The downsampling operation makes the features more focused, which is conducive to subsequent nonlinear activation. After being processed by the ReLU activation function, the module models the nonlinear relationship of the features. The ReLU activation function can effectively enhance the expressiveness of the model, while suppressing negative features, improving the sparsity and robustness of the features.

[0087] The activated features are restored through the upsampling projection layer, which remaps the reduced features back to the original feature space dimensions. The upsampling projection layer can introduce additional adaptive learning capabilities in the feature recovery process, so that the output features can be dynamically adjusted while retaining the original feature structure.

[0088] The second supporting image features and the second query image features are adaptively adjusted by the module, and the enhanced features and the third supporting image features F are output respectively. s_n * and the third query image feature F q_n * These features integrate category characteristics and contextual information at a higher level, further improving the accuracy of feature representation and classification ability.

[0089] S4. Input the third supporting image feature and the third query image feature into the target detection module for target detection to obtain a detection defect image dataset; and construct a loss function according to the supporting image dataset, the query image dataset, and the detection defect image dataset.

[0090] In a feasible implementation manner, in the present invention, the main function of the target detection module is to generate a detection result (c o , x o , y o , h o , w o ), including the category label c o , target center coordinates (x o , y o ) and the bounding box size (h o , w o ), the module structure is as follows Figure 4 shown.

[0091] The input feature map is processed by 1×1 convolution to extract high-level semantic information and target location information. These features are processed by the category prediction branch and the bounding box regression branch to output the final output result. Among them, the category prediction branch is responsible for determining the type of the target, using two 3×3 convolutions and one 1×1 convolution to extract the category information of the image, calculate the classification probability distribution of the target, and generate the H×W×C classification result; the bounding box regression branch generates the location and size parameters of the target through regression analysis to determine the exact location of the target in the query image; first, two 3×3 convolutions are used for feature screening, and then a 1×1 convolution is used to output H×W×4 location information and H×W×1 intersection-over-union score, which are organized as (c o , x o , y o ,h o , w o )’s defect detection output.

[0092] S5. According to the loss function, the prototype adaptive model is optimized to obtain an optimized prototype adaptive model.

[0093] In a feasible implementation, the loss function reflects the difference between the model's detection results for the query image and the support image and the corresponding target detection labels. The present invention splits the leather defect recognition task into two tasks: category prediction and bounding box regression. Therefore, the overall loss function of the present invention includes two parts: classification loss and regression loss, which are combined in a weighted manner, wherein the classification loss function is a binary cross entropy loss function, and the regression loss function is a complete intersection-over-union loss function. In the typical sample recognition task of belt defects, the present invention optimizes the back propagation of the prototype adaptive model based on the loss function, taking into account the accuracy of classification and positioning.

[0094] S6. Obtain a data set of belt images to be identified; perform belt defect identification based on the data set of belt images to be identified and a defect identification basic model and an optimized prototype adaptive model.

[0095] In a feasible implementation, the present invention focuses on the synergy of the support set and the query set in the training process, guides the model to learn the key features of the target category through a small number of labeled samples, and maximizes the generalization ability of the small sample category. In the inference stage, the model uses the category prototype features learned through training and the category features of a small amount of labeled support data to effectively identify the defect category in the new scenario.

[0096] The present invention proposes a cross-scenario belt defect recognition method based on adaptive typical sample learning technology. Aiming at the problem that the existing belt defect recognition basic model cannot be directly adapted to the new scene due to structural differences, a plug-and-play modular fine-tuning method is proposed, which can flexibly adapt to belt defect recognition algorithms of various structures, greatly improving the versatility and migration ability of the model; in order to solve the problem of excessive consumption of computing resources when the belt defect recognition model is fully fine-tuned, a learnable adaptive module is designed, which effectively reduces the amount of fine-tuning calculation by combining the down-sampling projection and up-sampling projection of the features, and accurately captures the belt defect features in the new scene; in view of the limitation of the belt defect recognition algorithm to be trained under a small amount of labeled data, a prototype enhancement module is proposed to construct a defect category prototype library, and realize efficient matching of feature extraction and category information through a similarity mechanism. While learning the new defect category information, the feature information of the original category is fully retained, and the detection performance in the typical sample scene is significantly improved. The present invention is an efficient and accurate cross-scenario belt defect recognition method based on typical sample learning that only requires a small amount of annotation.

[0097] Figure 5 The block diagram of a cross-scenario belt defect recognition device based on adaptive typical sample learning technology is shown according to an exemplary embodiment. The device is used in a cross-scenario belt defect recognition method based on adaptive typical sample learning technology. Figure 5 The device includes a data set construction module 510, a feature extraction module 520, a feature enhancement module 530, a loss function construction module 540, a model optimization module 550 and a belt defect recognition module 560. Among them:

[0098] A data set construction module 510 is used to capture images of belts in various industrial scenes through a camera, and to construct a supporting image data set and a query image data set;

[0099] A feature extraction module 520 is used to input the supporting image data set and the query image data set into the defect recognition basic model for feature extraction to obtain a first supporting image feature and a first query image feature;

[0100] A feature enhancement module 530 is used to input the first supporting image feature and the first query image feature into the prototype adaptive module for feature enhancement to obtain a third supporting image feature and a third query image feature;

[0101] Among them, the feature enhancement module is further used to:

[0102] Inputting the first supporting image feature and the first query image feature into the prototype enhancement module for feature enhancement to obtain the second supporting image feature and the second query image feature;

[0103] Inputting the second supporting image feature and the second query image feature into a learnable adaptive module for adaptive adjustment to obtain a third supporting image feature and a third query image feature;

[0104] A loss function construction module 540 is used to input the third supporting image feature and the third query image feature into the target detection module to perform target detection, and obtain a detection defect image dataset; and construct a loss function according to the supporting image dataset, the query image dataset, and the detection defect image dataset;

[0105] A model optimization module 550 is used to optimize the prototype adaptive model according to the loss function to obtain an optimized prototype adaptive model;

[0106] The belt defect recognition module 560 is used to obtain a belt image data set to be recognized; and perform belt defect recognition based on the belt image data set to be recognized based on a defect recognition basic model and an optimized prototype adaptive model.

[0107] Optionally, the data set construction module 510 is further configured to:

[0108] The camera is used to shoot belts in various industrial scenes to obtain a belt image dataset containing normal states and various defect states;

[0109] Label the belt image dataset to obtain a typical defect sample dataset;

[0110] The typical defect sample data set is divided according to a preset ratio to obtain a supporting image data set and a query image data set.

[0111] Among them, the basic model of defect recognition is a pre-trained deep neural network.

[0112] Optionally, the feature enhancement module 530 is further configured to:

[0113] Based on the preset target detection frame, the defect local details are cropped according to the supporting image data set to obtain the defect local image data set;

[0114] Performing element-by-element multiplication according to the first supporting image feature and the defect local image data set to obtain a first category prototype;

[0115] Based on multiple category prototypes in the defect category prototype warehouse, similarity measurement is performed according to the first category prototype to obtain the second category prototype;

[0116] Perform weighted calculation based on the first category prototype and the second category prototype to obtain the third category prototype;

[0117] Based on the third category prototype, feature enhancement is performed according to the first supporting image feature and the first query image feature to obtain the second supporting image feature and the second query image feature.

[0118] Optionally, the feature enhancement module 530 is further configured to:

[0119] Perform matrix multiplication according to the first supporting image feature and the third category prototype to obtain a supporting image similarity matrix; perform calculation according to the first supporting image feature and the supporting image similarity matrix to obtain an enhanced supporting image feature;

[0120] Perform matrix multiplication based on the first query image feature and the third category prototype to obtain a query image similarity matrix; perform calculation based on the first query image feature and the query image similarity matrix to obtain an enhanced query image feature;

[0121] Adding the first supporting image feature and the enhanced supporting image feature element by element to obtain a fused supporting image feature;

[0122] Adding the first query image feature and the enhanced query image feature element by element to obtain a fused query image feature;

[0123] The fused supporting image features and the fused query image features are input into the defect recognition basic model for high-dimensional feature extraction to obtain the second supporting image features and the second query image features.

[0124] Optionally, the feature enhancement module 530 is further configured to:

[0125] According to the second supporting image features and the second query image features, a dimensionality reduction mapping is performed through a downsampling projection layer to obtain a reduced dimensionality supporting image feature and a reduced dimensionality query image feature;

[0126] According to the reduced dimension support image features and the reduced dimension query image features, the ReLU activation function is used to process them to obtain the activated support image features and the activated query image features;

[0127] According to the activated supporting image features and the activated query image features, high-dimensional mapping is performed through an upsampling projection layer to obtain a third supporting image feature and a third query image feature.

[0128] The present invention proposes a cross-scenario belt defect recognition method based on adaptive typical sample learning technology. Aiming at the problem that the existing belt defect recognition basic model cannot be directly adapted to the new scene due to structural differences, a plug-and-play modular fine-tuning method is proposed, which can flexibly adapt to belt defect recognition algorithms of various structures, greatly improving the versatility and migration ability of the model; in order to solve the problem of excessive consumption of computing resources when the belt defect recognition model is fully fine-tuned, a learnable adaptive module is designed, which effectively reduces the amount of fine-tuning calculation by combining the down-sampling projection and up-sampling projection of the features, and accurately captures the belt defect features in the new scene; in view of the limitation of the belt defect recognition algorithm to be trained under a small amount of labeled data, a prototype enhancement module is proposed to construct a defect category prototype library, and realize efficient matching of feature extraction and category information through a similarity mechanism. While learning the new defect category information, the feature information of the original category is fully retained, and the detection performance in the typical sample scene is significantly improved. The present invention is an efficient and accurate cross-scenario belt defect recognition method based on typical sample learning that only requires a small amount of annotation.

[0129] Figure 6 is a schematic diagram of the structure of a cross-scenario belt defect recognition device provided by an embodiment of the present invention, such as Figure 6 As shown, the cross-scenario belt defect recognition device may include the above Figure 5 The cross-scenario belt defect recognition device based on the adaptive typical sample learning technology is shown. Optionally, the cross-scenario belt defect recognition device 610 may include a first processor 2001.

[0130] Optionally, the cross-scene belt defect identification device 610 may also include a memory 2002 and a transceiver 2003.

[0131] The first processor 2001, the memory 2002 and the transceiver 2003 may be connected via a communication bus.

[0132] Combine the following Figure 6 The various components of the cross-scenario belt defect recognition device 610 are introduced in detail:

[0133] The first processor 2001 is the control center of the cross-scenario belt defect recognition device 610, and can be a processor or a general term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention, such as one or more microprocessors (digital signal processors, DSPs), or one or more field programmable gate arrays (field programmable gate arrays, FPGAs).

[0134] Optionally, the first processor 2001 can perform various functions of the cross-scene belt defect identification device 610 by running or executing a software program stored in the memory 2002 and calling data stored in the memory 2002.

[0135] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 6 CPU0 and CPU1 are shown in FIG.

[0136] In a specific implementation, as an embodiment, the cross-scenario belt defect recognition device 610 may also include multiple processors, such as Figure 6 The first processor 2001 and the second processor 2004 are shown in FIG. Each of these processors may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).

[0137] The memory 2002 is used to store the software program for executing the solution of the present invention, and is controlled to be executed by the first processor 2001. The specific implementation method can refer to the above method embodiment, which will not be repeated here.

[0138] Optionally, the memory 2002 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 may be integrated with the first processor 2001, or may exist independently, and may be accessed through the interface circuit ( Figure 6 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.

[0139] The transceiver 2003 is used to communicate with a network device or a terminal device.

[0140] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 6 The receiver is used to implement a receiving function, and the transmitter is used to implement a sending function.

[0141] Optionally, the transceiver 2003 may be integrated with the first processor 2001, or may exist independently, and may be connected to the first processor 2001 through the interface circuit ( Figure 6 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.

[0142] It should be noted that Figure 6 The structure of the cross-scenario belt defect identification device 610 shown in the figure does not constitute a limitation on the router. The actual knowledge structure identification device may include more or fewer components than shown in the figure, or a combination of certain components, or a different arrangement of components.

[0143] In addition, the technical effects of the cross-scene belt defect recognition device 610 can refer to the technical effects of the cross-scene belt defect recognition method based on the adaptive typical sample learning technology described in the above method embodiment, and will not be repeated here.

[0144] It should be understood that the first processor 2001 in the embodiment of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0145] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0146] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware or any other combination. When implemented by software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.

[0147] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.

[0148] In the present invention, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0149] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0150] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0151] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0152] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0153] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0154] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0155] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program codes.

[0156] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. A cross-scenario belt defect recognition method based on adaptive typical sample learning technology, characterized in that: The method comprises: Use cameras to capture images of belts in various industrial scenarios, build supporting image datasets, and query image datasets; Inputting the supporting image data set and the query image data set into a defect recognition basic model for feature extraction to obtain a first supporting image feature and a first query image feature; Inputting the first supporting image feature and the first query image feature into a prototype adaptive module for feature enhancement to obtain a third supporting image feature and a third query image feature; The step of inputting the first supporting image feature and the first query image feature into the prototype adaptive module for feature enhancement to obtain the third supporting image feature and the third query image feature includes: Inputting the first supporting image feature and the first query image feature into a prototype enhancement module for feature enhancement to obtain a second supporting image feature and a second query image feature; Inputting the second supporting image feature and the second query image feature into a learnable adaptive module for adaptive adjustment to obtain a third supporting image feature and a third query image feature; Inputting the third supporting image feature and the third query image feature into a target detection module for target detection to obtain a detection defect image dataset; constructing a loss function according to the supporting image dataset, the query image dataset and the detection defect image dataset; According to the loss function, the prototype adaptive model is optimized to obtain an optimized prototype adaptive model; Acquire a belt image data set to be identified; perform belt defect identification based on the belt image data set to be identified and the defect identification basic model and the optimized prototype adaptive model.

2. The cross-scenario belt defect recognition method based on adaptive typical sample learning technology according to claim 1 is characterized in that: The method uses a camera to capture images of belts in various industrial scenes, builds a supporting image dataset and a query image dataset, including: The camera is used to shoot belts in various industrial scenes to obtain a belt image dataset containing normal states and various defect states; Annotating the belt image dataset to obtain a typical defect sample dataset; The typical defect sample data set is divided according to a preset ratio to obtain a supporting image data set and a query image data set.

3. The cross-scenario belt defect recognition method based on adaptive typical sample learning technology according to claim 1 is characterized in that: The defect recognition basic model is a pre-trained deep neural network.

4. The cross-scenario belt defect recognition method based on adaptive typical sample learning technology according to claim 1 is characterized in that: The step of inputting the first supporting image feature and the first query image feature into a prototype enhancement module for feature enhancement to obtain a second supporting image feature and a second query image feature includes: Based on a preset target detection frame, cropping the local details of the defect according to the supporting image dataset to obtain a local defect image dataset; Performing element-by-element multiplication on the first supporting image feature and the defect local image data set to obtain a first category prototype; Based on multiple category prototypes in the defect category prototype warehouse, similarity measurement is performed according to the first category prototype to obtain a second category prototype; Perform weighted calculation based on the first category prototype and the second category prototype to obtain a third category prototype; Based on the third category prototype, feature enhancement is performed according to the first supporting image feature and the first query image feature to obtain a second supporting image feature and a second query image feature.

5. The cross-scenario belt defect recognition method based on adaptive typical sample learning technology according to claim 4 is characterized in that: The step of performing feature enhancement based on the third category prototype and according to the first supporting image feature and the first query image feature to obtain the second supporting image feature and the second query image feature includes: Performing matrix multiplication according to the first supporting image feature and the third category prototype to obtain a supporting image similarity matrix; performing calculation according to the first supporting image feature and the supporting image similarity matrix to obtain an enhanced supporting image feature; Performing matrix multiplication according to the first query image feature and the third category prototype to obtain a query image similarity matrix; performing calculation according to the first query image feature and the query image similarity matrix to obtain an enhanced query image feature; Adding the first supporting image feature and the enhanced supporting image feature element by element to obtain a fused supporting image feature; Adding the first query image feature and the enhanced query image feature element by element to obtain a fused query image feature; The fused supporting image features and the fused query image features are input into the defect recognition basic model for high-dimensional feature extraction to obtain second supporting image features and second query image features.

6. The cross-scenario belt defect recognition method based on adaptive typical sample learning technology according to claim 1 is characterized in that: The step of inputting the second supporting image feature and the second query image feature into a learnable adaptive module for adaptive adjustment to obtain a third supporting image feature and a third query image feature includes: According to the second supporting image features and the second query image features, performing dimensionality reduction mapping through a downsampling projection layer to obtain reduced dimensionality supporting image features and reduced dimensionality query image features; According to the reduced dimension support image features and the reduced dimension query image features, processing is performed through a ReLU activation function to obtain activated support image features and activated query image features; According to the activated supporting image features and the activated query image features, high-dimensional mapping is performed through an upsampling projection layer to obtain a third supporting image feature and a third query image feature.

7. A cross-scene belt defect recognition device based on adaptive typical sample learning technology, the cross-scene belt defect recognition device based on adaptive typical sample learning technology is used to implement the cross-scene belt defect recognition method based on adaptive typical sample learning technology as claimed in any one of claims 1-6, characterized in that: The device comprises: The dataset construction module is used to capture images of belts in various industrial scenes through cameras, build supporting image datasets, and query image datasets; A feature extraction module, used for inputting the supporting image data set and the query image data set into the defect recognition basic model for feature extraction, so as to obtain a first supporting image feature and a first query image feature; A feature enhancement module, configured to input the first supporting image feature and the first query image feature into a prototype adaptive module for feature enhancement to obtain a third supporting image feature and a third query image feature; Wherein, the feature enhancement module is further used for: Inputting the first supporting image feature and the first query image feature into a prototype enhancement module for feature enhancement to obtain a second supporting image feature and a second query image feature; Inputting the second supporting image feature and the second query image feature into a learnable adaptive module for adaptive adjustment to obtain a third supporting image feature and a third query image feature; A loss function construction module is used to input the third supporting image feature and the third query image feature into a target detection module for target detection to obtain a detection defect image dataset; and to construct a loss function according to the supporting image dataset, the query image dataset and the detection defect image dataset; A model optimization module, used to optimize the prototype adaptive model according to the loss function to obtain an optimized prototype adaptive model; The belt defect recognition module is used to obtain a belt image data set to be recognized; and perform belt defect recognition based on the belt image data set to be recognized and based on the defect recognition basic model and the optimized prototype adaptive model.

8. The cross-scenario belt defect recognition device based on adaptive typical sample learning technology according to claim 1 is characterized in that: The feature enhancement module is further used to: Based on a preset target detection frame, cropping the local details of the defect according to the supporting image dataset to obtain a local defect image dataset; Performing element-by-element multiplication on the first supporting image feature and the defect local image data set to obtain a first category prototype; Based on multiple category prototypes in the defect category prototype warehouse, similarity measurement is performed according to the first category prototype to obtain a second category prototype; Perform weighted calculation based on the first category prototype and the second category prototype to obtain a third category prototype; Based on the third category prototype, feature enhancement is performed according to the first supporting image feature and the first query image feature to obtain a second supporting image feature and a second query image feature.

9. A cross-scenario belt defect recognition device, characterized in that: The cross-scenario belt defect recognition device includes: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 6 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, which can be called by a processor to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Defect prediction method and system for small sample knowledge transfer learning

    CN114663401A

  • Conveyor belt surface detection method and system based on machine vision

    CN115272980A

  • Method and system for detecting surface defects of thread bushing based on distillation learning

    CN115457042A

  • Defect detection method and device based on transfer learning and small sample learning, and medium

    CN116109627A

  • Metafeature enhancement-based small sample PCB defect detection method

    CN118587176A