Industrial visual anomaly identification method, system and device based on multi-modal feedback and medium

Through multimodal feedback methods, combined with deep learning and multimodal large models, the target detection model is adaptively trained, which solves the problems of poor adaptability of traditional methods and high cost of deep learning, and realizes efficient, accurate recognition and rapid adaptation of industrial visual inspection.

CN120635762APending Publication Date: 2025-09-12山东浪潮智能生产技术有限公司
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202510516982.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Traditional industrial visual inspection methods have poor adaptability, rely on manual rules or thresholds, and are difficult to cope with complex scenarios; deep learning methods require a large amount of labeled data, are costly and difficult to iterate and optimize, and have a high model misrecognition rate, making it difficult to meet the requirements of high precision and high recall rate.

Method used

A multimodal feedback method is adopted to collect video streams in real time through streaming protocols. A deep learning model is used to extract features and combined with a large multimodal model for anomaly identification. The target detection model is adaptively trained and the learning rate and sample ratio are dynamically adjusted to achieve rapid adaptation to changes in industrial scenarios.

Benefits of technology

It improves the accuracy and adaptability of abnormal image recognition, reduces the need for manual data annotation, reduces costs, ensures the real-time and reliability of recognition results, and adapts to the rapid changes in complex industrial environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635762A_ABST
    Figure CN120635762A_ABST
Patent Text Reader

Abstract

The invention provides an industrial visual anomaly recognition method, system and device based on multi-modal feedback and a medium, and belongs to the technical field of industrial visual detection.The method comprises the steps that a production line video stream is collected in real time, an ROI is automatically segmented, a picture sequence is generated, features of a set dimension are extracted, and the features of the set dimension are extracted; comparing a preset dynamic distance threshold with a feature library, and preliminarily identifying an abnormal image; a cue word of a detection requirement is constructed, the cue word and abnormal related information are input into the multi-modal large model, and a determined abnormal image is output; using the determined abnormal image to construct a data set to train a target detection model, and adjusting the learning rate in the training process; and inputting pictures captured in an industrial scene into the deep learning model, the multi-modal large model and the target detection model in sequence to obtain a recognition result, feeding back the recognition result to a user for confirmation, adding the recognition result to the data set, and optimizing training of the target detection model. According to the invention, accurate identification of abnormal images in industrial production is realized, and the identification efficiency is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of industrial visual inspection technology, and specifically relates to an industrial visual anomaly recognition method, system, equipment and medium based on multimodal feedback. Background Art

[0002] In industrial production, visual inspection has become a key link in ensuring product quality and production efficiency. With the development of intelligent and automated manufacturing, visual inspection of the production process has become an important part.

[0003] First, traditional industrial visual inspection relies primarily on manually defined rules or simple thresholds (such as edge detection and color thresholds). This approach suffers from a high false positive rate in complex and changing industrial production scenarios, such as fluctuating lighting conditions and products positioned at multiple angles. For example, changes in business scenarios can affect anomaly detection, necessitating manual intervention to adjust rules or thresholds. This is time-consuming and labor-intensive, and makes it difficult to adapt to new scenarios. Second, deep learning-based image anomaly detection is currently widely used in industrial production. Training deep neural network models can improve recognition accuracy and adaptability to a certain extent, but requires a large amount of labeled data for training, which is costly and time-consuming. A case study at a 3C electronics factory showed that achieving a 90% accuracy rate for a single defect required over 2,000 labeled samples, and the data preparation cycle took up to two months. Furthermore, as the environment changes, trained models are prone to false positives, making them difficult to meet user requirements for high precision and recall.

[0004] In summary, traditional methods in industrial visual inspection rely on manual rules or thresholds, have poor adaptability, and are difficult to cope with complex scenarios. Deep learning methods require a large amount of labeled data, which is costly and difficult to iteratively optimize. In addition, the model has a high misrecognition rate, making it difficult to continuously improve accuracy and recall rates. Moreover, the deployment cycle of new scenarios is long, making it difficult to quickly adapt to dynamically changing environmental requirements. Summary of the Invention

[0005] In a first aspect, an embodiment of the present application provides an industrial visual anomaly recognition method based on multimodal feedback, comprising the following steps: S1. Real-time video streams from the production line are collected through a streaming protocol, and a segmentation algorithm is used to automatically segment the ROI region to generate a time-stamped image sequence. S2. Use a deep learning model to extract features of a specified dimension from the image sequence. Use an indexing tool library to store the extracted features and establish a hierarchical index structure. Then, combine the extracted features with a pre-set dynamic distance threshold to perform preliminary identification of abnormal images. S3. Generate prompt words based on the detection requirements of the industrial scenario, and determine abnormality-related information for the initially identified abnormal image based on the industrial scenario. Then, input the prompt words and abnormality-related information into the multimodal large model, output the determined abnormal image, and set hard and easy sample labels for the determined abnormal image; S4. Use the identified abnormal images to construct a dataset to train the object detection model, and adjust the learning rate and the ratio of hard to easy examples during training; S5. Input the images captured in industrial scenes into the deep learning model, the multimodal large model, and the trained target detection model in sequence to obtain the recognition results of whether they are abnormal images and the type of abnormality. The images and recognition results are then fed back to the user for annotation and confirmation, and added to the dataset to optimize the training of the target detection model.

[0006] Furthermore, the specific steps of step S1 are as follows: S11. Access the video stream from the production line camera via the RTSP protocol and extract image frames at the set sampling rate; S12. Use the Mask R-CNN algorithm to automatically segment the ROI area of ​​the target business object and filter out background interference; S13. Unify the ROI area of ​​each frame image after segmentation to the target pixel size and perform grayscale normalization to complete the standardization process; S14. Add a timestamp, camera number, and production line station information as metadata to each frame of the image, and generate an image sequence in PNG format.

[0007] Furthermore, the specific steps of step S2 are as follows: S21. Use the EfficientNet-B7 convolutional neural network model to extract a feature vector of a specified dimension from the image sequence. S22. Construct a feature library using the vector index Faiss library to store the collected normal feature vectors and establish a hierarchical index structure; S23. Pre-set a dynamic distance threshold based on the industrial scenario, calculate the dynamic distance between each extracted feature vector and the existing feature vectors in the feature library, and when the dynamic distance is greater than the dynamic distance threshold, use the image to which the corresponding feature vector belongs as the initial identified abnormal image.

[0008] Furthermore, the specific steps of step S23 are as follows: S231. Determine the basic distance threshold, compensate the basic distance threshold according to the industrial scenario, and dynamically adjust it according to the set method during the detection process; S232. Compare the currently extracted feature vector with the feature library and determine whether the feature library is empty; If so, the current feature vector is used as the first seed of the feature library, saved in the hierarchical index structure, and a single-node graph is created, ending; If not, proceed to step S233; S233. Search the current feature vector from the top layer of the hierarchical index structure to locate the k candidate feature vectors with the smallest Manhattan distance to the current feature vector in the bottom layer; S234. Take the minimum distance from the Manhattan distances between the k candidate feature vectors and the current feature vector, and determine whether the minimum distance is greater than the dynamic distance threshold; If yes, go to step S235; If not, mark the current feature vector as a normal feature vector and add it to the feature library, and update the node graph of the hierarchical index structure, and end; S235. Use the image corresponding to the current feature vector as the abnormal image for preliminary identification.

[0009] Furthermore, the specific steps of step S3 are as follows: S31. Generate prompt word templates with variable parameters for each industrial scenario; S32 based on the detection needs of industrial scenarios, determine the value of the variable parameter in the prompt word template, generate a prompt word; S33 determines the feature vector, size information and metadata of the initially identified abnormal image as abnormal related information; S34. Integrate the prompt word and abnormality-related information into JSON format, and input it into the visual model together with the initially identified abnormal image, output the determined abnormal image, and determine the abnormality type and confidence level; S35. Set labels of difficult samples, samples to be tested, and easy samples for the determined abnormal image according to the confidence value of the abnormal type.

[0010] Furthermore, the specific steps of step S4 are as follows: S41. Set an initial sampling ratio of 1:n for hard examples to easy examples, where n>1, and construct a dataset by sampling from the identified abnormal images according to the initial sampling ratio until the number of samples reaches a set threshold. S42. Divide the dataset into a training set and a validation set according to a set ratio; S43. Use the PyTorch framework to build the YOLOv9 model as the target detection model, initialize the learning rate, use the deviation between the predicted anomaly type and the real-time anomaly type as the loss function, train the target detection model using the training set, and use the Adagrad optimizer to adjust the target detection model parameters according to the learning rate based on the loss function during training. S44. Perform validation using the validation set every N iterations during the training process, and calculate the loss value calculated in the target detection model using the validation set as input, as well as the recognition accuracy of the difficult examples. S45. When the loss value does not decrease in N iterations, reduce the learning rate according to the set method, and when the recognition accuracy of the difficult example is less than the set ratio threshold, increase the sampling ratio of the difficult example to the easy example; S46. When the loss value is less than the set threshold or the number of iterations exceeds the maximum value, training is stopped to obtain the final target detection model.

[0011] Furthermore, the specific steps of step S5 are as follows: S51. Capture the production line in the industrial scene and obtain real-time images; S52. If the real-time image is sequentially input into the deep learning model, the multimodal large model and the trained target detection model and identified as an abnormal screen, the abnormal screen is digitally signed and uploaded to the user management terminal and stored in the Mysql database; S53 receives the user's abnormal screen classification mark or correct the misjudgment on the user management terminal, and records the abnormal screen label; S54. Regularly add abnormal images and corresponding labels to the data set, and return to step S4 to train the target detection model until the accuracy and recall rate of the target detection model meet user requirements.

[0012] In a second aspect, an embodiment of the present application further provides an industrial visual anomaly recognition system based on multimodal feedback, comprising: The industrial dynamic acquisition module is used to collect production line video streams in real time through streaming protocols, and automatically segment the ROI area using a segmentation algorithm to generate a sequence of images with timestamps; The feature comparison module is used to extract features of a set dimension from an image sequence using a deep learning model, save the extracted features using an index tool library and establish a hierarchical index structure, and then perform preliminary identification of abnormal images based on the extracted features combined with a pre-set dynamic distance threshold; The large-scale model arbitration module is used to generate prompt words based on the detection requirements of industrial scenarios, and determine the abnormality-related information of the initially identified abnormal images based on the industrial scenarios. The prompt words and abnormality-related information are then input into the multimodal large-scale model, and the confirmed abnormal images are output. The hard and easy sample labels are set for the confirmed abnormal images. An adaptive training module is used to train the target detection model using a dataset constructed using identified abnormal images, and to adjust the learning rate and the ratio of hard to easy examples during training. The anomaly recognition and model optimization module is used to input images captured in industrial scenarios into the deep learning model, the multimodal large model, and the trained target detection model in sequence to obtain the recognition results of whether the image is abnormal and the type of anomaly. The images and recognition results are then fed back to the user for annotation and confirmation, and added to the dataset to optimize the training of the target detection model.

[0013] In a third aspect, an embodiment of the present application further provides an electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the industrial visual anomaly recognition method based on multimodal feedback as described in the first aspect are implemented.

[0014] In a fourth aspect, an embodiment of the present application further provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the industrial visual anomaly recognition method based on multimodal feedback as described in the first aspect are implemented.

[0015] It can be seen from the above technical solutions that this application has the following advantages: The industrial visual anomaly recognition method, system, equipment and medium based on multimodal feedback provided in this application can accurately identify abnormal images and their types and effectively reduce the misjudgment rate by combining deep learning models, multimodal large models and target detection models; through adaptive training, it can automatically optimize the model according to new abnormal image data and quickly adapt to changes in different industrial scenarios and abnormal types; reduce the need for manual data annotation, improve data processing efficiency through automated processes, and reduce labor costs; through the use of digital signatures, ensure the integrity and authenticity of data, and achieve data traceability and management; through real-time collection and processing of image data, quickly feedback abnormal recognition results, and optimize the model according to user annotation confirmation, providing real-time and reliability of abnormal recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solution of the present application, the following is a brief introduction to the drawings required for the description. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0017] Figure 1 Schematic diagram of the process of industrial visual anomaly recognition method based on multimodal feedback of the present invention.

[0018] Figure 2 Schematic diagram of the industrial visual anomaly recognition system based on multimodal feedback of the present invention. DETAILED DESCRIPTION

[0019] The various embodiments of the present disclosure will be described in more detail below in the specific steps of the industrial visual anomaly recognition method based on multimodal feedback. The present disclosure may have various embodiments, and adjustments and changes may be made therein. However, it should be understood that there is no intention to limit the various embodiments of the present disclosure to the specific embodiments disclosed herein, but rather that the present disclosure should be understood to encompass all adjustments, equivalents, and / or alternatives that fall within the spirit and scope of the various embodiments of the present disclosure.

[0020] For example, in industrial production, visual inspection technology has become a key means of ensuring product quality and improving production efficiency. As the manufacturing industry transitions toward intelligent and automated processes, the importance of visual inspection in the production process has become increasingly prominent. However, traditional industrial visual inspection methods primarily rely on manually defined rules or simple threshold judgments (such as edge detection and color thresholding). These methods suffer from high false positive rates when faced with complex and changing industrial production scenarios (such as fluctuating lighting conditions and products positioned at multiple angles). Once the business scenario changes, anomaly detection effectiveness is significantly reduced, requiring frequent manual intervention to adjust rules or thresholds. This is not only time-consuming and labor-intensive, but also difficult to adapt to the needs of new scenarios. In recent years, deep learning-based image anomaly detection technology has been widely used in industrial production. By training deep neural network models, recognition accuracy and adaptability can be improved to a certain extent. However, this approach also has significant limitations. For one thing, deep learning models require a large amount of labeled data for training, making data annotation costly and time-consuming. For example, in a 3C electronics factory, achieving a 90% accuracy rate required over 2,000 labeled samples for a single defect type, and the data preparation cycle took up to two months. On the other hand, as the production environment changes, the trained models are prone to misidentification problems and are unable to meet users' requirements for high precision and high recall.

[0021] In summary, traditional industrial visual inspection methods rely on manual rules or thresholds, have poor adaptability, and struggle to cope with complex scenarios. While deep learning methods have improved recognition performance to some extent, they require large amounts of labeled data, are costly, and make iterative optimization difficult. Furthermore, the models suffer from high false positive rates, making it difficult to continuously improve accuracy and recall. The deployment cycle for new scenarios is long, making it difficult to quickly adapt to dynamically changing environmental demands.

[0022] To address the above issues, this embodiment provides an industrial visual anomaly recognition method based on multimodal feedback. Through the steps of dynamic acquisition, feature extraction, multimodal large model arbitration, adaptive training and anomaly recognition, it achieves efficient and accurate recognition of abnormal images in industrial production.

[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0024] See also Figure 1 FIG. 1 is a flow chart of a method for identifying anomalies in industrial vision based on multimodal feedback in a specific embodiment, the method comprising the following steps: S1. Real-time video streams from the production line are collected through a streaming protocol, and a segmentation algorithm is used to automatically segment the ROI region to generate a time-stamped image sequence. It should be noted that the real-time acquisition of production line video streams through streaming protocols and the automatic segmentation of ROI areas to generate image sequences improves the efficiency of obtaining key image information of industrial production. Compared with manual collection and annotation, it improves data collection efficiency and improves data quality. S2. Use a deep learning model to extract features of a specified dimension from the image sequence. Use an indexing tool library to store the extracted features and establish a hierarchical index structure. Then, combine the extracted features with a pre-set dynamic distance threshold to perform preliminary identification of abnormal images. It should be noted that by fully utilizing the feature extraction capabilities of deep learning, using hierarchical indexing to accelerate feature retrieval, and using dynamic thresholds to adapt to changes in industrial scenarios, rapid initial screening of anomalies is achieved with high recognition accuracy, reducing the overall data processing volume; S3. Generate prompt words based on the detection requirements of the industrial scenario, and determine abnormality-related information for the initially identified abnormal image based on the industrial scenario. Then, input the prompt words and abnormality-related information into the multimodal large model, output the determined abnormal image, and set hard and easy sample labels for the determined abnormal image; It should be noted that the semantic understanding and image analysis capabilities of a multimodal large model, combined with the needs of industrial scenarios, can achieve accurate anomaly judgment. By setting sample labels, high-quality labeled data is provided for model training, which improves the accuracy of anomaly recognition compared to single-model judgment. S4. Use the identified abnormal images to construct a dataset to train the object detection model, and adjust the learning rate and the ratio of hard to easy examples during training; It should be noted that through adaptive training, the target detection model can quickly learn abnormal characteristics, optimize parameters and improve performance, enhance the detection ability of complex industrial anomalies, accelerate the convergence of model training and improve detection accuracy; S5. Images captured in industrial scenarios are sequentially fed into a deep learning model, a multimodal large model, and a trained object detection model to identify abnormal images and their types. The images and recognition results are then fed back to the user for annotation and confirmation before being added to the dataset to optimize the training of the object detection model. It should be noted that by inputting the captured images into the model in sequence to obtain recognition results, and optimizing model training after feedback from user annotations and confirmation, it can continuously absorb new data, adapt to changes in industrial scenarios, and improve detection accuracy. Compared with a mechanism without feedback, the target detection model improves operating accuracy and ensures industrial production quality.

[0025] This embodiment combines a multimodal large model with a target detection model to achieve accurate recognition and dynamic optimization of abnormal images. Compared with a single model, it improves recognition accuracy and ensures industrial production quality in various industrial scenarios.

[0026] Furthermore, as a refinement and extension of the specific implementation of the above embodiment, in order to fully illustrate the specific implementation process in this embodiment, another industrial visual anomaly recognition method based on multimodal feedback is provided, which includes the following steps: S1. Real-time video streams from the production line are collected using a streaming protocol, and a segmentation algorithm is used to automatically segment the ROI region to generate a sequence of images with timestamps. The specific steps of step S1 are as follows: S11. Access the video stream from the production line camera via the RTSP protocol and extract image frames at the set sampling rate; For example, the sampling rate can be set to a value from 1 to 10 frames per second; S12. Use the Mask R-CNN algorithm to automatically segment the ROI area of ​​the target business object and filter out background interference; S13. Unify the ROI area of ​​each frame image after segmentation to the target pixel size and perform grayscale normalization to complete the standardization process; Exemplarily, the target pixel size is set to 512×512 pixels; S14. Add a timestamp, camera number, and production line station information to each frame as metadata, and generate a PNG format image sequence. It should be noted that the video stream from the production line camera is accessed via the RTSP protocol, and image frames are extracted at a set sampling rate to ensure efficient and real-time data acquisition. The Mask R-CNN algorithm is used to automatically segment the ROI area and filter background interference, improving the accuracy and efficiency of image processing. The segmented images are standardized to unify pixel size and grayscale normalization, improving the consistency of image data. S2. Use a deep learning model to extract features of a set dimension from the image sequence, use an index tool library to save the extracted features and establish a hierarchical index structure, and then perform preliminary identification of abnormal images based on the extracted features and a pre-set dynamic distance threshold. The specific steps of step S2 are as follows: S21. Use the EfficientNet-B7 convolutional neural network model to extract a feature vector of a specified dimension from the image sequence. S22. Construct a feature library using the vector index Faiss library to store the collected feature vectors and establish a hierarchical index structure; S23. A dynamic distance threshold is pre-set based on the industrial scenario. Before each extracted feature vector enters the feature library, the dynamic distance with the feature vectors already in the feature library is calculated. When the dynamic distance is greater than the dynamic distance threshold, the image corresponding to the feature vector is used as the initial abnormal image for identification. For example, the basic distance threshold is set to 150, and the basic distance threshold is compensated according to the industrial scenario. For example, in a dark environment, the dynamic distance threshold is adjusted to 180. When the Manhattan distance between the extracted feature vector and the existing feature vector in the feature library is greater than 180, the corresponding image is used as an abnormal image for preliminary recognition. It should be noted that the EfficientNet-B7 model is used to extract feature vectors of a set dimension from image sequences to ensure the accuracy and efficiency of feature extraction. The Faiss library is used to build a feature library and establish a hierarchical index structure to improve the efficiency of feature storage and retrieval. A dynamic distance threshold is used to perform preliminary identification of abnormal images on the extracted features, improving adaptability to industrial scenarios and environments. The specific steps of step S23 are as follows: S231. Determine the basic distance threshold, compensate the basic distance threshold according to the industrial scenario, and dynamically adjust it according to the set method during the detection process; S232. Compare the currently extracted feature vector with the feature library and determine whether the feature library is empty; If so, the current feature vector is used as the first seed of the feature library, saved in the hierarchical index structure, and a single-node graph is created, ending; If not, proceed to step S233; S233. Search the current feature vector from the top layer of the hierarchical index structure to locate the k candidate feature vectors with the smallest Manhattan distance to the current feature vector in the bottom layer; For example, for the current feature vector Ft, starting from the top layer L0, the top 10 nodes are found from the 100 representative nodes, the top 30 nodes are found from the neighborhood of the top 10 nodes in the L1 layer, and the top 50 feature vectors are found from the neighborhood of the top 30 nodes in the bottom layer L2. The top 50 feature vectors with the smallest Manhattan distance between the current feature vector and the feature vector in the L2 layer are calculated, as well as 50 candidate feature vectors. S234. Take the minimum distance from the Manhattan distances between the k candidate feature vectors and the current feature vector, and determine whether the minimum distance is greater than the dynamic distance threshold; If yes, go to step S235; If not, mark the current feature vector as a normal feature vector and add it to the feature library, and update the node graph of the hierarchical index structure, and end; S235. The current feature vector corresponding to the picture is used as the initial identification of abnormal image; For example, the basic threshold θ0 is set to 150; Perform linear compensation based on the ambient light intensity of the industrial appearance inspection scene: θ=θ0×(1+0.05×(L-100) / 100); Among them, θ is the dynamic distance threshold, L is the illumination value; When the minimum distance D is detected 10 times in a row min ∈[θ-20,θ], automatically reduce θ by 5%; For example, in the appearance detection scenario, the initial distance threshold is set to 150±20% floating, and in the motion state scenario, the dynamic distance threshold is linearly adjusted with the device speed; It should be noted that the basic distance threshold is compensated according to the industrial scenario and dynamically adjusted during the detection process, which improves the adaptability of anomaly recognition. The hierarchical index structure and dynamic update of the node graph improve the efficiency of feature library construction and maintenance. By accurately marking abnormal images, a data foundation is provided for multimodal large model arbitration. S3. Generate prompt words based on the detection requirements of the industrial scenario, and determine abnormality-related information for the initially identified abnormal image based on the industrial scenario. Then, input the prompt words and abnormality-related information into the multimodal large model, output the determined abnormal image, and set hard and easy sample labels for the determined abnormal image. The specific steps of step S3 are as follows: S31. Generate prompt word templates with variable parameters for each industrial scenario; For example, in a parts inspection scenario in a machinery manufacturing plant, a prompt word template is generated: "Judge whether the parts in the image have {defecttype} defects, such as {example}, with a tolerance range of {tolerance}; S32 based on the detection needs of industrial scenarios, determine the value of the variable parameter in the prompt word template, generate a prompt word; For example, the variable parameter values ​​are determined as follows: {defecttype} is "scratch", {example} is "surface scratch length exceeds 2mm", and {tolerance} is "±0.5mm". The prompt word is generated: "Determine whether the component in the image has a scratch defect, such as if the surface scratch length exceeds 2mm, and the tolerance range is ±0.5mm." S33 determines the feature vector, size information and metadata of the initially identified abnormal image as abnormal related information; S34. Integrate the prompt word and abnormality-related information into JSON format, and input it into the visual model together with the initially identified abnormal image, output the determined abnormal image, and determine the abnormality type and confidence level; For example, after inputting the visual large model, the determined abnormal image is output, and the abnormality type is determined to be "scratch" with a confidence level of 0.95; S35. Setting labels for difficult samples, test samples, and easy samples for the determined abnormal images according to the confidence value of the abnormal type; For example, according to the confidence level of 0.95, the abnormal image is set as an easy example and marked as a "scratch" defect; It should be noted that by generating prompt word templates, integrating anomaly information input, and setting sample labels based on confidence, the multimodal large model can accurately understand industrial anomaly detection needs, output more reliable anomaly judgment results, improve the accuracy of anomaly type identification, and enhance the sample quality of model training. S4. Use the identified abnormal images to construct a dataset to train the object detection model. During the training process, adjust the learning rate and the ratio of hard examples to easy examples. The specific steps of step S4 are as follows: S41. Set an initial sampling ratio of 1:n for hard examples to easy examples, where n>1, and construct a dataset by sampling from the identified abnormal images according to the initial sampling ratio until the number of samples reaches a set threshold. Exemplarily, the initial sampling ratio of hard samples to easy samples is set to 1:3; S42. Divide the dataset into a training set and a validation set according to a set ratio; S43. Use the PyTorch framework to build the YOLOv9 model as the target detection model, initialize the learning rate, use the deviation between the predicted anomaly type and the real-time anomaly type as the loss function, train the target detection model using the training set, and use the Adagrad optimizer to adjust the target detection model parameters according to the learning rate based on the loss function during training. For example, the initial learning rate is 0.001; S44. Perform validation using the validation set every N iterations during the training process, and calculate the loss value calculated in the target detection model using the validation set as input, as well as the recognition accuracy of the difficult examples. S45. When the loss value does not decrease in N iterations, reduce the learning rate according to the set method, and when the recognition accuracy of the difficult example is less than the set ratio threshold, increase the sampling ratio of the difficult example to the easy example; For example, validation is performed every 5 iterations. If the loss value of the validation set does not decrease in 5 iterations, the learning rate is reduced to the original 0.5, that is, 0.005; if the recognition accuracy of the difficult sample is less than the set ratio threshold of 70% in 5 iterations, the sampling ratio of the difficult sample to the easy sample is increased to 2:3; S46. When the loss value is less than the set threshold or the number of iterations exceeds the maximum value, the training is stopped and the final target detection model is obtained; For example, if the validation loss value fluctuates by less than 1% for three consecutive times or reaches the maximum number of iterations of 100, the training is stopped and the final object detection model is output; It should be noted that by setting the sample sampling ratio, dynamically adjusting the learning rate and sample ratio, and optimizing the model parameters with the loss function, the model convergence speed is accelerated compared to fixed parameter training, the recognition accuracy of difficult samples is improved, and the generalization ability of the model is improved; S5. The images captured in the industrial scene are sequentially input into the deep learning model, the multimodal large model, and the trained target detection model to obtain the recognition results of whether they are abnormal images and the type of abnormality. The images and recognition results are then fed back to the user for annotation confirmation and added to the dataset to optimize the training of the target detection model. The specific steps of step S5 are as follows: The specific steps of step S5 are as follows: S51. Capture the production line in the industrial scene and obtain real-time images; S52. If the real-time image is sequentially input into the deep learning model, the multimodal large model and the trained target detection model and identified as an abnormal screen, the abnormal screen is digitally signed and uploaded to the user management terminal and stored in the Mysql database; S53 receives the user's abnormal screen classification mark or correct the misjudgment on the user management terminal, and records the abnormal screen label; S54. Regularly add abnormal images and corresponding labels to the data set, and return to step S4 to train the target detection model until the accuracy and recall rate of the target detection model meet user needs; It should be noted that by uploading abnormal images and providing user annotation feedback, regularly updating the data set and retraining the model, a closed loop of abnormality recognition and model optimization is built, so that the target detection model can continuously adapt to changes in industrial scenarios, improve the accuracy of abnormality recognition under long-term end operation, avoid model performance degradation, and ensure the long-term and reliable implementation of industrial production detection.

[0027] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0028] like Figure 2 As shown, the following is an embodiment of the industrial visual anomaly recognition system based on multimodal feedback provided by an embodiment of the present disclosure. This system and the industrial visual anomaly recognition method based on multimodal feedback of the above-mentioned embodiments belong to the same inventive concept. For details not fully described in the embodiment of the industrial visual anomaly recognition system based on multimodal feedback, please refer to the embodiment of the above-mentioned industrial visual anomaly recognition method based on multimodal feedback.

[0029] The system includes: The industrial dynamic acquisition module is used to collect production line video streams in real time through streaming protocols, and automatically segment the ROI area using a segmentation algorithm to generate a sequence of images with timestamps; The feature comparison module is used to extract features of a set dimension from an image sequence using a deep learning model, save the extracted features using an index tool library and establish a hierarchical index structure, and then perform preliminary identification of abnormal images based on the extracted features combined with a pre-set dynamic distance threshold; The large-scale model arbitration module is used to generate prompt words based on the detection requirements of industrial scenarios, and determine the abnormality-related information of the initially identified abnormal images based on the industrial scenarios. The prompt words and abnormality-related information are then input into the multimodal large-scale model, and the confirmed abnormal images are output. The hard and easy sample labels are set for the confirmed abnormal images. An adaptive training module is used to train the target detection model using a dataset constructed using identified abnormal images, and to adjust the learning rate and the ratio of hard to easy examples during training. The anomaly recognition and model optimization module is used to input images captured in industrial scenarios into the deep learning model, the multimodal large model, and the trained target detection model in sequence to obtain the recognition results of whether the image is abnormal and the type of anomaly. The images and recognition results are then fed back to the user for annotation and confirmation, and added to the dataset to optimize the training of the target detection model.

[0030] This embodiment achieves efficient and accurate recognition of abnormal images through the interactive collaboration of the industrial dynamic acquisition module, feature comparison module, large model arbitration module, adaptive training module, abnormality recognition and model optimization module.

[0031] The industrial visual anomaly recognition method based on multimodal feedback provided in the embodiment of the present application can be applied to electronic devices. Those skilled in the art will understand that the electronic device structure involved in the embodiment of the present invention does not constitute a limitation on the electronic device, and the electronic device may include more or fewer components than shown, or combine certain components, or arrange components differently. In the embodiment of the present invention, the electronic device includes but is not limited to a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of the present application described and / or required herein.

[0032] The electronic device may include a processor, an external memory interface, an internal memory, a universal serial bus (USB) interface, a charging management module, a power management module, a battery, a wireless communication module, an audio module, a speaker, a microphone, a sensor module, a button, a camera, a display, and a SIM card interface, etc.

[0033] It is understood that the structures illustrated in the embodiments of the present application do not constitute specific limitations on the electronic device. In other embodiments of the present application, the electronic device may include more or fewer components than shown, or combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0034] A processor may include one or more processing units, such as a central processing unit (CPU), an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors.

[0035] The processor can be the nerve center and command center of the electronic device. The controller can generate operation control signals based on the instruction opcode and timing signal to complete the control of instruction fetching and execution.

[0036] The processor may also include a memory for storing instructions and data. In some embodiments, the memory in the processor is a cache memory. This memory can store instructions or data that the processor has just used or is reusing. If the processor needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces processor latency, and thus improves system efficiency.

[0037] The above-mentioned electronic device realizes the industrial visual anomaly recognition method based on multimodal feedback of the present application, which collects the production line video stream in real time through the streaming protocol, and uses the segmentation algorithm to automatically segment the ROI area to generate a picture sequence with a timestamp; uses a deep learning model to extract features of a set dimension from the picture sequence, uses an index tool library to save the extracted features and establish a hierarchical index structure, and then combines the extracted features with a pre-set dynamic distance threshold to perform preliminary recognition of abnormal images; generates prompt words in combination with the detection needs of industrial scenes, and determines the abnormality-related information of the initially identified abnormal image according to the industrial scene, and then inputs the prompt words and the abnormality-related information into the multimodal large model, outputs the determined abnormal image, and provides the determined abnormal image with the output. Normal images are labeled with difficult and easy samples; determined abnormal images are used to construct a data set to train the target detection model, and the learning rate and the ratio of difficult and easy samples are adjusted during the training process; the pictures captured in the industrial scene are input into the deep learning model, the multimodal large model and the trained target detection model in sequence to obtain the recognition results of whether they are abnormal images and the type of abnormality, and then the pictures and recognition results are fed back to the user for annotation confirmation and added to the data set, and the technical solution for optimizing the training of the target detection model is achieved through the steps of dynamic collection, feature extraction, multimodal large model arbitration, adaptive training and abnormality recognition, so as to achieve the beneficial effect of efficient and accurate recognition of abnormal images in industrial production.

[0038] The storage medium provided in the present application stores a program product that can implement an industrial visual anomaly recognition method based on multimodal feedback.

[0039] The industrial visual anomaly recognition method based on multimodal feedback includes: real-time acquisition of production line video streams via a streaming protocol, automatic segmentation of the ROI region using a segmentation algorithm, and generation of a time-stamped image sequence; using a deep learning model to extract features of a set dimension from the image sequence, using an index tool library to save the extracted features and establish a hierarchical index structure, and then performing preliminary identification of abnormal images based on the extracted features and a pre-set dynamic distance threshold; generating prompt words based on the detection requirements of the industrial scene, and determining abnormality-related information for the initially identified abnormal images based on the industrial scene; then inputting the prompt words and abnormality-related information into a multimodal large model, outputting the determined abnormal images, and setting hard and easy sample labels for the determined abnormal images; using the determined abnormal images to construct a dataset to train the target detection model, and adjusting the learning rate and the ratio of hard and easy samples during the training process; inputting images captured in the industrial scene into the deep learning model, the multimodal large model, and the trained target detection model in sequence to obtain identification results of whether the image is abnormal and the type of anomaly; then feeding the images and recognition results back to the user for annotation confirmation and adding them to the dataset to optimize the training of the target detection model.

[0040] In some possible embodiments, the industrial visual anomaly recognition method based on multimodal feedback disclosed herein can be implemented in the form of a program product, which includes program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps of various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.

[0041] The storage medium of the present disclosure can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, a system, device or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0042] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for industrial visual anomaly recognition based on multimodal feedback, characterized in that: The steps include: S1. Real-time video streams from the production line are collected through a streaming protocol, and a segmentation algorithm is used to automatically segment the ROI region to generate a time-stamped image sequence. S2. Use a deep learning model to extract features of a specified dimension from the image sequence. Use an indexing tool library to store the extracted features and establish a hierarchical index structure. Then, combine the extracted features with a pre-set dynamic distance threshold to perform preliminary identification of abnormal images. S3. Generate prompt words based on the detection requirements of the industrial scenario, and determine abnormality-related information for the initially identified abnormal image based on the industrial scenario. Then, input the prompt words and abnormality-related information into the multimodal large model, output the determined abnormal image, and set hard and easy sample labels for the determined abnormal image; S4. Use the identified abnormal images to construct a dataset to train the object detection model, and adjust the learning rate and the ratio of hard to easy examples during training; S5. Input the images captured in industrial scenes into the deep learning model, the multimodal large model, and the trained target detection model in sequence to obtain the recognition results of whether they are abnormal images and the type of abnormality. The images and recognition results are then fed back to the user for annotation and confirmation, and added to the dataset to optimize the training of the target detection model.

2. The industrial visual anomaly recognition method based on multimodal feedback according to claim 1 is characterized in that: The specific steps of step S1 are as follows: S11. Access the video stream from the production line camera via the RTSP protocol and extract image frames at the set sampling rate; S12. Use the Mask R-CNN algorithm to automatically segment the ROI area of ​​the target business object and filter out background interference; S13. Unify the ROI area of ​​each frame image after segmentation to the target pixel size and perform grayscale normalization to complete the standardization process; S14. Add a timestamp, camera number, and production line station information as metadata to each frame of the image, and generate an image sequence in PNG format.

3. The industrial visual anomaly recognition method based on multimodal feedback according to claim 2 is characterized in that: The specific steps of step S2 are as follows: S21. Use the EfficientNet-B7 convolutional neural network model to extract a feature vector of a specified dimension from the image sequence. S22. Construct a feature library using the vector index Faiss library to store the collected normal feature vectors and establish a hierarchical index structure; S23. Pre-set a dynamic distance threshold based on the industrial scenario, calculate the dynamic distance between each extracted feature vector and the existing feature vectors in the feature library, and when the dynamic distance is greater than the dynamic distance threshold, use the image to which the corresponding feature vector belongs as the initial identified abnormal image.

4. The industrial visual anomaly recognition method based on multimodal feedback according to claim 3 is characterized in that: The specific steps of step S23 are as follows: S231. Determine the basic distance threshold, compensate the basic distance threshold according to the industrial scenario, and dynamically adjust it according to the set method during the detection process; S232. Compare the currently extracted feature vector with the feature library and determine whether the feature library is empty; If so, the current feature vector is used as the first seed of the feature library, saved in the hierarchical index structure, and a single-node graph is created, ending; If not, proceed to step S233; S233. Search the current feature vector from the top layer of the hierarchical index structure to locate the k candidate feature vectors with the smallest Manhattan distance to the current feature vector in the bottom layer; S234. Take the minimum distance from the Manhattan distances between the k candidate feature vectors and the current feature vector, and determine whether the minimum distance is greater than the dynamic distance threshold; If yes, go to step S235; If not, mark the current feature vector as a normal feature vector and add it to the feature library, and update the node graph of the hierarchical index structure, and end; S235. Use the image corresponding to the current feature vector as the abnormal image for preliminary identification.

5. The industrial visual anomaly recognition method based on multimodal feedback according to claim 3 is characterized in that: The specific steps of step S3 are as follows: S31. Generate prompt word templates with variable parameters for each industrial scenario; S32 based on the detection needs of industrial scenarios, determine the value of the variable parameter in the prompt word template, generate a prompt word; S33 determines the feature vector, size information and metadata of the initially identified abnormal image as abnormal related information; S34. Integrate the prompt word and abnormality-related information into JSON format, and input it into the visual model together with the initially identified abnormal image, output the determined abnormal image, and determine the abnormality type and confidence level; S35. Set labels of difficult samples, samples to be tested, and easy samples for the determined abnormal image according to the confidence value of the abnormal type.

6. The method for industrial visual anomaly recognition based on multimodal feedback according to claim 5, characterized in that: The specific steps of step S4 are as follows: S41. Set an initial sampling ratio of 1:n for hard examples to easy examples, where n>1, and construct a dataset by sampling from the identified abnormal images according to the initial sampling ratio until the number of samples reaches a set threshold. S42. Divide the dataset into a training set and a validation set according to a set ratio; S43. Use the PyTorch framework to build the YOLOv9 model as the target detection model, initialize the learning rate, use the deviation between the predicted anomaly type and the real-time anomaly type as the loss function, train the target detection model using the training set, and use the Adagrad optimizer to adjust the target detection model parameters according to the learning rate based on the loss function during training. S44. Perform validation using the validation set every N iterations during the training process, and calculate the loss value calculated in the target detection model using the validation set as input, as well as the recognition accuracy of the difficult examples. S45. When the loss value does not decrease in N iterations, reduce the learning rate according to the set method, and when the recognition accuracy of the difficult example is less than the set ratio threshold, increase the sampling ratio of the difficult example to the easy example; S46. When the loss value is less than the set threshold or the number of iterations exceeds the maximum value, training is stopped to obtain the final target detection model.

7. The method for industrial visual anomaly recognition based on multimodal feedback according to claim 6, characterized in that: The specific steps of step S5 are as follows: S51. Capture the production line in the industrial scene and obtain real-time images; S52. If the real-time image is sequentially input into the deep learning model, the multimodal large model and the trained target detection model and identified as an abnormal screen, the abnormal screen is digitally signed and uploaded to the user management terminal and stored in the Mysql database; S53 receives the user's abnormal screen classification mark or correct the misjudgment on the user management terminal, and records the abnormal screen label; S54. Regularly add abnormal images and corresponding labels to the data set, and return to step S4 to train the target detection model until the accuracy and recall rate of the target detection model meet user requirements.

8. An industrial visual anomaly recognition system based on multimodal feedback, characterized in that: include: The industrial dynamic acquisition module is used to collect production line video streams in real time through streaming protocols, and automatically segment the ROI area using a segmentation algorithm to generate a sequence of images with timestamps; The feature comparison module is used to extract features of a set dimension from an image sequence using a deep learning model, save the extracted features using an index tool library and establish a hierarchical index structure, and then perform preliminary identification of abnormal images based on the extracted features combined with a pre-set dynamic distance threshold; The large-scale model arbitration module is used to generate prompt words based on the detection requirements of industrial scenarios, and determine the abnormality-related information of the initially identified abnormal images based on the industrial scenarios. The prompt words and abnormality-related information are then input into the multimodal large-scale model, and the confirmed abnormal images are output. The hard and easy sample labels are set for the confirmed abnormal images. An adaptive training module is used to train the target detection model using a dataset constructed using identified abnormal images, and to adjust the learning rate and the ratio of hard to easy examples during training. The anomaly recognition and model optimization module is used to input images captured in industrial scenarios into the deep learning model, the multimodal large model, and the trained target detection model in sequence to obtain the recognition results of whether the image is abnormal and the type of anomaly. The images and recognition results are then fed back to the user for annotation and confirmation, and added to the dataset to optimize the training of the target detection model.

9. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the method for industrial visual anomaly recognition based on multimodal feedback as claimed in any one of claims 1 to 7 are implemented.

10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the industrial visual anomaly recognition method based on multimodal feedback are implemented as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Industrial anomaly detection method based on multi-agent prompt learning

    CN120852894A

  • Appearance quality detection method and system for ALC plate

    CN121027116A

  • Visual classification operation system and electronic equipment

    CN121170548A

  • Sample data centralized optimization processing method applied to image recognition

    CN121582701A

  • Abnormal behavior detection method and device, equipment and storage medium

    CN121582989A