Abnormality detection method, system, device and medium based on multi-modal visual large model
By employing a multimodal visual large model anomaly detection method, which utilizes image registration and large model recognition technology, the problems of environmental change and high cost in existing technologies are solved, achieving accurate detection of unknown types of anomalies and improving the accuracy and robustness of detection.
Patent Information
- Application Number
- CN202511144198.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-08-15
AI Technical Summary
Existing technologies are easily affected by changes in light and environment in the detection of abnormal objects, have poor generalization ability, and are expensive to detect based on 3D point clouds, making it difficult to adapt to complex scenes and the detection of unknown types of anomalies.
An anomaly detection method based on a multimodal visual large model is adopted. Images are acquired through inspection equipment and the template image and the image to be inspected are registered. An open set target detection large model is used to identify abnormal parts. Combined with a segmentation large model, image segmentation and similarity comparison are performed to eliminate false anomalies. Siamese network is used to calculate similarity to identify anomalies of unknown types.
It achieves accurate detection of unknown anomalies in complex scenarios, reduces environmental interference, improves detection accuracy and robustness, and avoids the limitations and high costs of traditional methods.
Smart Images

Figure CN120726398B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of foreign matter anomaly detection, and in particular to an anomaly detection method, system and device based on a multi-modal vision large model and a storage medium. BACKGROUND
[0002] Anomaly foreign matter detection is diverse and has multiple scenarios, and there is an urgent need in multiple fields. For example, in the field of rail transit, there is often a need for anomaly detection in scenarios such as train inspection, tunnel inspection, and wiring network inspection. In the field of power inspection, there is a need to detect bird nests, floating objects, and other foreign matter on power substation equipment and lines.
[0003] In the field of anomaly foreign matter detection, existing technologies mainly cover two directions: 2D image detection and 3D point cloud detection. 2D image detection includes traditional image processing-based methods and deep learning model-based methods. Deep learning model-based technology is further divided into CNN-based deep learning models and transformer architecture-based large models, which have gradually become the main technical support for current anomaly foreign matter detection.
[0004] In practical applications, 2D image detection-based methods are more common. For example, CN113673614B uses an image comparison-based method to obtain a difference image, and then uses image processing methods to obtain a foreign matter area. CN113989209B uses Faster R-CNN for target detection to detect known types of foreign matter. CN116342944A also uses target detection to divide targets into normal and foreign matter categories, and simulates various foreign matter images by means of field collection and PS to belong to the foreign matter category. 3D point cloud detection-based methods, such as CN108107444A, use a 3D camera to collect point cloud data, match the test point cloud with the model point cloud ICP, calculate the distance between the point pairs to screen out points beyond the distance, and then calculate the sparsity of the set of points beyond the distance to determine whether it is a foreign matter.
[0005] In addition, some existing patent technologies combine large model technology. For example, CN118982706A is based on SegmentAnything and open set classification, and CN117456477A supports multi-frame image input by modifying the network structure.
[0006] However, in actual use, the prior art has two difficulties in the anomaly detection process: on the one hand, methods based on 2D image detection, such as CN113673614B, are susceptible to light and environmental changes, leading to false detection; and methods based on target detection, such as CN113989209B and CN116342944A, still essentially rely on the recording of known or trained object features, and have poor generalization for untrained anomalies or foreign objects, and cannot solve the problem of numerous anomaly types that cannot be listed in advance; on the other hand, methods based on 3D point cloud detection, such as CN108107444A, require the use of 3D cameras to collect point cloud data, which is costly and complex to calculate, and has certain limitations in practical applications.
[0007] In addition, the detection processes and model architectures used by different patents are different, and the schemes of some patents differ in the focus and process of detecting anomalies, and have not formed a general detection scheme that can adapt to complex scene changes and does not require prior knowledge of anomaly or foreign object types.
[0008] Therefore, the present application proposes an anomaly detection method and system based on a multi-modal visual large model to solve the above technical problems. SUMMARY
[0009] The main purpose of the present application is to provide an anomaly detection method and system based on a multi-modal visual large model, which can identify unknown types of anomalies to solve the technical problems presented in the background art.
[0010] The present application solves the above technical problems by adopting the following technical solutions:
[0011] An anomaly detection method based on a multi-modal visual large model, comprising:
[0012] The inspection device moves to a specified point, and a specified magnification is used to capture images of the specified shooting scene at a preset shooting distance and angle;
[0013] The image in the normal state obtained by the inspection device during the initial inspection is used as a template image, and the image collected in real time during the subsequent inspection of the inspection device is used as a detection image, the inspection image and the template image are compared to perform an inspection image anomaly judgment operation;
[0014] The inspection image anomaly judgment operation determines whether there is an anomaly and the position of the anomaly in the image by the following steps:
[0015] S1. Register the template image and the image to be detected;
[0016] S2. Detecting targets in images without scene samples by using an open set target detection large model to identify various abnormal components;
[0017] S3. Using a segmentation large model to perform image segmentation within the detection box obtained by the target large model;
[0018] S4. For cases where the type of foreign matter needs to be identified, compare the contours generated by the segmentation of the template image and the image to be detected, eliminate part of the differences and analyze region by region, and set the contours with an IOU less than a specified threshold as possible abnormal places;
[0019] S5. Similarity comparison of the difference contours using a twin network calculation, eliminating contours with a similarity greater than a preset threshold, and the remaining being the abnormal area.
[0020] Preferably, the inspection equipment uses a navigation algorithm such as SLAM to navigate to a specified point and adjust the mechanical arm for shooting to a specified posture, to ensure that the camera mounted at the end of the mechanical arm takes pictures of the inspection points at a fixed angle in a fixed place.
[0021] Preferably, the matching algorithm such as lightglue is used in the S1 step to register the template image and the image to be detected, and the specific operations in the registration process include:
[0022] S11. Extracting image feature points using superpoint and generating feature point descriptors;
[0023] S12. Also using lightglue to optimally match the feature points;
[0024] S13. And according to the matching feature point data, using the RANSAC algorithm to calculate the homography transformation matrix of the template image and the image to be detected, for registering the two images.
[0025] Preferably, the specific operation process of the RANSAC algorithm includes:
[0026] S131. Based on the feature point descriptors, determine the key points in the matching feature point data, and randomly select four non-collinear samples to calculate the homography transformation matrix;
[0027] S132. Test all data through the homography transformation matrix, calculate the re-projection coordinates of the feature points in the first frame image in the second frame image according to the homography matrix, compare the distance between the re-projection coordinates and the matched feature point coordinates, if the distance is less than a specified threshold, it is considered as a correct matching point pair, otherwise it is considered as an incorrect matching, and the number of correct point pairs is recorded;
[0028] S133. Repeat the steps S131-S132 until the specified round of the cycle, compare the number of correctly matched point pairs after the cycle, take the case with the most number of correctly matched point pairs as the final result, eliminate the incorrect matches, and output the correctly matched pairs to achieve the screening of the feature point matching.
[0029] Preferably, the open set target detection large model in the S2 step has strong feature extraction capability and generalization ability due to its training on a large parameter-specific data set, and has the ability to recognize general targets. Therefore, using the open set target detection large model to detect targets in the image can recognize various abnormal components without the need for scene samples. The open set target detection large model adopts an open set target detection model based on a transfomer network structure, which includes:
[0030] a text backbone network and an image feature extraction backbone network, the text backbone network adopts BERT to extract text features, and the image feature extraction network adopts a swin Transformer network to extract multi-scale image features;
[0031] a text-image feature fusion module for inputting image features and text features into the fusion module for cross-modal feature fusion;
[0032] a language-guided query selection module for selecting fusion features related to the input text as the decoder query;
[0033] a multi-modal decoder for sending the cross-modal query to a self-attention layer to combine an image cross-attention layer for obtaining image features, a text cross-attention layer for text features, and an FFN layer in each cross-modal decoder layer;
[0034] an output layer for outputting the updated query and the reference point as the final box and class prediction.
[0035] Preferably, the specific operation process of the open set target detection large model in the S2 step for identifying abnormal components includes:
[0036] S21. Input the image into the feature extraction module to extract image features of different sizes from the input image to generate an image feature vector;
[0037] S22. Input the text into the text backbone network to extract text features to generate a text feature vector;
[0038] S23. Input the image features and the text features into the text-image feature fusion module to fuse the features, and use a cross-attention module to align the text and image features;
[0039] S24. The aligned image and text features are output to the language-guided query module, the similarity of the features is calculated, sorted, and the top n features are found as the query vector;
[0040] S25. The generated query is sent to the cross-modal decoder module, which sequentially performs cross-modal attention calculation with image features and text features to obtain the final decode output, and outputs the category text and coordinate frame position.
[0041] Preferably, the type and rectangular frame generated by the open set target detection large model in the S3 step are input into the segmentation large model to generate a segmentation mask, wherein the segmentation large model comprises the following modules:
[0042] An image encoder module for encoding an image and mapping the image to a feature space;
[0043] A prompt encoder module for position encoding and learning embedding of the generated prompts including points, frames, and text, and converting the prompt information into a feature vector;
[0044] A mask decoder module for integrating the feature vectors output by the image encoder module and the prompt encoder module, and decoding the final segmentation mask.
[0045] Preferably, the specific operation process of the segmentation large model for decoding to obtain the segmentation mask in the S3 step comprises:
[0046] S31. Map the image to the feature space and extract the image feature embedding vector;
[0047] S32. Positionally encode the coordinate frame information or mask information into a position description embedding vector;
[0048] S33. Input the image embedding vector and the position description embedding vector into the Transformer decoder, pass through multiple layers of self-attention and cross-attention modules, output the contour mask, and then pass through upsampling to obtain the mask prediction of the original image size;
[0049] S34. And predict the quality score of each mask through the MLP to evaluate the confidence of the mask.
[0050] Preferably, the specific operation process of the IOU comparison judgment in the S4 step comprises:
[0051] Comparing the contours generated by the segmentation of the template image and the image to be detected, and removing the specified difference part according to the IOU comparison;
[0052] Performing regional analysis on the segmented regions of the image and performing IOU judgment;
[0053] The calculation formula of IOU at this time is:
[0054]
[0055] Among them, is the area of the overlapping area of the two mask contours. is the total area of the combined area of the two mask contours;
[0056] IOU judgment is performed:
[0057] If the IOU is greater than a certain threshold, it is considered that the contour is the same part of the two images;
[0058] If the IOU is less than a certain threshold, it is considered that the contour is different between the two images, that is, it is a place where there may be an anomaly.
[0059] Preferably, the specific operation process of the S4 step comprises:
[0060] S41. Separate the template image from the contour mask on the inspection image one by one;
[0061] S42. Calculate the IOU of the contour mask on the inspection image and all contour masks on the template image one by one, calculate the highest IOU score and record it;
[0062] S43. Compare the highest IOU score with the threshold value, if it is higher than the threshold value, it is considered that the two contour masks have consistency, and the possibility of abnormality is excluded, if it is lower than the threshold value, it is considered that it is a contour that may exist an anomaly;
[0063] S44. Repeat the steps of S42-S43 for all contour masks on the inspection image to obtain the final list of contours that may exist an anomaly.
[0064] Preferably, the twin network in the S5 step adopts a multi-layer convolution, activation, and pooling layer, and finally extracts a feature vector, and the twin network adopts an ESSIM structural similarity formula obtained by improving an SSIM structural similarity formula as a loss function, and the loss function formula is obtained by image brightness, contrast, and structural similarity. The similarity quantity is:
[0065]
[0066] Among them, are the original image features of the template image and the image to be detected, respectively, indicates the feature and the ESSIM structural similarity evaluation result of the feature , is the feature Image features of the template image after processing by the Sobel operator Features Image features of the image to be inspected after processing by the Sobel operator. Features Image features of the template image after processing by the Laplacian operator Features Image features of the image to be inspected after processing by the Laplacian operator. , To specify the scaling factor, This indicates the calculation of SSIM structural similarity;
[0067] The specific steps for similarity comparison at this point include:
[0068] S51. Obtain the bounding rectangle of the contour mask of the suspected anomaly and get its vertex coordinates;
[0069] S52. Extract sub-image block pairs from the inspection image and the template image based on the coordinates of the circumscribed rectangle;
[0070] S53. Input the sub-images into the Siamese network for calculation and output the similarity score;
[0071] S54. Based on the preset threshold, those with similarity scores less than the preset threshold are set as the final extracted anomalies.
[0072] Preferably, in step S5, the Siamese network calculation is replaced by directly using the SSIM structural similarity algorithm to calculate the similarity between the inspection image and the template image. This is used to measure the similarity by comparing the brightness, contrast, and structure of the two images. The calculation formula is as follows:
[0073] in, for The mean, for The mean, , It is a constant. for variance for variance for , Covariance:
[0074] The specific steps for similarity comparison at this point include:
[0075] S51. Obtain the bounding rectangle of the contour mask of the suspected anomaly and get its vertex coordinates;
[0076] S52. Extracting a sub-image block pair from the patrol image and the template image according to the circumscribed rectangular coordinates;
[0077] S53. Performing structural similarity judgment on the sub-image pair by using an SSIM structural similarity algorithm to obtain a similarity score;
[0078] S54. According to a preset threshold, setting the sub-image block pair with a similarity score higher than the preset threshold as a pseudo anomaly, and setting the sub-image block pair with a smaller similarity score as a final extracted anomaly.
[0079] On the other hand, the present application also discloses an anomaly detection system based on a multi-modal visual large model, which is used to execute any of the above-mentioned anomaly detection methods based on a multi-modal visual large model, and comprises:
[0080] The patrol device is used to move to a specified point to take a specified scene image at a preset shooting distance, a specified angle, and a specified magnification;
[0081] The device navigation module is used to navigate and control the patrol device to move to the specified point for patrol;
[0082] The image registration module is used to register the template image and the to-be-detected image;
[0083] The open set target detection large model is used to identify and detect abnormal components in the image, and correspondingly generate type data and a rectangular frame;
[0084] The segmentation large model is used to generate a segmentation mask based on the type data and the rectangular frame generated by the open set target detection large model, so as to perform image segmentation on the abnormal components in the detection frame;
[0085] The contour comparison judgment module is used to compare the contours generated by the segmentation of the template image and the to-be-detected image, eliminate part of the differences, and analyze region by region;
[0086] The similarity comparison module is used to compare the similarity of the difference contours to obtain an abnormal region.
[0087] In another aspect, the present application also discloses a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to make the processor execute the steps of the above-mentioned method.
[0088] In still another aspect, the present application also discloses a computer device, which comprises a memory and a processor, and the memory stores a computer program, and the computer program is executed by the processor to make the processor execute the steps of the above-mentioned method.
[0089] From the above technical solution, the present application provides an anomaly detection method and system based on a multi-modal visual large model. Compared with the prior art, the present application has the following advantages:
[0090] 1. The application can recognize targets without scene samples by using Grounding-dino as an open set target detection large model to perform target detection operations, which can utilize its feature extraction and generalization ability trained on super large data sets to identify targets and achieve detection of unknown type abnormalities.
[0091] 2. The application can eliminate slight changes in the viewing angle of images taken at different times, facilitate better comparison of template images and image segmentation results to be detected, further reduce environmental factors that interfere with detection, and improve detection accuracy by performing image registration based on Lightglue matching algorithm before image comparison.
[0092] 3. The application can accurately generate segmentation masks by using SAM segmentation large model in the segmentation link and performing image segmentation within the detection box obtained by the target large model, and provide accurate segmentation basis for subsequent abnormality judgment by using its processing and segmentation ability of prompt information.
[0093] 4. The application can eliminate some difference pseudo-abnormal contours and achieve accurate determination of abnormal areas to improve the reliability and accuracy of abnormality detection by using IOU comparison of template image and image to be detected segmentation contour and similarity comparison based on SSIM structural similarity algorithm to perform abnormality judgment operation.
[0094] 5. The application can fully utilize the ability of large models to recognize and segment multiple types of targets by combining large model technologies such as open set target detection large model and SAM segmentation large model, without the need for prior perception of specific abnormal types, and can further utilize the powerful feature extraction capability of deep learning network to avoid interference from light environment and other factors in traditional image comparison algorithms.
[0095] 6. The twin network uses the ESSIM structural similarity formula as the loss function, and processes the original image features by introducing Sobel and Laplacian operators, which can strengthen the twin network's ability to capture image edge contours and other structural features, effectively improve the overall detection method's ability to distinguish abnormal areas, reduce pseudo-abnormal interference, make the similarity calculation result more in line with actual detection needs, and enhance the model's robustness and accuracy in abnormal detection scenarios.
[0096] It should be understood that the content described in this part is not intended to identify key or important features of embodiments of the application, nor is it intended to limit the scope of the application. Other features of the application will become apparent from the following description. Of course, any product implementing the application does not necessarily need to achieve all the advantages described above. BRIEF DESCRIPTION OF DRAWINGS
[0097] The accompanying drawings, which form a part of this specification, are included to provide a further understanding of the application, and are incorporated by reference in their entirety. The embodiments depicted herein are provided by way of example and are not intended as limitations on the present application. In the drawings:
[0098] Figure 1 The overall operation flowchart of embodiment one of the present application;
[0099] Figure 2 The twin network structure schematic diagram of embodiment one of the present application;
[0100] Figure 3 The registered template image and inspection image comparison schematic diagram of embodiment two of the present application;
[0101] Figure 4 The groudingDINO detection object frame schematic diagram of embodiment two of the present application;
[0102] Figure 5 The template image and inspection image contour segmentation comparison schematic diagram of embodiment two of the present application;
[0103] Figure 6 The detection anomaly schematic diagram of embodiment two of the present application;
[0104] Figure 7 The registered template image and inspection image comparison schematic diagram of embodiment three of the present application;
[0105] Figure 8 The manual frame selected part to be detected schematic diagram of embodiment three of the present application;
[0106] Figure 9 The template image and inspection image contour segmentation comparison schematic diagram of embodiment three of the present application;
[0107] Figure 10 The detection anomaly schematic diagram of embodiment three of the present application;
[0108] Figure 11 The registered template image and inspection image comparison schematic diagram of embodiment four of the present application;
[0109] Figure 12 The manual frame selected part to be detected schematic diagram of embodiment four of the present application;
[0110] Figure 13 The template image and inspection image contour segmentation comparison schematic diagram of embodiment four of the present application;
[0111] Figure 14 The detection anomaly schematic diagram of embodiment four of the present application. DETAILED DESCRIPTION
[0112] With reference to the drawings of the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. In the case of no conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative effort are within the protection scope of the present application.
[0113] In the embodiments, refer to the drawings in the embodiments of the present application for a detailed description. Figures 1 to 14 .
[0114] In the first embodiment, as shown in Figure 1 The abnormality detection method based on the multi-mode visual large model provided in the embodiments of the present application specifically includes the following steps.
[0115] Step one: image acquisition
[0116] When the inspection robot collects images, a navigation algorithm needs to be used to navigate the robot to a fixed point for shooting and teach the mechanical arm, so as to ensure that the camera carried at the end of the mechanical arm shoots the inspection point image at the fixed point at a fixed angle.
[0117] (1) Through the SLAM algorithm, the robot can construct a map in the environment and navigate to the fixed point at each inspection;
[0118] (2) Teach the mechanical arm, set the appropriate distance and angle of the camera shooting, and ensure that the image is clear and the mechanical arm moves to the fixed angle to shoot the train component each time.
[0119] Through the above two steps of image acquisition, it is ensured that the collected images can be compared subsequently.
[0120] Step two: match the template image and the image to be inspected.
[0121] In order to eliminate the slight view angle change of the images shot at different time periods and better compare the segmentation results of the template image and the image to be inspected, the template image and the image to be inspected are matched before comparison. The lightglue matching algorithm is used to register the images.
[0122] The lightglue matching includes the following sub-steps.
[0123] (1) superpoint is used to extract image feature points and generate feature point descriptors.
[0124] (2) lightglue is used to optimally match the feature points.
[0125] (3) According to the matching key points, a RANSAC algorithm is used to calculate the homography transformation matrix of the two images, and the two images are registered. The specific use steps of the RANSAC algorithm in the application are:
[0126] (a) First, four non-collinear samples are randomly selected from the matched key points, and a homography transformation matrix is calculated through the four points;
[0127] (b) Then, all the data are tested by the homography transformation matrix, the feature points in the first frame image are projected onto the second frame image according to the homography matrix, and the distance between the re-projection coordinates and the matched feature point coordinates is compared. If it is less than a certain threshold, it is considered to be a correct matching point pair, otherwise it is considered to be an incorrect matching, and the number of correct matching point pairs is recorded.
[0128] (c) Repeat steps (a) and (b), compare the number of correct matching point pairs after multiple cycles, and take the case with the most correct matching point pairs as the final result, eliminate the incorrect matching, and output the correct matching pairs, thereby realizing the screening of the feature point matching.
[0129] Step three: using an open set target detection large model to detect the target in the image.
[0130] The large model has strong feature extraction ability and generalization ability due to its training on a large parameter-specific data set, and has the ability to recognize general targets. In the rail transit field, without the need for scene samples, each component of the train can be recognized. At this time, Grouding-dino is an open set target detection model based on the transformer network structure. It includes the following modules:
[0131] Text backbone network and image feature extraction backbone network. The text backbone network uses BERT to extract text features. The image feature extraction network uses a swin Transformer network to extract multi-scale image features.
[0132] Text-image feature fusion module: after extracting the image and text features as described above, the two features are input into the fusion module for cross-modal feature fusion.
[0133] Language-guided query selection module: select the fusion features more related to the input text as the decoder query.
[0134] Multi-modal decoder: the cross-modal query is sent to the self-attention layer, the image cross-attention layer for combining image features, the text cross-attention layer for combining text features, and the FFN layer in each cross-modal decoder layer. The updated query and reference point output by the last layer are the final bounding box and class prediction.
[0135] The specific operation process of further detecting the target in the image includes
[0136] (1) The image is input to a feature extraction module, different size image features are extracted from the input image, and an image feature vector is generated.
[0137] (2) The text is input to a text backbone network, text features are extracted, and a text feature vector is generated.
[0138] (3) The image features of step 1 and the text features of step 2 are input to a text-image feature fusion module for feature fusion. A cross-attention module is used to align the features of the text and the image.
[0139] (4) The aligned image and text features are output to a language-guided query module, the similarity of the features is calculated, sorted, and the top n features are selected as the query vector.
[0140] (5) The generated query is sent to a cross-modal decoder module, which sequentially performs cross-modal attention calculation with the image features and the text features to obtain the final decode output, and outputs the category text and the coordinate frame position.
[0141] Step four: using a segmentation large model to segment the detection image. Image segmentation is performed within the detection frame obtained by the target large model.
[0142] The sam segmentation large model includes three modules: image encoder, prompt encoder, and mask decoder. The image encoder encodes the image and maps it to the feature space. The prompt encoder encodes the generated prompt (point, frame, text) and learns the embedding, converting the prompt information into a feature vector. The mask decoder module integrates the feature vectors output by the image encoder and the prompt, and decodes the final segmentation mask.
[0143] The type and rectangular frame generated by the Grouding-dino in step three are input to the sam segmentation model to generate a segmentation mask, and the specific operation process includes:
[0144] (1) Map the image to the feature space and extract the image feature embedding vector.
[0145] (2) The coordinate frame information or mask information is position coded into a position description embedding vector.
[0146] (3) The image embedding vector and the position description embedding vector are input into a Transformer decoder, and after passing through multiple layers of self-attention and cross-attention modules, an outline mask is output, and after upsampling, a mask prediction of the original image size is obtained. And through the MLP, the quality score of each mask is predicted to evaluate the confidence of the mask.
[0147] Step five: identifying the type of foreign matter needs to be taken, and the non-foreign matter class skips this step. The contours generated by the segmentation of the template image and the image to be inspected are compared, and part of the difference is removed according to the IOU comparison. The segmented regions of the two images in step four are analyzed region by region, and the IOU is judged. If the IOU is greater than a certain threshold, it is considered that the contour is the same part of the two images. If the IOU is less than a certain threshold, it is considered that the contour is different between the two images, that is, it is a place where there may be an anomaly.
[0148] The calculation formula of IOU is:
[0149]
[0150] wherein, is the area of the overlapping region of the two mask contours. is the total area of the combined region of the two mask contours.
[0151] Therefore, the specific operation process of identifying the type of foreign matter at this time includes:
[0152] (1) Separate the contour masks on the template image and the inspection image one by one.
[0153] (2) Calculate the IOU of the contour mask on the inspection image and all contour masks on the template image one by one, calculate the highest IOU score and record it.
[0154] (3) Compare the highest IOU score with the threshold value. If it is higher than the threshold value, it is considered that the two contour masks have consistency and exclude the possibility of abnormality. If it is lower than the threshold value, it is considered to be a contour that may have an anomaly.
[0155] (4) Repeat steps (2) and (3) for all contour masks on the inspection image. Get the final list of contours that may have anomalies.
[0156] Step six: comparing the similarity of the difference contours in step five, and removing the contours with similarity greater than the threshold value. The remaining ones are the abnormal regions.
[0157] In a specific embodiment, since the effect of the similarity model based on the twin network is better than the algorithm based on image processing, the similarity comparison is calculated by using the twin network calculation, the twin network adopts multiple convolution, activation and pooling layers, and finally extracts the feature vector, two pictures enter the network with the same structure and the same parameters, and the respective feature vectors are extracted, and the extracted feature vectors enter the measurement layer, the measurement layer adopts a full convolution network, and finally outputs the similarity measurement result.
[0158] Further, the model of the network is as shown in Figure 2 The output of the network is compared with the real label to calculate the loss of the contrast loss function, the contrast loss function of the twin network adopts the ESSIM structural similarity formula obtained by improving the SSIM structural similarity formula as the loss function, and the brightness, contrast and structural similarity of the image are balanced.
[0159] At this time, it should be noted that the common loss function of the twin network is the following contrast loss:
[0160]
[0161] Among them, W represents the parameter, Y represents the ratio label of whether sample one and sample two match, Y=1 represents that the two samples are similar, and Y=0 represents that the two samples are not similar. X1 and X2 represent sample one and sample two for comparison respectively, P represents the feature bit number of the sample, m represents the set threshold constant, such as 1.5, and N represents the number of samples, represents the Euclidean distance of the two samples.
[0162] Therefore, compared with the common contrast loss of the twin network, the contrast loss function mainly considers the Euclidean distance of the two samples, while the SSIM loss function mainly considers the brightness, contrast and structural similarity of the sample. For image samples, the SSIM loss function is more in line with the characteristics of the image, so it has more advantages in measuring the similarity of two images as a loss function. At this time, in order to make SSIM better reflect the changes in the structure of the image, especially the changes in the edge features of the image, the sobel operator and the Laplacian operator are added to the original image feature vector for processing, and the processed features better reflect the edge features of the image.
[0163] The sobel operator x direction is , The final gradient is .
[0164] The Laplacian operator is as follows: .
[0165] Therefore, the loss function formula is:
[0166]
[0167] in, These are the original image features of the template image and the original image features of the image to be inspected, respectively. Representation of features With features The ESSIM structural similarity assessment results, Features Image features of the template image after processing by the Sobel operator Features Image features of the image to be inspected after processing by the Sobel operator. Features Image features of the template image after processing by the Laplacian operator Features Image features of the image to be inspected after processing by the Laplacian operator. , To specify the scaling factor, This indicates the calculation of SSIM structural similarity.
[0168] Compared to the original SSIM, this method places greater emphasis on the edge contour features of the image, which can lay a good foundation for finding image pairs containing foreign objects.
[0169] The specific steps at this point include:
[0170] S51. Obtain the bounding rectangle of the contour mask of the suspected anomaly and get its vertex coordinates;
[0171] S52. Extract sub-image block pairs from the inspection image and the template image based on the coordinates of the circumscribed rectangle;
[0172] S53. Input the sub-images into the Siamese network for calculation and output the similarity score;
[0173] S54. Based on the preset threshold, those with similarity scores less than the preset threshold are set as the final extracted anomalies.
[0174] In summary, the method adopts a twin network, and simultaneously adopts an ESSIM structural similarity formula as a loss function. By introducing Sobel and Laplacian operators to process the original image features, the twin network can enhance the ability to capture structural features such as image edge contours. Compared with the traditional contrast loss function that only considers the Euclidean distance of samples, ESSIM measures the similarity of images from three dimensions of brightness, contrast, and structure, which is more in line with the characteristics of images, especially sensitive to image edge changes. This is crucial for accurately identifying areas with significant profile differences in anomaly detection, and can effectively improve the model's ability to identify abnormal areas, reduce false anomalies, make the similarity calculation results more in line with actual detection needs, and enhance the robustness and accuracy of the model in anomaly detection scenarios.
[0175] In another specific embodiment, the similarity comparison adopts an SSIM structural similarity algorithm.
[0176] SSIM measures the similarity of two images by comparing their brightness, contrast, and structure. Its calculation formula is as follows:
[0177]
[0178] wherein, is the mean of , is the mean of , , is a constant. is the variance of , is the variance of , is , the covariance of
[0179] The specific operation steps at this time include:
[0180] (1) Calculate the circumscribed rectangle of the suspected abnormal contour mask to obtain its vertex coordinates.
[0181] (2) Extract the sub-image block pair from the inspection image and the template image according to the circumscribed rectangle coordinates.
[0182] (3) Perform structural similarity judgment on the sub-image pair to obtain the similarity score.
[0183] (4) Set a threshold, and the larger the similarity score, the more false anomalies. The smaller the similarity score, the final extracted anomaly.
[0184] Further, in a specific embodiment, it can be further illustrated that, since the image statistical features are usually unevenly distributed in space, the image distortion is also variable in space. The SSIM of the local image is better than the global, therefore, a sliding window is used to calculate the statistical weighted SSIM, first, the image is divided into blocks, and the image is divided into M rows and N columns of sub-images. Then using a sliding window of size K, for each block, the mean, variance and covariance of each sliding window are calculated, and then the Gaussian coefficient is weighted to obtain the structural similarity of the block. Finally, the average value of the structural similarity of all blocks is taken to obtain the final structural similarity
[0185] In another aspect, the present application also discloses an anomaly detection system based on a multi-modal visual large model, which is used to execute the anomaly detection method based on the multi-modal visual large model proposed in the above embodiment, comprising:
[0186] The inspection device is used to move to a specified point to take a specified scene image at a preset shooting distance, a specified angle and a specified magnification;
[0187] The device navigation module is used to navigate and control the inspection device to move to the specified point for inspection;
[0188] The image registration module is used to register the template image and the to-be-detected image;
[0189] The open-set target detection large model is used to identify and detect abnormal components in the image, and generate type data and a rectangular frame correspondingly;
[0190] The segmentation large model is used to generate a segmentation mask based on the type data and the rectangular frame generated by the open-set target detection large model, so as to perform image segmentation on the abnormal components in the detection frame;
[0191] The contour comparison judgment module is used to compare the contours generated by the segmentation of the template image and the to-be-detected image, eliminate part of the differences and analyze region by region;
[0192] The similarity comparison module is used to compare the similarity of the difference contours, and obtain the abnormal region.
[0193] In summary, the present application is based on the large model technology, fully utilizes the ability of the large model to identify and segment all types of targets, so that it is unnecessary to perceive specific abnormal types in advance. And the powerful feature extraction ability of the deep learning network is utilized to avoid the interference of light environment and the like in the traditional image comparison algorithm.
[0194] In embodiment two, the present application is based on the above embodiment one, and further proposes an anomaly detection method and system based on a multi-modal visual large model, which is used to detect an air compressor case. In the specific use process, the operation process is as follows:
[0195] (1) Use a fixed camera to shoot images of the air compressor box of the train car.
[0196] Collect point cloud data through laser radar, angular velocity and acceleration data collected through IMU sensor, correct motion distortion of the point cloud, match the current point cloud with the constructed local subgraph through ICP algorithm, and finally fuse the point cloud node data of continuous scanning into the subgraph to complete the construction of the map.
[0197] After the map is constructed, the robot is positioned and navigated to the fixed inspection point through the map.
[0198] A template image and an image to be inspected are obtained at the same position, and image registration is performed. After registering the two images, as shown in Figure 3 , the left image is the template image collected for the first time, and the right image is the image collected during inspection. Through registration, the components are aligned.
[0199] (2) Use Grounding DINO open set target detection large model to detect the template image and the inspection image to obtain the target type and the detection box.
[0200] As shown in Figure 4 , the large model detects all identifiable components in the image, frames them and labels them. In this embodiment, the train air compressor mesh cover component is extracted and framed, and its label is equipment.
[0201] (3) Use a segmentation large model to segment the image.
[0202] Expand the rectangular frame coordinates of the air compressor mesh cover identified in step (2) by 20% to obtain new rectangular frame coordinates. Input the rectangular frame coordinates as a prompt into the SAM segmentation large model to obtain the segmentation result. As shown in Figure 5 , the left image is the segmentation effect of the template image, and the right image is the segmentation effect of the inspection image.
[0203] (4) Extract the mask part of each contour of the segmentation of the two images to obtain the mask image of each contour.
[0204] (5) Compare the IOU of each contour of the two images. If the IOU is greater than a certain threshold, it is considered to be the same part of the two images, and if the IOU is less than the threshold, it is considered to be a different part of the two images. The different part is a suspected abnormal place.
[0205] Pairwise analysis of the mask in step (4) is performed. After registering the inspection image with the mask, if the IOU is greater than a certain ratio (0.5), it is considered to be the same part of the two images and non-abnormal. If the IOU is less than the ratio, it is considered to be a suspected abnormality.
[0206] (6) The suspected anomaly obtained in step (5) is analyzed by geometric comparison (aspect ratio, etc.), similarity comparison and the like, and the following structural similarity comparison is adopted, and the corresponding similarity score is calculated, wherein the score less than a certain ratio is an anomaly.
[0207] After step (5) of IOU matching filtering, a plurality of suspected anomalies are obtained, and the similarity of a plurality of (11 groups) sub-images is judged, and the similarity matching score is calculated. First, the image is divided into 3*3 blocks, and the sliding window size is 1 / 4 of the image, and the final similarity score is calculated by sliding, which is 0.38, 0.65, 0.47, 0.34, 0.61, 0.42, 0.71, 0.35, 0.69, 0.31, and 0.85. Set 0.5 as the threshold. Then
[0208] The 0th, 2nd, 3rd, 5th, 7th and 9th groups are anomalies, and the remaining groups are non-anomalies. The final obtained abnormal position is as shown in Figure 6 .
[0209] In Example Three, the present application further proposes an anomaly detection method and system based on a multi-modal visual large model based on the above-mentioned Example Two, which is used to detect the air compressor mesh cover. In the specific use process, when performing anomaly detection, step (5) of Example Two is not used for IOU judgment. In step (2) of obtaining the detected part, the groudding-dino large model is not used to detect the air compressor mesh cover. Instead, a rectangular frame is manually framed on the template image during modeling and recorded. During registration, the template image is transformed by perspective transformation, and the image and the rectangular frame coordinates are transformed to the corresponding inspection image. The subsequent steps are the same.
[0210] The operation process is as follows:
[0211] (1) Collect images and perform image registration operation. The registered image is as shown in Figure 7 .
[0212] (2) Manually frame and label the equipment to be detected in the template image to obtain the target frame position information of the equipment to be detected and record it, as shown in Figure 8 .
[0213] (3) The rectangular frame generated in step (2) is input into the SAM segmentation large model as a prompt, and the segmentation result is obtained. As shown in Figure 9 , the left image is the segmentation effect diagram of the template image, and the right image is the segmentation effect diagram of the inspection image.
[0214] (4) Extract the contour mask of the two images, separate, and obtain the contour map of each part of the segmentation.
[0215] (5) Shape comparison is performed on each contour, and groups with large shape differences are removed, and 19 groups of sub-images are left. The similarity of these groups of sub-images is judged by grouping.
[0216] (6) The similarity matching scores of the 19 groups of suspected objects are calculated. Different block sizes and sliding window sizes are used for sub-images of different resolutions. The block resolution is 16*16. For example, if the image resolution is 48*48, the block size is 3*3, and if the image size is 64*64, the block size is 4*4. The final similarity scores are calculated in this way: 0.76, 0.87, 0.62, 0.73, 0.68, 0.69, 0.78, 0.87, 0.58, 0.63, 0.78, 0.62, 0.73, 0.76, 0.78, 0.61, 0.72, 0.31, 0.79. Set 0.5 as the threshold. The 17th group is abnormal, and the remaining groups are non-abnormal. The final abnormal position is shown in Figure 10 .
[0217] In Example Four, the present application further proposes an abnormality detection method and system based on a multi-modal visual large model based on the above-mentioned Example One, which is used to detect the air compressor mesh cover. In the specific use process, step (3) obtains the air compressor mesh cover part, and does not use the groudding-dino large model to detect the air compressor mesh cover. Instead, a rectangular frame is manually framed on the template image during modeling and recorded. During registration, the template image is transformed through perspective transformation, and the image and the rectangular frame coordinates are transformed to the corresponding inspection image. The subsequent steps are the same, and the operation process is as follows:
[0218] (1) Collect images and perform image registration operation. The registered image is shown in Figure 11 .
[0219] (2) Manually frame and label the equipment to be detected in the template image to obtain the target frame position information of the equipment to be detected and record it, as shown in Figure 12 .
[0220] (3) The rectangular frame generated in step (2) is input into the SAM segmentation large model as a prompt, and the segmentation result is obtained. As shown in Figure 13 , the left image is the segmentation effect diagram of the template image, and the right image is the segmentation effect diagram of the inspection image.
[0221] (4) Extract the contour mask of the two images, separate, and obtain the contour diagram of each part of the segmentation.
[0222] (5) IOU comparison is performed on each contour. After IOU matching filtering, the following 3 groups of suspected abnormalities are obtained, and the similarity of the 3 groups of sub-images is judged by grouping.
[0223] (6) The similarity matching scores of the three groups of suspected objects are calculated as 0.76, 0.42, and 0.47, respectively. Setting 0.5 as the threshold, the second and third groups are abnormal, and the remaining groups are non-abnormal. The final abnormal positions are as shown in Figure 14
[0224] In another aspect, the present application also discloses a computer readable storage medium, which stores a computer program. The computer program is executed by a processor, so that the processor executes the steps of the above method.
[0225] In another aspect, the present application also discloses a computer device, which comprises a memory and a processor. The memory stores a computer program. The computer program is executed by the processor, so that the processor executes the steps of the above method.
[0226] In another embodiment, the present application also provides a computer program product containing instructions, which, when executed on a computer, cause the computer to execute the above-mentioned method for anomaly detection based on a multi-modal visual large model.
[0227] It can be understood that the system provided by the embodiments of the present application corresponds to the method provided by the embodiments of the present application, and the explanation, examples and beneficial effects of related contents can refer to the corresponding parts in the above method.
[0228] The present application also provides an electronic device, which comprises a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory complete mutual communication through the communication bus,
[0229] The memory is used to store a computer program.
[0230] The processor is used to execute the program stored on the memory, so as to realize the above-mentioned method for anomaly detection based on a multi-modal visual large model.
[0231] The communication bus mentioned in the above electronic device can be a peripheral component interconnect bus or an extended industry standard architecture bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc.
[0232] The communication interface is used for communication between the above-mentioned electronic device and other devices.
[0233] The memory can include a random access memory and can also include a non-volatile memory, such as at least one disk memory. Optionally, the memory can also be at least one storage device located away from the above-mentioned processor.
[0234] The processor described above can be a general processor, including a central processing unit, a network processing unit, etc.; can also be a digital signal processor, an application specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component.
[0235] It should be further explained that the electronic device also includes a terminal device, which can also be referred to as a terminal, a user equipment, a mobile station, a mobile terminal, etc. The terminal device can be a mobile phone, a smart television, a wearable device, a tablet computer, a computer with wireless transceiving function, a virtual reality terminal device, an augmented reality terminal device, a wireless terminal in industrial control, a wireless terminal in unmanned driving, a wireless terminal in remote surgery, a wireless terminal in smart power grid, a wireless terminal in transportation safety, a wireless terminal in smart city, a wireless terminal in smart home, etc. The embodiments of the present application do not limit the specific technology and specific device form of the terminal device.
[0236] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable devices. The computer instructions can be stored in a computer readable storage medium or transmitted from one computer readable storage medium to another computer readable storage medium, for example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be magnetic media (such as floppy disk, hard disk, magnetic tape), optical media (such as DVD), or semiconductor media (such as solid state disk), etc.
[0237] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
[0238] In addition, it needs to be explained that if the embodiments of the present application involve directional indications (such as up, down, left, right, front, back, etc.), the directional indications are only used to explain the relative position relationship, movement condition, etc. between components in a certain posture, and if the certain posture changes, the directional indications will also change accordingly.
[0239] In addition, if the embodiments of the present application involve descriptions such as "first", "second", etc., the descriptions of "first", "second", etc. are only for description purposes and cannot be understood as indicating or implying the relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include at least one of the features. In addition, the meaning of "and / or" appearing throughout the text includes three parallel schemes. Taking "A and / or B" as an example, it includes A scheme, or B scheme, or A and B scheme. In addition, in the embodiments of the present application, "a plurality of" means two or more. In addition, the technical solutions of each embodiment can be combined with each other, but it must be based on the realization of the ordinary skilled in the art, when the combination of technical solutions appears contradictory or cannot be realized, it should be considered that the combination of technical solutions does not exist, nor is it within the scope of protection required by the present application.
Claims
1. A multi-modal vision large model-based anomaly detection method, characterized in that, Comprise: Move to the designated point by the inspection equipment, adopt the preset shooting distance and angle to the designated shooting scene to take the specified magnification to shoot and acquire the image; Take the image acquired by the inspection equipment in the initial inspection as the template image, take the image acquired by the inspection equipment in the subsequent inspection as the detection image, compare the inspection image with the template image, and judge the abnormality and the position of the abnormality in the image through the following steps: S1. Register the template image and the image to be detected; S2. Use the open set target detection large model to identify the target in the image without scene samples to detect various abnormal parts; S3. Use the segmentation large model to perform image segmentation in the detection box obtained by the open set target detection large model; S4. Compare the contours generated by the template image and the image to be detected, eliminate part of the difference and analyze region by region, and the contour with an IOU less than a specified threshold is set as a possible abnormal position; S5. The difference contour adopts a twin network calculation to compare the similarity, and the contour with a similarity greater than a preset threshold is eliminated, and the remaining is the abnormal area; The open set target detection large model in the S2 step adopts an open set target detection model based on a transfomer network structure, comprising: A text backbone network and an image feature extraction backbone network, the text backbone network adopts BERT to extract text features, and the image feature extraction network adopts a swin Transformer network to extract multi-scale image features; A text-image feature fusion module for inputting image features and text features into the fusion module for cross-modal feature fusion; A language-guided query selection module for selecting fusion features related to the input text as the decoder query; A multi-modal decoder for sending the cross-modal query to the self-attention layer to combine the image cross-attention layer for obtaining image features, the text cross-attention layer for text features, and the FFN layer in each cross-modal decoder layer; An output layer for outputting the updated query and reference point as the final box and class prediction; The twin network in the S5 step adopts multi-layer convolution, activation, and pooling to finally extract a feature vector, and the twin network adopts an ESSIM structural similarity formula obtained by improving the SSIM structural similarity formula as a loss function, which is a loss function formula based on the brightness, contrast, and structural similarity of the image. wherein, respectively are original image features of the template image and original image features of the image to be detected, denotes a feature and a feature an ESSIM structure similarity evaluation result, is a feature an image feature of the template image processed by a sobel operator, is a feature an image feature of the image to be detected processed by a sobel operator, is a feature an image feature of the template image processed by a laplacian operator, is a feature an image feature of the image to be detected processed by a laplacian operator, , is a specified proportionality coefficient, denotes an SSIM structure similarity calculation. 2.The multi-modal vision-based large model based anomaly detection method of claim 1, wherein, In the registration process of the S1 step, superpoint is used to extract image feature points and generate feature point descriptors, lightglue is also used to optimally match the feature points, and based on the matching feature point data, the RANSAC algorithm is used to calculate the homography transformation matrix of the template image and the image to be detected, which is used to register the two images, wherein the specific operation process of the RANSAC algorithm comprises: S131. Determine the key points in the matching feature point data based on the feature point descriptors, randomly extract four non-collinear samples to calculate the homography transformation matrix; S132. Test all data by the homography transformation matrix, calculate the re-projection coordinates of the feature points in the first frame image in the second frame image according to the homography matrix, compare the distance between the re-projection coordinates and the matched feature point coordinates, if less than a specified threshold, it is considered to be a correct matching point pair, otherwise it is considered to be a false matching, record the number of correct point pairs; S133. Repeat S131-S132 steps until the specified number of cycles, compare the number of correct matching point pairs after the cycle, take the case with the most correct matching point pairs as the final result, eliminate false matches, output correct matching pairs, to realize the screening of feature point matching. 3.The multi-modal vision-based large model based anomaly detection method of claim 1, wherein, In the S3 step, the type and rectangular frame generated by the open set target detection large model are input into the segmentation large model to generate a segmentation mask, wherein the segmentation large model comprises the following modules: An image encoder module for encoding an image, mapping the image to a feature space; A prompt encoder module for position encoding and learning embedding of the generated prompts including points, frames, and texts, converting the prompt information into a feature vector; A mask decoder module for integrating the feature vectors output by the image encoder module and the prompt encoder module, and decoding the final segmentation mask. 4.The multi-modal vision-based large model based anomaly detection method of claim 1, wherein, The specific operation steps of the S5 step include: S51. Calculate the circumscribed rectangle of the suspected abnormal contour mask and obtain its vertex coordinates; S52. Extract a sub-image block pair from the inspection image and the template image according to the circumscribed rectangle coordinates; S53. Input the sub-image into the twin network for calculation and output a similarity score; S54. According to the preset threshold, set the similarity score less than the preset threshold as the final extracted abnormality.
5. A multi-modal vision large model based anomaly detection system, characterized in that, An abnormality detection method based on a multi-modal visual large model according to any one of claims 1-4, comprising: An inspection device for moving to a specified point to take a specified scene image at a specified shooting distance, a specified angle, and a specified magnification; A device navigation module for navigating and controlling the inspection device to move to the specified point; An image registration module for registering the template image and the to-be-detected image; An open set target detection large model for identifying and detecting abnormal components in the image and corresponding generating type data and a rectangular frame; A segmentation large model for generating a segmentation mask based on the type data and the rectangular frame generated by the open set target detection large model, to perform image segmentation on the abnormal components in the detection frame; A contour comparison and judgment module for comparing the contours generated by the segmentation of the template image and the to-be-detected image, eliminating part of the differences and analyzing region by region; A similarity comparison module for comparing the similarity of the difference contours to obtain an abnormal region.
6. A computer-readable storage medium, characterized in that, A computer program is stored, which, when executed by a processor, causes the processor to perform the steps of the method of any one of claims 1 to 4.
7. A computer device, comprising: A memory and a processor are included, and the memory stores a computer program, which, when executed by the processor, causes the processor to perform the steps of the method of any one of claims 1 to 4.
Citation Information
Patent Citations
Laser data-based substation foreign object identification method
CN108107444A
Machine Vision-Based Foreign Object Intrusion Detection Device and Method for Subway Tunnels
CN113673614B
Power line foreign object detection method based on Faster R-CNN
CN113989209B
Method for detecting foreign matters in bullet train of unknown category
CN116342944A
SAM deep learning-based railway track foreign matter intrusion detection method and system, storage medium and terminal
CN117456477A