Industrial anomaly detection method and device based on multi-modal interaction

Through multimodal interaction and contrast learning methods, combined with two-dimensional images and three-dimensional point cloud data, the problems of scarcity of abnormal samples and difficult to capture depth structure abnormalities in three-dimensional industrial anomaly detection are solved, and efficient and accurate abnormal detection is achieved.

CN120125941APending Publication Date: 2025-06-10GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510062514.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Three-dimensional industrial anomaly detection faces the problems of scarce abnormal samples, disordered point cloud data and difficult to capture deep structure abnormalities, resulting in low detection efficiency and accuracy.

Method used

Using an industrial anomaly detection method based on multimodal interaction, we use the point cloud data of the object to be detected to obtain attention maps, mask feature blocks and reconstruct it, and combine two-dimensional image features with three-dimensional point cloud geometric characteristics to achieve the deep fusion of semantic information and geometric information.

Benefits of technology

The efficiency and accuracy of industrial anomaly detection are improved, especially when abnormal samples are scarce, accurate anomaly detection of complex industrial structures can be achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125941A_ABST
    Figure CN120125941A_ABST
Patent Text Reader

Abstract

The invention discloses an industrial anomaly detection method and device based on multi-modal interaction, and the method comprises the following steps: sampling the point cloud data of a to-be-detected article, and obtaining a feature block set; according to the point cloud data of the to-be-detected article, obtaining an attention image set through image coding; respectively covering the feature block sets through the attention image sets to obtain covering blocks; and reconstructing the covering block to obtain a reconstructed point cloud, and determining an abnormal state of the to-be-detected article according to the reconstructed point cloud and the point cloud data of the to-be-detected article. According to the invention, the industrial anomaly detection efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of anomaly detection, and particularly to an industrial anomaly detection method and device based on multimodal interaction. Background Art

[0002] With the rapid development of computer vision technology, industrial anomaly detection based on deep learning has been widely applied in fields such as industrial inspection, autonomous driving, and video surveillance, becoming a hot research direction. As a global leader in manufacturing, China faces the challenge of how to further improve product quality while ensuring production efficiency in the process of promoting high-quality development. Three-dimensional industrial anomaly detection, as an important branch of industrial inspection, has a wide range of application scenarios, covering multiple fields such as intelligent manufacturing, automotive industry, and aerospace. It can accurately detect geometric defects, assembly errors of components, and potential faults of key equipment, demonstrating irreplaceable value in the quality assurance of high-complexity components.

[0003] However, three-dimensional industrial anomaly detection faces many challenges, including the scarcity of anomaly samples, the disorder of point cloud data, and the difficulty of capturing deep structure anomalies. Traditional two-dimensional image-based detection methods perform well in detecting surface texture and color anomalies, but they cannot effectively handle deep structure anomalies, resulting in the omission of key defects. In addition, existing three-dimensional detection methods also have limitations. For example, the template matching method based on the memory bank mechanism has a high computational cost and poor scalability; although self-supervised learning methods have improved performance, sparse point cloud data is difficult to provide rich semantic information; and excellent two-dimensional pre-trained models still face technical difficulties in how to efficiently transfer knowledge to the three-dimensional field. These problems limit the efficiency and accuracy of three-dimensional industrial anomaly detection, and there is an urgent need to design more efficient solutions. Summary of the Invention

[0004] To overcome the above-mentioned defects of the prior art, the present invention provides an industrial anomaly detection method and device based on multimodal interaction, which can improve the efficiency of industrial anomaly detection.

[0005] An embodiment of the present invention provides an industrial anomaly detection method based on multimodal interaction, including the following steps:

[0006] Sampling the point cloud data of the item to be detected to obtain a set of feature blocks;

[0007] According to the point cloud data of the item to be detected, obtaining an attention map set through image encoding;

[0008] Masking the set of feature blocks respectively through the attention map set to obtain masked blocks;

[0009] Reconstruct the occlusion block to obtain a reconstructed point cloud, and determine the abnormal state of the item to be detected based on the reconstructed point cloud and the point cloud data of the item to be detected.

[0010] Further, sampling the point cloud data of the item to be detected to obtain a set of feature blocks, specifically including:

[0011] Perform downsampling on the point cloud data to obtain a downsampled point cloud set;

[0012] Calculate the overall change rate of each point in the downsampled point cloud set, and determine the points with the overall change rate ranking higher than the preset ranking threshold as sampling points;

[0013] Search for all the sampling points respectively by the k-nearest neighbor method to obtain k neighboring points corresponding to each sampling point, and aggregate the features of the k neighboring points through a preset mini-PointNet network to obtain the feature blocks corresponding to each sampling point, and finally aggregate all the feature blocks to obtain the set of feature blocks.

[0014] Preferably, calculating the overall change rate of each point in the downsampled point cloud set specifically includes:

[0015] Determine all the points adjacent to the position of the point to be calculated in the downsampled point cloud set as adjacent points, and extract all the adjacent points as the adjacent set corresponding to the point to be calculated;

[0016] Calculate the normal vector change rate and the curvature change rate between the point to be calculated and each adjacent point respectively, and calculate the overall change rate corresponding to the point to be calculated according to the normal vector change rate and the curvature change rate; wherein, the calculation formula of the overall change rate is specifically:

[0017]

[0018] Among them, R is the overall change rate, A is the total number of adjacent points, is the adjacent set, P i is the point to be calculated, P j is the adjacent point in the adjacent set, R norm (P i ,P j ) is the normal vector change rate between P i and P j , R curv (P i ,P j ) is the curvature change rate between P i and P j .

[0019] Preferably, calculating the normal vector change rate and the curvature change rate between the point to be calculated and each adjacent point respectively specifically includes:

[0020]

[0021] Among them, N i , N j are the normal vectors of point P i and point P j respectively, and v ij is the vector from point P i to point P j ;

[0022] R curv (P i , P j ) = |K i - K j |

[0023] Among them, K i and K j are the curvatures of point P i and point P j respectively, and the calculation formula of the curvature K i corresponding to point P i is specifically:

[0024]

[0025] Among them, λ 1 and λ 2 are the eigenvalues of the covariance matrix of the adjacent set corresponding to point P i , and ∈ is a constant greater than 0.

[0026] Furthermore, obtaining the attention map set by image encoding according to the point cloud data of the item to be detected specifically includes:

[0027] Generating a multi-view projection map set according to the point cloud data of the item to be detected; among them, the multi-view projection map set includes the projection maps obtained by projecting the point cloud data along the x, y, and z axes respectively;

[0028] Extracting all the projection maps through a preset two-dimensional model according to the multi-view projection map set to obtain the image features corresponding to the projection maps;

[0029] Performing max pooling on the image features corresponding to the projection maps to obtain the attention map corresponding to the projection maps, and combining the attention maps corresponding to all the projection maps into the attention map set.

[0030] Furthermore, masking the feature block set through the attention map set to obtain the masked block specifically includes:

[0031] Back-project all the attention maps in the attention map set into 3D space, and aggregate the back-projected attention maps to obtain an attention cloud; wherein, the corresponding positions of the points in the attention map in the attention cloud are determined by a point cloud index;

[0032] Calculate an attention score for each point in the attention cloud respectively according to the attention values of the points in the attention cloud in different attention maps; wherein, the specific calculation formula of the attention score is:

[0033]

[0034] wherein, S is the attention score, Softmax represents a preset Softmax function, I2P represents a back-projection operation, is the attention map set, P T represents the position coordinates of the point to be calculated in the attention cloud;

[0035] Sort all the attentions, and finally mask the points with the attention scores ranked in the last a% to obtain the masked block; wherein, a is a preset masking ratio.

[0036] Further, the reconstruction of the masked block to obtain a reconstructed point cloud specifically includes:

[0037] Encode the masked block through a preset encoder to obtain an encoded block;

[0038] After connecting the encoded block and a preset blank block, input them into a preset decoder to obtain a decoded block;

[0039] According to the decoded block, reconstruct the masked part of the masked block to obtain a reconstructed block, and combine the reconstructed block with the masked block to obtain the reconstructed point cloud.

[0040] Preferably, after obtaining the reconstructed point cloud, it further includes:

[0041] Back-project all the projection maps into 3D space to obtain back-projected maps;

[0042] Aggregate the reconstructed point cloud and the back-projected maps in a preset invariant space to respectively obtain an invariant point cloud and an invariant image;

[0043] Calculate the similarity between the invariant point cloud and the invariant image through a preset loss function, and optimize the preset decoder with the goal of maximizing the similarity; wherein, the specific preset loss function is:

[0044]

[0045] Among them, is the loss function value, N is the number of the invariant point clouds, is the k-th invariant point cloud, is the invariant image, s(·) represents the cosine similarity function, and τ is the preset temperature coefficient.

[0046] Furthermore, determining the abnormal state of the item to be detected according to the reconstructed point cloud and the point cloud data of the item to be detected specifically includes:

[0047] Calculating the cosine similarity between the reconstructed point cloud and the point cloud data of the item to be detected, and when the cosine similarity is less than the preset similarity threshold, determining that the abnormal state of the item to be detected is abnormal.

[0048] Another embodiment of the present invention provides an industrial anomaly detection device based on multimodal interaction, including: a sampling module, an encoding module, a masking module, and a reconstruction module;

[0049] The sampling module is used to sample the point cloud data of the item to be detected to obtain a set of feature blocks;

[0050] The encoding module is used to obtain an attention map set through image encoding according to the point cloud data of the item to be detected;

[0051] The masking module is used to mask the set of feature blocks through the attention map set to obtain masked blocks;

[0052] The reconstruction module is used to reconstruct the masked blocks to obtain a reconstructed point cloud, and determine the abnormal state of the item to be detected according to the reconstructed point cloud and the point cloud data of the item to be detected.

[0053] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0054] By combining two-dimensional image features with three-dimensional point cloud geometric characteristics and performing contrastive learning on the point cloud and multi-dimensional features, the method of the present invention can achieve deep fusion of semantic information and geometric information under the condition of limited industrial anomaly data, improve the robustness and generalization ability of the reconstructed point cloud representation, and further improve the efficiency of industrial anomaly inspection. The method of the present invention is particularly applicable to the problem of scarce abnormal samples in industrial scenarios and helps to achieve accurate anomaly detection of complex industrial structures. Figure 2 BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 is a schematic flowchart of an industrial anomaly detection method based on multimodal interaction provided by an embodiment of the present invention.

[0056] Figure 2 ​Schematic diagram of the structure of a point cloud encoding and decoding network provided by an embodiment of the present invention.

[0057] Figure 3 Schematic diagram of the process of a preferred embodiment of an industrial anomaly detection method based on multimodal interaction provided by an embodiment of the present invention.

[0058] Figure 4 Schematic diagram of the structure of an industrial anomaly detection device based on multimodal interaction provided by another embodiment of the present invention. Detailed implementation manners

[0059] The drawings are only for illustrative purposes and should not be construed as a limitation of this patent;

[0060] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0061] The technical solutions of the present invention will be further described below with reference to the drawings and embodiments.

[0062] Refer to Figure 1 , which is a schematic diagram of the process of an industrial anomaly detection method based on multimodal interaction provided by an embodiment of the present invention, and includes the following steps:

[0063] S1: Sample the point cloud data of the item to be detected to obtain a set of feature blocks;

[0064] S2: According to the point cloud data of the item to be detected, obtain an attention map set through image encoding;

[0065] S3: Mask the set of feature blocks respectively through the attention map set to obtain masked blocks;

[0066] S4: Reconstruct the masked blocks to obtain a reconstructed point cloud, and determine the abnormal state of the item to be detected according to the reconstructed point cloud and the point cloud data of the item to be detected.

[0067] For step S1, specifically, the sampling of the point cloud data of the item to be detected to obtain a set of feature blocks specifically includes:

[0068] Perform downsampling on the point cloud data to obtain a downsampled point cloud set;

[0069] Calculate the overall change rate of each point in the downsampled point cloud set, and determine the points with the overall change rate ranking higher than the preset ranking threshold as sampling points;

[0070] Search all the sampling points respectively by the k-nearest neighbor method to obtain k nearest points corresponding to each sampling point, and aggregate the features of the k nearest points through a preset mini-PointNet network to obtain a feature block corresponding to each sampling point. Finally, aggregate all the feature blocks to obtain the feature block set.

[0071] Preferably, calculating the overall change rate of each point in the downsampled point cloud specifically includes:

[0072] Determine all the points adjacent to the position of the point to be calculated in the downsampled point cloud as adjacent points, and extract all the adjacent points as the adjacent set corresponding to the point to be calculated;

[0073] Calculate the normal vector change rate and the curvature change rate between the point to be calculated and each adjacent point respectively, and calculate the overall change rate corresponding to the point to be calculated according to the normal vector change rate and the curvature change rate; wherein, the specific formula for the overall change rate is:

[0074]

[0075] wherein, R is the overall change rate, A is the total number of adjacent points, is the adjacent set, P i is the point to be calculated, P j is the adjacent point in the adjacent set, R norm (P i ,P j ) is the normal vector change rate between P i and P j , and R curv (P i ,P j ) is the curvature change rate between P i and P j .

[0076] Preferably, calculating the normal vector change rate and the curvature change rate between the point to be calculated and each adjacent point respectively specifically includes:

[0077]

[0078] wherein, N i , N j are the normal vectors of point P i and point P j respectively, and v ij is the vector from point P i to point P j ;

[0079] r curv (P i ,Pjj ) = |K i -K j |

[0080] where K i and K j are the curvatures of point P i and point P j respectively, and the calculation formula of the curvature K i corresponding to point P i is specifically as follows:

[0081]

[0082] where λ 1 and λ 2 are the eigenvalues of the covariance matrix of the adjacent set corresponding to point P i , and ∈ is a constant greater than 0.

[0083] In a preferred embodiment, in the point cloud downsampling stage, this preferred embodiment downsamples the number of points of the point cloud data of the item to be measured from N to M, denoted as Perceptual downsampling significantly improves the efficiency of the model in subsequent analysis by sparsifying the point cloud data, while avoiding the loss of key information that may be caused by random sampling in traditional methods.

[0084] Subsequently, by analyzing the normal vector and the local curvature change rate of the point cloud, points with higher geometric importance are preferentially selected for retention, which can enhance the sensitivity to abnormal regions. Among them, the curvature of a point can be calculated using the eigenvalues of the covariance matrix C i of the normal vector and its neighboring points, as shown in the following formula:

[0085]

[0086] where λ 1 and λ 2 are the eigenvalues of the covariance matrix C i , and ∈ is a small positive constant to prevent the denominator from being zero.

[0087] At the same time, in order to quantify the normal vector and the curvature change rate between two points (P i , P j ), the following formula is defined:

[0088]

[0089] R curv (P i , P j ) = |K i -K j |

[0090] where Ni , N j are the normal vectors of points P i and P j respectively, and v ij is the vector from point P i to point P j . K i and K j are the curvatures of points P i and P j respectively.

[0091] This preferred embodiment also introduces a change rate memory bank to capture local and global change rates in the point cloud. For each point P i , the memory bank stores the change rates of its adjacent points in terms of normal vector and curvature values. Thus, given the change rates R norm and R curv of the normal vector and curvature, the overall change rate R is calculated as the average of the change rates of its adjacent points:

[0092]

[0093] To preferentially select points with higher change rates, we sort the points based on the values in the memory bank . A lower rank indicates a higher change rate, which helps to capture significant abnormal structures. We select points with larger change rates by setting a threshold τ. By sampling from the points ranked from 1 to (where N is the total number of points), we can capture more points with significant change rates. The calculation of the final sampled point set (S) can be expressed as:

[0094] S = {P k | Rank(P k ) ≤ [τ · N]}

[0095] By introducing this geometric-aware point cloud sampling strategy, the accuracy and effectiveness of point cloud anomaly detection can be enhanced, thus ensuring that the sampled point set can represent abnormal structures with significant geometric features.

[0096] Finally, use k-nearest neighbors to search for the k neighbors of each sampled point, and aggregate the features of these neighbors through a mini-PointNet network to obtain the feature blocks of M points. Each feature block can represent a local spatial structure and perform long-range feature interaction with other feature blocks in subsequent encoding. We represent these feature blocks where C represents the feature dimension.

[0097] For step S2, specifically, the method of obtaining the attention map set through image encoding according to the point cloud data of the item to be detected specifically includes:

[0098] Generate a multi-view projection atlas based on the point cloud data of the item to be detected; wherein, the multi-view projection atlas includes projection maps obtained by projecting the point cloud data along the x, y, and z axes respectively;

[0099] Extract, through a preset two-dimensional model, all the projection maps according to the multi-view projection atlas to obtain image features corresponding to the projection maps;

[0100] Perform max pooling on the image features corresponding to the projection maps to obtain an attention map corresponding to the projection map, and combine the attention maps corresponding to all the projection maps into the attention atlas.

[0101] In a preferred embodiment, for the point cloud data of the item to be detected It is necessary to project the input point cloud P i onto three orthogonal views along the x, y, and z axes respectively. For each point, directly ignore one of its three coordinates and round the other two coordinates to obtain the two-dimensional position on the corresponding projection map. The multi-view projection atlas corresponding to the finally obtained point cloud P is denoted as

[0102] After obtaining the multi-view projection atlas, use the preset two-dimensional model (preferably WideResNet50) to extract multi-view image features, and each feature has C channels, denoted as where H and W represent the height and width of the feature map respectively. These two-dimensional features contain sufficient high-level semantic information learned from the large model image data. In addition, through encoding from different perspectives, the geometric information lost during the projection process can also be alleviated.

[0103] Subsequently, perform per-pixel max pooling on to reduce the feature dimension to one dimension and obtain a semantic attention map for each view. These single-channel attention maps represent the semantic importance of different regions of the image, and the attention atlas obtained by combining all the attention maps is denoted as where

[0104] For step S3, specifically, the masking of the feature block set by the attention atlas to obtain a masked block specifically includes:

[0105] Back-project all the attention maps in the attention atlas into 3D space and aggregate the back-projected attention maps to obtain an attention cloud; wherein, the corresponding positions of the points in the attention map in the attention cloud are determined by the point cloud index;

[0106] Calculate the attention scores for each point in the attention cloud respectively according to the attention values of the points in the attention cloud in different attention maps; wherein, the specific calculation formula of the attention score is as follows:

[0107]

[0108] wherein, S 3D is the attention score, Softmax represents a preset Softmax function, I2P represents a back-projection operation, is the attention map set, and P T represents the position coordinates of the point to be calculated in the attention cloud;

[0109] Sort all the attentions, and finally mask the points with the attention scores sorted in the last a% to obtain the masked block; wherein, a is a preset masking ratio.

[0110] In a preferred embodiment, the traditional masking strategy randomly samples the masked sample area according to a uniform distribution, which may prevent the encoder from seeing important spatial features and disrupt the decoder due to unimportant structures. Therefore, the method described in the embodiment of the present invention uses a 2D attention map to guide the masking of the point cloud, so as to sample more semantically important 3D parts for the encoder.

[0111] First, index through the coordinates of the point cloud tokens to back-project all the attention maps in the multi-view attention map set into 3D space and aggregate them into a 3D attention cloud Specifically, it includes:

[0112] Match the point cloud with the corresponding 2D attention map by performing I2P (image back-projection to point cloud operation) on the attention map (the attention map represents the attention scores of each region of the image, which can be used to analyze the regions that the model focuses on, or participate in subsequent calculations as additional weight information). By performing I2P on the attention maps from three different perspectives respectively, output the attention values of each point in the 2D projection map, and then average the spatial attention values calculated from the three perspectives to more comprehensively reflect the attention distribution of the point cloud.

[0113] Then, apply the Softmax function to normalize the M points in S 3D and regard the magnitude of each element as the visible probability of the corresponding point in the visible point cloud, specifically including:

[0114] The averaged spatial attention scores are softmax-normalized into a probability distribution (ranging from 0 to 1), and then torch.multinomial is used to shuffle these probability distributions, randomly selecting a certain number of elements (representing important features), while maintaining randomness in each run. Then, the shuffled order is restored through torch.argsort to ensure that the masking operation is carried out in the order of significance.

[0115] In summary, the specific formula for calculating the attention score is as follows:

[0116]

[0117] where I2P(·) represents the 2D to 3D back-projection operation.

[0118] With this 2D semantic prior, random masking becomes non-uniform sampling with different probabilities, where feature blocks covering more key 3D structures are more likely to be retained. This method enhances the representation learning of the encoder by paying more attention to important 3D geometric structures, and at the same time provides richer masked feature block clues for the decoder, thus achieving better reconstruction.

[0119] Finally, according to the preset masking ratio (which can be adjusted according to the actual situation), the number of elements to be masked is determined, and then a mask is constructed to mark some elements as masked (True means retained, False means masked), and finally the mask is returned. This process aims to guide the selection of the mask through significance information while maintaining randomness and robustness.

[0120] For step S4, specifically, the reconstruction of the masked block to obtain the reconstructed point cloud specifically includes:

[0121] Encoding the masked block through a preset encoder to obtain an encoded block;

[0122] After connecting the encoded block and a preset blank block, inputting them into a preset decoder to obtain a decoded block;

[0123] According to the decoded block, reconstructing the masked part of the masked block to obtain a reconstructed block, and combining the reconstructed block with the masked block to obtain the reconstructed point cloud.

[0124] In a preferred embodiment, referring to Figure 2 , is a schematic structural diagram of a point cloud encoding and decoding network provided by an embodiment of the present invention. By Figure 2It can be seen that the point cloud encoding and decoding network includes the two parts of the preset encoder and the preset decoder. The preset encoder part contains 3 layers, each layer contains 5 encoding blocks, each encoding block contains an attention layer and a feed-forward neural network, and learns the representation of the global 3D shape through the visible part; the preset decoder consists of 2 layers, and each layer contains 1 encoding block.

[0125] Input the visible part of the occlusion block into the preset encoder for encoding, and the encoding block can be obtained Subsequently, is connected with a group of learnable feature blocks (i.e., the preset blank block) and input into the lightweight decoder, where M mask represents the number of occlusion tokens, and M = M mask + M vis , M vis is the occluded part of the occlusion block.

[0126] In the decoder, the learnable feature blocks learn to capture spatial information clues from the visible feature blocks and reconstruct the occluded 3D coordinates. Based on the decoded decoding block , use to reconstruct the 3D coordinates of the occluded feature block and its k neighboring points. The true 3D coordinates of the masked points are predicted through the reconstruction head of a single-layer linear projection Finally, calculate the loss function through the Chamfer distance, and the formula is as follows:

[0127]

[0128] where, H 3D (·) represents the reconstruction head.

[0129] Preferably, after obtaining the reconstructed point cloud, it further includes:

[0130] Back-project all the projection maps into the 3D space to obtain the back-projected maps;

[0131] Aggregate the reconstructed point cloud and the back-projected maps in a preset invariant space to obtain the invariant point cloud and the invariant image respectively;

[0132] Calculate the similarity between the invariant point cloud and the invariant image through a preset loss function, and optimize the preset decoder with the goal of maximizing the similarity; where, the preset loss function is specifically:

[0133]

[0134] where, is the loss function value, N is the number of the invariant point clouds, is the invariant point cloud, is the invariant image, s(·) represents the cosine similarity function, and τ is the preset temperature coefficient.

[0135] In a preferred embodiment, in order to better understand the point cloud by using the image modality, this preferred embodiment proposes a cross-modal learning objective, specifically guiding the model to be more biased towards the understanding of 3D point clouds.

[0136] First, project the multi-view feature map back to the 3D space, and then aggregate the reconstructed point cloud and the back-projected image into an invariant space to obtain the invariant point cloud and the invariant image The final optimization objective is to maximize and similarity. Such a cross-modal learning strategy forces the model to learn from more difficult positive and negative samples, thereby improving the representation ability compared to learning only from within-modal alignment.

[0137] For step S4, further, the determining the abnormal state of the item to be detected according to the reconstructed point cloud and the point cloud data of the item to be detected specifically includes:

[0138] Calculate the cosine similarity between the reconstructed point cloud and the point cloud data of the item to be detected. When the cosine similarity is less than the preset similarity threshold, determine that the abnormal state of the item to be detected is abnormal.

[0139] In summary, referring to Figure 3 , it is a schematic flowchart of a preferred embodiment of an industrial anomaly detection method based on multi-modal interaction provided by an embodiment of the present invention. As can be seen from Figure 3 , the present invention proposes an industrial anomaly detection method based on multi-modal interaction and contrast learning, which can make full use of the multi-modal characteristics of two-dimensional images and three-dimensional point clouds. Through the combination of geometric-guided sampling and multi-modal contrast learning, the model's recognition ability for anomalies in complex industrial scenarios is significantly improved, while reducing the dependence on the annotation of abnormal samples. At the same time, the reconstruction of abnormal point clouds is realized through the masked autoencoder model, so as to realize accurate three-dimensional anomaly detection.

[0140] In the method of the present invention, a perceptual downsampling module is also designed to achieve efficient sampling through geometric feature analysis; based on the idea of multi-modal information fusion, the attention features of two-dimensional images are used to guide the selective masking operation of three-dimensional point clouds to ensure that the semantically important point clouds are visible. Compared with the traditional random masking method, this method pays more attention to key spatial clues and retains important three-dimensional structural information.

[0141] In addition, the present invention also introduces a cross-modal contrastive learning strategy, which maximizes the consistency between point clouds and images in the invariant space, while encouraging the invariance of changes in the point cloud modality, thereby further enhancing the collaborative representation ability between modalities. Compared with traditional methods, this method optimizes the joint learning of point cloud and two-dimensional image features through a cross-modal fusion strategy, makes full use of multi-modal data, effectively alleviates the problems of scarce abnormal samples and information mismatch between modalities, and significantly improves the detection accuracy of abnormal samples.

[0142] Referring to Figure 4 , which is a schematic structural diagram of an industrial anomaly detection device based on multi-modal interaction provided by another embodiment of the present invention, including: a sampling module 101, an encoding module 102, a masking module 103, and a reconstruction module 104;

[0143] The sampling module 101 is used to sample the point cloud data of the item to be detected to obtain a set of feature blocks;

[0144] The encoding module 102 is used to obtain an attention map set through image encoding according to the point cloud data of the item to be detected;

[0145] The masking module 103 is used to mask the set of feature blocks through the attention map set to obtain masked blocks;

[0146] The reconstruction module 104 is used to reconstruct the masked blocks to obtain reconstructed point clouds, and determine the abnormal state of the item to be detected according to the reconstructed point clouds and the point cloud data of the item to be detected.

[0147] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. An industrial anomaly detection method based on multimodal interaction, characterized in that: The steps include: Sampling the point cloud data of the object to be inspected to obtain a feature block set; Obtaining an attention atlas through image encoding according to the point cloud data of the object to be detected; The feature block set is respectively masked by the attention atlas to obtain masked blocks; The masked block is reconstructed to obtain a reconstructed point cloud, and the abnormal state of the object to be detected is determined according to the reconstructed point cloud and the point cloud data of the object to be detected.

2. The industrial anomaly detection method based on multimodal interaction according to claim 1, characterized in that: The point cloud data of the object to be detected is sampled to obtain a feature block set, which specifically includes: Downsampling the point cloud data to obtain a downsampled point cloud set; Calculate the overall change rate of each point in the downsampled point cloud set, and determine the point whose corresponding overall change rate ranking is higher than a preset ranking threshold as the sampling point; All the sampling points are searched separately by the k-nearest neighbor method to obtain k neighboring points corresponding to each sampling point, and the features of the k neighboring points are aggregated by a preset mini-PointNet network to obtain the feature blocks corresponding to each sampling point, and finally all the feature blocks are aggregated to obtain the feature block set.

3. The industrial anomaly detection method based on multimodal interaction according to claim 2, characterized in that: The calculating the overall change rate of each point in the down-sampled point cloud set specifically includes: Determine all points in the downsampled point cloud set that are adjacent to the position of the point to be calculated as adjacent points, and extract all the adjacent points as an adjacent set corresponding to the point to be calculated; The normal vector change rate and the curvature change rate between the point to be calculated and each adjacent point are calculated respectively, and the overall change rate corresponding to the point to be calculated is calculated according to the normal vector change rate and the curvature change rate; wherein the calculation formula of the overall change rate is specifically: Wherein, R is the overall change rate, A is the total number of adjacent points, is the adjacent set, P i is the point to be calculated, P j is the adjacent point in the adjacent set, R norm (P i ,P j ) is P i With P j The rate of change of the normal vector between curv (P i ,P j ) is P i With P j The rate of change of curvature between .

4. The industrial anomaly detection method based on multimodal interaction according to claim 3, characterized in that: The respectively calculating the normal vector change rate and the curvature change rate between the point to be calculated and each adjacent point specifically includes: Among them, N i 、N j Point P i and point P j The normal vector, v ij From point P i To point P j A vector of R curv (P i ,P j )=|K i -K j | Among them, K i and K j Point P i and point P j The curvature of point P i Corresponding curvature K i The specific calculation formula is: Where λ1 and λ2 are points P i The eigenvalue of the covariance matrix corresponding to the adjacent set, ∈ is a constant greater than 0.

5. The industrial anomaly detection method based on multimodal interaction according to claim 1, characterized in that: The step of obtaining an attention atlas through image coding based on the point cloud data of the object to be detected specifically includes: Generate a multi-view projection atlas according to the point cloud data of the object to be detected; wherein the multi-view projection atlas includes projection images obtained by projecting the point cloud data along the x, y, and z axes respectively; According to the multi-view projection atlas, all the projection images are extracted respectively by using a preset two-dimensional model to obtain image features corresponding to the projection images; The image features corresponding to the projection image are subjected to maximum pooling to obtain an attention map corresponding to the projection image, and the attention maps corresponding to all the projection images are combined into the attention map set.

6. The industrial anomaly detection method based on multimodal interaction according to claim 1, characterized in that: The step of masking the feature block set by using the attention atlas to obtain the masked block specifically includes: Back-projecting all attention maps in the attention map set into 3D space, and aggregating the back-projected attention maps to obtain an attention cloud; wherein the corresponding position of each point in the attention map in the attention cloud is determined by the point cloud index; According to the attention values ​​of each point in the attention cloud in different attention maps, an attention score is calculated for each point in the attention cloud; wherein the calculation formula of the attention score is specifically: Wherein, S is the attention score, Softmax represents the preset Softmax function, I2P represents the back-projection operation, is the attention atlas, P T Represents the position coordinates of the point to be calculated in the attention cloud; All the attentions are sorted, and finally the points with the corresponding attention scores sorted in the last a% are masked to obtain the masked block; wherein a is a preset masking ratio.

7. The industrial anomaly detection method based on multimodal interaction according to claim 5, characterized in that: The step of reconstructing the mask block to obtain a reconstructed point cloud specifically includes: Encoding the mask block by a preset encoder to obtain an encoded block; After connecting the coding block and the preset blank block, the coding block is input into a preset decoder to obtain a decoding block; According to the decoded block, the masked part of the mask block is reconstructed to obtain a reconstructed block, and the reconstructed block is combined with the mask block to obtain the reconstructed point cloud.

8. The industrial anomaly detection method based on multimodal interaction according to claim 7, characterized in that: After obtaining the reconstructed point cloud, the method further includes: Back-projecting all the projection images into the 3D space to obtain back-projection images; Aggregating the reconstructed point cloud and the back-projection image in a preset invariant space to obtain an invariant point cloud and an invariant image respectively; The similarity between the invariant point cloud and the invariant image is calculated by a preset loss function, and the preset decoder is optimized with the goal of maximizing the similarity; wherein the preset loss function is specifically: in, is the loss function value, N is the number of the invariant point clouds, is the kth invariant point cloud, is the invariant image, s(·) represents the cosine similarity function, and τ is a preset temperature coefficient.

9. The industrial anomaly detection method based on multimodal interaction according to claim 1, characterized in that: The determining the abnormal state of the object to be detected according to the reconstructed point cloud and the point cloud data of the object to be detected specifically includes: The cosine similarity between the reconstructed point cloud and the point cloud data of the object to be detected is calculated, and when the cosine similarity is less than a preset similarity threshold, it is determined that the abnormal state of the object to be detected is abnormal.

10. An industrial anomaly detection device based on multimodal interaction, characterized in that: include: Sampling module, encoding module, masking module and reconstruction module; The sampling module is used to sample the point cloud data of the object to be detected to obtain a feature block set; The encoding module is used to obtain an attention atlas through image encoding according to the point cloud data of the object to be detected; The masking module is used to mask the feature block set through the attention atlas to obtain a masked block; The reconstruction module is used to reconstruct the mask block to obtain a reconstructed point cloud, and determine the abnormal state of the object to be detected based on the reconstructed point cloud and the point cloud data of the object to be detected.