3D zero sample anomaly detection method and device based on 2D multi-mode large model

By adopting a 2D multimodal large model method in 3D anomaly detection, integrating 3D and 2D information, enhancing CLIP's 3D understanding ability, solving the problem of insufficient generalization ability in the zero sample situation in the existing technology, and achieving effective anomaly detection on unseen point cloud data.

CN119992236AInactive Publication Date: 2025-05-13ZHEJIANG UNIV

Patent Information

Application Number
CN202510469946.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing 3D anomaly detection methods have poor generalization capabilities when facing unknown semantics, especially in the case of zero samples, which is difficult to effectively detect anomaly point clouds.

Method used

The 3D zero-sample anomaly detection method based on the 2D multimodal big model is adopted. By integrating 3D and 2D information, the 3D understanding ability of CLIP is enhanced. The 2D global and local information representation is extracted using the visual encoder of the multimodal big model CLIP, and 3D global and local representations are generated by combining 2D information. The design can learn prompt text for mixed representation learning, and ultimately realize zero-sample anomaly detection.

Benefits of technology

It realizes the ability to detect abnormalities on unseen point cloud data, can handle different categories of exceptions, has strong innovation and industrial field application value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992236A_ABST
    Figure CN119992236A_ABST
Patent Text Reader

Abstract

The invention discloses a 3D zero sample anomaly detection method and device based on a 2D multi-mode large model. Based on a 2D multi-modal model CLIP, 3D anomaly detection is realized under the condition of not depending on a target domain sample, and point cloud information from 3D and 2D is integrated through the following modes: firstly, point cloud is rendered from multiple perspectives through the CLIP, and 2D representation of the point cloud is acquired; projecting the 2D representation back to the 3D space, thereby understanding the 3D representation of the point cloud; additional regularization is performed on the 2D representation to further enhance the understanding of the frame on the 3D representation. After the point information and the pixel information of the point cloud are represented, the invention provides a mixed representation learning method, generalization abnormal modes in the point cloud and the pixels are captured into learnable text prompts, and accurate detection of 3D anomalies is realized. According to the method, high-precision zero-sample anomaly detection for 3D data is realized for the first time, and the method has high innovativeness and industrial application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of anomaly detection technology, and in particular to a 3D zero-sample anomaly detection method and device based on a 2D multimodal large model. Background Art

[0002] 2D anomaly detection is widely used in industrial inspection and other fields. However, some anomalies in reality often manifest as abnormal spatial relationships, which are difficult to detect only by relying on RGB information, especially when the defects are similar to the background or foreground. In contrast, 3D anomaly detection captures subtle abnormal spatial relationships by characterizing spatial features. Current 3D anomaly detection methods usually rely on storing the features of normal point clouds and identifying anomalies by calculating the distance between test samples and normal features. These methods assume that the target point cloud is accessible and normal during the training phase. However, in practice, due to privacy protection or data loss, the target data may not be available. This makes these methods have poor generalization ability when facing unknown semantics. Although zero-shot anomaly detection has been studied in 2D images, zero-shot 3D anomaly detection remains a challenge. This problem requires the model to be able to detect anomalies on unseen point cloud data and handle different categories of anomalies. In recent years, visual language models, especially CLIP, have performed well in multiple tasks due to their strong generalization ability. Applying CLIP to 3D anomaly detection provides an effective idea for solving zero-shot 3D anomaly detection. Summary of the invention

[0003] The purpose of this invention is to further explore and improve the deficiencies of existing anomaly detection methods, and propose a 3D zero-sample anomaly detection method based on a 2D multimodal large model. This method enhances CLIP's 3D understanding ability by integrating information from 3D and 2D, and has strong innovation; at the same time, it realizes zero-sample detection of abnormal point clouds, which has great industrial field application value.

[0004] The object of the present invention is achieved through the following technical solutions: In a first aspect, the present invention provides a 3D zero-sample anomaly detection method based on a 2D multimodal large model, the method comprising the following steps:

[0005] Step 1, obtaining an auxiliary point cloud data set: obtaining an auxiliary point cloud data set for anomaly detection;

[0006] Step 2, multi-view rendering of point cloud: render each point cloud and its true value into a two-dimensional image from each corresponding viewpoint, and save the original 3D information;

[0007] Step 3, 2D global and local information representation: For each point cloud, use the visual encoder of the multimodal large model CLIP to extract 2D global and local information representation from its rendered image;

[0008] Step 4, representation of 3D global and local information: for each point cloud, after obtaining its 2D global and local representations, combine the 2D global representation into a 3D global representation; based on the 2D local representation and the occlusion relationship between the original points in each view after rendering, obtain the 3D local representation;

[0009] Step 5, hybrid representation learning: design learnable prompt texts for normal point clouds and abnormal point clouds; and encode them using a CLIP-based text encoder;

[0010] Step 6, calculation of loss function: the total loss function is composed of the sum of 3D global loss function, 3D local loss function, 2D global loss function and 2D local loss function

[0011] Step 7, training of the zero-shot anomaly detection framework: During the training process, the original parameters of the CLIP model are frozen, and the learnable prompt text is optimized by minimizing the hybrid loss function;

[0012] Step 8, the reasoning process of detection: after completing the training, perform zero-shot reasoning based on point cloud, or multimodal 3D reasoning with RGB color information added; judge the anomaly based on the anomaly score obtained by reasoning, and segment the abnormal area based on the anomaly score map.

[0013] Furthermore, the rendering method in step 2 is specifically as follows:

[0014] For the point cloud dataset obtained in step 1 , Indicates Point cloud, Indicates The true value of the number of point clouds, there are point cloud, the rendering matrix of the 𝑘th perspective is defined as , a total of views; simultaneously render the point cloud and its corresponding true value from different perspectives to obtain the corresponding 2D rendering; specifically and ;in, and Representative The point cloud corresponds to The rendered images and the corresponding pixel-level truth values ​​are labeled as 1 and normal pixels are labeled as 0.

[0015] Furthermore, in step 4, the 3D global representation and local representation are specifically:

[0016] (4.1) 3D global representation: For point cloud, through 2D global representations are combined into a 3D global representation, expressed as ;

[0017] (4.2) 3D local representation: For point cloud, through The 2D local representation and the occlusion relationship of the original point in each view after rendering are used to obtain the 3D local representation, which is recorded as , the occlusion relationship between points is recorded as , indicating the Point cloud The point is Whether it is visible in the rendering, 1 for visible and 0 for invisible.

[0018] Furthermore, in step 4, based on the encoded global feature vector, the 2D global features of the rendering image of each point cloud are combined through the rendering relationship to obtain the 3D global features. The specific calculation method is:

[0019]

[0020] The 2D local features of the rendering of each point cloud are combined to obtain a preliminary 3D local representation, which is calculated as follows:

[0021]

[0022] in Indicates Point cloud In the view The feature of the point, n is the number of points in the k-th view, and the three-dimensional coordinates of the point are ; is the rendering conversion between points and pixels, coming from , is the rendering matrix of the 𝑘th perspective, Indicates the actual distance from the camera to the object. Indicates the maximum distance, used for normalization;

[0023] definition For the In the point cloud The view is projected as The set of points at pixels; then:

[0024]

[0025]

[0026] in is an indicator function, Indicates Point cloud The three-dimensional coordinates of a point.

[0027] Finally, Point cloud The 3D local features of each view are calculated by the following formula: .

[0028] Furthermore, in step 5, the learnable prompt text of the normal point cloud and the learnable prompt text of the abnormal point cloud are defined respectively. and , the specific form is:

[0029]

[0030]

[0031] in, and is the Eth learnable embedding word, which updates the learnable vectors of normal text and abnormal text during training.

[0032] Furthermore, the loss function in step 6 is specifically:

[0033] (1) 3D global loss function, which uses cross entropy to represent the difference between the text representation and the global representation of the rendered image at each viewpoint, specifically:

[0034]

[0035] in, Indicates that normal text features, abnormal text features and The cosine similarity of is used to determine the probability of anomaly. Indicates normal semantic text or abnormal semantic text features;

[0036] (2) 3D local loss function, based on Dice Loss to penalize incorrect abnormal region segmentation, specifically:

[0037]

[0038] in , , It means calculating the cosine similarity between the abnormal point cloud prompt text and each point in the point cloud, that is, the abnormal score of each point. Indicates that the normal score of each point is calculated. is a matrix of all 1s, with shape and Consistency;

[0039] (3) 2D global loss function, which represents the difference between the text representation and each 2D global representation through cross entropy, specifically:

[0040]

[0041] (4) 2D local loss function, based on Dice Loss and Focal Loss to alleviate the category imbalance problem, specifically:

[0042]

[0043] in Indicates a connection operation;

[0044] (5) Total loss function:

[0045] .

[0046] Furthermore, the specific calculation process of the anomaly score graph and the anomaly score in the two reasoning processes in step 8 is:

[0047] (1) Zero-shot reasoning based on point cloud: After obtaining 2D and 3D local features from the input point cloud, cosine similarity is calculated with the encoded text to obtain an anomaly score map, and then a global anomaly score is obtained;

[0048] The anomaly score graph is:

[0049]

[0050] in represents a Gaussian filter, Indicates the cosine similarity between the abnormal point cloud prompt text and each point in the point cloud;

[0051] The global anomaly score combines global and local anomaly semantics, and the calculation formula is:

[0052]

[0053] in Represents a cosine similarity calculation operation.

[0054] (2) Multimodal 3D reasoning with RGB color information

[0055] When color information is available, the two-dimensional color image can be directly Extract 2D global and local information to obtain 2D features to fuse color information; then project the 2D features back to 3D space to calculate the color anomaly map:

[0056]

[0057] The color anomaly scores are:

[0058]

[0059] in Represents the anomaly score value of each point in the three-dimensional space.

[0060] The final multimodal anomaly score map and anomaly score are:

[0061]

[0062]

[0063] Furthermore, the abnormality determination and abnormal area segmentation in step 8 specifically refers to determining whether the detected object has an abnormality according to the abnormality score, and segmenting the abnormal area of ​​the detected object according to the abnormality score map.

[0064] In a second aspect, the present invention also provides a 3D zero-sample anomaly detection device based on a 2D multimodal large model, comprising a memory and one or more processors, wherein the memory stores executable code, and when the processor executes the executable code, the 3D zero-sample anomaly detection method based on a 2D multimodal large model is implemented.

[0065] In a third aspect, the present invention further provides a computer-readable storage medium having a program stored thereon, and when the program is executed by a processor, the 3D zero-sample anomaly detection method based on a 2D multimodal large model is implemented.

[0066] Compared with the prior art, the present invention has the following innovative advantages and significant effects:

[0067] 1) It uses offline high-precision rendering technology to integrate the 3D and 2D information of point clouds, enhances the 3D understanding ability of the 2D multimodal large model CLIP, and is highly innovative;

[0068] 2) A learnable prompt text is proposed, and with the help of hybrid representation learning, zero-shot anomaly detection and multimodal zero-shot anomaly detection of 3D point clouds are realized, which has high field application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 It is a schematic diagram of the overall process of the 3D anomaly detection framework of the present invention.

[0070] Figure 2 It is a schematic diagram of the zero-sample detection reasoning process based on point cloud of the present invention.

[0071] Figure 3 It is a schematic diagram of multimodal reasoning of the present invention.

[0072] Figure 4 It is a structural diagram of a 3D zero-sample anomaly detection device based on a 2D multimodal large model provided by the present invention. DETAILED DESCRIPTION

[0073] The specific implementation method and working principle of the present invention are described in detail below in conjunction with the accompanying drawings:

[0074] Example

[0075] This example conducts corresponding training and 3D point cloud zero-sample detection experiments on the industrial abnormal point cloud dataset MVTec3D-AD. The dataset contains a total of 4171 point cloud data of 10 objects, with RGB information and true value annotations of the point cloud. The details of the dataset are shown in Table 1:

[0076] Table 1 Detailed information of the MVTec3D-AD dataset

[0077]

[0078] In this embodiment, the 3D zero-sample anomaly detection method based on the 2D multimodal large model is experimented on the above data set. The method results are to achieve zero-sample anomaly detection and positioning of 3D point clouds and multimodal zero-sample anomaly detection and positioning. The detailed implementation steps are as follows:

[0079] Step 1: Re-divide the MVTec3D-AD point cloud dataset shown in Table 1 to obtain an auxiliary dataset. Specifically, the original point cloud data in the training set, validation set, and test set in the cookie category are used as auxiliary point clouds, and all other data are used for testing to test the effect of zero-sample detection. This is the auxiliary point cloud dataset used for training in this example, Indicates Point cloud, Indicates The true value of the number of point clouds, there are Point cloud, , there are 260 normal point clouds and 103 abnormal point clouds;

[0080] Step 2: For the auxiliary point cloud dataset obtained in step 1 For each point cloud in the image, we use offline high-precision rendering to save the original 3D information. With the help of the Open3D library, we rotate the point cloud along the X-axis to 9 different angles, and render a 336*336 image at each angle. The 9 angles are: . The point cloud image obtained by rendering the point cloud and the corresponding true value image are recorded as and , represents the corresponding point cloud, Indicates the corresponding rendering;

[0081] The rendering matrix for the 𝑘th view is defined as , a total of Render the point cloud and its corresponding ground truth from different perspectives to obtain the corresponding 2D rendering. and .in, and Representative The point cloud corresponds to The renderings and the corresponding pixel-level truth values ​​are shown in Figure 1. Abnormal pixels are marked as 1 and normal pixels are marked as 0.

[0082] Step 3: After obtaining the rendering image through step 2, for each point cloud, use the visual encoder of the original transformer architecture of the multimodal large model CLIP to encode the rendering image into a global feature vector. Specifically, Encoding global feature vector and the local eigenvector , from the rendered image Extract 2D global information representation and 2D local information representation ; The global feature vector reflects the common information of all points in the rendering, and the local feature vector reflects the information of each point in the rendering;

[0083] Step 4: For each point cloud, after obtaining the 2D features of the rendering image through step 3, the 9 2D global features of the 9 rendering images of each point cloud are combined through the rendering relationship to obtain the 3D global features of the 3D point cloud. ,Right now:

[0084]

[0085] Similarly, for The 3D local representation is obtained through the 2D local representation and the occlusion relationship of the original points in each view after rendering. Specifically, the 9 2D local features of the 9 renderings of each point cloud are combined to preliminarily obtain the 3D local features of the 3D point cloud. ,Right now:

[0086]

[0087] in Indicates Point cloud In the view The point in The local features in the k-th rendering, n is the number of points in the k-th view, and the 3D coordinates of the point are ; is the rendering conversion between points and pixels, coming from , is the rendering matrix of the 𝑘th perspective, Indicates the actual distance from the camera to the object. Indicates the maximum distance, which is used for normalization. Next, calculate the occlusion relationship of the point cloud in each rendering:

[0088]

[0089]

[0090] in is an indicator function. For the In the point cloud The view is projected as The set of points at pixels; Indicates Point cloud The three-dimensional coordinates of the points. =1, indicating the Point cloud The point is There is no obstruction in the rendering. =0 means it is blocked.

[0091] Final Point cloud The 3D local features of each view are calculated by the following formula: ;

[0092] Step 5, hybrid representation learning: define the learnable prompt text of the normal point cloud as:

[0093]

[0094] The learnable hint text that defines the abnormal point cloud is:

[0095]

[0096] in, is a learnable vector for normal text, is a learnable vector of abnormal text, which is updated during training. Indicates a specific description object and is not updated during training; the above-mentioned learnable prompt words and Adopting dynamic embedding structure, and is a trainable parameter vector of dimension d, where d is the same as the text encoder dimension of the pre-trained text-image model. The training phase is optimized by backpropagation and The value is to minimize the normal sample and and maximize its semantic distance with The difference between the two. The reasoning phase is frozen and And use it as static word embedding to participate in text encoding.

[0097] Step 6: Use the text encoder of CLIP's original transformer architecture to convert the text in step 5 and Encoded as vectors and ; Step 7, calculation of loss function: the total loss function is recorded as , which is composed of the sum of four sub-loss functions, namely 3D global loss function, 3D local loss function, 2D global loss function and 2D local loss function, specifically:

[0098] (1) 3D global loss function: Calculate the cosine similarity between the text feature and the global representation of the rendered image at each viewpoint, and use cross entropy to calculate the 3D global loss function, denoted as ;

[0099] in, Indicates that normal text features, abnormal text features and The cosine similarity of is used to determine the probability of anomaly. Indicates normal semantic text or abnormal semantic text features;

[0100] (2) 3D local loss function: Dice Loss helps to penalize incorrect segmentation, especially when the abnormal area is small or the categories are unbalanced. Therefore, we use it to construct a 3D local loss function, denoted as ;

[0101]

[0102] in , , It means calculating the cosine similarity between the abnormal point cloud prompt text and each point in the point cloud, that is, the abnormal score of each point. Indicates that the normal score of each point is calculated. is a matrix of all 1s, with shape and Consistency;

[0103] (3) 2D global loss function: Use cross entropy to quantify the difference between the text representation and each global two-dimensional feature, denoted as ;

[0104]

[0105] (4) 2D local loss function: Since the abnormal area is usually smaller than the normal area and the class imbalance problem often affects the performance of the model, Dice Loss and Focal Loss are used to alleviate this problem, especially when dealing with small-sized and unbalanced abnormal areas. The 2D local loss function is denoted as ;

[0106]

[0107] in Indicates a connection operation;

[0108] Finally, the total loss function is expressed as:

[0109]

[0110] Step 8. After designing the loss function of step 8, freeze the original parameters of the CLIP model, optimize the learnable text prompts by minimizing the hybrid loss function, and train the learnable text prompts so that the CLIP model can effectively capture three-dimensional generalized abnormal patterns across categories, and finally enable the model to understand the abnormal semantics in point clouds and pixels.

[0111] Step 9: After training the learnable text prompts in step 8, two reasoning tests are performed:

[0112] (1) Without RGB information, in the data set shown in Table 1, the point cloud data except the cookie category is used as the test sample. During the test, the specific implementation method is as follows Figure 2 As shown:

[0113] Using the method described in step 2, each point cloud is rendered into 9 images. Using the method described in step 3, global and local 2D visual features are extracted, and 3D global and local features are obtained by the method shown in step 4. Next, the trained prompt text is encoded using the method described in step 6. Finally, the anomaly score map and anomaly score are calculated using the following formula.

[0114] Specifically, the anomaly score map is:

[0115]

[0116] in represents a Gaussian filter, Represents the cosine similarity between the abnormal point cloud prompt text and each point in the point cloud.

[0117] The global anomaly score combines global and local anomaly semantics, and the calculation formula is:

[0118]

[0119] in Represents a cosine similarity calculation operation.

[0120] (2) Add RGB information to reasoning. The specific implementation method is as follows: Figure 3 As shown in the figure: the original RGB image is used to extract global and local 2D visual features using the method described in step 3, and 3D global and local features are obtained using the method shown in step 4. Next, the trained prompt text is encoded using the method described in step 6, and finally, the color anomaly score map and color anomaly score are calculated using the following formula.

[0121] Specific color abnormality map:

[0122]

[0123] The color anomaly scores are:

[0124]

[0125] in Represents the anomaly score value of each point in the three-dimensional space.

[0126] The final multimodal anomaly score map and anomaly scores are:

[0127]

[0128]

[0129] Step 10: determine whether the detected object is abnormal according to the abnormality score, and segment the abnormal area of ​​the detected object according to the abnormality score map. The inference results described in step 9 are statistically analyzed to finally obtain the test result. The indicator is , The index is 94.2%. After adding RGB information, The indicator is , The indicator is 96.1%.

[0130] At the same time, a heat map is drawn through the anomaly score graph, and it is observed that the abnormal area can be well segmented in the heat map.

[0131] The present invention is based on a 2D multimodal large model to perform anomaly detection on 3D point clouds. With the help of the powerful generalization ability of multimodal CLIP, the method realizes point cloud-based 3D zero-sample anomaly detection and multimodal zero-sample anomaly detection through the steps of obtaining auxiliary point cloud data sets, multi-view rendering of point clouds, characterization of 2D global and local information, characterization of 3D global and local information, hybrid representation learning, text encoding, loss function calculation, framework training, forward reasoning, determination and segmentation of abnormal areas. The entire embodiment is based on Figure 1 The process shown in , implements zero-shot anomaly detection. Figure 2 , Figure 3 This is the point cloud-based zero-shot detection reasoning process and the zero-shot multimodal zero-shot anomaly detection reasoning process of this method. After training, this method can use the anomaly score to determine whether the object under test is abnormal, and use the anomaly score map to segment the abnormal area, providing a guarantee for subsequent anomaly detection in industrial production processes.

[0132] Corresponding to the aforementioned embodiment of a 3D zero-sample anomaly detection method based on a 2D multimodal large model, the present invention also provides an embodiment of a 3D zero-sample anomaly detection device based on a 2D multimodal large model.

[0133] See also Figure 4 A 3D zero-sample anomaly detection device based on a 2D multimodal large model provided in an embodiment of the present invention includes a memory and one or more processors, wherein the memory stores executable code, and when the processor executes the executable code, it is used to implement a 3D zero-sample anomaly detection method based on a 2D multimodal large model in the above embodiment.

[0134] The embodiment of a 3D zero-sample anomaly detection device based on a 2D multimodal large model provided by the present invention can be applied to any device with data processing capabilities, and the device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented through software, or through hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor of any device with data processing capabilities in which it is located reading the corresponding computer program instructions in the non-volatile memory into the internal memory for execution. From a hardware perspective, if Figure 4 As shown, it is a hardware structure diagram of a 3D zero-sample anomaly detection device based on a 2D multimodal large model provided by the present invention, in which any device with data processing capability is located, except Figure 4 In addition to the processor, memory, network interface, and non-volatile memory shown, any device with data processing capabilities in which the apparatus in the embodiments is located may also include other hardware, generally based on the actual functions of the device with data processing capabilities, which will not be described in detail.

[0135] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.

[0136] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiment described above is only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of the present invention. Ordinary technicians in this field can understand and implement it without paying creative work.

[0137] An embodiment of the present invention further provides a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, a 3D zero-sample anomaly detection method based on a 2D multimodal large model in the above embodiment is implemented.

[0138] The computer-readable storage medium may be an internal storage unit of any device with data processing capability described in any of the aforementioned embodiments, such as a hard disk or a memory. The computer-readable storage medium may also be an external storage device of any device with data processing capability, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. equipped on the device. Furthermore, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capability. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capability, and may also be used to temporarily store data that has been output or is to be output.

[0139] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the 3D zero-sample anomaly detection method based on a 2D multimodal large model.

[0140] The above embodiments are used to illustrate the present invention rather than to limit the present invention. Any modification and change made to the present invention within the spirit of the present invention and the protection scope of the claims shall fall within the protection scope of the present invention.

Claims

1. A 3D zero-sample anomaly detection method based on a 2D multimodal large model, characterized in that: The method comprises the following steps: Step 1, obtaining an auxiliary point cloud data set: obtaining an auxiliary point cloud data set for anomaly detection; Step 2, multi-view rendering of point cloud: render each point cloud and its true value into a two-dimensional image from each corresponding viewpoint, and save the original 3D information; Step 3, 2D global and local information representation: For each point cloud, use the visual encoder of the multimodal large model CLIP to extract 2D global and local information representation from its rendered image; Step 4, representation of 3D global and local information: for each point cloud, after obtaining its 2D global and local representations, combine the 2D global representation into a 3D global representation; based on the 2D local representation and the occlusion relationship between the original points in each view after rendering, obtain the 3D local representation; Step 5, hybrid representation learning: design learnable prompt texts for normal point clouds and abnormal point clouds; and encode them using a CLIP-based text encoder; Step 6, calculation of loss function: the total loss function is composed of the sum of 3D global loss function, 3D local loss function, 2D global loss function and 2D local loss function Step 7, training of the zero-shot anomaly detection framework: During the training process, the original parameters of the CLIP model are frozen, and the learnable prompt text is optimized by minimizing the hybrid loss function; Step 8, the reasoning process of detection: after completing the training, perform zero-shot reasoning based on point cloud, or multimodal 3D reasoning with RGB color information added; judge the anomaly based on the anomaly score obtained by reasoning, and segment the abnormal area based on the anomaly score map.

2. The 3D zero-sample anomaly detection method based on a 2D multimodal large model according to claim 1, characterized in that: The rendering method in step 2 is specifically as follows: For the point cloud dataset containing the point cloud and its true value obtained in step 1, the point cloud and its corresponding true value are rendered from different perspectives to obtain the corresponding two-dimensional rendering image, where abnormal pixels are marked as 1 and normal pixels are marked as 0.

3. The 3D zero-sample anomaly detection method based on a 2D multimodal large model according to claim 1, characterized in that: In step 4, the 3D global representation and local representation are specifically: (4.1) 3D global representation: For i point cloud, through K 2D global representations are combined into a 3D global representation; (4.2) 3D local representation: For i point cloud, through K The 2D local representation and the occlusion relationship of the original points in each view after rendering are used to obtain the 3D local representation. The occlusion relationship between the points is expressed as i Point cloud j The point is k Whether it is visible in the rendering, 1 for visible and 0 for invisible.

4. The 3D zero-sample anomaly detection method based on a 2D multimodal large model according to claim 3 is characterized in that: In the step 4, based on the encoded global feature vector, the 2D global features of the rendering image of each point cloud are combined through the rendering relationship to obtain the 3D global features; The 2D local features of the rendering of each point cloud are combined to obtain a preliminary 3D local representation; the final 3D local representation is obtained through the 2D local representation and the occlusion relationship of the original points in each view after rendering.

5. The 3D zero-sample anomaly detection method based on a 2D multimodal large model according to claim 1, characterized in that: In the step 5, the learnable prompt words are mixed representation learning, and the learnable prompt texts of the normal point cloud and the learnable prompt texts of the abnormal point cloud are defined respectively; and the learnable vectors of the normal text and the abnormal text are updated during the training process.

6. The 3D zero-sample anomaly detection method based on a 2D multimodal large model according to claim 1, characterized in that: The loss function in step 6 is specifically: (1) A 3D global loss function that represents the difference between the text representation and the global representation of the rendered image at each viewpoint through cross entropy; (2) 3D local loss function, based on Dice Loss to penalize incorrect abnormal region segmentation; (3) 2D global loss function, which represents the difference between the text representation and each 2D global representation through cross entropy; (4) 2D local loss function, based on Dice Loss and Focal Loss to alleviate the category imbalance problem.

7. The 3D zero-sample anomaly detection method based on a 2D multimodal large model according to claim 1, characterized in that: The specific calculation process of the anomaly score graph and the anomaly score in the two reasoning processes in step 8 is: (1) Zero-shot reasoning based on point cloud: After obtaining 2D and 3D local features from the input point cloud, the cosine similarity is calculated with the encoded text to obtain an anomaly score map, and then a global anomaly score that combines global and local anomaly semantics is obtained; (2) Multimodal 3D reasoning with RGB color information: When color information is available, the 2D global and local information is directly extracted from the 2D color image to obtain 2D features, thereby fusing the color information. The 2D features are then projected back into 3D space to calculate the color anomaly map and the color anomaly score.

8. The 3D zero-sample anomaly detection method based on a 2D multimodal large model according to claim 1, characterized in that: The abnormality determination and abnormal area segmentation in step 8 specifically refers to determining whether the detected object has an abnormality according to the abnormality score, and segmenting the abnormal area of ​​the detected object according to the abnormality score map.

9. A 3D zero-sample anomaly detection device based on a 2D multimodal large model, comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that: When the processor executes the executable code, a 3D zero-sample anomaly detection method based on a 2D multimodal large model as described in any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, a 3D zero-sample anomaly detection method based on a 2D multimodal large model as described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Zero sample point cloud anomaly detection method and system considering prompt learning

    CN118052809A

  • Zero sample point cloud anomaly detection method and system considering multi-view projection

    CN118608447A

  • Point cloud anomaly detection method and device and medium

    CN119090887A

  • Three-dimensional object part segmentation using a machine learning model

    WO2024097470A1

Cited By

  • Industrial zero sample anomaly detection method and system based on cross-modal prompt learning

    CN120726400A

  • Industrial zero-shot anomaly detection method and system based on cross-modal prompt learning

    CN120726400B

  • 3D zero sample anomaly detection method based on 2D multi-mode large model

    CN120783112A