Abnormality detection method and device based on self-sensing fine tuning, equipment and medium

Through self-perception fine-tuning anomaly detection method, sketch anomaly mask is generated and prompt encoding is performed, target anomaly mask is optimized, and the performance degradation of the large model in the downstream field of abnormal detection is solved, achieving high accuracy and robust abnormal detection.

CN120298775APending Publication Date: 2025-07-11TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510358672.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing large models have a significant decline in performance when applied in downstream fields of abnormality detection, which fails to effectively solve the challenges of different abnormal forms in industrial product images and lacks high generalization capabilities.

Method used

By obtaining image features, segmenting prompt information and learnable features, sketch exception masks are generated and prompt encoding are performed, and decoding is combined with external prompt information and learnable features, target exception masks are generated, and the abnormality detection model is optimized using the self-perception fine-tuning mechanism.

Benefits of technology

It improves the accuracy and robustness of the anomaly detection model, can accurately segment abnormal areas in complex morphology, and is suitable for multiple industrial scenarios, enhancing the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298775A_ABST
    Figure CN120298775A_ABST
Patent Text Reader

Abstract

The invention provides an anomaly detection method and device based on self-sensing fine tuning, equipment and a medium, and aims to solve the problem of performance reduction when a model in the related technology is applied to anomaly detection. The method comprises the following steps: acquiring image features corresponding to an input image, segmenting prompt features corresponding to prompt information, and acquiring learnable features; decoding is carried out according to the image features, the prompt features and the learnable features, a sketch anomaly mask is obtained, and the sketch anomaly mask represents preliminary anomaly region prediction and is used for providing anomaly detection guidance; performing prompt coding on the sketch exception mask to obtain external prompt information; and decoding is carried out according to the external prompt information, the image features, the prompt features and the learnable features to obtain a target abnormal mask, and the target abnormal mask represents an abnormal detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and particularly to an anomaly detection method, apparatus, device, and medium based on self-aware fine-tuning. Background Art

[0002] Anomaly segmentation technology is dedicated to automatically identifying and isolating defective regions in industrial product images, which is crucial for improving the efficiency of the production process and ensuring the final quality of products. However, due to the wide variety of industrial products and diverse anomaly forms in the real world, the anomaly segmentation task faces huge challenges.

[0003] Existing large models are mainly trained on natural image datasets, but when applied to downstream fields of anomaly detection, there is usually a significant performance degradation problem. Models in related technologies usually directly apply the base model to the anomaly segmentation task, ignoring the key issue of domain shift between the pre-training dataset and the downstream task.

[0004] Therefore, developing an anomaly segmentation model with high generalization ability has become an important goal in this research field. Summary of the Invention

[0005] To overcome the problems in the related technologies, the present disclosure provides an anomaly detection method, apparatus, device, and medium based on self-aware fine-tuning. The technical solutions of the present disclosure are as follows:

[0006] According to the first aspect of the embodiments of the present disclosure, an anomaly detection method based on self-aware fine-tuning is provided, including:

[0007] Obtain the image features corresponding to the input image, the hint features corresponding to the segmentation hint information, and obtain learnable features;

[0008] Decode according to the image features, the hint features, and the learnable features to obtain a sketch anomaly mask, where the sketch anomaly mask represents a preliminary anomaly region prediction for providing anomaly detection guidance;

[0009] Perform hint encoding on the sketch anomaly mask to obtain external hint information;

[0010] Decode according to the external hint information, the image features, the hint features, and the learnable features to obtain a target anomaly mask, where the target anomaly mask represents the anomaly detection result.

[0011] Optionally, the hint feature includes a sparse embedding representing points or boxes, a dense embedding representing rough segmentation, and a position embedding representing positions; decoding is performed based on the external hint information, the image feature, the hint feature, and the learnable feature to obtain a target anomaly mask, including:

[0012] Replacing the dense embedding in the hint feature with the external hint information to obtain an external hint feature;

[0013] Obtaining the target anomaly mask through the external hint feature, the image feature, and the learnable feature.

[0014] Optionally, decoding is performed based on the image feature, the hint feature, and the learnable feature to obtain a sketch anomaly mask, including:

[0015] Obtaining a first processed representation through the image feature and the dense embedding;

[0016] Obtaining a second processed representation through the sparse embedding and the learnable feature;

[0017] Processing the first processed representation and the second processed representation using a bidirectional transformer to obtain a first completed representation and a second completed representation;

[0018] Embedding the first completed representation and the second completed representation into the input image through a cross-attention mechanism to obtain an embedded representation;

[0019] Multiplying the embedded representation by the upsampled first completed representation to obtain the sketch anomaly mask.

[0020] Optionally, obtaining the target anomaly mask through the external hint feature, the image feature, and the learnable feature, including:

[0021] Obtaining a third processed representation through the image feature and the external hint information in the external hint feature;

[0022] Obtaining a fourth processed representation through the sparse embedding and the learnable feature;

[0023] Processing the third processed representation and the fourth processed representation using a bidirectional transformer to obtain a third completed representation and a fourth completed representation;

[0024] Embedding the third completed representation and the fourth completed representation into the input image through a cross-attention mechanism to obtain an embedded representation;

[0025] Multiplying the embedded representation by the upsampled third completed representation to obtain the target anomaly mask.

[0026] Optionally, the first processing representation and the second processing representation are processed using a bidirectional transformer to obtain a first completion representation and a second completion representation, including:

[0027] Obtain a target similarity matrix of the input image; the target similarity matrix is used to improve the quality of anomaly segmentation;

[0028] Obtain a first relationship representation through the target similarity matrix and the initial output of the bidirectional transformer;

[0029] Obtain the target output of the bidirectional transformer through the first relationship representation and the initial output of the bidirectional transformer.

[0030] Optionally, obtaining a target similarity matrix of the input image includes:

[0031] Calculate the image features to obtain a similarity measurement result;

[0032] Process the similarity measurement result to obtain a preliminary similarity matrix;

[0033] Optimize the preliminary similarity matrix through a preset threshold to obtain the target similarity matrix of the input image.

[0034] Optionally, the method is applied to an anomaly detection network, and the anomaly detection network is determined according to the following steps:

[0035] Input the sample image into the preliminary anomaly detection network to obtain a sample sketch anomaly mask and a sample target anomaly mask;

[0036] Determine a first cross-entropy loss and a first Dice loss according to the sample sketch anomaly mask and the true anomaly mask of the sample image;

[0037] Determine a second cross-entropy loss and a second Dice loss according to the sample target anomaly mask and the true anomaly mask of the sample image;

[0038] Determine the total loss through the first cross-entropy loss, the first Dice loss, the second cross-entropy loss, and the second Dice loss;

[0039] Adjust the parameters of the fine-tuning module in the preliminary anomaly detection network through the total loss, and obtain the anomaly detection network when the adjustment end condition is satisfied.

[0040] According to the second aspect of the embodiments of the present disclosure, there is provided an anomaly detection device based on self-aware fine-tuning, including:

[0041] An acquisition module, configured to acquire image features corresponding to an input image, hint features corresponding to segmentation hint information, and acquire learnable features;

[0042] A decoding module, configured to decode according to the image features, the hint features, and the learnable features to obtain a sketch anomaly mask, where the sketch anomaly mask represents a preliminary anomaly region prediction for providing anomaly detection guidance;

[0043] An encoding module, configured to perform hint encoding on the sketch anomaly mask to obtain external hint information;

[0044] A detection module, configured to decode according to the external hint information, the image features, the hint features, and the learnable features to obtain a target anomaly mask, where the target anomaly mask represents an anomaly detection result.

[0045] According to a third aspect of the embodiments of the present disclosure, an electronic device is provided, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps of the anomaly detection method based on self-aware fine-tuning as described in the first aspect are implemented.

[0046] According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the steps of the anomaly detection method based on self-aware fine-tuning as described in the first aspect are implemented.

[0047] According to a fifth aspect of the embodiments of the present disclosure, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps of the anomaly detection method based on self-aware fine-tuning as described in the first aspect are implemented.

[0048] The present disclosure can obtain image features, prompt features, and learnable features, and generate a sketch anomaly mask through decoding. It can make a rough estimate through a preliminary sketch, providing a starting point for subsequent refined anomaly detection. The sketch anomaly mask is encoded to obtain external prompt information, and further decoded in combination with the image features, prompt features, and learnable features, so that the preliminary sketch can be refined into a target anomaly mask, thereby improving the accuracy of anomaly detection. By predicting the anomaly region through a rough sketch and then further adjusting it in combination with the prompt information and image features, the model can learn how to infer a more accurate anomaly region from the preliminary anomaly mask during the training process. Through the fine-tuning method proposed in the present disclosure, the model can handle anomaly detection tasks in different fields without relying on a large amount of labeled data, so that it can be widely applied to multiple different industrial scenarios and products in practical applications, improving the generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings required for the description of the embodiments of the present disclosure will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0050] Figure 1 FIG. is a schematic diagram of the steps of an anomaly detection method based on self-perception fine-tuning shown in the embodiments of the present disclosure;

[0051] Figure 2 FIG. is a schematic diagram of the structure of an anomaly detection network shown in the embodiments of the present disclosure;

[0052] Figure 3 FIG. is a schematic diagram of the working process of a decoder shown in the embodiments of the present disclosure;

[0053] Figure 4 FIG. is a schematic diagram of the working process of a visual relationship perception adapter shown in the embodiments of the present disclosure;

[0054] Figure 5 FIG. is a block diagram of an anomaly detection device based on self-perception fine-tuning shown in the embodiments of the present disclosure;

[0055] Figure 6 FIG. is a schematic diagram of an electronic device proposed in the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] Next, the technical solutions in the embodiments of the present disclosure will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present disclosure.

[0057] The terms "first", "second", etc. in the specification and claims of the present disclosure are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same category, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.

[0058] To solve the problems existing in the related art, the present disclosure provides an anomaly detection method based on self-perception fine-tuning. The method improves the accuracy, robustness, and generalization ability of the anomaly detection model, can accurately segment complex-shaped anomaly regions, and realizes efficient and automatic anomaly detection in different industrial scenarios.

[0059] Figure 1 is a schematic diagram of the steps of an anomaly detection method based on self-perception fine-tuning shown in the embodiments of the present disclosure. According to Figure 1 shown, the method may specifically include the following steps:

[0060] Step S11: Obtain the image features corresponding to the input image, the hint features corresponding to the segmentation hint information, and obtain learnable features.

[0061] The image features of the input image contain the detailed information in the image and can be used for subsequent detection of anomaly regions. The detailed information describes various aspects of the input image, such as shape, texture, color, edges, etc. The image features not only include the detailed information of each local region in the input image, but also can include the global structure and semantic information of the input image. A specific image encoder can be used to obtain the image features of the input image, and specifically, it can be carried out through the following formula:

[0062] e i mg = E img (I)

[0063] where, e img represents the image features of the input image, and E imgAn image encoder is denoted as, and an input image is denoted as I.

[0064] The hint features related to the image content are obtained through the segmentation hint information. The hint features may include the marks provided by the user or the manually specified areas to help the model understand the potential abnormal areas in the input image. For example, if the user marks a possible abnormal area in a certain image, the hint encoder will convert these hint information into features to assist the model in abnormal detection. A specific hint encoder can be used to obtain the hint features of the segmentation hint information, which can be specifically calculated through the following formula:

[0065] {e sparse ,e dense ,e pos} = E prompt (P)

[0066] where {e sparse ,e dense ,e pos} represents the hint features of the obtained segmentation hint information, E prompt represents the hint encoder, and P represents the segmentation hint information.

[0067] Learnable embeddings refer to the feature vectors that are obtained through learning during the model training process and can effectively represent the input information. Specifically, learnable embeddings are parameters that can be dynamically adjusted through training and can help the model better understand and represent the key features in the input data in a specific task. The learnable features in this disclosure include the IoU embedding e iou and the mask embedding e mask . The IoU embedding e iou helps the model capture the overlapping information between different regions in the input image. In abnormal detection, the IoU embedding helps understand the relationship between the abnormal region and other parts of the image and provides a reference for the subsequent generation of the abnormal mask. The mask embedding e mask is used to represent which regions in the input image are marked as abnormal or non-abnormal, and can provide more effective mask information for the abnormal detection model according to the features learned during the training process.

[0068] Step S12: Decode according to the image features, the hint features, and the learnable features to obtain a sketch abnormal mask, and the sketch abnormal mask represents the preliminary abnormal region prediction for providing abnormal detection guidance.

[0069] The sketch anomaly mask represents the possible abnormal regions in the image. It is a rough region prediction that provides a preliminary guidance for anomaly detection. The purpose is to provide a basic framework for subsequent refinement steps through a self-aware fine-tuning mechanism. The meaning of the self-aware fine-tuning mechanism is a strategy that allows the model to self-adjust and optimize according to its own preliminary prediction results. Specifically, the model not only relies on external annotation or hint information, but also uses its own preliminary prediction results for reverse adjustment to optimize the final segmentation effect. Thus, anomaly detection is more accurate and reliable in practical applications.

[0070] The rough sketch of the anomaly mask (i.e., the sketch anomaly mask) can be generated by the sketch decoder. Specifically, the sketch anomaly mask is obtained through the following formula:

[0071] m draft = D draft (e img , e sparse , e dense , e pos )

[0072] where D draft represents the sketch decoder, and e img represents the image feature, and {e sparse , e dense , e pos} represents the hint feature.

[0073] Step S13: Perform hint encoding on the sketch anomaly mask to obtain external hint information.

[0074] After obtaining the sketch anomaly mask through the sketch decoder, input the sketch anomaly mask into the hint encoder for hint encoding. After being processed by the hint encoder, external hint information is output. Thus, based on the information in the sketch anomaly mask, the model's perception of the abnormal region is optimized, enabling the model to process the hint information more precisely.

[0075] The external hint information is used to provide more guiding additional information for the generation process of the target anomaly mask, enabling the subsequent segmentation process to capture the abnormal region more precisely.

[0076] Step S14: Decode according to the external hint information, the image feature, the hint feature, and the learnable feature to obtain the target anomaly mask, and the target anomaly mask represents the anomaly detection result.

[0077] After obtaining the external prompt information, the external prompt information, the image features, the prompt features, and the learnable features are input into a mask decoder for decoding to generate a target anomaly mask. By combining the external prompt information corresponding to the sketch anomaly mask, the target anomaly mask is regenerated through decoding, realizing the refinement and optimization of the sketch anomaly mask. The finally obtained target anomaly mask can more accurately identify the anomaly regions in the input image.

[0078] Based on the target anomaly mask, a fine anomaly region segmentation result is obtained to ensure the accuracy of detection.

[0079] Adopting the embodiments of the present disclosure, the segmentation result is gradually refined through two decoding stages of sketch anomaly mask generation and target anomaly mask generation, ensuring the improvement of the accuracy of anomaly region segmentation. Through the staged decoding process, the complex anomaly detection task is effectively decomposed into multiple subtasks, and the generation of the initial sketch and the subsequent refinement process cooperate with each other, enabling higher efficiency in terms of computation. By adding learnable features and segmentation prompt information during the decoding process, the model can not only identify anomalies based on the image content but also better understand specific anomaly patterns in the image.

[0080] In practical applications, the anomaly detection method based on self-perception fine-tuning of this embodiment can be implemented through an anomaly detection network. Figure 2 It is a schematic structural diagram of an anomaly detection network shown in the embodiments of the present disclosure. As Figure 2 shown, the present disclosure includes an image encoder, a prompt encoder, a sketch decoder, and a mask decoder. Among them, the sketch decoder is a newly added decoder in the present disclosure. In the conventional technology, when detecting the anomaly regions of an input image, usually only the mask decoder is used for processing. The sketch decoder has the same structure as the original mask decoder and performs the same processing steps. And it is necessary to initialize the sketch decoder with the pre-trained weights of the mask decoder. The main purpose of initializing the sketch decoder with the pre-trained weights is to improve the training efficiency and performance of the model. Specifically, by borrowing the knowledge already learned by the mask decoder, it helps the sketch decoder to perform anomaly perception and segmentation more quickly.

[0081] In step S11, the image features and the prompt features can be obtained through the Figure 2 image encoder and the prompt encoder in

[0082] In step S12, the sketch decoder in Figure 2 is used to decode according to the image features, the prompt features, and the learnable features to obtain a sketch anomaly mask.

[0083] Since the sketch anomaly mask is the preliminary detection result of the anomaly region, it is necessary for the mask decoder to refine it based on the sketch anomaly mask to obtain the target anomaly mask.

[0084] First, step S13 needs to be executed to encode the sketch anomaly mask using the Figure 2 prompt encoder in it to obtain the external prompt information corresponding to the sketch anomaly mask.

[0085] Subsequently, step S14 is executed to use the Figure 2 mask decoder in it to combine the external prompt information, image features, and prompt features to obtain the target anomaly mask of the input image.

[0086] Among them, in an optional embodiment, the prompt features include sparse embeddings representing points or boxes, dense embeddings representing rough segmentations, and position embeddings representing positions; decoding according to the external prompt information, the image features, the prompt features, and the learnable features to obtain the target anomaly mask includes: replacing the dense embedding in the prompt features with the external prompt information to obtain the external prompt features; obtaining the target anomaly mask through the external prompt features, the image features, and the learnable features.

[0087] The prompt features include {e sparse , e dense , e pos}, where e sparse represents the sparse embedding, e dense represents the dense embedding, and e pos represents the position embedding.

[0088] The sparse embedding is mainly used to represent specific points or boxes in the image, such as clicks or annotations by users in the image; it can provide the rough position of the anomaly region, enabling the model to have a preliminary positioning, thus providing a direction for anomaly segmentation.

[0089] The dense embedding represents the rough segmentation information of the image. Through the dense embedding, the model can obtain a more detailed regional distribution, which helps to guide the model to perform a detailed segmentation of the anomaly region. The role of the dense embedding is to provide more features of the image region, thereby helping the model to consider a wider context information when performing anomaly detection.

[0090] The position embedding is used to indicate the spatial information of specific positions in the image. Through the position embedding, the model can understand the relative spatial relationship of each region in the image, thereby enhancing the model's perception ability of the internal structure of the image and finally achieving accurate positioning of the anomaly region.

[0091] The external prompt information can be expressed as e draft, which is the processing result obtained by inputting the sketch anomaly mask into the hint encoder and is derived from the preliminary prediction of the sketch anomaly mask. The external hint information can serve as an optimization guidance for the abnormal area and provide more instructive feedback for the subsequent decoding process.

[0092] By replacing the dense embedding in the hint feature with the external hint information, the external hint feature can be obtained. Thus, the information input during the decoding process of the model is updated, and the segmentation of the abnormal area is further optimized by utilizing the additional information provided by the external hint. Especially in the case of high image complexity, the external hint information can endow the model with stronger anomaly perception ability.

[0093] Specifically, the original hint feature is represented as {e sparse , e dense , e pos}. Then, the external hint feature after using e draft to replace e dense can be represented as {e sparse , e draft , e pos}. The external hint feature and the image feature are input into the mask decoder for decoding. The decoding process can be represented by the following formula:

[0094] m refine = D refine (ei mg , e sparse , ed raft , e pos )

[0095] where D refine represents the mask decoder, e img represents the image feature, and {e sparse , e draft , e pos} represents the external hint feature.

[0096] By combining the external hint information, the target anomaly mask obtained by the mask decoder is a finely optimized anomaly segmentation result. The target anomaly mask accurately marks the abnormal area in the image and provides the final anomaly detection result.

[0097] Adopting the embodiments of the present disclosure, by replacing the dense embedding with the external hint information, the model can utilize the external feedback to further optimize the segmentation of the abnormal area, enhance the adaptability of the model in complex environments, and this replacement mechanism effectively enhances the robustness of the model to cope with different hint information and visual inputs. The external hint information provides optimization guidance for the abnormal area, enabling the model to not only more accurately locate the abnormal area but also reduce misjudgment and missed judgment phenomena, improving the quality of the segmentation result.

[0098] Among them, in an optional embodiment, decoding is performed according to the image feature, the hint feature, and the learnable feature to obtain a sketch anomaly mask, including: obtaining a first processed representation through the image feature and the dense embedding; obtaining the second processed representation through the sparse embedding and the learnable feature; processing the first processed representation and the second processed representation using a bidirectional transformer to obtain a first completed representation and a second completed representation; embedding the first completed representation and the second completed representation into the input image through a cross-attention mechanism to obtain an embedded representation; multiplying the embedded representation by the upsampled first completed representation to obtain the sketch anomaly mask.

[0099] Combining the image feature and the dense embedding to obtain a first processed representation, so that the image can be initially encoded through the image content and the rough segmentation information to obtain a representation containing the main structural information of the image. The first processed representation can be denoted as t img , and specifically can be obtained through the following formula:

[0100] t img =e img +e dense

[0101] Among them, e img represents the image feature, and e dense represents the dense embedding.

[0102] Combining the sparse embedding and the learnable feature to obtain a second processed representation, and the second processed representation contains the potential positions of the abnormal regions in the image, especially the region information provided by sparse hints (such as manually marked points or boxes). The second processed representation can be denoted as t out , and specifically can be obtained through the following formula:

[0103] t out =[e sparse ,e iou ,e mask

[0104] Among them, e sparse represents the sparse embedding, e iou represents the IoU embedding, e mask represents the mask embedding, and [] represents the concatenation of embeddings.

[0105] ​The bidirectional Transformer can effectively capture the interrelationships and context information between features through the self-attention mechanism. After obtaining the first processed representation and the second processed representation, the first processed representation and the second processed representation are respectively input into the Transformer network. After being processed by the self-attention mechanism, the first completed representation and the second completed representation are obtained. The first completed representation and the second completed representation contain more context information and global information, providing more accurate features for subsequent sketch anomaly mask generation.

[0106] The processing process of the bidirectional Transformer for the first processed representation and the second processed representation can be expressed by the following formula:

[0107]

[0108] Where, represents the first completed representation, represents the second completed representation, and TW1, TW2 represent different bidirectional transformers.

[0109] After obtaining the first completed representation and the second completed representation, the first completed representation and the second completed representation are embedded into the input image through the cross-attention mechanism to obtain the embedded representation. Among them, the role of the cross-attention mechanism is to combine features from different sources (for example, the first completed representation and the second completed representation) with the input image, enabling the model to further adjust and optimize the detection of abnormal regions through the relationship between image features and prompt information, which helps to enhance the interaction between the image and the prompt features, enabling the model to better focus on the key abnormal regions in the image.

[0110] Specifically, the first completed representation and the second completed representation can be embedded into the input image through the following formula:

[0111]

[0112] Where, represents the embedded representation, C-Att represents the cross-attention embedded into the image, represents the first completed representation, represents the second completed representation.

[0113] After obtaining the embedded representation, the embedded representation is multiplied by the upsampled first completed representation to obtain the final sketch anomaly mask. Upsampling refers to the process of restoring lower-resolution features to higher resolution, enabling the anomaly mask to be refined in the spatial dimension. Multiplying with the upsampled first completed representation can effectively combine the global information from the first completed representation with the local information from the second completed representation to generate a more accurate anomaly mask. Specifically, it can be executed according to the following formula:

[0114]

[0115] Among them, m draft represents the sketch anomaly mask, represents the embedded representation, and UP represents upsampling, represents the first completed representation.

[0116] The finally obtained sketch anomaly mask is a preliminary prediction of the abnormal area. The sketch anomaly mask represents the preliminary location of the potential abnormal area in the image. It provides a rough guiding framework for subsequent anomaly detection and helps the model further optimize the recognition of the abnormal area.

[0117] By adopting the embodiments of the present disclosure, through the combination of image features, prompt features and learnable features, feature fusion is performed using a Transformer model, which improves the perception ability of the abnormal area. Among them, through the cross-attention mechanism, the model can capture more complex inter-feature dependencies, so as to effectively combine information from different sources, thereby optimizing the anomaly detection effect. The upsampling step further improves the spatial resolution of the mask and enhances the detection accuracy.

[0118] Among them, in an optional embodiment, the target anomaly mask is obtained through the external prompt feature, the image feature, and the learnable feature, including: obtaining a third processed representation through the external prompt information in the image feature and the external prompt feature; obtaining the fourth processed representation through the sparse embedding and the learnable feature; processing the third processed representation and the fourth processed representation using a bidirectional transformer to obtain a third completed representation and a fourth completed representation; embedding the third completed representation and the fourth completed representation into the input image through a cross-attention mechanism to obtain an embedded representation; multiplying the embedded representation by the upsampled third completed representation to obtain the target anomaly mask.

[0119] The external prompt information is the result of inputting the sketch anomaly mask into the prompt encoder. The external prompt information provides the model with a preliminary perception of the abnormal area, while the image feature provides a perception of the overall content of the input image. By combining the external prompt information and the external prompt information, the model can initially understand the global context of the abnormal area. The third processed representation can be determined by the following formula:

[0120] t img = e img + e draft

[0121] Among them, t img represents the third processed representation, e img represents the image feature, e draftIndicates external prompt information.

[0122] By combining sparse embeddings and learnable features, a fourth processed representation is generated. The sparse embeddings provide location information about the abnormal regions, while the learnable features fine-tune these locations through a neural network, thereby further enhancing the ability to accurately capture abnormal regions. The second processed representation is consistent with the fourth processed representation. The fourth processed representation can be determined by the following formula:

[0123] t out = [e sparse , e iou , e mask

[0124] where, e sparse represents the sparse embedding, e iou represents the IoU embedding, e mask represents the mask embedding, and [] represents the concatenation of embeddings.

[0125] After obtaining the third processed representation and the fourth processed representation, a bidirectional Transformer is used to further process the third processed representation and the fourth processed representation. Specifically, the bidirectional Transformer models the input features (the third processed representation and the fourth processed representation) through the self-attention mechanism to capture the long-term dependencies between the features, and enhances the understanding of the context information in a bidirectional (from left to right and from right to left) manner.

[0126] After being processed by the bidirectional Transformer, the third processed representation and the fourth processed representation will be transformed into the third completed representation and the fourth completed representation. The third completed representation and the fourth completed representation contain richer context information and provide refined feature representations for the subsequent generation of the target abnormal mask.

[0127] The processing process of the bidirectional Transformer for the third processed representation and the third processed representation can be expressed by the following formula:

[0128]

[0129] where, represents the third completed representation, represents the fourth completed representation, and TW1, TW2 represent different bidirectional transformers.

[0130] ​Subsequently, the first completion representation and the second completion representation are embedded into the input image through a cross-attention mechanism to obtain an embedded representation. The cross-attention mechanism can fuse information from different sources (such as image features, external prompt features, learnable features, etc.), enabling the model to accurately locate the abnormal region based on the interaction between the image features and the external prompt information. The fused embedded representation can improve the model's attention to the abnormal region, thereby obtaining an accurate target abnormal mask.

[0131] Specifically, the third completion representation and the fourth completion representation can be embedded into the input image through the following formula:

[0132]

[0133] where denotes the embedded representation, C-Att denotes the cross-attention embedded into the image, denotes the third completion representation, denotes the fourth completion representation.

[0134] The upsampled third completion representation is multiplied by the embedded representation to obtain the target abnormal mask. Multiplying the upsampled third completion representation by the embedded representation can combine the image features with the generated abnormal region features, further enhancing the ability to distinguish the abnormal region, and finally obtaining the target abnormal mask. Specifically, it can be executed according to the following formula:

[0135]

[0136] where m refine denotes the target abnormal mask, denotes the embedded representation, UP denotes upsampling, denotes the third completion representation.

[0137] Adopting the embodiments of the present disclosure, the external prompt features further refine the prompt information. The external prompt information obtained by using the sketch abnormal mask can replace the dense embedding in the original prompt features, providing more accurate and rich abnormal region location information. By performing prompt encoding on the sketch abnormal mask, the model further refines the abnormal mask through the external prompt information and other features to obtain the target abnormal mask. The target abnormal mask not only provides a more accurate position of the abnormal region but also can characterize important features such as the shape and size of the abnormality. The transformation process from the sketch abnormal mask to the target abnormal mask helps the model shift from a rough abnormal region prediction to a more accurate and refined abnormal detection, improving the accuracy of the detection result.

[0138] The output of the bidirectional transformer can be processed by a visual relation perception adapter, which enhances the decoder's performance by integrating visual relations between different regions into the decoding process. The relation perception adapter can utilize the object similarity matrix to improve the quality of anomaly segmentation.

[0139] The working process of the visual relation perception adapter is the same for the sketch decoder and the mask decoder. The following embodiments are the working process of the visual relation perception adapter for the sketch decoder. For the working process of the visual relation perception adapter for the mask decoder, reference can be made to the working process for the sketch decoder.

[0140] Among them, in an optional embodiment, the first processed representation and the second processed representation are processed using a bidirectional transformer to obtain a first completed representation and a second completed representation, including: obtaining the object similarity matrix of the input image; the object similarity matrix is used to improve the quality of anomaly segmentation; through the object similarity matrix and the initial output of the bidirectional transformer, a first relation representation is obtained; through the first relation representation and the initial output of the bidirectional transformer, the target output of the bidirectional transformer is obtained.

[0141] The object similarity matrix is a structure that describes the similarity or relationship between different regions in an image. It can capture the interaction between local and global features in the image and reflect the similarity between different regions. By calculating the similarity between different image regions, it helps the model understand which regions may belong to the same category and which regions may have a higher probability of being anomalies. Combining the object similarity matrix with the bidirectional transformer can enhance the quality of anomaly segmentation in the image and improve the ability to distinguish anomaly regions.

[0142] The object similarity matrix provides similarity information between image regions and can affect how relational modeling is performed on different regions of the image. By combining the object similarity matrix with the initial output of the bidirectional transformer, the transformer model can utilize this relational information to understand the dependencies between different regions. At this time, the model can identify which regions may be related and which regions have stronger mutual relationships.

[0143] The first relation representation is a feature representation of the inter-region relationship obtained through transformer modeling based on the similarity matrix and the initial output of the bidirectional transformer. The first relation representation helps the model learn the subtle differences between regions in the image, especially the regions that may represent anomalies, by focusing on the correlation between different regions of the image. The first relation representation can be obtained through the following formula:

[0144]

[0145] Among them, represents the first relationship representation, A represents the target similarity matrix, represents the initial output of the bidirectional transformer.

[0146] After obtaining the first relationship representation, through further processing of the first relationship representation and the initial output of the bidirectional transformer, the target output of the bidirectional transformer is obtained. The target output can be determined by the following formula:

[0147]

[0148] Among them, represents the target output, represents the initial output, represents the first relationship representation, β represents the weight coefficient, and β can be dynamically adjusted during the training process.

[0149] Due to the target similarity matrix bringing additional information about the relationships between regions, the target output of the transformer pays more attention to the regions with abnormal features, thus being able to accurately capture the abnormal regions in the image. The target output represents the final result of anomaly detection after the in-depth processing of the bidirectional transformer. These completed representations integrate the global and local information of image features, relationships between regions, and other input features, and can accurately reflect the specific location and morphological features of the abnormal regions.

[0150] Adopting the embodiments of the present disclosure, by combining the target similarity matrix and the bidirectional processing mechanism of the bidirectional transformer, not only the model's understanding of the similarity between regions is improved, but also the recognition ability of abnormal regions is further enhanced. The first relationship representation and the final target output through this processing flow have significantly improved the segmentation quality and accuracy of anomaly detection. Especially when dealing with complex or tiny anomalies, this method can effectively improve the robustness and accuracy of detection.

[0151] Before using the visual relationship perception adapter to integrate the visual relationships between different regions into the decoding process to enhance the mask decoder, it is first necessary to evaluate the visual relationships within the image, and then use the relationship perception adapter to introduce the visual relationship knowledge into the mask decoder. Specifically, the following embodiments can be used for visual relationship evaluation:

[0152] Among them, in an optional embodiment, obtaining the target similarity matrix of the input image includes:

[0153] Calculate the image features to obtain a similarity metric result; process the similarity metric result to obtain a preliminary similarity matrix; optimize the preliminary similarity matrix through a preset threshold to obtain the target similarity matrix of the input image.

[0154] In the feature space, cosine similarity can be used to measure the similarity degree of these regions. Specifically, the visual relationship between different regions of the input image can be measured through the following formula to obtain a similarity metric result:

[0155] S = cosine(e img , e img )

[0156] where cosine represents cosine similarity, and e img represents the image feature.

[0157] After obtaining the similarity metric result, since the similarity of a region with itself is the strongest, but this similarity is not helpful for anomaly detection and may instead interfere with the model's learning of the relationships between different regions, the diagonal elements (the similarity of each region with itself) are set to negative infinity, and then the softmax function is applied to the obtained similarity matrix A. The softmax function is usually used to convert a set of numerical values into a probability distribution, and the sum of all elements is 1.

[0158] Specifically, it can be implemented through the following formula:

[0159] A = softmax(S)

[0160] where S represents the similarity metric result.

[0161] After obtaining the similarity matrix A standardized by softmax, the relationship matrix is further optimized through a threshold mechanism to remove the region relationships with too low similarity, so that the final similarity matrix only retains the most important and effective relationships. Specifically, the similarity matrix A can be optimized according to the following formula to obtain the target similarity matrix A * :

[0162]

[0163] where the threshold is controlled by the hyperparameter α divided by the feature dimension d. d is the dimension of the feature. As the feature dimension increases, there may be more possible region relationships in the similarity matrix. Therefore, the threshold needs to be adjusted according to the feature dimension so that the model can optimize the relationship matrix according to different feature spaces.

[0164] By adopting the embodiments of the present disclosure, the relationship between different regions in the input image is measured by cosine similarity. Subsequently, the softmax function is used to normalize the similarity measurement result to obtain a similarity matrix, and the threshold mechanism is used to filter out the most important relationships to obtain the target similarity matrix, ensuring that the model can focus on the key regional relationships in the input image and improving the accuracy and robustness of anomaly detection.

[0165] Among them, in an optional embodiment, the method is applied to an anomaly detection network, and the anomaly detection network is determined according to the following steps: inputting a sample image into the preliminary anomaly detection network to obtain a sample sketch anomaly mask and a sample target anomaly mask; determining a first cross-entropy loss and a first Dice loss according to the sample sketch anomaly mask and the true anomaly mask of the sample image; determining a second cross-entropy loss and a second Dice loss according to the sample target anomaly mask and the true anomaly mask of the sample image; determining a total loss through the first cross-entropy loss, the first Dice loss, the second cross-entropy loss and the second Dice loss; adjusting the parameters of the fine-tuning module in the preliminary anomaly detection network through the total loss, and obtaining the anomaly detection network when the adjustment end condition is met.

[0166] The preliminary anomaly detection network includes a preliminary sketch decoder and a preliminary mask decoder. After inputting the sample image into the preliminary anomaly detection network, a sample sketch anomaly mask output by the preliminary sketch decoder and a sample target anomaly mask output by the preliminary mask decoder are obtained.

[0167] The cross-entropy loss measures the difference between the predicted probability distribution and the actual label distribution. Specifically, in an image segmentation or anomaly detection task, the cross-entropy loss is used to measure the difference between the classification of each pixel predicted by the decoder and the true label. For the sample sketch anomaly mask output by the preliminary sketch decoder and the sample target anomaly mask output by the preliminary mask decoder, the cross-entropy loss can calculate the difference between them and the true anomaly mask.

[0168] The Dice loss measures the performance of the model by calculating the overlap degree between the predicted region and the true region. The range of the Dice coefficient is from 0 (completely non-overlapping) to 1 (completely overlapping). The advantage of the Dice loss is that it is applicable to scenarios with imbalanced classes. In anomaly detection or image segmentation, the anomaly region may be relatively small, and the Dice loss can improve the model's detection ability for small-region anomalies by emphasizing the overlapping region in this case.

[0169] The total loss can be determined by the following formula:

[0170] L = L CE (m draft , Y) + Ldice (m draft ,Y)+L CE (m refine ,Y)+L dice (m refine ,Y)

[0171] Among them, L CE (m draft ,Y) represents the first cross-entropy loss, and L dice (m draft ,Y) represents the first Dice loss, and L CE (m refine ,Y) represents the second cross-entropy loss, and L dice (m refine ,Y) represents the second Dice loss.

[0172] By combining the cross-entropy loss and the Dice loss, it is possible to simultaneously focus on the rough regional structure and the precise pixel-level segmentation.

[0173] The fine-tuning module is a learnable PEFT module (Prompt-based Efficient Fine-Tuning Module). The PEFT module can introduce learnable prompts and, by fine-tuning a small number of parameters (i.e., the prompts), rather than comprehensively fine-tuning the entire network, achieve fine-tuning of the image encoder and the prompt encoder to improve the performance of the model in anomaly detection and segmentation tasks.

[0174] Fine-tune the parameters included in the PEFT module based on the total loss to further optimize the performance of the model in practical tasks.

[0175] By adopting the embodiments of the present disclosure and introducing the total loss (composed of the cross-entropy loss and the Dice loss), the training process of the network can be effectively balanced among various tasks. The total loss not only optimizes the generation process of the rough sketch but also optimizes the generation process of the fine segmentation mask, ensuring that the performance of the network in the overall task is more balanced and precise.

[0176] Figure 3 is a schematic diagram of the working process of a decoder shown in the embodiments of the present disclosure. Refer to Figure 3 as shown:

[0177] In the sketch stage: Use a sketch decoder to process the first processing representation and the second processing representation. The first processing representation and the second processing representation pass through two bidirectional transformers in sequence. A visual relationship perception adapter is connected behind each bidirectional transformer. The visual relationship perception adapter processes the initial output of the bidirectional transformer in combination with the first processing representation to obtain the target output of the bidirectional transformer, and inputs the target output into the next bidirectional transformer or uses it as the final result. After using the bidirectional transformer and the visual relationship perception adapter to process the first processing representation and the second processing representation to obtain the first processing result and the second processing result, a sketch anomaly mask is obtained through upsampling and the cross-attention mechanism.

[0178] Process the obtained sketch anomaly mask using a prompt encoder to obtain external prompt information, and use the external prompt information to replace the dense embedding in the prompt features to obtain a third processing representation and a fourth processing representation that is consistent with the second processing representation.

[0179] In the refinement stage: Use a mask decoder to process the third processing representation and the fourth processing representation. The third processing representation and the fourth processing representation pass through two bidirectional transformers in sequence. A visual relationship perception adapter is connected behind each bidirectional transformer. The visual relationship perception adapter processes the initial output of the bidirectional transformer in combination with the third processing representation to obtain the target output of the bidirectional transformer, and inputs the target output into the next bidirectional transformer or uses it as the final result. After using the bidirectional transformer and the visual relationship perception adapter to process the first processing representation and the second processing representation to obtain the first processing result and the second processing result, a target anomaly mask is obtained through upsampling and the cross-attention mechanism.

[0180] Figure 4 It is a schematic diagram of the working process of a visual relationship perception adapter shown in an embodiment of the present disclosure. According to Figure 4As shown: the target similarity matrix is obtained through the image features included in the first processing representation; the first relationship representation is obtained through the target similarity matrix and the initial output of the bidirectional transformer; finally, the weighted sum of the first relationship representation and the initial output of the bidirectional transformer is used to obtain the target output of the bidirectional transformer. For the target output of the first bidirectional transformer, it is used as the input to the next bidirectional transformer, and for the target output of the last bidirectional transformer, it is used as the final result of the two bidirectional transformers.

[0181] Based on the same technical concept, an embodiment of the present disclosure proposes an anomaly detection device based on self-aware fine-tuning. Figure 5 is a block diagram of an anomaly detection device based on self-aware fine-tuning shown in an embodiment of the present disclosure. According to Figure 5 as shown, the device includes:

[0182] An acquisition module 510, configured to acquire the image features corresponding to the input image, the hint features corresponding to the segmentation hint information, and acquire learnable features;

[0183] A decoding module 520, configured to decode according to the image features, the hint features, and the learnable features to obtain a sketch anomaly mask, where the sketch anomaly mask represents a preliminary anomaly region prediction for providing anomaly detection guidance;

[0184] An encoding module 530, configured to perform hint encoding on the sketch anomaly mask to obtain external hint information;

[0185] A detection module 540, configured to decode according to the external hint information, the image features, the hint features, and the learnable features to obtain a target anomaly mask, where the target anomaly mask represents the anomaly detection result.

[0186] Optionally, the hint features include sparse embeddings representing points or boxes, dense embeddings representing rough segmentation, and position embeddings representing positions; the detection module is specifically configured to perform:

[0187] Replace the dense embedding in the hint features with the external hint information to obtain external hint features;

[0188] Obtain the target anomaly mask through the external hint features, the image features, and the learnable features.

[0189] Optionally, the decoding module is specifically configured to perform:

[0190] Obtain a first processing representation through the image features and the dense embedding;

[0191] Obtain the second processed representation through the sparse embedding and the learnable features;

[0192] Process the first processed representation and the second processed representation using a bidirectional transformer to obtain a first completed representation and a second completed representation;

[0193] Embed the first completed representation and the second completed representation into the input image through a cross-attention mechanism to obtain an embedded representation;

[0194] Multiply the embedded representation by the upsampled first completed representation to obtain the sketch anomaly mask.

[0195] Optionally, the detection module is specifically configured to perform:

[0196] Obtain a third processed representation through the external prompt information in the image features and the external prompt features;

[0197] Obtain the fourth processed representation through the sparse embedding and the learnable features;

[0198] Process the third processed representation and the fourth processed representation using a bidirectional transformer to obtain a third completed representation and a fourth completed representation;

[0199] Embed the third completed representation and the fourth completed representation into the input image through a cross-attention mechanism to obtain an embedded representation;

[0200] Multiply the embedded representation by the upsampled third completed representation to obtain the target anomaly mask.

[0201] Optionally, the decoding module is specifically configured to perform:

[0202] Obtain a target similarity matrix of the input image; the target similarity matrix is used to improve the quality of anomaly segmentation;

[0203] Obtain a first relationship representation through the target similarity matrix and the initial output of the bidirectional transformer;

[0204] Obtain the target output of the bidirectional transformer through the first relationship representation and the initial output of the bidirectional transformer.

[0205] Optionally, the decoding module is specifically configured to perform:

[0206] Calculate the image features to obtain a similarity metric result;

[0207] Process the similarity measurement results to obtain a preliminary similarity matrix;

[0208] Optimize the preliminary similarity matrix through a preset threshold to obtain the target similarity matrix of the input image.

[0209] Optionally, the device includes a network training module for training an anomaly detection network; specifically, the network training module is configured to execute:

[0210] Input a sample image into the preliminary anomaly detection network to obtain a sample sketch anomaly mask and a sample target anomaly mask;

[0211] Determine a first cross-entropy loss and a first Dice loss according to the sample sketch anomaly mask and the true anomaly mask of the sample image;

[0212] Determine a second cross-entropy loss and a second Dice loss according to the sample target anomaly mask and the true anomaly mask of the sample image;

[0213] Determine a total loss through the first cross-entropy loss, the first Dice loss, the second cross-entropy loss, and the second Dice loss;

[0214] Adjust the parameters of the fine-tuning module in the preliminary anomaly detection network through the total loss, and obtain the anomaly detection network when the adjustment end condition is met.

[0215] An embodiment of the present disclosure also provides an electronic device. Refer to Figure 6 , Figure 6 is a schematic diagram of an electronic device proposed by an embodiment of the present disclosure. As Figure 6 shown, the electronic device 600 includes: a memory 610 and a processor 620. The memory 610 and the processor 620 are communicatively connected via a bus. A computer program is stored in the memory 610, and the computer program can run on the processor 620 to implement the steps in the anomaly detection method based on self-aware fine-tuning disclosed in the embodiments of the present disclosure.

[0216] An embodiment of the present disclosure also provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the steps in the anomaly detection method based on self-aware fine-tuning disclosed in the embodiments of the present disclosure are implemented.

[0217] An embodiment of the present disclosure also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps in the anomaly detection method based on self-aware fine-tuning disclosed in the embodiments of the present disclosure are implemented.

[0218] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other.

[0219] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, an apparatus, or a computer program product. Therefore, the embodiments of the present disclosure can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0220] The embodiments of the present disclosure are described with reference to the flowcharts and / or block diagrams of methods, apparatuses, electronic devices, and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0221] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0222] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0223] Although some embodiments of the present disclosure have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be interpreted to include the preferred embodiments as well as all changes and modifications that fall within the scope of the embodiments of the present disclosure.

[0224] The above provides a detailed introduction to an anomaly detection method, device, equipment, and medium based on self-perceived fine-tuning. Specific examples are used in this article to elaborate on the principles and implementation methods of the present disclosure. The description of the above embodiments is only used to help understand the method and its core idea of the present disclosure; at the same time, for those of ordinary skill in the art, according to the idea of the present disclosure, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be construed as a limitation to the present disclosure.

Claims

1. An anomaly detection method based on self-perception fine-tuning, characterized in that, Including: Obtaining image features corresponding to the input image, hint features corresponding to the segmentation hint information, and obtaining learnable features; Decoding according to the image features, the hint features, and the learnable features to obtain a sketch anomaly mask, where the sketch anomaly mask represents a preliminary anomaly region prediction for providing anomaly detection guidance; Performing hint encoding on the sketch anomaly mask to obtain external hint information; Decoding according to the external hint information, the image features, the hint features, and the learnable features to obtain a target anomaly mask, where the target anomaly mask represents the anomaly detection result.

2. The method according to claim 1, wherein The hint features include sparse embeddings representing points or boxes, dense embeddings representing rough segmentation, and position embeddings representing positions; Decoding according to the external hint information, the image features, the hint features, and the learnable features to obtain a target anomaly mask, including: Replacing the dense embedding in the hint features with the external hint information to obtain external hint features; Obtaining the target anomaly mask through the external hint features, the image features, and the learnable features.

3. The method according to claim 2, wherein Decoding according to the image features, the hint features, and the learnable features to obtain a sketch anomaly mask, including: Obtaining a first processed representation through the image features and the dense embedding; Obtaining the second processed representation through the sparse embedding and the learnable features; Processing the first processed representation and the second processed representation using a bidirectional transformer to obtain a first completed representation and a second completed representation; Embedding the first completed representation and the second completed representation into the input image through a cross-attention mechanism to obtain an embedded representation; Multiplying the embedded representation by the upsampled first completed representation to obtain the sketch anomaly mask.

4. The method according to claim 2, characterized in that, Obtaining the target anomaly mask through the external hint features, the image features, and the learnable features, including: Obtaining a third processed representation through the image features and the external hint information in the external hint features; Obtaining the fourth processed representation through the sparse embedding and the learnable features; Processing the third processed representation and the fourth processed representation using a bidirectional transformer to obtain a third completed representation and a fourth completed representation; Embedding the third completed representation and the fourth completed representation into the input image through a cross-attention mechanism to obtain an embedded representation; Multiplying the embedded representation by the upsampled third completed representation to obtain the target anomaly mask.

5. The method according to claim 3, characterized in that Processing the first processed representation and the second processed representation using a bidirectional transformer to obtain a first completed representation and a second completed representation, including: Obtaining a target similarity matrix of the input image; the target similarity matrix is used to improve the quality of anomaly segmentation; Obtaining a first relational representation through the target similarity matrix and the initial output of the bidirectional transformer; Obtain the target output of the bidirectional Transformer through the first relationship representation and the initial output of the bidirectional Transformer.

6. The method according to claim 5, wherein obtaining the target similarity matrix of the input image comprises: Calculate the image features to obtain a similarity measurement result; Process the similarity measurement result to obtain a preliminary similarity matrix; Optimize the preliminary similarity matrix through a preset threshold to obtain the target similarity matrix of the input image.

7. The method according to claim 1, wherein The method is applied to an anomaly detection network, and the anomaly detection network is determined according to the following steps: Input the sample image into the preliminary anomaly detection network to obtain a sample sketch anomaly mask and a sample target anomaly mask; Determine the first cross-entropy loss and the first Dice loss according to the sample sketch anomaly mask and the true anomaly mask of the sample image; Determine the second cross-entropy loss and the second Dice loss according to the sample target anomaly mask and the true anomaly mask of the sample image; Determine the total loss through the first cross-entropy loss, the first Dice loss, the second cross-entropy loss, and the second Dice loss; Adjust the parameters of the fine-tuning module in the preliminary anomaly detection network through the total loss, and obtain the anomaly detection network when the adjustment end condition is met.

8. An anomaly detection device based on self-perception fine-tuning, characterized in that, Comprising: An acquisition module for acquiring the image features corresponding to the input image, the hint features corresponding to the segmentation hint information, and acquiring learnable features; A decoding module for decoding according to the image features, the hint features, and the learnable features to obtain a sketch anomaly mask, where the sketch anomaly mask represents a preliminary anomaly region prediction for providing anomaly detection guidance; An encoding module for performing hint encoding on the sketch anomaly mask to obtain external hint information; A detection module for decoding according to the external hint information, the image features, the hint features, and the learnable features to obtain a target anomaly mask, where the target anomaly mask represents the anomaly detection result.

9. An electronic device, characterized in that, Comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps of the anomaly detection method based on self-perception fine-tuning according to any one of claims 1-7 are implemented.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, the steps of the anomaly detection method based on self-perception fine-tuning according to any one of claims 1-7 are implemented.