Abnormal sample generation method based on large model, computer equipment and computer readable storage medium
By combining small models and large models, real abnormal image samples similar to normal samples are generated, which solves the problem of insufficient authenticity of abnormal samples in the prior art and significantly improves the performance of abnormal detection.
Patent Information
- Application Number
- CN202510161336.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-06-03
AI Technical Summary
The abnormal samples generated in the prior art are very different from normal samples and lack authenticity, which leads to problems such as wasting computing resources and limited performance improvement when defining boundaries in the potential space.
Using a combination of small models and large models, we use small models to classify and evaluate image samples by training small models, use large models to generate text exception descriptions, and convert them into abnormal image samples to form an abnormal image sample set, which is used to update and retrain small models to filter out difficult abnormal image samples that are more similar to normal image samples.
The generated anomaly image samples are more realistic and can help the detector learn a clearer boundary between normal and abnormal, thereby significantly improving the performance of the visual anomaly detector.
Smart Images

Figure CN120088555A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of visual detection, and particularly relates to an abnormal sample generation method, a computer device, and a computer-readable storage medium based on a large model. Background Art
[0002] Unsupervised anomaly detection often has difficulty in distinguishing normal samples from abnormal samples, and there are various types of anomalies. Usually, it is difficult to obtain satisfactory results because it is difficult to learn the effective boundary between normal samples and abnormal samples. Therefore, how to improve the performance of anomaly detection is an important challenge for unsupervised anomaly detection. Unsupervised visual anomaly detection is one of the most important fields in anomaly detection, including industrial product defect detection, financial fraud detection, and video surveillance, etc. In any scenario, abnormal samples are scarce and unpredictable. Therefore, this challenge is particularly obvious in visual data. There is no available prior information about anomalies in these fields, and only normal samples are available for reference. Even images belonging to the same category may have significant differences, which further exacerbates the complexity of anomaly detection.
[0003] In recent years, to solve this problem, more and more research has focused on generating anomalies, and some anomaly detection methods based on generation have emerged to help detectors more effectively distinguish normal samples from abnormal ones. However, since the generated anomalies usually come from random factors, they often lack authenticity and deviate greatly from the real scenario. Especially in image samples, even samples of the same category may show significant differences, which will blur the boundary between normal samples and anomalies in the latent space. In addition, randomly generated anomalies usually have limited support in constructing an effective boundary. The excessive randomness in the generated anomalies makes the definition of the boundary lack effectiveness because most of them are quite different from normal samples and far from the boundary, resulting in waste of computing resources and limited performance improvement. Summary of the Invention
[0004] The purpose of the present invention is to provide an abnormal sample generation method, a computer device, and a computer-readable storage medium based on a large model to solve the problem that the generated abnormal samples in the prior art are quite different from normal samples and lack authenticity.
[0005] To solve the above technical problems, the present invention provides an abnormal sample generation method based on a large model, and the method includes:
[0006] 1) Training a small model using the current image sample set;
[0007] 2) Use the trained large model to generate a text normal description that can represent the normal distribution for normal image samples according to the classification results of the trained small model, and combine the text normal description and the set representation to generate a text abnormal description for the abnormal requirements; then convert the text abnormal description into abnormal image samples; group all the obtained abnormal image samples into an abnormal image sample set;
[0008] 3) Use the abnormal image sample set to update the current image sample set, retrain the small model using the updated current image sample set, and after retraining, use the small model to screen out the difficult abnormal image samples in the updated current image sample set. The difficult abnormal image samples are abnormal image samples that are relatively similar to normal image samples.
[0009] Further, the method further includes: 4) After retraining the small model, re - execute steps 2) - 3) for iteration, and integrate the difficult abnormal image samples screened out by the retrained small model each time after retraining the small model.
[0010] Further, before re - executing step 2) each time, first fine - tune the large model following direct preference optimization, and then use the fine - tuned large model to generate the text abnormal description.
[0011] Further, the way to use the abnormal image sample set to update the current image sample set in step 3) is: use the abnormal image sample set to increase the number of abnormal images in the current image sample set, or use the abnormal image sample set to replace the simple abnormal image samples in the current image sample set. The simple abnormal image samples are abnormal image samples that are significantly deviated from normal image samples.
[0012] Further, the small model is Deep SAD, and when retraining the small model in step 3), the training objective is to minimize the volume of the hypersphere enclosing the data in the latent space.
[0013] Further, the small model uses the Mahalanobis distance to classify the current image sample set.
[0014] Further, use a text - to - image model to convert the text abnormal description into abnormal image samples.
[0015] Further, the current image sample set in step 1) is initially a normal image sample set.
[0016] To solve the above - mentioned technical problems, the present invention also provides a computer device, including a processor, and the processor executes a computer program to implement the steps of the abnormal sample generation method based on a large model introduced above.
[0017] To solve the above technical problems, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-described method for generating abnormal samples based on a large model are implemented.
[0018] The present invention is an innovative invention, and its beneficial effects are as follows: The present invention combines a small model and a large model to generate more realistic abnormal image samples, called difficult abnormal image samples, that is, abnormal image samples that are relatively similar to normal image samples. The main functions of the small model include two aspects, namely, classifying abnormal image samples and normal image samples, and evaluating abnormal image samples to evaluate their similarity to normal image samples. The main function of the large model is to learn knowledge related to abnormalities to generate text abnormal descriptions. Based on this, the overall solution of the present invention is as follows: First, use the small model to screen out normal image samples, then use the large model to combine the text normal description representing the normal distribution and the set prompts to generate text abnormal descriptions, and then convert the text abnormal descriptions into abnormal image samples to obtain an abnormal image sample set. Then, use the abnormal image sample set to update the training set for training the small model to retrain the small model. After training, use the small model to screen out the required difficult abnormal image samples. The present invention uses the knowledge learned by the large model to retrain the small model so that the small model can learn a clearer boundary between normal and abnormal, and screen out more accurate "difficult abnormal image samples", which are more realistic abnormal image samples. Description of the Drawings
[0019] Figure 1 is a schematic diagram of unsupervised anomaly detection of the present invention and anomaly detection using the knowledge of LLMs;
[0020] Figure 2 is the overall flowchart of KAA of the present invention;
[0021] Figure 3 is a schematic diagram of different numbers of initial anomalies and their corresponding average precisions on the Oxford-102 dataset of the present invention;
[0022] Figure 4 is a schematic diagram of the anomalies generated by KAA of the present invention on the Oxford-102 dataset and their corresponding iterative processes;
[0023] Figure 5a 、 Figure 5b are respectively visualizations of the unsupervised method of learning the latent space and KAA of the present invention;
[0024] Figure 6a 、 Figure 6bThey are the Loss and AUC result graphs of the best baseline method ReContrast and the combination of ReContrast and KKA, respectively. Detailed implementation manners
[0025] The present invention combines a large model and a small model to generate abnormal image samples that are relatively similar to normal image samples. The overall process is as follows: First, the small model is used to screen normal image samples. Then, the large model is used to generate a normal text description representing the normal distribution based on the normal image samples, and an abnormal text description is generated by combining the normal text description and a prompt. Next, the abnormal text description is converted into abnormal image samples to obtain an abnormal image sample set. Furthermore, the abnormal image sample set is used to update the current image sample set to retrain the small model. After retraining, the small model is used to screen out difficult abnormal image samples. This method uses the large model to learn and extract indicators related to abnormalities to generate abnormal text descriptions, and retrains the small model based on the knowledge learned by the large model so that the small model can learn a clearer boundary between normal and abnormal, which helps to extract more accurate difficult abnormal image samples. Based on this, an abnormal sample generation method based on a large model, a computer device, and a computer-readable storage medium of the present invention can be realized.
[0026] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0027] An embodiment of an abnormal sample generation method based on a large model:
[0028] An abnormal sample generation method based on a large model of the present invention mainly aims to generate difficult abnormal samples with a high similarity to normal samples. These difficult abnormal samples are authentic and will be used to train a visual anomaly detector in the future, which can significantly improve the performance of the visual anomaly detector. The visual anomaly detector can be a product defect detector, a financial fraud detector, or a video surveillance detector. Among them, the product defect detector processes the input product image to identify the defect position, type, etc. in the product image. The input of the financial fraud detector is the user's financial transaction data, account behavior data, etc., and the output is whether there is an abnormal transaction and the possible types of fraud behaviors (such as forged identity, abnormal transfer, etc.). The input of the video surveillance detector is real-time or recorded video data, and the output is whether there is an abnormal event (such as intrusion, fighting, fire, etc.) and the time and location of the abnormal event.
[0029] Such as Figure 1As shown on the left side, in the context of unsupervised visual anomaly detection, for example, an airplane as an anomaly sample deviates significantly from normal samples such as cats. Therefore, the detector can more easily distinguish between anomaly and normal samples, and in this embodiment, such anomalies are referred to as "simple anomalies". On the contrary, when an anomaly sample (such as a dog) is similar to a normal sample, the detector will encounter difficulties in learning the effective boundary between normal samples and anomalies. Such anomalies are called "difficult anomalies". Difficult anomalies are usually more difficult to obtain due to their high similarity to normal samples, but have a greater impact on the detector's learning of the boundary. As Figure 1 As shown on the right side, with the help of the key knowledge enhanced by LLMs in the present invention, the detector can learn a clearer boundary between normal samples and anomalies.
[0030] The meanings of the symbols introduced in this embodiment are shown in Table 1.
[0031] Table 1
[0032]
[0033] The present invention adopts the idea of collaborative work of large and small models to generate difficult anomaly image samples. The small model mainly has two functions. The first is to classify image samples to distinguish whether they are anomaly image samples or normal image samples, and the second is to evaluate anomaly image samples to evaluate their similarity with normal image samples. The main function of the large model is to learn knowledge related to anomalies to generate text anomaly descriptions. Hereinafter, this method is called Key Knowledge Augmentation (KKA), and the entire process of KKA is summarized in Figure 2 Visually demonstrated, Table 2 shows the corresponding coded process.
[0034] Table 2
[0035]
[0036] The specific process is as follows:
[0037] Step 1, obtain the normal image sample set X n i , and use the normal image sample set as the current image sample set to train the small model.
[0038] Here, the normal image sample set is used first because of the sparsity and diversity of anomalies, so the normal image sample set is directly used to train the small model.
[0039] The small model uses an unsupervised learning-based visual anomaly detection method, including a feature extractor and an anomaly detector, which distinguishes normal samples from abnormal samples by learning the embedding distribution of normal samples. In this embodiment, the small model selects Deep SAD, including a feature extractor φ and a detector, which has the characteristic of binary classification. During testing, it distinguishes normal image samples and abnormal image samples by calculating the distance learned by the feature extractor φ in the latent space. The distance can be the Mahalanobis distance, thereby obtaining the score s of the abnormal image sample. k :
[0040]
[0041] where x n ∈X n is a normal image sample, z ∈ Z is the output of the feature extractor φ, which is the representation of the sample x in the embedding space, k represents different clusters in the latent space, μ k and Σ k are the sample mean and the sample covariance matrix respectively.
[0042] When the abnormal image sample deviates significantly from the normal distribution (i.e., a simple abnormal image sample, x ea ), the detector can effectively distinguish normal image samples and abnormal image samples through the above formula. However, when normal image samples and abnormal image samples are relatively similar (i.e., difficult abnormal image samples, i.e., x ha ), the performance of unsupervised anomaly detection is significantly reduced.
[0043] It should be noted that in addition to Deep SAD, the small model can also select other models in the prior art, such as SSD, SimpleNet, ReContrast, etc.
[0044] Step two, training. Generate a text normal description that can represent the normal distribution for the image samples classified as normal according to the classification results of the trained small model, and combine the text normal description and the set representation to generate a text abnormal description for the abnormal requirements; use the trained text-to-image model to convert the text abnormal description into an abnormal image sample; group all the obtained abnormal image samples into an abnormal image sample set.
[0045] In this embodiment, the large model used is LLMs, and the text-to-image model used is Stable Diffusion. Specifically:
[0046] 1) Use text to represent the knowledge in LLMs. The knowledge of normal image samples in the training data can be represented as Select text descriptions that can represent the normal distribution from :
[0047]
[0048] In the formula, K represents the number of normal image clusters, and s i is the anomaly score calculated by formula (1). The sample with the lowest score is the one closest to the normal distribution and can be used as the basis for generating normal descriptions. is the text description of the image that best matches the normal distribution, called the text normal description.
[0049] 2) Generate text anomaly descriptions using different representations according to the text normal description to generate prompts for anomaly requirements For example, Figure 2 in, the text normal description is "Gaura lindheimeri Engelmn.&A.Gray" (a kind of flower), and the prompt for generating anomalies is "Please generate new sentences that describe flowers. The sentences should follow the exact structure and style of the example below: ". The text anomaly descriptions generated by LLMs can be expressed as:
[0050]
[0051] In the formula, θ L generates the current token y based on the previous i - 1th token y i-1 and the prompt, i , represents the generation distribution defined by the large model parameters.
[0052] 3) Generate anomaly image samples through the text anomaly description and the text - to - image model
[0053]
[0054] In the formula, is the anomaly image sample generated corresponding to the text anomaly description . By running formulas (3) and (4) multiple times, a set of anomaly image samples
[0055] Step 3: Obtain a new current image sample set based on the abnormal image sample set and the current image sample set, retrain the small model using the latest current image sample set, and use the retrained small model to screen out the difficult abnormal image samples in the latest current image sample set. Specifically:
[0056] 1) Combine the obtained abnormal image sample set and the normal image sample set as the new current image sample set, and retrain the small model using this latest current image sample set.
[0057] The detector in the small model trained using the normal image sample set and the abnormal image sample set can distinguish normal image samples and abnormal image samples more effectively than unsupervised methods. However, the anomalies generated by formulas (3) and (4) include a large number of simple abnormal image samples because difficult abnormal image samples must satisfy specific distribution conditions and exhibit similarity to normal image samples. To obtain more difficult abnormal image samples use the normal image sample set and the abnormal image sample set to train the confusion evaluator integrated in the small model, which learns a transformation to minimize the volume of the hypersphere enclosing the data in the latent space Z:
[0058]
[0059] where W is the set of weights of the confusion evaluator θ C M is the number of anomalies, c is the predetermined center of the hypersphere, λ represents the weight parameter of the regularization term, and the larger this value, the stronger the constraint on the model complexity. represents the label, 0 or 1. Formula (5) tends to project normal image samples near the center of the hypersphere and anomalies far from the center.
[0060] 2) Calculate the distance of each sample to the center c in the latent space and select those abnormal image samples with a short distance to some normal image samples, that is, screen out the difficult abnormal image samples This process can be expressed as:
[0061]
[0062] where DF() is the distance function. The objective of formula (6) is to find difficult abnormal image samples that are easily confused with normal image samples.
[0063] In Step 4, after retraining the small model, fine-tune the large model, then return to Step 2 to regenerate the abnormal image sample set, and then execute Step 3 to retrain the small model, and use the retrained small model to screen out the difficult abnormal image samples in the latest current image sample set.
[0064] Among them, in order to extract more key knowledge about difficult anomalies from the LLMs, fine-tune the LLMs following Direct Preference Optimization (DPO), as shown in Equation (7). If not fine-tuned, the abnormal image samples generated by the large model each time may not vary much. After fine-tuning the large model, the abnormal image samples generated by the large model each time can be made to vary, and the proportion of the generated difficult abnormal image samples will also increase.
[0065]
[0066] In the formula, σ represents the sigmoid function, θ L ′ represents the fine-tuned LLMs, and θ L represents the original LLMs. After fine-tuning, more key knowledge can be extracted from the LLMs, and more difficult abnormal image samples can be generated in the following way:
[0067]
[0068] In the formula, represents the number of difficult abnormal images.
[0069] In each iteration of updating the abnormal data set, a new abnormal image sample set can be obtained through the newly acquired abnormal image samples each time to obtain a new abnormal image sample set To meet the different requirements of the detector for simple and difficult anomalies in different scenarios, in the process of re-executing Step 3, regarding the link of "obtaining a new current image sample set based on the abnormal image sample set and the current image sample set", it is divided into KKA (increase) and KKA (replace). KKA (increase) means continuously increasing the number of abnormal image samples (gradually increasing the proportion of difficult anomalies), that is, the number of abnormal image samples in the training set for retraining the small model is increased, and some normal image samples are replaced by abnormal image samples; while KKA (replace) keeps the number of abnormal image samples unchanged by continuously replacing simple abnormal image samples with more difficult abnormal image samples, that is, the number of abnormal image samples in the training set for retraining the small model remains unchanged, but some of the simple abnormal image samples are replaced by difficult abnormal image samples.
[0070] Step 5. During the entire loop process from Step 2 to Step 4, integrate the difficult abnormal image samples screened out in each pass of Step 3. The resulting samples are the final samples needed and are relatively realistic abnormal image samples.
[0071] Next, conduct experiments, evaluate the proposed method on multiple datasets, perform ablation studies, hyperparameter analysis, and convergence analysis.
[0072] 1) Dataset and parameter settings.
[0073] Conduct experiments on three different datasets. Each dataset contains class labels for classifying normal and abnormal samples.
[0074] CIFAR-100: The CIFAR-100 image classification dataset includes 20 superclasses, a total of 100 subclasses, with each subclass containing 600 images (500 for training and 100 for testing). Each image is assigned a fine label and a coarse label. In this experiment, one superclass (containing 5 subclasses) is designated as normal samples, while the remaining 19 superclasses (95 subclasses) are regarded as abnormal.
[0075] Oxford-102: This dataset contains 8,189 image-text pairs of flowers, with a total of 102 different classes. In this experiment, one flower class is selected as normal samples, and the remaining 101 classes are regarded as abnormal.
[0076] UCM-Caption: This dataset includes 21 land use image classes, and each image is described by five different sentences. Randomly select two classes as normal samples, and the remaining 19 classes as abnormal.
[0077] Implementation: In this embodiment, mainly focus on the extraction of anomaly-related knowledge and the generation of abnormal images for unsupervised anomaly detection.
[0078] Unsupervised baseline methods can be divided into three categories: (1) methods based on raw data, such as SSD; (2) methods that utilize cross-modal information to generate additional samples, such as CMDA; (3) methods that generate anomalies, such as SimpleNet. The proposed KKA belongs to the third category. For KKA, the learning rate of the feature extractor is set to 0.0001, the scheduler is adjusted at 50 epochs or 40 epochs, the batch size is 32, and the optimizer is Adam. The number of training epochs on the CIFAR-100, Oxford-102, and UCM-Caption datasets are 150, 200, and 300 respectively. Additionally, the generated anomaly datasets are updated 3 times, 2 times, and 2 times on CIFAR-100, Oxford-102, and UCM-Caption respectively.
[0079] 2) Result analysis.
[0080] Table 3 shows the performance of KKA in the addition and replacement modes, as well as the results of unsupervised methods with and without generated samples. KKA demonstrates strong generality and can be integrated into supervised anomaly detection methods (such as SAD) and generation-based anomaly detection methods (such as SimpleNet). Generally, introducing KKA significantly improves the detection performance. For example, on the CIFAR-100 dataset, KKA increases the AUC of SimpleNet from 74.62% to 84.04%, while generating only about 5% of the number of samples generated by SimpleNet. Additionally, the excellent performance of KKA based on SimpleNet indicates that extracting key knowledge from LLMs is more suitable for visual anomaly detection compared to generating anomaly samples through random noise. The results of Recontrast and ReContrast (without pre-training) clearly demonstrate the effectiveness of pre-training. In the absence of anomalies, the pre-trained model can integrate prior knowledge and roughly help the detector understand what is an anomaly. Based on the pre-trained model, Recontrast+KKA further introduces key knowledge for specific anomaly scenarios, bringing additional performance improvements. Overall, the results in Table 3 show that the samples generated by KKA are superior to the samples generated by SimpleNet using random noise, the multi-view samples generated by ReContrast, and the cross-modal information samples generated by CMDA. This also proves that the knowledge extracted from LLMs is more practically relevant for specific anomaly detection scenarios.
[0081] Table 3
[0082]
[0083] 3) Importance of key knowledge.
[0084] Some ablation experiments were conducted to demonstrate the importance of key knowledge (hard anomalies). To eliminate the performance improvement effect that may be brought about by the increase in data volume, the results of the KKA (replacement) mode were shown. KKA (replacement) keeps the number of anomalies fixed while gradually increasing the proportion of hard anomalies in each iteration. The relevant results are shown in Table 4, where (0) represents the initially generated anomalies, and (1), (2), and (3) represent the first, second, and third iterations of the updated anomaly dataset respectively. Higher numbers indicate more hard anomalies.
[0085] Table 4
[0086]
[0087] The relevant results show that as the proportion of hard anomalies increases (1, 2, 3), the performance of the detector improves significantly. For example, on the CIFAR-100 dataset, the AUC increases continuously with each iteration. In contrast, although the performance of the detector on the Oxford-102 and UCM-Caption datasets also improves, due to the difference in data volume, this improvement is not as stable as that on CIFAR-100: CIFAR-100 contains 2,500 normal samples, while Oxford-102 and UCM-Caption only contain 51 and 158 normal samples respectively. Too many hard anomalies may cause the detector to ignore fewer normal samples. At the same time, on the Oxford-102 and UCM-Caption datasets, the AUC of KKA (0) is 92.88% and 72.33% respectively, while after the second iteration, the AUC of the detector on these datasets increases to 94.31% and 74.60% respectively. By increasing the proportion of hard anomalies, the performance of the detector improves significantly, which clearly shows the key role of hard anomalies in anomaly detection.
[0088] Figure 3 The relationship between the number of generated anomalies and the number of hard anomalies was further demonstrated. In Figure 3 the orange dots represent KKA applied to increase hard anomalies, while the pink dots represent the initial number of anomalies (including more easy anomalies). The results show that increasing the initial number of anomalies has little effect on the performance of the detector. In addition, once the number of generated anomalies reaches 400, the performance of the detector tends to level off. Datasets with a higher proportion of hard anomalies can easily outperform those containing a large number of easy anomalies. This shows that hard anomalies contain more meaningful information in anomaly detection.
[0089] 4) Visualization of KKA.
[0090] The iterative process of anomaly generation by KKA on the Oxford-102 dataset is shown, and the results are as Figure 4As shown. In the initial generation stage, the anomalies exhibit a high degree of randomness, including flowers of various colors and shapes extracted from the prior knowledge of LLMs, ensuring the diversity of anomalies. In the first iteration, KKA extracts the key knowledge of LLMs based on the selected difficult anomalies, making the generated anomalies closer to normal samples. Therefore, the anomalies generated in the first time contain some smaller flowers and pink colors, similar to the characteristics of normal samples. In the second iteration, the influence of the key knowledge becomes more obvious in the generated anomalies. Although the anomalies still have distinguishable differences from normal samples, their characteristics are becoming more and more similar to normal samples, with a higher similarity than the anomalies in the first iteration.
[0091] In Figure 5a and Figure 5b the latent spaces learned by the unsupervised method SSD and the proposed KKA are visualized to demonstrate that the proposed method learns a more obvious boundary between normal samples and abnormal samples. As Figure 5a shown, the latent space learned by SSD is more chaotic. In contrast, as Figure 5b shown, the latent space learned by KKA can clearly distinguish most normal samples and abnormal samples because there are obvious distribution differences between blue points and red points. More specifically, although SSD can project some red points into a relatively concentrated area (i.e., easy anomalies), there are still a large number of red points interspersed among blue points (i.e., difficult anomalies).
[0092] 5) Different settings for anomaly detection.
[0093] In real-world scenarios, anomalies are inherently difficult to predict. They may not follow a certain pattern like the samples in the Flower-102 or UCM datasets, where normal samples and abnormal samples usually belong to the same category. Instead, anomalies are more likely to come from two aspects: similar distributions (i.e., the same supercategory but different subcategories) and different distributions (i.e., different supercategories). To further verify the effectiveness of KKA, the CIFAR-100 dataset was reconstructed, selecting one class as normal samples and treating the remaining 99 classes as anomalies. The results are shown in Table 5. Obviously, when only using one subclass as normal samples, the task is relatively easy because the distribution of normal samples is more uniform and the detector can focus more on the distribution of normal samples. Further, the KKA results in Table 1 and Table 4 show a consistent pattern: as the number of iterations increases, the performance of the detector gradually improves. This further emphasizes the important role of difficult anomalies in anomaly detection.
[0094] Table 5
[0095]
[0096] 6) Convergence analysis.
[0097] In Figure 6a , Figure 6b it is shown how the anomalies generated by KKA continuously improve the performance of the best baseline method, ReContrast. It should be noted that all settings of ReContrast and ReContrast with KKA are the same. As Figure 6a shown, with the more realistic anomaly samples generated by KKA, ReContrast can converge faster and more effectively. This is because the generated anomalies are no longer random and unrealistic content, but real-world images with the same "style" as normal samples. This is also why the combination of ReContrast and KKA shows better performance in Figure 6b . In Figure 6b , the AUC of ReContrast and KKA continues to increase over multiple epochs, while the increase in the AUC of ReContrast is relatively small, indicating that the anomaly samples generated by KKA help to accelerate convergence and improve performance.
[0098] An embodiment of a computer device:
[0099] An embodiment of a computer device according to the present invention includes a memory, a processor, an internal bus, a computer program stored in the memory, a communication interface, an input device, and a display screen. The processor and the memory communicate and interact with each other through the internal bus.
[0100] Among them, the processor is used to provide computing and control capabilities and execute the computer program to implement the steps of the method introduced in an embodiment of a method for generating anomaly samples based on a large model according to the present invention. It can be a microprocessor MCU, a programmable logic device FPGA, or other processing devices. The memory can be a non-volatile storage medium or an internal memory. The non-volatile storage medium stores an operating system and a computer program, and the internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0101] An embodiment of a computer-readable storage medium:
[0102] An embodiment of a computer-readable storage medium of the present invention stores a computer program thereon. When the program is executed by a processor, it implements the method for generating abnormal samples based on large models proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM for short), electrically erasable programmable read-only memory (EEPROM for short), erasable programmable read-only memory (EPROM for short), programmable read-only memory (PROM for short), read-only memory (ROM for short), magnetic memory, or flash memory.
[0103] In summary, the present invention first utilizes the extensive prior knowledge of LLMs to generate meaningful and reasonable abnormalities based on normal samples, that is, randomly extracting normal samples and designing prompt words to enable LLMs to generate knowledge related to abnormalities, and then using this knowledge to generate various abnormal images for detector training; then, according to their similarity to normal samples, the generated abnormalities are divided into simple abnormalities and difficult abnormalities; furthermore, the generated abnormalities are iteratively updated to gradually increase the proportion of difficult abnormalities, so that the detector can learn more effective boundaries. Among them, in order to further improve the performance of the detector, a confusion evaluator is trained to identify difficult abnormalities in the generated images. The experimental results show that this method significantly improves the performance of various visual abnormality detectors while maintaining a low generation cost.
[0104] The specific implementation manners are given above, but the present invention is not limited to the described implementation manners. The basic idea of the present invention lies in the above basic solution. For those of ordinary skill in the art, according to the teachings of the present invention, it does not require creative labor to design various deformed models, formulas, and parameters. Changes, modifications, substitutions, and variations made to the implementation manners without departing from the principle and spirit of the present invention still fall within the protection scope of the present invention.
Claims
1. A method for generating abnormal samples based on a large model, characterized in that: The steps include: 1) Use the current image sample set to train a small model; 2) Using the trained large model to generate a textual normal description that can represent the normal distribution for the normal image samples according to the classification results of the trained small model, and generating a textual abnormal description by combining the textual normal description and the set representation to generate the abnormal requirements; then converting the textual abnormal description into an abnormal image sample; and grouping all the obtained abnormal image samples into an abnormal image sample set; 3) Using the abnormal image sample set to update the current image sample set, using the updated current image sample set to retrain the small model, and after retraining, using the small model to screen out difficult abnormal image samples in the updated current image sample set, where the difficult abnormal image samples are abnormal image samples that are relatively similar to normal image samples.
2. The method for generating abnormal samples based on a large model according to claim 1, characterized in that: The method further includes: 4) After retraining the small model, re-execute steps 2) to 3) to iterate, and integrate the difficult abnormal image samples screened out by the retrained small model after each retraining of the small model.
3. The method for generating abnormal samples based on a large model according to claim 2, characterized in that: Before re-executing step 2) each time, the large model is first fine-tuned according to direct preference optimization, and then the fine-tuned large model is used to generate a textual anomaly description.
4. The method for generating abnormal samples based on a large model according to claim 1, characterized in that: In step 3), the method of updating the current image sample set using the abnormal image sample set is: using the abnormal image sample set to increase the number of abnormal images in the current image sample set, or using the abnormal image sample set to replace simple abnormal image samples in the current image sample set, where the simple abnormal image samples are abnormal image samples that deviate significantly from normal image samples.
5. The method for generating abnormal samples based on a large model according to claim 1, characterized in that: The small model is Deep SAD, and when the small model is retrained in step 3), the training goal is to minimize the volume of the hypersphere surrounding the data in the latent space.
6. The method for generating abnormal samples based on a large model according to claim 1, characterized in that: The small model uses Mahalanobis distance to classify the current image sample set.
7. The method for generating abnormal samples based on a large model according to any one of claims 1 to 6, characterized in that: The text-to-image model is used to convert text anomaly descriptions into anomaly image samples.
8. The method for generating abnormal samples based on a large model according to any one of claims 1 to 6, characterized in that: In step 1), the current image sample set is initially a normal image sample set.
9. A computer device comprising a processor, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.