Image generation model training method, sample generation method, and related device

WO2026179551A1PCT designated stage Publication Date: 2026-09-03TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/075383
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-28
Filing Date
2026-01-28
Publication Date
2026-09-03

Smart Images

  • Figure CN2026075383_03092026_PF_FP_ABST
    Figure CN2026075383_03092026_PF_FP_ABST
Patent Text Reader

Abstract

The present application provides an image generation model training method, a sample generation method, and a related device. The image generation model training method comprises: performing image cropping on a sample anomaly image on the basis of object anomaly position information to obtain an anomaly area sub-image; extracting sub-image feature information of the anomaly area sub-image and overall feature information of the sample anomaly image; performing image reconstruction on the sub-image feature information on the basis of object anomaly category information by means of an image generation model to obtain a reconstructed anomaly area sub-image; performing image reconstruction on the overall feature information on the basis of object feature description information, the object anomaly position information, and the object anomaly category information by means of the image generation model to obtain a reconstructed sample anomaly image; and training the image generation model on the basis of the anomaly area sub-image, the reconstructed anomaly area sub-image, the sample anomaly image, and the reconstructed sample anomaly image.
Need to check novelty before this filing date? Find Prior Art

Description

Image generation model training methods, sample generation methods, and related equipment

[0001] Cross-reference of related applications

[0002] This application is based on and claims priority to Chinese Patent Application No. 2025102403560, filed on February 28, 2025, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of computer technology, and in particular to an image generation model training method, a sample generation method, and related equipment. Background Technology

[0004] In industrial quality inspection scenarios, it is often necessary to perform industrial anomaly detection on parts on factory production lines to promptly identify and pinpoint abnormal parts and determine the types of defects, thereby avoiding adverse effects on product quality, safety, and user experience. However, in industrial settings, obtaining a large number of normal parts is easy, but obtaining defective parts is difficult. This lack of anomaly sample images for training makes it impossible to directly use conventional object detection algorithms to detect defects. Therefore, it is necessary to synthesize defects, i.e., generate anomaly sample images. Summary of the Invention

[0005] This application provides an image generation model training method, a sample generation method, and related equipment. The related equipment may include an image generation model training device, a sample generation device, electronic equipment, a computer-readable storage medium, and a computer program product. This allows the image generation model to simultaneously learn the feature information of local defects and the overall image information. This ensures the consistency between the overall abnormal image generated by the trained image generation model and the local abnormal image, thereby enabling the construction of diverse, realistic, and highly aligned abnormal image-mask data sample pairs, which is beneficial to improving the performance of downstream tasks such as anomaly detection, localization, and classification.

[0006] This application provides an image generation model training method, applied to electronic devices, including:

[0007] Acquire training data and an image generation model, wherein the training data includes at least one sample anomalous image, object feature description information in the sample anomalous image, object anomalous location information, and object anomalous category information;

[0008] Based on the abnormal location information of the object, the abnormal sample image is processed by image extraction to obtain an abnormal region sub-image;

[0009] Feature extraction processing is performed on the abnormal region sub-image and the sample abnormal image respectively to obtain the sub-image feature information corresponding to the abnormal region sub-image and the overall feature information corresponding to the sample abnormal image;

[0010] Using the image generation model, based on the object anomaly category information, image reconstruction processing is performed on the feature information of the sub-image to obtain the reconstructed anomaly region sub-image;

[0011] Using the image generation model, based on the object feature description information, the object abnormal location information, and the object abnormal category information, image reconstruction processing is performed on the overall feature information to obtain a reconstructed sample abnormal image;

[0012] Based on the loss information between the abnormal region sub-image and the reconstructed abnormal region sub-image, and the loss information between the sample abnormal image and the reconstructed sample abnormal image, the parameters of the image generation model are adjusted to obtain the trained image generation model.

[0013] This application provides a sample generation method applied to an electronic device, including:

[0014] Obtain object feature description information, object anomaly location information, and object anomaly category information for the target object;

[0015] Based on the object anomaly category information, an anomaly region sub-image is generated using an image generation model to obtain an anomaly region sub-image. The image generation model is trained using the image generation model training method provided in the embodiments of this application.

[0016] Using the image generation model, based on the object feature description information, the object anomaly location information, and the object anomaly category information, an overall image is generated to obtain an overall image of the anomaly.

[0017] Based on the abnormal region sub-image, the abnormal overall image is subjected to abnormal masking processing to obtain an abnormal mask image of the abnormal overall image. The sample pair composed of the abnormal overall image and the abnormal mask image is used to optimize the object anomaly detection model.

[0018] This application provides an image generation model training device, including:

[0019] The first acquisition unit is configured to acquire training data and an image generation model. The training data includes at least one sample abnormal image, object feature description information in the sample abnormal image, object abnormality location information, and object abnormality category information.

[0020] The image extraction unit is configured to perform image extraction processing on the sample abnormal image based on the abnormal position information of the object to obtain an abnormal region sub-image;

[0021] The feature extraction unit is configured to perform feature extraction processing on the abnormal region sub-image and the sample abnormal image respectively, to obtain the sub-image feature information corresponding to the abnormal region sub-image and the overall feature information corresponding to the sample abnormal image;

[0022] The first reconstruction unit is configured to perform image reconstruction processing on the feature information of the sub-image based on the object anomaly category information using the image generation model to obtain the reconstructed anomaly region sub-image.

[0023] The second reconstruction unit is configured to perform image reconstruction processing on the overall feature information based on the object feature description information, the object abnormal location information, and the object abnormal category information through the image generation model, to obtain a reconstructed sample abnormal image;

[0024] The parameter adjustment unit is configured to adjust the parameters of the image generation model based on the loss information between the abnormal region sub-image and the reconstructed abnormal region sub-image, and the loss information between the sample abnormal image and the reconstructed sample abnormal image, so as to obtain the trained image generation model.

[0025] This application provides a sample generation apparatus, including:

[0026] The second acquisition unit is configured to acquire object feature description information, object anomaly location information, and object anomaly category information for the target object.

[0027] The first generation unit is configured to generate a region sub-image based on the object anomaly category information using an image generation model, thereby obtaining an anomaly region sub-image. The image generation model is trained using the image generation model training method provided in the embodiments of this application.

[0028] The second generation unit is configured to generate an overall image based on the object feature description information, the object abnormal location information, and the object abnormal category information through the image generation model, thereby obtaining an abnormal overall image.

[0029] The masking unit is configured to perform anomaly masking processing on the overall abnormal image based on the abnormal region sub-image to obtain an anomaly mask image of the overall abnormal image. The sample pair formed by the overall abnormal image and the anomaly mask image is used to optimize the object anomaly detection model.

[0030] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0031] An electronic device provided in this application includes a processor and a memory. The memory stores multiple instructions, and the processor loads the instructions to execute the steps in the image generation model training method and sample generation method provided in this application.

[0032] This application also provides a computer-readable storage medium storing a computer program thereon, wherein the computer program, when executed by a processor, implements the steps in the image generation model training method and sample generation method provided in this application.

[0033] Furthermore, embodiments of this application also provide a computer program product, including a computer program or instructions, which, when executed by a processor, implement the steps in the image generation model training method and sample generation method provided in embodiments of this application.

[0034] The embodiments of this application have the following beneficial effects:

[0035] By sequentially learning the generation of anomalous region sub-images and the overall anomalous image during training, the image generation model can simultaneously learn the feature information of local anomalies and the overall image information. This ensures the consistency between the overall anomalous image and the local anomalous image generated by the trained image generation model, thereby constructing realistic and highly aligned anomalous image-mask data sample pairs. The method provided in this application can amplify a small amount of labeled anomalous defect data to generate a large amount of labeled anomalous defect data. The generated data has good realism and diversity, and can be directly used for supervised anomaly detection training, significantly improving the performance of downstream anomaly detection tasks, including defect detection, localization, and classification. Attached Figure Description

[0036] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0037] Figure 1a is a schematic diagram of a scenario for the sample generation method provided in an embodiment of this application;

[0038] Figure 1b is a flowchart of the image generation model training method provided in an embodiment of this application;

[0039] Figure 2a is a flowchart of the sample generation method provided in an embodiment of this application;

[0040] Figure 2b is an illustrative diagram of the image generation model training method and sample generation method provided in the embodiments of this application;

[0041] Figure 2c is another flowchart of the image generation model training method and sample generation method provided in the embodiments of this application;

[0042] Figure 2d is another illustrative diagram of the image generation model training method provided in the embodiment of this application;

[0043] Figure 2e is another illustrative diagram of the image generation model training method provided in the embodiment of this application;

[0044] Figure 3a is a schematic diagram of the image generation model training device provided in an embodiment of this application;

[0045] Figure 3b is a schematic diagram of the sample generation device provided in an embodiment of this application;

[0046] Figure 4 is a schematic diagram of the structure of the electronic device provided in an embodiment of this application. Detailed Implementation

[0047] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0048] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0049] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0050] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.

[0051] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.

[0052] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0053] 1) Diffusion Models are generative models that simulate the process of data samples being gradually covered by noise (forward diffusion process), and then learn how to reverse the noise to recover the original data (reverse denoising process).

[0054] 2) U-Net is a deep learning network architecture with an encoder-decoder structure. It combines high-level features with low-level features through skip connections, effectively capturing multi-scale features.

[0055] 3) Abnormal sample images refer to abnormal images of objects (such as images of objects with defects) contained in the training data used for model training. The objects shown in the images can be components, parts, or finished products in industrial production, and the defects of the objects can be manifested as abnormalities in size, material texture, surface finish, or assembly status.

[0056] 4) Object feature description information refers to textual information used to describe the inherent features or shooting characteristics of objects in sample abnormal images. Object feature description information may include object category information (such as component type "gear", "nut" or product type "watch"), and object shooting location information (such as the object's "front", "side", "top view" or other display posture or placement angle in the image, and can be uniquely identified using numbers such as P001, P002, etc.).

[0057] 5) Object anomaly location information refers to the location data used to accurately pinpoint the area where the object anomaly is located in the sample anomaly image. It can be represented by coordinates. For example, a rectangular area containing four values ​​[x_min, y_min, x_max, y_max] is used to define the rectangular area where the anomaly is located in the sample anomaly image. Here, (x_min, y_min) corresponds to the upper left corner of the rectangular area, and (x_max, y_max) corresponds to the lower right corner of the rectangular area.

[0058] 6) Object anomaly category information refers to textual information used to describe the type of object anomaly in the sample anomaly image. It can be a classification description of the nature of the defect. For example, for surface finish defects, the object anomaly category information can be "scratches", "burrs", "burns", etc.

[0059] This application provides an image generation model training method, a sample generation method, and related equipment. The related equipment may include an image generation model training device, a sample generation device, electronic equipment, a computer-readable storage medium, and a computer program product.

[0060] The sample generation device can be integrated into an electronic device, such as a terminal or server.

[0061] It is understood that the sample generation method in this application embodiment can be executed on a terminal, on a server, or jointly by a terminal and a server. The above examples should not be construed as limiting the embodiments of this application.

[0062] In one embodiment, as shown in FIG1a, the sample generation method is executed jointly by a terminal and a server. The sample generation system provided in this application embodiment includes a terminal 10 and a server 11, etc.; the terminal 10 and the server 11 are connected via a network, such as a wired or wireless network, etc., wherein the sample generation device can be integrated into the server.

[0063] Server 11 can be used to: acquire object feature description information, object anomaly location information, and object anomaly category information for a target object; generate anomaly region sub-maps based on the object anomaly category information using an image generation model; generate an overall image based on the object feature description information, object anomaly location information, and object anomaly category information using an image generation model; perform anomaly masking on the overall image based on the anomaly region sub-maps to obtain anomaly mask images of the overall image; and use the sample pairs consisting of the overall image and the anomaly mask images to optimize the object anomaly detection model. Server 11 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0064] Terminal 10 can be used to send object feature description information, object anomaly location information, and object anomaly category information of the target object to server 11, so that server 11 can generate corresponding anomaly region sub-images and anomaly overall images, thereby constructing sample pairs between the overall anomaly image and its anomaly mask image, and using the sample pairs to optimize the object anomaly detection model. Terminal 10 can include mobile phones, vehicle terminals, aircraft, tablet computers, laptops, or personal computers (PCs), etc. A client can also be set on terminal 10, which can be an application client or a browser client, etc.

[0065] The sample generation and other steps in the aforementioned server 11 can also be performed by the terminal 10.

[0066] The image generation model training device can be integrated into an electronic device, such as a terminal or a server. The image generation model training method in this application can be executed on a server, or it can be executed jointly by a terminal and a server; this application does not impose any limitations.

[0067] In one embodiment, an image generation model training method is implemented jointly by a terminal and a server. The image generation model training system provided in this application includes a terminal and a server, which are connected via a network, such as a wired or wireless network. The image generation model training device can be integrated into the server.

[0068] The server can be used to: acquire training data and an image generation model. The training data includes at least one sample anomalous image, object feature description information in the sample anomalous image, object anomalous location information, and object anomalous category information; perform image extraction processing on the sample anomalous image based on the object anomalous location information to obtain an anomalous region sub-image; perform feature extraction processing on the anomalous region sub-image and the sample anomalous image respectively to obtain sub-image feature information corresponding to the anomalous region sub-image and overall feature information corresponding to the sample anomalous image; perform image reconstruction processing on the sub-image feature information based on the object anomalous category information using the image generation model to obtain a reconstructed anomalous region sub-image; perform image reconstruction processing on the overall feature information based on the object feature description information, object anomalous location information, and object anomalous category information using the image generation model to obtain a reconstructed sample anomalous image; and adjust the parameters of the image generation model according to the loss information between the anomalous region sub-image and the reconstructed anomalous region sub-image, as well as the loss information between the sample anomalous image and the reconstructed sample anomalous image, to obtain a trained image generation model. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0069] The terminal can be used to: receive a trained image generation model sent by the server. This model generates sub-images of abnormal regions and their corresponding overall abnormal images, constructing sample pairs between the overall abnormal image and its mask image, thereby optimizing the object anomaly detection model. The terminal can include mobile phones, vehicle-mounted terminals, aircraft, tablets, laptops, or personal computers (PCs), etc. A client can also be configured on the terminal, which can be an application client or a browser client, etc.

[0070] The sample generation method and image generation model training method provided in the embodiments of this application relate to machine learning and computer vision technologies in the field of artificial intelligence.

[0071] The sample generation method and image generation model training method provided in this application can be applied to various scenarios that require object anomaly detection, as illustrated below.

[0072] 1) In industrial production line quality inspection scenarios, such as automated quality inspection processes for automotive parts and electronic components, the number of real defect samples (such as cracks, scratches, and missing solder joints) is small and diverse. In this case, quality inspection engineers can input parameters of the sample to be generated, such as "automotive piston," "side view P003," and "crack," into the server via a terminal. The server, using the sample generation method provided in this application, generates an image of a piston (an abnormal overall image) including cracks of specified types and locations, along with its corresponding abnormality mask image, through an image generation model (trained using the image generation model training method provided in this application). The sample pair consisting of the abnormal overall image and the abnormality mask image is used to optimize the object anomaly detection model. After deployment, the object anomaly detection model can quickly and accurately identify parts with real defects on the production line, thereby improving quality inspection efficiency and product yield.

[0073] 2) In product appearance inspection scenarios, such as the inspection of products with extremely high requirements for surface smoothness, like watches and smartphone casings, the shape, size, and location of minute defects (such as minor scratches and dents) vary infinitely, making it difficult to completely cover them with a limited number of real samples. In this case, engineers can input the parameters of the sample to be generated, such as "watch dial" or "P001 front," into the server via a terminal, specifying the categories and possible locations of various minute defects. The server, using the sample generation method provided in this application embodiment, generates diverse defect sample images (abnormal overall images) and their corresponding abnormality mask images through an image generation model (trained using the image generation model training method provided in this application embodiment). The sample pairs composed of the abnormal overall images and the abnormality mask images are used to optimize the object anomaly detection model. After deployment, the object anomaly detection model can have higher sensitivity and recognition accuracy for product surface defects, thereby ensuring that the product meets the quality standards for leaving the factory.

[0074] The following sections provide detailed descriptions of each example. It should be noted that the order in which the embodiments are described is not intended to limit the preferred order of the embodiments.

[0075] This application will describe the embodiments from the perspective of an image generation model training device, which can be integrated into an electronic device, such as a server or a terminal.

[0076] Taking an image generation model training device (terminal or server) for image generation model training as an example, the image generation model training method provided in this application embodiment is described below. Referring to Figure 1b, which is a flowchart of the image generation model training method provided in this application embodiment, the steps shown in Figure 1b will be explained in conjunction with the steps illustrated in Figure 1b.

[0077] In step 101, training data and an image generation model are obtained. The training data includes at least one sample abnormal image, object feature description information in the sample abnormal image, object abnormal location information, and object abnormal category information.

[0078] Among them, the image generation model is a neural network model used to generate abnormal images, such as an image generation model based on a diffusion model architecture.

[0079] The sample anomaly image can be an abnormal image of a certain object, which can be a component (such as an industrial part), assembly, or finished product. For example, an anomaly image can be an image corresponding to an object with defects. These defects can include flaws in the object's dimensions, material texture, surface finish, assembly condition, etc.

[0080] The object feature description information can include object category information and object shooting location information. For example, object category information could be the type of component, such as gears or nuts, or the product type, such as watches. Object shooting location information could include the object's display posture and placement angle, for example, shooting locations from different angles such as front, side, and top views. In practical applications, different shooting locations can be numbered, for example, using consecutive numbers according to the determined order of the shooting locations, such as P001, P002, P003, etc.

[0081] Among them, the abnormal object location information is the location of the abnormal part of the object in the sample abnormal image. For example, if the abnormal object location information is [1304, 1054, 1340, 1289], it can be said that the abnormal part of the object is located in a rectangular area with the upper left corner coordinates (1304, 1054) and the lower right corner coordinates (1340, 1289) in the image.

[0082] The object anomaly category information can include the defect type of the object's dimensions, material texture, surface finish, assembly status, etc. For example, for the surface finish of a component, the anomaly category could be scratches, burrs, burns, etc.

[0083] In some embodiments, training data can be obtained by capturing product images using an image acquisition device (such as a high-definition camera), then having domain experts (such as quality inspection engineers) manually annotate the captured images using annotation software to identify sample abnormal images and record their corresponding information, and finally storing the images and associated information (i.e., object feature description information, object abnormality location information, and object abnormality category information in the sample abnormal images) in a database or file system to form a training dataset.

[0084] In step 102, based on the abnormal object location information, the abnormal sample image is processed by image extraction to obtain an abnormal region sub-image.

[0085] Among them, image extraction processing is used to separate the part of the abnormal object location information in the sample abnormal image.

[0086] For example, an anomaly sub-image is a portion of a sample anomaly image, including the region within the sample anomaly image corresponding to the location information of the object anomaly. The anomaly sub-image, also known as a cropped defect image, can be called a cropped defect image.

[0087] For example, if the overall size of the abnormal image is 256×256 pixels, then the size of the abnormal region sub-image can be 64×64 pixels, meaning the size of the abnormal region sub-image is smaller than the size of the sample abnormal image.

[0088] For example, image matting is an operation that extracts local regions of an image based on location information. Taking the location information of an anomaly as represented by coordinates, such as the coordinates of the top-left and bottom-right corners of a rectangular bounding box (x_min, y_min, x_max, y_max), image matting uses these coordinates to segment the sample anomaly image, separating the pixel regions within the bounding box as anomaly region sub-images.

[0089] In some embodiments, based on the abnormal object location information, image extraction processing is performed on the sample abnormal image to obtain an abnormal region sub-image. This can be achieved in the following ways: by using the cropping function in an image processing library (such as OpenCV or Pillow), the position coordinates (i.e., the abnormal object location information) are taken as parameters and the sample abnormal image is sliced. For example, the sample abnormal image can be read into a NumPy array using a Python script, and then the corresponding pixel block can be extracted using the array index (such as image[y_min:y_max, x_min:x_max]) to obtain the abnormal region sub-image.

[0090] Image extraction processing can accurately separate abnormal regions from abnormal images of original samples. This not only eliminates interference from irrelevant background information, but also effectively reduces the amount of data for subsequent processing.

[0091] In step 103, feature extraction processing is performed on the abnormal region sub-image and the sample abnormal image respectively to obtain the sub-image feature information corresponding to the abnormal region sub-image and the overall feature information corresponding to the sample abnormal image.

[0092] In some embodiments, a neural network model can be used to extract features from the abnormal region sub-image and the sample abnormal image to obtain the sub-image feature information corresponding to the abnormal region sub-image and the overall feature information corresponding to the sample abnormal image.

[0093] For example, neural network models can be visual encoding models, such as ResNet (Residual Network), HRNet (High-Resolution Network), ViT (Vision Transformer) model, SwinT (Shifted Window Transformer) model, etc.

[0094] In this context, the feature information of the sub-image corresponding to the abnormal region focuses more on the abnormal part, while the feature information of the sample abnormal image focuses more on the image as a whole. The overall feature information of the sample abnormal image can include the color features, texture features, shape features, etc. of the objects in the sample abnormal image, and can also include feature information of the abnormal parts of the objects.

[0095] In step 104, the image generation model is used to perform image reconstruction processing on the feature information of the sub-image based on the object anomaly category information to obtain the reconstructed anomaly region sub-image.

[0096] In some embodiments, an image reconstruction process is performed on the feature information of a sub-image based on the object anomaly category information using an image generation model to obtain a reconstructed anomaly region sub-image. This can be achieved by: adding noise to the feature information of the sub-image to obtain noisy sub-image feature information; and then, using an image generation model, denoising the noisy sub-image feature information according to the object anomaly category information to obtain a reconstructed anomaly region sub-image.

[0097] In some embodiments, the subgraph feature information is denoised to obtain the denoised subgraph feature information. This can be achieved by using a noise addition strategy with a preset number of steps (e.g., T steps) to gradually inject noise into the subgraph feature information to obtain the denoised subgraph feature information.

[0098] For example, random perturbations are applied to the feature points in the subgraph feature information to introduce noise into the subgraph feature information; the introduced noise can be Gaussian noise, etc. For instance, noise is added to the subgraph feature information step by step in T steps, finally obtaining the noisy subgraph feature information after T noisy additions. During the T iterations of forward diffusion, the subgraph feature information gradually loses its unique and distinct characteristics, and finally, when T approaches ∞, the noisy subgraph feature information is equivalent to an isotropic Gaussian distributed noise (the added noise here can be, for example, Gaussian noise).

[0099] For example, the noise-adding process is a forward diffusion process, where x is diffused through a diffusion process at each time step t. t-1 Convert to x tAfter T time steps, the completely noisy data x is obtained. T x T That is, after T rounds of noise addition, the feature information of the sub-image can be denoised. The noise can follow a Gaussian distribution, and the denoising process can be a reverse diffusion process.

[0100] In some embodiments, content prompts for reconstructing anomaly region sub-maps can be determined based on object anomaly category information.

[0101] For example, the abnormal object category information is directly used as content prompt information.

[0102] For example, the abnormal object category information is fused with a preset template, and the fused template is used as content prompt information, thereby performing noise reduction processing based on the content prompt information.

[0103] For example, the preset template can be an anomaly category: []. The process of integrating the object anomaly category information with the preset template can be, for example, filling the object anomaly category information into [].

[0104] In some embodiments, an image generation model is used to denoise the feature information of the noisy sub-image based on the object anomaly category information to obtain a reconstructed anomaly region sub-image. This can be achieved by: using an image generation model to perform downsampling and upsampling processing on the feature information of the noisy sub-image at multiple scales based on the object anomaly category information to predict the noise residual information of the feature information of the noisy sub-image; and then denoising the feature information of the noisy sub-image based on the noise residual information to obtain a reconstructed anomaly region sub-image.

[0105] For example, object anomaly category information can be used to guide the denoising process, thereby reconstructing the anomaly region sub-image. The denoising process can involve multiple steps, such as T steps. In each denoising step, the image generation model predicts the noise contained in the feature information of the noisy sub-image. Utilizing the predicted current noise residual information, the noise in the features is gradually removed, thus restoring the anomaly region sub-image after T steps of denoising.

[0106] In some embodiments, an image generation model performs downsampling and upsampling processing on the noisy sub-image feature information at multiple scales based on object anomaly category information to predict the noise residual information of the noisy sub-image feature information. This can be achieved as follows: The image generation model extracts features from the object anomaly category information to obtain cue semantic feature information; attention processing is applied to the noisy sub-image feature information based on the cue semantic feature information to obtain attention fusion features; the attention fusion features are downsampled at multiple scales to obtain target features; the target features are upsampled at multiple scales to obtain upsampled target features; and attention processing is applied to the upsampled target features based on the cue semantic feature information to obtain the noise residual information of the noisy sub-image feature information.

[0107] In some embodiments, attention processing is performed on the feature information of the noisy subgraph based on the semantic feature information of the prompt to obtain attention fusion features. This can be achieved by: using the feature information of the noisy subgraph as a query vector, and using the semantic feature information of the prompt as a key vector and a value vector; determining the attention score based on the query vector and the key vector, and normalizing the attention score to obtain the attention weight; and performing a weighted summation of the value vector based on the attention weight to obtain the attention fusion features.

[0108] For example, the feature information of the noisy subgraph is taken as Q (query vector), and the semantic feature information of the prompt is taken as K (key vector) and V (value vector). The query vector and the key vector are then multiplied by a dot product and normalized by a scaling factor to obtain the attention score. Then, the attention score is normalized by the softmax activation function to obtain the attention weight. Finally, the value vector is weighted and summed according to the attention weight to obtain the attention fusion feature.

[0109] In some embodiments, attention processing is performed on the upsampled target features based on the prompt semantic feature information to obtain noise residual information of the subgraph feature information after adding noise. This can be achieved by: using the upsampled target features as a query vector and the prompt semantic feature information as a key vector and a value vector; determining the attention score based on the query vector and the key vector, and normalizing the attention score to obtain the attention weight; and performing a weighted summation of the value vector based on the attention weight to obtain the noise residual information.

[0110] For example, the upsampled target features are used as Q (query vector), and the semantic features of the prompts are used as K (key vector) and V (value vector). The query vector and key vector are then multiplied by a dot product and normalized by a scaling factor to obtain the attention score. The attention score is then normalized using the softmax activation function to obtain the attention weight. Finally, the value vector is weighted and summed according to the attention weight to obtain the noise residual information.

[0111] In some embodiments, attention mechanisms can be used to guide the model to focus on important parts of the input data, helping the model understand which features are relevant to the description of the content prompts, thereby improving the quality of the generated data.

[0112] For example, an image generation model may include an encoding part and a decoding part. The encoding part may include multiple downsampling modules, and the decoding part may include multiple upsampling modules. Each downsampling module may include an attention processing unit and a downsampling unit, and each upsampling module may include an upsampling unit and an attention processing unit. During each denoising step, the image generation model can predict the noise residual information of the current step based on the content cue information (i.e., cue semantic feature information) and the feature information of the current noisy sub-image.

[0113] For example, each step of denoising the feature information of the noisy sub-image can be achieved as follows: First, the downsampling module of the encoding part in the image generation model performs attention processing on the feature information of the noisy sub-image based on the cue semantic feature information. After attention processing, downsampling processing is performed. Then, it enters the next downsampling module of the encoding part, and performs attention processing through the attention processing unit of the next downsampling module. Then, it performs downsampling processing through the downsampling unit, and so on, until the last downsampling module is reached. Through the attention processing and downsampling processing of the last downsampling module, the target feature output by the encoding part is obtained. Then, the upsampling module of the decoding part performs upsampling processing on the target feature. After upsampling processing, attention processing is performed. Then, it enters the next upsampling module for upsampling processing and attention processing, and so on, to obtain the output of the decoding part, that is, the noise residual information of the feature information of the noisy sub-image.

[0114] In some embodiments, downsampling the attention fusion features at multiple scales to obtain the target features can be achieved as follows: downsampling the attention fusion features to obtain downsampled features; determining the downsampled features as new noisy subgraph feature information; returning to execute the step of performing attention processing on the noisy subgraph feature information based on the prompt semantic feature information to obtain attention fusion features, until the number of iterations meets the preset condition to obtain downsampled features at multiple scales; and determining the target features based on the downsampled features at the target scale.

[0115] For example, the downsampled features at the target scale could be the smallest downsampled features among multiple scales. The feature map of the downsampled features at the target scale has the lowest resolution but contains the richest semantic information.

[0116] For example, the preset conditions can be set according to the actual situation. The total number of iterations can be determined, for example, based on the number of downsampling modules in the encoding part of the image generation model.

[0117] In some embodiments, upsampling the target features at multiple scales to obtain upsampled target features can be achieved as follows: upsample the target features to obtain upsampled features; perform attention processing on the upsampled features based on the prompt semantic feature information to obtain attention-processed features; upsample the attention-processed features to obtain new upsampled features; return to the step of performing attention processing on the upsampled features based on the prompt semantic feature information to obtain attention-processed features, until the number of iterations meets a preset condition, and the upsampled target features are obtained.

[0118] For example, the preset conditions can be set according to the actual situation. The total number of iterations can be determined, for example, based on the number of upsampling modules in the decoding part of the image generation model.

[0119] For example, the upsampling process progressively generates higher-resolution image features. An attention mechanism is integrated throughout the downsampling and upsampling processes, allowing semantic information from cues to be incorporated into the image features. Simultaneously, through the attention mechanism, the model can reference content cues at each denoising step to ensure that the generated image matches the content cues.

[0120] In some embodiments, the denoising process of the feature information of the noisy sub-image is performed based on the noise residual information to obtain the reconstructed abnormal region sub-image. This can be achieved in the following way: the feature information of the noisy sub-image is denoised based on the noise residual information to obtain the denoised sub-image feature information; the denoised sub-image feature information is determined as the new noisy sub-image feature information; the process is then repeated to execute the step of performing downsampling and upsampling processing on the feature information of the noisy sub-image at multiple scales based on the object anomaly category information through the image generation model to predict the noise residual information of the noisy sub-image feature information, until the number of iterations meets the preset condition and the reconstructed abnormal region sub-image is obtained.

[0121] For example, in the forward diffusion process (i.e., the noise-adding process), the noisy subgraph feature information x at any time t. t It is x at time t-1 t-1 It is obtained by superimposing a Gaussian noise ε with a known variance. It can be expressed as formula (1):

[0122] Where, β t The variance scheduling parameter β is the noise added at time t. t =1-α t , where α t The decays with time step, such as α1 = 0.9, α2 = 0.89, α3 = 0.88.

[0123] In the reverse denoising process, the goal is to start from x. t Recover x t-1 Based on the above formula (1) and Bayes' theorem, it can be derived that from x... t and the predicted noise residual information ε θ (x t To estimate x using (c,t) t-1 The calculation formula for (i.e., the feature information of the denoised sub-image). For example, it is expressed as formula (2):

[0124] in, It is the cumulative variance parameter; ε θ (x t (c,t) represents the image generation model (such as U-Net) at time t, with the added noisy sub-image feature information x. t The noise residual is predicted using the object anomaly category information c as input; σ t It is a parameter that controls randomness (such as σ) t 2 =β t z is a standard Gaussian noise.

[0125] For example, each round of denoising can predict the current noise residual information through the image generation model, so as to gradually remove the noise in the feature information of the noisy sub-image and restore the abnormal region sub-image.

[0126] For example, the total number of iterations required could be T. For instance, noise could be added to the subgraph feature information T times to obtain the noisy subgraph feature information, and then the noisy subgraph feature information could be denoised T times.

[0127] By performing image reconstruction processing on the feature information of the sub-image based on the object anomaly category information (e.g., through a diffusion process of adding and removing noise), and with the help of an attention processing mechanism, a reconstructed anomaly region sub-image that is semantically closely related to the object anomaly category information can be generated. This allows the model to learn the essential features of the anomaly region corresponding to the object anomaly category, rather than simply copying them.

[0128] In step 105, an image generation model is used to perform image reconstruction processing on the overall feature information based on object feature description information, object anomaly location information, and object anomaly category information to obtain a reconstructed sample anomaly image.

[0129] In some embodiments, an image reconstruction process is performed on the overall feature information based on object feature description information, object anomaly location information, and object anomaly category information using an image generation model to obtain a reconstructed sample anomaly image. This can be achieved in the following ways: constructing content prompt information for reconstructing the sample anomaly image based on object feature description information, object anomaly location information, and object anomaly category information; adding noise to the overall feature information to obtain noisy overall feature information; and using an image generation model, denoising the noisy overall feature information according to the content prompt information to obtain the reconstructed sample anomaly image.

[0130] In some embodiments, content prompting information for reconstructing sample abnormal images can be constructed based on object feature description information, object abnormal location information, and object abnormal category information. This can be achieved by concatenating the object feature description information, object abnormal location information, and object abnormal category information to obtain the content prompting information for reconstructing sample abnormal images.

[0131] In other embodiments, content prompting information for reconstructing sample abnormal images can be constructed based on object feature description information, object abnormal location information, and object abnormal category information. This can also be achieved by fusing the object feature description information, object abnormal location information, and object abnormal category information with a preset template, using the fused template as content prompting information, and then performing noise reduction processing based on the content prompting information.

[0132] For example, the template format can be "A [object category] sample at pose [shooting point], which has a defect of a [defect category] at [defect coordinates]". The process of integrating the object feature description information, object abnormal location information, and object abnormal category information with the template can be, for example, by filling the corresponding information into the [] of the template. For instance, if the object feature description information includes the object category information "watch" and the object's shooting point information P007, the object abnormal location information is [1304,1054,1340,1289], and the object abnormal category information is HS (scratch), then the content prompt information obtained after integration is, for example: A watch sample at pose P007, which has a defect of a HS at [1304,1054,1340,1289], which means: A watch sample with a shooting point at P007 has a scratch defect, and the scratch location is at the [1304,1054,1340,1289] position in the image.

[0133] In some embodiments, the overall feature information is denoised to obtain the denoised overall feature information, which can be achieved by denoising the overall feature information multiple times to obtain the denoised overall feature information.

[0134] For example, random perturbations are applied to feature points in the overall feature information to introduce noise; the introduced noise can be Gaussian noise, etc. For instance, noise can be added to the overall feature information step by step in T steps, finally obtaining the noisy overall feature information after T iterations. During this process, the overall feature information gradually loses its unique and distinct characteristics in the T forward diffusion iterations. Eventually, when T approaches ∞, the noisy overall feature information is equivalent to an isotropic Gaussian distributed noise (the added noise here can be, for example, Gaussian noise).

[0135] For example, the noise-adding process is also a forward diffusion process, where x is diffused through a diffusion process at each time step t. t-1 Convert to x t After T time steps, the completely noisy data x is obtained. T x T That is, after T rounds of noise addition, the overall feature information shows that the noise follows a Gaussian distribution. Denoising is a reverse diffusion process.

[0136] In some embodiments, an image generation model is used to denoise the overall feature information after adding noise based on content prompts to obtain a reconstructed abnormal image of the sample. This can be achieved by: using an image generation model to perform downsampling and upsampling processing on the overall feature information after adding noise at multiple scales based on content prompts to predict the noise residual information of the overall feature information after adding noise; and then denoising the overall feature information after adding noise based on the noise residual information to obtain a reconstructed abnormal image of the sample.

[0137] For example, content prompts can be used to guide the denoising process, thereby reconstructing the abnormal sample image. This denoising process can involve multiple steps, such as a T-step process. In each denoising step, the image generation model predicts the noise contained in the overall feature information after adding noise. Utilizing the predicted current noise residual information, the noise in the features is gradually removed, thus restoring the abnormal sample image after T steps of denoising.

[0138] In some embodiments, an image generation model performs downsampling and upsampling processing on the overall feature information after adding noise at multiple scales based on content prompts to predict the noise residual information of the overall feature information after adding noise. This can be achieved as follows: using an image generation model, attention processing is performed on the overall feature information after adding noise based on content prompts to obtain attention fusion features; the attention fusion features are downsampled at multiple scales to obtain target features; the target features are upsampled at multiple scales to obtain upsampled target features; and attention processing is performed on the upsampled target features based on content prompts to obtain the noise residual information of the overall feature information after adding noise.

[0139] For example, features can be extracted from the content prompt information to obtain the prompt semantic feature information corresponding to the content prompt information. Then, attention processing can be performed on the overall feature information after adding noise based on the prompt semantic feature information, or attention processing can be performed on the upsampled target features based on the prompt semantic feature information.

[0140] In some embodiments, an image generation model performs attention processing on the overall feature information after adding noise based on the content prompt information to obtain attention fusion features. This can be achieved as follows: the overall feature information after adding noise is taken as Q (query vector), and the prompt semantic feature information corresponding to the content prompt information is taken as K (key vector) and V (value vector). The query vector and key vector are then multiplied by a dot product and normalized by a scaling factor to obtain an attention score. Then, the attention score is normalized using a softmax activation function to obtain attention weights. Finally, the value vectors are weighted and summed according to the attention weights to obtain the attention fusion features.

[0141] In some embodiments, attention processing is performed on the upsampled target features based on the content prompt information to obtain noise residual information of the overall feature information after adding noise. This can be achieved as follows: the upsampled target features are taken as Q (query vector), the prompt semantic features corresponding to the content prompt information are taken as K (key vector) and V (value vector), the query vector and key vector are then multiplied by a dot product and normalized by a scaling factor to obtain an attention score; then the attention score is normalized using a softmax activation function to obtain attention weights; finally, the value vectors are weighted and summed according to the attention weights to obtain noise residual information.

[0142] In some embodiments, the image generation model may include an encoding part and a decoding part. The encoding part may include multiple downsampling modules, and the decoding part may include multiple upsampling modules. Each downsampling module may include an attention processing unit and a downsampling unit, and each upsampling module may include an upsampling unit and an attention processing unit. During each denoising step, the image generation model can predict the noise residual information of the current step based on the content cue information (i.e., cue semantic feature information) and the current overall feature information after adding noise.

[0143] For example, each step of the denoising process for the overall feature information after adding noise can be implemented as follows: First, the downsampling module of the encoding part in the image generation model performs attention processing on the overall feature information after adding noise based on the cue semantic feature information. After attention processing, downsampling processing is performed. Then, it enters the next downsampling module of the encoding part, and performs attention processing through the attention processing unit of the next downsampling module. Then, it performs downsampling processing through the downsampling unit, and so on, until the last downsampling module is reached. Through the attention processing and downsampling processing of the last downsampling module, the target feature output by the encoding part is obtained. Then, the upsampling module of the decoding part performs upsampling processing on the target feature. After upsampling processing, attention processing is performed. Then, it enters the next upsampling module for upsampling processing and attention processing, and so on, to obtain the output of the decoding part, that is, the noise residual information of the overall feature information after adding noise.

[0144] In some embodiments, downsampling the attention fusion features at multiple scales to obtain the target features can be achieved as follows: downsampling the attention fusion features to obtain downsampled features; determining the downsampled features as new noisy overall feature information; returning to execute the step of performing attention processing on the noisy overall feature information according to the content prompt information to obtain attention fusion features, until the number of iterations meets the preset condition to obtain downsampled features at multiple scales; determining the target features based on the downsampled features at the target scale.

[0145] For example, the downsampled features at the target scale could be the smallest downsampled features among multiple scales. The feature map of the downsampled features at the target scale has the lowest resolution but contains the richest semantic information.

[0146] For example, the preset conditions can be set according to the actual situation. The total number of iterations can be determined based on the number of downsampling modules in the encoding part of the image generation model.

[0147] In some embodiments, upsampling the target features at multiple scales to obtain upsampled target features can be achieved as follows: upsample the target features to obtain upsampled features; perform attention processing on the upsampled features based on content prompts to obtain attention-processed features; upsample the attention-processed features to obtain new upsampled features; return to the step of performing attention processing on the upsampled features based on content prompts to obtain attention-processed features, until the number of iterations meets a preset condition, and the upsampled target features are obtained.

[0148] For example, the preset conditions can be set according to the actual situation. The total number of iterations can be determined based on the number of upsampling modules in the decoding part of the image generation model.

[0149] For example, the upsampling process progressively generates higher-resolution image features. An attention mechanism is integrated throughout the downsampling and upsampling processes, allowing semantic information from cues to be incorporated into the image features. Simultaneously, through the attention mechanism, the model can reference content cues at each denoising step to ensure that the generated image matches the content cues.

[0150] In some embodiments, the overall feature information after noise addition is denoised based on the noise residual information to obtain the reconstructed abnormal image of the sample. This can be achieved in the following way: the overall feature information after noise addition is denoised based on the noise residual information to obtain the denoised overall feature information; the denoised overall feature information is determined as the new overall feature information after noise addition; the process of performing downsampling and upsampling processing on the overall feature information after noise addition at multiple scales based on the content prompt information through the image generation model to predict the noise residual information of the overall feature information after noise addition is repeated until the number of iterations meets the preset condition to obtain the reconstructed abnormal image of the sample.

[0151] For example, in each round of denoising of the overall feature information after adding noise, the current noise residual information can be predicted by the image generation model to gradually remove noise from the overall feature information after adding noise and restore the abnormal sample image.

[0152] For example, the total number of iterations required could be T. For instance, noise could be added to the overall feature information T times to obtain the noisy overall feature information, and then denoising could be performed on the noisy overall feature information T times.

[0153] By constructing a content-based cue that integrates object feature descriptions, object anomaly location information, and object anomaly category information, the image generation model is guided to reconstruct the image using overall feature information. An attention processing mechanism, spanning both downsampling and upsampling processes, deeply fuses the content-based cue with the noisy overall feature information at different scales. This achieves a high degree of control over the reconstruction process of sample anomaly images, ensuring that the final reconstructed sample anomaly image not only maintains overall consistency with the original image but also that its anomaly features precisely match the requirements of the content-based cue in terms of category, morphology, and spatial location. This enables high-fidelity, controllable generation of specific anomalies within a complete scene.

[0154] In step 106, the parameters of the image generation model are adjusted based on the loss information between the abnormal region sub-image and the reconstructed abnormal region sub-image, as well as the loss information between the sample abnormal image and the reconstructed sample abnormal image, to obtain the trained image generation model.

[0155] In some embodiments, the parameters of the image generation model are adjusted based on the loss information between the abnormal region sub-image and the reconstructed abnormal region sub-image, and the loss information between the sample abnormal image and the reconstructed sample abnormal image, to obtain the trained image generation model. This can be achieved as follows: First loss information between the abnormal region sub-image and the reconstructed abnormal region sub-image is calculated based on the image similarity between them; second loss information between the sample abnormal image and the reconstructed sample abnormal image is calculated based on the image similarity between them; and the parameters of the image generation model are adjusted based on the first and second loss information to obtain the trained image generation model.

[0156] For example, the higher the image similarity, the lower the information loss; conversely, the lower the image similarity, the higher the information loss.

[0157] In some embodiments, the parameters of the image generation model are adjusted based on the first loss information and the second loss information to obtain the trained image generation model. This can be achieved by: determining the weight information corresponding to the first loss information and the second loss information respectively; fusing the first loss information and the second loss information according to the weight information to obtain the total loss information; and adjusting the parameters of the image generation model according to the total loss information to obtain the trained image generation model.

[0158] For example, the fusion of the first loss information and the second loss information can be achieved through weighted operations, summation operations, etc.

[0159] For example, the training process can begin by reconstructing the abnormal region sub-image to obtain the reconstructed abnormal region sub-image, followed by reconstructing the sample abnormal image to obtain the reconstructed sample abnormal image. Next, the image similarity between the abnormal region sub-image and the reconstructed abnormal region sub-image, as well as the image similarity between the sample abnormal image and the reconstructed sample abnormal image, is calculated. Then, the parameters of the image generation model are adjusted using the backpropagation algorithm. Based on the image similarity between the abnormal region sub-image and the reconstructed abnormal region sub-image, and the image similarity between the sample abnormal image and the reconstructed sample abnormal image, the parameters of the image generation model are optimized so that the reconstructed abnormal region sub-image closely approximates the original abnormal region sub-image, and the reconstructed sample abnormal image closely approximates the original sample abnormal image, resulting in a well-trained image generation model. For instance, the image similarity between the abnormal region sub-image and the reconstructed abnormal region sub-image can be greater than a first preset similarity, and the image similarity between the sample abnormal image and the reconstructed sample abnormal image can be greater than a second preset similarity. The first and second preset similarities can be set according to actual conditions.

[0160] In some embodiments, the parameters of the image generation model are adjusted according to the total loss information to obtain the trained image generation model. This can be achieved by: obtaining the gradient information of the total loss information for each parameter of the image generation model through the backpropagation algorithm; updating the parameters of the image generation model using the obtained gradient information according to the gradient descent optimization algorithm (e.g., batch gradient descent, stochastic gradient descent, etc.); repeating the above process until a certain number of iterations is reached or the image generation model converges, thereby obtaining the trained image generation model.

[0161] A dual-supervision mechanism is employed to adjust the parameters of the image generation model. It simultaneously calculates the first loss information between the anomalous region sub-image and its reconstructed anomalous region sub-image, and the second loss information between the sample anomalous image and its reconstructed sample anomalous image. By fusing these two loss information into a single total loss information and optimizing it using backpropagation, the image generation model is ensured to focus not only on the overall image reconstruction quality during training but also on accurately learning and reproducing the details of local anomalous regions. This joint optimization strategy, which considers both local and global aspects, results in a final trained image generation model with higher fidelity and more precise control over key anomalous features in reconstruction tasks.

[0162] Through steps 101 to 106, combining local fine-grained learning with globally controllable generation, a trained image generation model capable of reproducing specific abnormal scenes with high fidelity and accuracy is finally obtained. First, image extraction (step 102) focuses on the abnormal region sub-image and reconstructs it (steps 103-104), forcing the model to learn the essential semantics of the defect. Next, content-based cues guide the overall reconstruction of the abnormal sample image (step 105), ensuring the contextual accuracy and controllability of the generated content. Finally, during the training phase (step 106), the model parameters are adjusted by fusing the first loss information from local reconstruction and the second loss information from overall reconstruction to form the total loss information. This dual-supervision strategy, which considers both local details and global consistency, ensures that the final model, in the generation task, can both deeply understand and reproduce the fine features of the defect and accurately place it in the correct macroscopic scene, achieving highly controllable and accurate reproduction of specific abnormal samples.

[0163] In this embodiment, during training, the generation of small defect anomaly images and overall anomaly images are learned sequentially. This allows the model to simultaneously learn the feature information of local defects and the overall image information, ensuring the consistency between the overall anomaly images and local anomaly images generated by the trained image generation model. This enables the construction of realistic and highly aligned anomaly image-mask data sample pairs. The method provided in this embodiment can amplify a small amount of labeled anomaly defect data to generate a large amount of labeled anomaly defect data. The generated data exhibits good realism and diversity, and can be directly used for supervised anomaly detection training, significantly improving the performance of downstream anomaly detection tasks, including defect detection, localization, and classification.

[0164] This application will describe the embodiments from the perspective of a sample generation device, which may be integrated into an electronic device, such as a server or a terminal.

[0165] The following describes the sample generation method provided in this application embodiment, using a sample generation device (terminal or server) as an example. Referring to Figure 2a, which is a flowchart of the sample generation method provided in this application embodiment, the steps shown in Figure 2a will be explained in conjunction with the steps illustrated therein.

[0166] In step 201, the object feature description information, object anomaly location information, and object anomaly category information for the target object are obtained.

[0167] For example, the target object can be the object from which the abnormal image is to be generated. For example, the object can be a component (such as an industrial part), a component, or a finished product, and the abnormal image is an image of the object having defects. Defects can include defects in the object's size specifications, material texture, surface finish, assembly status, etc.

[0168] For example, object feature description information could be information describing the features of objects in the generated abnormal image. This could include object category information, object shooting location information, etc. Object category information could be, for example, the type of component, such as gears or nuts, or the type of product, such as watches. Object shooting location information could be, for example, the object's display posture, placement angle, etc., and could include shooting locations from different angles such as front, side, and top views. In practical applications, different shooting locations can be numbered, for example, using consecutive numbers according to the order in which the shooting locations are determined, such as P001, P002, P003, etc., so that the corresponding shooting location can be identified by the numerical number.

[0169] For example, the object anomaly location information refers to the position of the abnormal part of the object within the overall anomaly image to be generated. If the object anomaly location information is specified as [1304,1054,1340,1289], then the abnormal part of the object is located within a rectangular area with the upper left corner coordinates (1304,1054) and the lower right corner coordinates (1340,1289) in the generated overall anomaly image.

[0170] For example, object anomaly category information can include the type of defect such as the object's dimensions, material texture, surface finish, and assembly condition. For instance, for the surface finish of a component, the anomaly category could be scratches, burrs, burns, etc.

[0171] In step 202, an abnormal region sub-map is generated based on the object anomaly category information using an image generation model.

[0172] For example, the image generation model is trained using the image generation model training method provided in the embodiments of this application. The image generation model is a neural network model used to generate abnormal images, such as a diffusion model.

[0173] In some embodiments, an abnormal region sub-map is generated based on object anomaly category information using an image generation model. This can be achieved by sampling from a preset noise distribution to generate noise feature information; and by using an image generation model to denoise the noise feature information based on object anomaly category information to obtain the abnormal region sub-map.

[0174] For example, the preset noise distribution can be a Gaussian distribution, etc., and this application embodiment does not limit this. The noise feature information can be generated by random sampling from the preset noise distribution. This process can provide a starting point for the denoising process and create a feature map containing random noise.

[0175] For example, object anomaly category information can be used as content prompt information for generating anomaly region submaps.

[0176] In some embodiments, object anomaly category information can be directly used as content prompt information.

[0177] In other embodiments, object anomaly category information can be fused with a preset template, and the fused template can be used as content prompt information for noise reduction. The preset template can be anomaly category: []. The fusion process of object anomaly category information and preset template can be, for example, filling the object anomaly category information into [].

[0178] In some embodiments, an image generation model is used to denoise the noise feature information based on the object anomaly category information to obtain an anomaly region sub-image. This can be achieved by: using an image generation model to perform downsampling and upsampling processing on the noise feature information at multiple scales based on the object anomaly category information to predict the noise residual information of the noise feature information; and then denoising the noise feature information based on the noise residual information to obtain the anomaly region sub-image.

[0179] For example, object anomaly category information can be used to guide the denoising process, thereby generating an anomaly region sub-image. The denoising process can involve multiple steps. In each step, an image generation model predicts the noise contained in the noise features, and using the predicted current noise residual information, noise is gradually removed from the features. This process, involving T denoising steps, generates the anomaly region sub-image.

[0180] In some embodiments, an image generation model is used to perform downsampling and upsampling processing on noise feature information at multiple scales based on object anomaly category information to predict noise residual information of noise feature information. This can be achieved as follows: The image generation model extracts features from object anomaly category information to obtain cue semantic feature information; attention processing is applied to the noise feature information based on the cue semantic feature information to obtain attention fusion features; the attention fusion features are downsampled at multiple scales to obtain target features; the target features are upsampled at multiple scales to obtain upsampled target features; and attention processing is applied to the upsampled target features based on the cue semantic feature information to obtain noise residual information of noise feature information.

[0181] For example, attention processing of noisy features based on prompt semantic features to obtain attention fusion features can be achieved in the following way: Noisy features are treated as Q (query vector), and prompt semantic features are treated as K (key vector) and V (value vector). The query vector and key vector are then multiplied by a dot product and normalized using a scaling factor to obtain an attention score. The attention score is then normalized using a softmax activation function to obtain attention weights. Finally, the value vectors are weighted and summed according to the attention weights to obtain the attention fusion features.

[0182] For example, to obtain noise residual information of noise features by performing attention processing on the upsampled target features based on the semantic features of the prompts, this can be achieved as follows: The upsampled target features are used as Q (query vector), and the semantic features of the prompts are used as K (key vector) and V (value vector). The query vector and key vector are then multiplied by a dot product and normalized using a scaling factor to obtain an attention score. The attention score is then normalized using a softmax activation function to obtain attention weights. Finally, the value vectors are weighted and summed according to the attention weights to obtain the noise residual information.

[0183] In some embodiments, attention mechanisms can be used to guide the model to focus on important parts of the input data, helping the model understand which features are relevant to the description of the content prompts, thereby improving the quality of the generated data.

[0184] For example, an image generation model may include an encoding part and a decoding part. The encoding part may include multiple downsampling modules, and the decoding part may include multiple upsampling modules. Each downsampling module may include an attention processing unit and a downsampling unit, and each upsampling module may include an upsampling unit and an attention processing unit. During each denoising step, the image generation model can predict the noise residual information of the current step based on the content cue information (i.e., cue semantic feature information) and the current noise feature information.

[0185] For example, each step of the denoising process for noise feature information can be implemented as follows: First, the downsampling module of the encoding part in the image generation model performs attention processing on the noise feature information based on the cue semantic feature information. After attention processing, downsampling processing is performed. Then, the next downsampling module of the encoding part performs attention processing through its attention processing unit, and then downsampling processing is performed again through the downsampling unit. This process continues until the last downsampling module is reached. Through the attention processing and downsampling processing of the last downsampling module, the target feature output by the encoding part is obtained. Then, the upsampling module of the decoding part performs upsampling processing on the target feature. After upsampling processing, attention processing is performed. Then, the next upsampling module performs upsampling processing and attention processing, and so on, to obtain the noise residual information of the noise feature information, which is the output of the decoding part.

[0186] In some embodiments, downsampling the attention fusion features at multiple scales to obtain the target features can be achieved as follows: downsampling the attention fusion features to obtain downsampled features; identifying the downsampled features as new noise feature information; returning to execute the step of performing attention processing on the noise feature information based on the prompt semantic feature information to obtain attention fusion features, until the number of iterations meets a preset condition to obtain downsampled features at multiple scales; and determining the target features based on the downsampled features at the target scale.

[0187] For example, the downsampled features at the target scale could be the smallest downsampled features among multiple scales. The feature map of the downsampled features at the target scale has the lowest resolution but contains the richest semantic information.

[0188] For example, the preset conditions can be set according to the actual situation. The total number of iterations can be determined based on the number of downsampling modules in the encoding part of the image generation model.

[0189] In some embodiments, upsampling the target features at multiple scales to obtain upsampled target features can be achieved as follows: upsample the target features to obtain upsampled features; perform attention processing on the upsampled features based on the prompt semantic feature information to obtain attention-processed features; upsample the attention-processed features to obtain new upsampled features; return to the step of performing attention processing on the upsampled features based on the prompt semantic feature information to obtain attention-processed features, until the number of iterations meets a preset condition, and the upsampled target features are obtained.

[0190] For example, the preset conditions can be set according to the actual situation. The total number of iterations can be determined based on the number of upsampling modules in the decoding part of the image generation model.

[0191] For example, the upsampling process progressively generates higher-resolution image features. An attention mechanism is integrated throughout the downsampling and upsampling processes, allowing semantic information from cues to be incorporated into the image features. Simultaneously, through the attention mechanism, the model can reference content cues at each denoising step to ensure that the generated image matches the content cues.

[0192] In some embodiments, the noise feature information is denoised based on the noise residual information to obtain an abnormal region sub-image. This can be achieved by: denoising the noise feature information based on the noise residual information to obtain denoised feature information; determining the denoised feature information as new noise feature information; returning to the step of performing downsampling and upsampling processing on the noise feature information at multiple scales based on the object anomaly category information using an image generation model to predict the noise residual information of the noise feature information, until the number of iterations meets a preset condition to obtain an abnormal region sub-image.

[0193] For example, each iteration is a time step. At each time step, the model predicts a denoised feature based on the object anomaly category information and the current noise feature information. It continuously adjusts the feature values ​​of feature points in the feature map, gradually removing noise from the noise feature information, so that the feature gradually approaches the target. By repeating the above steps multiple times, through multiple rounds of refinement, as the time steps progress, the image gradually transforms from pure noise into a meaningful image related to the content prompt information.

[0194] In step 203, an overall image is generated using an image generation model based on object feature description information, object anomaly location information, and object anomaly category information to obtain an overall image of the anomaly.

[0195] In some embodiments, an overall image of the anomaly is generated by an image generation model based on object feature description information, object anomaly location information, and object anomaly category information. This can be achieved by: constructing content prompt information for image generation based on object feature description information, object anomaly location information, and object anomaly category information; sampling from a preset noise distribution to generate noise feature information; and using the image generation model to denoise the noise feature information based on the content prompt information to obtain the overall image of the anomaly.

[0196] For example, the preset noise distribution can be a Gaussian distribution, etc., and this application embodiment does not limit this. The noise feature information can be generated by random sampling from the preset noise distribution. This process can provide a starting point for the denoising process and create a feature map containing random noise.

[0197] In some embodiments, object feature description information, object anomaly location information, and object anomaly category information can be concatenated to obtain content prompt information for generating an overall image of the anomaly.

[0198] In other embodiments, object feature description information, object abnormal location information, and object abnormal category information can be fused with a preset template, and the fused template can be used as content prompt information, thereby performing noise reduction processing based on the content prompt information.

[0199] For example, the template format could be "A [object category] sample at pose [shooting point], which has defect of a [defect category] at [defect coordinates]". The process of integrating the object feature description information, object abnormal location information, and object abnormality category information with the template could be, for example, filling the corresponding information into the [] of the template. For instance, if the object feature description information includes the object category information watch and the object's shooting point information P007, the object abnormal location information is [1304,1054,1340,1289], and the object abnormality category information is HS (scratch), then the content prompt information obtained after integration could be, for example: A watch sample at pose P007, which has defect of a HS at [1304,1054,1340,1289].

[0200] In some embodiments, an abnormal overall image is obtained by denoising the noise feature information based on the content prompt information using an image generation model. This can be achieved by: using an image generation model to perform downsampling and upsampling processing on the noise feature information at multiple scales based on the content prompt information to predict the noise residual information of the noise feature information; and then denoising the noise feature information based on the noise residual information to obtain the abnormal overall image.

[0201] For example, content-based prompts can be used to guide the denoising process, thereby generating an overall image of the anomaly. The denoising process can involve multiple steps. In each step, an image generation model predicts the noise contained in the noise features, and using the predicted current noise residual information, noise is gradually removed from the features. This process, involving T steps of denoising, generates the overall image of the anomaly.

[0202] In some embodiments, the noise feature information is downsampled and upsampled at multiple scales according to the content prompt information by an image generation model to predict the noise residual information of the noise feature information. This can be achieved by: using an image generation model to perform attention processing on the noise feature information according to the content prompt information to obtain attention fusion features; performing downsampling processing on the attention fusion features at multiple scales to obtain target features; performing upsampling processing on the target features at multiple scales to obtain upsampled target features; and performing attention processing on the upsampled target features according to the content prompt information to obtain the noise residual information of the noise feature information.

[0203] For example, features can be extracted from the content prompt information to obtain the prompt semantic feature information corresponding to the content prompt information. Then, attention processing can be performed on the noise feature information based on the prompt semantic feature information, or attention processing can be performed on the upsampled target features based on the prompt semantic feature information.

[0204] In some embodiments, attention processing is performed on the noise feature information based on the content prompt information to obtain attention fusion features. This can be achieved as follows: the noise feature information is taken as Q (query vector), and the prompt semantic feature information corresponding to the content prompt information is taken as K (key vector) and V (value vector). The query vector and key vector are then multiplied by a dot product and normalized by a scaling factor to obtain an attention score. Then, the attention score is normalized using a softmax activation function to obtain attention weights. Finally, the value vectors are weighted and summed according to the attention weights to obtain the attention fusion features.

[0205] In some embodiments, attention processing is performed on the upsampled target features based on the content prompt information to obtain noise residual information of the noise feature information. This can be achieved as follows: the upsampled target features are taken as Q (query vector), the prompt semantic feature information corresponding to the content prompt information is taken as K (key vector) and V (value vector), the query vector and the key vector are then multiplied by a dot product and normalized by a scaling factor to obtain an attention score; then the attention score is normalized using a softmax activation function to obtain attention weights; finally, the value vectors are weighted and summed according to the attention weights to obtain the noise residual information.

[0206] For example, attention processing can be used to integrate content cues into the image generation process, ensuring that the generated image and the content cues are consistent.

[0207] In some embodiments, the image generation model may include an encoding part and a decoding part. The encoding part may include multiple downsampling modules, and the decoding part may include multiple upsampling modules. Each downsampling module may include an attention processing unit and a downsampling unit, and each upsampling module may include an upsampling unit and an attention processing unit. During each denoising step, the image generation model can predict the noise residual information of the current step based on the content cue information (i.e., cue semantic feature information) and the current noise feature information.

[0208] For example, each step of the denoising process for the noise feature information can be implemented as follows: First, the downsampling module of the encoding part in the image generation model performs attention processing on the noise feature information based on the cue semantic feature information. After attention processing, downsampling processing is performed. Then, it enters the next downsampling module of the encoding part, and performs attention processing through the attention processing unit of the next downsampling module. Then, it performs downsampling processing through the downsampling unit, and so on, until the last downsampling module is reached. Through the attention processing and downsampling processing of the last downsampling module, the target feature output by the encoding part is obtained. Then, the upsampling module of the decoding part performs upsampling processing on the target feature. After upsampling processing, attention processing is performed. Then, it enters the next upsampling module for upsampling processing and attention processing, and so on, to obtain the noise residual information of the noise feature information, which is the output of the decoding part.

[0209] In some embodiments, downsampling the attention fusion features at multiple scales to obtain target features can be achieved as follows: downsampling the attention fusion features to obtain downsampled features; identifying the downsampled features as new noise feature information; returning to execute the step of performing attention processing on the noise feature information according to the content prompt information to obtain attention fusion features, until the number of iterations meets the preset condition to obtain downsampled features at multiple scales; and determining the target features based on the downsampled features at the target scale.

[0210] For example, the downsampled features at the target scale could be the smallest downsampled features among multiple scales. The feature map of the downsampled features at the target scale has the lowest resolution but contains the richest semantic information.

[0211] For example, the preset conditions can be set according to the actual situation. The total number of iterations can be determined based on the number of downsampling modules in the encoding part of the image generation model.

[0212] In some embodiments, upsampling the target features at multiple scales to obtain upsampled target features can be achieved as follows: upsample the target features to obtain upsampled features; perform attention processing on the upsampled features based on content prompts to obtain attention-processed features; upsample the attention-processed features to obtain new upsampled features; return to the step of performing attention processing on the upsampled features based on content prompts to obtain attention-processed features, until the number of iterations meets a preset condition, and the upsampled target features are obtained.

[0213] For example, the preset conditions can be set according to the actual situation. The total number of iterations can be determined based on the number of upsampling modules in the decoding part of the image generation model.

[0214] For example, the upsampling process progressively generates higher-resolution image features. An attention mechanism is integrated throughout the downsampling and upsampling processes, allowing semantic information from cues to be incorporated into the image features. Simultaneously, through the attention mechanism, the model can reference content cues at each denoising step to ensure that the generated image matches the content cues.

[0215] In some embodiments, the noise feature information is denoised based on the noise residual information to obtain an abnormal overall image. This can be achieved by: denoising the noise feature information based on the noise residual information to obtain denoised feature information; determining the denoised feature information as new noise feature information; returning to the step of using an image generation model to perform downsampling and upsampling processing on the noise feature information at multiple scales based on content prompts to predict the noise residual information of the noise feature information, until the number of iterations meets a preset condition to obtain an abnormal overall image.

[0216] For example, each iteration is a time step. At each time step, the model predicts a denoised feature based on the content prompts and the current noise features, continuously adjusting the feature values ​​of the feature points in the feature map, gradually removing noise from the noise features, so that the features gradually approach the target.

[0217] Repeat the above steps multiple times. Through multiple rounds of refinement, as time progresses, the image gradually transforms from pure noise into a meaningful image that is relevant to the content prompts.

[0218] In step 204, based on the abnormal region sub-image, the abnormal overall image is subjected to abnormal masking processing to obtain the abnormal mask image of the abnormal overall image. The sample pair composed of the abnormal overall image and the abnormal mask image is used to optimize the object anomaly detection model.

[0219] For example, compared to the overall image of the anomaly, the anomaly region sub-image is a smaller image. For instance, if the overall image of the anomaly is 256×256 pixels, then the anomaly region sub-image would be 64×64 pixels. Understandably, after the image generation model is trained as described above, it can output anomaly region sub-images based on the input object anomaly category information.

[0220] Among them, the object anomaly detection model is the model used downstream for anomaly detection.

[0221] In some embodiments, the abnormal overall image is subjected to abnormal masking processing based on the abnormal region sub-image to obtain an abnormal mask image of the abnormal overall image. This can be achieved in the following way: performing abnormal masking processing on the abnormal region sub-image to obtain an abnormal mask image of the abnormal region sub-image; generating an initial mask image based on the image size information of the abnormal overall image; and copying the abnormal mask image corresponding to the abnormal region sub-image to the corresponding position in the initial mask image based on the position information of the abnormal region sub-image in the abnormal overall image to obtain the abnormal mask image of the abnormal overall image.

[0222] For example, since most pixels in the generated abnormal region sub-image correspond to the defect region itself, this embodiment of the application can first obtain the defect mask of the abnormal region sub-image, that is, the abnormal mask image of the abnormal region sub-image.

[0223] For example, the abnormality mask processing of the abnormal region subgraph can be achieved by segmenting the abnormal region subgraph to obtain the defect mask in the abnormal region subgraph.

[0224] For example, image segmentation can be performed on defect masks in abnormal region sub-images by extracting a segmentation model from the mask. Segmentation models can be, for example, SAM (Segment Anything Model), SAM2, FastSAM, etc.

[0225] For example, image segmentation can be achieved as follows: Image features are extracted from the abnormal region sub-image using a segmentation model. Then, based on given prompts (e.g., abnormality type), the image features are decoded to determine the extent of the abnormal portion. Finally, the segmentation result is output as a mask. The pixels in the mask indicate which parts belong to the segmented target object and which belong to the background or other parts, thus obtaining the defect mask. Pixels in the defect mask that are not abnormal have a value of 0, while pixels in abnormal parts retain their original pixel values. It is understood that the size of this defect mask is the same as the size of the abnormal region sub-image.

[0226] After obtaining the anomaly mask image of the anomaly region sub-image, a completely black mask (i.e., the initial mask image) can be generated based on the width and height of the overall anomaly image. Then, based on the position coordinates of the generated anomaly region sub-image in the overall anomaly image, the anomaly mask image from the anomaly region sub-image is copied to the completely black mask at the corresponding positions, thus generating the defect mask of the overall anomaly image. In this way, a sample pair of generated samples, namely the overall anomaly image and its defect mask, is obtained. This sample pair can then be used for training downstream tasks such as detection, localization, and classification.

[0227] For example, the defect mask of the overall anomaly image is the full anomaly image with a defect mask image, that is, the anomaly mask image of the overall anomaly image, which contains the location information of the anomaly. In the anomaly mask image of the overall anomaly image, the pixel value of non-anomaly pixels is 0, while the pixel values ​​of anomaly pixels are retained. It can be understood that the anomaly mask image of the overall anomaly image is the same size as the overall anomaly image.

[0228] In some embodiments, the object anomaly detection model can be optimized based on sample pairs consisting of anomaly overall images and anomaly mask images. This can be achieved as follows: the anomaly overall image from the sample pair is used as input and fed into the object anomaly detection model to be optimized for forward inference to obtain a predicted anomaly mask; then, the difference between the predicted mask and the anomaly mask image serving as the ground truth in the sample pair is calculated using a preset loss function (e.g., cross-entropy loss, mean squared error loss, etc.); finally, based on the calculated loss value, the parameters of the object anomaly detection model are updated using a backpropagation algorithm and an optimizer (e.g., Adam). By iterating this process on multiple sample pairs, the gap between the prediction and the ground truth is continuously reduced, thereby improving the object anomaly detection model's ability to accurately identify and segment abnormal regions.

[0229] Steps 201 to 204 generate anomaly region sub-images (step 202) and anomaly overall images (step 203), ensuring both realism and controllability in local details and global scene representation. Subsequently, the generated anomaly region sub-images are used to directly perform anomaly masking on the anomaly overall image (step 204), thereby efficiently and accurately annotating pixel-level anomaly locations automatically. This integrated generation and annotation process greatly simplifies the training data preparation process, enabling the batch production of sample pairs containing anomaly overall images and anomaly mask images, providing rich and high-quality data support for model optimization of object anomaly detection models.

[0230] Compared to methods that directly generate full-image anomalies, the embodiments of this application can improve the generation effect of anomaly images and also provide highly accurate defect mask images. For methods that directly generate full-image anomalies, the small area occupied by defects leads to poor generation of the defect portion of the full-image anomaly, and accurate defect mask images cannot be obtained when directly generating full-image anomalies.

[0231] In this embodiment, an abnormal image can be generated using a serial diffusion model, and an abnormal mask can be obtained simultaneously, generating realistic and highly aligned abnormal image-mask data pairs. For example, the image generation model to be trained can be a pre-trained diffusion model, which serially generates abnormal region sub-images and the overall abnormal image, allowing the model to simultaneously learn the feature information of local defects and the overall image information, ensuring the consistency between the generated overall abnormal image and the local abnormal images.

[0232] Here, "serial" refers to training the image generation model of the abnormal region sub-image and the image generation model of the overall abnormal image in series. In fact, these two models are the same model (i.e., the image generation model mentioned above), but they are trained separately using the abnormal region sub-image and the overall abnormal image.

[0233] The trained image generation model provided in this application embodiment can be used in industrial quality inspection tasks, generating a large number of relatively realistic defects and assisting downstream defect detection tasks. The generated defect effect image is shown in Figure 2b. The black lines in Figure 2b are defect annotation boxes generated based on the method provided in this application embodiment. These defect annotation boxes can be used as annotation boxes for downstream defect detection tasks. The defects within the black lines are the generated defects. Since Figure 2b is a line diagram, the defect morphology within the defect annotation boxes is only schematically represented. For example, the morphology within defect annotation boxes-1,-3, and-6 can represent scratches or wear defects, while the morphology within defect annotation boxes-2,-4, and-5 can represent pits or pockmarks defects.

[0234] Figure 2c shows the flowchart of the small sample defect generation algorithm based on the serial diffusion model. It includes five modules: training data preprocessing module, defect map generation module, full map generation module, loss function calculation module, and mask generation module.

[0235] The training data preprocessing module is used to preprocess the training data to obtain sample pairs of overall abnormal images and abnormal region sub-images. Furthermore, corresponding text prompts (i.e., content hints) need to be added to both the overall abnormal image (i.e., the sample abnormal image in the above embodiment) and the abnormal region sub-images. For example, the original data is an overall abnormal image with a label file containing the coordinates of the defect, the category of the defect, the object category of the image, and the shooting location of the image. First, the abnormal image can be cut out using the coordinate information in the label file to obtain the corresponding abnormal region sub-image. Because the information in the abnormal region sub-image is relatively prioritized, the defect category is directly used as the content hint for the abnormal region sub-image, with a format such as: [Defect Category]; for example: HS (Scratch). For the content hint for the overall abnormal image, the format can be set as: A [Object Category] sample at pose [Shooting Location], which has defect of a [Defect Category] at [Defect Coordinates]. This training data preprocessing method yields four elements for a sample pair: the overall abnormal image, the content hints of the overall abnormal image, the abnormal region sub-image, and the content hints of the abnormal region sub-image.

[0236] For the defect map generation module, the model is first trained using anomaly region sub-images to enable it to generate such sub-images. A schematic diagram of the defect map generation module algorithm is shown in Figure 2d. For the input anomaly region sub-image, its visual features z are first extracted using a visual encoding model. Then, z is input into a pre-trained image generation model, such as a stable diffusion (SD) model. The defect name corresponding to the anomaly region sub-image is used as the text prompt (i.e., content hint information) input (e.g., HS). The SD model is used to reconstruct the visual features z, aiming to make the output as close as possible to the original input anomaly region sub-image. This reconstruction process can be, for example, by first adding noise to the visual features using the SD model's noise-adding module to obtain the noisy sub-image feature information, and then using the U-Net module in the SD model for noise prediction, followed by denoising based on the prediction results. The input to U-Net is the noisy latent features (noisy sub-image feature information) and the corresponding content hint information. The task of U-Net is to predict the noise residual in the current noise distribution. By using the noise residuals predicted by U-Net at each time step, the model can gradually remove noise and gradually approach a clear image. In each step, the model adjusts the current feature distribution based on the predicted noise residuals, thereby gradually reducing noise.

[0237] After passing through the defect image generation module, the model is able to generate local defects. This application can further train the same model serially with the overall abnormal image through the full image generation module, enabling the model to generate overall abnormal images. A schematic diagram of the full image generation module algorithm is shown in Figure 2e. For the input overall abnormal image, a visual feature z is extracted through a visual coding model. Then, z is input into the defect image generation module to call the same stable diffusion model. The content prompt information of the complete abnormal image obtained in the data preprocessing stage (e.g., "A watch sample at pose P007, which has defect of a HS at [1304,1054,1340,1289]") is used as input. The SD model is used to reconstruct the visual feature z, with the goal of making the output as close as possible to the original input overall abnormal image. Similarly, this reconstruction process can be, for example, by first adding noise to the visual features through the noise-adding module of the SD model to obtain the noise-adding overall feature information, and then performing noise prediction through the U-Net module in the SD model, and then performing denoising based on the prediction results. The input to U-Net is latent features with noise (overall feature information after adding noise) and corresponding content cues. The task of U-Net is to predict the noise residual in the current noise distribution. By using the noise residual predicted by U-Net at each time step, the model can gradually remove noise and gradually approach a clear image. In each step, the model adjusts the current feature distribution based on the predicted noise residual, thereby gradually reducing noise.

[0238] After training the defect map generation module and the full image generation module alternately with samples of all overall abnormal images and abnormal region sub-images, the trained model can generate corresponding images when inputting relevant prompt information. For example, when the input text prompt (i.e., the prompt information) is in the form of "[defect category]", the generated image is an abnormal region sub-image. When the input prompt information in the trained model is in the form of "A[object category] sample at pose[shooting point], which has defect of a[defect category] at[defect coordinates]", the generated image is a complete abnormal image. Furthermore, because the model has learned the ability to generate abnormal region sub-images, the defect portion of the generated complete abnormal image also performs very well.

[0239] For the loss function calculation module, since the model has two loss functions during training, namely LD and LS, where LD is the reconstruction loss function when generating the abnormal region sub-image (i.e., the first loss information in the above embodiment), and LS is the reconstruction loss function when generating the whole image (i.e., the second loss information in the above embodiment), during training, the loss function calculation module can sum the two reconstruction losses by weight, which is the total loss function of the model, as shown in formula (3): L=wL S +(1-w) L D (3)

[0240] Where w is the weighting parameter and L is the total loss.

[0241] After training the image generation model, this embodiment of the application can also obtain the defect mask of the overall abnormal image based on the generated abnormal region sub-image through the mask generation module. The defect mask obtained in this way has high accuracy and can be used for downstream detection tasks.

[0242] After training the image generation model, this embodiment of the application can also obtain the defect mask of the overall abnormal image based on the generated abnormal region sub-image through the mask generation module. The defect mask obtained in this way has high accuracy and can be used for downstream detection tasks.

[0243] The embodiments of this application can generate a large number of labeled defect images. In industrial AI (artificial intelligence) quality inspection, defect images that were originally unavailable can be obtained, and the labeling cost can be greatly reduced, thereby significantly improving the performance of downstream anomaly detection tasks, including defect detection, location, and classification.

[0244] As can be seen from the above, the embodiments of this application can obtain object feature description information, object anomaly location information, and object anomaly category information for a target object; generate a region sub-image based on the object anomaly category information using an image generation model to obtain an anomaly region sub-image; generate an overall image based on the object feature description information, object anomaly location information, and object anomaly category information using an image generation model to obtain an anomaly overall image; perform anomaly masking processing on the anomaly overall image based on the anomaly region sub-image to obtain an anomaly mask image of the anomaly overall image; and use the sample pair composed of the anomaly overall image and the anomaly mask image to optimize the object anomaly detection model.

[0245] This application embodiment utilizes an image generation model trained by serially generating defect anomaly sub-images and overall anomaly images to generate anomaly region sub-images and overall anomaly images, ensuring consistency between the anomaly region sub-images and the overall anomaly images, thereby generating realistic and highly aligned anomaly image-mask data sample pairs. Through the method of this application embodiment, a small amount of labeled anomaly defect data can be amplified to generate a large amount of labeled anomaly defect data, and the generated data has good realism and diversity. This data can be directly used for supervised anomaly detection training, significantly improving the performance of downstream anomaly detection tasks, including defect detection, localization, and classification.

[0246] To better implement the above methods, this application also provides an image generation model training and sample generation system. This system includes an image generation model training device 31 and a sample generation device 32. As shown in FIG3a, the image generation model training device 31 may include a first acquisition unit 3101, an image extraction unit 3102, a feature extraction unit 3103, a first reconstruction unit 3104, a second reconstruction unit 3105, and a parameter adjustment unit 3106; as shown in FIG3b, the sample generation device 32 may include a second acquisition unit 3201, a first generation unit 3202, a second generation unit 3203, and a masking unit 3204, as follows:

[0247] A. Image generation model training device 31

[0248] (1) First acquisition unit 3101;

[0249] The first acquisition unit is configured to acquire training data and an image generation model. The training data includes at least one sample abnormal image, object feature description information in the sample abnormal image, object abnormal location information, and object abnormal category information.

[0250] (2) Image extraction unit 3102;

[0251] The image extraction unit is configured to perform image extraction processing on the sample abnormal image based on the abnormal position information of the object to obtain an abnormal region sub-image.

[0252] (3) Feature extraction unit 3103;

[0253] The feature extraction unit is configured to perform feature extraction processing on the abnormal region sub-image and the sample abnormal image respectively, to obtain the sub-image feature information corresponding to the abnormal region sub-image and the overall feature information corresponding to the sample abnormal image.

[0254] (4) First reconstruction unit 3104;

[0255] The first reconstruction unit is configured to perform image reconstruction processing on the feature information of the sub-image based on the object anomaly category information using the image generation model, and obtain the reconstructed anomaly region sub-image.

[0256] In some embodiments, the first reconstruction unit includes a first noise-adding subunit and a first noise-reducing subunit, as follows:

[0257] The first noise-adding subunit is configured to add noise to the sub-graph feature information to obtain noise-adding sub-graph feature information;

[0258] The first denoising subunit is configured to perform denoising processing on the feature information of the noisy sub-image based on the object anomaly category information using the image generation model, to obtain the reconstructed anomaly region sub-image.

[0259] In some embodiments, the first denoising subunit is configured to perform downsampling and upsampling processing on the noisy sub-image feature information at multiple scales based on the object anomaly category information using the image generation model, so as to predict the noise residual information of the noisy sub-image feature information; and perform denoising processing on the noisy sub-image feature information based on the noise residual information to obtain the reconstructed abnormal region sub-image.

[0260] (5) Second reconstruction unit 3105;

[0261] The second reconstruction unit is configured to perform image reconstruction processing on the overall feature information based on the object feature description information, the object abnormal location information, and the object abnormal category information through the image generation model, to obtain a reconstructed sample abnormal image.

[0262] In some embodiments, the second reconstruction unit may include a prompting construction subunit, a second noise-adding subunit, and a second noise-reducing subunit, as follows:

[0263] The prompt construction subunit is configured to construct content prompt information for reconstructing the sample abnormal image based on the object feature description information, the object abnormal location information, and the object abnormal category information;

[0264] The second noise-adding subunit is configured to add noise to the overall feature information to obtain noise-added overall feature information.

[0265] The second denoising subunit is configured to perform denoising processing on the overall feature information after noise addition based on the content prompt information through the image generation model, so as to obtain the reconstructed abnormal sample image.

[0266] In some embodiments, the second denoising subunit is configured to perform downsampling and upsampling processing on the noisy overall feature information at multiple scales according to the content prompt information through the image generation model, so as to predict the noise residual information of the noisy overall feature information; and perform denoising processing on the noisy overall feature information according to the noise residual information to obtain the reconstructed sample abnormal image.

[0267] In some embodiments, the image generation model performs downsampling and upsampling processing on the noisy overall feature information at multiple scales based on the content prompt information to predict the noise residual information of the noisy overall feature information. This can be achieved in the following way: the image generation model performs attention processing on the noisy overall feature information based on the content prompt information to obtain attention fusion features; the attention fusion features are downsampled at multiple scales to obtain target features; the target features are upsampled at multiple scales to obtain upsampled target features; and attention processing is performed on the upsampled target features based on the content prompt information to obtain the noise residual information of the noisy overall feature information.

[0268] In some embodiments, the denoising process of the noisy overall feature information based on the noise residual information to obtain the reconstructed abnormal sample image can be achieved in the following way: denoising the noisy overall feature information based on the noise residual information to obtain denoised overall feature information; determining the denoised overall feature information as the new noisy overall feature information; returning to the step of performing downsampling and upsampling processing on the noisy overall feature information at multiple scales based on the content prompt information using the image generation model to predict the noise residual information of the noisy overall feature information, until the number of iterations meets a preset condition to obtain the reconstructed abnormal sample image.

[0269] (6) Parameter adjustment unit 3106;

[0270] The parameter adjustment unit is configured to adjust the parameters of the image generation model based on the loss information between the abnormal region sub-image and the reconstructed abnormal region sub-image, and the loss information between the sample abnormal image and the reconstructed sample abnormal image, so as to obtain the trained image generation model.

[0271] B. Sample generation device 32

[0272] (7) Second acquisition unit 3201;

[0273] The second acquisition unit is configured to acquire object feature description information, object anomaly location information, and object anomaly category information for the target object.

[0274] (8) First generation unit 3202;

[0275] The first generation unit is configured to generate a region sub-map based on the object anomaly category information using an image generation model, thereby obtaining an anomaly region sub-map.

[0276] In some embodiments, the first generation unit may include a first sampling subunit and a first denoising processing subunit, as follows:

[0277] The first sampling subunit is configured to sample from a preset noise distribution to generate noise feature information;

[0278] The first denoising processing subunit is configured to perform denoising processing on the noise feature information based on the object anomaly category information using an image generation model to obtain an anomaly region sub-image.

[0279] In some embodiments, the first denoising processing subunit is configured to perform downsampling and upsampling processing on the noise feature information at multiple scales based on the object anomaly category information using an image generation model, so as to predict the noise residual information of the noise feature information; and perform denoising processing on the noise feature information according to the noise residual information to obtain an anomaly region sub-image.

[0280] (9) Second generation unit 3203;

[0281] The second generation unit is configured to generate an overall image based on the object feature description information, the object abnormal location information, and the object abnormal category information using the image generation model, thereby obtaining an overall abnormal image.

[0282] In some embodiments, the second generation unit may include a construction subunit, a second sampling subunit, and a second denoising processing subunit, as follows:

[0283] The construction subunit is configured to construct content prompt information for image generation based on the object feature description information, the object abnormal location information, and the object abnormal category information;

[0284] The second sampling subunit is configured to sample from a preset noise distribution to generate noise feature information;

[0285] The second denoising processing subunit is configured to perform denoising processing on the noise feature information based on the content prompt information using the image generation model to obtain an abnormal overall image.

[0286] (10) Mask unit 3204;

[0287] The masking unit is configured to perform anomaly masking processing on the overall abnormal image based on the abnormal region sub-image to obtain an anomaly mask image of the overall abnormal image. The sample pair formed by the overall abnormal image and the anomaly mask image is used to optimize the object anomaly detection model.

[0288] In some embodiments, the masking unit may include a masking subunit, a generating subunit, and a copying subunit, as follows:

[0289] The masking sub-unit is configured to perform anomaly masking processing on the anomaly region sub-image to obtain an anomaly mask image of the anomaly region sub-image.

[0290] A sub-unit is configured to generate an initial mask image based on the image size information of the overall abnormal image;

[0291] The copying subunit is configured to copy the anomaly mask image corresponding to the anomaly region sub-image to the corresponding position in the initial mask image based on the position information of the anomaly region sub-image in the overall anomaly image, thereby obtaining the anomaly mask image of the overall anomaly image.

[0292] As can be seen from the above, in this embodiment of the application, the first acquisition unit 3101 acquires training data and an image generation model. The training data includes at least one sample abnormal image, object feature description information in the sample abnormal image, object abnormal location information, and object abnormal category information. The image extraction unit 3102 performs image extraction processing on the sample abnormal image based on the object abnormal location information to obtain an abnormal region sub-image. The feature extraction unit 3103 performs feature extraction processing on the abnormal region sub-image and the sample abnormal image respectively to obtain sub-image feature information corresponding to the abnormal region sub-image and overall feature information corresponding to the sample abnormal image. The first reconstruction unit 3104 then... The image generation model, based on the object anomaly category information, performs image reconstruction processing on the feature information of the sub-image to obtain a reconstructed anomaly region sub-image; the second reconstruction unit 3105, through the image generation model, performs image reconstruction processing on the overall feature information based on the object feature description information, the object anomaly location information, and the object anomaly category information to obtain a reconstructed sample anomaly image; the parameter adjustment unit 3106 adjusts the parameters of the image generation model according to the loss information between the anomaly region sub-image and the reconstructed anomaly region sub-image, and the loss information between the sample anomaly image and the reconstructed sample anomaly image, to obtain a trained image generation model;

[0293] Alternatively, the second acquisition unit 3201 acquires object feature description information, object anomaly location information, and object anomaly category information for the target object; the first generation unit 3202 generates a region sub-image based on the object anomaly category information using an image generation model to obtain an anomaly region sub-image; the second generation unit 3203 generates an overall image based on the object feature description information, the object anomaly location information, and the object anomaly category information using the image generation model to obtain an anomaly overall image; the masking unit 3204 performs anomaly masking processing on the anomaly overall image based on the anomaly region sub-image to obtain an anomaly mask image of the anomaly overall image, and the sample pair composed of the anomaly overall image and the anomaly mask image is used to optimize the object anomaly detection model.

[0294] This application embodiment allows for the sequential learning of small defect anomaly images and overall anomaly images during training. This enables the model to simultaneously learn the feature information of local defects and the overall image information, ensuring consistency between the overall and local anomaly images generated by the trained image generation model. This allows for the construction of realistic and highly aligned anomaly image-mask data sample pairs. The method of this application embodiment can amplify a small amount of labeled anomaly defect data to generate a large amount of labeled anomaly defect data. The generated data exhibits good realism and diversity, and can be directly used for supervised anomaly detection training, significantly improving the performance of downstream anomaly detection tasks, including defect detection, localization, and classification.

[0295] This application also provides an electronic device, as shown in FIG4, which illustrates the structural schematic diagram of the electronic device involved in this application embodiment. The electronic device may be a terminal or a server, etc.

[0296] The electronic device may include components such as a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, and an input unit 404. Those skilled in the art will understand that the electronic device structure shown in FIG. 4 does not constitute a limitation on the electronic device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:

[0297] The processor 401 is the control center of the electronic device, connecting various parts of the device via various interfaces and lines. It executes software programs and / or modules stored in the memory 402, and calls data stored in the memory 402, to perform various functions and process data. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 401.

[0298] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.

[0299] The electronic device also includes a power supply 403 that supplies power to the various components. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 403 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0300] The electronic device may also include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0301] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. In the embodiments of this application, the processor 401 in the electronic device loads the executable files corresponding to the processes of one or more applications into the memory 402 according to the following instructions, and the processor 401 runs the applications stored in the memory 402 to realize various functions, as follows:

[0302] The process involves acquiring training data and an image generation model. The training data includes at least one sample anomalous image, object feature description information, object anomalous location information, and object anomalous category information within the sample anomalous image. Based on the object anomalous location information, image segmentation is performed on the sample anomalous image to obtain an anomalous region sub-image. Feature extraction is then performed on both the anomalous region sub-image and the sample anomalous image to obtain sub-image feature information corresponding to the anomalous region sub-image and overall feature information corresponding to the sample anomalous image. Using the image generation model, image reconstruction is performed on the sub-image feature information based on the object anomalous category information to obtain a reconstructed anomalous region sub-image. Using the image generation model, image reconstruction is performed on the overall feature information based on the object feature description information, the object anomalous location information, and the object anomalous category information to obtain a reconstructed sample anomalous image. Finally, the parameters of the image generation model are adjusted based on the loss information between the anomalous region sub-image and the reconstructed anomalous region sub-image, as well as the loss information between the sample anomalous image and the reconstructed sample anomalous image, to obtain a trained image generation model.

[0303] Alternatively, acquire object feature description information, object anomaly location information, and object anomaly category information for the target object; generate anomaly region sub-images based on the object anomaly category information using an image generation model; generate an overall image based on the object feature description information, the object anomaly location information, and the object anomaly category information using the same image generation model; perform anomaly masking on the overall anomaly image based on the anomaly region sub-images to obtain anomaly mask images of the overall anomaly image; the sample pair consisting of the overall anomaly image and the anomaly mask image is used to optimize the object anomaly detection model.

[0304] The implementation methods for each of the above operations can be found in the description of the preceding embodiments.

[0305] As can be seen from the above, the embodiments of this application can acquire training data and an image generation model. The training data includes at least one sample abnormal image, object feature description information, object abnormal location information, and object abnormal category information in the sample abnormal image. Based on the object abnormal location information, image extraction processing is performed on the sample abnormal image to obtain an abnormal region sub-image. Feature extraction processing is performed on the abnormal region sub-image and the sample abnormal image respectively to obtain sub-image feature information corresponding to the abnormal region sub-image and overall feature information corresponding to the sample abnormal image. Through the image generation model, based on the object abnormal category information, image reconstruction processing is performed on the sub-image feature information to obtain a reconstructed abnormal region sub-image. Through the image generation model, based on the object feature description information, the object abnormal location information, and the object abnormal category information, image reconstruction processing is performed on the overall feature information to obtain a reconstructed sample abnormal image. According to the loss information between the abnormal region sub-image and the reconstructed abnormal region sub-image, and the loss information between the sample abnormal image and the reconstructed sample abnormal image, the parameters of the image generation model are adjusted to obtain a trained image generation model.

[0306] Alternatively, acquire object feature description information, object anomaly location information, and object anomaly category information for the target object; generate anomaly region sub-images based on the object anomaly category information using an image generation model; generate an overall image based on the object feature description information, the object anomaly location information, and the object anomaly category information using the same image generation model; perform anomaly masking on the overall anomaly image based on the anomaly region sub-images to obtain anomaly mask images of the overall anomaly image; the sample pair consisting of the overall anomaly image and the anomaly mask image is used to optimize the object anomaly detection model.

[0307] In this embodiment, during training, the generation of anomalous region sub-images and the overall anomalous image are learned sequentially. This allows the image generation model to simultaneously learn the feature information of local anomalies and the overall image information, ensuring the consistency between the overall anomalous image and the local anomalous image generated by the trained image generation model. This enables the construction of realistic and highly aligned anomalous image-mask data sample pairs. The method provided in this embodiment can amplify a small amount of labeled anomalous defect data to generate a large amount of labeled anomalous defect data. The generated data exhibits good realism and diversity, and can be directly used for supervised anomaly detection training, significantly improving the performance of downstream anomaly detection tasks, including defect detection, localization, and classification.

[0308] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0309] Therefore, embodiments of this application provide a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute steps in any of the image generation model training methods and sample generation methods provided in embodiments of this application. For example, the instructions can execute the following steps:

[0310] The process involves acquiring training data and an image generation model. The training data includes at least one sample anomalous image, object feature description information, object anomalous location information, and object anomalous category information within the sample anomalous image. Based on the object anomalous location information, image segmentation is performed on the sample anomalous image to obtain an anomalous region sub-image. Feature extraction is then performed on both the anomalous region sub-image and the sample anomalous image to obtain sub-image feature information corresponding to the anomalous region sub-image and overall feature information corresponding to the sample anomalous image. Using the image generation model, image reconstruction is performed on the sub-image feature information based on the object anomalous category information to obtain a reconstructed anomalous region sub-image. Using the image generation model, image reconstruction is performed on the overall feature information based on the object feature description information, the object anomalous location information, and the object anomalous category information to obtain a reconstructed sample anomalous image. Finally, the parameters of the image generation model are adjusted based on the loss information between the anomalous region sub-image and the reconstructed anomalous region sub-image, as well as the loss information between the sample anomalous image and the reconstructed sample anomalous image, to obtain a trained image generation model.

[0311] Alternatively, acquire object feature description information, object anomaly location information, and object anomaly category information for the target object; generate anomaly region sub-images based on the object anomaly category information using an image generation model; generate an overall image based on the object feature description information, the object anomaly location information, and the object anomaly category information using the same image generation model; perform anomaly masking on the overall anomaly image based on the anomaly region sub-images to obtain anomaly mask images of the overall anomaly image; the sample pair consisting of the overall anomaly image and the anomaly mask image is used to optimize the object anomaly detection model.

[0312] For details on how to implement each of the above operations, please refer to the description of the embodiments above.

[0313] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0314] Since the instructions stored in the computer-readable storage medium can execute the steps of any of the image generation model training methods and sample generation methods provided in the embodiments of this application, the beneficial effects that any of the image generation model training methods and sample generation methods provided in the embodiments of this application can achieve can be realized, as detailed in the preceding embodiments.

[0315] This application provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps in the image generation model training method and sample generation method provided in this application.

[0316] The foregoing has provided a detailed description of an image generation model training method, a sample generation method, and related equipment provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of the embodiments of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the methods and core ideas of the embodiments of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the embodiments of this application. Therefore, the content of this specification should not be construed as a limitation on the embodiments of this application.

Claims

1. An image generation model training method, applied to electronic devices, comprising: Acquire training data and an image generation model, wherein the training data includes at least one sample anomalous image, object feature description information in the sample anomalous image, object anomalous location information, and object anomalous category information; Based on the abnormal location information of the object, the abnormal sample image is processed by image extraction to obtain an abnormal region sub-image; Feature extraction processing is performed on the abnormal region sub-image and the sample abnormal image respectively to obtain the sub-image feature information corresponding to the abnormal region sub-image and the overall feature information corresponding to the sample abnormal image; Using the image generation model, based on the object anomaly category information, image reconstruction processing is performed on the feature information of the sub-image to obtain the reconstructed anomaly region sub-image; Using the image generation model, based on the object feature description information, the object abnormal location information, and the object abnormal category information, image reconstruction processing is performed on the overall feature information to obtain a reconstructed sample abnormal image; Based on the loss information between the abnormal region sub-image and the reconstructed abnormal region sub-image, and the loss information between the sample abnormal image and the reconstructed sample abnormal image, the parameters of the image generation model are adjusted to obtain the trained image generation model.

2. The method according to claim 1, wherein, The step of performing image reconstruction processing on the sub-image feature information based on the object anomaly category information using the image generation model to obtain the reconstructed anomaly region sub-image includes: The subgraph feature information is subjected to noise addition processing to obtain the noisy subgraph feature information; Using the image generation model, the feature information of the noisy sub-image is denoised based on the object anomaly category information to obtain the reconstructed anomaly region sub-image.

3. The method according to any one of claims 1 to 2, wherein, The step of using the image generation model to denoise the feature information of the noisy sub-image based on the object anomaly category information to obtain the reconstructed anomaly region sub-image includes: Based on the object anomaly category information, the image generation model performs downsampling and upsampling processing on the noisy sub-image feature information at multiple scales to predict the noise residual information of the noisy sub-image feature information. Based on the noise residual information, the feature information of the noisy sub-image is denoised to obtain the reconstructed abnormal region sub-image.

4. The method according to any one of claims 1 to 3, wherein, The step of using the image generation model to perform downsampling and upsampling processing on the noisy sub-image feature information at multiple scales based on the object anomaly category information, in order to predict the noise residual information of the noisy sub-image feature information, includes: The image generation model is used to extract features from the abnormal category information of the object to obtain prompt semantic feature information. Based on the semantic feature information of the prompt, attention processing is performed on the feature information of the noisy subgraph to obtain attention fusion features; The attention fusion features are downsampled at multiple scales to obtain the target features; The target features are upsampled at multiple scales to obtain upsampled target features; Based on the semantic feature information of the prompt, attention processing is performed on the upsampled target features to obtain the noise residual information of the noise-added sub-image feature information.

5. The method according to any one of claims 1 to 4, wherein, The upsampling process on the target features at multiple scales to obtain the upsampled target features includes: The target features are upsampled to obtain the upsampled features; Based on the semantic feature information of the prompt, attention processing is performed on the upsampled features to obtain attention-processed features; The attention-processed features are upsampled to obtain new upsampled features; Return to the step of performing attention processing on the upsampled features based on the semantic feature information of the prompt to obtain attention-processed features, until the number of iterations meets the preset condition to obtain the upsampled target features.

6. The method according to any one of claims 1 to 5, wherein, The step of performing attention processing on the upsampled target features based on the prompt semantic feature information to obtain the noise residual information of the noise feature information includes: The upsampled target features are used as the query vector, and the prompt semantic features are used as the key vector and value vector; Attention scores are determined based on the query vector and the key vector, and attention weights are obtained by normalizing the attention scores. The noise residual information is obtained by weighting and summing the value vector according to the attention weights.

7. The method according to claim 1, wherein, The step of using the image generation model to perform image reconstruction processing on the overall feature information based on the object feature description information, the object anomaly location information, and the object anomaly category information to obtain a reconstructed sample anomaly image includes: Based on the object feature description information, the object anomaly location information, and the object anomaly category information, content prompt information for reconstructing the sample anomaly image is constructed. The overall feature information is subjected to noise addition processing to obtain the noise-added overall feature information; Using the image generation model, and based on the content prompts, the overall feature information after noise addition is denoised to obtain the reconstructed abnormal sample image.

8. The method according to any one of claims 1 or 7, wherein, The step of using the image generation model to denoise the overall feature information after adding noise, based on the content prompt information, to obtain the reconstructed abnormal sample image includes: The image generation model performs downsampling and upsampling processing on the noisy overall feature information at multiple scales based on the content prompt information, in order to predict the noise residual information of the noisy overall feature information. Based on the noise residual information, the overall feature information after noise addition is denoised to obtain the reconstructed abnormal image of the sample.

9. The method according to any one of claims 1, 7 to 8, wherein, The step of using the image generation model to perform downsampling and upsampling processing on the noisy overall feature information at multiple scales based on the content prompt information, in order to predict the noise residual information of the noisy overall feature information, includes: Using the image generation model, attention processing is performed on the overall feature information after adding noise, based on the content prompt information, to obtain attention fusion features; The attention fusion features are downsampled at multiple scales to obtain the target features; The target features are upsampled at multiple scales to obtain upsampled target features; Based on the content prompt information, attention processing is performed on the upsampled target features to obtain the noise residual information of the overall feature information after adding noise.

10. The method according to any one of claims 1, 7 to 9, wherein, The step of denoising the overall feature information after noise addition based on the noise residual information to obtain the reconstructed sample anomaly image includes: Based on the noise residual information, the noise-added overall feature information is denoised to obtain the denoised overall feature information. The denoised overall feature information is determined as the new denoised overall feature information; Return to the step of performing the image generation model, based on the content prompt information, to perform downsampling and upsampling processing on the overall feature information after adding noise at multiple scales to predict the noise residual information of the overall feature information after adding noise, until the number of iterations meets the preset condition, and the reconstructed sample abnormal image is obtained.

11. A sample generation method, applied to an electronic device, comprising: Obtain object feature description information, object anomaly location information, and object anomaly category information for the target object; An abnormal region sub-image is generated by using an image generation model based on the object anomaly category information. The image generation model is trained using the method described in any one of claims 1 to 10. Using the image generation model, based on the object feature description information, the object anomaly location information, and the object anomaly category information, an overall image is generated to obtain an overall image of the anomaly. Based on the abnormal region sub-image, the abnormal overall image is subjected to abnormal masking processing to obtain an abnormal mask image of the abnormal overall image. The sample pair composed of the abnormal overall image and the abnormal mask image is used to optimize the object anomaly detection model.

12. The method according to claim 11, wherein, The step of generating an anomaly region sub-map using an image generation model based on the object anomaly category information includes: Noise feature information is generated by sampling from a preset noise distribution; By using an image generation model, the noise feature information is denoised based on the object anomaly category information to obtain an anomaly region sub-image.

13. The method according to any one of claims 11 to 12, wherein, The step of using an image generation model to denoise the noise feature information based on the object anomaly category information to obtain an anomaly region sub-image includes: Using an image generation model, based on the object anomaly category information, the noise feature information is downsampled and upsampled at multiple scales to predict the noise residual information of the noise feature information; Based on the noise residual information, the noise feature information is denoised to obtain an abnormal region sub-image.

14. The method according to claim 11, wherein, The step of generating an overall image based on the object feature description information, the object anomaly location information, and the object anomaly category information using the image generation model to obtain an overall anomaly image includes: Based on the object feature description information, the object abnormality location information, and the object abnormality category information, content prompt information for image generation is constructed. Noise feature information is generated by sampling from a preset noise distribution; Using the image generation model, and based on the content prompt information, the noise feature information is denoised to obtain an abnormal overall image.

15. The method according to any one of claims 11 to 14, wherein, The step of performing anomaly masking processing on the overall abnormal image based on the abnormal region sub-image to obtain the anomaly mask image of the overall abnormal image includes: Anomaly masking processing is performed on the abnormal region sub-image to obtain the anomaly mask image of the abnormal region sub-image; Based on the image size information of the overall abnormal image, an initial mask image is generated; Based on the position information of the abnormal region sub-image in the overall abnormal image, the abnormal mask image corresponding to the abnormal region sub-image is copied to the corresponding position in the initial mask image to obtain the abnormal mask image of the overall abnormal image.

16. An image generation model training device, comprising: The first acquisition unit is configured to acquire training data and an image generation model. The training data includes at least one sample abnormal image, object feature description information in the sample abnormal image, object abnormality location information, and object abnormality category information. The image extraction unit is configured to perform image extraction processing on the sample abnormal image based on the abnormal position information of the object to obtain an abnormal region sub-image; The feature extraction unit is configured to perform feature extraction processing on the abnormal region sub-image and the sample abnormal image respectively, to obtain the sub-image feature information corresponding to the abnormal region sub-image and the overall feature information corresponding to the sample abnormal image; The first reconstruction unit is configured to perform image reconstruction processing on the feature information of the sub-image based on the object anomaly category information using the image generation model to obtain the reconstructed anomaly region sub-image. The second reconstruction unit is configured to perform image reconstruction processing on the overall feature information based on the object feature description information, the object abnormal location information, and the object abnormal category information through the image generation model, to obtain a reconstructed sample abnormal image; The parameter adjustment unit is configured to adjust the parameters of the image generation model based on the loss information between the abnormal region sub-image and the reconstructed abnormal region sub-image, and the loss information between the sample abnormal image and the reconstructed sample abnormal image, so as to obtain the trained image generation model.

17. A sample generation apparatus, comprising: The second acquisition unit is configured to acquire object feature description information, object anomaly location information, and object anomaly category information for the target object. The first generation unit is configured to generate a region sub-map based on the object anomaly category information using an image generation model, thereby obtaining an anomaly region sub-map. The image generation model is trained using the method described in any one of claims 1 to 10. The second generation unit is configured to generate an overall image based on the object feature description information, the object abnormal location information, and the object abnormal category information through the image generation model, thereby obtaining an abnormal overall image. The masking unit is configured to perform anomaly masking processing on the overall abnormal image based on the abnormal region sub-image to obtain an anomaly mask image of the overall abnormal image. The sample pair formed by the overall abnormal image and the anomaly mask image is used to optimize the object anomaly detection model.

18. An electronic device comprising a memory and a processor; the memory storing an application program, the processor being configured to run the application program within the memory to perform steps in the image generation model training method of any one of claims 1 to 10 or the sample generation method of any one of claims 11 to 15.

19. A computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to perform steps in the image generation model training method of any one of claims 1 to 10 or the sample generation method of any one of claims 11 to 15.

20. A computer program product comprising a computer program or instructions that, when executed by a processor, implement the steps of the image generation model training method of any one of claims 1 to 10 or the sample generation method of any one of claims 11 to 15.