Data generation method, training data generation method and related equipment
By performing noise processing on the first image and adjusting the neural network model, the problem of uncontrollable image background in the prior art is solved, the consistency between the original image and the generated image background is achieved, and the adaptability and picture quality of the data generation technology are improved.
Patent Information
- Application Number
- CN202311584881.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-23
- Publication Date
- 2025-05-23
AI Technical Summary
Existing data generation technology is difficult to accurately control the image background, resulting in the generated background being uncontrollable and the consistency between the original image and the generated image cannot be maintained.
By performing noise processing on the first image extracted image features, combining the target description information and the neural network model, the noise features are adjusted to achieve background control, and a second image maintaining consistency is generated.
It realizes precise control of the image background, maintains the consistency of the background between the original image and the generated image, expands the adaptability of data generation technology in different application scenarios, and improves the quality of generated images.
Smart Images

Figure CN120032199A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology. Specifically, the present application relates to a data generation method, a training data generation method and related equipment. Background Art
[0002] Data generation is a process of generating new data by analyzing and processing existing data using computer algorithms and artificial intelligence techniques.
[0003] In the field of image processing, the existing data generation technology is to obtain the required clear image by performing noise estimation and then denoising on the pure noise image. However, this technology cannot effectively control the image background. For example, a rabbit wearing sunglasses can be generated through data generation technology, but in the generation process, since each input is random noise, the background generated each time is variable and uncontrollable. For example, the background generated for the first time may be green grass, and the background generated for the second time may be a river. Therefore, it is difficult to accurately control the background through the existing data generation technology, and it is also impossible to maintain the consistency of the background between the original image and the generated image. Summary of the invention
[0004] In order to solve at least one of the above technical problems, the present application embodiment provides a data generation method, a training data generation method and related devices. The technical solution is as follows:
[0005] In a first aspect, an embodiment of the present application provides a data generation method, comprising:
[0006] Performing noise processing on the image features extracted from the first image to obtain the first image features with noise;
[0007] Extract features from target description information to obtain description features; the target description information includes first description information for describing the content of the target image;
[0008] Obtaining a noise feature based on the first image feature and the description feature through a neural network model;
[0009] Performing denoising processing on the first image feature based on the noise feature to obtain a denoised second image feature;
[0010] A second image including the target image content is generated based on the second image feature.
[0011] In a feasible embodiment, obtaining the noise feature based on the first image feature and the description feature through a neural network model includes:
[0012] Acquire target position information, where the information is used to indicate the position of the target image content in the second image;
[0013] The first image feature and the description feature are processed by a neural network model, and the processing result is adjusted based on the target position information to obtain a noise feature.
[0014] In a feasible embodiment, the target location information includes at least one of the following:
[0015] Identification information extracted from the first image and used to identify a location where the target image content is generated;
[0016] Preset coordinate information;
[0017] a third image with a region logo;
[0018] The second description information is extracted from the target description information and is used to indicate a generation position of the target image content.
[0019] In a feasible embodiment, the data generation method is performed by a non-trained diffusion model including the neural network model;
[0020] The processing of the first image feature and the description feature by the neural network model and adjusting the processing result based on the target position information to obtain the noise feature includes:
[0021] Through the neural network model, attention processing is performed on the first image feature and the description feature based on the attention mechanism, and the attention processing result is adjusted based on the target position information.
[0022] In a feasible embodiment, the neural network model includes at least one attention module, and the attention module includes a cross-attention unit and a self-attention unit connected in series;
[0023] The performing attention processing on the first image feature and the description feature based on the attention mechanism through the neural network model, and adjusting the attention processing result based on the target position information, includes:
[0024] Determining, by the cross-attention unit, a first probability value corresponding to an element in a cross-attention feature matrix based on a first query matrix corresponding to the first image feature and a first key matrix corresponding to the descriptive feature; the first probability value indicates a probability of generating the target image content at a position corresponding to the corresponding element;
[0025] Determining, by the self-attention unit, a second probability value corresponding to an element in a self-attention feature matrix based on a second query matrix corresponding to the first image feature and a second key matrix corresponding to the first image feature, wherein the second probability value indicates a probability of generating the target image content at a position corresponding to the corresponding element;
[0026] Adjust at least one of the first probability value and the second probability value based on the target position information to obtain an adjusted attention processing result;
[0027] Noise prediction is performed based on the adjusted attention processing result to obtain noise features.
[0028] In a feasible embodiment, adjusting at least one of the first probability value and the second probability value based on the target location information includes at least one of the following:
[0029] Adjusting a first probability value corresponding to an element in the cross-attention feature matrix based on a position indicated by the target position information;
[0030] Adjusting a second probability value corresponding to an element in the self-attention feature matrix based on the position indicated by the target position information;
[0031] The first probability value corresponding to the element in the cross-attention feature matrix is adjusted based on the target position information and a preset loss function.
[0032] In a feasible embodiment, the step of adjusting the first probability value corresponding to the element in the cross-attention feature matrix based on the position indicated by the target position information, and the step of adjusting the second probability value corresponding to the element in the self-attention feature matrix based on the position indicated by the target position information, include at least one of the following:
[0033] For an element in the attention feature matrix that is located in the target position information indication area, increasing the probability value of the element;
[0034] For an element in the attention feature matrix that is outside the area indicated by the target position information, the probability value of the element is reduced.
[0035] In a feasible embodiment, the step of adjusting the first probability value corresponding to the element in the cross-attention feature matrix based on the position indicated by the target position information, and the step of adjusting the second probability value corresponding to the element in the self-attention feature matrix based on the position indicated by the target position information, include:
[0036] For an element in the attention feature matrix that is located within the target position information indication area, adjusting the probability value of the element based on a preset first coefficient;
[0037] For an element in the attention feature matrix that is outside the target position information indication area, adjusting the probability value of the element based on a preset second coefficient;
[0038] The first coefficient is greater than the second coefficient.
[0039] In a feasible embodiment, adjusting the first probability value corresponding to the element in the cross-attention feature matrix based on the target position information and a preset loss function includes:
[0040] Based on the target position information and a preset loss function, adjusting the first probability value corresponding to the element in the cross-attention feature matrix by a gradient descent method;
[0041] Among them, the loss function indicates the ratio of the sum of the first probability values of the elements in the cross-attention feature matrix located in the target position information indication area to the sum of the first probability values of all elements in the cross-attention feature matrix.
[0042] In a second aspect, an embodiment of the present application provides a method for generating training data, comprising:
[0043] Acquire a target image used to generate a sample image and annotation information corresponding to the target image;
[0044] By using the data generation method provided in the first aspect, a sample image including a sample object is generated based on the target image and the annotation information;
[0045] Obtaining a training data pair based on the sample image and the annotation information;
[0046] The annotation information includes category information and position information for annotating sample objects in the sample image; and the target image content includes the sample objects.
[0047] In a feasible embodiment, the target image includes a defect-free image, and the sample image is an image with defects generated based on the target image; the category information of the sample object includes information indicating the defect type to which the sample object belongs; and the method further includes:
[0048] An image processing model for quality inspection is trained using the training data.
[0049] In a third aspect, an embodiment of the present application provides a data generating device, including:
[0050] A feature noise adding module, used for performing noise adding processing on the image features extracted from the first image to obtain the first image features with noise;
[0051] A text processing module, used to extract features from target description information to obtain description features; the target description information includes first description information used to describe the content of the target image;
[0052] A network prediction module, configured to obtain a noise feature based on the first image feature and the description feature through a neural network model;
[0053] A feature denoising module, configured to perform denoising processing on the first image feature based on the noise feature to obtain a denoised second image feature;
[0054] An image generating module is used to generate a second image including the target image content based on the second image feature.
[0055] In a fourth aspect, an embodiment of the present application provides a device for generating training data, including:
[0056] An image acquisition module, used to acquire a target image used to generate a sample image and annotation information corresponding to the target image;
[0057] A sample generating module, configured to generate a sample image including a sample object based on the target image and the annotation information through the data generating device provided by the third aspect;
[0058] A training data generation module, used to obtain a training data pair based on the sample image and the annotation information;
[0059] The annotation information includes category information and position information for annotating sample objects in the sample image; and the target image content includes the sample objects.
[0060] In a fifth aspect, an embodiment of the present application provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory, and the processor executes the computer program to implement the steps of the method provided in the first aspect or the second aspect above.
[0061] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the method provided in the first aspect or the second aspect are implemented.
[0062] In a seventh aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the method provided in the first or second aspect above.
[0063] The beneficial effects of the technical solution provided by the embodiment of the present application are:
[0064] In the first aspect, the embodiment of the present application provides a data generation method. Specifically, firstly, the image features extracted from the first image are subjected to noise processing to obtain the first image features with noise, that is, the present application can add noise to the first image to obtain image information with noise data; then, features can be extracted from the target description information including the first description information (used to describe the target image content) to obtain description features, and then the first image features and the description features can be processed through a neural network model to obtain noise features, and then the first image can be subjected to denoising processing based on the noise features to obtain second image features that keep the image background consistent, and finally a second image including the target image content can be generated based on the second image features. Compared with the prior art, the present application can accurately control the background and maintain the consistency of the background between the original image and the generated image.
[0065] In the second aspect, the embodiment of the present application provides a method for generating training data. Specifically, when a target image used to generate a sample image and annotation information corresponding to the target image are obtained, the data generation method provided in the first aspect can be used to generate a sample image including a sample object based on the target image and the annotation information, and then a training data pair is obtained based on the sample image and the annotation information; wherein the annotation information includes category information and position information for annotating the sample object in the sample image, and the target image content in the first aspect includes the sample object. Compared with the prior art, the implementation of the present application can maintain the consistency of the background between the target image and the sample image, so that the training data pair composed of the sample image and the annotation information can be used as a training sample for processing downstream model training tasks, thereby realizing rapid generation of training samples in a short time and reducing the cost of annotating the target image. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in describing the embodiments of the present application are briefly introduced below.
[0067] Figure 1 A flowchart of a data generation method provided in an embodiment of the present application;
[0068] Figure 2 A flowchart of a method for generating training data provided in an embodiment of the present application;
[0069] Figure 3 An operation flow chart based on a diffusion model provided in an embodiment of the present application;
[0070] Figure 4 A computational visualization diagram of an attention mechanism provided in an embodiment of the present application;
[0071] Figure 5A schematic diagram of a defect-free image provided by an embodiment of the present application;
[0072] Figure 6 A schematic diagram of a defective image provided by an embodiment of the present application;
[0073] Figure 7a A schematic diagram of a first image including identification information provided in an embodiment of the present application;
[0074] Figure 7b A schematic diagram of a third image with a region identification provided in an embodiment of the present application;
[0075] Figure 8a A schematic diagram of an image feature matrix provided in an embodiment of the present application;
[0076] Figure 8b A schematic diagram of an attention feature matrix provided in an embodiment of the present application;
[0077] Fig. 9 A schematic diagram of a data generating device provided in an embodiment of the present application;
[0078] Fig.10 A schematic diagram of a training data generation device provided in an embodiment of the present application;
[0079] Fig.11 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0080] The embodiments of the present application are described below in conjunction with the drawings in the present application. It should be understood that the implementation methods described below in conjunction with the drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions of the embodiments of the present application.
[0081] It will be understood by those skilled in the art that, unless specifically stated, the singular forms "one", "said", and "the" used herein may also include plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present application refer to that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements and / or components, but do not exclude the implementation as other features, information, data, steps, operations, elements, components and / or combinations thereof supported by the technical field. It should be understood that when we say that an element is "connected" or "coupled" to another element, the one element may be directly connected or coupled to the other element, or it may refer to that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term, for example, "A and / or B" may be implemented as "A", or as "B", or as "A and B".
[0082] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0083] The embodiments of the present application involve artificial intelligence (AI), which is a theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science, which attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that the machines have the functions of perception, reasoning and decision-making. Artificial intelligence technology is a comprehensive discipline that involves a wide range of fields, including both hardware-level technology and software-level technology. Basic artificial intelligence technologies generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction system, mechatronics, etc. Among them, pre-trained models are also called large models and basic models. After fine-tuning, they can be widely used in downstream tasks in various major directions of artificial intelligence. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0084] Specifically, embodiments of the present application relate to computer vision technology (Computer Vision, CV) and machine learning (Machine Learning, ML) technology.
[0085] Among them, computer vision is a science that studies how to make machines "see". To put it more specifically, it refers to machine vision such as using cameras and computers to replace human eyes to identify and measure targets, and further perform graphic processing to make computer processing into images that are more suitable for human eye observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish an artificial intelligence system that can obtain information from images or multidimensional data. Large model technology has brought important changes to the development of computer vision technology. Pre-trained models in the visual field such as swin-transformer, ViT, V-MOE, and MAE can be quickly and widely applied to downstream specific tasks after fine tuning. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous positioning and map construction, and other technologies, as well as common biometric recognition technologies such as face recognition and fingerprint recognition.
[0086] Among them, machine learning is a multi-disciplinary interdisciplinary subject, involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning by teaching. The pre-trained model is the latest development of deep learning, which integrates the above technologies.
[0087] The present application specifically relates to data generation technology. In the existing data generation technology, a pure noise image is subjected to noise estimation and then denoised to obtain a clear image. Since the noise is predicted and denoised based on random Gaussian noise, the consistency of the background between the original image and the generated image cannot be guaranteed, resulting in limited application scenarios of the generated data and no improvement in the quality of the generated image.
[0088] In response to at least one technical problem existing in the above-mentioned prior art, an embodiment of the present application proposes a data generation scheme, which can add noise to the original image to obtain image information with noise data, and then when the image is subsequently denoised, an image with a consistent background can be generated, thereby expanding the adaptability of the generated data in different application scenarios and effectively improving the quality of the generated images.
[0089] The following describes several exemplary embodiments to illustrate the technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application. It should be noted that the following embodiments can refer to, draw on or combine with each other, and the same terms, similar features and similar implementation steps in different embodiments will not be described repeatedly.
[0090] The following describes the data generation method in the embodiment of the present application.
[0091] Specifically, the execution subject of the method provided in the embodiment of the present application may be a terminal or a server; the terminal (also referred to as a device) may be a smart phone, a tablet computer, a laptop computer, a desktop computer, an intelligent voice interaction device (such as an intelligent speaker), a wearable electronic device (such as a smart watch), a vehicle terminal, an intelligent home appliance (such as a smart TV), an AR / VR device, etc., but is not limited thereto. The server may be an independent physical server, or a server cluster or a distributed system (such as a distributed cloud storage system) composed of multiple physical servers, or a cloud server that provides cloud computing and cloud storage services.
[0092] Specifically, Figure 1 As shown, the data generation method includes steps S101 to S105:
[0093] Step S101: performing noise processing on the image features extracted from the first image to obtain the noisy first image features.
[0094] The first image may be any image (also called a picture); and it may also be a specific image to adapt to different scenarios, which is described in conjunction with the following scenario examples (only as an example, not limiting the embodiments of the present application):
[0095] For example, in the scenario of constructing training data, if it is necessary to obtain a sample image for model training based on the first image, the first image can be an original image to be processed and does not have the target object that the model needs to recognize compared with the sample image. If the processing task of the model is to identify whether the product shown in the image has defects, the first image can be a defect-free image, and the sample image is an image with defects generated based on the first image.
[0096] For example, in an image editing scenario, if it is necessary to add a target object based on a first image to obtain an edited image, the first image may be the image to be edited.
[0097] Alternatively, if Figure 3As shown, for the first input image, image features can be extracted through a picture editor. Considering that if a higher-definition image is required in data generation, high-dimensional features can be extracted through the picture editor. High-dimensional features can be more structured and complex features combined on the basis of representing the image from aspects such as image texture, color, and shape. It can be understood that the richer the information contained in the image features, the higher the accuracy and stability of the results obtained from image processing.
[0098] Optionally, the image features extracted from the first image may form an image feature matrix, in which each element represents the image content of one or more pixels in the image. Figure 8a To illustrate with an example, Figure 8a A 9*9 image feature matrix is shown, in which the feature indicated by each element corresponds to a pixel in the first image. Assuming that the image feature with green color is currently extracted from the first image, the element 1 in the third row and third column of the feature matrix indicates that the color of the pixel corresponding to the position is green, and the element 0 in the third row and first column of the feature matrix indicates that the color of the pixel corresponding to the position is non-green.
[0099] Optionally, when performing noise processing on the extracted image features, such as Figure 3 As shown, Gaussian noise (which may be Gaussian noise of a digital image) may be input into the image feature extracted from the first image to obtain the noisy first image feature. Other noise addition techniques may also be used to perform noise addition processing, which is not limited in this application.
[0100] Step S102: extracting features from target description information to obtain description features; the target description information includes first description information for describing target image content.
[0101] The target description information may include first description information input for describing the target image content to be generated, and the expression form may be text, voice, etc.; if the target description information is text information, such as Figure 3 As shown, feature extraction can be performed on the input text through a text editor to obtain descriptive features; optionally, if the target description information is voice information, feature extraction can be performed through a text editor after converting the voice information into text information, or feature extraction can be performed using a voice editor to obtain descriptive features.
[0102] Optionally, the target description information may also be an image, such as describing the image content to be generated in the form of an image, which may be implemented by inputting an image, and then extracting features through an image editor, and knowing the target image content to be generated from the feature representation. Assuming that the first image is a rabbit on the grass, the target description information may be an image including carrots. Accordingly, the target description information is used to guide the model to generate an image including carrots and rabbits on the grass on the first image.
[0103] The target image content includes content intended to be generated on the first image, such as Figure 5 and Figure 6 The target image content may include the content relative to Figure 5 For Figure 6 Optionally, when the target description information is text information, it can be a text description such as "there is a running Bichon Frise on the grass", and if the first image is a grass image, the target description information can guide the model to generate a running Bichon Frise on the original first image.
[0104] Step S103: Obtain noise features based on the first image features and the description features through a neural network model.
[0105] Alternatively, if Figure 3 As shown, the neural network model can be independent or a part of the overall network. Optionally, it can be a neural network model based on the attention mechanism, such as Unet, or other models for noise estimation, which is not limited in this application.
[0106] Among them, Unet adopts a U-shaped network structure, and the backbone can be divided into a feature extraction network (encoder, which obtains feature maps of each level through downsampling such as convolution and pooling) and a feature fusion network (decoder, which fuses feature maps of each level with feature maps obtained by deconvolution through jump connection). The last layer performs noise prediction by calculating the loss. Unet can fuse shallow and deep features of the image.
[0107] Alternatively, if Figure 3 As shown, the input of the neural network may include the first image feature and the description feature, and the output may be the predicted noise. During the processing, the network model may be iterated, looped, etc.
[0108] Step S104: performing denoising processing on the first image feature based on the noise feature to obtain a denoised second image feature.
[0109] Optionally, the denoising process may be performed using relevant technologies, which is not limited in this application. The denoised second image feature may be a high-dimensional feature, such as Figure 3As shown, it is suitable for extracting high-dimensional features. After noise estimation and denoising, high-dimensional features for outputting generated images can be obtained (at this time, the feature represents the content corresponding to the description feature).
[0110] Step S105: generating a second image including the target image content based on the second image feature.
[0111] Alternatively, if Figure 3 As shown, the second image feature can be decoded by a picture decoder to output a second image including the target image content. Accordingly, the second image is equivalent to an image generated based on the first image.
[0112] In an embodiment of the present application, the data generation method provided can be implemented by a diffusion model. Diffusion refers to continuously adding noise to a picture to eventually generate pure noise data, and the data generation principle of the diffusion model is the inverse process of diffusion, that is, continuously denoising the pure noise data to eventually obtain a clear picture. In one example, the diffusion model can generate clear picture data similar to that obtained by shooting from pure noise data. Among them, the diffusion method can achieve the effect of generating pictures from text. For example, given a description as a generation guide condition, the diffusion model can continuously denoise the noisy data to generate the corresponding picture.
[0113] Compared with the prior art, the embodiment of the present application adds noise to an existing image to obtain image information with noisy data. On this basis, when the image is subsequently denoised, the consistency of the image background can be maintained, which can effectively expand the scenarios used by data generation technology. For example, the solution can be used to quickly generate training data and reduce the cost of manual annotation of images; it can also be used in image editing scenarios to improve the quality and efficiency of image generation. Among them, background consistency means that the background part in the second image is as consistent as possible with the content of the first image, except for the location where the target image content is generated.
[0114] In a feasible embodiment, in step S103, a noise feature is obtained based on the first image feature and the description feature through a neural network model, including steps A1-A2:
[0115] Step A1: Acquire target position information, where the target position information is used to indicate the position of the target image content in the second image.
[0116] Among them, the target location information may include at least one of identification information extracted from the first image for identifying the target image content generation location, preset coordinate information, a third image with an area identification, and second description information extracted from the target description information for indicating the target image content generation location.
[0117] Optionally, the target position information is used to guide the model to generate target image content at a specified position. The generation position indicated by the target position information (which may be one or more points, areas, etc.) may be determined in the following multiple ways, as follows:
[0118] Method 1: Use the first image to indicate the generating position; Figure 7a As shown, it is possible to Figure 5 As shown in FIG. 1 , the identification information of the generated position is generated on the image. For example, the operating object can determine the identification information by selecting an area in the image. Optionally, the identification information can be as follows: Figure 7a The rectangular frame shown may also be an area mark of any shape, a point mark of any position, etc. The same first image may include one or more identification information. When obtaining the target position information in this way, feature extraction may be performed through an image editor, and the extracted identification features are used to indicate the generation position of the target image content.
[0119] Method 2, indicating the generation position through preset coordinate information; if the center point coordinates and the width and height (x, y, w, h) of the corresponding area are given, the coordinate information can guide the model to generate the target image content in the corresponding area in the form of a rectangular box (in the field of image detection, the target object can be annotated by the annotation format of the rectangular box). Optionally, the generation position can also be determined by giving the corresponding coordinate information of the point located on the diagonal line. That is, in an embodiment of the present application, the method of indicating the generation position through coordinate information is not limited. For example, in one example, it may be necessary to generate multiple small black dots, then the preset coordinate information may include multiple independent coordinate information, that is, the target position information and the target image content to be generated may have an associated relationship.
[0120] Method 3: indicating the generation location by a third image with a region identifier; Figure 7b As shown, the image size can be changed by comparing the image size with the first image (such as Figure 5 Region marking is performed in an image that is consistent with the first image (as shown), and an image with a region identifier is obtained as input as the target position information indicating the generation position. Optionally, the region identifier can be of any shape, such as a polygon or an irregular shape. One or more region identifiers can also be included in the same image. Optionally, when the target position information is provided in this way, feature extraction can be performed through a picture editor, and the extracted features are used to indicate the generation position of the target image content. Optionally, the third image with the region identifier is different in size from the first image. The size of the image with the region identifier can be adjusted to be consistent with the size of the first image by normalization, image cropping, image ratio adjustment, etc., with the first image as a reference, and then the generation position indicated by the corresponding region identifier is determined.
[0121] Method 4: Indicating the generation position through target description information; Optionally, the target description information may include second description information for indicating the generation position of the target image content. For example, when the target description information is text information, it may be "generate curtains at coordinates (x, y, w, h)", and feature information describing the generation position may also be extracted from the target description information. Optionally, the target description information does not necessarily need to indicate the generation position through coordinate information, and it may also indicate the generation position of the target image content by describing the relative position relationship with the existing objects in the first image. For example, if the first image is an image showing a rabbit on the grass, the target description information may be "generate a carrot smaller than the rabbit at the position of the rabbit's mouth", and it can be determined that the generation position referred to by the target description information is the rabbit's mouth. Optionally, when providing the target position information in this way, feature extraction may be performed through the same editor, that is, features including descriptions of the target image content and descriptions of the generation position may be extracted from the target description information. Features describing different contents may also be extracted through different editors.
[0122] Step A2: Processing the first image feature and the description feature through a neural network model, and adjusting the processing result based on the target position information to obtain a noise feature.
[0123] Optionally, in step A2, a fusion calculation is performed on three inputs, namely, the first image feature, the description feature and the target position information (such as the coordinates of the rectangular frame), and the output is the estimated noise. Subsequently, the estimated noise can be continuously removed from the noisy feature until meaningful feature data is obtained, and finally the image is restored using a picture decoder to obtain a second image.
[0124] Optionally, the self-attention feature matrix and the cross-attention feature matrix can be used to process the first image features (noisy data) and the description features, and then the values of the self-attention feature matrix and the cross-attention feature matrix can be modified using the target position information to achieve regional control.
[0125] In an embodiment of the present application, the region where the target image content is generated can be controlled by inputting the target position information, that is, the position where the target image content is generated in the image can be effectively controlled; specifically, region controllable means that the position of the generated target in the picture can be controlled, for example, in an industrial parts product picture, a dent defect is specified to be generated at the edge of the part in the lower left corner of the picture.
[0126] In a feasible embodiment, the data generation method provided by the embodiment of the present application is performed by a non-trained diffusion model including a neural network model.
[0127] In step A2, the first image feature and the description feature are processed by a neural network model, and the processing result is adjusted based on the target position information to obtain a noise feature, including step A21:
[0128] Step A21: Perform attention processing on the first image feature and the description feature based on the attention mechanism through a neural network model, and adjust the attention processing result based on the target position information.
[0129] Among them, non-training means that there is no need to train the model. The diffusion model is directly used without the need for training data or modifying the training parameters. By modifying the calculation method of the model, regional controllable data generation can be achieved.
[0130] Among them, the attention mechanism is a calculation method that gives three features, namely query Q, key K and value V, and outputs Assuming that Q, K and V come from the same data source, such as they are all image features, the corresponding attention is called self-attention; assuming that Q, K and V come from different data sources, such as Q represents the first image feature and K and V represent description features, the corresponding attention is called cross-attention.
[0131] Optionally, a visualization of the attention mechanism computation, such as Figure 4 As shown in the figure, Q and K are high-dimensional feature matrices. Using matrix multiplication, we get the output result A = QxK. Then we perform a softmax operation on A, that is, we normalize the matrix numerically. The value of softmax(A) represents the probability value. On this basis, the adjustment of the attention processing result based on the target position information can be understood as adjusting the operation result (probability value) of softmax(A).
[0132] Alternatively, if Figure 3 As shown, the diffusion model adopted in the embodiment of the present application may include network structures such as an image editor (used to extract feature information of an input image), a text editor (used to extract descriptive features of an input text), a neural network (such as Unet), and an image decoder (used to decode high-dimensional features after denoising to generate a final output image).
[0133] Optionally, the neural network model includes at least one attention module, and the attention module may include a cross attention unit and a self attention unit connected in series. In step A21, the neural network model is used to perform attention processing on the first image feature and the description feature based on the attention mechanism, and the attention processing result is adjusted based on the target position information to obtain the noise feature, including steps A211-A214:
[0134] Step A211: Through the cross-attention unit, based on the first query matrix corresponding to the first image feature and the first key matrix corresponding to the descriptive feature, determine the first probability value corresponding to the element in the cross-attention feature matrix; the first probability value indicates the probability of generating the target image content at the position corresponding to the corresponding element.
[0135] Optionally, if the first query matrix Q represents an image feature matrix (the data source is the first image feature), and the first key matrix K represents a description feature matrix (the data source is the description feature), then the above matrix softmax(A) can be called a cross-attention feature matrix. It can be understood that the first probability value is a probability corresponding to the calculation result of Q and K after the softmax operation, which is used to indicate the probability of generating the target image content at the position corresponding to the element of the cross-attention feature matrix.
[0136] Among them, the elements in the cross-attention feature matrix indicate the feature values calculated based on the input features in the cross-attention calculation. Optionally, in an embodiment of the present application, the probability values obtained by the above-mentioned softmax (A) processing can constitute the cross-attention feature matrix, that is, the elements in the current cross-attention feature matrix can be the first probability values. In one example, the position corresponding to the element can be determined based on the position of each first image feature in the first image. For example, the cross-attention feature matrix is obtained by processing the first query matrix and the first key matrix by matrix multiplication. Accordingly, the first query matrix corresponding to the first image feature obtained by the attention mechanism can determine the position of each different first image feature in the cross-attention feature matrix.
[0137] Step A212: Through the self-attention unit, based on the second query matrix corresponding to the first image feature and the second key matrix corresponding to the first image feature, determine the second probability value corresponding to the element in the self-attention feature matrix; the second probability value indicates the probability of generating the target image content at the position corresponding to the corresponding element.
[0138] Optionally, if the second query matrix Q and the second key matrix K represent different image feature matrices (the data source is the first image feature), the above matrix softmax(A) can be called a self-attention feature matrix. It can be understood that the second probability value is the corresponding probability after the calculation of Q and K and the calculation result is subjected to the softmax operation, which is used to indicate the probability of generating the target image content at the position corresponding to the element of the self-attention feature matrix.
[0139] Among them, the elements in the self-attention feature matrix indicate the feature values calculated based on the input features in the self-attention calculation. Optionally, in an embodiment of the present application, the probability values obtained by the above-mentioned softmax (A) processing can constitute the self-attention feature matrix, that is, the elements in the current self-attention feature matrix can be second probability values. In one example, the position corresponding to the element can be determined based on the position of each first image feature in the first image. For example, the self-attention feature matrix is obtained by processing the second query matrix and the second key matrix by matrix multiplication. Accordingly, the second query matrix and the second key matrix corresponding to the first image features obtained by the attention mechanism can determine the position of each different first image feature in the self-attention feature matrix.
[0140] Step A213: Adjust at least one of the first probability value and the second probability value based on the target position information to obtain an adjusted attention processing result.
[0141] Optionally, in the embodiment of the present application, in addition to adjusting the cross-attention feature matrix, the self-attention feature matrix can also be adjusted to increase the probability of generating the target image content at the position indicated by the target position information. Different adjustment schemes are provided to adapt to different attention modes. This will be described in detail in subsequent embodiments.
[0142] For example, for the above cross-attention feature matrix and self-attention feature matrix, refer to Figure 8b The attention feature matrix diagram shown in FIG. 1 , in which each element indicates the probability of generating the target image content at its corresponding position; Figure 8b In the illustrated case, the probability of generating the target image content within the region indicated by the bold dashed box is greater than the probability of generating the target image content outside the region.
[0143] The attention feature matrix is determined based on the first image features, that is, the image feature matrix formed by the features extracted from the first image has a one-to-one correspondence with the elements in the attention feature matrix. Figure 8a and Figure 8b As shown, Figure 8a The element 0 in the first row and first column of corresponds to pixel A in the first image. Figure 8b The element 0.1 in the first row and first column of corresponds to pixel A in the first image; Figure 8a The element 1 in the third row and fourth column of corresponds to pixel B in the first image. Figure 8b The element 0.96 in the third row and fourth column corresponds to pixel B in the first image; and so on.
[0144] Step A214: Perform noise prediction based on the adjusted attention processing result to obtain noise features.
[0145] In the embodiment of the present application, the target position information can be used to indicate the generation position of the target image content, so as to realize regional controllability. It can be understood that the above adjustment of the first probability value and the second probability value is based on Q and K, and in the calculation of the attention mechanism, it can essentially refer to That is, the softmax operation is performed on the result calculated based on Q and K, and then V is combined to obtain the final attention processing result. Optionally, the neural network finally outputs the predicted noise information.
[0146] In a feasible embodiment, adjusting at least one of the first probability value and the second probability value based on the target location information in step A23 includes at least one of steps A231 to A233:
[0147] Step A231: Adjust the first probability value corresponding to the element in the cross-attention feature matrix based on the position indicated by the target position information.
[0148] Optionally, in the adjustment of the first probability value, in order to better guide the model to generate the target image content at the position indicated by the target position information, at least one of the probability values corresponding to the elements in the cross-attention feature matrix that are located in the area indicated by the target position information and the probability values corresponding to the elements outside the area can be adjusted. In one example, the adjustment of the first probability value can be achieved by increasing the probability value corresponding to the elements in the area and reducing at least one of the probability values corresponding to the elements outside the area. This operation can increase the gap between the probability values corresponding to the elements in the area and outside the area, so as to more clearly distinguish the possibility of generating the target image content at different positions, thereby guiding the model to generate the target image content at a position with a higher probability value.
[0149] Compared with the prior art, the embodiments of the present application can adjust the value (first probability value) of the cross-attention feature matrix by directly multiplying the coefficient, which can effectively guide the model to generate the target image content at the position specified by the target position information while improving the processing efficiency.
[0150] Optionally, in step A231, adjusting the first probability value corresponding to the element in the cross-attention feature matrix based on the position indicated by the target position information includes at least one of steps A231a-A231b:
[0151] Step A231a: For the elements in the cross-attention feature matrix that are located in the target position information indication area, adjust the first probability value of the element based on a preset first coefficient.
[0152] Step A231b: For the elements in the cross-attention feature matrix that are outside the target position information indication area, adjust the first probability value of the element based on a preset second coefficient.
[0153] The first coefficient is greater than 1, and the second coefficient is less than 1.
[0154] In one example, assuming that the target position information is reflected as a rectangular box in the first image, the value inside the rectangular box (first probability value) can be multiplied by a first coefficient a (a>1), and the value outside the rectangular box (first probability value) can be multiplied by a second coefficient b (b<1). By adjusting the value of the cross-attention matrix cross-attention, the difference between the probability values corresponding to the elements inside the rectangular box and the probability values corresponding to the elements outside the rectangular box is increased, thereby better guiding the model to generate the target image content in the specified rectangular box (at the position with a large probability value).
[0155] Optionally, the probability value may be adjusted by addition, subtraction, etc. For example, for the probability value corresponding to the element in the target position information indication area, the first value may be added to the original probability value to obtain the adjusted probability value (the probability value may be limited to 1 at most); for the probability value corresponding to the element outside the target position information indication area, the second value may be subtracted from the original probability value to obtain the adjusted probability value (the probability value may be limited to 0 at most).
[0156] Step A232: Adjust the second probability value corresponding to the element in the self-attention feature matrix based on the position indicated by the target position information.
[0157] Optionally, in the adjustment of the second probability value, in order to better guide the model to generate the target image content at the position indicated by the target position information, at least one of the probability values corresponding to the elements in the area indicated by the target position information and the probability values corresponding to the elements outside the area in the self-attention feature matrix can be adjusted. In one example, the adjustment of the second probability value can be achieved by increasing the probability value corresponding to the elements in the area and reducing the probability value corresponding to the elements outside the area. This operation can increase the gap between the probability values corresponding to the elements in the area and outside the area, so as to more clearly distinguish the possibility of generating the target image content at different positions, thereby guiding the model to generate the target image content at a position with a higher probability value.
[0158] Optionally, step A232 may also adjust the second probability value by directly multiplying by a coefficient, adding, subtracting, etc. as in step A231, such as adjusting the second probability value of the element in the self-attention feature matrix located within the target position information indication area based on a preset first coefficient (greater than 1); adjusting the second probability value of the element in the self-attention feature matrix located outside the target position information indication area based on a preset second coefficient (less than 1). It is understandable that the preset coefficient used to adjust the cross-attention feature matrix and the preset coefficient used to adjust the self-attention feature matrix can be different values, as long as the first coefficient used is an arbitrary value greater than 1 and the second coefficient is an arbitrary value less than 1.
[0159] Step A233: Adjust the first probability value corresponding to the element in the cross-attention feature matrix based on the target position information and a preset loss function.
[0160] Compared with the prior art, the embodiment of the present application uses a gradient descent algorithm (BatchGradient Descent) through a preset function to adjust the cross-attention feature matrix, so that the probability of generating the target image content at the position indicated by the target position information increases, and guides the model to generate the content corresponding to the target description information at the corresponding position. Among them, the gradient descent algorithm is an optimization algorithm in machine learning, which continuously adjusts the model parameters in an iterative manner to minimize the loss function. Optionally, the gradient descent algorithm can include batch gradient descent, stochastic gradient descent, mini-batch gradient descent, momentum gradient descent, adaptive gradient descent, etc.
[0161] Optionally, in step A233, adjusting the first probability value corresponding to the element in the cross-attention feature matrix based on the target position information and a preset loss function includes step A233a:
[0162] Step A233a: Based on the target position information and a preset loss function, adjust the first probability value corresponding to the element in the cross-attention feature matrix by a gradient descent method.
[0163] Among them, the loss function indicates the ratio of the sum of the first probability values of the elements in the cross-attention feature matrix located in the target position information indication area to the sum of the first probability values of all elements in the cross-attention feature matrix.
[0164] Optionally, the target position information (such as a rectangular box) can be used to guide the probability value of the self-attention feature matrix, which can be expressed as shown in the following formula (1):
[0165]
[0166] In formula (1), The numerator in represents the addition of all probability values in the area indicated by the target position information corresponding to the self-attention matrix, and the denominator represents the addition of all probability values in the self-attention feature matrix; B represents the area indicated by the target position information, such as the rectangular box (specific area in the matrix) that it can indicate; u in the numerator belongs to B, that is, data is selected as the numerator within the area indicated by B, and there is no position restriction in the denominator, that is, all data in the entire matrix are added; where u indicates any element of the plane (there may be multiple planes), and i is used to distinguish which plane it corresponds to. It should be noted that the "plane" referred to here is a unit content determined based on the description feature, and each word corresponds to a page. Example description: "A cat is catching a mouse in the upper left corner of the image [x, y, w, h]", then "a", "cat", "catch" and "mouse" will generate a page, and through the iteration of the loss function, the model can be guided to generate a cat at the [x, y, w, h] position of the page corresponding to the "cat", and there will be no restriction on the [x, y, w, h] of other pages. Among them, γ indicates the number of layers. In the network structure of the embodiment of the present application, there may be multiple self-attention layers to form a block (such as stacking 12 layers), and multiple cross-attention layers (such as stacking 12 layers) to form a block, and γ is used to indicate a cross-attention layer among the 12 layers.
[0167] Among them, the value of the self-attention feature matrix can be changed by taking the derivative of the above loss function and using the gradient descent method, as shown in the following formula (2):
[0168]
[0169] In the above formula (2), the purpose is to adjust the feature Z t , so as to minimize the part shown in formula (1), that is, maximum, When the matrix value is the largest, it is concentrated in the area indicated by the target position information; t η represents the learning rate, which is multiplied by the gradient to get the change, feature Z t Subtract the change to get the new Z t , so that the value of formula (1) is minimized.
[0170] Among them, the diffusion model can continuously reduce the loss value during multiple iterations, which means that the value of the specified generation position will increase, thereby guiding the model to generate the target image content indicated by the target description information in the area indicated by the target position information.
[0171] Optionally, the loss function may also be obtained by referring to related technologies, and the present application is not limited to the loss function calculation method provided in the above embodiment.
[0172] Combine the following Figure 2 The method for generating training data provided in the embodiment of the present application is described.
[0173] Considering that the existing methods cannot maintain the consistency of the background and the position of the generated target is greatly different from the given position, it is impossible to realize the generation of data sets through data generation technology. In the embodiment of the present application, the data generation method provided in the above embodiment can achieve the consistency of the background and improve the effect of regional control in a non-training way. On this basis, the embodiment of the present application also proposes a downstream solution based on the above data generation method, such as obtaining a training set based on the data generation method, reducing the cost of constructing a training set, and improving the development efficiency of downstream tasks. Among them, downstream tasks can be detection, segmentation, classification and other tasks in the field of computer vision, such as predicting the rectangular box coordinates of a specific target in a picture, predicting the area of a specific target in a picture, predicting the category of a specific target in a picture, etc.; processing this downstream task requires a large amount of data to train the neural network model parameters.
[0174] Specifically, the execution subject of the method provided in the embodiment of the present application may be a terminal or a server; the terminal (also referred to as a device) may be a smart phone, a tablet computer, a laptop computer, a desktop computer, an intelligent voice interaction device (such as an intelligent speaker), a wearable electronic device (such as a smart watch), a vehicle terminal, an intelligent home appliance (such as a smart TV), an AR / VR device, etc., but is not limited thereto. The server may be an independent physical server, or a server cluster or a distributed system (such as a distributed cloud storage system) composed of multiple physical servers, or a cloud server that provides cloud computing and cloud storage services.
[0175] Specifically, the method for generating training data includes the following steps S201 to S203:
[0176] Step S201: obtaining a target image for generating a sample image and annotation information corresponding to the target image.
[0177] Step S202: Generate a sample image including a sample object based on the target image and the annotation information through a data generation method.
[0178] Step S203: obtaining a training data pair based on the sample image and the annotation information.
[0179] The annotation information includes category information and position information for annotating sample objects in the sample image; and the target image content includes the sample objects.
[0180] When constructing a training data set, the data set may include pictures (sample images) and annotation information. The annotations for different downstream tasks are different. In the detection field, the annotation format may be a rectangular box (or an area of any other shape), that is, given the coordinates of the center point of the target, the corresponding width, height and target category (x, y, w, h, c). Since the rectangular box and description information are given, the information (x, y, w, h, c) of the corresponding target (sample object) is known and can be used as annotation information. Because the pictures in the data set can be obtained using the data generation method provided in the above embodiment, the embodiments of the present application can obtain pictures and annotation results. It can be understood that multiple pictures and annotation data pairs constitute a data set.
[0181] Optionally, in an application scenario, such as industrial product quality inspection, there are many types of defects, such as scratches, bruises, bruises, coarse wires, collapsed edges, glue holes, missed milling, dents, etc. on the product surface, but the amount of data for each type is very small. Therefore, when downstream tasks in the detection field rely heavily on the amount of data, the prediction accuracy is very low. In order to reduce the development cycle of downstream tasks, reduce annotation costs, and improve prediction accuracy, the embodiments of the present application can use the data generation method provided in the above embodiments to generate training data in a short period of time. Accordingly, the target image can include a defect-free image, and the sample image can be a defective image generated based on the target image; the category information of the sample object can include information indicating the defect type to which the sample object belongs.
[0182] In addition, the above-mentioned method for generating training data also includes: training an image processing model for quality inspection using the training data.
[0183] In the field of industrial AI quality inspection, such as the quality inspection scenario of equipment, equipment such as laptop shells, tablet shells, electronic watch shells, etc., it is inevitable that the shell will be bruised, crushed, scratched, etc. during the industrial production process. Generally, the shell is manually inspected to see if it has defects, or a machine quality inspection method is used to identify whether the shell has defects. The embodiment of the present application takes into account that the detection algorithm requires a large amount of defect data for continuous learning, thereby continuously improving the detection accuracy. However, the problem of small amount of defect data and uneven types of defects leads to the need for long-term material collection, continuous enrichment of the amount of data in the training set and the diversity of defects, and most of the data sets formed in the prior art need to be manually labeled, and the labeling efficiency is low and the labeling quality is poor. On this basis, if the data generation method provided in the above embodiment is adopted, any defect data of a specified location can be generated in a short period of time, saving the time spent on collecting data. Under the condition of sufficient data volume, the accuracy of defect detection and prediction can be effectively improved, and the development cycle of downstream tasks can be reduced. In addition, the quality of labeling can be effectively guaranteed and the cost of labeling can be reduced.
[0184] It should be noted that in the optional embodiments of the present application, the data involved (such as the first image, target description information, target location information and other related data), when the above embodiments of the present application are applied to specific products or technologies, need to obtain permission or consent from the user, and the collection, use and processing of the relevant data need to comply with the relevant laws, regulations and standards of the relevant countries and regions. In other words, if the embodiments of the present application involve data related to the object, these data need to be obtained with the authorization and consent of the object and in compliance with the relevant laws, regulations and standards of the country and region.
[0185] The present application embodiment provides a data generating device, such as Fig. 9 As shown, the data generating device 100 may include: a feature denoising module 101, a text processing module 102, a network prediction module 103, a feature denoising module 104 and an image generating module 105.
[0186] Among them, the feature noise addition module 101 is used to perform noise addition processing on the image features extracted from the first image to obtain noisy first image features; the text processing module 102 is used to extract features from the target description information to obtain description features; the target description information includes first description information for describing the target image content; the network prediction module 103 is used to obtain noise features based on the first image features and the description features through a neural network model; the feature denoising module 104 is used to perform denoising processing on the first image features based on the noise features to obtain denoised second image features; the image generation module 105 is used to generate a second image including the target image content based on the second image features.
[0187] In a feasible embodiment, when the network prediction module 103 is used to obtain the noise feature based on the first image feature and the description feature through the neural network model, it is specifically used to:
[0188] Acquire target position information, where the information is used to indicate the position of the target image content in the second image;
[0189] Processing the first image feature and the description feature through a neural network model, and adjusting the processing result based on the target position information to obtain a noise feature;
[0190] In a feasible embodiment, the target location information includes at least one of the following:
[0191] Identification information extracted from the first image and used to identify a location where the target image content is generated;
[0192] Preset coordinate information;
[0193] a third image with a region logo;
[0194] The second description information is extracted from the target description information and is used to indicate a generation position of the target image content.
[0195] In a feasible embodiment, the data generation method is performed by a non-trained diffusion model including the neural network model;
[0196] The processing of the first image feature and the description feature by the neural network model and adjusting the processing result based on the target position information to obtain the noise feature includes:
[0197] Through the neural network model, attention processing is performed on the first image feature and the description feature based on the attention mechanism, and the attention processing result is adjusted based on the target position information.
[0198] In a feasible embodiment, the neural network model includes at least one attention module, and the attention module includes a cross-attention unit and a self-attention unit connected in series.
[0199] When the network prediction module 103 is used to process the first image feature and the description feature through a neural network model and adjust the processing result based on the target position information to obtain the noise feature, it is specifically used to:
[0200] Determining, by the cross-attention unit, a first probability value corresponding to an element in a cross-attention feature matrix based on a first query matrix corresponding to the first image feature and a first key matrix corresponding to the descriptive feature; the first probability value indicates a probability of generating the target image content at a position corresponding to the corresponding element;
[0201] Determining, by the self-attention unit, a second probability value corresponding to an element in a self-attention feature matrix based on a second query matrix corresponding to the first image feature and a second key matrix corresponding to the first image feature, wherein the second probability value indicates a probability of generating the target image content at a position corresponding to the corresponding element;
[0202] Adjust at least one of the first probability value and the second probability value based on the target position information to obtain an adjusted attention processing result;
[0203] Noise prediction is performed based on the adjusted attention processing result to obtain noise features.
[0204] In a feasible embodiment, the network prediction module 103 is used to adjust at least one of the first probability value and the second probability value based on the target location information, including at least one of the following:
[0205] Adjusting a first probability value corresponding to an element in the cross-attention feature matrix based on a position indicated by the target position information;
[0206] Adjusting a second probability value corresponding to an element in the self-attention feature matrix based on the position indicated by the target position information;
[0207] The first probability value corresponding to the element in the cross-attention feature matrix is adjusted based on the target position information and a preset loss function.
[0208] In a feasible embodiment, the network prediction module 103 is used to perform the step of adjusting the first probability value corresponding to the element in the cross-attention feature matrix based on the position indicated by the target position information, and the step of adjusting the second probability value corresponding to the element in the self-attention feature matrix based on the position indicated by the target position information, including at least one of the following:
[0209] For an element in the attention feature matrix that is located in the target position information indication area, increasing the probability value of the element;
[0210] For an element in the attention feature matrix that is outside the area indicated by the target position information, the probability value of the element is reduced.
[0211] In a feasible embodiment, the network prediction module 103 is used to perform the step of adjusting the first probability value corresponding to the element in the cross-attention feature matrix based on the position indicated by the target position information, and the step of adjusting the second probability value corresponding to the element in the self-attention feature matrix based on the position indicated by the target position information, including:
[0212] For an element in the attention feature matrix that is located within the target position information indication area, adjusting the probability value of the element based on a preset first coefficient;
[0213] For an element in the attention feature matrix that is outside the target position information indication area, adjusting the probability value of the element based on a preset second coefficient;
[0214] The first coefficient is greater than the second coefficient.
[0215] In a feasible embodiment, the network prediction module 103 is used to adjust the first probability value corresponding to the element in the cross-attention feature matrix based on the target position information and the preset loss function, including:
[0216] Based on the target position information and a preset loss function, adjusting the first probability value corresponding to the element in the cross-attention feature matrix by a gradient descent method;
[0217] Among them, the loss function indicates the ratio of the sum of the first probability values of the elements in the cross-attention feature matrix located in the target position information indication area to the sum of the first probability values of all elements in the cross-attention feature matrix.
[0218] The present application embodiment provides a device for generating training data, such as Fig.10 As shown, the training data generating device 200 may include: an image acquisition model 201 , a sample generating module 202 , and a training data generating module 203 .
[0219] The image acquisition module 201 is used to acquire a target image for generating a sample image and annotation information corresponding to the target image; the sample generation module 202 is used to generate a sample image including a sample object based on the target image and the annotation information through a data generation device; the training data generation module 203 is used to obtain a training data pair based on the sample image and the annotation information;
[0220] The annotation information includes category information and position information for annotating sample objects in the sample image; and the target image content includes the sample objects.
[0221] In a feasible embodiment, the target image includes a defect-free image, and the sample image is an image with defects generated based on the target image; the category information of the sample object includes information indicating the defect type to which the sample object belongs; and the apparatus 200 further includes:
[0222] The training module is used to train an image processing model for quality inspection using the training data.
[0223] The device of the embodiments of the present application can execute the method provided by the embodiments of the present application, and the implementation principles are similar. The actions performed by each module in the device of each embodiment of the present application correspond to the steps in the method of each embodiment of the present application. For the detailed functional description of each module of the device, please refer to the description in the corresponding method shown in the previous text, which will not be repeated here.
[0224] The modules involved in the embodiments described in this application can be implemented by software. The name of the module does not limit the module itself in some cases. For example, the feature noise addition module can also be described as "a module for performing noise addition processing on the feature information extracted from the first image to obtain the first image feature with noise", "a first module", etc.
[0225] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0226] In an embodiment of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to implement the steps of the data generation method. Compared with the related art, the following can be achieved:
[0227] In the first aspect, the embodiment of the present application provides a data generation method. Specifically, firstly, the image features extracted from the first image are subjected to noise processing to obtain the first image features with noise, that is, the present application can add noise to the first image to obtain image information with noise data; then, features can be extracted from the target description information describing the target image content to be generated to obtain description features, and then the first image features and the description features can be processed by a neural network to obtain noise features, and then the first image features can be subjected to denoising processing based on the noise features to obtain second image features that keep the image background consistent, and finally a second image including the target image content can be generated based on the second image features. Compared with the prior art, the present application can accurately control the background and maintain the consistency of the background between the original image and the generated image.
[0228] In the second aspect, the embodiment of the present application provides a method for generating training data. Specifically, when a target image used to generate a sample image and annotation information corresponding to the target image are obtained, the data generation method provided in the first aspect can be used to generate a sample image including a sample object based on the target image and the annotation information, and then a training data pair is obtained based on the sample image and the annotation information; wherein the annotation information includes category information and position information for annotating the sample object in the sample image, and the target image content in the first aspect includes the sample object. Compared with the prior art, the implementation of the present application can maintain the consistency of the background between the target image and the sample image, so that the training data pair composed of the sample image and the annotation information can be used as a training sample for processing downstream model training tasks, thereby realizing rapid generation of training samples in a short time and reducing the cost of annotating the target image.
[0229] In an alternative embodiment, an electronic device is provided, such as Fig.11 As shown, Fig.11 The electronic device 4000 shown includes: a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, such as through a bus 4002. Optionally, the electronic device 4000 may also include a transceiver 4004, which may be used for data interaction between the electronic device and other electronic devices, such as data transmission and / or data reception. It should be noted that in actual applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present application.
[0230] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It may implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of this application. Processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0231] The bus 4002 may include a path to transmit information between the above components. The bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.11 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0232] The memory 4003 may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compressed optical disk, laser disk, optical disk, digital versatile disk, Blu-ray disk, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, without limitation herein.
[0233] The memory 4003 is used to store the computer program for executing the embodiment of the present application, and the execution is controlled by the processor 4001. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the steps shown in the above method embodiment.
[0234] Among them, electronic equipment includes but is not limited to: terminals and servers.
[0235] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps and corresponding contents of the aforementioned method embodiment can be implemented.
[0236] The embodiment of the present application also provides a computer program product, including a computer program, which can implement the steps and corresponding contents of the aforementioned method embodiment when executed by a processor.
[0237] The terms "first", "second", "third", "fourth", "1", "2", etc. (if any) in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than that shown or described in the drawings.
[0238] It should be understood that, although each operation step is indicated by arrows in the flowchart of the embodiment of the present application, the implementation order of these steps is not limited to the order indicated by the arrows. Unless clearly stated herein, in some implementation scenarios of the embodiment of the present application, the implementation steps in each flowchart can be performed in other orders according to demand. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on actual implementation scenarios. Some or all of these sub-steps or stages may be executed at the same time, and each sub-step or stage in these sub-steps or stages may also be executed at different times respectively. In different scenarios of execution time, the execution order of these sub-steps or stages may be flexibly configured according to demand, and the embodiment of the present application does not limit this.
[0239] The above is only an optional implementation method for some implementation scenarios of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the technical concept of the solution of the present application, other similar implementation methods based on the technical ideas of the present application are also within the protection scope of the embodiments of the present application.
Claims
1. A data generation method, It is characterized in that include: Performing noise processing on the image features extracted from the first image to obtain the first image features with noise; Extract features from target description information to obtain description features; The target description information includes first description information for describing the target image content; Obtaining a noise feature based on the first image feature and the description feature through a neural network model; Performing denoising processing on the first image feature based on the noise feature to obtain a denoised second image feature; A second image including the target image content is generated based on the second image feature.
2. The method according to claim 1, It is characterized in that The obtaining of the noise feature based on the first image feature and the description feature by using the neural network model includes: Acquire target position information, where the information is used to indicate the position of the target image content in the second image; The first image feature and the description feature are processed by a neural network model, and the processing result is adjusted based on the target position information to obtain a noise feature.
3. The method according to claim 2, It is characterized in that The data generation method is performed by a non-trained diffusion model including the neural network model; The processing of the first image feature and the description feature by the neural network model and adjusting the processing result based on the target position information to obtain the noise feature includes: Through the neural network model, attention processing is performed on the first image feature and the description feature based on the attention mechanism, and the attention processing result is adjusted based on the target position information.
4. The method according to claim 3, It is characterized in that The neural network model includes at least one attention module, wherein the attention module includes a cross-attention unit and a self-attention unit connected in series; The performing attention processing on the first image feature and the description feature based on the attention mechanism through the neural network model, and adjusting the attention processing result based on the target position information, includes: Determining, by the cross-attention unit, a first probability value corresponding to an element in a cross-attention feature matrix based on a first query matrix corresponding to the first image feature and a first key matrix corresponding to the descriptive feature; the first probability value indicates a probability of generating the target image content at a position corresponding to the corresponding element; Determining, by the self-attention unit, a second probability value corresponding to an element in a self-attention feature matrix based on a second query matrix corresponding to the first image feature and a second key matrix corresponding to the first image feature, wherein the second probability value indicates a probability of generating the target image content at a position corresponding to the corresponding element; Adjust at least one of the first probability value and the second probability value based on the target position information to obtain an adjusted attention processing result; Noise prediction is performed based on the adjusted attention processing result to obtain noise features.
5. The method according to claim 4, It is characterized in that The adjusting at least one of the first probability value and the second probability value based on the target position information includes at least one of the following: Adjusting a first probability value corresponding to an element in the cross-attention feature matrix based on a position indicated by the target position information; Adjusting a second probability value corresponding to an element in the self-attention feature matrix based on the position indicated by the target position information; The first probability value corresponding to the element in the cross-attention feature matrix is adjusted based on the target position information and a preset loss function.
6. The method according to claim 5, It is characterized in that The step of adjusting the first probability value corresponding to the element in the cross-attention feature matrix based on the position indicated by the target position information, and the step of adjusting the second probability value corresponding to the element in the self-attention feature matrix based on the position indicated by the target position information, include at least one of the following: For an element in the attention feature matrix that is located in the target position information indication area, increasing the probability value of the element; For an element in the attention feature matrix that is outside the area indicated by the target position information, the probability value of the element is reduced.
7. The method according to claim 5, It is characterized in that The adjusting the first probability value corresponding to the element in the cross-attention feature matrix based on the target position information and a preset loss function includes: Based on the target position information and a preset loss function, adjusting the first probability value corresponding to the element in the cross-attention feature matrix by a gradient descent method; Among them, the loss function indicates the ratio of the sum of the first probability values of the elements in the cross-attention feature matrix located in the target position information indication area to the sum of the first probability values of all elements in the cross-attention feature matrix.
8. A method for generating training data, It is characterized in that include: Acquire a target image used to generate a sample image and annotation information corresponding to the target image; By the method of any one of claims 1 to 7, a sample image including a sample object is generated based on the target image and the annotation information; Obtaining a training data pair based on the sample image and the annotation information; The annotation information includes category information and position information for annotating sample objects in the sample image; and the target image content includes the sample objects.
9. The method according to claim 8, It is characterized in that The target image includes a defect-free image, and the sample image is an image with defects generated based on the target image; the category information of the sample object includes information indicating the defect type to which the sample object belongs; The method further comprises: An image processing model for quality inspection is trained using the training data.
10. A data generating device, It is characterized in that include: A feature noise adding module, used for performing noise adding processing on the image features extracted from the first image to obtain the first image features with noise; A text processing module is used to extract features from target description information to obtain description features; The target description information includes first description information for describing the target image content; A network prediction module, configured to obtain a noise feature based on the first image feature and the description feature through a neural network model; A feature denoising module, used for denoising the first image feature according to the noise feature to obtain a denoised second image feature; An image generating module is used to generate a second image including the target image content based on the second image feature.
11. A device for generating training data, It is characterized in that include: An image acquisition module, used to acquire a target image used to generate a sample image and annotation information corresponding to the target image; a sample generation module, configured to generate a sample image including a sample object based on the target image and the annotation information through the apparatus according to claim 10; A training data generation module, used to obtain a training data pair based on the sample image and the annotation information; The annotation information includes category information and position information for annotating sample objects in the sample image; and the target image content includes the sample objects.
12. An electronic device comprising a memory, a processor and a computer program stored in the memory, It is characterized in that The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 9.
13. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.
14. A computer program product comprising a computer program, It is characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.