Image Processing Method, Apparatus, Electronic Device, and Computer-Readable Storage Medium
By combining defect type information in image processing for feature fusion and model training, the problem of unpredictable image completion results and single application scenarios in the prior art is solved, and higher image processing accuracy is achieved.
Patent Information
- Application Number
- CN202210395623.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-14
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-04-14
AI Technical Summary
The existing image processing methods lack control over the diversity domain in the image completion process, resulting in unpredictable results and relatively single application scenarios, which reduces the accuracy of image processing.
By obtaining defect image samples corresponding to the original image samples, defect instances, and defect type information, add a preset mask to generate the target image sample. Then, the defective image sample and the target image sample are featured using a preset image processing model, and converted into defect condition features with defective type information, and feature fusion is performed to generate the target defect image. Based on this image, the image processing model is converged, the trained model is obtained, and the defect instance is completed in the image to be completed.
By controlling the generation of defect types, the domain types of completed defect images are controlled, and example application scenarios for completing different domain contents are added, thereby improving the accuracy of image processing.
Smart Images

Figure CN115115536B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technologies, and particularly to an image processing method, apparatus, electronic device, and computer-readable storage medium. Background Art
[0002] In recent years, with the popularity of neural network technologies in the field of artificial intelligence, the application of neural networks in the field of image processing has also made great progress, especially in the aspect of applying neural networks to image completion. Existing image processing methods often directly map an instance image and an image processed by a mask to the same manifold space for feature fusion, and map the fused features to the original image space, so as to obtain a target instance image after completing the instance.
[0003] In the process of researching and practicing the existing technologies, the inventors of the present invention found that when completing an instance into an image processed by a mask, although diverse results can be generated through the diversity of the instance, there is a lack of control over the diversity domain, making the diverse results unpredictable. In addition, since the domain content of the completed instance is the same as the domain content of the original image, the application scenario of this image completion method is also relatively single. Therefore, the accuracy of image processing is relatively low. Summary of the Invention
[0004] Embodiments of the present invention provide an image processing method, apparatus, electronic device, and computer-readable storage medium, which can improve the accuracy of image processing.
[0005] An image processing method includes:
[0006] Obtaining an original image sample, as well as a defect image sample and defect type information corresponding to at least one defect instance, and adding a preset mask to the original image sample to obtain a target image sample;
[0007] Respectively extracting features from the defect image sample and the target image sample by using a preset image processing model to obtain defect features of the defect image sample and image features of the target image sample;
[0008] Converting the defect type information into defect conditional features, and adding the defect conditional features to the defect features to obtain target defect features;
[0009] Fusing the target defect features and the image features to obtain a target defect image, where the target defect image is an image after completing the defect instance in the mask area of the target image sample;
[0010] Based on the target defect image, defect image samples, target image samples, original image samples, defect features, and image features, converge the preset image processing model to obtain a trained image processing model, and use the trained image processing model to complete the defect instances in the image to be completed.
[0011] Correspondingly, an embodiment of the present invention provides an image processing apparatus, including:
[0012] An acquisition unit, configured to acquire original image samples, defect image samples corresponding to at least one defect instance, and defect type information, and add a preset mask to the original image samples to obtain target image samples;
[0013] An extraction unit, configured to respectively extract features from the defect image samples and target image samples by using a preset image processing model to obtain defect features of the defect image samples and image features of the target image samples;
[0014] An addition unit, configured to convert the defect type information into defect condition features, and add the defect condition features to the defect features to obtain target defect features;
[0015] A fusion unit, configured to fuse the target defect features and image features to obtain a target defect image, where the target defect image is an image after completing the defect instances in the mask area of the target image samples;
[0016] A completion unit, configured to converge the preset image processing model based on the target defect image, defect image samples, target image samples, original image samples, defect features, and image features to obtain a trained image processing model, and use the trained image processing model to complete the defect instances in the image to be completed.
[0017] Optionally, in some embodiments, the addition unit may specifically be configured to determine the splitting number of the defect features according to the dimension information of the defect features; split the defect features into defect sub-features corresponding to the splitting number to obtain a defect sub-feature set; add the defect condition features to each defect sub-feature in the defect sub-feature set to obtain target defect features.
[0018] Optionally, in some embodiments, the fusion unit may specifically be configured to extract position features from the target image samples, add the position features to the image features to obtain target image features; fuse the target image features and target defect features to obtain fused defect features; add the fused defect features to the image features to obtain fused image features, and generate a target defect image based on the fused image features.
[0019] Optionally, in some embodiments, the fusion unit may be specifically configured to convert the target image feature into a key feature, and convert the target defect feature into a query feature and a value feature; fuse the key feature and the query feature, and determine the attention weight of the value feature according to the fused feature; weight the value feature based on the attention weight to obtain a fused defect feature.
[0020] Optionally, in some embodiments, the complementing unit may be specifically configured to determine the image loss information of the preset image processing model according to the target defect image, the defect image sample, the target image sample, the original image sample, and the defect feature; determine the feature distribution loss information of the preset image processing model based on the defect feature and the image feature; fuse the image loss information and the feature distribution loss information, and converge the preset image processing model based on the fused loss information to obtain a trained image processing model.
[0021] Optionally, in some embodiments, the complementing unit may be specifically configured to determine the reconstruction loss information of the preset image processing model based on the target defect image, the target image sample, the defect image sample, and the defect feature; determine the domain loss information of the preset image processing model according to the target defect image, the original image sample, and the defect image sample; fuse the reconstruction loss information and the domain loss information to obtain the image loss information of the preset image processing model.
[0022] Optionally, in some embodiments, the complementing unit may be specifically configured to perform mask processing on the target defect image according to the preset mask to obtain a masked defect image, and compare the masked defect image with the target image sample to obtain image reconstruction loss information; reconstruct the defect image based on the defect feature to obtain a reconstructed defect image, and compare the reconstructed defect image with the defect image sample to obtain defect reconstruction loss information; fuse the image reconstruction loss information and the image reconstruction loss information to obtain the reconstruction loss information of the preset image processing model.
[0023] Optionally, in some embodiments, the complementing unit may be specifically configured to extract image features from the masked defect image to obtain a first sample image feature; extract image features from the target image sample to obtain a second sample image feature; calculate the feature distance between the first sample image feature and the second sample image feature to obtain image reconstruction loss information.
[0024] Optionally, in some embodiments, the completion unit may specifically be configured to identify an image of a masked area in the target defect image to obtain a current defect image; determine image domain loss information of the preset image processing model according to the target defect image and the original image sample; determine defect domain loss information of the preset image processing model based on the current defect image and the defect image sample; and fuse the image domain loss information and the defect domain loss information to obtain domain loss information of the preset image processing model.
[0025] Optionally, in some embodiments, the completion unit may specifically be configured to perform domain feature extraction on the current defect image to obtain a current defect domain feature, and add a defect condition feature to the current defect domain feature to obtain a target defect domain feature; determine a current defect type of the current defect image and a type probability corresponding to the current defect type based on the target defect domain feature; and calculate adversarial loss information between the current defect type and the defect type of the defect image sample based on the type probability to obtain defect domain loss information.
[0026] Optionally, in some embodiments, the completion unit may specifically be configured to perform convolution processing on the defect feature to obtain a processed defect feature; perform convolution processing on the image feature to obtain a processed image feature; sample a target latent feature from a latent feature set of a preset normal distribution; and calculate adversarial loss information between the target latent feature, the processed defect feature, and the processed image feature to obtain feature distribution loss information of the preset image processing model.
[0027] Optionally, in some embodiments, the completion unit may specifically be configured to receive an image processing request, where the image processing request carries an image to be processed and a defect identifier corresponding to the image to be processed; add a preset mask to the image to be processed to obtain an image to be completed, and perform feature extraction on the image to be completed by using the trained image processing model to obtain a current image feature; and complete a target defect instance corresponding to the defect identifier in the image to be completed according to the latent feature set of the preset normal distribution, the current image feature, and the defect condition feature to obtain a target image.
[0028] Optionally, in some embodiments, the completion unit may specifically be configured to sample a current latent feature from the latent feature set of the preset normal distribution, and filter out a target defect condition feature corresponding to the defect identifier from the defect condition features; add the target defect condition feature to the current latent feature to obtain a defect feature to be completed; fuse the defect feature to be completed and the current image feature, and generate the target image based on the completed image feature.
[0029] Optionally, in some embodiments, the image processing device may further include a detection unit. Specifically, the detection unit may be configured to label the defect identifier in the target image to obtain a current image sample; train a preset defect detection model based on the current image sample to obtain a trained defect detection model; and use the trained defect detection model to detect defects in the image to be detected.
[0030] In addition, an embodiment of the present invention further provides an electronic device, including a processor and a memory. The memory stores an application program, and the processor is configured to run the application program in the memory to implement the image processing method provided by the embodiment of the present invention.
[0031] In addition, an embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps in any one of the image processing methods provided by the embodiments of the present invention.
[0032] In the embodiment of the present application, after obtaining the original image sample, the defect image sample corresponding to at least one defect instance, and the defect type information, and adding a preset mask to the original image sample to obtain a target image sample, a preset image processing model is used to extract features from the defect image sample and the target image sample respectively to obtain the defect features of the defect image sample and the image features of the target image sample. Then, the defect type information is converted into defect condition features, and the defect condition features are added to the defect features to obtain target defect features. Then, the target defect features and the image features are fused to obtain a target defect image. Based on the target defect image, the defect image sample, the target image sample, the original image sample, the defect features, and the image features, the preset image processing model is converged to obtain a trained image processing model, and the trained image processing model is used to complete the defect instance in the image to be completed. Since this solution converts the defect type information into defect condition features and controls the generation of defect types by using the defect condition features as conditional bits, the domain type diversity of the completed defect image is ensured to be controllable. In addition, the domain content between the original image sample and the defect instance is different, so the application scenarios of completing instances with different domain contents can be increased. Therefore, the accuracy of image processing can be improved. Description of the Drawings
[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.
[0034] Figure 1It is a schematic diagram of the scenario of the image processing method provided by the embodiment of the present invention;
[0035] Figure 2 It is a schematic flowchart of the image processing method provided by the embodiment of the present invention;
[0036] Figure 3 It is a schematic diagram of the target image sample after masking the part image sample provided by the embodiment of the present invention;
[0037] Figure 4 It is a schematic diagram of the image of the inverted mask provided by the embodiment of the present invention;
[0038] Figure 5 It is a schematic diagram of training the preset image processing model provided by the embodiment of the present invention;
[0039] Figure 6 It is a schematic diagram of comparing the generated defects with the real defects provided by the embodiment of the present invention;
[0040] Figure 7 It is a schematic diagram of comparing the diverse crack defects with the real crack defects provided by the embodiment of the present invention;
[0041] Figure 8 It is another schematic flowchart of the image processing method provided by the embodiment of the present invention;
[0042] Figure 9 It is a schematic diagram of the structure of the image processing device provided by the embodiment of the present invention;
[0043] Figure 10 It is another schematic diagram of the structure of the image processing device provided by the embodiment of the present invention;
[0044] Figure 11 It is a schematic diagram of the structure of the electronic device provided by the embodiment of the present invention. Detailed implementation manners
[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.
[0046] The embodiment of the present invention provides an image processing method, device, electronic device and computer-readable storage medium. Among them, the image processing device can be integrated in the electronic device, and the electronic device can be a server or a terminal device, etc.
[0047] Among them, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, network acceleration services (Content Delivery Network, CDN), and big data and artificial intelligence platforms. The terminal includes, but is not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, etc. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here. The embodiments of the present invention can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, etc.
[0048] For example, referring to Figure 1 , taking the example that the image processing device is integrated in an electronic device, after the electronic device obtains the original image sample, as well as the defect image sample and defect type information corresponding to at least one defect instance, and adds a preset mask to the original image sample to obtain the target image sample, it uses a preset image processing model to extract features from the defect image sample and the target image sample respectively to obtain the defect features of the defect image sample and the image features of the target image sample. Then, it converts the defect type information into defect condition features, adds the defect condition features to the defect condition features to obtain the target defect features, fuses the target defect features and the image features to obtain the target defect image. Then, based on the target defect image, the defect image sample, the target image sample, the original image sample, the defect features and the image features, it converges the preset image processing model to obtain the trained image processing model, and uses the trained image processing model to complete the defect instances in the image to be completed, thereby improving the accuracy of image processing.
[0049] Among them, the image processing method provided by the embodiments of this application relates to the computer vision direction in the field of artificial intelligence. The embodiments of this application can extract features from the target image sample and the defect image sample, etc.
[0050] Among them, Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject that involves a wide range of fields, including both hardware-level and software-level technologies. Among them, artificial intelligence software technology mainly includes directions such as computer vision technology and machine learning / deep learning.
[0051] Among them, Computer Vision (CV) is a science that studies how to enable machines to "see". Further, it refers to machine vision that uses a computer to replace the human eye to identify and measure targets, and further performs image processing to make the image more suitable for human eyes to observe or be transmitted to an instrument for detection after being processed by the computer. As a scientific discipline, computer vision studies related theories and technologies and attempts to establish an artificial intelligence system that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing and image recognition.
[0052] Among them, it can be understood that in the specific implementation manners of this application, relevant data such as the original image samples, defect image samples, and images to be completed of the object are involved. When the following embodiments of this application are applied to specific products or technologies, permission or consent is required, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions.
[0053] The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.
[0054] This embodiment will be described from the perspective of an image processing device, which can be specifically integrated in an electronic device. The electronic device can be a server or a terminal device, etc.; among them, the terminal can include devices such as a tablet computer, a notebook computer, a personal computer (PC), a wearable device, a virtual reality device, or other intelligent devices that can perform image processing.
[0055] An image processing method includes:
[0056] Obtain an original image sample, as well as defect image samples corresponding to at least one defect instance and defect type information, and add a preset mask to the original image sample to obtain a target image sample. Use a preset image processing model to extract features from the defect image samples and the target image sample respectively to obtain defect features of the defect image samples and image features of the target image sample. Convert the defect type information into defect condition features, and add the defect condition features to obtain target defect features. Fuse the target defect features and the image features to obtain a target defect image, which is an image after complementing the defect instance in the masked area of the target image sample. Based on the target defect image, defect image samples, target image sample, original image sample, defect features, and image features, converge the preset image processing model to obtain a trained image processing model, and use the trained image processing model to complement the defect instance in the image to be complemented.
[0057] As Figure 2 shown, the specific process of this image processing method is as follows:
[0058] 101. Obtain an original image sample, as well as defect image samples corresponding to at least one defect instance and defect type information, and add a preset mask to the original image sample to obtain a target image sample.
[0059] Among them, the original image sample can be an image sample without defect instances that has not been masked, and there can be various types of the original image sample. For example, in an industrial application scenario, it can be a normal part or component without defect instances. In other application scenarios, it can be an image sample without defect instances. A defect instance can be understood as an instance corresponding to a defect existing on an object. For example, taking the object as a part, the defect instance can be a defect that may appear on the part, and there can be various types of this defect, such as cracks, pockmarks, protrusions, grooves, or other types of defects, etc. The defect image sample corresponding to the defect instance can be an image sample containing the defect instance, and the defect type information can be information indicating the type of the defect instance, such as a label indicating the defect instance or other types of information.
[0060] Among them, there can be various ways to obtain the original image sample, as well as defect image samples corresponding to at least one defect instance and defect type information. Specifically, it can be as follows:
[0061] For example, it is possible to directly receive the original image sample uploaded by the terminal, as well as the defect image sample and defect type information corresponding to at least one defect instance. Alternatively, it is possible to receive the original image sample uploaded by the terminal and the defect image sample corresponding to at least one defect instance, and identify the defect type information of the defect instance in the defect image sample. Or, it is also possible to randomly select an original image sample from the image sample library, identify the image type in the original image sample, and based on this image type, filter out at least one defect instance corresponding to this image type from the defect instance set, and obtain the defect image sample and defect type information corresponding to this defect instance.
[0062] Among them, it should be noted that there may be an association relationship between the original image sample and the defect instance. For example, taking the original image sample as a part image, the defect instance can be the possible defect of the part in this part image. In addition, there can be multiple defects in one part, and one type of defect can also exist on multiple parts.
[0063] After obtaining the original image sample, as well as the defect image sample and defect type information corresponding to at least one defect instance, a preset mask can be added to the original image sample to perform mask processing. There can be various ways of mask processing. For example, a preset mask can be obtained, added to the original image sample, and the original image sample is masked by this preset mask to obtain the target image sample.
[0064] Among them, the mask can be a picture (usually in binary format) storing information about where to be masked. If it is in binary format, 1 represents the part that needs to be masked, and 0 represents the part that should not be masked. (The meanings of 0 and 1 can also be reversed, depending on the specific usage method). This means that there is a masked area in the target image sample, and this area can be the mask area. Taking the original image sample as a part image sample as an example, the target image sample can be as Figure 3 shown.
[0065] 102. Use a preset image processing model to extract features from the defect image sample and the target image sample respectively, to obtain the defect features corresponding to the defect image sample and the image features of the target image sample.
[0066] Among them, the defect features are used to characterize the feature information of the defect instance in the defect image sample, and the image features are used to characterize the feature information of the target image sample.
[0067] Among them, there can be various ways to use a preset image processing model to extract features from the defect image sample and the target image sample respectively. Specifically, it can be as follows:
[0068] For example, the defect latent encoding network of a preset image processing model can be used to encode a defect image sample to obtain a defect latent encoding, and this defect latent encoding is used as the defect feature of the defect image sample. The image latent encoding network of the preset image processing model is used to encode a target image sample to obtain an image latent encoding, and this image latent encoding is used as the image feature of the target image sample.
[0069] Among them, the dimensions of the hard encodings output by the defect latent encoding network and the image latent encoding network can be the same, both can be tensors of n*m. In addition, the network structures of the defect latent encoding network and the image latent encoding network can be the same. For example, both can be autoencoders, or they can be different. When the defect latent encoding network is an autoencoder, the defect feature can also be reconstructed into a defect image sample through the decoder in the autoencoder.
[0070] 103. Convert the defect type information into a defect conditional feature, and add the defect conditional feature to the defect feature to obtain a target defect feature.
[0071] Among them, the defect conditional feature can be understood as the feature information that guides the diversity of the target instance image for completion. Taking the dimension of the defect feature as n*m as an example, this defect conditional feature can be a vector of m dimensions. This defect conditional feature is understood as a conditional bit, that is, it can be used to control the generation of different defect types by controlling this conditional bit, so as to achieve the generation of diverse real samples. The real sample here can be to add a defect instance to a real image sample to obtain an image sample with a real defect instance. This real sample is used as the training sample of the defect detection model, so as to improve the detection accuracy of the defect detection model and further improve the accuracy of defect detection.
[0072] Among them, there are various ways to convert the defect type information into a defect conditional feature. Specifically, it can be as follows:
[0073] For example, type feature extraction can be performed on the defect type information, and the extracted type feature (embedding) is used as the defect conditional feature. Or, the type information can be extracted from the defect type information and converted into the defect conditional feature. Or, one-hot encoding or other encoding methods can be performed on the defect type information to obtain a defect type encoding, and this defect type encoding is used as the defect conditional feature.
[0074] After converting the defect individual information into defect condition features, the defect condition features can be added to the defect features in various ways. For example, according to the dimensionality information of the defect features, the number of splits of the defect features can be determined, and the defect features can be split into pairs of defect sub-features corresponding to the number of splits to obtain a set of defect sub-features. Then, the defect condition features are added to each defect sub-feature in the set of defect sub-features to obtain the target defect features.
[0075] Among them, the number of splits can be the number required for splitting the defect features. There are various ways to determine the number of splits of the defect features according to the dimensionality information of the defect features. For example, the dimensionality information of the defect features is obtained, the target dimension is selected from this dimensionality information, and this target dimension is used as the number of splits. For instance, taking the dimension of the defect features as n*m as an example, n can be used as the number of splits.
[0076] After determining the number of splits of the defect features, the defect features can be split into defect sub-features corresponding to the number of splits in various ways. For example, taking the number of splits as n as an example, the defect features can be split into n defect tokens of m dimensions, and the defect tokens of m dimensions are used as defect sub-features, thereby obtaining a set of defect sub-features.
[0077] After splitting out the set of defect sub-features, the defect condition features can be added to each defect sub-feature in the set of defect sub-features in various ways. For example, the m-dimensional defect condition features can be added to each defect token, thereby obtaining new n defect tokens of m bits, and these are used as the target defect features (defect latent tokens), which can be represented by Dtokens(n*m).
[0078] 104. Fuse the target defect features and the image features to obtain the target defect image.
[0079] Among them, the target defect image can be the image after adding defect instances to the masked area of the target image sample. Taking the part image after masking the target image sample as an example, the target defect image can be the defective part image, which is used to represent the image of the part with defects.
[0080] Among them, there are various ways to fuse the target defect features and the image features. Specifically, it can be as follows:
[0081] For example, position features are extracted from the target image sample, and the position features are added to the image features to obtain the target image features. The target image features and the target defect features are fused to obtain the fused defect features, the fused defect features are added to the image features to obtain the fused image features, and based on the fused image features, the target defect image is generated.
[0082] Among them, there are various ways to add the position features to the image features. For example, according to the dimensionality information of the image features, the number of splits of the image features can be determined, the image features are split into image sub-features corresponding to the number of splits to obtain an image sub-feature set, and the position features are added to each image sub-feature in the image sub-feature set to obtain the target image features. The method of splitting and adding the position features can refer to the method of generating the target defect features. For example, the image features (masked image latent code) can be regarded as n m-dimensional image tokens, and each image token is added with its position feature (embeding) to obtain new n m-dimensional image tokens, and these new n m-dimensional image tokens are used as the target image features.
[0083] After obtaining the target image features, the target image features and the target defect features can be fused. There are various fusion methods. For example, the target image features can be converted into key features, and the target defect features are converted into query features and value features. The key features and the query features are fused, and according to the fused features, the attention weights of the value features are determined, and based on the attention weights, the value features are weighted to obtain the fused defect features.
[0084] Among them, the key features, query features, and value features can be abstract features of the embedding vectors in different subspaces. There are various ways to convert the target image features into key features and convert the target defect features into query features and value features. For example, attention parameters can be obtained. The attention parameters include key parameters, query parameters, and value parameters. The target image features and the key parameters are fused to obtain the key features, the target defect features and the query parameters are fused to obtain the query features, and the target defect features and the value parameters are fused to obtain the value features.
[0085] Among them, there are various ways to fuse the target image features and the key parameters. For example, taking the target image features as M tokens(n*m) , and the key parameters as W K (m * m) as an example, directly multiplying the target image features and the key parameters can obtain the key features, which can be specifically shown in the following formula (1):
[0086] K n*m = M tokens * WK (1)
[0087] Among them, K n*m is the key feature, M tokens is the target image feature, and W K is the key parameter.
[0088] Among them, the method of fusing the target defect feature with the query parameter and the value feature respectively can refer to the fusion process of the target image feature and the key parameter above. Taking the target defect feature as D tokens(n*m) , query parameter W K (m*m) and value parameter W V (m*m) as an example, multiplying the target defect feature directly by the query parameter can obtain the query feature, which can be specifically shown in formula (2):
[0089] Q n*m = D tokens * W Q (2)
[0090] Among them, Q n*m is the query feature, D tokens is the target defect feature, and W Q is the parameter parameter.
[0091] Multiplying the target defect feature directly by the value parameter can obtain the value feature, which can be specifically shown in formula (3):
[0092] V n*m = D tokens * W V (3)
[0093] Among them, V n*m is the value feature, D tokens is the target defect feature, and W V is the value parameter.
[0094] After converting out the key feature, query feature and value feature, the key feature and the query feature can be fused. There are various fusion methods. For example, the current dimension information of the key feature can be obtained, and the target dimension value can be extracted from the current dimension information. The key feature is converted in format to obtain the key feature in the target format. The key feature in the target format is fused with the query feature to obtain the initially fused feature. Then, the ratio between the initially fused feature and the target dimension value is calculated to obtain the fused feature. For example, taking the key feature as K n*m , and the query feature as Q n*m as an example, the target dimension value d k = m is extracted from the key feature, and the key feature K is transposed to obtain the key feature K in the target format T , the key feature K in the target formatT Multiply with the query feature Q to obtain the initially fused feature Q*K T Calculate the initially fused feature Q*K T and the target dimension value d k to obtain the fused feature
[0095] After fusing the key feature and the query feature, the attention weight of the value feature can be determined according to the fused feature, and the attention weight is used to indicate the importance degree of the value feature. There are multiple ways to determine the attention weight. For example, the fused feature can be normalized to obtain the attention weight of the value feature. There are multiple ways of normalization, and the softmax function or other normalization algorithms can be used for processing.
[0096] After determining the attention weight, the value feature can be weighted based on the attention weight. There are multiple ways of weighting. For example, the attention weight can be weighted with n m-dimensional features in the value feature respectively, and the n weighted m-dimensional features are fused to obtain the fused defect feature, which can be specifically shown in formula (4):
[0097]
[0098] where Attention(Q, K, V) is the fused defect feature, Q is the query feature, K T is the key feature in the target format, d k is the target dimension value in the key feature, and V is the value feature.
[0099] After obtaining the fused defect feature, the fused defect feature can be added to the image feature to obtain the fused image feature. There are multiple ways to add the fused defect feature to the image feature. For example, the fused defect feature can be directly added to the image feature to obtain the fused image feature, or the fused defect feature and the image feature can be concatenated to obtain the fused image feature, or the weighting coefficients of the fused defect feature and the image feature can be obtained respectively, and the fused defect feature and the image feature are weighted based on the weighting coefficients, and the weighted fused defect feature and the weighted image feature are added to obtain the fused image feature.
[0100] After obtaining the fused image features, a target defect image can be generated based on the fused image features. There can be multiple ways to generate the target defect image. For example, the image mapping network (Generator) of a preset image processing model can be used to map the fused image features back to the image space of the original image sample (h*w*c, where h is the height of the original image sample, w is the width of the original image sample, and c is the number of channels of the original image sample), thereby obtaining the target defect image.
[0101] 105. Based on the target defect image, defect image sample, target image sample, original image sample, defect features, and image features, converge the preset image processing model to obtain a trained image processing model, and use the trained image processing model to complete defect instances in the image to be completed.
[0102] Among them, the image to be completed is an image that needs to complete defect instances in the masked area.
[0103] Among them, there can be multiple ways to converge the preset image processing model based on the target defect image, defect image sample, target image sample, original image sample, defect features, and image features. Specifically, it can be as follows:
[0104] For example, based on the target defect image, defect image sample, target image sample, original image sample, and defect features, determine the image loss information of the preset image processing model. Based on the defect features and image features, determine the feature distribution loss information of the preset image processing model. Fuse the image loss information and the feature distribution loss information, and converge the preset image processing model based on the fused loss information to obtain a trained image processing model.
[0105] Among them, the image loss information can be used to indicate the loss information of the image reconstruction loss and the image domain type loss of the generated target defect image. There can be multiple ways to determine the image loss information. For example, based on the target defect image, target image sample, defect image sample, and defect features, determine the reconstruction loss information of the preset image processing model. According to the target defect image, original image sample, and defect image sample, determine the domain loss information of the preset image processing model. Fuse the reconstruction loss information and the domain loss information to obtain the image loss information of the preset image processing model.
[0106] Among them, the reconstruction loss information is used to indicate the loss information of the reconstruction loss of the target image sample and the reconstruction loss of the defective image sample. There are various ways to determine the reconstruction loss information. For example, a preset mask is added to the target defective image to obtain the defective image after masking, and the defective image after masking is compared with the defective image sample to obtain the image reconstruction loss information. The defective image is reconstructed based on the defect features to obtain the reconstructed defective image, and the reconstructed defective image and the defective image sample are compared to obtain the defect reconstruction loss information. The image reconstruction loss information and the defect reconstruction loss information are fused to obtain the reconstruction loss information of the preset image processing model.
[0107] Among them, the image reconstruction loss information is used to indicate the reconstruction error between the generated target defective image and the target image sample. Mask processing can be understood as adding a mask to the target defective image for masking. There are various ways to perform mask processing on the target defective image. For example, the preset mask is inverted to obtain the inverted mask, and the image of the inverted mask can be as Figure 4 shown. The inverted mask is added to the target defective image to obtain the defective image after masking.
[0108] Among them, there are various ways to perform the inversion process. For example, taking the mask in the binary format as an example, 1 in the mask represents the part to be masked, and 0 represents the part that should not be masked. The inversion process is to invert 0 and 1 in the preset mask. Taking the target mask M after processing the preset mask as an example, the preset mask mask = 1 - M.
[0109] After obtaining the defective image after masking, there are various ways to compare the defective image after masking with the target image sample. For example, feature extraction can be performed on the defective image after masking to obtain the first sample image feature, feature extraction can be performed on the target image sample to obtain the second sample image feature, and the feature distance between the first sample image feature and the second sample image feature is calculated to obtain the image reconstruction loss information, which can be specifically shown in formula (5):
[0110]
[0111] Among them, Loss1 is the image reconstruction loss information, M is the target mask after the inversion process of the preset mask, U(I i , m ) is the target defective image generated by the preset image processing network, and I m is the target image sample.
[0112] Among them, the defect reconstruction loss information is used to indicate the error between the defect instance reconstructed based on the defect features and the defect instance in the defect image sample. There are various ways to reconstruct the defect image based on the defect features. For example, the decoder in the autoencoder network can be used to decode the defect features to obtain the reconstructed defect image. Or, the defect instance can also be reconstructed based on the defect features to obtain the reconstructed defect image.
[0113] After obtaining the reconstructed defect image, the method of comparing the reconstructed defect image with the defect image sample is the same as the process of comparing the masked defect image with the target image sample, and the defect reconstruction loss information is also obtained through the L1 loss, which will not be elaborated here one by one. This step mainly ensures that the encoding of the defect basically retains all the information of the defect image sample.
[0114] Among them, the domain loss information is used to indicate the loss information of the authenticity of the generated target defect image and the error between the defect and the. There are various ways to determine the domain loss information. For example, the image in the masked area is identified in the target defect image to obtain the current defect image. According to the target defect image and the original image sample, the image domain loss information of the preset image processing model is determined. Based on the current defect image and the defect image sample, the defect domain loss information of the preset image processing model is determined. The image domain loss information and the defect domain loss information are fused to obtain the domain loss information of the preset image processing model.
[0115] Among them, the defect domain loss information is used to indicate the error between the domain type of the defect area in the target defect image and the domain type of the defect image sample. There are various ways to determine the defect domain loss information. For example, feature extraction is performed on the current defect image to obtain the current defect domain feature, and the defect condition feature is added to the current defect domain feature to obtain the target defect domain feature. Based on the target defect domain feature, the current defect type of the current defect image and the type probability corresponding to the current defect type are determined. Based on the type probability, the adversarial loss information between the current defect type and the defect type of the defect image sample is calculated to obtain the defect domain loss information, which can be specifically shown in formula (6):
[0116]
[0117] Among them, L adv1 is the defect domain loss information, D(i) is the type probability, Pdata is the real domain (real defect type), Pgenerate is the generated domain (current defect type), and the adversarial loss information is determined by the discriminator (D).
[0118] Among them, the image domain loss information is used to indicate the error between the image features of the generated target defect image and the image type of the original image sample. The method for determining the image domain loss information can refer to the calculation process of the defect domain loss information above. However, it should be noted that when determining the defect domain loss information, the defect condition feature needs to be added to the current defect domain feature, while this is not required for the image domain loss information. In addition, different discriminators can be used to determine these two adversarial loss information.
[0119] Among them, the feature distribution loss information is used to indicate the distribution error between the defect feature and the image feature, which are two latent encodings. There are various ways to determine the feature distribution loss information. For example, the defect feature is convolved to obtain the processed defect feature, the image feature is convolved to obtain the processed image feature, a target latent feature is sampled from the set of latent features with a preset normal distribution, and the adversarial loss information among the target latent feature, the processed defect feature, and the processed image feature is calculated to obtain the feature distribution loss information of the preset image processing model.
[0120] Among them, there are various ways to convolve the defect feature. For example, the defect feature can be convolved through a 1*1 convolution to transform from an n*m vector to an n*1 vector, and this n*1 vector is used as the processed defect feature. The way to convolve the image feature is the same as that of the defect feature, and will not be elaborated here one by one.
[0121] Among them, the set of latent features with a preset normal distribution can be a set of latent space encodings randomly generated by the program, and the distribution of the latent features in this set conforms to the normal distribution. After sampling the target latent feature from the set of latent features with a preset normal distribution, the adversarial loss information among the target latent feature, the processed defect feature, and the processed image feature can be calculated. There are various ways to calculate the adversarial loss information. For example, the discriminator (D) can be used to calculate the adversarial loss information between the target latent feature and the processed defect feature to obtain the defect feature distribution loss information, and the discriminator (D) can be used to calculate the adversarial loss information between the target latent feature and the processed image feature to obtain the image feature distribution loss information. The defect feature distribution loss information and the image feature distribution loss information are fused to obtain the feature distribution loss information.
[0122] Among them, for the feature distribution loss information, the distribution distance function (such as the KL divergence or other distribution distance functions, etc.) can also be used to calculate the feature distribution loss information among the target latent feature, the processed defect feature, and the processed image feature.
[0123] After determining the image loss information and the feature distribution loss information, the image loss information and the feature distribution loss information can be fused. There are various fusion methods. For example, since the image loss information includes the reconstruction loss information and the domain loss information, the weighted coefficients of the reconstruction loss information, the domain loss information, and the feature distribution loss information can be obtained. Based on the weighted coefficients, the reconstruction loss information, the domain loss information, and the feature distribution loss information are weighted respectively, and the weighted reconstruction loss information, the weighted domain loss information, and the weighted feature distribution loss information are fused to obtain the fused loss information, which can be specifically shown in formula (7):
[0124] Loss=λ rec *Loss rec +λ dis *Loss dis +λ adv *Loss adv (7)
[0125] Among them, Loss is the fused loss information, λ rec is the weighted coefficient of the reconstruction loss information, Loss rec is the reconstruction loss information, Loss dis is the feature distribution loss information, λ dis is the weighted coefficient of the feature distribution loss information, Loss adv is the domain loss information, λ adv is the weighted coefficient of the domain loss information, λ rec 、λ dis 、λ adv are all hyperparameters.
[0126] After determining the fused loss information, the preset image processing model can be converged based on the fused loss information. There are various convergence methods. For example, based on the fused loss information, the network parameters of the preset image processing model can be updated based on the gradient descent algorithm to obtain the trained image processing model. Or, other network parameter update algorithms can also be used to update the network parameters of the preset image processing model based on the fused loss information to obtain the trained image processing model.
[0127] Among them, taking the original image sample as the part image sample as an example, the process of training the preset image processing model can be as Figure 5 shown. Use the defect hidden coding network E1 of the preset image processing model to perform hidden coding on the defect image sample to obtain defect features, and convert the defect type information of the defect instance into defect conditional features. Add a preset mask to the part image sample to obtain the target part image sample (I m) Use the image latent encoding network E2 of the preset image processing model to perform latent encoding on the target part image sample to obtain part features. Input the defect condition features, defect features, and part features into the attention network of the preset image processing model to obtain the fused image features. Then, map the fused image features back to the original image space through the generator (G) in the preset image processing model, thereby obtaining the defective part image (I g ). When training the preset image processing model, the supervision signal part can be divided into three parts, specifically as follows:
[0128] (1) Reconstruction error:
[0129] The reconstruction error can include two parts. One is the L1 loss between the defective part image (I g ) * mask and the target part image sample (I m ). The other is the reconstruction error of the defect latent encoding network E1 (encode) and the defect latent decoding network D1 (decode), which is obtained from the L1 loss between the reconstructed image decoded and the defective image input into the encode. This step mainly ensures that the encoding of the defect basically retains all the information of the original part image. The sum of the above two losses is denoted as Loss rec .
[0130] (2) Distribution error:
[0131] After the two latent space encodings (defect features and image features) pass through a 1*1 convolution, from an n*m to an n*1 vector, and then are sent together with samples sampled from the normal distribution into D3 (discriminator 3) to obtain the adversarial loss, which is denoted as Loss dis .
[0132] (3) The authenticity of the generated defective part image and the error in the defect domain
[0133] Finally, the generated defective part image should be consistent with the domain of the original part image, and the domain of the defective area of the defective part image should be the same as the domain of the defective image sample. They are respectively judged by different discriminators (the entire defective part image is judged by D1, and the defective area is judged by D2). The difference is that when judging whether they are in the same domain for the defective area, a conditional vector (embedding of the conditional bit) needs to be added to ensure that the current defective instance in the defective area of the generated defective part image matches the defective instance c. This loss is denoted as Loss adv .
[0134] Fuse the loss information corresponding to these three errors, and converge the preset image processing model based on the fused loss information, so as to obtain the trained image processing model.
[0135] After converging the preset image processing model, the defective instances of the image to be completed can be complemented based on the trained image processing. There are various ways to complete the image to be completed. For example, an image processing request can be received. The image processing request carries the image to be processed and the defect identifier corresponding to the image to be processed. Add a preset mask to the image to be processed to obtain the image to be completed, and use the trained image processing model to extract the features of the image to be completed to obtain the current image features. According to the set of latent features of the preset normal distribution, the current image features, and the defect condition features, complete the target defect instance corresponding to the defect identifier in the image to be completed to obtain the target image.
[0136] Among them, there are various ways to complete the target defect instance corresponding to the defect identifier in the image to be completed according to the set of latent features of the preset normal distribution, the current image features, and the defect condition features. For example, sample the current latent feature from the set of latent features of the preset normal distribution, and screen out the target defect condition feature corresponding to the defect identifier from the defect condition features. Add the target defect condition feature to the current latent feature to obtain the defect feature to be completed, fuse the defect feature to be completed and the current image features, and generate the target image based on the completed image features.
[0137] Among them, it can be found that after training the preset image processing model, in the inference stage (model application stage), sampling can be directly performed in the set of latent features of the preset normal distribution, and the target defect condition feature (conditional code) corresponding to the defect to be generated can be selected, and cooperate with the image to be completed (I m ) Pass through the latent encoding network E2 to perform latent encoding on the image to be completed to obtain the current image features. Send the current image features, the target defect condition features, and the sampled current latent features into the attention module of the trained image processing model to obtain the fused representation result, and finally the generator (G) of the trained image processing model is used to generate the final target image (the completed defective image).
[0138] Among them, taking the image to be processed as a part image as an example, after masking the part image, the target part image is obtained, and the defect instance corresponding to the defect identifier is added to the target part image, so that a diverse defective part image can be obtained. The generated defective part image can contain different types of realistic defects. The comparison between the generated defects and the real defects can be as Figure 6As shown, taking a defect as a crack for example, defect part images containing diverse cracks can be generated. The comparison between the generated diverse crack defects and real crack defects can be as Figure 7 shown.
[0139] After completing the target defect instance corresponding to the defect identifier for the image to be completed, the obtained target image can also be used as a training sample for a preset defect detection model, thereby solving the problem that the detection index of the defect detection model cannot be achieved and the robustness is not strong due to insufficient real defect data, and further improving the accuracy of defect detection. Therefore, there can be various ways to detect the image to be detected. For example, defect identifiers can be marked in the target image to obtain the current image sample, and the preset defect detection model can be trained based on the current image sample to obtain the trained defect detection model, and the trained defect detection model can be used to detect the defects in the image to be detected.
[0140] Among them, there can be various ways to train the preset defect model based on the current image sample. For example, the preset defect detection model can be used to extract features from the current image sample to obtain the current defect features corresponding to the current image sample, and based on the current defect features, the defect type of the current image sample can be predicted to obtain the predicted defect type. Based on the predicted defect type and the marked defect identifier, the preset defect detection model can be converged to obtain the trained defect detection model. Or, the preset defect detection model can also be used to identify the current defect region in the current image sample, and feature extraction can be performed on the defect image corresponding to the current defect region to obtain the current defect features, and based on the current defect features, the defect type of the current image sample can be predicted to obtain the predicted defect type. Based on the predicted defect type and the marked defect identifier, the preset defect detection model can be converged to obtain the trained defect detection model.
[0141] After training the preset defect detection model, there can be various ways to detect the defects in the image to be detected based on the trained defect detection model. For example, at least one image to be detected can be obtained, and the trained defect detection model can be used to extract features from the image to be detected to obtain the defect features corresponding to the image to be detected, and based on the defect features, the defect detection result of the image to be detected can be determined. Or, at least one image to be detected can be obtained, and the trained defect detection model can be used to identify the candidate defect region in the image to be detected, and feature extraction can be performed on the image corresponding to the candidate defect region to obtain the defect features corresponding to the image to be detected, and based on the defect features, the defect detection result of the image to be detected can be determined.
[0142] Among them, this solution can be applied to the scenario of modern industrial automation vision quality inspection. In the modern industrial automation vision quality inspection solution, a large number of defective pictures need to be collected to train the model so that the model can have good generalization ability, so as to meet the requirements of the part overkill rate and part omission rate of the production line. However, in reality, due to reasons such as not yet in mass production, material management, and the long-tail distribution of defect categories, not all defects that are emphasized on the production line can provide stable and sufficient training samples. Therefore, how to obtain more and more real defective pictures has become a key link in improving the algorithm index. This solution is a technology that can cover part of the area on the normal part drawing with a mask and finally generate a generated picture with the required defect at the masked place, thereby solving the problem of data volume shortage in the actual optimization algorithm.
[0143] Among them, the numerator of the part overkill rate is the number of qualified (ok) parts misjudged as unqualified (ng) parts by the algorithm model, and the denominator is the total number of parts detected or the total number of qualified parts detected. The numerator of the part omission rate is the number of unqualified (ng) parts misjudged as qualified (ok) parts by the algorithm model. The denominator is the total number of parts detected or the total number of all unqualified parts detected.
[0144] As can be seen from the above, the embodiment of this application obtains the original image sample, as well as the defect image sample and defect type information corresponding to at least one defect instance, adds a preset mask to the original image sample to obtain a target image sample, and then uses a preset image processing model to extract features from the defect image sample and the target image sample respectively to obtain the defect features of the defect image sample and the image features of the target image sample. Then, the defect type information is converted into defect condition features, and the defect condition features are added to the defect features to obtain target defect features. Then, the target defect features and the image features are fused to obtain a target defect image. Based on the target defect image, the defect image sample, the target image sample, the original image sample, the defect features and the image features, the preset image processing model is converged to obtain a trained image processing model, and the trained image processing model is used to complete the defect instance in the image to be completed; since this solution converts the defect type information into defect condition features and controls the generation of defect types by using the defect condition features as conditional bits, it ensures that the domain type diversity of the completed defect images is controllable. In addition, the domain content between the original image sample and the defect instance is not the same, so the application scenarios of completing instances with different domain contents can be increased. Therefore, the accuracy of image processing can be improved.
[0145] According to the method described in the above embodiments, the following will give further detailed examples.
[0146] In this embodiment, it will be described by taking the image processing device specifically integrated in an electronic device, the electronic device being a server, the original image sample being a part image sample, and the target defect image being a defective part image as an example.
[0147] As Figure 8 shown, an image processing method has the following specific process:
[0148] 201. The server obtains a part image sample, as well as a defect image sample and defect type information corresponding to at least one defect instance.
[0149] For example, the server can directly receive the part image sample, as well as the defect image sample and defect type information corresponding to at least one defect instance uploaded by the terminal. Or, it can receive the part image sample uploaded by the terminal and the defect image sample corresponding to at least one defect instance, and identify the defect type information of the defect instance in the defect image sample. Or, it can also randomly select a part image sample from the image sample library, identify the image type in the part image sample, and based on this image type, screen out at least one defect instance corresponding to this image type from the defect instance set, and obtain the defect image sample and defect type information corresponding to this defect instance.
[0150] 202. The server adds a preset mask to the part image sample to obtain a target part image sample.
[0151] For example, the server can obtain the preset mask, add the preset mask to the part image sample, and perform mask covering processing on the part image sample through the preset mask, so as to obtain the target image sample.
[0152] 203. The server uses a preset image processing model to extract features from the defect image sample and the target part image sample respectively, to obtain the defect features corresponding to the defect image sample and the part features of the target part image sample.
[0153] For example, the server can use the defect hidden coding network of the preset image processing model to encode the defect image sample to obtain a defect hidden coding (n*m), and use this defect hidden coding as the defect features of the defect image sample. Use the image hidden coding network of the preset image processing model to encode the target part image sample to obtain an image hidden coding (n*m), and use this image hidden coding as the part features of the target part image sample.
[0154] 204. The server converts the defect type information into defect conditional features.
[0155] For example, the server can perform type feature extraction on the defect type information, and use the extracted type features (embedding) as defect condition features. Or, the server can extract type information from the defect type information and convert this type information into defect condition features. Or, the server can also perform one-hot encoding or other encoding methods on the defect type information to obtain a defect type encoding, and use this defect type encoding as a defect condition feature.
[0156] 205. The server adds the defect condition features to the defect features to obtain target defect features.
[0157] For example, the server obtains the dimension information of the defect features, filters out the target dimension n from this dimension information, and uses the target dimension n as the splitting number. The server splits the defect features into n defect tokens of m dimensions, and uses the m-dimensional defect tokens as defect sub-features, thereby obtaining a set of defect sub-features. Add the m-dimensional defect condition features to each defect token, thereby obtaining new n m-bit defect tokens, and use these as target defect features (defect latent tokens), which can be represented by Dtokens(n*m).
[0158] 206. The server fuses the target defect features and the part features to obtain a defective part image.
[0159] For example, the server extracts position features from the target part image samples, treats the part features (masked image latent code) as n image tokens of m dimensions, and adds their position features (embeding) to each image token to obtain new n image tokens of m dimensions. Use these new n m-dimensional image tokens as target part features. Obtain attention parameters, which include key parameters, query parameters, and value parameters. Directly multiply the target part features by the key parameters to obtain key features, which can be specifically shown in formula (1) below. Directly multiply the target defect features by the query parameters to obtain query features, which can be specifically shown in formula (2) below. Directly multiply the target defect features by the value parameters to obtain value features, which can be specifically shown in formula (3) below.
[0160] The server extracts the target dimension value d k = m from the key features, transposes the key features K to obtain key features K in the target format T , and multiplies the key features K in the target format T by the query features Q to obtain the initially fused feature Q*K T . Calculate the initially fused feature Q*K T and the target dimension value d kThe ratio between them is obtained to get the fused feature.
[0161] The server uses the softmax function or other normalization algorithms to normalize the fused feature, so as to obtain the attention weights of the value features. The attention weights are weighted with the n m-dimensional features in the value features respectively, and the weighted n m-dimensional features are fused to obtain the fused defect feature, which can be specifically shown in formula (4).
[0162] The server can directly add the fused defect feature and the part feature to obtain the fused part feature. Or, the fused defect feature and the part feature can be concatenated to obtain the fused part feature. Or, the weighted coefficients of the fused defect feature and the part feature can be obtained respectively, and the fused defect feature and the part feature are weighted based on the weighted coefficients, and the weighted fused defect feature and the weighted part feature are added to obtain the fused part feature.
[0163] The server uses the image mapping network (Generator) of the preset image processing model to map the fused part feature back to the image space of the part image sample (h*w*c, where h is the height of the part image sample, w is the width of the part image sample, and c is the number of channels of the part image sample), so as to obtain the defective part image.
[0164] 207. Based on the defective part image, the defective image sample, the target part image sample, the part image sample, the defect feature and the part feature, the preset image processing model is converged to obtain the trained image processing model.
[0165] For example, the server performs an inversion process on the preset mask to obtain the inverted mask, adds the inverted mask to the defective part image to obtain the masked defective image, extracts features from the masked defective image to obtain the first sample part feature, extracts features from the target part image sample to obtain the second sample part feature, and calculates the feature distance between the first sample part feature and the second sample part feature to obtain the part reconstruction loss information, which can be specifically shown in formula (5). The server uses the decoder in the autoencoder network to decode the defect feature to obtain the reconstructed defective image, or, the defective instance can also be reconstructed based on the defect feature to obtain the reconstructed defective image. The reconstructed defective image is compared with the defective image sample to obtain the defect reconstruction loss information. The part reconstruction loss information and the defect reconstruction loss information are fused to obtain the reconstruction loss information of the preset image processing model.
[0166] The server identifies the image of the masked area in the defective part image to obtain the current defective image, extracts features from the current defective image to obtain the current defective domain features, adds defective condition features to the current defective domain features to obtain the target defective domain features, and based on the target defective domain features, determines the current defective type of the current defective image and the type probability corresponding to the current defective type. Based on the type probability, calculates the adversarial loss information between the current defective type and the defective type of the defective image sample to obtain the defective domain loss information, which can be specifically shown in formula (6). According to the defective part image and the part image sample, determines the part domain loss information of the preset image processing model. Fuses the part domain loss information and the defective domain loss information to obtain the domain loss information of the preset image processing model. Fuses the reconstruction loss information and the domain loss information to obtain the image loss information of the preset image processing model.
[0167] The server can perform 1*1 convolution processing on the defective features to convert from an n*m vector to an n*1 vector, and use this n*1 vector as the processed defective features. The way of performing convolution processing on the part features is the same as that of the defective features. Sample the target latent features from the set of latent features in the preset normal distribution, and use the discriminator (D) to calculate the adversarial loss information between the target latent features and the processed defective features to obtain the defective feature distribution loss information. Use the discriminator (D) to calculate the adversarial loss information between the target latent features and the processed part features to obtain the part feature distribution loss information. Fuse the defective feature distribution loss information and the part feature distribution loss information to obtain the feature distribution loss information.
[0168] The server obtains the weighting coefficients of the reconstruction loss information, the domain loss information, and the feature distribution loss information, and based on the weighting coefficients, weights the reconstruction loss information, the domain loss information, and the feature distribution loss information respectively, and fuses the weighted reconstruction loss information, the weighted domain loss information, and the weighted feature distribution loss information to obtain the fused loss information, which can be specifically shown in formula (7).
[0169] The server updates the network parameters of the preset image processing model based on the fused loss information using the gradient descent algorithm to obtain the trained image processing model. Or, other network parameter update algorithms can also be used to update the network parameters of the preset image processing model based on the fused loss information to obtain the trained image processing model.
[0170] 208. The server uses the trained image processing model to complete the defective instance in the image to be completed.
[0171] For example, the server can receive an image processing request that carries the part image to be processed and the defect identifier corresponding to the part image to be processed, add a preset mask to the part image to be processed to obtain the part image to be completed, sample from the set of hidden features of a preset normal distribution, and select the target defect conditional feature (conditional code) corresponding to the defect to be generated. Combine it with the part image to be completed (I m ) Pass through the hidden encoding network E2 to perform hidden encoding on the part image to be completed, obtaining the current part feature. Send the current part feature, the target defect conditional feature, and the sampled current hidden feature into the attention module of the trained image processing model to obtain the fused representation result, and finally use the generator (G) of the trained image processing model to generate the final target part image (the defect part image after completion).
[0172] Optionally, after completing the target defect instance corresponding to the defect identifier for the part image to be completed, the server can also mark the defect identifier in the target part image to obtain the current part image sample, use a preset defect detection model to extract features from the current part image sample to obtain the current defect feature corresponding to the current part image sample, and based on the current defect feature, predict the defect type of the current part image sample to obtain the predicted defect type. Based on the predicted defect type and the marked defect identifier, converge the preset defect detection model to obtain the trained defect detection model. Or, it can also use the preset defect detection model to identify the current defect area in the current part image sample, extract features from the defect image corresponding to the current defect area to obtain the current defect feature, and based on the current defect feature, predict the defect type of the current part image sample to obtain the predicted defect type. Based on the predicted defect type and the marked defect identifier, converge the preset defect detection model to obtain the trained defect detection model.
[0173] The server can obtain at least one part image to be detected, use the trained defect detection model to extract features from the part image to be detected to obtain the defect feature corresponding to the part image to be detected, and based on the defect feature, determine the defect detection result of the part image to be detected. Or, it can also obtain at least one part image to be detected, use the trained defect detection model to identify the candidate defect area in the part image to be detected, extract features from the image corresponding to the candidate defect area to obtain the defect feature corresponding to the part image to be detected, and based on the defect feature, determine the defect detection result of the part image to be detected.
[0174] As can be seen from the above, after the server in this embodiment obtains the part image sample, the defect image sample corresponding to at least one defect instance, and the defect type information, and adds a preset mask to the part image sample to obtain the target part image sample, the preset image processing model is used to extract features from the defect image sample and the target part image sample respectively, so as to obtain the defect features of the defect image sample and the part features of the target part image sample. Then, the defect type information is converted into defect condition features, and the defect condition features are added to the defect features to obtain the target defect features. Then, the target defect features and the part features are fused to obtain the defective part image. Based on the defective part image, the defect image sample, the target part image sample, the part image sample, the defect features, and the part features, the preset image processing model is converged to obtain the trained image processing model, and the trained image processing model is used to complete the defect instance in the part image to be completed. Since this solution converts the defect type information into defect condition features and controls the generation of the defect type by using the defect condition features as condition bits, the domain type diversity of the completed defect image is ensured to be controllable. In addition, the domain content between the part image sample and the defect instance is different, so the application scenarios of completing instances with different domain contents can be increased. Therefore, the accuracy of image processing can be improved.
[0175] To better implement the above method, an embodiment of the present invention further provides an image processing device. The image processing device can be integrated in an electronic device, such as a server or a terminal. The terminal can include a tablet computer, a notebook computer, and / or a personal computer, etc.
[0176] For example, as Figure 9 shown, the image processing device can include an acquisition unit 301, an extraction unit 302, an addition unit 303, a fusion unit 304, and a completion unit 305, as follows:
[0177] (1) Acquisition unit 301;
[0178] The acquisition unit 301 is configured to acquire an original image sample, a defect image sample corresponding to at least one defect instance, and defect type information, and add a preset mask to the original image sample to obtain a target image sample.
[0179] For example, the obtaining unit 301 can be specifically configured to receive the original image sample uploaded by the terminal, as well as the defect image sample and defect type information corresponding to at least one defect instance. Alternatively, it can receive the original image sample uploaded by the terminal and the defect image sample corresponding to at least one defect instance, and identify the defect type information of the defect instance in the defect image sample. Or, it can also randomly select an original image sample from the image sample library, identify the image type in the original image sample, and based on this image type, filter out at least one defect instance corresponding to this image type from the defect instance set, and obtain the defect image sample and defect type information corresponding to the defect instance. Obtain a preset mask, add the preset mask to the original image sample, and perform mask covering processing on the original image sample through the preset mask to obtain a target image sample.
[0180] (2) Extraction unit 302;
[0181] The extraction unit 302 is configured to respectively extract features from the defect image sample and the target image sample by using a preset image processing model, so as to obtain the defect features of the defect image sample and the image features of the target image sample.
[0182] For example, the extraction unit 302 can be specifically configured to encode the defect image sample by using the defect hidden coding network of the preset image processing model to obtain a defect hidden code, and use the defect hidden code as the defect features of the defect image sample. Encode the target image sample by using the image hidden coding network of the preset image processing model to obtain an image hidden code, and use the image hidden code as the image features of the target image sample.
[0183] (3) Addition unit 303;
[0184] The addition unit 303 is configured to convert the defect type information into defect condition features, and add the defect condition features to the defect features to obtain target defect features.
[0185] For example, the addition unit 303 can be specifically configured to extract the type information from the defect type information and convert the type information into defect condition features. Or, it can also perform one-hot encoding or other encoding methods on the defect type information to obtain a defect type code, and use the defect type code as the defect condition features. According to the dimension information of the defect features, determine the splitting quantity of the defect features, split the defect features into pairs of defect sub-features with the splitting quantity, obtain a set of defect sub-features, and add the defect condition features to each defect sub-feature in the set of defect sub-features to obtain target defect features.
[0186] (4) Fusion unit 304;
[0187] The fusion unit 304 is configured to fuse the target defect feature and the image feature to obtain a target defect image, which is an image obtained by complementing the defect instance in the mask region of the target image sample.
[0188] For example, the fusion unit 304 may specifically be configured to extract a position feature from the target image sample, add the position feature to the image feature to obtain a target image feature. Fuse the target image feature and the target defect feature to obtain a fused defect feature, add the fused defect feature to the image feature to obtain a fused image feature, and generate a target defect image based on the fused image feature.
[0189] (5) Completion unit 305;
[0190] The completion unit 305 is configured to converge a preset image processing model based on the target defect image, the defect image sample, the target image sample, the original image sample, the defect feature, and the image feature to obtain a trained image processing model, and use the trained image processing model to complete the defect instance in the image to be completed.
[0191] For example, the completion unit 305 may specifically be configured to determine the image loss information of the preset image processing model according to the target defect image, the defect image sample, the target image sample, the original image sample, and the defect feature, determine the feature distribution loss information of the preset image processing model based on the defect feature and the image feature, fuse the image loss information and the feature distribution loss information, and converge the preset image processing model based on the fused loss information to obtain a trained image processing model. Use the trained image processing model to complete the defect instance in the image to be completed.
[0192] Optionally, in some embodiments, the image processing apparatus may further include a detection unit 306, as Figure 10 shown, specifically as follows:
[0193] The detection unit 306 is configured to train a preset defect detection model based on the target image, and use the trained defect detection model to detect defects in the image to be detected.
[0194] For example, the detection unit 306 may specifically be configured to label the defect identifier in the target image to obtain a current image sample, train a preset defect detection model based on the current image sample to obtain a trained defect detection model, and use the trained defect detection model to detect defects in the image to be detected.
[0195] In specific implementation, the above units may be implemented as independent entities, or may be arbitrarily combined and implemented as the same or several entities. For the specific implementation of the above units, reference may be made to the foregoing method embodiments, which will not be elaborated herein.
[0196] As can be seen from the above, in the embodiment of the present application, after the acquisition unit 301 acquires the original image sample, as well as the defect image sample and defect type information corresponding to at least one defect instance, and adds a preset mask to the original image sample to obtain the target image sample, the extraction unit 302 uses a preset image processing model to extract features from the defect image sample and the target image sample respectively, to obtain the defect features of the defect image sample and the image features of the target image sample. Then, the addition unit 303 converts the defect type information into defect condition features, and adds the defect condition features to the defect features to obtain the target defect features. Then, the fusion unit 304 fuses the target defect features and the image features to obtain the target defect image. The completion unit 305 converges the preset image processing model based on the target defect image, the defect image sample, the target image sample, the original image sample, the defect features and the image features, to obtain the trained image processing model, and uses the trained image processing model to complete the defect instance in the image to be completed; since this solution converts the defect type information into defect condition features, and controls the generation of the defect type by using the defect condition features as conditional bits, thereby ensuring that the domain type diversity of the completed defect image is controllable. In addition, the domain content between the original image sample and the defect instance is not the same, so that the application scenarios of completing instances with different domain contents can be increased. Therefore, the accuracy of image processing can be improved.
[0197] The embodiment of the present invention also provides an electronic device, as Figure 11 shown, which shows a schematic structural diagram of the electronic device involved in the embodiment of the present invention. Specifically:
[0198] The electronic device may include a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, an input unit 404 and other components. Those skilled in the art can understand that Figure 11 the structure of the electronic device shown in
[0199] The processor 401 is the control center of the electronic device, connecting various parts of the entire electronic device through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 402, and by invoking the data stored in the memory 402, it executes various functions of the electronic device and processes data. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 401 either.
[0200] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, image playback function, etc.); the data storage area can store the data created according to the use of the electronic device. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.
[0201] The electronic device further includes a power supply 403 for powering each component. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 403 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0202] The electronic device may further include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.
[0203] Although not shown, the electronic device may further include a display unit, etc., which will not be elaborated here. Specifically, in this embodiment, the processor 401 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402 to realize various functions as follows:
[0204] Obtain the original image sample, as well as the defect image sample and defect type information corresponding to at least one defect instance, add a preset mask to the original image sample to obtain the target image sample, use a preset image processing model to extract features from the defect image sample and the target image sample respectively to obtain the defect features of the defect image sample and the image features of the target image sample, convert the defect type information into defect condition features, and add the defect condition features to obtain the target defect features, fuse the target defect features and the image features to obtain the target defect image, which is the image after complementing the defect instance in the mask area of the target image sample, and based on the target defect image, the defect image sample, the target image sample, the original image sample, the defect features and the image features, converge the preset image processing model to obtain the trained image processing model, and use the trained image processing model to complement the defect instance in the image to be complemented.
[0205] For example, the electronic device receives the original image sample uploaded by the terminal, as well as the defect image sample and defect type information corresponding to at least one defect instance. Alternatively, it can receive the original image sample uploaded by the terminal and the defect image sample corresponding to at least one defect instance, and identify the defect type information of the defect instance in the defect image sample. Alternatively, it can also randomly select an original image sample from the image sample library, identify the image type in the original image sample, and based on this image type, screen out at least one defect instance corresponding to this image type from the defect instance set, and obtain the defect image sample and defect type information corresponding to this defect instance. Obtain a preset mask, add this preset mask to the original image sample, and perform mask covering processing on the original image sample through this preset mask to obtain a target image sample. Use the defect hidden coding network of the preset image processing model to encode the defect image sample to obtain a defect hidden code, and use this defect hidden code as the defect feature of the defect image sample. Use the image hidden coding network of the preset image processing model to encode the target image sample to obtain an image hidden code, and use this image hidden code as the image feature of the target image sample. Extract the type information from the defect type information and convert this type information into a defect condition feature. Alternatively, it can also perform one-hot encoding or other encoding methods on the defect type information to obtain a defect type code, and use this defect type code as the defect condition feature. According to the dimension information of the defect feature, determine the splitting quantity of the defect feature, split the defect feature into pairs of defect sub-features with the splitting quantity to obtain a set of defect sub-features, and add the defect condition feature to each defect sub-feature in the set of defect sub-features to obtain a target defect feature. Extract the position feature from the target image sample, and add the position feature to the image feature to obtain a target image feature. Fuse the target image feature and the target defect feature to obtain a fused defect feature, add the fused defect feature to the image feature to obtain a fused image feature, and generate a target defect image based on the fused image feature. Determine the image loss information of the preset image processing model according to the target defect image, defect image sample, target image sample, original image sample, and defect feature. Based on the defect feature and the image feature, determine the feature distribution loss information of the preset image processing model. Fuse the image loss information and the feature distribution loss information, and converge the preset image processing model based on the fused loss information to obtain a trained image processing model. Use the trained image processing model to complete the defect instance in the image to be completed to obtain a target image. Mark the defect identifier in the target image to obtain the current image sample, train the preset defect detection model based on the current image sample to obtain a trained defect detection model, and use the trained defect detection model to detect the defects in the image to be detected.
[0206] For the specific implementation of each of the above operations, reference can be made to the previous embodiments, which will not be elaborated here.
[0207] As can be seen from the above, in the embodiment of the present application, after obtaining the part image sample, as well as the defect image samples and defect type information corresponding to at least one defect instance, and adding a preset mask to the part image sample to obtain the target part image sample, a preset image processing model is used to extract features from the defect image sample and the target part image sample respectively, obtaining the defect features of the defect image sample and the part features of the target part image sample. Then, the defect type information is converted into defect conditional features, and the defect conditional features are added to the defect features to obtain the target defect features. Then, the target defect features and the part features are fused to obtain a defective part image. Based on the defective part image, the defect image sample, the target part image sample, the part image sample, the defect features, and the part features, the preset image processing model is converged to obtain a trained image processing model, and the trained image processing model is used to complete the defect instance in the part image to be completed. Since this solution converts the defect type information into defect conditional features and controls the generation of the defect type by using the defect conditional features as conditional bits, the domain type diversity of the completed defect images is ensured to be controllable. In addition, the domain content between the part image sample and the defect instance is not the same, so the application scenarios for completing instances with different domain contents can be increased. Therefore, the accuracy of image processing can be improved.
[0208] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions or by controlling related hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0209] Therefore, an embodiment of the present invention provides a computer-readable storage medium, which stores multiple instructions that can be loaded by a processor to execute the steps in any one of the image processing methods provided by the embodiments of the present invention. For example, the instructions can execute the following steps:
[0210] Obtain the original image sample, as well as the defect image sample and defect type information corresponding to at least one defect instance, add a preset mask to the original image sample to obtain the target image sample, use a preset image processing model to extract features from the defect image sample and the target image sample respectively to obtain the defect features of the defect image sample and the image features of the target image sample, convert the defect type information into defect condition features, and add the defect condition features to obtain the target defect features, fuse the target defect features and the image features to obtain the target defect image, which is the image after complementing the defect instance in the mask area of the target image sample. Based on the target defect image, the defect image sample, the target image sample, the original image sample, the defect features and the image features, converge the preset image processing model to obtain the trained image processing model, and use the trained image processing model to complement the defect instance in the image to be complemented.
[0211] For example, receive the original image sample uploaded by the receiving terminal, as well as the defect image sample and defect type information corresponding to at least one defect instance. Alternatively, receive the original image sample uploaded by the receiving terminal and the defect image sample corresponding to at least one defect instance, and identify the defect type information of the defect instance in the defect image sample. Alternatively, randomly select an original image sample from the image sample library, identify the image type in the original image sample, and based on this image type, screen out at least one defect instance corresponding to this image type from the defect instance set, and obtain the defect image sample and defect type information corresponding to this defect instance. Obtain a preset mask, add this preset mask to the original image sample, and perform mask covering processing on the original image sample through this preset mask to obtain a target image sample. Use the defect hidden coding network of the preset image processing model to encode the defect image sample to obtain a defect hidden coding, and use this defect hidden coding as the defect feature of the defect image sample. Use the image hidden coding network of the preset image processing model to encode the target image sample to obtain an image hidden coding, and use this image hidden coding as the image feature of the target image sample. Extract the type information from the defect type information and convert this type information into a defect condition feature. Alternatively, one-hot encode the defect type information or use other encoding methods to obtain a defect type encoding, and use this defect type encoding as the defect condition feature. According to the dimension information of the defect feature, determine the number of splits of the defect feature, split the defect feature into split number pairs of defect sub-features to obtain a set of defect sub-features, and add the defect condition feature to each defect sub-feature in the set of defect sub-features to obtain a target defect feature. Extract the position feature from the target image sample, and add the position feature to the image feature to obtain a target image feature. Fuse the target image feature and the target defect feature to obtain a fused defect feature, add the fused defect feature to the image feature to obtain a fused image feature, and generate a target defect image based on the fused image feature. Determine the image loss information of the preset image processing model according to the target defect image, defect image sample, target image sample, original image sample, and defect feature. Based on the defect feature and the image feature, determine the feature distribution loss information of the preset image processing model. Fuse the image loss information and the feature distribution loss information, and converge the preset image processing model based on the fused loss information to obtain a trained image processing model. Use the trained image processing model to complete the defect instance in the image to be completed to obtain a target image. Mark the defect identifier in the target image to obtain the current image sample, train the preset defect detection model based on the current image sample to obtain a trained defect detection model, and use the trained defect detection model to detect defects in the image to be detected.
[0212] For the specific implementation of each of the above operations, please refer to the previous embodiments and will not be elaborated here.
[0213] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.
[0214] Since the instructions stored in the computer-readable storage medium can execute the steps in any one of the image processing methods provided by the embodiments of the present invention, the beneficial effects achievable by any one of the image processing methods provided by the embodiments of the present invention can be realized. For details, see the previous embodiments and will not be elaborated here.
[0215] Among them, according to one aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in various optional implementation manners in the above-mentioned image processing aspect or image completion aspect.
[0216] The above has introduced in detail an image processing method, apparatus, electronic device, and computer-readable storage medium provided by the embodiments of the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those skilled in the art, based on the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. An image processing method, characterized in that, Including: Obtain an original image sample, as well as a defect image sample and defect type information corresponding to at least one defect instance, and add a preset mask to the original image sample to obtain a target image sample; Use a preset image processing model to extract features from the defect image sample and the target image sample respectively, to obtain defect features of the defect image sample and image features of the target image sample; Convert the defect type information into defect condition features, and add the defect condition features to the defect features to obtain target defect features; Fuse the target defect features and the image features to obtain a target defect image, where the target defect image is an image after complementing the defect instance in the masked area of the target image sample; Determine the image loss information of the preset image processing model according to the target defect image, the defect image sample, the target image sample, the original image sample, and the defect features; Determine the feature distribution loss information of the preset image processing model based on the defect features and the image features, including: performing convolution processing on the defect features to obtain processed defect features, performing convolution processing on the image features to obtain processed image features, sampling target latent features from a latent feature set of a preset normal distribution, and calculating the adversarial loss information among the target latent features, the processed defect features, and the processed image features to obtain the feature distribution loss information of the preset image processing model; Fuse the image loss information and the feature distribution loss information, and converge the preset image processing model based on the fused loss information to obtain a trained image processing model, and use the trained image processing model to complement the defect instance in the image to be complemented.
2. The image processing method according to claim 1, wherein The adding the defect condition features to the defect features to obtain target defect features includes: Determine the splitting number of the defect features according to the dimension information of the defect features; Split the defect features into defect sub-features corresponding to the splitting number to obtain a set of defect sub-features; Add the defect condition features to each defect sub-feature in the set of defect sub-features to obtain target defect features.
3. The image processing method according to claim 2, wherein The fusing the target defect features and the image features to obtain a target defect image includes: Extract position features from the target image sample, and add the position features to the image features to obtain target image features; Fuse the target image features and the target defect features to obtain fused defect features; Add the fused defect features to the image features to obtain fused image features, and generate a target defect image based on the fused image features.
4. The image processing method according to claim 3, wherein The fusing the target image features and the target defect features to obtain fused defect features includes: Convert the target image features into key features, and convert the target defect features into query features and value features; Fuse the key features and the query features, and determine the attention weights of the value features according to the fused features; Weight the value features based on the attention weights to obtain the fused defect features.
5. The image processing method according to claim 1, characterized in that, Determining the image loss information of the preset image processing model according to the target defect image, defect image samples, target image samples, original image samples, and defect features includes: Determine the reconstruction loss information of the preset image processing model based on the target defect image, target image samples, defect image samples, and defect features; Determine the domain loss information of the preset image processing model according to the target defect image, original image samples, and defect image samples; Fuse the reconstruction loss information and the domain loss information to obtain the image loss information of the preset image processing model.
6. The image processing method according to claim 5, wherein The determining the reconstruction loss information of the preset image processing model based on the target defect image, target image samples, defect image samples, and defect features includes: Perform masking processing on the target defect image according to the preset mask to obtain the masked defect image, and compare the masked defect image with the target image sample to obtain image reconstruction loss information; Reconstruct a defect image based on the defect features to obtain the reconstructed defect image, and compare the reconstructed defect image with the defect image sample to obtain defect reconstruction loss information; Fuse the image reconstruction loss information and the defect reconstruction loss information to obtain the reconstruction loss information of the preset image processing model.
7. The image processing method according to claim 6, characterized in that, The comparing the masked defect image with the target image sample to obtain image reconstruction loss information includes: Extract image features from the masked defect image to obtain first sample image features; Extract image features from the target image sample to obtain second sample image features; Calculate the feature distance between the first sample image features and the second sample image features to obtain image reconstruction loss information.
8. The image processing method according to claim 5, characterized in that The determining the domain loss information of the preset image processing model according to the target defect image, original image samples, and defect image samples includes: Identify the image in the masked area in the target defect image to obtain the current defect image; Determine the image domain loss information of the preset image processing model according to the target defect image and the original image sample; Determine the defect domain loss information of the preset image processing model based on the current defect image and the defect image sample; Fuse the image domain loss information and the defect domain loss information to obtain the domain loss information of the preset image processing model.
9. The image processing method according to claim 8, wherein The determining the defect domain loss information of the preset image processing model based on the current defect image and the defect image sample includes: Extract domain features from the current defect image to obtain current defect domain features, and add defect conditional features to the current defect domain features to obtain target defect domain features; Determine the current defect type of the current defect image and the type probability corresponding to the current defect type based on the target defect domain features; Calculate the adversarial loss information between the current defect type and the defect type of the defect image sample based on the type probability to obtain the defect domain loss information.
10. The image processing method according to claim 1, wherein Completing the defect instance in the image to be completed using the trained image processing model includes: Receiving an image processing request, where the image processing request carries the image to be processed and the defect identifier corresponding to the image to be processed; Adding a preset mask to the image to be processed to obtain an image to be completed, and using the trained image processing model to extract features from the image to be completed to obtain the current image features; Completing the target defect instance corresponding to the defect identifier in the image to be completed according to the set of latent features of the preset normal distribution, the current image features, and the defect condition features to obtain a target image.
11. The image processing method according to claim 10, wherein The step of completing the target defect instance corresponding to the defect identifier in the image to be completed according to the set of latent features of the preset normal distribution, the current image features, and the defect condition features to obtain a target image includes: Sampling a current latent feature from the set of latent features of the preset normal distribution, and screening out the target defect condition features corresponding to the defect identifier from the defect condition features; Adding the target defect condition features to the current latent feature to obtain a defect feature to be completed; Fusing the defect feature to be completed and the current image features, and generating the target image based on the completed image features.
12. The image processing method according to claim 10, wherein After completing the target defect instance corresponding to the defect identifier in the image to be completed according to the set of latent features of the preset normal distribution, the current image features, and the defect condition features to obtain a target image, it further includes: Annotating the defect identifier in the target image to obtain a current image sample; Training a preset defect detection model based on the current image sample to obtain a trained defect detection model; Using the trained defect detection model to detect defects in the image to be detected.
13. An image processing apparatus, characterized in that, It includes: An acquisition unit for acquiring an original image sample, as well as a defect image sample and defect type information corresponding to at least one defect instance, and adding a preset mask to the original image sample to obtain a target image sample; An extraction unit for using a preset image processing model to extract features from the defect image sample and the target image sample respectively to obtain the defect features of the defect image sample and the image features of the target image sample; An addition unit for converting the defect type information into defect condition features and adding the defect condition features to the defect features to obtain target defect features; A fusion unit for fusing the target defect features and the image features to obtain a target defect image, where the target defect image is an image after completing the defect instance in the masked area of the target image sample; A completion unit for determining the image loss information of the preset image processing model according to the target defect image, the defect image sample, the target image sample, the original image sample, and the defect features; Based on the defect features and the image features, determining the feature distribution loss information of the preset image processing model includes: performing convolution processing on the defect features to obtain processed defect features, performing convolution processing on the image features to obtain processed image features, sampling a target latent feature from a set of latent features in a preset normal distribution, calculating the adversarial loss information among the target latent feature, the processed defect features, and the processed image features, to obtain the feature distribution loss information of the preset image processing model; fusing the image loss information and the feature distribution loss information, and converging the preset image processing model based on the fused loss information to obtain a trained image processing model, and using the trained image processing model to complete the defect instance in the image to be completed.
14. An electronic device, characterized in that, It includes a processor and a memory, the memory stores an application program, and the processor is used to run the application program in the memory to execute the steps in the image processing method according to any one of claims 1 to 12.
15. A computer program product comprising computer programs / instructions, characterized in that, When the computer program / instructions are executed by the processor, the steps in the image processing method according to any one of claims 1 to 12 are implemented.
16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by the processor to execute the steps in the image processing method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Image processing method and device, electronic equipment and storage medium
CN110348537A
Image editing method and device, electronic equipment and storage medium
CN111814566A
Defect detection method and device, electronic equipment and readable storage medium
CN112801047A
Defect image generation method based on generative adversarial network
CN114022586A