Training Method of Model, Image Processing Method and Device, and Medium

By segmenting and preprocessing the semantic images to be processed and training with generators and discriminators, the problem that the existing generative adversarial network model is difficult to apply to conditional generation tasks is solved, and the model performance is improved.

CN115424013BActive Publication Date: 2025-06-24PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210820156.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-13
Publication Date
2025-06-24
Estimated Expiration
2042-07-13

AI Technical Summary

Technical Problem

The existing generative adversarial network model is difficult to apply to conditional generative tasks, and the model performance is poor.

Method used

A model training method is proposed. By obtaining the semantic images to be processed, segmenting and preprocessing them, and using the generator and discriminator of the initial network model for training, the generative adversarial network model can be applied to conditional generation tasks.

Benefits of technology

The performance of the generative adversarial network model is effectively improved, making it better applicable to conditional generation tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115424013B_ABST
    Figure CN115424013B_ABST
Patent Text Reader

Abstract

This embodiment provides a training method, an image processing method and device, and a medium for a model, belonging to the field of artificial intelligence technology, including: performing segmentation processing on a semantic image to be processed to obtain semantic image blocks; performing data preprocessing on the semantic image blocks to obtain feature vectors; inputting the feature vectors into an encoder in a generator of an initial network model for data processing to obtain output vectors; inputting the output vectors into a multi-layer perceptron in the generator for data mapping processing to obtain target image blocks; performing image reorganization processing on the target image blocks in the generator to obtain target images; inputting a preset comparison image and the target images into a discriminator of the initial network model for discrimination processing to obtain image discrimination values; performing training processing on the initial network model according to a preset loss function and the image discrimination values to obtain a generative adversarial network model; this generative adversarial network model can be applied to conditional generation tasks and effectively improve the model performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of artificial intelligence technology, and in particular, to a method for training a model, an image processing method and device, and a medium. Background Art

[0002] With the development of artificial intelligence technology, the usage rate of generative adversarial network models has gradually increased. In related technologies, generative adversarial network models are usually applied to unconditional generation tasks. However, in the field of image processing, they often correspond to conditional generation tasks, and current generative adversarial network models are difficult to be applied to the above-mentioned generation tasks, resulting in poor model performance. Summary of the Invention

[0003] The main purpose of the embodiments of the present application is to propose a method for training a model, an image processing method and device, and a medium, such that the trained generative adversarial network model can be applied to conditional generation tasks and effectively improve the model performance.

[0004] To achieve the above object, a first aspect of the embodiments of the present application proposes a method for training a model, the training method including:

[0005] Obtain a semantic image to be processed, perform segmentation processing on the semantic image to be processed to obtain a plurality of semantic image blocks;

[0006] Perform data preprocessing on the semantic image blocks to obtain a feature vector corresponding to each semantic image block;

[0007] Input the feature vector into an encoder in a generator of an initial network model for data processing to obtain an output vector, wherein a fast Fourier transform module is provided in the encoder, and the fast Fourier transform module is used to extract features from the feature vector;

[0008] Input the output vector into a multi-layer perceptron in the generator of the initial network model for data mapping processing to obtain a target image block;

[0009] Perform image reorganization processing on the target image block in the generator of the initial network model to obtain a target image;

[0010] Input a preset comparison image and the target image into a discriminator of the initial network model for discrimination processing to obtain an image discrimination value;

[0011] Train the initial network model according to a preset loss function and the image discrimination value to obtain a generative adversarial network model.

[0012] In some embodiments, the encoder includes a first encoding processing module and a second encoding processing module, and a fast Fourier transform module is provided in both the first encoding processing module and the second encoding processing module;

[0013] Processing the feature vector by the encoder in the generator of the initial network model to obtain an output vector includes:

[0014] Inputting the feature vector into the first encoding processing module for feature encoding to obtain a first vector, wherein the fast Fourier transform module in the first encoding processing module is used for feature extraction of the feature vector;

[0015] Inputting the first vector into the second encoding processing module for feature encoding to obtain an output vector, wherein the fast Fourier transform module in the second encoding processing module is used for feature extraction of the first vector.

[0016] In some embodiments, a multi-head self-attention module is further provided in the first encoding processing module;

[0017] Inputting the feature vector into the first encoding processing module for feature encoding to obtain a first vector includes:

[0018] Inputting the feature vector into the multi-head self-attention module of the first encoding processing module for feature processing to obtain a second vector;

[0019] Inputting the feature vector into the fast Fourier transform module of the first encoding processing module for feature extraction to obtain a third vector;

[0020] Performing residual sum processing and normalization processing on the feature vector, the second vector, and the third vector to obtain a first vector.

[0021] In some embodiments, a fully connected layer is further provided in the second encoding processing module;

[0022] Inputting the first vector into the second encoding processing module for feature encoding to obtain an output vector includes:

[0023] Inputting the first vector into the fully connected layer of the second encoding processing module for classification processing to obtain a fourth vector;

[0024] Inputting the first vector into the fast Fourier transform module of the second encoding processing module for feature extraction to obtain a fifth vector;

[0025] Perform residual summation processing and normalization processing on the first vector, the fourth vector, and the fifth vector to obtain an output vector.

[0026] In some embodiments, the fast Fourier transform module of the second encoding processing module includes a first Fourier transform unit, a first convolution unit, a first activation layer, a second convolution unit, and a second Fourier transform unit;

[0027] The step of inputting the first vector into the fast Fourier transform module of the second encoding processing module for feature extraction to obtain a fifth vector includes:

[0028] Input the first vector into the first Fourier transform unit for feature extraction to obtain first vector feature data;

[0029] Successively pass the first vector feature data through the convolution processing of the first convolution unit, the activation processing of the first activation layer, and the convolution processing of the second convolution unit to obtain first target vector data;

[0030] Input the first target vector data into the second Fourier transform unit for feature extraction to obtain a fifth vector.

[0031] In some embodiments, the step of training the initial network model according to a preset loss function and the image discrimination value to obtain a generative adversarial network model includes:

[0032] Update the parameters corresponding to the generator and the discriminator in the initial network model according to the preset loss function to obtain an updated image discrimination value;

[0033] When the updated image discrimination value is greater than or equal to a preset value, train to obtain the generative adversarial network model.

[0034] In some embodiments, the step of performing data preprocessing on the semantic image block to obtain a feature vector corresponding to each semantic image block includes:

[0035] Perform position encoding processing on the semantic image block to obtain position encoding data corresponding to each semantic image block;

[0036] Input the position encoding data into a preset linear layer for linear processing to obtain a feature vector corresponding to each semantic image block.

[0037] A second aspect of the embodiments of the present application proposes an image processing method, including:

[0038] Obtain an original semantic segmentation image;

[0039] Input the original semantic segmentation image into a generative adversarial network model for image processing to obtain a target semantic image, where the generative adversarial network model is trained according to the training method described in any one of the embodiments of the first aspect of this application.

[0040] A third aspect of the embodiments of this application provides a computer device, which includes a memory and a processor. Among them, a program is stored in the memory, and when the program is executed by the processor, the processor is used to execute the method described in any one of the embodiments of the first aspect of this application or the method described in any one of the embodiments of the second aspect of this application.

[0041] A fourth aspect of the embodiments of this application provides a storage medium, which is a computer-readable storage medium. The storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute the method described in any one of the embodiments of the first aspect of this application or the method described in any one of the embodiments of the second aspect of this application.

[0042] The training method, image processing method, device, and medium of the model proposed in the embodiments of this application obtain a semantic image to be processed, perform segmentation processing on the semantic image to be processed to obtain a number of semantic image blocks; perform data preprocessing on the semantic image blocks to obtain a feature vector corresponding to each semantic image block; input the feature vector into the encoder in the generator of the initial network model for data processing to obtain an output vector, where a fast Fourier transform module is provided in the encoder, and the fast Fourier transform module is used to extract features from the feature vector; input the output vector into the multi-layer perceptron in the generator of the initial network model for data mapping processing to obtain a target image block; perform image reorganization processing on the target image block in the generator of the initial network model to obtain a target image; input a preset comparison image and the target image into the discriminator of the initial network model for discrimination processing to obtain an image discrimination value; perform training processing on the initial network model according to a preset loss function and the image discrimination value to obtain a generative adversarial network model. The embodiments of this application train the initial network model by obtaining a semantic image to be processed, so that the trained generative adversarial network model can be applied to conditional generation tasks, effectively improving the model performance. Description of the Drawings

[0043] Figure 1 is a schematic flowchart of the training method of the model provided by the embodiments of this application;

[0044] Figure 2 is Figure 1 a sub-flowchart of step S200 in

[0045] Figure 3 is Figure 1 a schematic diagram of the sub - process of step S300 in

[0046] Figure 4 is Figure 3 a schematic diagram of the sub - process of step S310 in

[0047] Figure 5 is Figure 3 a schematic diagram of the sub - process of step S320 in

[0048] Figure 6 is Figure 5 a schematic diagram of the sub - process of step S322 in

[0049] Figure 7 is Figure 1 a schematic diagram of the sub - process of step S700 in

[0050] Figure 8 is a schematic diagram of the semantic image block provided by the embodiment of the present application;

[0051] Figure 9 is a schematic diagram of the structure of the encoder in the generator provided by the embodiment of the present application;

[0052] Figure 10 is a schematic diagram of the fast Fourier transform module provided by the embodiment of the present application;

[0053] Figure 11 is a schematic diagram of the process of the image processing method provided by the embodiment of the present application;

[0054] Figure 12 is a block diagram of the module structure of the training device of the model provided by the embodiment of the present application;

[0055] Figure 13 is a block diagram of the module structure of the image processing device provided by the embodiment of the present application;

[0056] Figure 14 is a schematic diagram of the hardware structure of the computer device provided by the embodiment of the present application. Detailed implementation manners

[0057] In order to make the purpose, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0058] It should be noted that although the functional modules are divided in the device schematic diagram and the logical sequence is shown in the flowchart, in some cases, the steps shown or described may be executed in a different module division in the device or a different order in the flowchart. Terms such as "first", "second", etc. in the specification, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence.

[0059] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0060] In addition, the described features, structures or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known methods, devices, implementations or operations are not shown or described in detail to avoid obscuring aspects of the present disclosure.

[0061] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.

[0062] The flowcharts shown in the drawings are only exemplary illustrations and do not necessarily include all the contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.

[0063] First, several terms involved in this application are analyzed:

[0064] Artificial Intelligence (AI): It is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. Artificial intelligence is a branch of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing, and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. It is also a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0065] Natural Language Processing (NLP): It is a discipline that studies the language problems of human-computer interaction. According to the different difficulties of technology implementation, such systems can be divided into three types: simple matching type, fuzzy matching type, and paragraph understanding type. The simple matching type tutoring and answering system mainly realizes the matching of the questions raised by students and the relevant response entries in the answer library through simple keyword matching technology, so as to automatically answer questions or provide relevant tutoring. The fuzzy matching type tutoring and answering system adds the matching of synonyms and antonyms on this basis. In this way, even if the questions raised by students cannot find direct matching answers in the answer library according to the original keywords, but if the words synonymous or antonymous with the keywords can be matched, relevant response entries can still be found in the answer library. The paragraph understanding type tutoring and answering system is the most ideal and truly intelligent tutoring and answering system (the simple matching type and the fuzzy matching type can strictly speaking only be called "automatic tutoring and answering systems" rather than "intelligent tutoring and answering systems").

[0066] Generative Adversarial Nets (GAN): It is a deep generative model based on adversarial learning. The generative adversarial network model generates quite good outputs through the mutual game learning of (at least) two structures in the framework: the generator and the discriminator, that is, the generator and the discriminator are trained at the same time and compete in the minimax algorithm. This adversarial method avoids some difficulties of some traditional generative models in practical applications and cleverly approximates some unsolvable loss functions through adversarial learning, and has a wide range of applications in the generation of data such as images, videos, natural languages, and music.

[0067] Transformer: As an attention-based encoder-decoder architecture, it has not only revolutionized the field of natural language processing but also made some pioneering work in the field of computer vision. Compared with Convolutional Neural Networks (CNNs), Vision Transformers have excellent modeling capabilities and performance. A Transformer consists of an encoding component, a decoding component, and the connection between them. Among them, the encoding component is composed of several encoders, and the decoding component is also composed of the same number of decoders (i.e., corresponding to the number of encoders).

[0068] Encoder-Decoder: It is a common model framework in deep learning. Many common applications are designed using the encoding-decoding framework. The Encoder and Decoder parts can be any text, speech, image, video data, etc. Based on Encoder-Decoder, various models can be designed.

[0069] Encoder: Encoding is to transform the input sequence into a fixed-length vector.

[0070] Decoder is to transform the previously generated fixed vector back into an output sequence. Among them, the input sequence can be text, speech, image, video; the output sequence can be text, image.

[0071] Loss Function: It is a function that maps the values of random events or their related random variables to non-negative real numbers to represent the "risk" or "loss" of the random event. In applications, the loss function is usually associated with the learning criterion and the optimization problem, that is, the model is solved and evaluated by minimizing the loss function. For example, it is used for parameter estimation of models in statistics and machine learning, for risk management and decision-making in macroeconomics, and for optimal control theory in control theory.

[0072] PIL (Python Image Library, Pillow): It is a Python (computer programming language) image processing library. It contains many encapsulations and has rich functions. It is one of the several commonly used image processing libraries in Python.

[0073] Multiple Perceptron (MLP) is a type of feedforward neural network. An MLP contains at least three node layers. Except for the input nodes, each node is a neuron using a non-linear activation function.

[0074] With the development of artificial intelligence technology, the usage rate of generative adversarial network models has gradually increased. In related technologies, generative adversarial network models are usually applied to unconditional generation tasks. However, in the field of image processing, they often correspond to conditional generation tasks, and current generative adversarial network models are difficult to be applied to the above-mentioned generation tasks, with poor model performance.

[0075] Based on this, the embodiments of the present application propose a model training method, an image processing method and device, and a medium. By obtaining a semantic image to be processed, segmenting the semantic image to be processed to obtain a plurality of semantic image blocks; performing data preprocessing on the semantic image blocks to obtain a feature vector corresponding to each semantic image block; inputting the feature vector into an encoder in a generator of an initial network model for data processing to obtain an output vector, wherein a fast Fourier transform module is provided in the encoder, and the fast Fourier transform module is used for feature extraction of the feature vector; inputting the output vector into a multi-layer perceptron in the generator of the initial network model for data mapping processing to obtain a target image block; performing image reorganization processing on the target image block in the generator of the initial network model to obtain a target image; inputting a preset comparison image and the target image into a discriminator of the initial network model for discrimination processing to obtain an image discrimination value; training the initial network model according to a preset loss function and the image discrimination value to obtain a generative adversarial network model. The embodiments of the present application train the initial network model by obtaining a semantic image to be processed, so that the trained generative adversarial network model can be applied to conditional generation tasks, effectively improving the model performance.

[0076] The embodiments of the present application provide a model training method, an image processing method and device, a computer device, and a storage medium, which are specifically described through the following embodiments. First, the model training method in the embodiments of the present application is described.

[0077] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0078] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0079] The training method of the model provided by the embodiments of the present application relates to the field of artificial intelligence technology. The training method of the model provided by the embodiments of the present application can be applied to a terminal, can also be applied to a server side, or can also be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart watch, etc.; the server side can be configured as an independent physical server, can also be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application for implementing the training method of the model, etc., but is not limited to the above forms.

[0080] The embodiments of the present application can be used in many general or specific computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, small computers, large computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0081] Referring to Figure 1 , according to the training method of the model of the first aspect embodiment of the present application, it includes but is not limited to steps S100 to S700.

[0082] Step S100: Obtain a semantic image to be processed, perform segmentation processing on the semantic image to be processed, and obtain a plurality of semantic image blocks;

[0083] Step S200: Perform data preprocessing on the semantic image blocks to obtain a feature vector corresponding to each semantic image block;

[0084] Step S300: Input the feature vector into an encoder in a generator of an initial network model for data processing to obtain an output vector, where a fast Fourier transform module is provided in the encoder, and the fast Fourier transform module is used to extract features from the feature vector;

[0085] Step S400: Input the output vector into the multi-layer perceptron in the generator of the initial network model for data mapping processing to obtain a target image patch;

[0086] Step S500: Perform image reorganization processing on the target image patch in the generator of the initial network model to obtain a target image;

[0087] Step S600: Input the preset comparison image and the target image into the discriminator of the initial network model for discrimination processing to obtain an image discrimination value;

[0088] Step S700: Train the initial network model according to the preset loss function and the image discrimination value to obtain a generative adversarial network model.

[0089] In step S100 of some embodiments, the semantic image to be processed is segmented to obtain a plurality of semantic image patches. The segmentation process is to perform a cutting operation on the semantic image to be processed, which belongs to the data preprocessing part. Specifically, it can be implemented through a Python package, such as through PIL to implement the segmentation process. Exemplarily, the size of the obtained semantic image to be processed should be consistent with the size of the output target image. In some embodiments, step S100 can also be executed in the framework of the GAN model.

[0090] In step S200 of some embodiments, since the semantic image patches cannot be recognized by the initial network model, it is necessary to perform data preprocessing on the semantic image patches to obtain a feature vector corresponding to each semantic image patch. This feature vector can be recognized by a computer and / or the initial network model to meet the requirements of model training. Therefore, before inputting the segmented semantic image patches into the initial network model, it is necessary to perform data preprocessing on the semantic image patches to obtain a feature vector that meets the requirements of model training. Exemplarily, each semantic image patch corresponds to a feature vector.

[0091] In step S300 of some embodiments, the generator of the initial network model includes an encoder. All feature vectors are input into the encoder in the generator of the initial network model for data processing to obtain an output vector. Step S300 is specifically implemented through the encoder in the generator of the initial network model. Exemplarily, the generator includes multiple encoders, and each encoder is provided with a fast Fourier transform module. After all feature vectors are sequentially processed by each encoder, a final output vector is obtained. By setting the fast Fourier transform module to perform feature extraction on the feature vectors, it is convenient to obtain the output vector. The fast Fourier transform module can facilitate the acquisition of more global information, thereby refining more and more accurate features, which is beneficial to improving the accuracy of the output target image.

[0092] In step S400 of some embodiments, the generator of the initial network model includes a multi-layer perceptron. The output vector is processed by the multi-layer perceptron for data mapping to obtain a target image block. The multi-layer perceptron set can map multiple input output vectors to the output target image block, and the multi-layer perceptron can identify data that cannot be linearly separated. Exemplarily, the target image block is the real image block obtained after processing the semantic image to be processed.

[0093] In step S500 of some embodiments, in the generator of the initial network model, image reorganization processing (corresponding to segmentation processing) can be performed on each target image block by using data preprocessing, and then a complete target image can be obtained. Correspondingly, steps S300 to S500 are the processing flow of the generator of the initial network model.

[0094] In step S600 of some embodiments, it is the processing flow of the discriminator of the initial network model. By inputting a preset comparison image and a target image into the discriminator of the initial network model for discrimination processing, an image discrimination value is obtained. The purpose of the discriminator set is to distinguish the difference between the target image generated by the generator and the comparison image, that is, the set real image. That is, there are two inputs to the discriminator, one is the target image and the other is the comparison image. The structure of the discriminator is similar to that of the generator. The discriminator also includes multiple encoders. Different from the generator, the output of the generator is an image, while the output of the discriminator is a value, that is, the image discrimination value. The image discrimination value represents the authenticity between the target image and the comparison image. That is, the discriminator in the embodiment of the present application corresponds to the generated target image and the preset comparison image, and its purpose is to distinguish the authenticity of the generated target image to better guide the generator to generate a more ideal target image. It should be noted that a target image corresponds to an image discrimination value. The larger the image discrimination value, the better the generated target image and the more real the image. Exemplarily, the numerical range of the image discrimination value is [0,1]. It should be noted that steps S100 to S600 are the forward process.

[0095] In step S700 of some embodiments, in the training process of the initial network model, the parameters corresponding to the generator and the discriminator are updated through the backward process, and the parameter update is achieved through multiple preset loss functions. Exemplarily, by minimizing the preset loss function to update the parameters of the model, the purpose of training can be achieved, that is, a generative adversarial network model is trained.

[0096] It should be noted that for the discriminator, in the embodiments of the present application, it is implemented by designing the decoder of the Transformer. Specifically, by introducing a class feature (i.e., the image discrimination value) as the output of the discriminator, this class feature is similar to the class feature of the BERT model, and its function is to score the authenticity of the generated target image, so as to obtain the authenticity score of the discriminator for the generated target image. The discriminator in the embodiments of the present application not only plays the role of discriminating the clarity of the target image, but also plays the role of discriminating the authenticity of the generated target image. If the generated target image is not the desired image, then generating a clearer image is meaningless.

[0097] In the related art, the generative adversarial network model based on the Transformer mainly designs the model itself. Assuming that the input is some random vectors (such as Gaussian noise), which can be regarded as sequence information in natural language processing in essence, so the encoder of the traditional Transformer can be migrated. For the input of the discriminator, which is the generated image, the structure of ViT (Vision Transformer) is used to discriminate the difference between the generated image and the real image. However, these works only solve the problem of traditional unconditional generation tasks. In practical applications, such as in the field of image processing, the application of the generative adversarial network model has requirements for the input. For example, the input is a semantic image, so that it can be applied to more downstream tasks (such as image editing, style transfer, etc.). Therefore, through the training method of the model in the embodiments of the present application, the trained generative adversarial network model can be applied to conditional generation tasks, effectively improving the model performance. Specifically, by designing a generative adversarial network model suitable for semantic image input (this generative adversarial network model includes a generator, a discriminator and a preset loss function), it not only maintains the function of the conditional generative adversarial network model, but also can give full play to the unique advantages of the Transformer.

[0098] As Figure 2 shown, it can be understood that step S200 is to perform data preprocessing on the semantic image blocks to obtain the feature vectors corresponding to each semantic image block, which specifically includes but is not limited to steps S210 to S220:

[0099] Step S210, perform position encoding processing on the semantic image blocks to obtain the position encoding data corresponding to each semantic image block;

[0100] Step S220, input the position encoding data into a preset linear layer for linear processing to obtain the feature vectors corresponding to each semantic image block.

[0101] Since the semantic image blocks obtained through segmentation processing cannot yet be recognized by the initial network model, it is necessary to perform data preprocessing on the semantic image blocks. Specifically, position encoding processing is performed on each semantic image block after segmentation processing. Exemplarily, position encoding processing is to add a position information recognizable by a computer, that is, the position encoding data corresponding to each semantic image block, to each semantic image block, and then perform a preprocessing, namely linear processing, on each position encoding data to obtain a feature vector recognizable by the computer. The way of this preprocessing is to pass through a preset linear layer. Specifically, the position encoding data is input into the preset linear layer for linear processing to obtain the feature vector corresponding to each semantic image block. This feature vector can be recognized by the computer and / or the initial network model to meet the requirements of model training. The set linear layer is such that each neuron is connected to all neurons in the previous layer to achieve linear combination and linear transformation of the previous layer.

[0102] As Figure 3 shown, it can be understood that the encoder includes a first encoding processing module and a second encoding processing module, and fast Fourier transform modules are provided in both the first encoding processing module and the second encoding processing module; step S300 is to input the feature vector into the encoder in the generator of the initial network model for data processing to obtain an output vector, specifically including but not limited to steps S310 to S320:

[0103] Step S310, input the feature vector into the first encoding processing module for feature encoding to obtain a first vector, where the fast Fourier transform module in the first encoding processing module is used to extract features from the feature vector;

[0104] Step S320, input the first vector into the second encoding processing module for feature encoding to obtain an output vector, where the fast Fourier transform module in the second encoding processing module is used to extract features from the first vector.

[0105] Specifically, the encoder includes a first encoding processing module and a second encoding processing module, and performs feature encoding through the first encoding processing module and the second encoding processing module to obtain more global information. The fast Fourier transform module in the first encoding processing module provided extracts features from the feature vector, and the fast Fourier transform module in the second encoding processing module extracts features from the first vector, so as to facilitate paying attention to more global information through the fast Fourier transform module. Exemplarily, the fast Fourier transform module maps the output features (i.e., the feature vector, the first vector) of the semantic image block with position encoding data into the frequency domain, and the frequency domain pays attention to both the real part and the imaginary part at the same time to enrich the information it contains. The embodiment of the present application obtains more global information by setting the fast Fourier transform module, thereby helping the model to extract more rich image representations, so as to improve the clarity of the generated target image.

[0106] As Figure 4 shown, it can be understood that a multi-head self-attention module is also provided in the first encoding processing module; step S310 is to input the feature vector into the first encoding processing module for feature encoding to obtain a first vector, which specifically includes but is not limited to steps S311 to S313:

[0107] Step S311: Input the feature vector into the multi-head self-attention module of the first encoding processing module for feature processing to obtain a second vector;

[0108] Step S312: Input the feature vector into the fast Fourier transform module of the first encoding processing module for feature extraction to obtain a third vector;

[0109] Step S313: Perform residual summation processing and normalization processing on the feature vector, the second vector, and the third vector to obtain a first vector.

[0110] In order to obtain more global information, the embodiments of the present application designed the attention mechanism of the Transformer, that is, the embodiments of the present application designed a brand-new residual Fourier self-attention block. Specifically, a multi-head self-attention module and a fast Fourier transform module are provided in the first encoding processing module. By inputting the feature vector into the multi-head self-attention module for feature processing, a second vector is obtained. Exemplarily, the attention mechanism generally exists widely in deep learning network structures, which can improve the learning effect of the model. The multi-head self-attention module set in the embodiments of the present application can group (head) each attention operation, so as to extract feature information from multiple dimensions. Then, the feature vector is input into the fast Fourier transform module for feature extraction to obtain a third vector. The fast Fourier transform module can focus on more global information, thereby improving the accuracy of the output target image. By performing residual summation processing and normalization processing on the feature vector, the second vector, and the third vector, a first vector is obtained. The embodiments of the present application set the residual summation processing, which can prevent gradient explosion and network degradation (the key to making the network deeper).

[0111] As Figure 5 shown, it can be understood that a fully connected layer is also provided in the second encoding processing module; step S320 is to input the first vector into the second encoding processing module for feature encoding to obtain an output vector, which specifically includes but is not limited to steps S321 to S323:

[0112] Step S321: Input the first vector into the fully connected layer of the second encoding processing module for classification processing to obtain a fourth vector;

[0113] Step S322: Input the first vector into the fast Fourier transform module of the second encoding processing module for feature extraction to obtain a fifth vector;

[0114] Step S323: Perform residual summation processing and normalization processing on the first vector, the fourth vector, and the fifth vector to obtain an output vector.

[0115] Specifically, a fully connected layer and a second encoding processing module are provided in the second encoding processing module of the embodiments of the present application. By inputting the first vector into the fully connected layer of the second encoding processing module for classification processing, a fourth vector is obtained. Exemplarily, in a fully connected layer, each node is connected to all nodes in the previous layer, which is used to synthesize the previously extracted features (i.e., the first vector). Due to its fully connected characteristic, generally, the fully connected layer has the most parameters. Similar to an MLP, each neuron in the fully connected layer is fully connected to all neurons in its previous layer. The fully connected layer can integrate local information with class discrimination. To improve the performance of the model network, the activation function of each neuron in the fully connected layer generally uses the ReLU (Rectified Linear Activation Function). The output value of the last fully connected layer is passed to an output, that is, the fourth vector. Then, the first vector is input into the fast Fourier transform module of the second encoding processing module for feature extraction to obtain a fifth vector. The fast Fourier transform module can focus on more global information, thereby improving the accuracy of the output target image. By performing residual summation processing and normalization processing on the first vector, the fourth vector, and the fifth vector, an output vector is obtained. The embodiments of the present application set up residual summation processing, which can prevent gradient explosion and network degradation (the key to making the network deeper).

[0116] As Figure 6 shown, it can be understood that the fast Fourier transform module of the second encoding processing module includes a first Fourier transform unit, a first convolution unit, a first activation layer, a second convolution unit, and a second Fourier transform unit; step S322 is to input the first vector into the fast Fourier transform module of the second encoding processing module for feature extraction to obtain a fifth vector, which specifically includes but is not limited to steps S324 to S326:

[0117] Step S324: Input the first vector into the first Fourier transform unit for feature extraction to obtain first vector feature data;

[0118] Step S325: After sequentially passing the first vector feature data through the convolution processing of the first convolution unit, the activation processing of the first activation layer, and the convolution processing of the second convolution unit, first target vector data is obtained;

[0119] Step S326: Input the first target vector data into the second Fourier transform unit for feature extraction to obtain a fifth vector.

[0120] In the embodiment of the present application, a fast Fourier transform module is provided in the second encoding processing module to be able to focus on more global information, thereby improving the accuracy of the output target image. Specifically, the fast Fourier transform module of the second encoding processing module includes a first Fourier transform unit, a first convolution unit, a first activation layer, a second convolution unit, and a second Fourier transform unit. Among them, both the first Fourier transform unit and the second Fourier transform unit perform feature extraction to facilitate focusing on more global information. Exemplarily, by mapping the first vector (first target vector data) to the frequency domain, the frequency domain simultaneously focuses on the real part and the imaginary part to enrich the information it contains, thereby helping the model to extract more rich image features, so as to improve the clarity of the generated target image.

[0121] In step S325 of some embodiments, after the first vector feature data is successively subjected to the convolution processing of the first convolution unit, the activation processing of the first activation layer, and the convolution processing of the second convolution unit, the first target vector data is obtained. The convolution processing of the first convolution unit and the convolution processing of the second convolution unit that are provided are both used for feature extraction of the first vector feature data. The convolution unit is the core of the model. Through convolution operations, two important purposes of dimensionality reduction processing and feature extraction can be achieved; and the activation processing of the first activation layer that is provided, the role of the first activation layer is to perform activation processing on the linear output of the previous layer through a non-linear activation function, which can be used to simulate any function, thereby enhancing the representation ability of the network.

[0122] It can be understood that the fast Fourier transform module of the first encoding processing module has the same structure as the fast Fourier transform module of the second encoding processing module. Specifically, the fast Fourier transform module of the first encoding processing module includes a third Fourier transform unit, a third convolution unit, a second activation layer, a fourth convolution unit, and a fourth Fourier transform unit; step S312 is to input the feature vector into the fast Fourier transform module of the first encoding processing module for feature extraction to obtain a third vector, which specifically includes but is not limited to the following steps:

[0123] Input the feature vector into the third Fourier transform unit for feature extraction to obtain second vector feature data; after the second vector feature data is successively subjected to the convolution processing of the third convolution unit, the activation processing of the second activation layer, and the convolution processing of the fourth convolution unit, the second target vector data is obtained; input the second target vector data into the fourth Fourier transform unit for feature extraction to obtain a third vector.

[0124] In the embodiment of the present application, a fast Fourier transform module is set in the first encoding processing module to be able to focus on more global information, thereby improving the accuracy of the output target image. Specifically, both the third Fourier transform unit and the fourth Fourier transform unit perform feature extraction, facilitating the focus on more global information. Exemplarily, by mapping the feature vector (the second target vector data) into the frequency domain, the frequency domain simultaneously focuses on the real part and the imaginary part to enrich the information it contains, thereby helping the model to extract more abundant image representations, and thus improving the clarity of the generated target image. The convolution processing of the set third convolution unit and the convolution processing of the fourth convolution unit are both used for feature extraction of the second vector feature data. The convolution unit is the core of the model. Through convolution operations, two important purposes of dimensionality reduction processing and feature extraction can be achieved; and the activation processing of the set second activation layer, the role of the second activation layer is to perform activation processing on the linear output of the previous layer through a non-linear activation function, which can be used to simulate any function, thereby enhancing the representation ability of the network.

[0125] Exemplarily, referring to Figure 9 , the figure shows the encoder in the generator of the embodiment of the present application. The encoder includes a first encoding processing module and a second encoding processing module. It can be understood that the legend only represents one encoder, while in the actual network (i.e., the generator), there are multiple encoders.

[0126] Exemplarily, referring to Figure 10 , the figure shows the fast Fourier transform module of an embodiment of the present application. Specifically, Figure 10 in: the real and imaginary part Fourier transform corresponds to the first Fourier transform unit (the third Fourier transform unit) of the embodiment of the present application, the first 1×1 convolution from top to bottom corresponds to the first convolution unit (the third convolution unit) of the embodiment of the present application, the activation layer corresponds to the first activation layer (the second activation layer) of the embodiment of the present application, the second 1×1 convolution from top to bottom corresponds to the second convolution unit (the fourth convolution unit) of the embodiment of the present application, and the real part fast Fourier transform corresponds to the second Fourier transform unit (the fourth Fourier transform unit) of the embodiment of the present application.

[0127] As Figure 7 shown, it can be understood that step S700 is to train the initial network model according to the preset loss function and the image discrimination value to obtain a generative adversarial network model, specifically including but not limited to steps S710 to S720:

[0128] Step S710, according to the preset loss function, update the parameters corresponding to the generator and the discriminator in the initial network model to obtain an updated image discrimination value;

[0129] Step S720: When the updated image discrimination value is greater than or equal to the preset value, a generative adversarial network model is trained.

[0130] Exemplarily, the L1 loss function, cross-entropy loss function, and perceptual loss function are used as the preset loss functions in the embodiments of the present application, and the initial network model is trained using the preset loss functions and the image discrimination value. Specifically, the parameters corresponding to the generator and discriminator of the initial network model are updated through the preset loss function. When the preset loss function tends to converge, the training is completed, that is, a generative adversarial network model is obtained. That is, according to the preset loss function, the parameters corresponding to the generator and discriminator in the initial network model are updated to obtain the updated image discrimination value. When the updated image discrimination value is greater than or equal to the preset value, a generative adversarial network model is trained.

[0131] Exemplarily, the calculation formula of the L1 loss function is as follows: L1 = |P t -P g |;

[0132] The calculation formula of the cross-entropy loss function is as follows:

[0133]

[0134] The calculation formula of the perceptual loss function is as follows:

[0135] where P t represents the comparison image, P g represents the generated target image, N represents the number of semantic image patches, represents the feature after the i-th ReLU layer of the trained VGG-19 (number of parameters).

[0136] In the embodiments of the present application, the Loss (loss value) is calculated by using the comparison image (i.e., the real image), and the preset loss functions such as the L1 loss function, cross-entropy loss function, and perceptual loss function are used to constrain the generation of a target image as similar as possible to the comparison image, so as to train the initial network model until convergence.

[0137] The L1 loss function and cross-entropy loss function adopted in the embodiments of the present application are used to make the generated target image as clear as possible, while the perceptual loss function is to maintain consistency at the feature level, and its role is to make the generated target image more real and conform to human perception.

[0138] It can be understood that the step of segmenting the to-be-processed semantic image in step S100 to obtain a plurality of semantic image patches includes, but is not limited to, the following steps: segmenting the to-be-processed semantic image according to a preset segmentation ratio to obtain a plurality of semantic image patches.

[0139] First, the input semantic image to be processed is segmented, and the specific preset segmentation ratio can be freely selected. Refer to Figure 8 , for example, the input semantic image to be processed is segmented to obtain a 7*7 semantic image patch. In some embodiments, the method of Swin Transformer can be referred to, and specifically, it can be adjusted through experiments.

[0140] After that, all semantic image patches are preprocessed to obtain the feature vector corresponding to each semantic image patch. For example, all semantic image patches are used as the input sequence and fed into an encoder similar to Swin Transformer for position encoding processing; the feature vector can represent the feature information of the image. For example, the position encoding processing can be represented by sin and cos, and its dimension only needs to be consistent with each semantic image patch.

[0141] In some embodiments, the embodiment of the present application can use an architecture similar to Swin Transformer as the generator of the initial network model or the generative adversarial network model.

[0142] The training method of the model proposed by the embodiment of the present application includes: obtaining the semantic image to be processed, segmenting the semantic image to be processed to obtain a plurality of semantic image patches; preprocessing the semantic image patches to obtain the feature vector corresponding to each semantic image patch; inputting the feature vector into the encoder in the generator of the initial network model for data processing to obtain an output vector, where a fast Fourier transform module is provided in the encoder for feature extraction of the feature vector; inputting the output vector into a multi-layer perceptron in the generator of the initial network model for data mapping processing to obtain a target image patch; performing image reorganization processing on the target image patch in the generator of the initial network model to obtain a target image; inputting the preset comparison image and the target image into the discriminator of the initial network model for discrimination processing to obtain an image discrimination value; training the initial network model according to the preset loss function and the image discrimination value to obtain a generative adversarial network model. The embodiment of the present application trains the initial network model by obtaining the semantic image to be processed, so that the trained generative adversarial network model can be applied to conditional generation tasks, effectively improving the model performance.

[0143] The embodiments of the present application provide a conditional generative adversarial network model based on Transformer, which can implement a generative adversarial network model with conditional image input (such as a semantic image to be processed); the embodiments of the present application also perform segmentation processing on the semantic image to be processed to achieve image serialization. In addition, a brand-new residual Fourier self-attention block is designed in the segmentation processing, and a fast Fourier transform module is added to the traditional self-attention block to obtain more global information, thereby helping the model extract more abundant image representations, so as to improve the clarity of the generated target image.

[0144] In related technologies, with the rapid development of Transformer, from its popularity in the NLP field to its application in almost all fields now, it shows the powerfulness of Transformer. It extracts the features of interest through a self-attention-based module and then uses these features to complete various downstream tasks. Since it is very friendly to sequence information, Transformer is indispensable in various tasks in the NLP field. Based on its excellent performance in NLP, many current researchers are committed to applying Transformer to the visual field. From ViT initially used for classification to Swin Transformer that can be applied to downstream tasks such as semantic segmentation and object detection, and many effects have exceeded those of traditional convolutional neural network models. However, there are currently few Transformers applied to generative tasks. Although there is a ViT GAN model for generative tasks, this model only successfully applies ViT to GAN, and its effect can only be approximated to that of traditional CNN-based models. Moreover, most current attention is paid to unconditional generative tasks, while conditional GAN models can be better applied to most downstream tasks, such as image style transfer and image editing.

[0145] Based on this, the embodiments of the present application provide a conditional GAN model based on Transformer, which improves the existing GAN model to further enhance its performance while ensuring the conditional GAN model.

[0146] Refer to Figure 11 , the embodiments of the present application also propose an image processing method, including but not limited to steps S800 to S900:

[0147] Step S800, obtaining an original semantic segmentation image;

[0148] Step S900: Input the original semantic segmentation image into the generative adversarial network model for image processing to obtain the target semantic image, where the generative adversarial network model is trained according to the training method of any embodiment of the first aspect of this application.

[0149] For the generative adversarial network model, the input of the embodiment of this application is: the original semantic segmentation image, and the output is the target semantic image corresponding to the original semantic segmentation image. The generative adversarial network model of the embodiment of this application is a conditional model based on Transformer, which can not only ensure the function of generating the target semantic image, but also realize image editing based on multiple prior images (such as the original semantic segmentation image). Compared with the related technology, the embodiment of this application not only uses the better Transformer architecture in the related technology, that is, the encoder part in Swin Transformer to design the generative adversarial network model, but also designs the self-attention module of the Transformer architecture to make its ability to focus on global information stronger, so that the obtained features are more informative, which is beneficial to generating a clearer target semantic image.

[0150] The embodiment of this application is to generate a target semantic image that meets the requirements. Different from the ViT GAN model in the related technology, the ViT GAN model generates a series of different images by giving noise, and these images are not necessarily the ones finally needed. In the embodiment of this application, by giving certain conditions (such as the parsing graph, semantic segmentation graph), a target semantic image that meets the requirements is generated, so as to achieve the purpose of controlling the generative adversarial network model to generate an ideal image. It can be understood that the clarity of the embodiment of this application has also been greatly improved. In practical applications, the generative adversarial network model of the embodiment of this application can be used to generate the desired image, that is, the target semantic image, rather than a random image that simply conforms to a certain distribution. In addition, the input image can also be the semantic segmentation graph or hand-drawn image of the desired generated image, and the output corresponds to the real image. The embodiment of this application can be applied to application scenarios such as drawing and restoration.

[0151] It should be noted that after the generative adversarial network model is trained by the training method of any embodiment of the first aspect of this application, that is, after the generator and discriminator are trained, it enters the testing or verification or application stage. During the testing, verification and application stages, the discriminator is not needed, that is, only the generator of the generative adversarial network model is used to generate the target semantic image.

[0152] The embodiment of this application also provides a training device for the model, refer to Figure 12, a training method of the above model can be implemented. The training device includes: a segmentation processing module 100, a preprocessing module 200, a data processing module 300, a mapping processing module 400, a rearrangement processing module 500, a discrimination processing module 600, and a training processing module 700.

[0153] Specifically, the segmentation processing module 100 is configured to obtain a semantic image to be processed, perform segmentation processing on the semantic image to be processed, and obtain a plurality of semantic image blocks; the preprocessing module 200 is configured to perform data preprocessing on the semantic image blocks to obtain a feature vector corresponding to each semantic image block; the data processing module 300 is configured to input the feature vector into an encoder in a generator of an initial network model for data processing to obtain an output vector. Among them, a fast Fourier transform module is provided in the encoder, and the fast Fourier transform module is configured to perform feature extraction on the feature vector; the mapping processing module 400 is configured to input the output vector into a multi-layer perceptron in a generator of the initial network model for data mapping processing to obtain a target image block; the rearrangement processing module 500 is configured to perform image rearrangement processing on the target image block in the generator of the initial network model to obtain a target image; the discrimination processing module 600 is configured to input a preset comparison image and the target image into a discriminator of the initial network model for discrimination processing to obtain an image discrimination value; the training processing module 700 is configured to perform training processing on the initial network model according to a preset loss function and the image discrimination value to obtain a generative adversarial network model.

[0154] The training device of the model in the embodiment of the present application is used to execute the training method of the model in the above embodiment, and its specific processing process is the same as that of the training method of the model in the above embodiment, and will not be repeated here one by one.

[0155] The embodiment of the present application further provides an image processing device. Referring to Figure 13 , the above image processing method can be implemented. The image processing device includes: an image acquisition module 800 and an image processing module 900. Specifically, the image acquisition module 800 is configured to acquire an original semantic segmentation image; the image processing module 900 is configured to input the original semantic segmentation image into a generative adversarial network model for image processing to obtain a target semantic image, where the generative adversarial network model is trained according to the training method in the first aspect embodiment of the present disclosure.

[0156] The image processing device of the embodiment of the present application is used to execute the image processing method in the above embodiment, and its specific processing process is the same as that of the image processing method in the above embodiment, and will not be repeated here one by one.

[0157] The embodiment of the present application further provides a computer device, including:

[0158] At least one processor, and,

[0159] A memory communicatively connected to at least one processor; wherein,

[0160] The memory stores instructions that are executed by at least one processor so that when the at least one processor executes the instructions, the training method of the first aspect embodiment of the present application or the image processing method of the second aspect embodiment of the present application is implemented.

[0161] The following will combine Figure 14 to elaborate in detail on the hardware structure of the computer device. The computer device includes: a processor 510, a memory 520, an input / output interface 530, a communication interface 540, and a bus 550.

[0162] The processor 510 can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;

[0163] The memory 520 can be implemented in forms such as a ROM (Read Only Memory), a static storage device, a dynamic storage device, or a RAM (Random Access Memory). The memory 520 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 520 and are called by the processor 510 to execute the model training method of the embodiments of the present application or execute the image processing method of the embodiments of the present application;

[0164] The input / output interface 530 is used to implement information input and output;

[0165] The communication interface 540 is used to implement communication interaction between this device and other devices, and can implement communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.); and

[0166] The bus 550 transmits information between the various components of the device (such as the processor 510, the memory 520, the input / output interface 530, and the communication interface 540);

[0167] Among them, the processor 510, the memory 520, the input / output interface 530, and the communication interface 540 are communicatively connected to each other inside the device through the bus 550.

[0168] An embodiment of the present application also provides a storage medium, which is a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the model training method or the image processing method of the embodiments of the present application.

[0169] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely located relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0170] The model training method, the image processing method, the device, and the medium proposed in the embodiments of the present application obtain a semantic image to be processed, perform segmentation processing on the semantic image to be processed to obtain a plurality of semantic image blocks; perform data preprocessing on the semantic image blocks to obtain a feature vector corresponding to each semantic image block; input the feature vector into an encoder in a generator of an initial network model for data processing to obtain an output vector, wherein a fast Fourier transform module is provided in the encoder for performing feature extraction on the feature vector; input the output vector into a multi-layer perceptron in the generator of the initial network model for data mapping processing to obtain a target image block; perform image reorganization processing on the target image block in the generator of the initial network model to obtain a target image; input a preset comparison image and the target image into a discriminator of the initial network model for discrimination processing to obtain an image discrimination value; and perform training processing on the initial network model according to a preset loss function and the image discrimination value to obtain a generative adversarial network model. The embodiments of the present application train the initial network model by obtaining a semantic image to be processed, so that the trained generative adversarial network model can be applied to conditional generation tasks, effectively improving the model performance.

[0171] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0172] Those skilled in the art can understand that Figures 1 to 7 , Figure 11 the technical solutions shown in

[0173] do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown, or combine some steps, or different steps.

[0174] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.

[0175] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0176] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can represent: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (one) of the following" or its similar expressions refer to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0177] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the above units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical or other forms.

[0178] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0179] In addition, each functional unit in various embodiments of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0180] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of this application. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs and other various media that can store programs.

[0181] The preferred embodiments of the embodiments of this application have been described above with reference to the drawings, but this does not limit the scope of the rights of the embodiments of this application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of this application shall be within the scope of the rights of the embodiments of this application.

Claims

1. A training method for a model, characterized in that, The training method includes: Obtain a semantic image to be processed, perform segmentation processing on the semantic image to be processed, and obtain a plurality of semantic image blocks; Perform data preprocessing on the semantic image blocks to obtain a feature vector corresponding to each semantic image block; Input the feature vector into the encoder in the generator of the initial network model for data processing to obtain an output vector; Input the output vector into the multi-layer perceptron in the generator of the initial network model for data mapping processing to obtain a target image block; Perform image reconstruction processing on the target image block in the generator of the initial network model to obtain a target image; Input a preset comparison image and the target image into the discriminator of the initial network model for discrimination processing to obtain an image discrimination value; According to a preset loss function, update the parameters corresponding to the generator and the discriminator in the initial network model to obtain an updated image discrimination value; When the updated image discrimination value is greater than or equal to a preset value, train to obtain a generative adversarial network model; Wherein, the encoder includes a first encoding processing module and a second encoding processing module, and a fast Fourier transform module is provided in both the first encoding processing module and the second encoding processing module; the step of inputting the feature vector into the encoder in the generator of the initial network model for data processing to obtain an output vector includes: Input the feature vector into the first encoding processing module for feature encoding to obtain a first vector, wherein the fast Fourier transform module in the first encoding processing module is used for feature extraction of the feature vector; Input the first vector into the second encoding processing module for feature encoding to obtain an output vector, wherein the fast Fourier transform module in the second encoding processing module is used for feature extraction of the first vector.

2. The training method according to claim 1, wherein A multi-head self-attention module is further provided in the first encoding processing module; The step of inputting the feature vector into the first encoding processing module for feature encoding to obtain a first vector includes: Input the feature vector into the multi-head self-attention module of the first encoding processing module for feature processing to obtain a second vector; Input the feature vector into the fast Fourier transform module of the first encoding processing module for feature extraction to obtain a third vector; Perform residual summation processing and normalization processing on the feature vector, the second vector, and the third vector to obtain a first vector.

3. The training method according to claim 1, wherein A fully connected layer is further provided in the second encoding processing module; The step of inputting the first vector into the second encoding processing module for feature encoding to obtain an output vector includes: Input the first vector into the fully connected layer of the second encoding processing module for classification processing to obtain a fourth vector; Input the first vector into the fast Fourier transform module of the second encoding processing module for feature extraction to obtain a fifth vector; Perform residual summation processing and normalization processing on the first vector, the fourth vector, and the fifth vector to obtain an output vector.

4. The training method according to claim 3, characterized in that The fast Fourier transform module of the second encoding processing module includes a first Fourier transform unit, a first convolution unit, a first activation layer, a second convolution unit, and a second Fourier transform unit; Inputting the first vector into the fast Fourier transform module of the second encoding processing module for feature extraction to obtain a fifth vector includes: Inputting the first vector into the first Fourier transform unit for feature extraction to obtain first vector feature data; After the first vector feature data is successively subjected to convolution processing by the first convolution unit, activation processing by the first activation layer, and convolution processing by the second convolution unit, first target vector data is obtained; Inputting the first target vector data into the second Fourier transform unit for feature extraction to obtain a fifth vector.

5. The training method according to any one of claims 1 to 4, characterized in that, Preprocessing the semantic image block to obtain a feature vector corresponding to each semantic image block includes: Performing position encoding processing on the semantic image block to obtain position encoding data corresponding to each semantic image block; Inputting the position encoding data into a preset linear layer for linear processing to obtain a feature vector corresponding to each semantic image block.

6. An image processing method, characterized in that, Including: Obtaining an original semantic segmentation image; Inputting the original semantic segmentation image into a generative adversarial network model for image processing to obtain a target semantic image, where the generative adversarial network model is trained according to the training method described in any one of claims 1 to 5.

7. A computer device, characterized in that, The computer device includes a memory and a processor, where a computer program is stored in the memory, and when the computer program is executed by the processor, the processor is used to execute: The training method described in any one of claims 1 to 5; or The image processing method described in claim 6.

8. A storage medium, the storage medium being a computer-readable storage medium, characterized in that, The computer-readable storage stores a computer program, and when the computer program is executed by a computer, the computer is used to execute: The training method described in any one of claims 1 to 5; or The image processing method described in claim 6.

Citation Information

Patent Citations

  • Model generation method and device, anomaly detection method and device and electronic equipment

    CN114400019A

  • Image segmentation method and apparatus, computer device, and storage medium

    WO2022105125A1