Face Image Enhancement Method, System and Medium Based on Self-Attention Network

By using a self-attention network in the generative adversarial network, the problem of local receptive field limitation in image enhancement is solved, and more efficient face image enhancement is achieved, ensuring the integrity and accuracy of the image.

CN114387181BActive Publication Date: 2025-06-20E SURFING IOT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111611284.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-27
Publication Date
2025-06-20
Estimated Expiration
2041-12-27

AI Technical Summary

Technical Problem

In the prior art, the generative adversarial network model built on convolutional neural networks has local receptive field limitations in the image enhancement process, resulting in the loss of deep details of the image, affecting the integrity and accuracy of the image.

Method used

The self-attention network is used as the basic unit for generating adversarial models, instead of the traditional underlying architecture of neural networks, so that the generative adversarial model pays more attention to the connections of various parts of the image, thereby generating enhanced images that better reflect the overall details of the face image data. The model is trained twice through the face database and actual scene data to improve the model's data enhancement ability and generalization ability.

Benefits of technology

The efficiency of face image enhancement is improved, the integrity and accuracy of face image is guaranteed, and the generated enhanced images can better reflect the overall details of face image data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114387181B_ABST
    Figure CN114387181B_ABST
Patent Text Reader

Abstract

The present invention discloses a face image enhancement method, system and medium based on a self-attention network. The method includes: constructing an initial generative adversarial model based on the self-attention network; obtaining a first face image through a face database, and training the initial generative adversarial model according to the first face image to obtain a first generative adversarial model; obtaining a second face image of an actual scene, and training the first generative adversarial model according to the second face image to obtain a second generative adversarial model; obtaining a third face image to be enhanced, and inputting the third face image into the second generative adversarial model to obtain an enhanced fourth face image. The present invention enables the output enhanced image to better reflect the overall details of the face image data, ensures the integrity and accuracy of the face image, and improves the efficiency of face image enhancement. The present invention can be widely applied to the field of image processing technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a face image enhancement method, system and medium based on a self-attention network. Background Art

[0002] In the process of building a smart campus, a safe campus is a very important part, which includes face detection and recognition, the deployment of personnel on campus, and the mapping of the nodes they pass through on campus. All these requirements need to be captured by cameras and then analyzed and identified by artificial intelligence in the background. However, in actual engineering projects, many times due to the movement or occlusion of objects, or the poor capture angle or dim light, or simply too few samples, it is often impossible to capture enough high-quality images to complete the efficient training of the recognition model. Therefore, a method for data enhancement based on existing image data is needed.

[0003] Existing common data enhancement solutions: In business scenarios where data enhancement is required, a generative adversarial network is generally used to enhance image data. The specific process is that after inputting real data pictures into the model, the model will automatically generate enhanced data images to achieve the purpose of enhancing the ability of the original recognition model. The internal core of the generative adversarial network model is divided into two parts, namely a generator and a discriminator. When training the generative adversarial model, these two sub-models are usually trained alternately, that is, using the game confrontation between them, and finally seeking a Nash equilibrium for data enhancement. Specifically, the generator is responsible for generating as realistic data pictures as possible to make the discriminator think they are real. Its input is a noise vector, and its output is as realistic image data as possible. The discriminator is responsible for trying to distinguish which inputs are real data and which data are generated by the generator. Its input is image data, and its output is true or false.

[0004] In the prior art, the generator and discriminator in a generative adversarial network model usually based on a convolutional neural network output the generated image data after several upsampling, fully connected layers, convolutional layers, and regularization processes. However, the generative adversarial model based on a convolutional neural network has only a local receptive field, and the deep details of the image will be lost as the model is trained, affecting the integrity and accuracy of the image. Summary of the Invention

[0005] The purpose of the present invention is to solve at least to some extent one of the technical problems existing in the prior art.

[0006] To this end, an object of an embodiment of the present invention is to provide a face image enhancement method based on a self-attention network. In this embodiment of the present invention, a self-attention network is used as a basic unit of a generative adversarial model, replacing the original underlying architecture of a neural network, enabling the generative adversarial model to pay more attention to the connections between various parts of an image, so that the output enhanced image can better reflect the overall details of face image data, ensuring the integrity and accuracy of the face image; the generative adversarial model is trained twice with face data in a face database and in an actual scenario, enabling the generative adversarial model to have a stronger face image data enhancement ability and a certain subsequent generalization ability, thereby improving the efficiency of face image enhancement.

[0007] Another object of an embodiment of the present invention is to provide a face image enhancement system based on a self-attention network.

[0008] To achieve the above technical objectives, the technical solutions adopted in the embodiments of the present invention include:

[0009] In a first aspect, an embodiment of the present invention provides a face image enhancement method based on a self-attention network, including the following steps:

[0010] Construct an initial generative adversarial model based on a self-attention network;

[0011] Obtain a first face image from a face database, and train the initial generative adversarial model according to the first face image to obtain a first generative adversarial model;

[0012] Obtain a second face image in an actual scenario, and train the first generative adversarial model according to the second face image to obtain a second generative adversarial model;

[0013] Obtain a third face image to be enhanced, and input the third face image into the second generative adversarial model to obtain an enhanced fourth face image.

[0014] Further, in an embodiment of the present invention, the step of constructing an initial generative adversarial model based on a self-attention network specifically includes:

[0015] Construct a self-attention network, which includes an input embedding layer, an attention mechanism module, and a feed-forward module;

[0016] Construct a generator and a discriminator according to the self-attention network;

[0017] Construct the initial generative adversarial model according to the generator and the discriminator.

[0018] Further, in an embodiment of the present invention, the generator includes a multi-layer perceptron, a plurality of upsampling modules, and a first linear flattening layer, and the upsampling module includes the self-attention network and an upsampling layer.

[0019] Further, in an embodiment of the present invention, the discriminator includes a plurality of the self-attention networks and a second linear flattening layer.

[0020] Further, in an embodiment of the present invention, the step of obtaining a first face image through a face database and training the initial generative adversarial model according to the first face image to obtain a first generative adversarial model is specifically as follows:

[0021] Obtain a first face image through a face database, input the first face image into the initial generative adversarial model, and alternately train the generator and the discriminator until the generator and the discriminator reach a Nash equilibrium to obtain the first generative adversarial model.

[0022] Further, in an embodiment of the present invention, the step of obtaining a second face image of an actual scene and training the first generative adversarial model according to the second face image to obtain a second generative adversarial model specifically includes:

[0023] Capture in the actual scene of the face recognition system to obtain the second face image;

[0024] Input the second face image into the first generative adversarial model, and alternately train the generator and the discriminator until the generator and the discriminator reach a Nash equilibrium to obtain the second generative adversarial model.

[0025] Further, in an embodiment of the present invention, the step of inputting the third face image into the second generative adversarial model to obtain an enhanced fourth face image specifically includes:

[0026] Input the third face image into the generator of the second generative adversarial model;

[0027] Perceive the third face image through the multi-layer perceptron to generate self-attention tokens;

[0028] Encode the self-attention tokens through the self-attention network and then perform upsampling through the upsampling layer to obtain a first enhanced data;

[0029] Fuse the third face image and the first enhanced data through the first linear flattening layer to obtain an enhanced fourth face image.

[0030] In a second aspect, an embodiment of the present invention provides a face image enhancement system based on a self-attention network, including:

[0031] A generative adversarial model construction module, configured to construct an initial generative adversarial model based on the self-attention network;

[0032] A first model training module, configured to obtain a first face image from a face database, and train the initial generative adversarial model according to the first face image to obtain a first generative adversarial model;

[0033] A second model training module, configured to obtain a second face image of an actual scene, and train the first generative adversarial model according to the second face image to obtain a second generative adversarial model;

[0034] An image enhancement module, configured to obtain a third face image to be enhanced, and input the third face image into the second generative adversarial model to obtain an enhanced fourth face image.

[0035] In a third aspect, an embodiment of the present invention provides a face image enhancement device based on a self-attention network, including:

[0036] At least one processor;

[0037] At least one memory, configured to store at least one program;

[0038] When the at least one program is executed by the at least one processor, the at least one processor is caused to implement the above-mentioned face image enhancement method based on a self-attention network.

[0039] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, in which a program executable by a processor is stored, and the program executable by the processor is used to execute the above-mentioned face image enhancement method based on a self-attention network when executed by the processor.

[0040] The advantages and beneficial effects of the present invention will be partially given in the following description, partially will become obvious from the following description, or will be understood through the practice of the present invention:

[0041] In the embodiment of the present invention, the self-attention network is used as the basic unit of the generative adversarial model, replacing the original neural network underlying architecture, so that the generative adversarial model will pay more attention to the connection of each part of the image, so that the output enhanced image can better reflect the overall details of the face image data, ensuring the integrity and accuracy of the face image; the generative adversarial model is trained twice with the face data in the face database and the actual scene, so that the generative adversarial model has a stronger face image data enhancement ability and a certain subsequent generalization ability, thereby improving the efficiency of face image enhancement. Brief Description of the Drawings

[0042] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following introduces the drawings required to be used in the embodiments of the present invention. It should be understood that the drawings in the following introduction are only for conveniently and clearly presenting some embodiments of the technical solutions in the present invention. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0043] Figure 1 It is a flowchart of the steps of a face image enhancement method based on a self-attention network provided by an embodiment of the present invention;

[0044] Figure 2 It is a schematic structural diagram of the generator of the generative adversarial model provided by an embodiment of the present invention;

[0045] Figure 3 It is a schematic structural diagram of the discriminator of the generative adversarial model provided by an embodiment of the present invention;

[0046] Figure 4 It is a block diagram of the structure of a face image enhancement system based on a self-attention network provided by an embodiment of the present invention;

[0047] Figure 5 It is a block diagram of the structure of a face image enhancement device based on a self-attention network provided by an embodiment of the present invention. Detailed Description of the Embodiments

[0048] The following details the embodiments of the present invention. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention. For the step numbers in the following embodiments, they are only set for the convenience of elaboration and explanation, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0049] In the description of the present invention, the meaning of "a plurality" is two or more. If the first and second are described, it is only for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence of the indicated technical features. In addition, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field of the present invention.

[0050] Refer to Figure 1, embodiments of the present invention provide a face image enhancement method based on a self-attention network, which specifically includes the following steps:

[0051] S101. Construct an initial generative adversarial model based on the self-attention network;

[0052] S102. Obtain a first face image through a face database, and train the initial generative adversarial model according to the first face image to obtain a first generative adversarial model;

[0053] S103. Obtain a second face image of an actual scene, and train the first generative adversarial model according to the second face image to obtain a second generative adversarial model;

[0054] S104. Obtain a third face image to be enhanced, and input the third face image into the second generative adversarial model to obtain an enhanced fourth face image.

[0055] Specifically, embodiments of the present invention use a self-attention network as the basic unit of the generative adversarial model, replacing the original underlying neural network architecture, enabling the generative adversarial model to pay more attention to the connections between various parts of the image, so that the output enhanced image can better reflect the overall details of the face image data, ensuring the integrity and accuracy of the face image; through the face database and the face data of the actual scene, the generative adversarial model is trained twice, enabling the generative adversarial model to have a stronger face image data enhancement ability and a certain subsequent generalization ability, thereby improving the efficiency of face image enhancement.

[0056] Further as an optional implementation manner, step S101 of constructing the initial generative adversarial model based on the self-attention network specifically includes:

[0057] S1011. Construct a self-attention network, which includes an input embedding layer, an attention mechanism module, and a feed-forward module;

[0058] S1012. Construct a generator and a discriminator according to the self-attention network;

[0059] S1013. Construct the initial generative adversarial model according to the generator and the discriminator.

[0060] Specifically, embodiments of the present invention use a self-attention network that has been very successful in natural language processing. It consists of an input embedding layer and multiple modules composed of two sub-parts. Both of these two sub-parts have a residual structure, namely an attention mechanism module with multiple heads and a feed-forward module.

[0061] As a further optional implementation, the generator includes a multi-layer perceptron, a plurality of upsampling modules, and a first linear flattening layer, and the upsampling module includes the self-attention network and an upsampling layer.

[0062] Specifically, as Figure 2 shown in the structural schematic diagram of the generator of the generative adversarial network provided by the embodiment of the present invention, first, a multi-layer perceptron is used to generate tokens to be used by the self-attention network, then a plurality of upsampling modules composed of the self-attention network and an upsampling layer, and finally a linear flattening layer is used to generate the final image data.

[0063] As a further optional implementation, the discriminator includes a plurality of the self-attention networks and a second linear flattening layer.

[0064] Specifically, as Figure 2 shown in the structural schematic diagram of the discriminator of the generative adversarial network provided by the embodiment of the present invention, the input image data is segmented, and then the determination result of the discriminator for the input can be obtained directly through several layers of self-attention networks and a linear flattening layer.

[0065] As a further optional implementation, the step of obtaining a first face image through a face database and training the initial generative adversarial model according to the first face image to obtain a first generative adversarial model is specifically as follows:

[0066] Obtain a first face image through a face database, input the first face image into the initial generative adversarial model, and alternately train the generator and the discriminator until the generator and the discriminator reach Nash equilibrium to obtain the first generative adversarial model.

[0067] Specifically, several common face databases can be used for the face database. Training the generative adversarial model with the first face image in the face database can enable the generative adversarial model to have the general ability to generate and recognize face images.

[0068] As a further optional implementation, the step of obtaining a second face image of an actual scene and training the first generative adversarial model according to the second face image to obtain a second generative adversarial model specifically includes:

[0069] Capture in the actual scene of the face recognition system to obtain the second face image;

[0070] Input the second face image into the first generative adversarial model, and alternately train the generator and the discriminator until the generator and the discriminator reach Nash equilibrium to obtain the second generative adversarial model.

[0071] Specifically, in the actual application scenarios of the face recognition system (such as campuses and communities), capturing faces to obtain the second face images for training the generative adversarial model can enable the generative adversarial model to have the ability to generate and recognize specific face images corresponding to the application scenarios.

[0072] Further as an optional implementation manner, the step of inputting the third face image into the second generative adversarial model to obtain the enhanced fourth face image specifically includes:

[0073] Inputting the third face image into the generator of the second generative adversarial model;

[0074] Generating self-attention tokens by perceiving the third face image through the multi-layer perceptron;

[0075] Encoding the self-attention tokens through the self-attention network and then performing upsampling through the upsampling layer to obtain the first enhanced data;

[0076] Fusing the third face image and the first enhanced data through the first linear flattening layer to obtain the enhanced fourth face image.

[0077] Specifically, first generate the self-attention tokens to be used by the self-attention network through the multi-layer perceptron. To reduce calculations, instead of using pixel-by-pixel calculations, block-based modeling can be considered. Then, alternately perform encoding and upsampling through several layers of self-attention networks and upsampling layers to obtain the enhanced data for image enhancement. Pixelshuffle method can be used for upsampling. Finally, fuse the enhanced data with the original image to obtain the enhanced face image.

[0078] The above has described the method steps of the embodiments of the present invention. It can be understood that the embodiments of the present invention address the problem of too little high-quality image data in actual engineering projects, perform data enhancement based on the existing image data, so as to efficiently train the face recognition model and improve its performance. The embodiments of the present invention mainly relate to the method of data enhancement of image data at the algorithm level, and avoid hardware-related upgrades and transformations to obtain higher-quality images, and are easy to be directly implemented and promoted in actual engineering projects.

[0079] Refer to Figure 4 , the embodiments of the present invention provide a face image enhancement system based on a self-attention network, including:

[0080] A generative adversarial model construction module, configured to construct an initial generative adversarial model based on a self-attention network;

[0081] The first model training module is used to obtain the first face image from a face database and train the initial generative adversarial model according to the first face image to obtain the first generative adversarial model;

[0082] The second model training module is used to obtain the second face image of an actual scene and train the first generative adversarial model according to the second face image to obtain the second generative adversarial model;

[0083] The image enhancement module is used to obtain the third face image to be enhanced, input the third face image into the second generative adversarial model, and obtain the enhanced fourth face image.

[0084] The content in the above method embodiments is applicable to the system embodiments of the present invention. The functions specifically implemented in the system embodiments of the present invention are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0085] Refer to Figure 5 , an embodiment of the present invention provides a face image enhancement device based on a self-attention network, including:

[0086] At least one processor;

[0087] At least one memory for storing at least one program;

[0088] When the above at least one program is executed by the above at least one processor, the above at least one processor implements the above-mentioned face image enhancement method based on a self-attention network.

[0089] The content in the above method embodiments is applicable to the device embodiments of the present invention. The functions specifically implemented in the device embodiments of the present invention are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0090] An embodiment of the present invention also provides a computer-readable storage medium, in which a program executable by a processor is stored. The program executable by the processor is used to execute the above-mentioned face image enhancement method based on a self-attention network when executed by the processor.

[0091] A computer-readable storage medium according to an embodiment of the present invention can execute a face image enhancement method based on a self-attention network provided by an embodiment of the method of the present invention, can execute any combination of implementation steps of the method embodiment, and has the corresponding functions and beneficial effects of the method.

[0092] An embodiment of the present invention also discloses a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device may read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes Figure 1 the method shown.

[0093] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order mentioned in the operation diagrams. For example, depending on the functions / operations involved, two consecutive blocks shown may actually be executed substantially simultaneously or the above-mentioned blocks can sometimes be executed in the reverse order. In addition, the embodiments presented and described in the flowcharts of the present invention are provided by way of example for the purpose of providing a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated, in which the order of various operations is changed and sub-operations described as part of a larger operation are executed independently.

[0094] In addition, although the present invention has been described in the context of functional modules, it should be understood that one or more of the above functions and / or features may be integrated in a single physical device and / or software module unless otherwise stated to the contrary, or one or more functions and / or features may be implemented in separate physical devices or software modules. It can also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present invention. More precisely, considering the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skills of an engineer. Therefore, those skilled in the art can implement the present invention as set forth in the claims without undue experimentation. It can also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.

[0095] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the above method in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs, Read-Only Memories), random access memories (RAMs, Random Access Memories), magnetic disks, or optical discs.

[0096] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or used in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.

[0097] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part with one or more wirings (electronic device), a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the above program can be printed, because the above program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or processing it in other suitable ways when necessary, and then storing it in a computer memory.

[0098] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0099] In the foregoing description of this specification, descriptions with reference to the terms "one embodiment / example", "another embodiment / example", or "certain embodiments / examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0100] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.

[0101] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A face image enhancement method based on a self-attention network, characterized in that, It includes the following steps: Construct an initial generative adversarial model based on a self-attention network; Obtain a first face image from a face database, and train the initial generative adversarial model according to the first face image to obtain a first generative adversarial model; Obtain a second face image of an actual scene, and train the first generative adversarial model according to the second face image to obtain a second generative adversarial model; Obtain a third face image to be enhanced, and input the third face image into the second generative adversarial model to obtain an enhanced fourth face image; The step of constructing the initial generative adversarial model based on the self-attention network specifically includes: Construct a self-attention network, and the self-attention network includes an input embedding layer, an attention mechanism module, and a feed-forward module; Construct a generator and a discriminator according to the self-attention network. The generator includes a multi-layer perceptron, a plurality of upsampling modules, and a first linear flattening layer. The upsampling module includes the self-attention network and an upsampling layer. The discriminator includes a plurality of the self-attention networks and a second linear flattening layer; Construct the initial generative adversarial model according to the generator and the discriminator; The step of inputting the third face image into the second generative adversarial model to obtain an enhanced fourth face image specifically includes: Input the third face image into the generator of the second generative adversarial model; Perceive the third face image through the multi-layer perceptron to generate self-attention tokens; Encode the self-attention tokens through the self-attention network and then perform upsampling through the upsampling layer to obtain first enhanced data; Fuse the third face image and the first enhanced data through the first linear flattening layer to obtain an enhanced fourth face image.

2. The face image enhancement method based on a self-attention network according to claim 1, characterized in that, The step of obtaining the first face image from the face database and training the initial generative adversarial model according to the first face image to obtain the first generative adversarial model is specifically: Obtain a first face image from the face database, input the first face image into the initial generative adversarial model, and alternately train the generator and the discriminator until the generator and the discriminator reach Nash equilibrium to obtain the first generative adversarial model.

3. The face image enhancement method based on a self-attention network according to claim 1, characterized in that, The step of obtaining the second face image of the actual scene and training the first generative adversarial model according to the second face image to obtain the second generative adversarial model specifically includes: Capture in the actual scene of the face recognition system to obtain the second face image; Input the second face image into the first generative adversarial model, and alternately train the generator and the discriminator until the generator and the discriminator reach Nash equilibrium to obtain the second generative adversarial model.

4. A face image enhancement system based on a self-attention network, characterized in that, It includes: A generative adversarial model construction module for constructing an initial generative adversarial model based on a self-attention network; A first model training module for obtaining a first face image from a face database and training the initial generative adversarial model according to the first face image to obtain a first generative adversarial model; The second model training module is used to obtain the second face image of the actual scene and train the first generative adversarial model according to the second face image to obtain the second generative adversarial model; The image enhancement module is used to obtain the third face image to be enhanced, input the third face image into the second generative adversarial model, and obtain the enhanced fourth face image; The initial generative adversarial model constructed based on the self-attention network specifically includes: Construct a self-attention network, and the self-attention network includes an input embedding layer, an attention mechanism module, and a feed-forward module; Construct a generator and a discriminator according to the self-attention network. The generator includes a multi-layer perceptron, a plurality of upsampling modules, and a first linear flattening layer. The upsampling module includes the self-attention network and an upsampling layer. The discriminator includes a plurality of the self-attention networks and a second linear flattening layer; Construct the initial generative adversarial model according to the generator and the discriminator; Inputting the third face image into the second generative adversarial model to obtain the enhanced fourth face image specifically includes: Input the third face image into the generator of the second generative adversarial model; Perceive the third face image through the multi-layer perceptron to generate self-attention tokens; Encode the self-attention tokens through the self-attention network and then perform upsampling through the upsampling layer to obtain the first enhanced data; Fuse the third face image and the first enhanced data through the first linear flattening layer to obtain the enhanced fourth face image.

5. A face image enhancement device based on a self-attention network, characterized in that, It includes: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements a face image enhancement method based on a self-attention network according to any one of claims 1 to 3.

6. A computer-readable storage medium storing a program executable by a processor, characterized in that, The program executable by the processor, when executed by the processor, is used to execute a face image enhancement method based on a self-attention network according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Face image complementing method based on self-attention deep generative adversarial network

    CN110288537A

  • Generative adversarial mechanism and attention mechanism-based standard face generation method

    WO2020168731A1