An adaptive five - sense fusion method, device and equipment for retaining user characteristics
Through the adaptive five-feature fusion method, the five-feature sub-network and the fusion network are used, combined with the SDFEF module and the GAN network, the problem of the inability to retain the details of the user's facial features in the existing technology is solved, and the natural and high-definition effects of the five-feature fusion are achieved.
Patent Information
- Application Number
- CN202210119316.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-08
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-02-08
AI Technical Summary
The existing five-feature fusion algorithm cannot effectively retain the user's own detailed features of the facial features, resulting in the inadequate naturalness after the fusion process.
Adaptive five-feature fusion method is adopted, and the source face image and model reference image are preprocessed, and feature extraction and fusion are used for feature extraction and fusion, including SDFEF module, wavelet, CBAM, self-attention and soft-mask, combined with GAN network for training, and an adaptive fusion parameter α optimization loss function is introduced.
The natural retention and high-definition fusion of facial features are achieved, which enhances the controllability of the fusion effect, can better control the fusion or retain features, and the generated result map is more natural and clear.
Smart Images

Figure CN114596237B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to an adaptive five - sense organ fusion method, device, and equipment that retain user features. Background Art
[0002] For any given user image and a reference image, the algorithm for fusing the five - sense organs of the reference Figure 5 onto the user's face is called the five - sense organ fusion algorithm. However, most of the existing five - sense organ fusion algorithms are more similar to "face - swapping" algorithms, that is, they make the five - sense organ features of the fused image as similar as possible to those of the reference image, and rarely consider retaining the user's five - sense organ features. This limits the existing algorithms to the scenarios of face - swapping gameplay. However, when users are retouching pictures in daily life, they more hope to incorporate some more beautiful features of the reference image while retaining their own five - sense organ features, so as to achieve the purpose of naturally becoming more beautiful. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to propose an adaptive five - sense organ fusion method, device, and equipment that retain user features, aiming to solve the problem that the existing five - sense organ fusion methods cannot retain the user's own five - sense organ detail features, resulting in an unnatural fusion process.
[0004] To achieve the above purpose, the present invention provides an adaptive five - sense organ fusion method that retains user features, and the method includes:
[0005] Pre - process the source face image and the model reference image to obtain a first face image and a second face image. Among them, the pre - processing includes face alignment and face cropping, and the model reference image is used to fuse the five - sense organ parts in the model reference image onto the reference image of the source face image;
[0006] Input the first face image and the second face image into a pre - trained five - sense organ fusion model to obtain a fused face result image. Among them, the five - sense organ fusion model includes five - sense organ sub - networks corresponding to each five - sense organ part and a fusion network, and the five - sense organ sub - networks include an SDFEF module for feature extraction and feature fusion of the first face image and the second face image.
[0007] Preferably, the step of inputting the first face image and the second face image into a pre - trained five - sense organ fusion model to obtain a fused face result image includes:
[0008] Extract each five - sense organ part of the first face image and the second face image respectively;
[0009] Input at least one pair of facial feature parts into the corresponding facial feature sub-network to obtain a fused part map of the corresponding facial feature parts, where the pair of facial feature parts is the same facial feature of the first face image and the second face image;
[0010] Input each of the fused part maps into the fusion network to obtain a first result map, and paste the first result map back to the source face image to obtain the face result map.
[0011] Preferably, the SDFEF module includes wavelet and CBAM for feature extraction, and self-attention and soft-mask for feature fusion.
[0012] Preferably, the facial feature fusion model is constructed based on the GAN network; the network training process of the facial feature fusion model includes:
[0013] Align and crop the acquired image training data based on face points to obtain face training data, where the face training data includes user images and reference images;
[0014] Use the face points to extract each facial feature part of the user image and the reference image and the facial feature mask corresponding to each facial feature part;
[0015] After splicing the extracted facial feature parts of the user image, the facial feature parts of the reference image, and the facial feature masks of each facial feature part based on the same facial feature part, input them into each corresponding facial feature sub-network for training.
[0016] Preferably, it further includes:
[0017] Use the loss function Loss to perform supervised optimization on the user image and the reference image during the network training process, and balance the weights of the user image and the reference image by introducing an adaptive fusion parameter α in the loss function Loss.
[0018] Preferably, the loss function Loss includes:
[0019] Loss = L gan + L mse + α * (L id-u + L fd-u ) + (1 - α) * (L id-r + L fd-r ),
[0020] where L gan and L mse respectively represent ganloss and mse loss, L id-u 、Lfd-u respectively represent the id loss and face landmark loss of the generated image and the user image, L id-r and L fd-r respectively represent the id loss and face landmark loss of the generated image and the reference image.
[0021] To achieve the above object, the present invention further provides an adaptive five - sense organ fusion device for retaining user features, and the device includes:
[0022] A pre - processing unit, configured to pre - process the source face image and the model reference image to obtain a first face image and a second face image. Wherein, the pre - processing includes face alignment and face cropping, and the model reference image is used to fuse the five - sense organ parts in the model reference image into the reference image of the source face image;
[0023] A five - sense organ fusion unit, configured to input the first face image and the second face image into a pre - trained five - sense organ fusion model to obtain a fused face result image. Wherein, the five - sense organ fusion model includes five - sense organ sub - networks corresponding to each five - sense organ part and a fusion network, and the five - sense organ sub - network includes an SDFEF module for feature extraction and feature fusion of the first face image and the second face image.
[0024] To achieve the above object, the present invention further proposes a device, including a processor, a memory, and a computer program stored in the memory, and the computer program is executed by the processor to implement the steps of an adaptive five - sense organ fusion method for retaining user features as described in the above embodiments.
[0025] To achieve the above object, the present invention further proposes a computer - readable storage medium, on which a computer program is stored, and the computer program is executed by a processor to implement the steps of an adaptive five - sense organ fusion method for retaining user features as described in the above embodiments.
[0026] Advantageous effects:
[0027] In the above solution, the SDFEF module designed in the five - sense organ fusion model can more reasonably extract features and retain the details of the five - sense organs, realizing a more natural fusion of the features of the source face image and the model reference image.
[0028] In the above solution, by splitting the five - sense organs into multiple five - sense organ sub - networks corresponding to each five - sense organ part to generate the fusion results of each region respectively, it is possible to better control the five - sense organ features to be fused or retained, enhancing the controllability of the overall fusion effect. Brief Description of the Drawings
[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0030] Figure 1 Schematic diagram of the result of artifacts in the facial features after the integration of the existing solutions.
[0031] Figure 2 Schematic flow chart of an adaptive facial feature integration method for retaining user characteristics provided by an embodiment of the present invention.
[0032] Figure 3 Schematic overall flow chart provided by an embodiment of the present invention.
[0033] Figure 4 Schematic structural diagram of the facial feature sub-network of the eye part provided by an embodiment of the present invention.
[0034] Figure 5 Schematic structural diagram of the SDFEF module provided by an embodiment of the present invention.
[0035] Figure 6 Schematic diagram of the effect of using wavelet to extract the eye part provided by an embodiment of the present invention.
[0036] Figure 7 Schematic annotation diagram corresponding to the facial feature parts provided by an embodiment of the present invention.
[0037] Figure 8 Schematic structural diagram of an adaptive facial feature integration device for retaining user characteristics provided by an embodiment of the present invention.
[0038] The realization of the invention purpose, functional characteristics and advantages will be further described in conjunction with the embodiments and with reference to the accompanying drawings. Detailed implementation manners
[0039] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0040] In the description of the present invention, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features.
[0041] The content of the present invention will be elaborated in detail below in conjunction with embodiments.
[0042] In addition, some existing five - sense organ fusion algorithms based on deep learning also have other problems: Since the supervision of the features of the reference image is mainly achieved through ID loss, although it generally looks similar to the face in the reference image, it is difficult to specifically learn the specific five - sense organ features (such as double eyelids, nose shape, eyebrow shape, etc.) that the user wants, with low controllability, and the resulting image is prone to problems such as low clarity and unnatural artifacts. Refer to Figure 1 as shown.
[0043] Based on this, the present invention proposes an adaptive five - sense organ fusion method that retains user features, which can retain the details of the five - sense organs and achieve a more natural, higher - clarity, and more realistic result in the generated image after fusion.
[0044] Refer to Figure 2 as shown in the flowchart of an adaptive five - sense organ fusion method that retains user features provided by an embodiment of the present invention.
[0045] In this embodiment, the method includes:
[0046] S11, pre - process the source face image and the model reference image to obtain a first face image and a second face image, where the pre - processing includes face alignment and face cropping, and the model reference image is used to fuse the five - sense organ parts in the model reference image into the reference image of the source face image.
[0047] S12. Input the first face image and the second face image into a pre-trained facial feature fusion model to obtain a fused face result image. The facial feature fusion model includes facial feature sub-networks corresponding to each facial feature part and a fusion network. The facial feature sub-network includes an SDFEF module for feature extraction and feature fusion of the first face image and the second face image.
[0048] Further, the step of inputting the first face image and the second face image into a pre-trained facial feature fusion model to obtain a fused face result image includes:
[0049] S12-1. Extract each facial feature part of the first face image and the second face image respectively;
[0050] S12-2. Input at least one pair of facial feature parts into the corresponding facial feature sub-network to obtain a fused part image of the corresponding facial feature part. The pair of facial feature parts are the same facial feature of the first face image and the second face image;
[0051] S12-3. Input each of the fused part images into the fusion network to obtain a first result image, and paste the first result image back to the source face image to obtain the face result image.
[0052] As Figure 3 shown. In this embodiment, by performing face alignment on the source face image and the model reference image, and cropping the image to the face range, the first face image and the second face image are obtained. In a specific implementation, by using face landmarks to draw each facial feature part and the corresponding mask of the first face image and the second face image. The facial feature parts include eyebrows, eyes, nose, and mouth. The four parts are respectively passed through the corresponding facial feature sub-networks to obtain regional fusion results. For example, taking Eye-Net as an example, as Figure 4 shown is the schematic structural diagram of the facial feature sub-network of the eye part. The structures and processing flows of the facial feature sub-networks of other facial feature parts are the same as it. Further, input the eye parts and masks cropped from the first face image and the second face image into the SDFEF module, extract features and perform feature fusion on the eye parts of the first face image and the second face image respectively, and then pass the fused features through resblocks, upsampling, and convolution to obtain the fused eye image. Finally, fuse the facial feature fusion results of each facial feature sub-network through the fusion network to obtain a complete facial effect, and finally paste it back to the original image to obtain the final result. Further, in the usage stage, the facial feature sub-networks can be selectively used to achieve the effect of controlling the fusion of partial regions of the face. For example, if only the eye part is desired to be migrated, only use the facial feature sub-network of the eye part to obtain the eye image, and then combine it with the user's own eyebrows, nose, and mouth Figure 1Input the fusion network together to get the final effect image.
[0053] The five sense organs fusion model is constructed based on the GAN network; the network training process of the five sense organs fusion model includes:
[0054] Performing face alignment and face cropping on the acquired image training data based on face points to obtain face training data, wherein the face training data includes a user image and a reference image;
[0055] Using the face points, extract each facial feature of the user image and the reference image and a facial feature mask corresponding to each facial feature;
[0056] The extracted facial features of the user image, the facial features of the reference image, and the facial masks of the facial features are spliced based on the same facial features, and then input into each corresponding facial feature sub-network for training.
[0057] The SDFEF module includes wavelet and CBAM for feature extraction, and self-attention and soft-mask for feature fusion.
[0058] In this embodiment, a batch of high-definition face image data is selected as a training set, face alignment is performed according to a unified standard face point, faces at different angles in the image data set are straightened, and the image is cropped to the face range; four masks of eyebrows, eyes, nose, and mouth of the user image and the reference image are drawn through the face points, and the corresponding areas are cut out through the masks. For example, taking the eyes as an example, the user image eyes (3 channels), the reference image eyes (3 channels) and the mask (1 channel) are spliced together to form a 7-channel image input corresponding to the eye facial features sub-network for training.
[0059] This embodiment uses GAN (Generative Adversarial Network) as the basic structure to build the network, and splits the facial features into four corresponding facial features sub-networks for separate generation. In addition, a self-attention facial features detail feature extraction and fusion module (SDFEF, Self-attention Detail Feature Extraction and Fusion) is provided in each facial features sub-network. Figure 5 The schematic diagram of the structure of the SDFEF module is shown in Figure 1. Specifically, in the feature extraction part, the structure combined with wavelet and CBAM is mainly used for feature extraction. The above structure can extract the detailed features of the image through wavelet, and the introduction of CBAM can redistribute the weights of the channels and spatial scales of the features through the attention mechanism. Through the combination of the two, the detailed features of the image can be learned more reasonably and the clarity of the result image can be improved.Figure 6 Schematic diagram of the effect of using wavelet to extract the eye area. Further, in the feature fusion part, self-attention and soft-mask are combined. First, self-attention is used to enhance the attention to the spatial scale, reassign weights to the features, and then through a self-learning soft-mask, the features corresponding to each pixel are individually weighted and fused on the spatial scale. In the general feature fusion method, the features are directly concatenated, which does not consider the spatial relationship of the features and the fusion is very rough. Therefore, compared with the traditional feature splicing scheme, the above structure can further assist the network in distinguishing the features that the reference image and the user image need to pay attention to, and can more naturally fuse the features of the reference image and the user image on the spatial scale.
[0060] Further, the loss function Loss is used to supervise and optimize the user image and the reference image during the network training process, and an adaptive fusion parameter α is introduced in the loss function Loss to balance the weights of the user image and the reference image.
[0061] Among them, the loss function Loss includes:
[0062] Loss = L gan + L mse + α * (L id-u + L fd-u ) + (1 - α) * (L id-r + L fd-r ), where L gan and L mse respectively represent ganloss and mse loss, L id-u , L fd-u respectively represent the idloss and face point loss of the generated image and the user image, L id-r , L fd-r respectively represent the idloss and face point loss of the generated image and the reference image.
[0063] In this embodiment, based on the soft-mask, a fully connected layer is used to learn an adaptive fusion parameter α, and α is used to balance the supervision weights of the reference image and the user image in the training loss. The formula is as follows:
[0064] Loss = L gan + L mse + α * (L id-u + L fd-u ) + (1 - α) * (L id-r + L fd-r )
[0065] Among them, L gan and L mse respectively represent ganloss and mse loss, and L id-u and L fd-u respectively represent the idloss and face point loss between the generated image and the user image. L id-r and L fd-r respectively represent the idloss and face point loss between the generated image and the reference image. The generated image refers to the final full-face generated image obtained after the fusion of the facial feature sub-network and the fusion network. The adaptive fusion parameter α can better guide the network to retain more user facial feature by directly adjusting the loss supervision during training, and can generate a more natural fusion effect.
[0066] In this embodiment, through four independent facial feature sub-networks, the eyebrows, eyes, nose, and mouth are generated respectively, and then the transferred facial features and the user image are fused through a simple fusion network, which can obtain a more delicate facial feature generation effect than the general GAN algorithm. Moreover, most existing algorithms basically perform supervised training on the entire face, and it is difficult to ensure the fusion effect of each facial feature part. Therefore, in this embodiment, a series of features that significantly affect the appearance are defined for different facial feature parts, and these features are integrated into the supervised loss of the facial feature sub-network for further targeted learning. The targeted features of each facial feature sub-network are shown in Table 1 below and Figure 7 the corresponding schematic diagram of the facial feature part. Where f represents the pixel coordinates of each point, and the subscript y represents the y-axis coordinate. Through the supervision of these targeted features, the local feature similarity between the fused result image and the reference image can be significantly improved, and it provides an idea for artificially selecting facial features, which can better control the facial features to be fused or retained, and enhances the controllability of the overall fusion effect.
[0067]
[0068] Table 1
[0069] In summary, in specific applications, facial feature transfer can be achieved for any user image and reference image through this facial feature fusion model, without restrictions such as angles, lighting, and skin colors, and the application scenarios are more extensive. Moreover, the facial feature transfer results are naturally fused and can well adapt to both the case of transferring partial facial feature regions and the case of transferring all facial features. When fusing the reference Figure 5 facial features, the overall aesthetic effect of the user image is ensured. Furthermore, by generating each region separately through the facial feature sub-networks, the high definition and real details of the transferred facial features can be guaranteed, and the facial feature regions to be transferred can be freely selected, supporting the separate transfer and free combination of the eyebrows, eyes, nose, and mouth, with high regional controllability.
[0070] Refer toFigure 8 The figure shows a schematic structural diagram of an adaptive five - sense organ fusion device for retaining user characteristics provided by an embodiment of the present invention.
[0071] In this embodiment, the device 80 includes:
[0072] A pre - processing unit 81, configured to pre - process the source face image and the model reference image to obtain a first face image and a second face image. Among them, the pre - processing includes face alignment and face cropping, and the model reference image is used to fuse the five - sense organ parts in the model reference image into the reference image of the source face image;
[0073] A five - sense organ fusion unit 82, configured to input the first face image and the second face image into a pre - trained five - sense organ fusion model to obtain a fused face result image. Among them, the five - sense organ fusion model includes five - sense organ sub - networks corresponding to each five - sense organ part and a fusion network, and the five - sense organ sub - networks include an SDFEF module for feature extraction and feature fusion of the first face image and the second face image.
[0074] Further, the five - sense organ fusion unit 82 includes:
[0075] A five - sense organ extraction unit, configured to extract each five - sense organ part of the first face image and the second face image respectively;
[0076] A first processing unit, configured to input at least one pair of five - sense organ parts into the corresponding five - sense organ sub - network to obtain a fused part image of the corresponding five - sense organ part, where the pair of five - sense organ parts are the same five - sense organ of the first face image and the second face image;
[0077] A second processing unit, configured to input each of the fused part images into the fusion network to obtain a first result image, and paste the first result image back to the source face image to obtain the face result image.
[0078] Further, the SDFEF module includes a wavelet and a CBAM for the feature extraction, and a self - attention and a soft - mask for the feature fusion.
[0079] Further, the five - sense organ fusion model is constructed based on a GAN network; the network training process of the five - sense organ fusion model includes:
[0080] Aligning and cropping the acquired image training data based on face points to obtain face training data, where the face training data includes user images and reference images;
[0081] Extract each facial feature part of the user image and the reference image and the facial feature mask corresponding to each facial feature part by using the facial points respectively.
[0082] After splicing the extracted facial feature parts of the user image, the facial feature parts of the reference image, and the facial feature masks of each facial feature part based on the same facial feature part respectively, input them into each corresponding facial feature sub-network for training.
[0083] Furthermore, it further includes:
[0084] Use the loss function Loss to supervise and optimize the user image and the reference image during network training, and balance the weights of the user image and the reference image by introducing the adaptive fusion parameter α in the loss function Loss.
[0085] Furthermore, the loss function Loss includes:
[0086] Loss = L gan + L mse + α * (L id-u + L fd-u ) + (1 - α) * (L id-r + L fd-r ),
[0087] wherein, L gan and L mse respectively represent ganloss and mse loss, L id-u , L fd-u respectively represent the idloss and face point loss of the generated image and the user image, L id-r , L fd-r respectively represent the idloss and face point loss of the generated image and the reference image.
[0088] Each unit module of the device 80 can respectively execute the corresponding steps in the above method embodiments, so the unit modules will not be elaborated here. For details, please refer to the descriptions of the above corresponding steps.
[0089] The embodiment of the present invention also provides a device, which includes the adaptive facial feature fusion device for retaining user features as described above. Among them, the adaptive facial feature fusion device for retaining user features can adopt Figure 8 the structure of the embodiment, and correspondingly, it can execute Figure 2 the technical solutions of the method embodiments shown, and its implementation principle and technical effects are similar. For details, please refer to the relevant records in the above embodiments, which will not be elaborated here.
[0090] The device includes: devices with a photographing function such as mobile phones, digital cameras or tablet computers, or devices with an image processing function, or devices with an image display function. The device may include components such as a memory, a processor, an input unit, a display unit, and a power supply.
[0091] Among them, the memory can be used to store software programs and modules. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as an image playback function, etc.); the data storage area can store data created according to the use of the device, etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory can also include a memory controller to provide access to the memory by the processor and the input unit.
[0092] The input unit can be used to receive input digital or character or image information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls. Specifically, the input unit of this embodiment, in addition to including a camera, may also include a touch-sensitive surface (such as a touch display screen) and other input devices.
[0093] The display unit can be used to display information input by the user or information provided to the user and various graphical user interfaces of the device. These graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof. The display unit may include a display panel. Optionally, the display panel can be configured in the form of an LCD (Liquid Crystal Display), an OLED (Organic Light-Emitting Diode), etc. Further, the touch-sensitive surface can cover the display panel. When the touch-sensitive surface detects a touch operation on or near it, it is transmitted to the processor to determine the type of touch event. Subsequently, the processor provides a corresponding visual output on the display panel according to the type of touch event.
[0094] The embodiment of the present invention also provides a computer-readable storage medium. The computer-readable storage medium can be the computer-readable storage medium included in the memory in the above embodiment; it can also exist separately and be a computer-readable storage medium not assembled into the device. At least one instruction is stored in the computer-readable storage medium, and the instruction is loaded and executed by the processor to implement Figure 2 the adaptive facial feature fusion method for retaining user features as shown. The computer-readable storage medium can be a read-only memory, a magnetic disk, an optical disc, etc.
[0095] It should be noted that the various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the device embodiments, equipment embodiments and storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the corresponding descriptions in the method embodiments.
[0096] Also, in this text, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0097] The above description shows and describes the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications and environments, and can be changed within the scope of the inventive concept herein through the above teachings or the techniques or knowledge in the relevant field. And any changes and modifications made by those skilled in the art without departing from the spirit and scope of the present invention shall fall within the protection scope of the appended claims of the present invention.
Claims
1. An adaptive five - sense organ fusion method for retaining user characteristics, characterized in that, The method includes: Preprocessing the source face image and the model reference image to obtain a first face image and a second face image. Among them, the preprocessing includes face alignment and face cropping, and the model reference image is used to fuse the facial feature parts in the model reference image into the reference image of the source face image; Inputting the first face image and the second face image into a pre-trained facial feature fusion model to obtain a fused face result image. Among them, the facial feature fusion model includes facial feature sub-networks corresponding to each facial feature part and a fusion network, and the facial feature sub-network includes an SDFEF module for feature extraction and feature fusion of the first face image and the second face image; Further, the SDFEF module includes a wavelet and a CBAM for the feature extraction, and a self-attention and a soft-mask for the feature fusion. Among them, the wavelet is used to extract the detailed features of the image, the CBAM is used to reassign the weights of the channels and spatial scales of the features, the self-attention is used to enhance the attention to the spatial scale and reassign the weights of the features, and the soft-mask is used to perform weighted fusion on the features corresponding to each pixel at the spatial scale separately.
2. The adaptive five - sense fusion method for retaining user characteristics according to claim 1, wherein, The step of inputting the first face image and the second face image into a pre-trained facial feature fusion model to obtain a fused face result image includes: Extracting each facial feature part of the first face image and the second face image respectively; Inputting at least one pair of facial feature parts into the corresponding facial feature sub-network to obtain a fused part image of the corresponding facial feature part, where the pair of facial feature parts are the same facial feature of the first face image and the second face image; Inputting each of the fused part images into the fusion network to obtain a first result image, and pasting the first result image back to the source face image to obtain the face result image.
3. An adaptive five - sense fusion method for retaining user characteristics according to claim 1, characterized in that, The facial feature fusion model is constructed based on a GAN network; The network training process of the facial feature fusion model includes: Aligning and cropping the acquired image training data based on face points to obtain face training data, where the face training data includes user images and reference images; Using the face points to extract each facial feature part of the user image and the reference image and the facial feature mask corresponding to each facial feature part respectively; After splicing the extracted facial feature parts of the user image, the facial feature parts of the reference image, and the facial feature masks of each facial feature part based on the same facial feature part, inputting them into the corresponding facial feature sub-network for training respectively.
4. An adaptive five - sense fusion method for retaining user characteristics according to claim 3, characterized in that, It further includes: Using a loss function Loss to perform supervised optimization on the user image and the reference image during the network training process, and introducing an adaptive fusion parameter α in the loss function Loss to balance the weights of the user image and the reference image.
5. An adaptive five - sense fusion method for retaining user characteristics according to claim 4, characterized in that, The loss function Loss includes: Loss=L gan +L mse +α*(L id-u +L fd-u )+(1-α)*(L id-r +L fd-r ), Among them, L gan and L mse respectively represent the gan loss and the mse loss. L id-u and L fd-u respectively represent the id loss and the face point loss between the generated image and the user image. L id-r and L fd-r respectively represent the id loss and the face point loss between the generated image and the reference image.
6. An adaptive five - sense organ fusion device for retaining user characteristics, characterized in that, The device includes: A preprocessing unit for preprocessing the source face image and the model reference image to obtain a first face image and a second face image, wherein the preprocessing includes face alignment and face cropping, and the model reference image is used to fuse the facial feature parts in the model reference image into the reference image of the source face image; A facial feature fusion unit for inputting the first face image and the second face image into a pre-trained facial feature fusion model to obtain a fused face result image, wherein the facial feature fusion model includes facial feature sub-networks corresponding to each facial feature part and a fusion network, and the facial feature sub-networks include an SDFEF module for feature extraction and feature fusion of the first face image and the second face image; Further, the SDFEF module includes a wavelet and a CBAM for the feature extraction, and a self-attention and a soft-mask for the feature fusion; wherein the wavelet is used to extract the detailed features of the image, the CBAM is used to reassign weights to the channels and spatial scales of the features, the self-attention is used to enhance the attention to the spatial scales, reassign weights to the features, and the soft-mask is used to perform weighted fusion of the features corresponding to each pixel separately on the spatial scale.
7. A device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory, and the computer program is executed by the processor to implement the steps of an adaptive facial feature fusion method for retaining user features as described in any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and the computer program is executed by a processor to implement the steps of an adaptive facial feature fusion method for retaining user features as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Real-time skin makeup migration method and device, electronic equipment and readable storage medium
CN111815534A
Image fusion method and device, equipment and storage medium
CN112967261A
Face fusion method and device, equipment and storage medium
CN113486944A