An image processing method, apparatus, electronic device, and storage medium

By presenting the image transformation function item in the view interface, obtaining and generating the target image of the target image template, and using the image processing model to adjust the color and margin mode of the texture feature set, the problems of facial deformation and hardware resource occupation in the image face-changing operation are solved, and efficient image processing and user experience improvement are achieved.

CN112750176BActive Publication Date: 2025-10-17TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010946434.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-10
Publication Date
2025-10-17
Estimated Expiration
2040-09-10

AI Technical Summary

Technical Problem

In Internet social applications, image face-swapping operations can easily cause significant deformation of the target user's face. There is a large gap between the newly synthesized face and the face to be replaced, which reduces the image's recognition. In addition, existing technologies fail to effectively solve the problem of hardware resource usage.

Method used

By presenting the image transformation function item in the view interface, the image to be processed is obtained and the target image of the target image template is generated. The texture feature set of the object is obtained by using the image processing model for overlay rendering. The color and margin mode are adjusted to match the image to be processed and the target image template to achieve the replacement of object parts.

Benefits of technology

It retains the characteristics of the image to be processed, avoids the distortion of the target image, realizes batch processing of different images to be processed by users, reduces the occupation of hardware resources, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112750176B_ABST
    Figure CN112750176B_ABST
Patent Text Reader

Abstract

The application provides an image processing method, comprising: presenting an image transformation function item in a view interface, in response to a triggering operation on the image transformation function item, acquiring and presenting a to-be-processed image containing a first object; in response to a transformation determination operation triggered based on the to-be-processed image, generating and presenting a target image of a target image template. The application also provides an image processing device, an electronic device and a storage medium. The application can retain the characteristics of the to-be-processed image, avoid distortion of the replaced target image, realize batch processing of different to-be-processed images of the user, reduce the occupation of hardware resources in the image processing process, reduce the increase of the cost of hardware devices, and improve the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the information processing technology field, and particularly relates to an image processing method and device, electronic equipment and a storage medium. BACKGROUND

[0002] In the Internet social application, the face changing operation of the image can increase the interesting of the social application, and the target part needs to be replaced to the corresponding part of the other object in the target image template while keeping the situation of the target part of the person in the original image (such as a picture or a video frame). For this purpose, the artificial intelligence technology provides a scheme of training a proper image processing model to support the above-mentioned application, but due to the various color changes of the target image, it is easy to cause a large deformation of the target user face (the face to be replaced), and there is a large gap between the new combined face and the face to be replaced, which reduces the recognition of the image. SUMMARY

[0003] Therefore, the embodiments of the present application provide an image processing method and device, electronic equipment and a storage medium, which can retain the characteristics of the image to be processed, avoid distortion of the target image after replacement, realize batch processing of different images to be processed by the user, and reduce the occupation of hardware resources in the image processing process, reduce the increase of hardware device cost, and improve the user experience.

[0004] The technical scheme of the embodiments of the present application is as follows:

[0005] The embodiments of the present application provide an image processing method, comprising:

[0006] presenting an image transformation function item in a view interface;

[0007] in response to a trigger operation on the image transformation function item, obtaining and presenting an image to be processed containing a first object;

[0008] in response to a transformation determination operation triggered based on the image to be processed, generating and presenting a target image of a target image template, wherein the target image of the target image template is an image generated by replacing a corresponding part of a second object in the target image template with a target part of the first object in the image to be processed.

[0009] The embodiments of the present application also provide an image processing device, comprising:

[0010] an information display module configured to present an image transformation function item in a view interface;

[0011] an information transmission module configured to obtain and present an image to be processed containing a first object in response to a trigger operation on the image transformation function item;

[0012] The information display module is used to generate and present a target image of the target image template in response to a transformation determination operation triggered based on the image to be processed; wherein, the target image of the target image template is an image generated by replacing the corresponding part of the second object in the target image template with the target part of the first object in the image to be processed.

[0013] In the above solution, the information display module is used to mark the image template indicated by the template switching operation in the image template switching interface.

[0014] In the above solution, the information display module is used to present an image sharing function item for sharing the target image;

[0015] The information transmission module is configured to share the target image in response to a triggering operation on the image sharing function item.

[0016] In the above solution, the information display module is used to trigger the image processing model and obtain the first texture feature set corresponding to the target part of the first object in the image to be processed through the image processing model;

[0017] The information display module is used to obtain a second texture feature set that matches a corresponding part of a second object in the target image template;

[0018] The information display module is used to overlay and render the first texture feature set and the second texture feature set. In the target image template after overlay rendering, the first texture feature set covers the second texture feature set on the corresponding part of the second object.

[0019] In the above solution, the information display module is configured to, when the target part is a face, adjust the color mode of the superimposed rendering of the first texture feature set and the second texture feature set so that the face of the first object in the image to be processed matches the color of the second object in the target image template;

[0020] The information display module is used to adjust the margin mode of the superimposed rendering of the first texture feature set and the second texture feature set when the target part is the hairstyle part, so that the hairstyle part of the first object in the image to be processed is adapted to the facial features of the second object in the target image template.

[0021] In the above solution, the information display module is configured to, in response to a viewing operation on the image transformation function item, present a content page including the image to be processed and the image template, and present at least one interactive function item on the content page, the interactive function item being configured to enable interaction with the image to be processed;

[0022] The information transmission module is configured to receive an interactive operation triggered based on the interactive function item for the image to be processed, so as to execute a corresponding interactive instruction.

[0023] In the above solution, the information display module is configured to present first interactive prompt information in the content page, where the first interactive prompt information is configured to prompt that interactive content corresponding to the interactive operation can be presented in the view interface.

[0024] The information display module is configured to switch page content to the view interface in response to an operation of switching to the view interface.

[0025] In the above solution, the information display module is configured to present second interactive prompt information in the content page, where the second interactive prompt information is configured to prompt that interactive content corresponding to the interactive operation can be presented in a target image template library interface corresponding to a target image template.

[0026] The content page is switched to the target image template library interface in response to an instruction of switching to the target image template library interface.

[0027] Embodiments of the present application also provide an electronic device, which comprises:

[0028] a memory configured to store executable instructions;

[0029] a processor configured to execute the executable instructions stored in the memory, so as to implement the image processing method described in the foregoing.

[0030] Embodiments of the present application also provide a computer readable storage medium storing executable instructions, which are executed by a processor to implement the image processing method described in the foregoing.

[0031] Embodiments of the present application present an image transformation function item in a view interface, and in response to a triggering operation on the image transformation function item, an image to be processed containing a first object is acquired and presented; in response to a transformation determination operation triggered based on the image to be processed, a target image of a target image template is generated and presented, where the target image of the target image template is an image generated by replacing a corresponding part of a second object in the target image template with a target part of the first object in the image to be processed; thus, the features of the image to be processed can be retained, the replaced target image can be prevented from being distorted, batch processing of different images to be processed by a user can be implemented, the occupation of hardware resources in the image processing process is reduced, the increase of hardware device cost is reduced, and the use experience of the user is improved. BRIEF DESCRIPTION OF DRAWINGS

[0032] Figure 1A use scenario diagram of the image processing method provided by the embodiment of the present application is shown in FIG. 1.

[0033] Figure 2 A composition structure diagram of the electronic device provided by the embodiment of the present application is shown in FIG. 2.

[0034] Figure 3 A processing diagram for generating a face changing image in a conventional scheme is shown in FIG. 3.

[0035] Figure 4 A principle diagram for face changing by an image processing model in the related technology of the present application is shown in FIG. 4.

[0036] Figure 5 A face changing effect diagram of the related technology in the embodiment of the present application is shown in FIG. 5.

[0037] Figure 6 An optional flow diagram of the image processing method provided by the embodiment of the present application is shown in FIG. 6.

[0038] Figure 7A A diagram of the image processing interface provided by the embodiment of the present application is shown in FIG. 7.

[0039] Figure 7B A diagram of the image processing interface provided by the embodiment of the present application is shown in FIG. 8.

[0040] Figure 8 A diagram of the image processing model in the embodiment of the present application is shown in FIG. 9.

[0041] Figure 9 A diagram of the image processing model in the embodiment of the present application is shown in FIG. 10.

[0042] Figure 10 A diagram of the image processing model in the embodiment of the present application is shown in FIG. 11.

[0043] Figure 11 A diagram of the image processing model in the embodiment of the present application is shown in FIG. 12.

[0044] Figure 12 A diagram of the image processing model in the embodiment of the present application is shown in FIG. 13.

[0045] Figure 13 A diagram of the image processing model in the embodiment of the present application is shown in FIG. 14.

[0046] Figure 14 A diagram of the image processing in the embodiment of the present application is shown in FIG. 15.

[0047] Figure 15 A diagram of the image processing interface provided by the embodiment of the present application is shown in FIG. 16.

[0048] Figure 16 The schematic diagram of the image processing interface provided by the embodiment of the present application. DETAILED DESCRIPTION

[0049] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be described in further detail below with reference to the drawings, and the described embodiments should not be regarded as limiting the present application, and all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0050] In the following description, “some embodiments” are related to a subset of all possible embodiments, but it can be understood that “some embodiments” can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0051] Before the embodiments of the present application are described in further detail, the terms and phrases related to the embodiments of the present application are explained, and the terms and phrases related to the embodiments of the present application are applicable to the following explanations.

[0052] 1) In response to a condition or state indicating that the operation performed depends on, when the dependent condition or state is met, one or more operations performed can be real-time or have a set delay; in the absence of special instructions, there is no restriction on the execution order of multiple operations performed.

[0053] 2) Down-sampling processing, taking a sample once every few samples in a sample sequence, so that the new sequence is the down-sampling of the original sequence, for example: for an image I with size M*N, s times down-sampling is performed, that is, a resolution image with size (M / s)*(N / s) is obtained, where s should be the greatest common divisor of M and N

[0054] 3) Convolutional Neural Networks (CNN) is a class of feedforward neural networks containing convolutional computation and having a deep structure, which is one of the representative algorithms of deep learning. Convolutional Neural Networks have representation learning ability and can perform shift-invariant classification on input information according to its hierarchical structure.

[0055] 4) Client, a carrier for implementing specific functions in a terminal, for example, a mobile client (APP) is a carrier for implementing specific functions in a mobile terminal, for example, a program for performing user gesture recognition.

[0056] 5) Component, is the function module of the view of the applet, also called front-end component, buttons, titles, tables, sidebars, contents and footers in the page, etc., the component includes modularized code to facilitate reuse in different pages of the applet.

[0057] 6) Mini Program, is a program developed based on a front-end-oriented language (such as JavaScript) to implement services in a Hyper Text Markup Language (HTML) page, downloaded by a client (such as a browser or any client with an embedded browser core) via a network (such as the Internet), and interpreted and executed in the browser environment of the client, saving the step of installation in the client. For example, through voice instructions to wake up the applet in the terminal to realize various services such as image editing and face changing in a social network client.

[0058] 7) RGB, a three-primary color encoding method, also known as RGB color mode, is an industry standard for color, which obtains various colors through the changes of red (R), green (G), and blue (B) three color channels and their mutual superposition. RGB represents the three channels of red, green, and blue. This standard includes almost all colors that can be perceived by human vision, and is one of the most widely used color systems.

[0059] 8) Face changing, replacing the corresponding part in the different objects of other images with the target part of the object in the image to be processed, which is referred to as the corresponding part.

[0060] Figure 1 The use scenario diagram of the image processing method provided by the embodiments of the present application is shown in Figure 1 The terminal (including terminal 10-1 and terminal 10-2) is provided with a client or applet of image processing software, and to realize a supporting example application, the image processing device 30 of the embodiments of the present application can be a server, and the terminal running various clients can be used to display the image processing result of the image processing device, and the two are connected through a network 300, wherein the network 300 can be a wide area network or a local area network, or a combination of the two, and data transmission is realized using a wireless link. The terminal 10 submits the image to be processed, and the image processing device 30 responds to the triggering operation of the image transformation function item to realize image processing, and the terminal 10 acquires and presents the image to be processed containing the first object and presents the target image corresponding to the target image template.

[0061] In some embodiments of the present application, a video client can be run in the terminal 10, which can submit a corresponding image processing request to the server according to the to-be-replaced face 120 and the target face 110 indicated by the user in the playing interface through various human-computer interaction modes (such as gestures, voice, etc.), so as to realize the image processing method provided by the present application when the executable instructions in the storage medium of the server are executed by the processor, and achieve the corresponding face replacement effect. Further, the above image processing process can also be migrated to the server, and the different frame images after replacement are re-encoded by means of the hardware resources of the server to form a video with the face replacement effect, so as to be called by the WeChat applet or shared to different application processes of the user terminal 10.

[0062] Among them, the image processing method provided by the embodiments of the present application can be realized based on artificial intelligence. Artificial intelligence (AI) is the theory, method, technology and application system for using digital computers or machine controlled by digital computers to simulate, extend and expand human intelligence, perceive environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which tries to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.

[0063] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software level technology. Artificial intelligence basic technology generally includes technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics, etc. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, etc.

[0064] In the embodiments of the present application, the artificial intelligence software technologies mainly involved include the above-mentioned speech processing technologies and machine learning and the like. For example, the speech recognition technology (Automatic Speech Recognition, ASR) in the speech technology can be involved, which includes speech signal preprocessing, speech signal frequency analyzing, speech signal feature extraction, speech signal feature matching / recognition, speech training and the like.

[0065] For example, machine learning (Machine learning, ML) can be involved, which is a multi-field interdisciplinary subject involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and the like. It is specially studied how the computer simulates or realizes the learning behavior of human beings to obtain new knowledge or skills, reorganizes the existing knowledge structure to constantly improve the performance. Machine learning is the core of artificial intelligence and is the fundamental approach to making computers intelligent, and its application is widespread in various fields of artificial intelligence. Machine learning usually includes technologies such as deep learning, and deep learning includes artificial neural networks such as convolutional neural networks (Convolutional Neural Network, CNN), recurrent neural networks (Recurrent Neural Network, RNN), deep neural networks (Deep neural network, DNN) and the like.

[0066] It can be understood that the image processing method and the speech processing provided by the present application can be applied to an intelligent device. The intelligent device can be any device with a speech instruction recognition function, for example, it can be an intelligent terminal, an intelligent home device (such as an intelligent sound box, an intelligent washing machine and the like), an intelligent wearable device (such as an intelligent watch), a vehicle-mounted intelligent central control system (which wakes up the small programs in the terminal for executing different tasks through speech instructions) or an AI intelligent medical device (which is awakened and triggered through speech instructions) and the like.

[0067] The structure of the electronic device of the embodiments of the present application will be described in detail below. The electronic device can be implemented in various forms, such as a special terminal with a speech recognition model training function, or a server provided with a speech recognition model training function, for example, the server 200 in the foregoing Figure 1 . Figure 2 The schematic diagram of the constituent structure of the electronic device provided by the embodiments of the present application can be understood as follows: Figure 2 Only exemplary structures of the electronic device are shown, not all structures. The partial structures or all structures shown can be implemented as needed. Figure 2

[0068] The electronic device provided by the embodiments of the present application includes at least one processor 201, a memory 202, a user interface 203, and at least one network interface 204. The various components in the electronic device are coupled together through a bus system 205. It can be understood that the bus system 205 is used to realize the connection and communication between the components. In addition to the data bus, the bus system 205 also includes a power bus, a control bus, and a status signal bus. However, for the purpose of clear illustration, all kinds of buses are marked as the bus system 205 in the Figure 2 .

[0069] The user interface 203 can include a display, a keyboard, a mouse, a trackball, a click wheel, a key, a button, a touchpad, or a touch screen, etc.

[0070] It can be understood that the memory 202 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. The memory 202 in the embodiments of the present application can store data to support the operation of the terminal (such as 10-1). Examples of these data include any computer programs for operating on the terminal (such as 10-1), such as an operating system and an application program. The operating system contains various system programs, such as a framework layer, a core library layer, a driver layer, etc., for realizing various basic services and processing hardware-based tasks. The application program can include various application programs.

[0071] ​In some embodiments, the electronic device provided by the embodiments of the present application can be implemented in a combination of software and hardware. For example, the electronic device provided by the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to perform the image processing method provided by the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can use one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), or other electronic elements.

[0072] As an example of the electronic device provided by the embodiments of the present application being implemented in a combination of software and hardware, the electronic device provided by the embodiments of the present application can be directly embodied as a combination of software modules executed by the processor 201. The software modules can be located in the storage medium, and the storage medium is located in the memory 202. The processor 201 reads executable instructions included in the software modules in the memory 202, and in combination with the necessary hardware (for example, including the processor 201 and other components connected to the bus 205), the image processing method provided by the embodiments of the present application is completed.

[0073] As an example, the processor 201 can be an integrated circuit chip that has the processing capability of a signal, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0074] As an example of the electronic device provided in the embodiments of the present application implemented by hardware, the apparatus provided in the embodiments of the present application can be directly implemented by a hardware coding processor form of the processor 201, for example, by one or more application specific integrated circuits (ASIC), DSP, programmable logic devices (PLD), complex programmable logic devices (CPLD), field programmable gate arrays (FPGA) or other electronic elements to implement the image processing method provided in the embodiments of the present application.

[0075] The memory 202 in the embodiments of the present application is used to store various types of data to support the operation of the electronic device. Examples of these data include any executable instructions for operating on the electronic device, such as executable instructions, programs implementing the image processing method provided in the embodiments of the present application can be included in the executable instructions.

[0076] In other embodiments, the electronic device provided in the embodiments of the present application can be implemented in software, Figure 2 The electronic device stored in the memory 202 is shown, which can be software in the form of programs and plug-ins, and includes a series of modules. As an example of the program stored in the memory 202, the electronic device can include the following software modules in the image processing module of the electronic device: information display module 2081, information transmission module 2082. When the software modules in the electronic device are read into the RAM by the processor 201 and executed, the image processing method provided in the embodiments of the present application will be implemented, and the functions of each software module in the electronic device in the embodiments of the present application will be introduced below.

[0077] The information display module 2081 is configured to present an image transformation function item in a view interface, and the image transformation function item is configured to implement replacement of a corresponding part of a second object in a target image template for a target part of a first object in a to-be-processed image.

[0078] The information transmission module 2082 is configured to acquire and present a to-be-processed image containing the first object in response to a triggering operation on the image transformation function item.

[0079] The information display module 2081 is configured to generate and present a target image of a target image template in response to a transformation determination operation triggered based on the to-be-processed image; and wherein the target image of the target image template is an image generated by replacing a corresponding part of a second object in the target image template with a target part of a first object in the to-be-processed image.

[0080] According to Figure 2 As shown in the electronic device, in an aspect of the present application, the present application also provides a computer program product or computer program, which comprises computer instructions stored in a computer readable storage medium. The processor of the computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions, so that the computer device executes different embodiments and combinations of embodiments provided in various optional implementations of the image processing method.

[0081] In combination Figure 2 The image processing apparatus shown illustrates the image processing method provided by the embodiments of the present application. Before introducing the image processing method provided by the embodiments of the present application, first introduce the process of generating a corresponding image processing result by the image processing model in the present application, Figure 3 The processing schematic diagram for generating a face changing image in the conventional scheme, the target part of the first object in the to-be-processed image is replaced with the corresponding part of the second object in the target image template, wherein the image processing model comprises an encoder and a decoder. The decoder is one-to-one corresponding to the single target face used for replacing the single to-be-replaced face. Among them, the single to-be-replaced face can be understood as: the to-be-replaced face A in the original image set, the target face B, replacing the to-be-replaced face in the original image set with the target face having the same style characteristics as the to-be-replaced face, the to-be-replaced face A has the style characteristics of the target face B. Therefore, the number of decoders in the single image processing model depends on the number of different single target faces (such as different faces) that the single image processing model needs to process. For example, when the single image processing model needs to replace a single to-be-replaced face in a video with 2 different target faces, the single image processing model needs to set decoders corresponding to the 2 different target faces.

[0082] Figure 4This is a schematic diagram of the principle of face-swapping using an image processing model in the related art of this application. After the encoder and decoder are trained, the encoder extracts style features from the face to be replaced in the original image set (that is, encodes the face to be replaced). The style features are input into the decoder for decoding. This decoding process is a face conversion, forming a new face-swapping image that includes the facial features of the target face and the style of the face to be replaced, such as expression and demeanor. The above-mentioned object can be any creature with facial features (including humans and animals). Taking the human face as an example, the image processing method provided in the embodiment of this application will be further described.

[0083] Furthermore, Figure 3 and Figure 4 The "face swap" process shown here uses machine learning to locate facial key points and transfer the user's facial expressions to the target face. This allows the generated image to have both the user's and the specific image's appearance characteristics. Specifically, after the encoder and decoder are trained, the encoder extracts style features from the face to be replaced in the original image set (that is, encodes the face to be replaced). The style features are then input into the decoder for decoding. This decoding process is a facial transformation, resulting in a new face-swapped image that includes the facial features of the target face and the style of the face to be replaced, such as expression and demeanor.

[0084] But reference Figure 5 , Figure 5 This is a schematic diagram of the face-swapping effect of the related technology in the embodiment of this application. Specifically, the drawback of the related technology is that the face-swapping scheme of expression transfer easily causes a large deformation of the target user's face (the face to be replaced), and the newly synthesized face looks very different from the face to be replaced. In addition, the newly generated face does not retain the user's original hairstyle, which reduces recognition. Figure 5 The face of the boy on the left side of the middle is covered with the hairstyle of the girl in the middle, which makes the image on the right side distorted after the face change, affecting the user experience.

[0085] In order to solve such problems, an optional flow chart of an image processing method provided in an embodiment of the present application is as follows: Figure 6 As shown, this method can be applied to the field of terminal image processing, and the target parts between two objects in different images can be replaced through an image processing process or an image processing applet, wherein, Figure 6 The steps shown can be performed by various electronic devices running image processing devices, such as dedicated terminals, servers or server clusters with image processing functions. Figure 6 The steps shown are explained.

[0086] Step 601: The image processing apparatus presents an image transformation function item in a viewing interface.

[0087] The image transformation function item is configured to replace a target part of a first object in the image to be processed with a corresponding part of a second object in a target image template.

[0088] In some embodiments of the present application, the terminal is provided with an application client, such as a video application client, a news application client, an instant messaging client, etc. The terminal can present a new image processed by face replacement through the application client. The face replacement between the image to be processed and the target image in the image template can be performed by a small program in the instant messaging client. Correspondingly, the target image corresponding to the target image template generated by the small program in the instant messaging client can also be saved in the storage medium of the terminal.

[0089] Step 602: The image processing device acquires and presents an image to be processed containing a first object in response to a trigger operation for the image transformation function item.

[0090] In some embodiments of the present application, before the target image corresponding to the target image template is presented, the following operations can also be performed:

[0091] An image template selection interface is presented, and at least one image template is presented in the image template selection interface. In response to a selection operation of an image template triggered based on the image template selection interface, the image template corresponding to the selection operation is taken as the target image template. Wherein, for reference Figure 7A , Figure 7A The schematic diagram of the image processing interface provided by the embodiments of the present application is shown in the figure. When the user performs face replacement on the image to be processed through the image processing function APP, the image processing APP can provide different image templates to the user in the image template selection interface. These image templates have different display styles. In some embodiments, the selection operation of the image template can be triggered by a trigger operation on the image template selection interface, such as obtaining a mouse click operation or a screen touch operation of the user on the played video content. In some examples, the selection operation of the image template can also be triggered by a selection shortcut key set on the terminal, such as obtaining a click operation of the user on the Print Screen key of the keyboard; or by obtaining a click operation of the user on the selection shortcut key of the image template on a mobile intelligent device such as a mobile phone or a tablet computer.

[0092] Step 603: The image processing device generates and presents a target image of a target image template in response to a transformation determination operation triggered based on the image to be processed.

[0093] The target image of the target image template is an image generated by replacing a corresponding part of a second object in the target image template with a target part of a first object in the image to be processed.

[0094] Specifically, in some embodiments of the present application, reference is made to Figure 7B , Figure 7B The schematic diagram of the image processing interface provided by the embodiments of the present application is shown in FIG. 1, in which the user can input different images to be replaced and select corresponding target image templates through the terminal. The number of images to be replaced and the number of target image templates can be adjusted through the user terminal, for example, each image to be processed corresponds to a target image template, or all images to be processed correspond to the same target image template, M*N groups of images to be replaced and target image templates are processed, wherein the values of M and N can be adjusted according to different use requirements. As shown in FIG. 1, the target part of the first object A / B / C in the three images to be processed input by the user can be replaced with the corresponding part of the second object A1 / B1 / C1 in the three target image templates respectively through the image processing method provided by the present application, and the image generated by replacing the corresponding part of the second object A1 / B1 / C1 in the target image template is generated and presented. Figure 7B

[0095] In some embodiments of the present application, the target image can also be shared in response to the triggering operation of the image sharing function item. The image sharing function item can be associated with a default sharing path, for example, a function item for sharing to an instant messaging software or a social software, or a sharing interface containing at least two sharing path selection items can be presented in response to the triggering of the image sharing function item, and the target image can be shared to different social application processes or image interception application processes through the selected sharing path in response to the sharing path selection operation triggered based on the sharing interface, so as to share or intercept the generated new image.

[0096] In some embodiments of the present application, the target image corresponding to the target image template is generated and presented, including:

[0097] ​trigger an image processing model, and obtain, by the image processing model, a first texture feature set corresponding to a target part of a first object in the image to be processed; obtain a second texture feature set matched with a corresponding part of a second object in the target image template; and perform superimposed rendering on the first texture feature set and the second texture feature set, so that the first texture feature set covers the second texture feature set on the corresponding part of the second object in the target image template after superimposed rendering. In the embodiments, the image processing model can use various existing convolutional neural network structures (for example, DenseBox, VGGNet, ResNet, SegNet, etc.). In practice, a convolutional neural network (CNN) is a kind of feedforward neural network, and the artificial neurons thereof can respond to surrounding units in a certain coverage range, and have excellent performance for image processing, so that the convolutional neural network can be used to process images. The convolutional neural network can include a convolutional layer, a pooling layer, a feature fusion layer, a fully connected layer, etc. The convolutional layer can be used to extract image features. The pooling layer can be used to downsample the input information. The feature fusion layer can be used to fuse the image features (for example, in the form of a feature vector) corresponding to each frame. For example, the feature values at the same positions in the feature vectors or feature matrices corresponding to different images can be averaged to perform feature fusion, to generate a fused feature vector or feature matrix. The fully connected layer can be used to further process the obtained features, and the vector output after processing is a feature vector corresponding to the image to be processed.

[0098] wherein the reference Figure 8 and Figure 9 , Figure 8 is a schematic diagram of a neural network processing an image in the image processing model in the embodiments of the present application, Figure 9 is a schematic diagram of a neural network reducing an image level in the image processing model in the embodiments of the present application. In some embodiments of the present application, the neural network of the image processing model can be a convolutional neural network, and the face and hairstyle regions in the body shown in the image to be processed or the target image are calibrated by a large data training sample.

[0099] Generally, the convolutional neural network is composed of three parts: the convolutional layer is responsible for extracting local features in the image; the pooling layer is used to greatly reduce the parameter magnitude (to realize dimension reduction); and the fully connected layer is used to output the processing result of the image processing model.

[0100] In the operation process of the convolutional layer, the local features in the picture can be extracted by filtering of the convolution kernel of the convolutional layer. The pooling layer can more effectively reduce the data dimension than the convolutional layer, and can greatly reduce the amount of calculation.

[0101] Reference Figure 10 and Figure 11 , Figure 10 is a schematic diagram of a neural network processing an image for an image processing model in embodiments of the present application, Figure 11 is a schematic diagram of image segmentation effect of a neural network of an image processing model in embodiments of the present application, specifically,

[0102] In a complex video frame image, the video frame image also needs to be segmented. In some embodiments of the present application, a feasible architecture for image segmentation problems can use the SegNet network shown in Figure 10 to segment the body of the target object into different parts. Of course, in the process of identifying and processing the image to be processed, the image processing apparatus encapsulated in the small program process can also detect different objects in the image to be processed or the image template selected by the user through the deep learning image segmentation network to segment different parts of these objects. When segmenting and identifying different parts of the object, at least one of the following can be used: target detection algorithm (SSD Single Shot Multi Box Detector), R-CNN algorithm (Region-Convolutional Neural Networks); Fast R-CNN algorithm (Fast Region-Convolutional Neural Networks), SPP-NET algorithm (Spatial Pyramid Pooling Network), YOLO algorithm (You Only Look Once), FPN algorithm (Feature Pyramid Networks), and variable convolution algorithm DCN (Deformable ConvNets).

[0103] Reference Figure 11 When the user triggers the face changing function of the small program, for the image A to be processed presented in the view interface, the body of the target object is segmented into different parts through the face changing function of the small program SegNet network, as shown in image B, which realizes the segmentation of the body of the target object into arms, torso, neck, face and hairstyle, and is represented by different image shadows. Finally, as shown in image C, through the selection instruction of the user, the face to be replaced is selected as the target part of the first object in the image to be processed. It should be noted that when the user triggers the face changing function of the small program, the other body parts of the first object can also be replaced according to the different part recognition results of the target object shown in image B, and it is not limited to the target part being the face.

[0104] Reference Figure 12 , Figure 12For the image segmentation effect diagram of the neural network of the image processing model in the embodiments of the present application, taking a convolutional neural network as an example, the convolutional neural network of the image processing model is used to process the target image. When the target part is a face, the color mode of the superimposed rendering of the first texture feature set and the second texture feature set is adjusted, and the face of the first object in the to-be-processed image matches the color of the second object in the target image template. Or, when the target part is a hairstyle part (for example Figure 11 The area a shown in the figure is the hairstyle of the first object as the target part, Figure 12 The area b shown in the figure is the hairstyle of the second object as the target part), the margin mode of the superimposed rendering of the first texture feature set and the second texture feature set is adjusted, so that the hairstyle part of the first object in the to-be-processed image matches the facial feature of the second object in the target image template. Wherein, in the process of the image processing model trained by artificial intelligence technology for face replacement of the target object, due to the variety of color changes of the target image, it is easy to cause large deformation of the color and hairstyle contour of the face to be replaced by the target user, and there is a large gap between the new human face and the human face to be replaced, which reduces the recognition of the image. Therefore, by adjusting the color mode of the superimposed rendering of the first texture feature set and the second texture feature set, the face of the first object is matched with the color of the second object in the target image template, which can avoid the distortion caused by the too large difference between the face color of the generated new image and the target image template.

[0105] Further, since different users have different needs when implementing the face replacement function, for example, it can be desired that the face of the first object in the to-be-processed image matches the overall color of the second object in the target image template, and it can also be desired that the face of the first object in the to-be-processed image matches the face color of the second object in the target image template to increase the interest of the face replacement function. Therefore, in some embodiments of the present application, when the target part is a face, the color mode of the superimposed rendering of the first texture feature set and the second texture feature set can be adjusted according to the selection instruction of the user, so that the face of the first object in the to-be-processed image matches the face color of the second object in the target image template.

[0106] Further, referring to Figure 13 , Figure 13For the image segmentation effect diagram of the neural network of the image processing model in the embodiments of the present application, continuing to take the convolutional neural network as an example, when the first object in the to-be-processed image is a girl (the target part is long hair), and the second object in the target image template is a boy (the corresponding part is short hair), the edge margin mode of the superposition and rendering of the first texture feature set and the second texture feature set is adjusted, so that the hairstyle part of the first object in the to-be-processed image is adapted to the facial feature of the second object in the target image template, and the distortion of the face changing result caused by the inconsistency of the hairstyle of the first object and the second object in the face changing process is avoided.

[0107] Specifically, the texture a can be established in the to-be-processed image, the hairstyle and the face area in the face A in the to-be-processed image are identified, and the texture a1 is established through matting, and the determined texture information is represented based on the corresponding gray value. Further, the target image template corresponds to the server that saves different template information, and the user can select different target image templates according to the user's own preference, determine the face of the person to be replaced, and establish the texture b according to the target image template, identify the hairstyle and the face area of the face B in the template image, and establish the texture b1 through matting. In the process of executing the image processing method provided in the present application, due to the diversity of the physical properties of the surface of the object in the image, the gray or color information representing a certain specific surface feature is different due to the diversity of the physical properties, and different physical surfaces will produce different texture images, so that the texture of the corresponding image can be accurately represented through matting. Further, different face recognition algorithms can be used to detect the face area in the to-be-processed image, for example, the supervised descent method (SDM Supervised Descent Method), to realize the supplementary detection of the face area in the to-be-processed image, and to further improve the accuracy of the detection of the face area in the to-be-processed image.

[0108] Continuing to refer to Figure 14 , Figure 14 For the effect diagram of image processing in the embodiments of the present application, the above four textures a, a1, b, and b1 are superimposed and rendered in the OpenGL environment, the texture b1 pixel is changed to be close to the background color. The texture a1 is transformed so that the face A position and size adapt to the face B. And a1 is superimposed on the texture b, for reference to formula 1:

[0109] ResultColor.rgb = mix (modelColor.rgb, userColor.rgb, userColor.r) Formula 1

[0110] Further, the position and size transformation process of the five features of the target face includes:

[0111] The chin position of the target face key point can be determined as a reference point for position offset, and face A is offset so as to be completely aligned with the chin of face B. All face key points of face A and face B are determined as a minimum rectangular selection, and the area ratio of the selection is taken as the scale of scaling. The difference between the rotation angles roll1 and roll2 of the face can also be obtained through the corresponding face key points, which is the rotation angle that needs to be corrected.

[0112] In some embodiments of the present application, the RGB difference value of the superimposed texture can also be adjusted to achieve skin color balance of the face changing target.

[0113] In some embodiments of the present application, the RGB difference value of the superimposed texture can also be adjusted to achieve skin color balance of the face changing target. Figure 12 In some embodiments of the present application, the RGB difference value of the superimposed texture can also be adjusted to achieve skin color balance of the face changing target. Figure 13 The RGB difference value of the superimposed texture can also be adjusted to achieve skin color balance of the face changing target. The RGB difference value of the superimposed texture can also be adjusted to achieve skin color balance of the face changing target.

[0114] The RGB difference value of the superimposed texture can also be adjusted to achieve skin color balance of the face changing target. The RGB difference value of the superimposed texture can also be adjusted to achieve skin color balance of the face changing target.

[0115] The RGB difference value of the superimposed texture can also be adjusted to achieve skin color balance of the face changing target. The RGB difference value of the superimposed texture can also be adjusted to achieve skin color balance of the face changing target. The RGB difference value of the superimposed texture can also be adjusted to achieve skin color balance of the face changing target.

[0116] The RGB difference value of the superimposed texture can also be adjusted to achieve skin color balance of the face changing target. Figure 10 The RGB difference value of the superimposed texture can also be adjusted to achieve skin color balance of the face changing target. Figure 11 The image information processing method provided by the present application can retain the face key point features and hairstyle of the user without causing face expression distortion, better adapt to the face skin color of the template picture, and improve the recognition degree of the user's face after face changing, thereby improving the user's use experience.

[0117] In some embodiments of the present application, the method further comprises:

[0118] In response to a viewing operation on the image transformation function item, a content page including the to-be-processed image and the image template is presented, and at least one interactive function item is presented in the content page, the interactive function item being used to realize interaction with the to-be-processed image; an interaction operation on the to-be-processed image triggered based on the interactive function item is received, so as to execute a corresponding interaction instruction. Wherein, Figure 15The schematic diagram of the image processing interface provided by the embodiment of the present application is shown in the figure. When the user triggers the face changing function, in response to the viewing operation on the image transformation function item, a content page including the to-be-processed image and the image template is presented. Since the to-be-processed image is filtered from the existing gallery information of the user terminal, it cannot meet the real-time needs of the user. Therefore, at least one interactive function item is presented in the content page, the interactive operation on the to-be-processed image triggered based on the interactive function item is received, and the corresponding interactive instruction is executed. The shooting process of the user terminal can be switched to, the collection of the real-time image of the user is realized, and the first object in the to-be-processed image is collected according to the real-time environment of the user. The acquisition of the face image data is realized through the interactive instruction shown in the embodiment. The following methods can be used: a face detection algorithm is used to frame the face position of the target object; a feature point such as the eyes, mouth, nose, and the like of the face is marked using a feature point positioning algorithm; the face image is intercepted according to the detected face position, and the intercepted face image is aligned based on the feature point (for example, the eyes). The exemplary resolution of the face image can be 1024*1024 (pixels).

[0119] In some embodiments of the present application, the method further comprises:

[0120] The first interactive prompt information is presented in the content page, the first interactive prompt information is used to prompt that the interactive content corresponding to the interactive operation can be presented in the view interface, and the content page is switched to the view interface in response to the operation of switching to the view interface. When the user triggers the face changing function applet, if the to-be-processed image presented in the view interface is not adaptive, the user terminal can re-collect the first object in the real-time environment, the user confirms through the first interactive prompt information that the first object in the real-time environment re-collected by the terminal can be presented in the view interface for the face changing process, and the content page is switched to the view interface, so as to realize the adjustment of the first object to be processed according to the real-time environment of the user and enrich the selection categories of the user.

[0121] In some embodiments of the present application, the method further comprises:

[0122] The second interactive prompt information is presented in the content page, the second interactive prompt information is used to prompt that the interactive content corresponding to the interactive operation can be presented in the target image template library interface corresponding to the target image template, and the content page is switched to the target image template library interface in response to the instruction of switching to the target image template library interface. Specifically, reference is made to Figure 16 , Figure 16The schematic diagram of the image processing interface provided by the embodiment of the present application is shown in the figure. Since the different target image templates in the target image template library of the face changing process are all pre-processed, the texture feature sets corresponding to different target parts are determined, and the types of the templates are fixed. The user can confirm that the second object in the real-time environment re-collected by the terminal can be presented in the target image template library interface through the second interactive prompt information, so as to configure the image templates in the target image template library by the user according to the different real-time environments of the user.

[0123] The embodiment of the present application presents an image transformation function item in the view interface; in response to a triggering operation on the image transformation function item, a to-be-processed image containing the first object is acquired and presented; in response to a transformation determination operation triggered based on the to-be-processed image, a target image of a target image template is generated and presented, wherein the target image of the target image template is an image generated by replacing a corresponding part of a second object in the target image template with a target part of the first object in the to-be-processed image; thus, the characteristics of the to-be-processed image can be retained, the distortion of the replaced target image can be avoided, batch processing of different to-be-processed images of the user can be realized, the occupation of hardware resources in the image processing process is reduced, the increase of the cost of hardware devices is reduced, and the use experience of the user is improved.

[0124] The above is only an embodiment of the present application, and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. An image processing method, characterized in that: The method comprises: Presenting image transformation function items in the view interface; In response to a triggering operation on the image transformation function item, acquiring and presenting an image to be processed containing a first object; In response to a transformation determination operation triggered based on the image to be processed, a first texture of the image to be processed is established, and a hairstyle and a face area in the image to be processed are identified, and a second texture of the corresponding area is established by cutout; Establishing a third texture of the target image template, identifying the hairstyle and face area in the target image template, and establishing a fourth texture of the corresponding area by cutting out the image; Overlaying and rendering the first texture, the second texture, the third texture, and the fourth texture to obtain a target image of a target image template; wherein the second texture is overlaid on the third texture; During the overlay rendering process, adjusting the color difference between the second texture and the fourth texture to balance the skin color of the target image; A target image of the target image template is presented.

2. The method according to claim 1, characterized in that The method comprises: Presenting an image template selection interface, and presenting at least one image template in the image template selection interface; In response to an image template selection operation triggered based on the image template selection interface, the image template corresponding to the selection operation is used as the target image template.

3. The method according to claim 1, characterized in that After presenting the target image of the target image template, the method includes: Presenting an image template switching interface, and presenting at least one image template to be switched in the image template switching interface; In response to a template switching operation triggered based on the image template switching interface, presenting an image in the image template indicated by the template switching operation; The image is generated by using the target part of the first object in the image to be processed to replace the corresponding part of the second object in the image template indicated by the template switching operation.

4. The method according to claim 3, characterized in that The method further comprises: In the image template switching interface, the image template indicated by the template switching operation is marked.

5. The method according to claim 1, wherein The method further comprises: Presenting an image sharing function item for sharing the target image; In response to a triggering operation on the image sharing function item, the target image is shared.

6. The method according to claim 1, characterized in that The method further comprises: During the overlay rendering process, a margin mode of the overlay rendering is adjusted with respect to the hairstyle part, so that the hairstyle part of the first object in the to-be-processed image is adapted to the facial features of the second object in the target image template.

7. The method according to claim 1, characterized in that The method further comprises: In response to a viewing operation on the image transformation function item, presenting a content page including the image to be processed and the image template, and presenting at least one interactive function item in the content page, the interactive function item being used to implement interaction with the image to be processed; An interactive operation on the image to be processed triggered by the interactive function item is received to execute a corresponding interactive instruction.

8. The method according to claim 7, characterized in that The method further comprises: Presenting first interactive prompt information on the content page, where the first interactive prompt information is used to prompt that the interactive content corresponding to the interactive operation can be presented on the view interface; In response to the operation of switching to the view interface, the content page is switched to the view interface.

9. The method according to claim 7, characterized in that The method further comprises: Presenting second interactive prompt information on the content page, wherein the second interactive prompt information is used to prompt that the interactive content corresponding to the interactive operation can be presented in the target image template library interface corresponding to the target image template; In response to an instruction to switch to the target image template library interface, the content page is switched to the target image template library interface.

10. An image processing device, characterized in that: The device comprises: An information display module, used to present image transformation function items in the view interface; An information transmission module, configured to acquire and present an image to be processed containing a first object in response to a triggering operation on the image transformation function item; The information display module is configured to establish a first texture of the image to be processed in response to a transformation determination operation triggered based on the image to be processed, identify a hairstyle and a face region in the image to be processed, and establish a second texture of the corresponding region by cutout; Establishing a third texture of the target image template, identifying the hairstyle and face area in the target image template, and establishing a fourth texture of the corresponding area by cutting out the image; Overlaying and rendering the first texture, the second texture, the third texture, and the fourth texture to obtain a target image of a target image template; wherein the second texture is overlaid on the third texture; During the overlay rendering process, adjusting the color difference between the second texture and the fourth texture to balance the skin color of the target image; A target image of the target image template is presented.

11. The device according to claim 10, characterized in that The information display module is configured to present an image template selection interface and present at least one image template in the image template selection interface; The information display module is configured to respond to an image template selection operation triggered based on the image template selection interface and use the image template corresponding to the selection operation as the target image template.

12. The device according to claim 10, characterized in that The information display module is configured to present an image template switching interface and present at least one image template to be switched in the image template switching interface; The information display module is configured to, in response to a template switching operation triggered based on the image template switching interface, present an image in the image template indicated by the template switching operation; The image is generated by using the target part of the first object in the image to be processed to replace the corresponding part of the second object in the image template indicated by the template switching operation.

13. An electronic device, characterized in that: The electronic device comprises: a memory for storing executable instructions; A processor, configured to implement the image processing method according to any one of claims 1 to 9 when running the executable instructions stored in the memory.

14. A computer-readable storage medium storing executable instructions, characterized in that: When the executable instructions are executed by a processor, the image processing method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Training method of fusion image processing model, image processing method, device and storage medium

    CN110826593A