Image processing method and device, electronic equipment and storage medium

By using a pre-trained stylized model to process images, the target style image and key point information after the stylized processing of the target object is solved, the problem of inaccurate detection of key points of the target object is achieved, the accurate application of target special effects is achieved, and the display effect of special effects data is improved.

CN120298201APending Publication Date: 2025-07-11BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410039696.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-10
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, the key point detection of the target object after image stylization is inaccurate, resulting in the special effect tasks being unable to accurately act on the target object, resulting in poor display effect of the special effect data.

Method used

The image is processed through the pre-trained stylized model, and the target style image and key point information after the stylized processing of the target object are obtained, and the target style image is processed based on the pre-set target task and key point information to obtain the target special effects corresponding to the target task.

Benefits of technology

The key point information of the target object after stylized processing is accurately determined, the accuracy and effect of key point detection is improved, and the matching degree between the target task and the key point information of the target object is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298201A_ABST
    Figure CN120298201A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image processing method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring a to-be-processed image comprising a target object; processing the to-be-processed image based on a pre-trained stylization model to obtain a target style image for stylizing the target object and key point information of the stylized target object; and processing the target style image according to a preset target task and key point information to obtain a target special effect corresponding to the target task. According to the technical scheme, the effect of accurately determining the key point information of the target object after stylization processing is achieved, and the effect of deeply coupling image stylization processing and object key point detection is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of image processing technologies, and in particular, to an image processing method, apparatus, electronic device, and storage medium. Background Art

[0002] Currently, when performing a special effect task on an image, usually the image is stylized, or a special effect prop is used to perform special effect processing on the image to obtain a special effect image. Further, the special effect data is obtained by processing the special effect image according to the special effect task.

[0003] However, when processing an image based on the above method, the target object in the obtained special effect image usually undergoes deformation. Further, when processing the special effect image according to the special effect task, there may be a problem that the special effect task cannot be accurately applied to the target object, resulting in a poor display effect of the special effect data, that is, there is a problem of poor matching between the special effect and the target object. Summary of the Invention

[0004] The present disclosure provides an image processing method, apparatus, electronic device, and storage medium to achieve the effect of accurately determining the key point information of the target object after stylization processing, achieving the effect of deeply coupling the image stylization processing and object key point detection, and improving the matching degree between the target task and the key point information of the target object.

[0005] In a first aspect, an embodiment of the present disclosure provides an image processing method, which includes:

[0006] Obtain a to-be-processed image including a target object;

[0007] Process the to-be-processed image based on a pre-trained stylization model to obtain a target stylized image of the target object after stylization processing and the key point information of the target object after stylization processing;

[0008] Process the target stylized image according to a preset target task and the key point information to obtain a target special effect corresponding to the target task.

[0009] In a second aspect, an embodiment of the present disclosure further provides an image processing apparatus, which includes:

[0010] An image acquisition module, configured to obtain a to-be-processed image including a target object;

[0011] An image processing module, configured to process the to-be-processed image based on a pre-trained stylization model to obtain a target stylized image of the target object after stylization processing and the key point information of the target object after stylization processing;

[0012] An effect determination module, configured to process the target style image according to a preset target task and the key point information, so as to obtain a target effect corresponding to the target task.

[0013] In a third aspect, an embodiment of the present disclosure further provides an electronic device, which includes:

[0014] One or more processors;

[0015] A storage device, configured to store one or more programs,

[0016] When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the image processing method according to any one of the embodiments of the present disclosure.

[0017] In a fourth aspect, an embodiment of the present disclosure further provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute the image processing method according to any one of the embodiments of the present disclosure when executed by a computer processor.

[0018] The technical solution of the embodiment of the present disclosure is to obtain a to-be-processed image including a target object, and further, process the to-be-processed image based on a pre-trained stylization model to obtain a target style image obtained by stylizing the target object and key point information of the target object after stylization processing. Finally, according to a preset target task and key point information, the target style image is processed to obtain a target effect corresponding to the target task, which solves the problems in the related art that when processing an image, the special effect task cannot be accurately applied to the target object, resulting in poor display effects of special effect data, etc. The effect of accurately determining the key point information of the target object after stylization processing is achieved, and the effect of deeply coupling image stylization processing and object key point detection is achieved. Furthermore, the accuracy and effect of key point detection are improved, and the effect of executing a target task on the target object according to the key point information to obtain a target effect is achieved, and the matching degree between the target task and the key point information of the target object is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Combined with the accompanying drawings and referring to the following specific embodiments, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more obvious. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the original elements and elements are not necessarily drawn to scale.

[0020] Figure 1 It is a schematic flow chart of an image processing method provided by an embodiment of the present disclosure;

[0021] Figure 2Schematic flowchart of another image processing method provided by an embodiment of the present disclosure;

[0022] Figure 3 Schematic diagram of a stylization model provided by an embodiment of the present disclosure;

[0023] Figure 4 Schematic flowchart of another image processing method provided by an embodiment of the present disclosure;

[0024] Figure 5 Schematic structural diagram of an image processing apparatus provided by an embodiment of the present disclosure;

[0025] Figure 6 Schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners

[0026] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0027] It should be understood that the various steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0028] As used herein, the term "including" and its variations are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0029] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependent relationships.

[0030] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".

[0031] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of these messages or information.

[0032] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to users in an appropriate manner and user authorization should be obtained in accordance with relevant laws and regulations.

[0033] For example, when responding to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that executes the operation of the technical solution of the present disclosure based on the prompt message.

[0034] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user can be, for example, a pop-up window manner, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0035] It can be understood that the above process of notifying and obtaining user authorization is only illustrative and does not constitute a limitation on the implementation manner of the present disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0036] It can be understood that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of corresponding laws, regulations and related provisions.

[0037] Before introducing the present technical solution, an exemplary description of the application scenario can be made first. This technical solution can be applied to any scenario where an effect task needs to be performed on a stylized image after stylization processing. Exemplarily, when a to-be-processed image including a target object is obtained, the to-be-processed image can be stylized to obtain a stylized effect image. At this time, the target object in the image can be an object after stylization processing. Further, the key points of the target object after stylization processing can be determined, and the stylized effect image can be processed according to the effect task and the key point information to obtain the target effect. In the related art, the key points of the target object are usually determined through a key point detection algorithm. However, the key point detection algorithm is trained based on the real data of the target object. When performing key point detection on the target object after stylization processing, there may be a problem that the detection result is inaccurate, and further, the effect elements cannot be accurately attached to the target object.

[0038] At this time, based on the technical solution of the embodiments of the present disclosure, after obtaining a to-be-processed image including a target object, the to-be-processed image can be input into a stylization model to perform stylization and key point determination processing on the to-be-processed image based on the stylization model. Further, a target stylized image of the target object after stylization processing and key point information of the target object after stylization processing can be obtained. Further, according to a preset target task and the key point information, the target stylized image is processed to obtain a target special effect corresponding to the target task. Thus, the effect of both stylizing the image and determining the key points of the stylized image is achieved. Further, the effect of accurately determining the key point positions of the object after stylization processing is achieved. Moreover, the effect of performing a target task on the target object according to the key point information to obtain a target special effect is achieved, improving the matching degree between the target task and the key point information of the target object.

[0039] Figure 1 FIG. is a schematic flowchart of an image processing method provided by an embodiment of the present disclosure. The embodiments of the present disclosure are applicable to the situation of determining key points for an object after stylization processing and performing a special effect task on the object in a special effect image according to the key point information. This method can be executed by an image processing device, which can be implemented in the form of software and / or hardware. Optionally, it is implemented by an electronic device, which can be a mobile terminal, a PC, or a server, etc.

[0040] As Figure 1 shown, the method of this embodiment may specifically include:

[0041] S110. Obtain a to-be-processed image including a target object.

[0042] In this embodiment, the to-be-processed image can be understood as an image that needs to be processed with special effects. Optionally, the to-be-processed image can be a default template image, an image collected based on a terminal device, an image obtained from a target storage space (such as an image library of an application software or a local terminal album) in response to a user's trigger operation, or an image received from an external device upload. Among them, the terminal device can refer to an electronic device with an image shooting function such as a camera, a smart phone, and a tablet computer. Correspondingly, the to-be-processed image may include a target object. The target object can be understood as an object to be processed with special effects. The target object can be any type of object included in the image. Optionally, the target object includes a human object, an animal object, or a certain limb part associated with the user, etc.

[0043] It should be noted that the number of target objects included in the image to be processed can be one or more. Whether it is one or more, the technical solutions provided by the embodiments of the present disclosure can be used to process them.

[0044] In practical applications, when performing special effect processing on any image and / or object in the image, an image to be processed including the target object can be obtained first. Furthermore, subsequent image processing procedures can be continued on the obtained image to be processed.

[0045] It should be noted that the terminal device for obtaining the image to be processed can be a terminal that supports special effect processing on images. For example, it can be a user terminal registered in an application software with special effect processing functions; or a user terminal registered in an application software with special effect prop production functions. The embodiments of the present disclosure do not make specific limitations in this regard.

[0046] S120. Process the image to be processed based on a pre-trained stylization model to obtain a target stylized image of the stylized processing of the target object and the key point information of the target object after the stylized processing.

[0047] In this embodiment, the stylization model can be a neural network model that takes an image as an input object, performs stylization processing on the image and determines key point information, and outputs an image of a specific stylization type and key point information. The stylization model can include a stylization processing unit for performing stylization processing on the image and a key point extraction unit for detecting and extracting key points of the image after the stylization processing. Among them, the stylization processing unit can be a neural network model with any structure. Optionally, it can be a Generative Adversarial Network (GAN). The key point extraction unit can be any neural network model capable of detecting and extracting key points of the target object in the image. Optionally, the key point extraction unit can include a downsampling layer, a flattening layer, and a linear layer. It should be noted that the stylization model can also include an encoder. Furthermore, based on the encoder, feature extraction can be performed on the image to be processed to obtain rich and non-redundant image features. Further, the extracted image features can be used as the inputs of the stylization processing unit and the key point extraction unit respectively, and a stylized image corresponding to the image to be processed and the key point information of the target object after the stylized processing can be obtained. The advantage of setting an encoder in the stylization model is that redundant image features in the image to be processed can be removed, thereby improving the processing efficiency and key point extraction accuracy of the stylization model and enhancing the image effect of the stylized image. It should also be noted that the stylization model and the stylization special effect are in one-to-one correspondence, that is, the stylization model corresponding to any stylization special effect can only output a target stylized image corresponding to the stylization special effect.

[0048] In this embodiment, after obtaining the image to be processed including the target object, the image to be processed can be input into the stylization model, and the image to be processed can be processed based on the stylization model to obtain the target stylized image of the stylized processing of the target object and the key point information of the target object after the stylized processing.

[0049] Among them, the target stylized image can be understood as a special effect image obtained by adding a stylized special effect to the target object in the image. It should be noted that the dimension of the target stylized image can be associated with the stylized processing unit in the stylization model. Exemplarily, if the stylized processing unit is a deformation stylized processing unit, the dimension of the target stylized image can be 6D; if the stylized processing unit is a non-deformation stylized processing unit, the dimension of the target stylized image can be 3D or 4D.

[0050] Among them, the key point information can be understood as information characterizing the position of the key points on the target object. The key points on the target object can be pre-defined key points, and the number of these key points can be one or more. The key point information can include any coordinate information that can characterize the position of the key points in the coordinate system where they are located. Optionally, the key point information includes the pixel coordinates of the pre-defined key points after the deformation of the target object; or, the key point information includes the pixel coordinates and depth coordinates of the pre-defined key points after the deformation of the target object. The pre-defined key points can be understood as pre-determined key points corresponding to the target object. The number of key points can be one or more. Among them, the deformation of the target object can be understood as the stylized processing of the target object to cause the target object to deform. The pixel coordinates can be the position coordinates of the key points in the image, and the information stored in the coordinates is the key point pixel information. The depth coordinates can be the coordinate values characterizing the depth information of the key points in the image. Among them, the depth information can be the shooting distance between the target object in the image and the shooting device. When the key point information includes the pixel coordinates of the key points after the deformation of the target object, the key point can be a two-dimensional key point; when the key point information includes the pixel coordinates and depth coordinates of the key points after the deformation of the target object, the key point can be a three-dimensional key point. It should be noted that the advantage of the key point information including pixel coordinates or including pixel coordinates and depth coordinates is that: it can enable the finally obtained key point information to perform both the task of mounting two-dimensional special effect objects and the task of mounting three-dimensional special effect objects, improving the flexibility of the key point information.

[0051] It should be noted that whether the key point information finally output by the stylization model includes pixel coordinates or both pixel coordinates and depth coordinates is associated with the target task to be performed based on the key point information. In the case where the target task is a task that can be performed without depth information, the key point information may include the pixel coordinates of the key points after the deformation of the target object; in the case where the target task is a task that requires depth information to be performed, the key point information may include the pixel coordinates and depth coordinates of the key points after the deformation of the target object.

[0052] In the actual application process, the acquired to-be-processed image including the target object can be input into the stylization model. Furthermore, the to-be-processed image can be subjected to feature extraction based on the encoder in the stylization model to obtain to-be-processed features. After that, the to-be-processed features can be processed based on the stylization processing unit to perform stylization processing on the target object and obtain the target style image. Meanwhile, the to-be-processed features can be processed based on the key point extraction unit to detect the pre-defined key points after the deformation of the target object and extract the key point information. Furthermore, the key point information of the target object after stylization processing can be obtained.

[0053] S130. Process the target style image according to the pre-set target task and key point information to obtain a target special effect corresponding to the target task.

[0054] It should be noted that for the target style image, after obtaining the target style image, the target style image can be directly displayed on the display interface; for the key point information, the key point information can be associated with the subsequent target task to be performed and can provide an execution basis for the subsequent target task to be performed. Therefore, after obtaining the key point information, before receiving the target task, the pre-defined key points may not be displayed based on the key point information. In the case of receiving the target task, the target task can be performed on the target style image based on the key point information.

[0055] In this embodiment, the target task can be understood as a special effect task performed based on the key point information. The target task may be a task of mounting a special effect object at the key points of the target object. Optionally, the target task may include a task of mounting a two-dimensional special effect object and / or a three-dimensional special effect object for the target object based on the key point information. For example, the target task may be to mount an earring special effect at the key points of the target object. The two-dimensional special effect object can be understood as a two-dimensional special effect element that can be mounted on the target object. The three-dimensional special effect object can be understood as a three-dimensional special effect element that can be mounted on the target object. The target special effect can be understood as a special effect image corresponding to the target style image and with the special effect object mounted on the target object.

[0056] It should be noted that the target task may include both two-dimensional special effect objects and three-dimensional special effect objects, or may include two-dimensional special effect objects or three-dimensional special effect objects. Moreover, the types of special effect objects (including two-dimensional special effect objects and / or three-dimensional special effect objects) included in the target task may also be one or more, and the embodiments of the present disclosure do not make specific limitations thereto.

[0057] In practical applications, if there is a preset target task, the special effect objects to be mounted (including two-dimensional special effect objects and / or three-dimensional special effect objects) can be determined based on the target task. Further, the determined special effect objects can be mounted on the target object according to the key point information to process the target style image. Furthermore, the target special effect corresponding to the target task can be obtained. The advantage of such a setting is that: it realizes the effect of mounting special effect objects on the target object according to the key point information to obtain the target special effect, and improves the adaptability between the special effect objects to be mounted and the target object.

[0058] Exemplarily, assume that the target task is to mount special effect earrings on the target object according to the key point information, and the determined key point information may include the ear key point information of the target object. After obtaining the target stylized image and the key point information, the special effect earrings can be mounted on the ears of the target object according to the ear key point information. Furthermore, the target stylized image of the target object with the special effect earrings mounted can be obtained, and this image can be used as the target special effect corresponding to the target task.

[0059] The technical solution of the embodiments of the present disclosure obtains the image to be processed including the target object. Further, the image to be processed is processed based on a pre-trained stylization model to obtain the target stylized image of the stylization processing of the target object and the key point information of the target object after the stylization processing. Finally, according to the preset target task and the key point information, the target stylized image is processed to obtain the target special effect corresponding to the target task, which solves the problems in the related art that when processing an image, the special effect task cannot be accurately applied to the target object, resulting in poor display effects of the special effect data, etc. It realizes the effect of accurately determining the key point information of the target object after the stylization processing, achieves the effect of deeply coupling the image stylization processing and the object key point detection, and further improves the accuracy and effect of the key point detection. Moreover, it realizes the effect of performing the target task on the target object according to the key point information to obtain the target special effect, and improves the matching degree between the target task and the key point information of the target object.

[0060] Figure 2Schematic flowchart of another image processing method provided by an embodiment of the present disclosure. Based on the technical solution of the above embodiment, the method in this embodiment can input the image to be processed into a stylization model, and process the image to be processed based on the encoder, stylization processing unit, and key point extraction unit in the stylization model. Furthermore, a target stylized image and key point information can be obtained. For specific implementation manners, reference may be made to the description of this embodiment. Among them, the same or similar technical features as those in the foregoing embodiments will not be described herein again.

[0061] As Figure 2 shown, the method of this embodiment may specifically include:

[0062] S210. Obtain an image to be processed including a target object.

[0063] S220. Determine the feature to be processed of the image to be processed based on the encoder in the stylization model, and process the feature to be processed based on the stylization processing unit and the key point extraction unit in the stylization model respectively to obtain a target stylized image and key point information.

[0064] Among them, the encoder can be understood as a neural network model that processes an image into features of a corresponding feature dimension. In this embodiment, the encoder in the stylization model can be a neural network model composed of at least one convolutional layer and a downsampling layer, and this neural network model can encode the image to be processed into a feature image of a specific dimension. The stylization processing unit can be understood as a neural network model that can perform stylization processing on a target object. The stylization processing unit can be a neural network model based on a Generative Adversarial Network (GAN). The stylization processing unit can be a neural network model composed of at least one convolutional layer and an upsampling layer. The key point extraction unit can be understood as a neural network model that can perform key point detection and key point information extraction. The key point extraction unit can be a neural network model composed of a downsampling layer, a flattening layer, and a linear layer.

[0065] In practical applications, after inputting the image to be processed into the stylization model, the encoder in the stylization model can be used to extract features from the image to be processed, obtaining the features to be processed of the image to be processed. The features to be processed output by the encoder can be local image features in the image to be processed. Compared with global image features, local image features have the characteristics of rich quantity contained in the image, small correlation between features, and not being affected by the disappearance of some features in the case of occlusion, etc. Therefore, the features to be processed can also be understood as the local expression of image features, reflecting the local features of the image to be processed, and these features can be applied to application scenarios such as image matching and processing. The advantage of such a setting is that it can improve the efficiency of subsequent feature extraction to accelerate the model calculation speed.

[0066] Furthermore, the features to be processed can be processed based on the stylization processing unit in the stylization model. Furthermore, a target stylized image of the target object can be obtained. And, the features to be processed can be processed based on the key point extraction unit in the stylization model. Furthermore, the key point information of the target object after stylization processing can be obtained. Next, the processing process of the stylization processing unit for the features to be processed and the processing process of the key point extraction unit for the features to be processed will be specifically described.

[0067] Optionally, processing the features to be processed respectively based on the stylization processing unit and the key point extraction unit in the stylization model to obtain the target stylized image and the key point information, including: processing the features to be processed based on at least one convolutional layer and an upsampling layer in the stylization processing unit to obtain the target stylized image of the deformation of the target object; and, processing the features to be processed sequentially based on the downsampling layer, flattening layer, and linear layer in the key point extraction unit to obtain the key point information of the target object after deformation.

[0068] Among them, the flattening layer can be understood as a neural network model for flattening features. The flattening layer is used to convert multi-dimensional input data into a one-dimensional array, usually used between convolutional neural networks and fully connected layers, and does not change the total amount of data itself, but only changes the shape of the data. Exemplarily, assuming there is a three-dimensional feature map, after being processed by the flattening layer, this feature map will be converted into a one-dimensional array as the input of the fully connected layer. The linear layer is also called the fully connected layer. Each node in the fully connected layer is connected to all nodes in the previous layer, and is used to synthesize the local features extracted previously, and its output is a feature value.

[0069] In practical applications, after obtaining the feature to be processed, at least one convolutional layer and upsampling layer in the stylization processing unit can be used to process the feature to be processed. After processing the feature to be processed based on at least one convolutional layer, the processed feature to be processed can be restored to an image of the original size based on the upsampling layer, and this image can be used as the target style image for the deformation of the target object. Moreover, after obtaining the feature to be processed, the downsampling layer in the key point extraction unit can be used to process the feature to be processed to reduce the feature size of the feature to be processed and lower the feature dimension of the feature to be processed. Furthermore, the processed feature to be processed can be flattened based on the flattening layer, and finally, the flattened feature can be mapped based on the linear layer. Thus, the key point information after the deformation of the target object can be obtained. The advantage of such a setting is that it realizes the effect of simultaneously performing stylization processing and key point determination on the image to be processed based on the stylization model, and improves the accuracy of key point detection.

[0070] Exemplarily, it can be combined with Figure 3 to illustrate the process of processing the image to be processed based on the stylization model. Assume that the image size of the image to be processed is 512×512×3. After inputting the image to be processed into the stylization model, the encoder can be used to extract features from the image to be processed to obtain the feature to be processed. At this time, the feature size of the obtained feature to be processed is 4×4. Further, at least one convolutional layer and upsampling layer in the stylization processing unit can be used to process the feature to be processed to encode and restore the feature to be processed into an image with an image size of 512×512×6, and this image can be used as the target style image. Also, the downsampling layer in the key point extraction unit can be used to process the feature to be processed to reduce the 4×4-sized feature to be processed to a 1×1-sized feature. After that, the feature can be flattened based on the flattening layer, and finally, through a linear layer, the flattened feature can be mapped into the final required key point coordinates, and these key point coordinates can be used as the key point information after the deformation of the target object.

[0071] S230. Process the target style image according to the preset target task and key point information to obtain the target special effect corresponding to the target task.

[0072] The technical solution of the embodiment of the present disclosure obtains a to-be-processed image including a target object. Further, based on the encoder in the stylization model, the to-be-processed features of the to-be-processed image are determined, and the to-be-processed features are processed respectively by the stylization processing unit and the key point extraction unit in the stylization model to obtain a target stylized image and key point information. Finally, according to a preset target task and the key point information, the target stylized image is processed to obtain a target special effect corresponding to the target task, realizing the effect of stylizing the image and determining key points based on the stylization model, achieving the effect of deeply coupling image stylization processing and object key point detection. Furthermore, the accuracy and efficiency of key point detection are improved.

[0073] Figure 4 The following is a schematic flowchart of another image processing method provided by the embodiment of the present disclosure. Based on the technical solution of the above embodiment, before processing the to-be-processed image based on the stylization model, a sample image can be obtained, and a three-dimensional reconstruction model corresponding to the target object in the sample image is determined, and the grid point information of at least one predefined key point in the three-dimensional reconstruction model is determined. Further, the three-dimensional reconstruction model is deformed according to the deformation parameter to obtain an actual stylized image, and the actual key point information corresponding to the grid point information under the action of the deformation parameter is determined to construct a training sample. Then, the stylization model can be trained based on the training sample. The specific implementation manner can refer to the description of this embodiment. Among them, the technical features that are the same as or similar to the foregoing embodiment are not described herein again.

[0074] As Figure 4 shown, the method of this embodiment may specifically include:

[0075] S310. Obtain a plurality of sample images.

[0076] Among them, the sample image can be an image captured by a camera device, or an image reconstructed by an image reconstruction model, or an image pre-stored in a storage space. At the same time, the sample image may include one or more objects, and the object included in the image can be used as the target object.

[0077] In practical applications, before training the stylization model, a plurality of training samples can be obtained first to train the model based on the training samples. To improve the accuracy of the model, the training samples can be constructed as many and rich as possible. To construct the training samples, a plurality of sample images can be obtained. Then, the sample images can be processed to obtain the training samples for training the stylization model.

[0078] S320. Determine the three-dimensional reconstruction model corresponding to the target object in the sample image, and determine the grid point information of at least one predefined key point in the three-dimensional reconstruction model.

[0079] It should be noted that for each sample image, the S320 method can be used to determine the corresponding 3D reconstruction model and the grid point information of the key points in the 3D reconstruction model.

[0080] Among them, the 3D reconstruction model can be understood as a 3D model constructed based on the target object, and the model structure of this model corresponds to the object size and object ratio of the target object. In practical applications, when the target object is detected in the sample image, the pre-generated or real-time generated 3D reconstruction model corresponding to the target object can be retrieved. For example, when the target object is detected in the sample image, a 3D mesh reflecting the limb characteristics of the target object can be constructed in real time using multiple patches. Furthermore, the 3D mesh is used as the 3D reconstruction model corresponding to the target object. Further, after the 3D reconstruction model is constructed, the application can also annotate the model and associate it with the object identifier of the target object. Based on this, if the target object is detected again in the sample image in a subsequent process, the already constructed 3D mesh can be directly called as the 3D reconstruction model. In this embodiment, the 3D reconstruction model can be composed of at least one patch, and each patch can be composed of three vertices. The patch can be used as the mesh model, and the vertices can be used as the model vertices. Therefore, the grid point information of the key points in the 3D reconstruction model can be understood as the model vertex information corresponding to the key points, that is, the information characterizing the position of the key points on the 3D reconstruction model. The grid point information can be spatial position information. In practical applications, after obtaining multiple sample images, for each sample image, when the target object is detected in the sample image, the 3D reconstruction model corresponding to the target object can be determined. Further, the grid point information of each key point in the 3D reconstruction model can be determined according to the pre-defined position information of each key point on the target object.

[0081] S330. Perform deformation processing on the 3D reconstruction model according to the deformation parameters corresponding to the target style to obtain the actual stylized image, and determine the actual key point information corresponding to the grid point information under the action of the deformation parameters.

[0082] In this embodiment, the target style can be understood as a special effect that can perform stylized special effects processing on the target object. The target style and the stylized model to be trained can be in one-to-one correspondence, that is, one target style can correspond to one stylized model, and this stylized model can output a target style image corresponding to the target style. Optionally, the target style can include a deformation style and / or a non-deformation style. Among them, the deformation style can be understood as a special effect that performs deformation stylization processing on the target object to change the object shape of the target object. The non-deformation style can be understood as a stylized special effect that does not change the object shape of the target object. Optionally, the non-deformation style can include a material style, that is, a special effect that can change the surface properties of the target object, and a special effect that can change the skin texture characteristics and / or skin display color of the target object in the image to be processed. The deformation parameter can be a parameter used to indicate the stylized deformation effect finally presented by the target object. The deformation parameter corresponds to the target style and is a parameter characterizing the deformation degree corresponding to the corresponding target style. The actual stylized image can be understood as a stylized image obtained by stylizing the target object in the sample image based on the target style. The actual key point information can be understood as the key point information of the target object after stylization processing.

[0083] It should be noted that the deformation parameter can be a parameter determined during the construction of the target style. These parameters can be parameters obtained after multiple edits. Furthermore, the finally determined deformation parameter can present the stylized special effect corresponding to the target style to the greatest extent.

[0084] In practical applications, after determining the three-dimensional reconstruction model corresponding to the target object, the deformation parameter corresponding to the target style can be obtained. Further, the three-dimensional reconstruction model can be deformed according to the deformation parameter of the target style to obtain a deformed three-dimensional reconstruction model. Furthermore, the deformed three-dimensional reconstruction model can be rendered to obtain an actual stylized image corresponding to the sample image, and this stylized image matches the target style.

[0085] Further, the actual key point information corresponding to the mesh point information under the action of the deformation parameter can be determined. The actual key point information can be determined by projecting the mesh point information under the action of the deformation parameter onto the screen coordinate system.

[0086] Optionally, determining the actual key point information corresponding to the mesh point information under the action of the deformation parameter includes: determining the coordinate data corresponding to the mesh point information under the action of the deformation parameter; and determining the actual pixel coordinates of the mesh point information in the sample image based on the coordinate data, the model matrix, the camera matrix, and the projection matrix, and using them as the actual key point information.

[0087] In this embodiment, the coordinate data can be understood as the coordinate data of the key points on the three-dimensional reconstruction model under the action of the deformation parameters. The Model matrix can be a matrix representing the position information of the target object in the world coordinate system. The View matrix can be a matrix representing the position information of the camera in the world coordinate system. The Projection matrix can be the projection transformation matrix of the camera. The actual pixel coordinates can be understood as the coordinate data of the grid point information in the sample image, and the coordinate information stored in the coordinate data is pixel information.

[0088] In practical applications, after performing deformation processing on the three-dimensional reconstruction model based on the deformation parameters, the coordinate data corresponding to the grid point information under the action of the deformation parameters can be determined according to the deformation parameters and the grid point information. Further, the Model matrix, the View matrix, and the Projection matrix can be determined, and matrix transformation processing can be performed on the coordinate data according to the Model matrix, the View matrix, and the Projection matrix. Furthermore, the actual pixel coordinates of the grid point information in the sample image can be obtained, and the actual pixel coordinates can be used as the actual key point information. The advantage of such a setting is that: the effect of accurately annotating the positions of the key points in the training data is achieved, and the effect of improving the preparation efficiency of the key point training data while reducing the preparation cost of the training data is achieved.

[0089] S340. Determine the training samples for training the stylization model according to the sample image, the actual stylized image of the sample image, and the actual key point information.

[0090] In practical applications, after obtaining the sample image, the corresponding actual stylized image of the sample image, and the actual key point information, training samples can be constructed according to the sample image, the corresponding actual stylized image of the sample image, and the actual key point information. So that each training sample includes the sample image, the actual stylized image corresponding to the sample image, and the actual key point information.

[0091] It should be noted that if the target task performed based on the key point information includes the task of attaching a three-dimensional special effect object to the target object according to the key point information, in order to improve the adaptability between the attached three-dimensional special effect object and the target object, the depth coordinates of the grid point information can also be used as the actual key point information, so that the actual key point information corresponding to the grid point information under the action of the deformation parameters includes the actual pixel coordinates and the depth coordinates of the grid point information.

[0092] Based on this, on the basis of the above technical solutions, it further includes: in the case where the target task includes attaching a three-dimensional special effect object to the target object, determining the depth coordinates of the grid point information, and updating the training samples based on the depth coordinates.

[0093] In this embodiment, the depth coordinate can represent the distance between the grid point and the camera, that is, the depth information of the grid point.

[0094] In practical applications, when the target task includes attaching a three-dimensional special effect object to the target object, the distance between the grid point information and the camera can be determined. Furthermore, the depth coordinate of the grid point information can be obtained. Further, the training sample can be updated according to the depth coordinate, so that the training sample can include the depth coordinate of the grid point information.

[0095] It should be noted that after obtaining the actual key point information in the training sample, the actual key point information can be normalized so that both the actual pixel coordinate and the depth coordinate in the actual key point information are normalized to the interval from -1 to 1.

[0096] S350. Train the stylization model based on the training sample.

[0097] In this embodiment, after obtaining multiple training samples, the stylization model can be trained based on the training samples. It should be noted that for each training sample, the following process can be used to train it, so as to obtain the stylization model.

[0098] Optionally, training the stylization model based on the training sample includes: inputting the sample image in the training sample into the stylization model for stylization and key point determination processing, and outputting a predicted stylized image and predicted key point information; determining a loss value based on the actual stylized image, actual key point information, predicted stylized image, and predicted key point information of the training sample, and correcting the model parameters in the stylization model based on the loss value until the loss function in the stylization model converges.

[0099] In this embodiment, the predicted stylized image can be the stylized image output after stylizing the target object by inputting the sample image into the stylization model. The predicted key point information can be the key point information of the target object after stylization output by inputting the sample image into the stylization model. The loss value can be a value representing the degree of difference between the predicted output and the actual output. The loss function can be determined based on the loss value and is used to represent the function of the degree of difference between the predicted output and the actual output.

[0100] In practical applications, for each training sample, the sample image in the training sample can be input into the stylization model, and the sample image can be feature-extracted based on the encoder in the stylization model to obtain the sample feature to be processed of the sample image. Furthermore, the sample feature to be processed can be respectively subjected to stylization and key point determination processing based on the stylization processing unit and the key point extraction unit in the stylization model. Thus, a predicted stylized image and predicted key point information can be output.

[0101] Further, the predicted stylized image can be compared with the actual stylized image in the training samples, and the predicted key point information can be compared with the actual key point information in the training samples to determine the loss value. Furthermore, the model parameters in the stylized model can be corrected according to the loss value. After that, the training error of the loss function in the stylized model, that is, the loss parameter, can be used as the condition for detecting whether the current loss function has converged. For example, whether the training error is less than the preset error, or whether the error change trend tends to be stable, or whether the current model iteration number is equal to the preset number, etc. If it is detected that the convergence condition is met, for example, the training error of the loss function is less than the preset error or the error change tends to be stable, it indicates that the training of the stylized model is completed. At this time, the iterative training can be stopped. If it is detected that the current convergence condition is not met, other training samples can be further obtained to train the stylized model until the training error of the loss function is within the preset range. When the training error of the loss function reaches convergence, the trained stylized model can be obtained.

[0102] S360. Obtain a to-be-processed image including a target object.

[0103] S370. Process the to-be-processed image based on a pre-trained stylized model to obtain a target stylized image of the target object after stylization processing and the key point information of the target object after stylization processing.

[0104] S380. Process the target stylized image according to a preset target task and key point information to obtain a target special effect corresponding to the target task.

[0105] In the technical solution of the embodiment of the present invention, by obtaining a plurality of sample images, further determining a three-dimensional reconstruction model corresponding to the target object in the sample images, and determining the grid point information of at least one predefined key point in the three-dimensional reconstruction model. Then, performing deformation processing on the three-dimensional reconstruction model according to the deformation parameters corresponding to the target style to obtain an actual stylized image, and determining the actual key point information corresponding to the grid point information under the action of the deformation parameters. Then, according to the sample image, the actual stylized image of the sample image, and the actual key point information, determining a training sample for training the stylized model. Further, training the stylized model based on the training sample. Then, obtaining a to-be-processed image including the target object, processing the to-be-processed image based on the pre-trained stylized model to obtain a target style image of the stylized target object and the key point information of the target object after stylized processing. According to the predefined target task and the key point information, processing the target style image to obtain a target special effect corresponding to the target task, which realizes the effect of training a stylized model including a stylized processing unit and a key point extraction unit based on the constructed training data. Further, it realizes the effect of stylizing the image and determining the key points based on the stylized model, achieving the effect of deeply coupling the image stylization processing and the object key point detection. Furthermore, it improves the accuracy and efficiency of key point detection.

[0106] Figure 5 FIG. is a schematic structural diagram of an image processing device provided by an embodiment of the present disclosure, as Figure 5 shown, the device includes: an image acquisition module 410, an image processing module 420, and a special effect determination module 430.

[0107] Among them, the image acquisition module 410 is configured to acquire a to-be-processed image including a target object; the image processing module 420 is configured to process the to-be-processed image based on a pre-trained stylized model to obtain a target style image of the stylized target object and the key point information of the target object after stylized processing; the special effect determination module 430 is configured to process the target style image according to a predefined target task and the key point information to obtain a target special effect corresponding to the target task.

[0108] Based on the above optional technical solutions, optionally, the image processing module 420 is specifically configured to determine the to-be-processed features of the to-be-processed image based on the encoder in the stylized model, and respectively process the to-be-processed features based on the stylized processing unit and the key point extraction unit in the stylized model to obtain the target stylized image and the key point information.

[0109] Based on the above optional technical solutions, optionally, the image processing module 420 includes: a target style image determination unit and a key point information determination unit.

[0110] The target style image determination unit is configured to process the to-be-processed feature based on at least one convolutional layer and upsampling layer in the stylization processing unit to obtain a target style image of the deformation of the target object; and,

[0111] The key point information determination unit is configured to process the to-be-processed feature based on the downsampling layer, flattening layer, and linear layer in the key point extraction unit in sequence to obtain the key point information after the deformation of the target object.

[0112] Based on the above optional technical solutions, optionally, the key point information includes the pixel coordinates of the predefined key points after the deformation of the target object; or, the key point information includes the pixel coordinates and depth coordinates of the predefined key points after the deformation of the target object.

[0113] Based on the above optional technical solutions, optionally, the target task includes the task of attaching two-dimensional special effect objects and / or three-dimensional special effect objects to the target object according to the key point information.

[0114] Based on the above optional technical solutions, optionally, the device further includes: a sample image acquisition module, a model determination module, a model processing module, and a training sample determination module.

[0115] The sample image acquisition module is configured to acquire a plurality of sample images;

[0116] The model determination module is configured to determine the three-dimensional reconstruction model corresponding to the target object in the sample image and determine the grid point information of at least one predefined key point in the three-dimensional reconstruction model;

[0117] The model processing module is configured to perform deformation processing on the three-dimensional reconstruction model according to the deformation parameters corresponding to the target style to obtain an actual stylized image and determine the actual key point information corresponding to the grid point information under the action of the deformation parameters;

[0118] The training sample determination module is configured to determine the training samples for training the stylization model according to the sample image, the actual stylized image of the sample image, and the actual key point information.

[0119] Based on the above optional technical solutions, optionally, the model processing module includes: a coordinate data determination unit and a pixel coordinate determination unit.

[0120] A coordinate data determination unit for determining the coordinate data corresponding to the grid point information under the action of the deformation parameters;

[0121] A pixel coordinate determination unit for determining the actual pixel coordinates of the grid point information in the sample image based on the coordinate data, the model matrix, the camera matrix, and the projection matrix, and using them as the actual key point information.

[0122] Based on the above optional technical solutions, optionally, the device further includes: a depth coordinate determination module.

[0123] The depth coordinate determination module is used to determine the depth coordinates of the grid point information and update the training sample based on the depth coordinates when the target task includes attaching a three-dimensional special effect object to the target object.

[0124] Based on the above optional technical solutions, optionally, the device further includes: a model training module.

[0125] The model training module is used to train the stylization model based on the training sample;

[0126] The model training module includes: a sample image processing unit and a model parameter correction unit.

[0127] The sample image processing unit is used to input the sample image in the training sample into the stylization model for stylization and key point determination processing, and output a predicted stylized image and predicted key point information;

[0128] The model parameter correction unit is used to determine a loss value based on the actual stylized image, the actual key point information, the predicted stylized image, and the predicted key point information of the training sample, and correct the model parameters in the stylization model based on the loss value until the loss function in the stylization model converges.

[0129] The technical solution of the embodiment of the present disclosure obtains a to-be-processed image including a target object. Further, the to-be-processed image is processed based on a pre-trained stylization model to obtain a target stylized image of the target object after stylization processing and key point information of the target object after stylization processing. Finally, according to a preset target task and the key point information, the target stylized image is processed to obtain a target special effect corresponding to the target task, which solves the problems in the related art that when processing an image, the special effect task cannot be accurately applied to the target object, resulting in a poor display effect of the special effect data, etc., realizes the effect of accurately determining the key point information of the target object after stylization processing, achieves the effect of deeply coupling image stylization processing and object key point detection, and further improves the accuracy and effect of key point detection. Moreover, it realizes the effect of executing the target task on the target object according to the key point information to obtain the target special effect, and improves the matching degree between the target task and the key point information of the target object.

[0130] The image processing device provided by the embodiment of the present disclosure can execute the image processing method provided by any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the method.

[0131] It should be noted that the various units and modules included in the above device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the embodiments of the present disclosure.

[0132] Figure 6 This is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Next, refer to Figure 6 , which shows a schematic structural diagram of an electronic device (such as the terminal device or server in Figure 6 ) 500 suitable for implementing the embodiment of the present disclosure. The terminal device in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 6 The electronic device shown is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present disclosure.

[0133] As Figure 6As shown, the electronic device 500 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 501, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An editing / output (I / O) interface 505 is also connected to the bus 504.

[0134] Generally, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 6 an electronic device 500 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.

[0135] Specifically, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.

[0136] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0137] The electronic device provided by the embodiment of the present disclosure and the image processing method provided by the above embodiment belong to the same inventive concept. The technical details not described in detail in this embodiment may be referred to in the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0138] The embodiment of the present disclosure provides a computer storage medium, on which a computer program is stored, and when the program is executed by a processor, the image processing method provided by the above embodiment is implemented.

[0139] It should be noted that the above-mentioned computer-readable medium in the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0140] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (Hyper Text Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed network.

[0141] The above-mentioned computer-readable medium may be included in the above-mentioned electronic device; or it may exist separately without being assembled into the electronic device.

[0142] The above computer-readable medium carries one or more programs which, when executed by the electronic device, cause the electronic device to: obtain a to-be-processed image including a target object; process the to-be-processed image based on a pre-trained stylization model to obtain a target stylized image of the target object after stylization processing and key point information of the target object after stylization processing; and process the target stylized image according to a preset target task and the key point information to obtain a target special effect corresponding to the target task.

[0143] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by connecting through the Internet using an Internet service provider).

[0144] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0145] The units described in the embodiments of the present disclosure may be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation on the unit itself in some cases. For example, the first acquisition unit may also be described as "the unit for acquiring at least two Internet protocol addresses".

[0146] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on a Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0147] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a Random Access Memory (RAM), a Read Only Memory (ROM), an Erasable Programmable Read Only Memory (EPROM or Flash memory), an optical fiber, a portable compact disc read only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0148] The above description is only a preferred embodiment of the present disclosure and an illustration of the applied technical principles. Those skilled in the art should understand that the scope of the disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, a technical solution formed by mutually replacing the above features with (but not limited to) technical features having similar functions disclosed in the present disclosure.

[0149] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0150] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms for implementing the claims.

Claims

1. An image processing method, characterized in that, Including: Obtaining a to-be-processed image including a target object; Processing the to-be-processed image based on a pre-trained stylization model to obtain a target stylized image of stylization processing for the target object and key point information of the target object after stylization processing; Processing the target stylized image according to a preset target task and the key point information to obtain a target special effect corresponding to the target task.

2. The method according to claim 1, characterized in that, The processing the to-be-processed image based on a pre-trained stylization model to obtain a target stylized image of stylization processing for the target object and key point information of the target object after stylization processing includes: Determining to-be-processed features of the to-be-processed image based on an encoder in the stylization model, and respectively processing the to-be-processed features based on a stylization processing unit and a key point extraction unit in the stylization model to obtain the target stylized image and the key point information.

3. The method according to claim 2, characterized in that, The respectively processing the to-be-processed features based on a stylization processing unit and a key point extraction unit in the stylization model to obtain the target stylized image and the key point information includes: Processing the to-be-processed features based on at least one convolutional layer and an upsampling layer in the stylization processing unit to obtain a target stylized image of deformation of the target object; and Processing the to-be-processed features in sequence based on a downsampling layer, a flattening layer and a linear layer in the key point extraction unit to obtain key point information after deformation of the target object.

4. The method according to any one of claims 1-3, characterized in that, The key point information includes pixel coordinates of predefined key points after deformation of the target object; or, the key point information includes pixel coordinates and depth coordinates of predefined key points after deformation of the target object.

5. The method according to claim 1, characterized in that The target task includes a task of attaching a two-dimensional special effect object and / or a three-dimensional special effect object to the target object according to the key point information.

6. The method according to claim 1, wherein Also including: Obtaining a plurality of sample images; Determining a three-dimensional reconstruction model corresponding to a target object in the sample images, and determining grid point information of at least one predefined key point in the three-dimensional reconstruction model; Deforming the three-dimensional reconstruction model according to deformation parameters corresponding to a target style to obtain an actual stylized image, and determining actual key point information corresponding to the grid point information under the action of the deformation parameters; Determining training samples for training the stylization model according to the sample images, the actual stylized images of the sample images and the actual key point information.

7. The method according to claim 6, wherein The determining actual key point information corresponding to the grid point information under the action of the deformation parameters includes: Determining coordinate data corresponding to the grid point information under the action of the deformation parameters; Determining actual pixel coordinates of the grid point information in the sample images according to the coordinate data, a model matrix, a camera matrix and a projection matrix, and using the actual pixel coordinates as the actual key point information.

8. The method according to claim 6, wherein Also including: In a case where the target task includes attaching a three-dimensional special effect object to the target object, determining depth coordinates of the grid point information, and updating the training samples based on the depth coordinates.

9. The method according to claim 6, characterized in that Also including: Train the stylization model based on the training samples; The training of the stylization model based on the training samples includes: Input the sample images in the training samples into the stylization model for stylization and key point determination processing, and output the predicted stylized images and predicted key point information; Determine the loss value according to the actual stylized images, actual key point information, predicted stylized images and predicted key point information of the training samples, and correct the model parameters in the stylization model based on the loss value until the loss function in the stylization model converges.

10. An image processing apparatus, characterized in that, It includes: An image acquisition module for acquiring a to-be-processed image including a target object; An image processing module for processing the to-be-processed image based on a pre-trained stylization model to obtain a target stylized image of the target object after stylization and key point information of the target object after stylization; An special effect determination module for processing the target stylized image according to a preset target task and the key point information to obtain a target special effect corresponding to the target task.

11. An electronic device, characterized in that, The electronic device includes: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the image processing method according to any one of claims 1-9.

12. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions are used to execute the image processing method according to any one of claims 1-9 when executed by a computer processor.