Image processing method and device, model training method and device, medium and program product

By acquiring images and generating target description text from deformation labels, and training a model using the ControlNet network, the problem of distortion in deformation processing in human body beautification algorithms is solved, thereby improving the human body beautification effect and making image processing more convenient.

CN120833255APending Publication Date: 2025-10-24BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410495412.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-23
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Existing human body beautification algorithms suffer from distortion during deformation processing, and the complexity of human body postures and the diversity of clothing and accessory textures increase the difficulty of beautification editing.

Method used

By acquiring the original image and labels of deformation capabilities, target description text is generated, and a target deformation processing model is trained using the ControlNet network. Combined with the image description generation module and the deformation combination module, the free combination of human body beautification effects can be achieved.

Benefits of technology

It improves the reliability of human body beautification effects and the convenience of image processing, and can quickly generate deformation results that satisfy users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120833255A_ABST
    Figure CN120833255A_ABST
Patent Text Reader

Abstract

The invention relates to an image processing method and device, a model training method and device, a medium and a program product, and the image processing method comprises the steps: obtaining an original image and a label used for representing the deformation capability, and the original image comprises an object which deforms according to the deformation capability; acquiring an image description text of the original image; generating a target description text based on the image description text and the label; generating a target image based on the original image and the target description text, wherein the target image comprises the object deformed according to the deformation capability; according to the technical scheme, the deformation effect after deformation capability processing can be obtained through the original image and the label used for representing the deformation capability, so that reliable technical support can be provided for human body beautification, the human body beautification effect can be further improved, and the convenience of image processing can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of image processing, and in particular, to an image processing method, a model training method, an apparatus, a medium and a program product. BACKGROUND

[0002] The beautification of the human body has always been one of the focuses of photo editing software. However, compared with face beautification, the posture of the human body is more complex, and the texture of clothing and accessories is more diverse, so it is more difficult to use a neural network to edit the human body. Therefore, there is no mature body beautification solution at present, and the existing body beautification algorithms in the related art also mostly have the problem of effect distortion in the morphing process. SUMMARY

[0003] This summary is provided to introduce a selection of concepts, which will be described in more detail in the detailed description section. This summary is not intended to identify key or essential features of the claimed technology, nor is it intended to limit the scope of the claimed technology.

[0004] In a first aspect, the present disclosure provides an image processing method, comprising:

[0005] obtaining an original image, and obtaining an original image, and a label for representing a morphing capability, the original image comprising an object to be morphed according to the morphing capability;

[0006] obtaining an image description text of the original image;

[0007] generating a target description text based on the image description text and the label;

[0008] generating a target image based on the original image and the target description text, the target image comprising the object after morphing according to the morphing capability.

[0009] In a second aspect, the present disclosure provides a model training method, comprising:

[0010] obtaining a plurality of sets of morphing image samples, each set of the morphing image samples comprising an original sample image, at least one label, and a morphing effect image of the original sample image corresponding to the at least one label;

[0011] training a first preset initial ControlNet (a kind of image morphing algorithm based on control points) network with the plurality of sets of morphing image samples as training data to obtain a target morphing processing model.

[0012] In a third aspect, the present disclosure provides a model training method, comprising:

[0013] a plurality of groups of first image sample data, each group of the first image sample data comprising an original sample image and description text label data corresponding to the original sample image;

[0014] training, as training data, a plurality of groups of first image sample data, each group of the first image sample data comprising an original sample image and description text label data corresponding to the original sample image, to obtain an image description generation module;

[0015] a plurality of groups of second image sample data, each group of the second image sample data comprising an original sample image, label text data, and a morphing effect image corresponding to the original sample image, the label text data comprising description text label data corresponding to the original sample image and at least one label;

[0016] training, as training data, a plurality of groups of second image sample data, each group of the second image sample data comprising an original sample image, label text data, and a morphing effect image corresponding to the original sample image, the label text data comprising description text label data corresponding to the original sample image and at least one label, to obtain a morphing combination module;

[0017] combining the image description generation module and the morphing combination module into a target morphing processing model.

[0018] In a fourth aspect, the present disclosure provides an image processing device, the device comprising:

[0019] a first obtaining module configured to obtain an original image and a label for indicating a morphing capability, the original image comprising an object to be morphed according to the morphing capability;

[0020] a determining module configured to obtain an image description text of the original image, generate a target description text based on the image description text and the label, and generate a target image based on the original image and the target description text, the target image comprising the object morphed according to the morphing capability.

[0021] In a fifth aspect, the present disclosure provides a model training device, the device comprising:

[0022] a second obtaining module configured to obtain a plurality of groups of morphing image samples, each group of the morphing image samples comprising an original sample image, at least one label, and a morphing effect image of the original sample image corresponding to the at least one label;

[0023] a first training module configured to train, as training data, a plurality of groups of morphing image samples, each group of the morphing image samples comprising an original sample image, at least one label, and a morphing effect image of the original sample image corresponding to the at least one label, to obtain the target morphing processing model.

[0024] In a sixth aspect, the present disclosure provides a model training device, the device comprising:

[0025] The third obtaining module is configured to obtain a plurality of groups of first image sample data, each group of the first image sample data comprising an original sample image and description text label data corresponding to the original sample image;

[0026] The second training module is configured to train a preset initial image description algorithm by taking the plurality of groups of first image sample data as training data, to obtain the image description generation module.

[0027] The fourth obtaining module is configured to obtain a plurality of groups of second image sample data, each group of the second image sample data comprising an original sample image, label text data, and a morphing effect image corresponding to the original sample image, the label text data comprising description text label data corresponding to the original sample image and at least one label.

[0028] The second training module is configured to train a second preset initial ControlNet network by taking the plurality of groups of second image sample data as training data, to obtain the morphing combination module.

[0029] The coupling module is configured to combine the image description generation module and the morphing combination module into the target morphing processing model.

[0030] In a seventh aspect, the present disclosure provides a computer readable medium having a computer program stored thereon, the program being executed by a processing device to implement the steps of the method of the first aspect, the second aspect, or the third aspect.

[0031] In an eighth aspect, the present disclosure provides an electronic device comprising:

[0032] A storage device having a computer program stored thereon;

[0033] A processing device configured to execute the computer program in the storage device to implement the steps of the method of the first aspect, the second aspect, or the third aspect.

[0034] In a ninth aspect, the present disclosure provides a computer program product comprising a computer program, the computer program being executed by a processor to implement the steps of the method of the first aspect, the second aspect, or the third aspect.

[0035] The technical scheme above, by acquiring an original image and a label used for representing a deformation capability, the original image including an object to be deformed according to the deformation capability; acquiring an image description text of the original image; generating a target description text based on the image description text and the label; and generating a target image based on the original image and the target description text, the target image including the object after being deformed according to the deformation capability, the deformation effect after being processed by the deformation capability can be obtained through the original image and the label used for representing the deformation capability, thereby providing reliable technical support for human beautification, which is beneficial to further improve the human beautification effect and effectively improve the convenience of image processing.

[0036] Other features and advantages of the present disclosure will be described in detail in the following detailed description section. BRIEF DESCRIPTION OF DRAWINGS

[0037] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent by describing in detail exemplary embodiments thereof with reference to the attached drawings in which:

[0038] Figure 1 is a flowchart of an image processing method according to an exemplary embodiment of the present disclosure;

[0039] Figure 2 is a flowchart of an image processing method according to an exemplary embodiment of the present disclosure; Figure 1

[0040] Figure 3 is a flowchart of a model training method according to an exemplary embodiment of the present disclosure;

[0041] Figure 4 is a flowchart of a model training method according to another exemplary embodiment of the present disclosure;

[0042] Figure 5 is a block diagram of an image processing apparatus according to an exemplary embodiment of the present disclosure;

[0043] Figure 6 is a block diagram of a model training apparatus according to an exemplary embodiment of the present disclosure;

[0044] Figure 7 is a block diagram of a model training apparatus according to another exemplary embodiment of the present disclosure;

[0045] Figure 8 is a block diagram of an electronic device according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION​

[0046] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein; rather, these embodiments are provided so that the present disclosure can be more thoroughly and completely understood.

[0047] It should be understood that each of the steps recited in the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0048] As used herein, the term "comprises" and its variations are intended to mean "including but not limited to". The term "based on" means "based, at least in part, on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Related definitions will be given in the description below.

[0049] It should be noted that the terms "first", "second", and the like in the present disclosure are only used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units.

[0050] It should be noted that the terms "one", "multiple" in the present disclosure are illustrative and not restrictive, and those skilled in the art should understand that unless otherwise explicitly stated in the context, it should be understood as "one or more".

[0051] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.

[0052] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type, scope of use, use scenario, etc. of the personal information involved in the present disclosure should be informed to the user and the authorization of the user should be obtained in accordance with relevant laws and regulations.

[0053] For example, in response to receiving an active request of a user, a prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will need to acquire and use personal information of the user. Thus, the user can autonomously select whether to provide the personal information to the software or hardware such as an electronic device, an application program, a server or a storage medium, etc. performing the operation of the technical solution of the present disclosure according to the prompt information.

[0054] As an optional but non-limiting implementation manner, in response to receiving an active request of a user, the manner of sending a prompt information to the user may, for example, be a pop-up window manner, in which the prompt information can be presented in a text manner. In addition, the pop-up window can also carry a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0055] It can be understood that the above notification and acquisition of user authorization process is only illustrative and does not limit the implementation manner of the present disclosure, and other manners meeting relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0056] Meanwhile, it can be understood that the data (including but not limited to the data itself, acquisition or use of the data) involved in the technical solution should comply with the requirements of relevant laws and regulations and relevant provisions.

[0057] Figure 1 is a flowchart of an image processing method according to an example embodiment of the present disclosure; as shown in Figure 1 the image processing method can include:

[0058] In step 101, an original image and a label for representing a deformation capability are acquired.

[0059] One or more labels can be acquired, each of the labels can be an identifier of a deformation capability, and the original image includes a specified deformation object, which can be understood as an object to be deformed according to the deformation capability.

[0060] It should be noted that the specified morphing object can be a human body image, can also be a human face, can also be an animal body or other specified object, in the case of the specified morphing object being a human body, the morphing capability can be lengthening the neck, slimming the waist, lengthening the legs, slimming the legs, lengthening the arms, slimming the arms, adjusting the ratio of the upper body and the lower body of the body, etc., wherein the identifier of each morphing capability can be any specified symbol, which can be regarded as a Token, different Tokens can be authorized different morphing capability operation permissions, for example, the identifier of slimming the arms can be SBBX, &slim-arm&, 001 or “slim-arm”, etc., the identifier of lengthening the neck can be “swan neck”, TEJ or 002, or other specified letters, numbers, words and other symbols. In the case of the specified morphing object being a human face, the morphing capability can be whitening, double eyelid, eyebrow shape modification, lip shape modification, etc., and the label can be the corresponding Token.

[0061] Step 102, obtaining an image description text of the original image.

[0062] Wherein, the image description text is used to describe the image content in the original image, for example, the image description text of an original image can be “A girl in a red dress is dancing”.

[0063] Step 103, generating a target description text based on the image description text and the label.

[0064] Wherein, the image description text and one or more labels for representing morphing capabilities can be spliced to obtain the target description text.

[0065] It should be noted that the splicing can be performed according to a pre-set splicing format, for example, a specified separator symbol can be added between different labels, for example, the separator symbol can be a space, a comma, a semicolon, a vertical line, etc.

[0066] For example, if the image description text is “A girl in a red dress is dancing”, the user input label is “swan neck” and “&slim-arm&”, and the target description text can be “A girl in a red dress is dancing swan neck &slim-arm&”.

[0067] Step 104, generating a target image based on the original image and the target description text, the target image comprising the object after morphing according to the morphing capability.

[0068] The target image includes an object after deformation processing according to deformation capability corresponding to one or more labels used to represent the deformation capability.

[0069] In an implementation, the original image and the target description text can be input into a preset deformation combination network to obtain the target image output by the deformation combination module. The deformation combination network can present deformation effects after free combination of deformation capability corresponding to one or more labels used to represent the deformation capability.

[0070] In another implementation, the above steps 102 to 104 can be implemented by inputting the original image and one or more labels used to represent deformation capability into a preset target deformation processing model to obtain a target image output by the target deformation processing model.

[0071] The target deformation processing model can be a neural network structure model based on a ControlNet algorithm.

[0072] For example, when the labels input by the user are a label for lengthening the neck, a label for lengthening the legs, and a label for thinning the arms, the human body in the target image can be an object after lengthening the neck, lengthening the legs, and thinning the arms.

[0073] The training of the target deformation processing model can have two implementations. In a first implementation, the training process of the target deformation processing model can include: obtaining a plurality of sets of deformation image samples, each set of deformation image samples including an original sample image, at least one label, and a deformation effect image of the original sample image corresponding to the at least one label; using the plurality of sets of deformation image samples as training data, training a first preset initial ControlNet network to obtain the target deformation processing model.

[0074] It should be noted that, due to the zero-shot learning capability of the ControlNet network itself, existing training data can be fully utilized to identify and process classes or tasks that have not been seen before. Therefore, in the training process of the target deformation processing model, the model does not need to be trained for all label combinations, and can directly use existing deformation capability combinations to predict other combinations. Generally, when the plurality of sets of deformation image samples collectively include n deformation capability labels, the target deformation processing model obtained by training can have 2n nThe target morphing processing model can present the morphing effect after freely combining the morphing capabilities corresponding to the one or more labels for representing morphing capabilities, thereby effectively improving the combination efficiency of various morphing capabilities of the model, thereby facilitating rapid acquisition of image processing effects satisfactory to the user.

[0075] In the second implementation, the training process of the target morphing processing model can include: obtaining a plurality of groups of first image sample data, each group of the first image sample data including an original sample image and description text annotation data corresponding to the original sample image; training a preset initial image description algorithm using the plurality of groups of first image sample data as training data to obtain the image description generation module; obtaining a plurality of groups of second image sample data, each group of the second image sample data including an original sample image, annotation text data, and a morphing effect image corresponding to the original sample image, the annotation text data including description text annotation data corresponding to the original sample image and at least one label; training a second preset initial ControlNet network using the plurality of groups of second image sample data as training data to obtain the morphing combination module; and combining the image description generation module and the morphing combination module into the target morphing processing model.

[0076] The description text annotation data is a description text of image content in a specified language, for example, the description text annotation data for a certain image can be "a girl is dancing ballet". The first preset initial ControlNet network and the second preset initial ControlNet network can be the same or different, for example, the first preset initial ControlNet network and the second preset initial ControlNet network can be ControlNet algorithms with the same hyperparameters, or two ControlNet algorithms with some or all different hyperparameters.

[0077] In addition, it should be noted that if the plurality of groups of second image sample data collectively include n labels of morphing capabilities, the target morphing processing model obtained by training can have the ability to present 2 n combinations.

[0078] The above technical solutions can obtain a morphing effect formed by combining morphing capabilities corresponding to one or more labels for representing morphing capabilities by inputting an original image and one or more labels, thereby providing reliable technical support for human beautification, facilitating further improvement of human beautification effect, and effectively improving the convenience of image processing.

[0079] Figure 2is according to Figure 1 A flowchart of an image processing method according to the embodiment shown in FIG. 1 can optionally include the following steps: Figure 1 In an embodiment, the steps 102 to 104 can be implemented by inputting the original image and one or more labels representing the deformation capability into a target deformation processing model, which can include an image description generation module and a deformation combination module coupled with the image description generation module. Based on the target deformation processing model, an image description text of the original image is first obtained, and then a target description text is generated based on the image description text and the labels. Then, a target image is generated based on the original image and the target description text. The specific implementation is as shown in FIG. 2. Figure 2 The embodiment shown in FIG. 2 can include the following steps:

[0080] In step 1021, the original image is input into the image description generation module to obtain an image description text corresponding to the original image output by the image description generation module.

[0081] In this step, the image description text corresponding to the original image can be generated by the image description generation module. The image description generation module can be a network based on NIC (Neural Image Caption) algorithm, which is used to convert image content into natural language description. It can also be a visual attention-based image description network, a SCA-CNN (Spatial and Channel-wise Attention in Convolutional Networks) based image description network, or a Transformer-based image description generation algorithm. It can also be other neural network algorithms in the prior art.

[0082] In step 1022, the image description text and the one or more labels representing the deformation capability are spliced to obtain a target description text.

[0083] In the splicing process, the splicing format can be set in advance, for example, a specified separator symbol can be added between different labels, such as space, comma, semicolon, vertical bar, etc.

[0084] For example, if the image description text is "A girl in a red dress is dancing", and the user input labels are "swan neck" and "&slim-arm&", the target description text can be "A girl in a red dress is dancing swan neck &slim-arm&".

[0085] Step 1023, input the target description text and the original image into the morphing combination module to obtain the target image output by the morphing combination module.

[0086] Wherein, the morphing combination module can output the morphing effect after combining the morphing capabilities corresponding to the multiple labels under the condition that the user inputs the multiple labels.

[0087] It should be noted that under the condition that the user inputs the one or more labels for representing morphing capabilities, the target morphing processing model can present the morphing effect after freely combining the morphing capabilities corresponding to the one or more labels for representing morphing capabilities, which can effectively improve the combination efficiency of various morphing capabilities of the model, thereby facilitating the rapid obtaining of the image processing effect satisfactory to the user.

[0088] The above technical solution can splice the image description text and the one or more labels for representing morphing capabilities to obtain a target description text, input the target description text and the original image into the morphing combination module to obtain the target image output by the morphing combination module, and can improve the accurate recognition of the model on the image content through the input of the image description text, thereby providing further data basis for the morphing processing and morphing capability combination of the model, and facilitating the further improvement of the accuracy and reliability of the output result of the model.

[0089] Figure 3 is a flowchart of a model training method according to an example embodiment of the present disclosure, as shown in Figure 3 The method can include:

[0090] S11, obtaining multiple sets of morphing image samples, each set of morphing image samples including an original sample image, at least one label, and a morphing effect image of the original sample image corresponding to the at least one label.

[0091] Wherein, each set of morphing image samples can include one label or multiple labels, and in the case of including multiple labels, the morphing effect image of the original sample image is an effect image processed by multiple morphing capabilities corresponding to the multiple labels.

[0092] It should be noted that during the training process, the original sample image is used as the control condition of the ControlNet network, the at least one label is used as the text input of the network, and the deformation effect diagram of the original sample image corresponding to the at least one label is used to supervise the ControlNet network. Due to the zero-sample learning ability of the ControlNet network itself, it can make full use of the existing training data to identify and process those categories or tasks that have not been seen. Therefore, even if there is no combination of labels in the training data, the target deformation processing model can directly use the existing deformation ability combination to predict other combinations. Usually, when the multiple groups of deformable image samples include a total of n types of deformation ability labels, the trained target deformation processing model can have the ability to present 2 n Ability to combine.

[0093] S12: Using the multiple groups of deformed image samples as training data, training a first preset initial ControlNet network to obtain a target deformation processing model.

[0094] Among them, the target deformation processing model can present the deformation effect after freely combining the deformation capabilities corresponding to the one or more labels for representing the deformation capabilities when the user inputs the one or more labels for representing the deformation capabilities, which can effectively improve the model's combination efficiency of various deformation capabilities, thereby facilitating the rapid acquisition of user-satisfied image processing effects.

[0095] For example, when multiple sets of deformable image samples include labels indicating neck lengthening, leg lengthening, and arm thinning, the trained target deformation processing model can be capable of any combination of the three deformation capabilities of neck lengthening, leg lengthening, and arm thinning. That is, the target deformation processing model can provide images processed with neck lengthening, leg lengthening, and arm thinning, or deformed objects processed with a combination of neck lengthening and leg lengthening, or a combination of neck lengthening and arm thinning, or deformed objects processed with leg lengthening and arm thinning, or images processed with only one of the three deformation capabilities. The user's needs can be determined by input labels. For example, if the user inputs neck lengthening and leg lengthening, the user's need is determined to be neck lengthening and leg lengthening; if the user inputs leg lengthening, the user's need is determined to be leg lengthening only.

[0096] It should be noted that the specific implementation of training the first preset initial ControlNet network with the plurality of sets of deformation image samples as training data can refer to the training process of the ControlNet network in the prior art, and the present disclosure does not limit this.

[0097] The above technical solution can train a model with more deformation capability combinations by a limited label combination mode in limited deformation image samples, and the target deformation processing model can realize that a user obtains an image with a deformation capability combination corresponding to a plurality of labels by inputting the plurality of labels, thereby effectively improving the convenience of image processing and facilitating further improvement of the user experience of the model.

[0098] Figure 4 is a flowchart of a model training method according to another example embodiment of the present disclosure, as shown in Figure 4 The method can include the following steps:

[0099] S21, a plurality of sets of first image sample data are obtained, and each set of the first image sample data includes an original sample image and description text annotation data corresponding to the original sample image.

[0100] S22, an initial image description algorithm is trained with the plurality of sets of first image sample data as training data, to obtain an image description generation module.

[0101] The initial image description algorithm can be an initial NIC algorithm, an initial Visual Attention image description network, an initial SCA-CNN algorithm, or an initial Transformer algorithm, or other initial neural network algorithms in the prior art.

[0102] It should be noted that the specific implementation of training the first preset initial ControlNet network with the plurality of sets of deformation image samples as training data can refer to the training process of the ControlNet network in the prior art, and the present disclosure does not limit this.

[0103] S23, a plurality of sets of second image sample data are obtained, and each set of the second image sample data includes an original sample image, annotation text data, and a deformation effect image corresponding to the original sample image, wherein the annotation text data includes description text annotation data corresponding to the original sample image and at least one label.

[0104] The at least one label in the annotation text data can be pre-annotated by a user, and the annotation text data can splice the pre-annotated label of the user and the description text generated by the image description generation module in the S22 step.

[0105] For example, if the description text generated by the image description generation module is "a girl is dancing ballet", and the labels input by the user are "swan neck", "&slim-arm&", and "leg lengthening", the annotation text data can be "a girl is dancing ballet, swan neck, &slim-arm&, leg lengthening".

[0106] S24, training the second preset initial ControlNet network with the plurality of sets of second image sample data as training data to obtain a morphing combination module.

[0107] The first preset initial ControlNet network and the second preset initial ControlNet network can be the same or different, for example, the first preset initial ControlNet network and the second preset initial ControlNet network can be ControlNet algorithms with the same hyperparameters, or two ControlNet algorithms with some or all different hyperparameters.

[0108] S25, combining the image description generation module and the morphing combination module into a target morphing processing model.

[0109] In this step, the output end of the image description generation module can be coupled with the first input end of the morphing combination module, the input end of the image description generation module and the second input end of the morphing combination module can be used as the input end of the target morphing processing model, and the output end of the morphing combination module can be used as the output end of the target morphing processing model.

[0110] It should be noted that if the plurality of sets of second image sample data includes n kinds of morphing ability labels, the target morphing processing model obtained by training can have a combination of 2 n It can be understood that even if the combination mode of the label does not appear in the training data, the target morphing processing model can directly use the existing morphing ability combination to predict other combinations. In this way, when the user inputs one or more labels representing morphing ability, the target morphing processing model can present the morphing effect after freely combining the morphing ability corresponding to the one or more labels representing morphing ability, thereby effectively improving the model's combination efficiency for various morphing abilities, and providing reliable data basis for quickly obtaining user-satisfactory images.

[0111] It should be noted that the plurality of sets of second image sample data as training data, and the specific implementation of training the second preset initial ControlNet network can refer to the training process of the ControlNet network in the prior art, and the present disclosure does not limit it.

[0112] The above technical solution integrates the image description generation module in the target morphing processing model, so that the target morphing processing model can further improve the accuracy of image content recognition through the input of the image description text, thereby providing further data basis for the combination of morphing processing and morphing ability of the model, and facilitating further improvement of the accuracy and reliability of the model output result.

[0113] Figure 5 is a block diagram of an image processing device according to an example embodiment of the present disclosure; as Figure 5 shown, the image processing device can include:

[0114] The first acquisition module 501 is configured to acquire an original image and a label for representing morphing ability, the original image including an object to be morphed according to the morphing ability.

[0115] The determination module 502 is configured to acquire an image description text of the original image, generate a target description text based on the image description text and the label, and generate a target image based on the original image and the target description text, the target image including the object after morphing according to the morphing ability.

[0116] The above technical solution can obtain a morphing effect formed by the combination of one or more morphing abilities corresponding to one or more labels for representing morphing ability through input of an original image and one or more labels, thereby providing reliable technical support for human beautification, and facilitating further improvement of human beautification effect.

[0117] Optionally, the training process of the target morphing processing model can include:

[0118] Acquiring a plurality of sets of morphing image samples, each set of morphing image samples including an original sample image, at least one label, and a morphing effect image of the original sample image corresponding to the at least one label;

[0119] Using the plurality of sets of morphing image samples as training data, a first preset initial ControlNet network is trained to obtain the target morphing processing model.

[0120] Optionally, the target morphing processing model includes an image description generation module and a morphing combination module coupled with the image description generation module.

[0121] The determination module 502 is configured to:

[0122] inputting the original image and one or more labels for representing the deformation capability into a target deformation processing model, the target deformation processing model comprising an image description generation module, and a deformation combination module coupled with the image description generation module;

[0123] generating an image description text corresponding to the original image by the image description generation module;

[0124] splicing the image description text and the one or more labels for representing the deformation capability to obtain a target description text;

[0125] inputting the target description text and the original image into the deformation combination module to obtain the target image output by the deformation combination module.

[0126] Optionally, the target deformation processing model is obtained by training in the following manner:

[0127] obtaining a plurality of groups of first image sample data, each group of the first image sample data comprising an original sample image and description text annotation data corresponding to the original sample image;

[0128] training a preset initial image description algorithm by taking the plurality of groups of first image sample data as training data to obtain the image description generation module;

[0129] obtaining a plurality of groups of second image sample data, each group of the second image sample data comprising an original sample image, annotation text data, and a deformation effect image corresponding to the original sample image, the annotation text data comprising description text annotation data corresponding to the original sample image and at least one label;

[0130] training a second preset initial ControlNet network by taking the plurality of groups of second image sample data as training data to obtain the deformation combination module;

[0131] combining the image description generation module and the deformation combination module into the target deformation processing model.

[0132] Optionally, the combining the image description generation module and the deformation combination module into the target deformation processing model comprises:

[0133] coupling an output end of the image description generation module with a first input end of the deformation combination module, an input end of the image description generation module and a second input end of the deformation combination module serving as input ends of the target deformation processing model, and an output end of the deformation combination module serving as an output end of the target deformation processing model.

[0134] The above technical scheme integrates the image description generation module in the target morphing processing model, so that the target morphing processing model can further improve the accuracy of image content recognition through the input of the image description text, thereby providing further data basis for the combination of morphing processing and morphing capability of the model, and facilitating the further improvement of the accuracy and reliability of the model output result.

[0135] Figure 6 is a block diagram of a model training device according to an example embodiment of the present disclosure, which can include:

[0136] The second acquisition module 601 is configured to acquire a plurality of sets of morphing image samples, each set of the morphing image samples including an original sample image, at least one label, and a morphing effect image corresponding to the at least one label of the original sample image.

[0137] The first training module 602 is configured to train a first preset initial ControlNet network using the plurality of sets of morphing image samples as training data, to obtain a target morphing processing model.

[0138] The above technical scheme can train a model that presents more combinations of morphing capabilities using a limited number of label combination modes in limited morphing image samples, and the target morphing processing model can realize that a user inputs a plurality of labels to obtain an image with a combination of morphing capabilities corresponding to the plurality of labels, thereby effectively improving the convenience of image processing and facilitating the further improvement of the user experience of the model.

[0139] Figure 7 is a block diagram of a model training device according to another example embodiment of the present disclosure, which can include:

[0140] The third acquisition module 701 is configured to acquire a plurality of sets of first image sample data, each set of the first image sample data including an original sample image and description text annotation data corresponding to the original sample image.

[0141] The second training module 702 is configured to train a preset initial image description algorithm using the plurality of sets of first image sample data as training data, to obtain the image description generation module.

[0142] The fourth acquisition module 703 is configured to acquire a plurality of sets of second image sample data, each set of the second image sample data including an original sample image, annotation text data, and a morphing effect image corresponding to the original sample image, the annotation text data including description text annotation data corresponding to the original sample image and at least one label.

[0143] A second training module 704 is configured to train a second preset initial ControlNet network using the plurality of sets of second image sample data as training data to obtain a deformation combination module;

[0144] The coupling module 705 is configured to combine the image description generation module and the deformation combination module into a target deformation processing model.

[0145] Optionally, the coupling module 705 is configured to:

[0146] The output end of the image description generation module is coupled to the first input end of the deformation combination module, the input end of the image description generation module and the second input end of the deformation combination module serve as the input end of the target deformation processing model, and the output end of the deformation combination module serves as the output end of the target deformation processing model.

[0147] The above technical solution, by integrating the image description generation module into the target deformation processing model, enables the target deformation processing model to further improve the accuracy of image content recognition through the input of the image description text, thereby providing further data basis for the combination of the model's deformation processing and deformation capabilities, which is conducive to further improving the accuracy and reliability of the model's output results.

[0148] Reference below Figure 8 , which shows a schematic structural diagram of an electronic device 800 suitable for implementing the embodiments of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 8 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0149] like Figure 8 As shown, the electronic device 800 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the electronic device 800 are also stored in the RAM 803. The processing device 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0150] In general, the following devices can be connected to the I / O interface 805: input devices 806, including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; output devices 807, including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; storage devices 808, including, for example, a magnetic tape, a hard disk, and the like; and communication devices 809. The communication devices 809 can allow the electronic device 800 to communicate wirelessly or wired with other devices to exchange data. Although Figure 8 The electronic device 800 is shown with various devices, but it is understood that all of the shown devices are not required to be implemented or present. More or less devices can alternatively be implemented or present.

[0151] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program according to embodiments of the present disclosure. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 809, or installed from the storage devices 808, or installed from the ROM 802. When the computer program is executed by the processing devices 801, the above-mentioned functions defined in the methods of embodiments of the present disclosure are performed.

[0152] It should be noted that the computer-readable medium described above can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in baseband or propagated as a carrier wave in a propagated data signal, in which the computer-readable program code is contained. Such a propagated data signal can take many forms, including but not limited to, an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium that can send, propagate or transfer the program for use by or in connection with the instruction execution system, apparatus or device. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to, wire, cable, RF (radio frequency), etc., or any suitable combination of the above.

[0153] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.

[0154] The computer-readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device, and is not assembled into the electronic device.

[0155] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to: acquire an original image and a label for representing a deformation capability, the original image comprising an object to be deformed according to the deformation capability; acquire an image description text of the original image; generate a target description text based on the image description text and the label; and generate a target image based on the original image and the target description text, the target image comprising the object deformed according to the deformation capability.

[0156] Computer program code for carrying out operations of the present disclosure can be written in any one or more programming languages, including object oriented programming languages such as Java, Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0157] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or operations, or combinations of special purpose hardware and computer instructions.

[0158] The modules described in the embodiments of the present disclosure can be implemented by software, or by hardware. In some cases, the name of a module does not constitute a limitation on the module itself, for example, a first acquisition module can also be described as an "acquisition of an original image and a label for representing a deformation capability".

[0159] The functionality described herein above can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, an example type of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.

[0160] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0161] According to one or more embodiments of the present disclosure, example 1 provides an image processing method, the method comprising: obtaining an original image, and a label for representing a deformation capability, the original image comprising an object to be deformed according to the deformation capability; obtaining an image description text of the original image; generating a target description text based on the image description text and the label; generating a target image based on the original image and the target description text, the target image comprising the object deformed according to the deformation capability.

[0162] According to one or more embodiments of the present disclosure, example 2 provides the image processing method of example 1, the obtaining an image description text of the original image, the generating a target description text based on the image description text and the label, and the generating a target image based on the original image and the target description text are implemented by a target deformation processing model, a training process of the target deformation processing model can comprise:

[0163] obtaining a plurality of sets of deformation image samples, each set of the deformation image samples comprising an original sample image, at least one label, and a deformation effect image of the original sample image corresponding to the at least one label;

[0164] The first preset initial ControlNet network is trained by taking the plurality of sets of deformation image samples as training data, so as to obtain the target deformation processing model.

[0165] According to one or more embodiments of the present disclosure, example 3 provides the image processing method of example 1, and the image description text of the original image is obtained; the target description text is generated based on the image description text and the label, and the target image is generated based on the original image and the target description text, comprising:

[0166] The original image and one or more labels for representing deformation capability are input into a target deformation processing model, the target deformation processing model comprising an image description generation module and a deformation combination module coupled with the image description generation module;

[0167] The image description generation module generates an image description text corresponding to the original image;

[0168] The image description text and the one or more labels for representing deformation capability are spliced to obtain a target description text;

[0169] The target description text and the original image are input into the deformation combination module to obtain the target image output by the deformation combination module.

[0170] According to one or more embodiments of the present disclosure, example 4 provides the image processing method of example 3, and the target deformation processing model is obtained by the following method:

[0171] A plurality of sets of first image sample data are obtained, each set of the first image sample data comprising an original sample image and description text annotation data corresponding to the original sample image;

[0172] The preset initial image description algorithm is trained by taking the plurality of sets of first image sample data as training data, so as to obtain the image description generation module;

[0173] A plurality of sets of second image sample data are obtained, each set of the second image sample data comprising an original sample image, annotation text data, and a deformation effect image corresponding to the original sample image, the annotation text data comprising description text annotation data corresponding to the original sample image and at least one label;

[0174] The second preset initial ControlNet network is trained by taking the plurality of sets of second image sample data as training data, so as to obtain the deformation combination module;

[0175] The image description generation module and the deformation combination module are combined into the target deformation processing model.

[0176] According to one or more embodiments of the present disclosure, example 5 provides the image processing method of example 4, wherein the image description generation module and the morphing combination module are combined into the target morphing processing model, including:

[0177] The output end of the image description generation module is coupled with the first input end of the morphing combination module, the input end of the image description generation module and the second input end of the morphing combination module are input ends of the target morphing processing model, and the output end of the morphing combination module is an output end of the target morphing processing model.

[0178] According to one or more embodiments of the present disclosure, example 6 provides a model training method, including:

[0179] Obtaining a plurality of sets of morphing image samples, each set of the morphing image samples including an original sample image, at least one label, and a morphing effect image corresponding to the at least one label of the original sample image;

[0180] Training a first preset initial ControlNet network using the plurality of sets of morphing image samples as training data to obtain a target morphing processing model.

[0181] According to one or more embodiments of the present disclosure, example 7 provides a model training method, including:

[0182] Obtaining a plurality of sets of first image sample data, each set of the first image sample data including an original sample image and description text annotation data corresponding to the original sample image;

[0183] Training a preset initial image description algorithm using the plurality of sets of first image sample data as training data to obtain an image description generation module;

[0184] Obtaining a plurality of sets of second image sample data, each set of the second image sample data including an original sample image, annotation text data, and a morphing effect image corresponding to the original sample image, the annotation text data including description text annotation data and at least one label corresponding to the original sample image;

[0185] Training a second preset initial ControlNet network using the plurality of sets of second image sample data as training data to obtain a morphing combination module;

[0186] Combining the image description generation module and the morphing combination module into a target morphing processing model.

[0187] According to one or more embodiments of the present disclosure, example 8 provides the method of example 7, characterized in that the combining the image description generation module and the morphing combination module into the target morphing processing model comprises:

[0188] coupling an output end of the image description generation module with a first input end of the morphing combination module, an input end of the image description generation module and a second input end of the morphing combination module serving as input ends of the target morphing processing model, and an output end of the morphing combination module serving as an output end of the target morphing processing model.

[0189] According to one or more embodiments of the present disclosure, example 9 provides an image processing apparatus, the apparatus comprising:

[0190] a first acquisition module configured to acquire an original image and a label for representing a morphing capability, the original image comprising an object to be morphed according to the morphing capability;

[0191] a determination module configured to acquire an image description text of the original image, generate a target description text based on the image description text and the label, and generate a target image based on the original image and the target description text, the target image comprising the object morphed according to the morphing capability.

[0192] According to one or more embodiments of the present disclosure, example 10 provides a model training apparatus, the apparatus comprising:

[0193] a second acquisition module configured to acquire a plurality of sets of morphing image samples, each set of the morphing image samples comprising an original sample image, at least one label, and a morphing effect image of the original sample image corresponding to the at least one label;

[0194] a first training module configured to train a first preset initial ControlNet network using the plurality of sets of morphing image samples as training data to obtain a target morphing processing model.

[0195] According to one or more embodiments of the present disclosure, example 11 provides a model training apparatus, the apparatus comprising:

[0196] a third acquisition module configured to acquire a plurality of sets of first image sample data, each set of the first image sample data comprising an original sample image and description text annotation data corresponding to the original sample image;

[0197] a second training module configured to train a preset initial image description algorithm using the plurality of sets of first image sample data as training data to obtain an image description generation module;

[0198] a fourth obtaining module, configured to obtain a plurality of sets of second image sample data, each set of the second image sample data comprising an original sample image, labeled text data, and a morphing effect image corresponding to the original sample image, the labeled text data comprising description text label data and at least one label corresponding to the original sample image;

[0199] a second training module, configured to train a second preset initial ControlNet network by taking the plurality of sets of second image sample data as training data, to obtain a morphing combination module;

[0200] a coupling module, configured to combine the image description generation module and the morphing combination module into a target morphing processing model.

[0201] According to one or more embodiments of the present disclosure, example 12 provides a computer readable medium having a computer program stored thereon, wherein the program, when executed by a processing apparatus, implements the steps of the method of any one of examples 1-8.

[0202] According to one or more embodiments of the present disclosure, example 13 provides an electronic device, comprising:

[0203] a storage device having a computer program stored thereon;

[0204] a processing apparatus configured to execute the computer program in the storage device to implement the steps of the method of any one of examples 1-8.

[0205] According to one or more embodiments of the present disclosure, example 14 provides a computer program product comprising a computer program which, when executed by a processor, implements the steps of the method of any one of examples 1-8.

[0206] The above description is merely preferred embodiments of the present disclosure and a description of the principles of the technology applied. It should be understood by those skilled in the art that the disclosed scope of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combinations of the above technical features or equivalent features without departing from the disclosed concept. For example, the above features can be replaced with other technical features disclosed in the present disclosure (but not limited to) having similar functions to form technical solutions.

[0207] Moreover, while operations have been depicted in a particular order, this should not be understood as requiring such an order nor limiting it to only those operations shown and described. One of ordinary skill in the art will recognize that many of the operations can be performed in a differing order, or be performed concurrently, that some operations can be performed in any order or omitted, and that some operations can be performed in parallel. Similarly, while several specific implementation details have been discussed in the context of the above discussion, these should not be construed as limitations on the scope of the present disclosure. Certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination.

[0208] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims. With respect to the devices in the above-described embodiments, in which various modules perform operations, the specific manner in which the various modules perform the operations has been described in detail in the embodiments relating to the method. Here, no detailed explanation will be given.

Claims

1. An image processing method, characterized by, The method comprises: obtaining an original image and a label for representing a morphing capability, the original image comprising an object to be morphed according to the morphing capability; obtaining an image description text of the original image; generating a target description text based on the image description text and the label; generating a target image based on the original image and the target description text, the target image comprising the object morphed according to the morphing capability.

2. The image processing method according to claim 1, wherein: The obtaining of the image description text of the original image, the generating of the target description text based on the image description text and the label, and the generating of the target image based on the original image and the target description text are implemented by a target morphing processing model, and a training process of the target morphing processing model can comprise: obtaining a plurality of sets of morphing image samples, each set of the morphing image samples comprising an original sample image, at least one label, and a morphing effect image corresponding to the at least one label and the original sample image; training a first preset initial ControlNet network by taking the plurality of sets of morphing image samples as training data, to obtain the target morphing processing model.

3. The image processing method of claim 1, wherein, The obtaining of the image description text of the original image, the generating of the target description text based on the image description text and the label; The generating of the target image based on the original image and the target description text, the target image comprising the object morphed according to the morphing capability, comprises: inputting the original image and one or more labels for representing a morphing capability into a target morphing processing model, the target morphing processing model comprising an image description generation module and a morphing combination module coupled with the image description generation module; generating an image description text corresponding to the original image by the image description generation module; splicing the image description text and the one or more labels for representing a morphing capability to obtain a target description text; inputting the target description text and the original image into the morphing combination module to obtain the target image output by the morphing combination module.

4. The image processing method of claim 3, wherein, The target morphing processing model is trained in the following manner: obtaining a plurality of sets of first image sample data, each set of the first image sample data comprising an original sample image and description text annotation data corresponding to the original sample image; training a preset initial image description algorithm by taking the plurality of sets of first image sample data as training data, to obtain the image description generation module; obtaining a plurality of sets of second image sample data, each set of the second image sample data comprising an original sample image, annotation text data, and a morphing effect image corresponding to the original sample image, the annotation text data comprising description text annotation data corresponding to the original sample image and at least one label; training a second preset initial ControlNet network by taking the plurality of sets of second image sample data as training data, to obtain the morphing combination module; combining the image description generation module and the morphing combination module into the target morphing processing model.

5. The image processing method of claim 4, wherein, The image description generation module and the deformation combination module are combined into the target deformation processing model, including: The output end of the image description generation module is coupled with the first input end of the deformation combination module, the input end of the image description generation module and the second input end of the deformation combination module serve as the input end of the target deformation processing model, and the output end of the deformation combination module serves as the output end of the target deformation processing model.

6. A model training method, comprising: The method comprises: Obtaining a plurality of sets of deformation image samples, each set of deformation image samples comprising an original sample image, at least one label, and a deformation effect image of the original sample image corresponding to the at least one label; Using the plurality of sets of deformation image samples as training data, training a first preset initial ControlNet network to obtain a target deformation processing model.

7. A model training method, comprising: The method comprises: Obtaining a plurality of sets of first image sample data, each set of first image sample data comprising an original sample image and description text annotation data corresponding to the original sample image; Using the plurality of sets of first image sample data as training data, training a preset initial image description algorithm to obtain an image description generation module; Obtaining a plurality of sets of second image sample data, each set of second image sample data comprising an original sample image, annotation text data, and a deformation effect image corresponding to the original sample image, the annotation text data comprising description text annotation data corresponding to the original sample image and at least one label; Using the plurality of sets of second image sample data as training data, training a second preset initial ControlNet network to obtain a deformation combination module; Combining the image description generation module and the deformation combination module into a target deformation processing model.

8. The method of claim 7, wherein, The image description generation module and the deformation combination module are combined into the target deformation processing model, including: The output end of the image description generation module is coupled with the first input end of the deformation combination module, the input end of the image description generation module and the second input end of the deformation combination module serve as the input end of the target deformation processing model, and the output end of the deformation combination module serves as the output end of the target deformation processing model.

9. An image processing apparatus characterized by comprising: The device comprises: A first obtaining module configured to obtain an original image and a label for representing a deformation capability, the original image comprising an object to be deformed according to the deformation capability; A determining module configured to obtain an image description text of the original image, generate a target description text based on the image description text and the label, and generate a target image based on the original image and the target description text, the target image comprising the object deformed according to the deformation capability.

10. A model training apparatus, comprising: The device comprises: A second obtaining module configured to obtain a plurality of sets of deformation image samples, each set of deformation image samples comprising an original sample image, at least one label, and a deformation effect image of the original sample image corresponding to the at least one label; The first training module is configured to train the first preset initial ControlNet network by taking the plurality of sets of morphing image samples as training data, to obtain a target morphing processing model.

11. A model training apparatus, comprising: The device comprises: The third acquisition module is configured to acquire a plurality of sets of first image sample data, each set of the first image sample data comprising an original sample image and description text label data corresponding to the original sample image; The second training module is configured to train a preset initial image description algorithm by taking the plurality of sets of first image sample data as training data, to obtain an image description generation module; The fourth acquisition module is configured to acquire a plurality of sets of second image sample data, each set of the second image sample data comprising an original sample image, label text data, and a morphing effect image corresponding to the original sample image, the label text data comprising description text label data corresponding to the original sample image and at least one label; The second training module is configured to train a second preset initial ControlNet network by taking the plurality of sets of second image sample data as training data, to obtain a morphing combination module; The coupling module is configured to combine the image description generation module and the morphing combination module into a target morphing processing model.

12. A computer readable medium having stored thereon a computer program, characterized in that, The program, when executed by a processing device, implements the steps of the method of any one of claims 1-8.

13. An electronic device, comprising: Comprise: A storage device having a computer program stored thereon; A processing device for executing the computer program in the storage device to implement the steps of the method of any one of claims 1-8.

14. A computer program product comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the steps of the method of any one of claims 1-8. The computer program, when executed by a processor, implements the steps of the method of any one of claims 1-8.