Image Processing Method, Apparatus, Electronic Device, and Storage Medium

通过接收图像修改描述信息进行语义编码和风格特征调整,解决了现有技术中图像编辑效率低的问题,实现了自动化的高效图像风格变换。

CN114266840BActive Publication Date: 2025-07-08BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111572467.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-21
Publication Date
2025-07-08
Estimated Expiration
2041-12-21

AI Technical Summary

Technical Problem

In the prior art, image editing methods are inefficient in processing, manual editing is time-consuming and labor-intensive, and deep learning model editing requires a lot of resources and is inconvenient to use.

Method used

By receiving image modification description information, perform semantic coding, obtaining style feature change information, adjusting the original style feature set, generating style transformed images, and using pre-trained semantic coding models and style feature recognition models to achieve automatic image style adjustment.

Benefits of technology

Improve image editing efficiency, reduce labor costs and model training resource requirements, and support rapid style transformation of multiple image types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114266840B_ABST
    Figure CN114266840B_ABST
Patent Text Reader

Abstract

The present disclosure relates to an image processing method, apparatus, electronic device, and storage medium. The method includes: receiving image modification description information corresponding to an original image to be subjected to style transformation; performing semantic encoding on the image modification description information to obtain corresponding semantic encoding information; obtaining style feature change information corresponding to the semantic encoding information; obtaining an original style feature set of the original image; adjusting the original style features in the original style feature set based on the style feature change information to obtain a target style feature set; and generating a style transformation image corresponding to the image modification description information based on the target style feature set. In this solution, the original style features are adjusted based on the semantic encoding information corresponding to the image modification description information, and a style transformation image is generated, without the need for manual image adjustment, effectively improving the editing efficiency, and the image style can be modified by adjusting the style features, which can reduce the image processing cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computers, and in particular, to an image processing method, apparatus, electronic device, and storage medium. Background Art

[0002] With the development of Internet technology, a large number of image resources, such as pictures or videos composed of multiple frames of pictures, flood the network. To improve the quality of image resources or for the need of playability, users often edit images.

[0003] In related technologies, users can manually modify the image content to be beautified in the image one by one with the help of editing software; or they can train a customized model based on a deep learning model to achieve the effect of beautifying the image through the customized model.

[0004] However, manual image editing is time-consuming and laborious, requiring a large amount of labor costs; image editing based on a deep learning model requires combining multiple separately trained models, spending a large amount of resources on model training, and is inconvenient to use. It can be seen that the above image editing methods have low processing efficiency. Summary of the Invention

[0005] The present disclosure provides an image processing method, apparatus, electronic device, and storage medium to at least solve the problem of low processing efficiency in the related art image editing methods. The technical solutions of the present disclosure are as follows:

[0006] According to a first aspect of an embodiment of the present disclosure, an image processing method is provided, including:

[0007] Receiving image modification description information corresponding to an original image to be subjected to style transformation;

[0008] Semantically encoding the image modification description information to obtain semantic encoding information corresponding to the image modification description information;

[0009] Obtaining style feature change information corresponding to the semantic encoding information, where the style feature change information is feature change information corresponding to an image style;

[0010] Obtaining an original style feature set corresponding to the original image;

[0011] Adjusting the original style features in the original style feature set based on the style feature change information to obtain a target style feature set;

[0012] Generating a style transformation image corresponding to the image modification description information based on the target style feature set.

[0013] In an exemplary embodiment, adjusting the original style features in the original style feature set based on the style feature change information to obtain a target style feature set includes:

[0014] Adjusting a first original style feature in the original style feature set based on the style feature change information to obtain an adjusted first style feature; the first original style feature is the original style feature corresponding to the style feature change information;

[0015] Combining a second original style feature and the adjusted first style feature to obtain the target style feature set, where the second original style feature is the original style feature other than the first original style feature in the original style feature set.

[0016] In an exemplary embodiment, obtaining the style feature change information corresponding to the semantic encoding information includes:

[0017] Obtaining a mapping relationship between a pre-determined image change feature space and a text semantic feature space;

[0018] Taking the semantic encoding information as the text semantic feature in the text semantic feature space, and obtaining the image change feature corresponding to the text semantic feature in the image change feature space based on the mapping relationship;

[0019] Taking the image change feature as the style feature change information.

[0020] In an exemplary embodiment, performing semantic encoding on the image modification description information to obtain the semantic encoding information corresponding to the image modification description information includes:

[0021] Inputting the image modification description information into a pre-trained semantic encoding model to encode the image modification description information based on the semantic encoding parameters in the semantic encoding model, so as to obtain the semantic encoding information corresponding to the image modification description information;

[0022] The semantic encoding model is trained based on paired training texts and training images.

[0023] In an exemplary embodiment, the steps of training the semantic encoding model include:

[0024] Obtaining training samples, where the training samples include first training images and training texts paired with the first training images;

[0025] Inputting the training texts into a semantic encoding model to be trained for encoding to obtain text encoding features of the training texts;

[0026] Input the first training image into an image encoding model for encoding to obtain the image encoding features of the first training image;

[0027] Determine a first similarity between each of the first training images and each of the training texts according to the image encoding features of the first training image and the text encoding features of the training text; the first similarity characterizes the similarity between the encoding features corresponding to the image modality and the encoding features corresponding to the text modality;

[0028] Determine a target loss value corresponding to the semantic encoding model to be trained according to the first similarity;

[0029] Adjust the model parameters of the semantic encoding model to be trained according to the target loss value until the training end condition is satisfied to obtain the trained semantic encoding model.

[0030] In an exemplary embodiment, there are multiple training samples, and determining the target loss value corresponding to the semantic encoding model to be trained according to the first similarity includes:

[0031] Obtain a first loss value corresponding to the semantic encoding model to be trained according to the first similarity, and the first loss value has a negative correlation with the first similarity;

[0032] Obtain a second loss value corresponding to the semantic encoding model to be trained according to a second similarity, and the second loss value has a positive correlation with the second similarity; the second similarity is the similarity between the training texts in different training samples and a first training text;

[0033] Obtain the target loss value corresponding to the semantic encoding model to be trained based on the first loss value and the second loss value.

[0034] In an exemplary embodiment, the original style feature set is determined by a pre-trained style feature recognition model based on the input original image, and the steps of training the style feature recognition model include:

[0035] Obtain a second training image, and input the second training image into the style feature recognition model to be trained so as to obtain a predicted style feature set corresponding to the second training image through the style feature recognition model to be trained;

[0036] Generate a corresponding predicted image based on the predicted style feature set;

[0037] Obtain a third loss value corresponding to the style feature recognition model to be trained according to the difference between the predicted image and the second training image, where the third loss value is positively correlated with the difference;

[0038] Adjust the model parameters of the style feature recognition model to be trained according to the third loss value until the training end condition is met, and obtain the trained style feature recognition model.

[0039] In an exemplary embodiment, before receiving the image modification description information corresponding to the original image to be subjected to style transformation, it further includes:

[0040] Obtain a video to be processed, and determine a target video frame in the video as the original image to be subjected to style transformation;

[0041] After generating the style transformation image corresponding to the image modification description information based on the target style feature set, it further includes:

[0042] Based on the style transformation image and other video frames in the video, obtain a target video; the other video frames are video frames in the video other than the target video frame.

[0043] According to a second aspect of the embodiments of the present disclosure, there is provided an image processing apparatus, including:

[0044] A description information acquisition unit configured to receive image modification description information corresponding to an original image to be subjected to style transformation;

[0045] A semantic encoding unit configured to perform semantic encoding on the image modification description information to obtain semantic encoding information corresponding to the image modification description information;

[0046] A style transformation determination unit configured to obtain style feature change information corresponding to the semantic encoding information, where the style feature change information is feature change information corresponding to an image style;

[0047] An original feature set acquisition unit configured to obtain an original style feature set corresponding to the original image;

[0048] A target feature set acquisition unit configured to adjust the original style features in the original style feature set based on the style feature change information to obtain a target style feature set;

[0049] A style transformation image acquisition unit configured to generate a style transformation image corresponding to the image modification description information based on the target style feature set.

[0050] In an exemplary embodiment, the target feature set obtaining unit is configured to perform:

[0051] Adjust the first original style feature in the original style feature set based on the style feature change information to obtain an adjusted first style feature; the first original style feature is the original style feature corresponding to the style feature change information;

[0052] Combine the second original style feature and the adjusted first style feature to obtain the target style feature set, where the second original style feature is the original style feature other than the first original style feature in the original style feature set.

[0053] In an exemplary embodiment, the style transformation determining unit is configured to perform:

[0054] Obtain a pre-determined mapping relationship between the image change feature space and the text semantic feature space;

[0055] Use the semantic encoding information as the text semantic feature in the text semantic feature space, and obtain the image change feature corresponding to the text semantic feature in the image change feature space based on the mapping relationship;

[0056] Use the image change feature as the style feature change information.

[0057] In an exemplary embodiment, the semantic encoding unit is configured to perform:

[0058] Input the image modification description information into a pre-trained semantic encoding model to encode the image modification description information based on the semantic encoding parameters in the semantic encoding model, and obtain the semantic encoding information corresponding to the image modification description information;

[0059] The semantic encoding model is trained based on paired training texts and training images.

[0060] In an exemplary embodiment, the apparatus further includes:

[0061] A first training image obtaining unit configured to obtain a training sample, where the training sample includes a first training image and a training text paired with the first training image;

[0062] A training text encoding unit configured to input the training text into a semantic encoding model to be trained for encoding to obtain the text encoding feature of the training text;

[0063] The first training image encoding unit is configured to perform encoding the first training image by inputting it into an image encoding model to obtain the image encoding features of the first training image;

[0064] The first similarity obtaining unit is configured to perform determining the first similarity between each of the first training images and each of the training texts according to the image encoding features of the first training image and the text encoding features of the training text; the first similarity characterizes the similarity between the encoding features corresponding to the image modality and the encoding features corresponding to the text modality;

[0065] The target loss value obtaining unit is configured to perform determining the target loss value corresponding to the semantic encoding model to be trained according to the first similarity;

[0066] The first parameter adjustment unit is configured to perform adjusting the model parameters of the semantic encoding model to be trained according to the target loss value until the training end condition is satisfied, and obtaining the trained semantic encoding model.

[0067] In an exemplary embodiment, there are multiple training samples, and the target loss value obtaining unit is configured to perform:

[0068] Obtaining a first loss value corresponding to the semantic encoding model to be trained according to the first similarity, where the first loss value has a negative correlation with the first similarity;

[0069] Obtaining a second loss value corresponding to the semantic encoding model to be trained according to a second similarity, where the second loss value has a positive correlation with the second similarity; the second similarity is the similarity between the training texts in different training samples and a first training text;

[0070] Based on the first loss value and the second loss value, obtaining the target loss value corresponding to the semantic encoding model to be trained.

[0071] In an exemplary embodiment, the original style feature set is determined by a pre-trained style feature recognition model based on the input original image, and the apparatus further includes:

[0072] The second training image obtaining unit is configured to perform obtaining a second training image, and inputting the second training image into the style feature recognition model to be trained, so as to obtain the predicted style feature set corresponding to the second training image through the style feature recognition model to be trained;

[0073] The predicted image generating unit is configured to perform generating a corresponding predicted image based on the predicted style feature set;

[0074] A third loss value obtaining unit, configured to obtain a third loss value corresponding to the style feature recognition model to be trained according to a difference between the predicted image and the second training image, where the third loss value is positively correlated with the difference;

[0075] A second parameter adjustment unit, configured to adjust model parameters of the style feature recognition model to be trained according to the third loss value until a training end condition is met, to obtain the trained style feature recognition model.

[0076] In an exemplary embodiment, the apparatus further includes:

[0077] A video obtaining unit, configured to obtain a video to be processed and determine a target video frame in the video as an original image to be subjected to style transformation;

[0078] The apparatus further includes:

[0079] A video updating unit, configured to obtain a target video based on the style-transformed image and other video frames in the video; the other video frames are video frames in the video other than the target video frame.

[0080] According to a third aspect of embodiments of the present disclosure, there is provided an electronic device, including:

[0081] A processor;

[0082] A memory for storing instructions executable by the processor;

[0083] Wherein, the processor is configured to execute the instructions to implement the image processing method as described in any one of the above.

[0084] According to a fourth aspect of embodiments of the present disclosure, there is provided a computer-readable storage medium, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the image processing method as described in any one of the above.

[0085] According to a fifth aspect of embodiments of the present disclosure, there is provided a computer program product, where the computer program product includes instructions, and when the instructions are executed by a processor of an electronic device, enabling the electronic device to execute the image processing method as described in any one of the above.

[0086] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:

[0087] In the solution of the present disclosure, by modifying the description letter of the input image, the original style features associated in the image can be adjusted based on the corresponding semantic coding information, and the corresponding style transformation image can be generated. There is no need to manually adjust the image content, which effectively improves the image editing efficiency. At the same time, by adjusting the style features of the image, the image style can be modified, reducing the image processing cost.

[0088] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0089] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0090] Figure 1 is an application environment diagram of an image processing method shown according to an exemplary embodiment.

[0091] Figure 2 is a flowchart of an image processing method shown according to an exemplary embodiment.

[0092] Figure 3 is a flowchart of a method for training a semantic coding model shown according to an exemplary embodiment.

[0093] Figure 4 is a schematic diagram of a feature matrix shown according to an exemplary embodiment.

[0094] Figure 5 is a flowchart of another image processing method shown according to an exemplary embodiment.

[0095] Figure 6 is a block diagram of an image processing apparatus shown according to an exemplary embodiment.

[0096] Figure 7 is a block diagram of an electronic device shown according to an exemplary embodiment.

[0097] Figure 8 is a block diagram of another electronic device shown according to an exemplary embodiment. DETAILED DESCRIPTION

[0098] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0099] It should be noted that the terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present disclosure are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0100] It should also be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in the present disclosure are all information and data authorized by the user or fully authorized by all parties.

[0101] With the development of Internet technology, a large number of image resources, such as pictures or videos composed of multiple frames of pictures, are flooding the network. In order to improve the quality of image resources or for the need of playability, users often edit images. For example, users can modify the facial features or body characteristics of the people in the picture, or modify the background of the picture; similarly, users can modify the video frames in the video to achieve the effect of modifying the original video.

[0102] In the related art, users can modify the image content to be beautified in the image one by one based on manual methods with the help of editing software. Or based on a deep learning model, corresponding customized models can be trained for different image types, and the effect of beautifying the image can be achieved through the customized models, such as training a model specifically for faces.

[0103] However, manual image editing is time-consuming and laborious, and requires a relatively high labor cost; image editing based on a deep learning model requires combining multiple separately trained models, which requires a large amount of resources for model training and is relatively inconvenient to use. For image processing of different types or different parts, a matching model needs to be called for processing. It can be seen that the above-mentioned image editing methods have low processing efficiency.

[0104] An image processing method provided by the present disclosure can be applied to an application environment as Figure 1 shown. In this application environment, the terminal 110 interacts with the server 120 through the network. The terminal 110 can send the image to be processed to the server 120, and the server 120 executes the image processing method provided by the present disclosure to process the received image. Of course, the image processing method in the present disclosure can also be applied to the terminal 110, that is, the terminal 110 can execute the image processing method of the present disclosure to process the image stored in the terminal 110.

[0105] As an example, the terminal 110 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. Among them, the portable wearable device can be a smart watch, a smart bracelet, a head-mounted device, etc. The server 120 can be an independent physical server or a server cluster composed of multiple physical servers, and can be a cloud server that provides basic cloud computing services such as cloud servers, cloud databases, cloud storage, and CDN.

[0106] Figure 2 is a flowchart of an image processing method shown according to an exemplary embodiment. As Figure 2 shown, when this method is used in the server 120, it includes the following steps.

[0107] In step S210, receive the image modification description information corresponding to the original image to be subjected to style transformation.

[0108] As an example, style transformation can refer to modifying the image style of the original image. Among them, the image style can be the visible artistic style and / or image characteristics presented by the whole image or the objects in the image.

[0109] The image modification description information can be information represented by natural language. Specifically, the image modification description information can be text information or voice information input by the user based on natural language. The image modification description information can include the image style to be presented in the image.

[0110] In practical applications, the server can obtain the original image to be subjected to style transformation and receive the image modification description information for the original image.

[0111] Specifically, for example, after the user determines the image style to be presented after the style transformation of the original image, the user can send a corresponding image modification instruction with a modification intention to the terminal 110 by means of text input or voice input. For example, if the original image is an image of a cat and the user hopes that the animal in the picture can have the style of "lovely" after the style transformation, the user can input text or voice "lovely cat" to the terminal. In response to the user's operation, the terminal 110 can send the original image and the image modification instruction to the server 120.

[0112] After receiving the original image and the image modification instruction, if the image modification instruction is an instruction generated by text input, the server can obtain image modification description information based on the text in the image modification instruction; if the image modification instruction is the voice input by the user, the server can perform speech recognition on the received voice and obtain image modification description information based on the recognized text content. In the present disclosure, the user only needs to input the image modification description information to trigger the human-computer interactive image modification, without the user having professional image processing knowledge or knowing the responsible image modification parameters, which greatly reduces the processing threshold of image style transformation, and is simple and fast at the same time, effectively improving the image processing efficiency.

[0113] In step S220, semantic encoding is performed on the image modification description information to obtain semantic encoding information corresponding to the image modification description information.

[0114] As an example, the semantic encoding information may be information representing the semantic change situation.

[0115] In a specific implementation, an image can be described and summarized by text, that is, the information contained in the image can be represented by text; correspondingly, the currently obtained original image and the image obtained by performing style transformation on the original image can also be described by different texts respectively. For the same object or the same content, the difference in text expression between the original image and the image after style transformation before and after the style transformation can be determined as the style transformation direction of the original image, that is, the difference in text content can be associated with the situation of image style transformation.

[0116] Based on this, after obtaining the image modification description information, semantic encoding can be performed on the image modification description information, and based on the result of the semantic encoding, semantic encoding information corresponding to the image modification description information can be obtained.

[0117] Specifically, since it may be that some or all of the content in the image modification description information represents the user's modification intention for the original image. For example, for "lovely cat", the modification intention may be "lovely", may be "cat", or may be the whole "lovely cat". When the server obtains the semantic encoding information corresponding to the image modification description information, it can obtain a reference information. After performing semantic encoding on the image modification description information to obtain the corresponding semantic encoding result, by comparing the semantic encoding result with the encoding result corresponding to the reference information, the semantic encoding information corresponding to the image modification description information can be obtained. When there are multiple objects in the image, the reference information for multiple objects can also be input in advance. Then, after obtaining the image modification description information, by comparing the speech encoding result corresponding to the image modification description information with the encoding result corresponding to the reference information, the semantic encoding information for one or more of the multiple objects can be determined according to the difference.

[0118] In step S230, obtain the style feature change information corresponding to the semantic coding information.

[0119] Among them, the style feature change information is the feature change information corresponding to the image style, which can be the feature change information of the same style feature or the feature change information of multiple style features.

[0120] Specifically, the image style can be determined by one or more style features corresponding to the image. Among them, the style feature can be the feature of the visible attribute corresponding to the content or object in the image, and there can be a one-to-one correspondence between the style feature and the visible attribute, that is, one style feature represents the feature of one visible attribute. As an example, the style feature can be the feature of the visible attribute corresponding to the object in the image. Taking "human face" as an example, the visible attributes can include at least one of the following: face shape, expression, face orientation, hairstyle, human face skin color, human face illumination; or, the style feature can also be the feature of the visible attribute corresponding to the whole image or the image background, such as attributes such as the lines, colors used, or composition of the image.

[0121] Since the style features of the image are directly related to the image content, by changing the style features of the image, the visible attributes of the image can be adjusted, and thus the effect of adjusting the image style can be achieved. Therefore, after obtaining the semantic coding information, the style feature change information corresponding to the semantic coding information can be obtained.

[0122] In step S240, obtain the set of original style features corresponding to the original image.

[0123] In step S250, adjust the original style features in the set of original style features based on the style feature change information to obtain a set of target style features.

[0124] As an example, the set of original style features can be a set composed of multiple original style features corresponding to the original image.

[0125] In a specific implementation, the set of original style features corresponding to the original image can be obtained, and based on the determined style feature change information, the original style features in the set of original style features can be adjusted to obtain a set of target style features.

[0126] In step S260, generate a style transformation image corresponding to the image modification description information based on the set of target style features.

[0127] Specifically, since the style features are associated with the image content, after obtaining the adjusted set of target style features, a style transformation image corresponding to the image modification description information can be generated based on the set of target style features.

[0128] In the above image processing method, it is possible to receive image modification description information corresponding to an original image to be subjected to style transformation, perform semantic encoding on the image modification description information to obtain semantic encoding information corresponding to the image modification description information, and obtain style feature change information corresponding to the semantic encoding information, where the style feature change information is feature change information corresponding to the image style; after obtaining the original style feature set corresponding to the original image, it is possible to adjust the original style features in the original style feature set based on the style feature change information to obtain a target style feature set, and generate a style transformation image corresponding to the image modification description information based on the target style feature set. In the solution of the present disclosure, by inputting the image modification description information, it is possible to adjust the original style features associated in the image based on the corresponding semantic encoding information and generate a corresponding style transformation image, without the need to manually adjust the image content, effectively improving the image editing efficiency. At the same time, by adjusting the style features of the image, the image style can be modified, reducing the image processing cost.

[0129] In an exemplary embodiment, in step S250, the adjusting the original style features in the original style feature set based on the style feature change information to obtain a target style feature set may include:

[0130] Adjusting a first original style feature in the original style feature set based on the style feature change information to obtain an adjusted first style feature; combining a second original style feature and the adjusted first style feature to obtain the target style feature set.

[0131] Wherein, the first original style feature is the original style feature corresponding to the style feature change information; the second original style feature is the original style feature other than the first original style feature in the original style feature set.

[0132] Specifically, the transformation of the image style may be achieved by changing one style feature, or may require changing multiple style features. For example, the modification of the length or color of hair can be achieved by changing the style feature corresponding to the hairstyle, while the transformation of the image style of "aging" may involve style features corresponding to multiple visible attributes such as hairstyle, skin color, and skin texture.

[0133] Based on this, the style feature change information determined based on the semantic encoding information may be the change information corresponding to the first style feature associated with the semantic encoding information in the original style feature set. Furthermore, based on the style feature transformation information, the first original style feature in the original style feature set can be adjusted to obtain an adjusted first style feature.

[0134] After adjusting the first original style feature, the second original style feature in the original style feature set and the adjusted first style feature can be combined to obtain a target style feature set.

[0135] In the present disclosure, by adjusting the first original style feature related to the style feature change information in the original style feature set and keeping the second original style feature unrelated to the style feature change information unchanged, a target style feature set is generated, which can accurately adjust the specified image style while retaining the image content that is not indicated to be adjusted in the image, and avoid changing other image styles of the image.

[0136] In one exemplary embodiment, in step S230, the obtaining of the style feature change information corresponding to the semantic encoding information may include:

[0137] Obtaining a mapping relationship between a pre-determined image change feature space and a text semantic feature space; using the semantic encoding information as the text semantic feature in the text semantic feature space, and obtaining the image change feature corresponding to the text semantic feature in the image change feature space based on the mapping relationship; using the image change feature as the style feature change information.

[0138] As an example, the image change feature space may be the feature space where the style feature is located; the text semantic feature space may be the feature space where the text semantic feature is located.

[0139] In practical applications, there is a mapping relationship between the image change feature space and the text semantic feature space, that is, when the text semantic feature changes by Δt, the image it describes will also change, and this change is reflected by corresponding changes in one or more style features in the image. Therefore, when the text semantic feature changes by Δt, the style feature of the image will also change by Δs.

[0140] In practical applications, the mapping relationship between the image change feature space and the text semantic feature space can be pre-determined. In one example, when constructing the mapping relationship, the original style feature s1 can be obtained. The original style feature s1 can generate the corresponding image i1, and the image i1 can have the corresponding text semantic feature t1. After adjusting the original style feature s1 by Δs to obtain the adjusted style feature s2, the image i2 and the text semantic feature t2 corresponding to the style feature s2 can also be obtained, and Δt can be obtained based on t1 and t2. Furthermore, based on Δs and Δt, the mapping relationship between the image change feature space and the text semantic feature space can be obtained.

[0141] After obtaining the semantic encoding information, a pre-determined mapping relationship can be obtained, and the semantic encoding information can be used as the text semantic feature in the text semantic feature space. Then, based on the pre-determined mapping relationship, the image change feature corresponding to the text semantic feature in the image change feature space can be obtained, and the image change feature can be used as the style feature change information.

[0142] In the present disclosure, the semantic encoding information can be used as the text semantic feature in the text semantic feature space, and the image change feature corresponding to the text semantic feature in the image change feature space can be obtained based on the mapping relationship. Then, the image change feature can be used as the style feature change information, and the change in the style feature of the image can be determined according to the text semantics input by the user, providing an accurate feature change indication for the style change of the image.

[0143] Moreover, in the traditional technology, relevant filtering or pixel point conversion algorithms can be designed based on pixel points to process the key parts of the human body in the picture to achieve beautification effects such as face slimming and skin smoothing. However, this method can only modify specified types of images or specific parts in the image, with low efficiency and limitations. In the present disclosure, since style features reflecting the image style can be extracted from different images, and during the image processing process, the original style feature set is modified based on the style feature change information, that is, a style transformation image can be generated based on the modified target style feature set, enabling different images to be processed using the method of the present disclosure, quickly obtaining the style transformation image, greatly improving the generalization of the model processing, without the need to train a model specifically for different types of pictures or different object parts, and being able to support multiple retouching tasks without retraining the model.

[0144] In an exemplary embodiment, the text semantic feature space can be the feature space corresponding to the semantic encoding model, that is, each text semantic feature in the text semantic feature space can be determined by the semantic encoding model.

[0145] In an exemplary embodiment, in step S220, the semantic encoding of the image modification description information to obtain the semantic encoding information corresponding to the image modification description information may include:

[0146] Inputting the image modification description information into a pre-trained semantic encoding model to encode the image modification description information based on the semantic encoding parameters in the semantic encoding model, so as to obtain the semantic encoding information corresponding to the image modification description information.

[0147] Among them, the semantic encoding model is trained based on paired training texts and training images.

[0148] In practical applications, after obtaining the image modification description information, the image modification description information can be input into a pre-trained semantic encoding model. The semantic encoding model is trained based on paired training texts and training images. By training the semantic encoding model based on the paired training texts and training images, the information in the text modality can be incorporated into the image modality.

[0149] After inputting the image modification description information, the image modification description information can be encoded based on the semantic encoding parameters in the semantic encoding model to obtain the semantic encoding information corresponding to the image modification description information.

[0150] Specifically, for example, after inputting the image modification description information into the semantic encoding model, the first text encoding feature corresponding to the image modification description information can be obtained through the semantic encoding model. The first text encoding feature can include the word vectors corresponding to multiple word segments in the image modification description information. And, a reference information can be pre-input into the semantic encoding model in advance to obtain the second text encoding feature corresponding to the reference information through the semantic encoding model. The second text encoding feature can include the word vectors corresponding to multiple word segments in the reference information. Then, the semantic encoding information can be obtained according to the difference between the first text encoding feature and the second text encoding feature. Among them, the reference information can be the reference information input by the user. For example, for the input image modification description information "lovely cat", the user can also input "cat" or "lovely person" as the reference information. Of course, the user can also not input the reference information, then the information associated with the image modification description information stored in advance can be used as the reference information. Of course, a null value can also be used as the reference information.

[0151] In the present disclosure, since the semantic encoding model is trained based on paired training texts and training images, the semantic encoding information with the image modality can be obtained through the semantic encoding model, enabling the text modality information to participate in the modification of the image modality information, providing a basis for subsequently determining the corresponding style feature change information based on the semantic encoding information.

[0152] In an exemplary embodiment, as Figure 3 shown, the steps of training the semantic encoding model may include:

[0153] In step S310, training samples are obtained.

[0154] Among them, the training samples may include the first training image and the training text paired with the first training image.

[0155] In practical applications, a first training image can be obtained, and text associated with the first training image can be obtained as training text paired with the first training image. Thus, the first training image and its paired training text can be used as training samples.

[0156] Specifically, when obtaining the training text paired with the first training image, text associated with the first training image can be crawled from network resources provided by network forums or social platforms, or image resources or video resources uploaded by users to the social platform as the paired training text. For example, the title corresponding to the first training image can be used as the paired training text, or the first training image can be obtained from the video resource, and the subtitle or title corresponding to the video resource can be used as the paired training text. Of course, the paired training text can also be extracted from the copywriting corresponding to the first training image.

[0157] In step S320, the training text is input into the semantic encoding model to be trained for encoding, and the text encoding features of the training text are obtained.

[0158] As an example, the text encoding features can be features representing the semantics of the text. For example, they can be the encoding feature sequence corresponding to the training text, and this encoding feature sequence can contain the feature vectors corresponding to each word segment of the training text. The semantic encoding model to be trained can be constructed based on a CNN network (Convolutional Neural Networks) or an RNN network (Recurrent Neural Network).

[0159] After obtaining the training text, the training text can be input into the semantic encoding model to be trained for encoding, and the text encoding features corresponding to the training text are obtained.

[0160] In step S330, the first training image is input into the image encoding model for encoding, and the image encoding features of the first training image are obtained.

[0161] As an example, the image encoding features can be the feature vectors corresponding to the training image, and the image encoding features can represent multiple image features of the training image. The image encoding features are different from the style features of the image. The multiple image features represented by the image encoding features can be related to each other, with a relatively high coupling degree and mutual influence.

[0162] Specifically, after obtaining the first training image, the first training image can be input into the image encoding model for encoding, and the image encoding features corresponding to the first training image are obtained.

[0163] In step S340, based on the image encoding features of the first training image and the text encoding features of the training text, determine the first similarity between each of the first training images and each of the training texts.

[0164] Among them, the first similarity can characterize the similarity between the encoding features corresponding to the image modality and the encoding features corresponding to the text modality. Specifically, the same content can be represented by information in multiple modalities. For example, when expressing a content, it can be expressed through text, speech, or image. In this embodiment, the first similarity can characterize the similarity between the encoding features obtained in the image modality and the encoding features obtained in the text modality.

[0165] After obtaining the image encoding features of the first training image and the text encoding features of the training text, the first similarity between each first training image and each training text can be determined according to the image encoding features and the text encoding features. Specifically, for example, the image encoding features and the text encoding features can be feature vectors, and the inner product between the image encoding features and the text encoding features can be obtained as the first similarity between the first training image and the training text.

[0166] In step S350, based on the first similarity, determine the target loss value corresponding to the semantic encoding model to be trained.

[0167] After determining the first similarity, the target loss value corresponding to the semantic encoding model to be trained can be determined according to the first similarity. Specifically,

[0168] In step S360, adjust the model parameters of the semantic encoding model to be trained according to the target loss value until the training end condition is satisfied, and obtain the trained semantic encoding model.

[0169] After determining the target loss value, the model parameters of the semantic encoding model to be trained can be adjusted according to the target loss value to reduce the target loss value corresponding to the next training process. Repeat the model training process until the training end condition is satisfied, such as the current target loss value tends to be stable and the fluctuation of the target loss value is less than the threshold, then the trained semantic encoding model can be obtained.

[0170] In the present disclosure, the first similarity between each first training image and each training text can be determined according to the image encoding features of the first training images and the text encoding features of the training texts, and the target loss value corresponding to the semantic encoding model to be trained can be determined according to the first similarity. Furthermore, the model parameters of the semantic encoding model to be trained can be adjusted according to the target loss value until the training end condition is satisfied, and the trained semantic encoding model can be obtained. Through the training method in the present disclosure, the similarity between the text encoding features and the image encoding features generated by the semantic encoding model is appropriately increased, so as to integrate the information in the text modality into the information in the image modality, providing a basis for subsequent semantic encoding of the modified description information of the image and obtaining the style feature change information corresponding to the semantic encoding information.

[0171] In an exemplary embodiment, there can be multiple training samples, and the number (batch) of training samples can be adjusted during the model training process. In practical applications, multiple images and the texts paired with each image can be obtained, and then the currently obtained images and their paired texts can be subjected to data cleaning. For example, cleaning can be performed in units of image-text pairs. When the image quality is lower than the preset quality requirement, such as when the image size or clarity is too low, or when the text length is too long or too short (such as the number of characters in the text is greater than or less than the preset threshold) or contains special characters (such as containing preset characters or sensitive keywords), the image and its paired text can be filtered out together. Based on the filtered images and paired texts that meet the data specifications, they are used as the first training images and paired training texts. Through data cleaning, the reliability of the first training images and their paired training texts can be improved, providing a basis for the accuracy of the semantic encoding model.

[0172] Of course, in order to improve the model expression ability of the semantic encoding model to be trained, data augmentation can also be performed on the images and paired texts obtained after data cleaning.

[0173] Specifically, if the image is obtained from a video, such as a video cover, frame extraction can be performed on the same video. For example, the following at least one frame extraction strategy can be used to obtain images similar to the obtained image from the video: uniform frame extraction, frame extraction at a fixed frame interval, and frame extraction according to the inter-frame difference. By performing frame extraction on the video, multiple video images including the video cover can be obtained, and the multiple video images can correspond to the same video title.

[0174] If the image is directly crawled, the following at least one transformation operation can be performed on the image to obtain the transformed image: rotation operation, flipping operation, scaling transformation, translation transformation, scale transformation, noise perturbation, color transformation, occlusion operation. By performing different transformation operations on the image, multiple images with the same or high similarity in image content can be obtained.

[0175] For the text paired with the image, data augmentation of the text can be achieved through at least one of the following operations: replacing the corresponding content in the text with synonyms, randomly permuting adjacent characters, replacing with Chinese equivalent characters, translation conversion, sentence pattern transformation (such as changing to an inverted sentence). Through the above operations, multiple texts with the same semantics but different actual expressions can be obtained.

[0176] For each pair of paired images and texts, after data augmentation, the multiple images and texts obtained after data augmentation can be paired to obtain multiple training samples. When there are multiple pairs of paired images and texts, a large number of training samples can be quickly obtained in this way, saving manpower and material resources. While expanding the scale of training data, the quality of training data is improved, thereby improving the image quality of the style transformation image obtained in subsequent image processing.

[0177] In step S350, the determining of the target loss value corresponding to the semantic encoding model to be trained according to the first similarity may include:

[0178] Obtaining a first loss value corresponding to the semantic encoding model to be trained according to the first similarity; obtaining a second loss value corresponding to the semantic encoding model to be trained according to a second similarity; and obtaining the target loss value corresponding to the semantic encoding model to be trained based on the first loss value and the second loss value.

[0179] Among them, the first loss value is negatively correlated with the first similarity. The second similarity is the similarity between the training text in different training samples and the first training text; the second loss value is positively correlated with the second similarity.

[0180] In a specific implementation, after obtaining the first similarity, the first loss value corresponding to the semantic encoding model to be trained can be determined according to the first similarity. And, the second loss value corresponding to the semantic encoding model to be trained can be determined according to the similarity between the training text in different training samples and the first training text, that is, according to the second similarity. Furthermore, the target loss value corresponding to the semantic encoding model to be trained can be obtained based on the first loss value and the second loss value.

[0181] Specifically, the training text belonging to the same training sample is paired with the first training image, which can be regarded as a positive sample. Since the two are paired, the semantic encoding features corresponding to the first training text in the positive sample and the text encoding features corresponding to the training text can have a relatively high similarity; while the training text that does not belong to different training samples is not paired with the first training image, which can be regarded as a negative sample. Since the two are not paired, the similarity between the training text in different training samples and the first training text is relatively low. For example, Figure 4As shown, a feature matrix can be generated based on multiple training samples. After inputting training text 1, training text 2... training text n into the semantic encoding model, corresponding text encoding features T1, T2... Tn can be obtained; after inputting the first training image 1, the first training image 2... the first training image n into the semantic encoding model, corresponding image encoding features I1, I2... In can be obtained. By calculating the inner products corresponding to each text encoding feature and semantic encoding feature, a corresponding matrix can be generated.

[0182] When training the semantic encoding model, the corresponding objective function can maximize the similarity of positive samples, that is, maximize the first similarity. At the same time, it can minimize the similarity of negative samples, that is, minimize the second similarity. Thus, the target loss value can be determined based on the first loss value corresponding to the first similarity and the second loss value corresponding to the second similarity.

[0183] In the present disclosure, based on the first similarity, the first loss value corresponding to the semantic encoding model to be trained can be obtained, and the first loss value is negatively correlated with the first similarity; at the same time, based on the second similarity, the second loss value corresponding to the semantic encoding model to be trained can be obtained. The second similarity is the similarity between the training text in different training samples and the first training text, and the second loss value is positively correlated with the second similarity; furthermore, the target loss value corresponding to the semantic encoding model to be trained can be obtained based on the first loss value and the second loss value. In the solution of the present disclosure, since there already exists a paired or unpaired relationship between the training text and the first training text, through the way of contrastive learning, self-supervised model training of the semantic encoding model can be realized, and it can automatically construct supervision information among a large number of paired or unpaired first training images and training texts, align the text modality and the image modality, and obtain a reliable semantic encoding model.

[0184] In an exemplary embodiment, the image encoding model in step S330 can be a pre-trained model with high confidence, that is, the image encoding model and the semantic encoding model can be trained separately. After training the image encoding model, based on the assistance of the image encoding model, the semantic encoding model is trained to align the text modality corresponding to the semantic encoding model with the image modality corresponding to the image encoding model.

[0185] In another example, the image encoding model can be jointly trained with the semantic encoding model. Taking the image encoding model as a CNN network as an example, the first training image can be input into the image encoding model to be trained for encoding, and based on the output of the image encoding model to be trained, the image encoding features corresponding to the first training image can be obtained. Then, based on the corresponding target loss value, the model parameters of the semantic encoding model and the image encoding model are adjusted respectively. When the training end condition is met, the trained semantic encoding model and image encoding model with aligned text modality and image modality are obtained.

[0186] In an exemplary embodiment, the set of original style features may be determined by a pre-trained style feature recognition model based on the input original image. The steps of training the style feature recognition model may include:

[0187] Obtain a second training image, input the second training image into the style feature recognition model to be trained, so as to obtain a set of predicted style features corresponding to the second training image through the style feature recognition model to be trained; generate a corresponding predicted image based on the set of predicted style features; obtain a third loss value corresponding to the style feature recognition model to be trained according to the difference between the predicted image and the second training image; adjust the model parameters of the style feature recognition model to be trained according to the third loss value until the training end condition is satisfied, and obtain the trained style feature recognition model.

[0188] Among them, the second training image may be an image for training the style feature recognition model. The set of predicted style features may include style features predicted by the style feature recognition model to be trained. The third loss value is positively correlated with the difference between the predicted image and the second training image.

[0189] In practical applications, a second training image may be obtained and input into the style feature recognition model to be trained, where the style feature recognition model to be trained may be a neural network model. After inputting the second training image, the style feature recognition model to be trained may analyze the second training image, predict the style features corresponding to the second training image, and obtain a set of predicted style features corresponding to the second training image. The set of predicted style features may include the style features corresponding to the second training image predicted by the style feature recognition model to be trained.

[0190] After obtaining the set of predicted style features, a corresponding predicted image may be generated based on the set of predicted style features. Furthermore, a third loss value corresponding to the style feature recognition model to be trained may be obtained according to the difference between the predicted image and the second training image. The second loss value is positively correlated with the difference, that is, the greater the difference, the greater the third loss value, and the smaller the difference, the smaller the third loss value. After obtaining the third loss value, the model parameters of the style feature recognition model to be trained may be adjusted according to the third loss value, so that the difference between the predicted image generated by the style feature recognition model to be trained and the second training image becomes smaller and smaller until the training end condition is satisfied, and the trained style feature recognition model may be obtained.

[0191] In the present disclosure, by training a style feature recognition model to be trained, a pre-trained style feature recognition model is obtained, which can quickly recognize the set of original style features corresponding to the original image after obtaining the original image, decouple the image data features of the original image to obtain multiple independent style features, and provide a basis for quickly and accurately adjusting the specified style of the original image.

[0192] In an exemplary embodiment, when generating the style transformation image corresponding to the image modification description information based on the target style feature set, the target style feature set can be input into a pre-trained image generation model, and the image generation model generates the style transformation image based on each style feature in the target style feature set.

[0193] In practical applications, the image generation model can be the generator in a generative adversarial network (GAN, Generative Adversarial Networks). Specifically, multiple training images for training the generative adversarial network can be obtained in advance. For example, various types of open-source image datasets such as face, sky, or landscape images can be crawled. After obtaining the image dataset, data augmentation can also be performed on the image dataset to obtain the augmented image data, and then the augmented image data and the image dataset can be used as the training images for training the generative adversarial network.

[0194] After obtaining the training images, a generator and a discriminator can be defined in advance. Among them, the generator can generate a corresponding image based on the input random noise distribution vector. For example, a random noise distribution that conforms to common distribution rules can be used as the input; and the discriminator can be used to determine whether the image generated by the generator is a real image, that is, whether it is the same as the training image. As an example, the generator and / or the discriminator can be at least one of the following types of neural networks: CNN, RNN, or fully connected neural network.

[0195] When training the generative adversarial network to be trained, M samples can be extracted from multiple training images, and the generator uses the input random noise distribution to generate M samples. At the beginning of the training, the generator can be fixed and the discriminator can be trained to make it identify the authenticity of the input image as much as possible. Here, the loss value of the discriminator is determined based on the recognition result, and the model parameters of the discriminator are adjusted according to this loss value.

[0196] After the discriminator is cyclically updated K times, the discriminator can be fixed and the generator can be trained to make the discriminator unable to distinguish the authenticity of the images generated by the generator as much as possible. Based on this, the loss value of the generator is determined, and the model parameters of the generator are adjusted. When the discriminator cannot distinguish the authenticity of the images generated by the generator, that is, the discrimination probability is 0.5, the training can be stopped, and the current generator is used as the image generation model.

[0197] In this embodiment, the generator can be a style generator, and a mapping network can be set in the style generator to map the input random noise distribution into style features that control the image style. Furthermore, the style generator can generate corresponding images based on the style features obtained after mapping. Through this mapping network, the style generator can generate a style feature that does not need to follow the random noise distribution, effectively reducing the correlation between image data features and achieving feature decoupling.

[0198] Furthermore, the style generator can also include a style module to achieve the effect of adaptive instance normalization. Moreover, during the process of generating images, random noise can be added to each input layer of the style generator so that the style generator can generate more diverse random details.

[0199] After obtaining the trained generator, when training the style feature recognition model to be trained, the predicted style feature set can also be input into this generator, and the corresponding predicted images can be generated through this generator.

[0200] In an exemplary embodiment, before receiving the image modification description information corresponding to the original image to be subjected to style transformation, it may further include:

[0201] Obtain the video to be processed, and determine the target video frame in the video as the original image to be subjected to style transformation.

[0202] In practical applications, the video to be processed can be obtained, such as a video to be subjected to style change or content modification. Furthermore, the target video in the video can be determined as the original image to be subjected to style transformation.

[0203] Specifically, after the server 120 obtains the video to be processed, it can perform frame extraction on the video to be processed at a preset interval, and determine the extracted video frame as the target video frame. Alternatively, the user can send the video to be processed to the server 120 through the terminal 110. When sending the video to be processed, a video modification instruction carrying time information can be sent to the server 120 at the same time. Then, after receiving the video modification instruction, the server 120 can determine the target video frame in the video to be processed according to the time information in the video modification instruction.

[0204] After generating the style-transformed image corresponding to the image modification description information based on the target style feature set, the following steps may further be included:

[0205] Based on the style-transformed image and other video frames in the video, obtain a target video.

[0206] Wherein, the other video frames are video frames other than the target video frame in the video.

[0207] After obtaining the style-transformed image, it is possible to re-assemble based on the style-transformed image and other video frames in the video to obtain a target video.

[0208] In the present disclosure, by obtaining a video to be processed and determining a target video frame in the video as the original image to be subjected to style transformation, and then, after obtaining the style-transformed image, a target video can be obtained based on the style-transformed image and other video frames in the video. Through a human-computer interaction method, the style of the video content is automatically transformed without manual intervention throughout the process, avoiding the need for the user to edit a large number of videos in the video one by one, and effectively improving the video modification efficiency.

[0209] To enable those skilled in the art to better understand the above steps, the following exemplarily illustrates the embodiments of the present application through an example, but it should be understood that the embodiments of the present application are not limited thereto.

[0210] As Figure 5 shown, after obtaining the original image, such as an image containing a "cat", the original image can be input into a pre-trained style feature recognition model, and the original style feature set S corresponding to the original image can be obtained through the style feature recognition model. If the currently obtained original style feature set S is input into the trained image generation model, the original image can be obtained.

[0211] Meanwhile, the image modification description information input by the user, for example, "lovely cat" as the image modification description information, can be input into the semantic encoding model. Through the semantic encoding model, the semantic encoding information Δt corresponding to the image modification description information can be obtained, and based on the mapping relationship between the pre-determined image change feature space and the text semantic feature space, the style feature change information Δs corresponding to Δt can be obtained.

[0212] After obtaining the original style feature set S and the style feature change information Δs, the original style feature set S can be adjusted based on the style feature change information Δs to obtain a target style feature set, that is, S + Δs. Furthermore, the target style feature set can be input into the image generation model to obtain a style-transformed image.

[0213] It should be understood that although Figure 2 and Figure 3The steps in the flowchart are shown in sequence according to the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 2 and Figure 3 at least some of the steps in

[0214] can include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least some of the steps or stages in other steps or other steps.

[0215] Figure 6 is a block diagram of an image processing apparatus 600 shown according to an exemplary embodiment. Referring to Figure 6 , the apparatus includes a description information acquisition unit 601, a semantic encoding unit 602, a style transformation determination unit 603, an original feature set acquisition unit 604, a target feature set acquisition unit 605, and a style transformation image acquisition unit 606.

[0216] The description information acquisition unit 601 is configured to receive image modification description information corresponding to an original image to be subjected to style transformation;

[0217] The semantic encoding unit 602 is configured to perform semantic encoding on the image modification description information to obtain semantic encoding information corresponding to the image modification description information;

[0218] The style transformation determination unit 603 is configured to obtain style feature change information corresponding to the semantic encoding information, where the style feature change information is feature change information corresponding to an image style;

[0219] The original feature set acquisition unit 604 is configured to obtain an original style feature set corresponding to the original image;

[0220] The target feature set acquisition unit 605 is configured to adjust the original style features in the original style feature set based on the style feature change information to obtain a target style feature set;

[0221] The style transformation image acquisition unit 606 is configured to generate a style transformation image corresponding to the image modification description information based on the target style feature set.

[0222] In an exemplary embodiment, the target feature set acquisition unit is configured to perform:

[0223] Adjust the first original style feature in the original style feature set based on the style feature change information to obtain an adjusted first style feature; the first original style feature is the original style feature corresponding to the style feature change information;

[0224] Combine the second original style feature and the adjusted first style feature to obtain the target style feature set, where the second original style feature is the original style feature other than the first original style feature in the original style feature set.

[0225] In an exemplary embodiment, the style transformation determination unit is configured to perform:

[0226] Obtain the mapping relationship between the pre-determined image change feature space and the text semantic feature space;

[0227] Use the semantic encoding information as the text semantic feature in the text semantic feature space, and obtain the image change feature corresponding to the text semantic feature in the image change feature space based on the mapping relationship;

[0228] Use the image change feature as the style feature change information.

[0229] In an exemplary embodiment, the semantic encoding unit is configured to perform:

[0230] Input the image modification description information into a pre-trained semantic encoding model to encode the image modification description information based on the semantic encoding parameters in the semantic encoding model, and obtain the semantic encoding information corresponding to the image modification description information;

[0231] The semantic encoding model is trained based on paired training texts and training images.

[0232] In an exemplary embodiment, the device further includes:

[0233] The first training image acquisition unit is configured to acquire training samples, where the training samples include a first training image and a training text paired with the first training image;

[0234] The training text encoding unit is configured to input the training text into the semantic encoding model to be trained for encoding to obtain the text encoding feature of the training text;

[0235] A first training image encoding unit, configured to perform encoding the first training image into an image encoding model to obtain an image encoding feature of the first training image;

[0236] A first similarity obtaining unit, configured to perform determining a first similarity between each of the first training images and each of the training texts according to the image encoding feature of the first training image and the text encoding feature of the training text; the first similarity characterizes the similarity between the encoding feature corresponding to the image modality and the encoding feature corresponding to the text modality;

[0237] A target loss value obtaining unit, configured to perform determining a target loss value corresponding to the semantic encoding model to be trained according to the first similarity;

[0238] A first parameter adjustment unit, configured to perform adjusting model parameters of the semantic encoding model to be trained according to the target loss value until a training end condition is satisfied, to obtain the trained semantic encoding model.

[0239] In an exemplary embodiment, there are multiple training samples, and the target loss value obtaining unit is configured to perform:

[0240] Obtaining a first loss value corresponding to the semantic encoding model to be trained according to the first similarity, where the first loss value has a negative correlation with the first similarity;

[0241] Obtaining a second loss value corresponding to the semantic encoding model to be trained according to a second similarity, where the second loss value has a positive correlation with the second similarity; the second similarity is the similarity between the training texts in different training samples and a first training text;

[0242] Based on the first loss value and the second loss value, obtaining the target loss value corresponding to the semantic encoding model to be trained.

[0243] In an exemplary embodiment, the original style feature set is determined by a pre-trained style feature recognition model based on the input original image, and the apparatus further includes:

[0244] A second training image obtaining unit, configured to perform obtaining a second training image, inputting the second training image into the style feature recognition model to be trained, so as to obtain a predicted style feature set corresponding to the second training image through the style feature recognition model to be trained;

[0245] A predicted image generating unit, configured to perform generating a corresponding predicted image based on the predicted style feature set;

[0246] A third loss value obtaining unit, configured to obtain a third loss value corresponding to the style feature recognition model to be trained according to a difference between the predicted image and the second training image, where the third loss value has a positive correlation with the difference;

[0247] A second parameter adjustment unit, configured to adjust model parameters of the style feature recognition model to be trained according to the third loss value until a training end condition is met, so as to obtain the trained style feature recognition model.

[0248] In an exemplary embodiment, the apparatus further includes:

[0249] A video obtaining unit, configured to obtain a video to be processed and determine a target video frame in the video as an original image to be subjected to style transformation;

[0250] The apparatus further includes:

[0251] A video updating unit, configured to obtain a target video based on the style transformation image and other video frames in the video; the other video frames are video frames in the video other than the target video frame.

[0252] Regarding the apparatus in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0253] Figure 7 It is a block diagram of an electronic device 700 for performing an image processing method according to an exemplary embodiment. For example, the electronic device 700 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0254] Referring to Figure 7 , the electronic device 700 may include one or more of the following components: a processing component 702, a memory 704, a power component 706, a multimedia component 708, an audio component 710, an input / output (I / O) interface 712, a sensor component 714, and a communication component 716.

[0255] The processing component 702 generally controls the overall operation of the electronic device 700, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 702 may include one or more processors 720 to execute instructions to complete all or part of the steps of the above-mentioned methods. In addition, the processing component 702 may include one or more modules to facilitate the interaction between the processing component 702 and other components. For example, the processing component 702 may include a multimedia module to facilitate the interaction between the multimedia component 708 and the processing component 702.

[0256] The memory 704 is configured to store various types of data to support the operation of the electronic device 700. Examples of such data include instructions for any application or method operating on the electronic device 700, contact data, phone book data, messages, pictures, videos, etc. The memory 704 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disks, optical disks, or graphene memory.

[0257] The power component 706 provides power to various components of the electronic device 700. The power component 706 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the electronic device 700.

[0258] The multimedia component 708 includes a screen that provides an output interface between the electronic device 700 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may not only sense the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 708 includes a front camera and / or a rear camera. When the electronic device 700 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each of the front camera and the rear camera may be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0259] The audio component 710 is configured to output and / or input audio signals. For example, the audio component 710 includes a microphone (MIC), which is configured to receive external audio signals when the electronic device 700 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 704 or transmitted via the communication component 716. In some embodiments, the audio component 710 further includes a speaker for outputting audio signals.

[0260] The I / O interface 712 provides an interface between the processing component 702 and a peripheral interface module, which may be a keyboard, a click wheel, buttons, etc. These buttons may include, but are not limited to: a home button, a volume button, a power button, and a lock button.

[0261] The sensor component 714 includes one or more sensors for providing an assessment of the state of the electronic device 700 in various aspects. For example, the sensor component 714 can detect the on / off state of the electronic device 700, the relative positioning of components, such as the display and keypad of the electronic device 700. The sensor component 714 can also detect a change in the position of the electronic device 700 or an electronic device 700 component, the presence or absence of user contact with the electronic device 700, the orientation or acceleration / deceleration of the device 700, and a change in the temperature of the electronic device 700. The sensor component 714 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 714 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 714 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0262] The communication component 716 is configured to facilitate communication between the electronic device 700 and other devices in a wired or wireless manner. The electronic device 700 can access a wireless network based on communication standards, such as WiFi, a carrier network (such as 2G, 3G, 4G, or 5G), or a combination thereof. In an exemplary embodiment, the communication component 716 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 716 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0263] In an exemplary embodiment, the electronic device 700 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.

[0264] In an exemplary embodiment, there is also provided a computer-readable storage medium including instructions, such as a memory 704 including instructions, and the above instructions can be executed by a processor 720 of the electronic device 700 to complete the above method. For example, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0265] In an exemplary embodiment, there is also provided a computer program product, and the computer program product includes instructions, and the above instructions can be executed by a processor 720 of the electronic device 700 to complete the above method.

[0266] Figure 8 FIG. 800 is a block diagram of an electronic device 800 for performing an image processing method according to an exemplary embodiment. For example, the electronic device 800 may be a server. Referring to Figure 8 , the electronic device 800 includes a processing component 820, which further includes one or more processors, and memory resources represented by a memory 822 for storing instructions executable by the processing component 820, such as application programs. The application programs stored in the memory 822 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 820 is configured to execute instructions to perform the above method.

[0267] The electronic device 800 may further include: a power component 824 configured to perform power management of the electronic device 800, a wired or wireless network interface 826 configured to connect the electronic device 800 to a network, and an input / output (I / O) interface 828. The electronic device 800 may operate based on an operating system stored in the memory 822, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, or the like.

[0268] In an exemplary embodiment, there is also provided a computer-readable storage medium including instructions, such as a memory 822 including instructions, and the above instructions can be executed by a processor of the electronic device 800 to complete the above method. The storage medium may be a computer-readable storage medium. For example, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0269] In an exemplary embodiment, a computer program product is further provided. The computer program product includes instructions that can be executed by a processor of the electronic device 800 to implement the above method.

[0270] It should be noted that the above-mentioned device, electronic device, computer-readable storage medium, computer program product, etc. may also include other implementation manners according to the description of the method embodiments. The specific implementation manners can refer to the description of the relevant method embodiments and will not be elaborated here one by one.

[0271] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.

[0272] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. An image processing method, characterized in that, Including: Receiving image modification description information corresponding to an original image to be subjected to style transformation; Inputting the image modification description information into a semantic encoding model to encode the image modification description information based on semantic encoding parameters in the semantic encoding model, so as to obtain semantic encoding information corresponding to the image modification description information; the semantic encoding model is trained based on paired training texts and training images; Obtaining style feature change information corresponding to the semantic encoding information, where the style feature change information is feature change information corresponding to an image style; Obtaining an original style feature set corresponding to the original image by a style feature recognition model; the style feature recognition model is obtained by adjusting model parameters of a style feature recognition model to be trained according to a third loss value, the third loss value is determined according to the difference between a predicted image and a second training image and is positively correlated with the difference, the predicted image is generated according to a predicted style feature set, and the predicted style feature set is extracted from the second training image by the style feature recognition model to be trained; Adjusting original style features in the original style feature set based on the style feature change information to obtain a target style feature set; Generating a style transformation image corresponding to the image modification description information based on the target style feature set.

2. The method according to claim 1, wherein The adjusting original style features in the original style feature set based on the style feature change information to obtain a target style feature set includes: Adjusting a first original style feature in the original style feature set based on the style feature change information to obtain an adjusted first style feature; the first original style feature is the original style feature corresponding to the style feature change information; Combining a second original style feature and the adjusted first style feature to obtain the target style feature set, where the second original style feature is an original style feature other than the first original style feature in the original style feature set.

3. The method according to claim 1, wherein The obtaining style feature change information corresponding to the semantic encoding information includes: Obtaining a pre-determined mapping relationship between an image change feature space and a text semantic feature space; Taking the semantic encoding information as text semantic features in the text semantic feature space, and obtaining image change features corresponding to the text semantic features in the image change feature space based on the mapping relationship; Taking the image change features as the style feature change information.

4. The method according to claim 1, wherein The steps of training to obtain the semantic encoding model include: Obtaining training samples, where the training samples include first training images and training texts paired with the first training images; Inputting the training texts into a semantic encoding model to be trained for encoding to obtain text encoding features of the training texts; Inputting the first training images into an image encoding model for encoding to obtain image encoding features of the first training images; Determine the first similarity between each of the first training images and each of the training texts according to the image coding features of the first training images and the text coding features of the training texts; the first similarity characterizes the similarity between the coding features corresponding to the image modality and the coding features corresponding to the text modality; Determine the target loss value corresponding to the semantic coding model to be trained according to the first similarity; Adjust the model parameters of the semantic coding model to be trained according to the target loss value until the training end condition is satisfied, and obtain the trained semantic coding model.

5. The method according to claim 4, characterized in that, There are multiple training samples, and determining the target loss value corresponding to the semantic coding model to be trained according to the first similarity includes: Obtain the first loss value corresponding to the semantic coding model to be trained according to the first similarity, and the first loss value is negatively correlated with the first similarity; Obtain the second loss value corresponding to the semantic coding model to be trained according to the second similarity, and the second loss value is positively correlated with the second similarity; the second similarity is the similarity between the training texts in different training samples and the first training text; Based on the first loss value and the second loss value, obtain the target loss value corresponding to the semantic coding model to be trained.

6. The method according to claim 1, wherein The steps of training the style feature recognition model include: Obtain a second training image, and input the second training image into the style feature recognition model to be trained, so as to obtain a predicted style feature set corresponding to the second training image through the style feature recognition model to be trained; Generate a corresponding predicted image based on the predicted style feature set; Obtain the third loss value corresponding to the style feature recognition model to be trained according to the difference between the predicted image and the second training image, and the third loss value is positively correlated with the difference; Adjust the model parameters of the style feature recognition model to be trained according to the third loss value until the training end condition is satisfied, and obtain the trained style feature recognition model.

7. The method according to claim 1, characterized in that Before receiving the image modification description information corresponding to the original image to be subjected to style transformation, it further includes: Obtain the video to be processed, and determine the target video frame in the video as the original image to be subjected to style transformation; After generating the style transformation image corresponding to the image modification description information based on the target style feature set, it further includes: Obtain the target video based on the style transformation image and other video frames in the video; the other video frames are video frames in the video other than the target video frame.

8. An image processing apparatus, characterized in that, It includes: A description information acquisition unit configured to receive the image modification description information corresponding to the original image to be subjected to style transformation; A semantic coding unit configured to input the image modification description information into the semantic coding model, and encode the image modification description information based on the semantic coding parameters in the semantic coding model to obtain the semantic coding information corresponding to the image modification description information; The semantic encoding model is trained based on paired training texts and training images; A style transformation determination unit, configured to execute obtaining style feature change information corresponding to the semantic encoding information, where the style feature change information is feature change information corresponding to an image style; An original feature set obtaining unit, configured to execute obtaining an original style feature set corresponding to the original image by a style feature recognition model; the style feature recognition model is obtained by adjusting model parameters of a style feature recognition model to be trained according to a third loss value, the third loss value is determined according to the difference between a predicted image and a second training image and is in a positive correlation relationship with the difference, the predicted image is generated according to a predicted style feature set, and the predicted style feature set is extracted from the second training image by the style feature recognition model to be trained; A target feature set obtaining unit, configured to execute adjusting original style features in the original style feature set based on the style feature change information to obtain a target style feature set; A style transformation image obtaining unit, configured to execute generating a style transformation image corresponding to the image modification description information based on the target style feature set.

9. The device according to claim 8, wherein, The target feature set obtaining unit is configured to execute: Adjusting a first original style feature in the original style feature set based on the style feature change information to obtain an adjusted first style feature; the first original style feature is the original style feature corresponding to the style feature change information; Combining a second original style feature and the adjusted first style feature to obtain the target style feature set, where the second original style feature is the original style feature other than the first original style feature in the original style feature set.

10. The device according to claim 8, characterized in that, The style transformation determination unit is configured to execute: Obtaining a mapping relationship between a pre-determined image change feature space and a text semantic feature space; Taking the semantic encoding information as text semantic features in the text semantic feature space, and obtaining image change features corresponding to the text semantic features in the image change feature space based on the mapping relationship; Taking the image change features as the style feature change information.

11. The device according to claim 8, characterized in that, The apparatus further includes: A first training image obtaining unit, configured to execute obtaining a training sample, where the training sample includes a first training image and a training text paired with the first training image; A training text encoding unit, configured to execute inputting the training text into a semantic encoding model to be trained for encoding to obtain text encoding features of the training text; A first training image encoding unit, configured to execute inputting the first training image into an image encoding model for encoding to obtain image encoding features of the first training image; The first similarity obtaining unit is configured to determine a first similarity between each of the first training images and each of the training texts according to the image encoding features of the first training images and the text encoding features of the training texts; the first similarity characterizes the similarity between the encoding features corresponding to the image modality and the encoding features corresponding to the text modality; The target loss value obtaining unit is configured to determine a target loss value corresponding to the semantic encoding model to be trained according to the first similarity; The first parameter adjustment unit is configured to adjust the model parameters of the semantic encoding model to be trained according to the target loss value until the training end condition is satisfied, and obtain the trained semantic encoding model.

12. The device according to claim 8, characterized in that There are multiple training samples, and the target loss value obtaining unit is configured to perform: According to the first similarity, obtain a first loss value corresponding to the semantic encoding model to be trained, and the first loss value is negatively correlated with the first similarity; According to a second similarity, obtain a second loss value corresponding to the semantic encoding model to be trained, and the second loss value is positively correlated with the second similarity; the second similarity is the similarity between the training texts in different training samples and the first training text; Based on the first loss value and the second loss value, obtain the target loss value corresponding to the semantic encoding model to be trained.

13. The device according to claim 8, characterized in that The apparatus further includes: The second training image obtaining unit is configured to obtain a second training image, and input the second training image into the style feature recognition model to be trained, so as to obtain a set of predicted style features corresponding to the second training image through the style feature recognition model to be trained; The predicted image generating unit is configured to generate a corresponding predicted image based on the set of predicted style features; The third loss value obtaining unit is configured to obtain a third loss value corresponding to the style feature recognition model to be trained according to the difference between the predicted image and the second training image, and the third loss value is positively correlated with the difference; The second parameter adjustment unit is configured to adjust the model parameters of the style feature recognition model to be trained according to the third loss value until the training end condition is satisfied, and obtain the trained style feature recognition model.

14. The device according to claim 8, characterized in that, The apparatus further includes: The video obtaining unit is configured to obtain a video to be processed, and determine a target video frame in the video as the original image to be subjected to style transformation; The apparatus further includes: The video updating unit is configured to obtain a target video based on the style-transformed image and other video frames in the video; the other video frames are video frames in the video other than the target video frame.

15. An electronic device, characterized in that, Includes: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the image processing method according to any one of claims 1 to 7.

16. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the image processing method according to any one of claims 1 to 7.

17. A computer program product comprising instructions, characterized in that, When the instructions are executed by a processor of an electronic device, the electronic device is enabled to execute the image processing method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Interactive image editing method and device, readable storage medium and electronic equipment

    CN113448477A