Image editing method and device, computer equipment, storage medium and computer program product

By adjusting the visual impact of pixels in the image so that it is negatively correlated with the impact of the text, the problems of text editing accuracy and natural boundary transition in local image editing are solved, achieving a more delicate and realistic editing effect.

CN120765804APending Publication Date: 2025-10-10BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510646215.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

It is difficult for existing technologies to balance text editing accuracy and natural boundary transitions in local image editing.

Method used

By obtaining the text influence of the local editing description text on multiple pixels in the original image, the visual influence of the pixels in the original image is adjusted so that the target visual influence is negatively correlated with the text influence, and the target image is generated.

Benefits of technology

It improves the accuracy of text editing and the naturalness of boundary transitions, enhances the natural transition between local editing content and original image content, and improves the authenticity and fidelity of editing results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120765804A_ABST
    Figure CN120765804A_ABST
Patent Text Reader

Abstract

The invention relates to an image editing method and device, electronic equipment, a storage medium and a computer program product, relates to the field of artificial intelligence, and can give consideration to text editing accuracy and boundary transition naturalness when an image is locally edited. The method comprises the following steps: acquiring an original image to be locally edited and a local editing description text; determining the text influence degree of the local editing description text on the subsequent editing of a plurality of pixels in the original image; according to the text influence degree, adjusting the visual influence degree of the original image on subsequent editing of a plurality of pixels in the original image to obtain a target visual influence degree, and according to the original image and the target visual influence degree, determining image overall features of the original image; the image overall feature is used for generating a target image after the original image is locally edited, and the target visual influence degree is in negative correlation with the text influence degree; and generating a target image after the original image is locally edited according to the image overall features.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of artificial intelligence, and particularly relates to an image editing method and device, computer equipment, a storage medium and a computer program product. BACKGROUND

[0002] With the development of computer technology and artificial intelligence technology, various images can be conveniently generated through text instructions, for example, a specified manner of editing a certain specific region in an image can be limited by using a text instruction.

[0003] However, in practice, it is found that in the related art, when performing local editing, there may be a situation that the editing effect is different from the text instruction, or a situation that the redrawn content matches the text instruction but the transition boundary is unnatural. It can be seen that the related art is difficult to effectively balance the text editing accuracy and the natural boundary transition of the local editing region. SUMMARY

[0004] The present disclosure provides an image editing method, device, computer equipment, storage medium and computer program product to at least solve the problem that the related art is difficult to balance the text editing accuracy and the natural boundary transition. The technical solutions of the present disclosure are as follows:

[0005] According to a first aspect of an embodiment of the present disclosure, an image editing method is provided, comprising:

[0006] obtaining an original image to be locally edited and a local editing description text;

[0007] determining a text influence degree of the local editing description text on subsequent editing of a plurality of pixels in the original image;

[0008] According to the text influence degree, adjusting a visual influence degree of the original image on subsequent editing of a plurality of pixels in the original image to obtain a target visual influence degree, and determining an image overall feature of the original image according to the original image and the target visual influence degree; the image overall feature is used to generate a target image after local editing of the original image, and the target visual influence degree is negatively correlated with the text influence degree;

[0009] generating the target image according to the image overall feature.

[0010] In one embodiment, the determination of the text influence degree of the local editing description text on the subsequent editing of the plurality of pixels in the original image comprises:

[0011] Acquire semantic unit text features corresponding to each of the plurality of semantic units in the partial edit description text, and determine overall text features corresponding to the partial edit description text based on the semantic unit text features;

[0012] The image editing features corresponding to the original image are obtained, and a trained text influence degree prediction module is used to determine the text influence degree of each of multiple pixels in the original image based on the feature splicing results of the image editing features and the overall text features; the image editing features are features determined from the edited image during the process of pre-editing the original image according to the local editing description text.

[0013] In one embodiment, obtaining the image editing features corresponding to the original image includes:

[0014] Obtaining a randomly generated noise image and a mask image obtained by masking the local edited region marked in the original image, and inputting the noise image, the mask image, and the local edit description text into a trained image generation module; the image generation module includes a plurality of cascaded feature processing units, and the image generation module is configured to perform denoising processing on the noise image based on the mask image and the local edit description text to obtain an image in which the local edited region is edited according to the local edit description text;

[0015] Each of the feature processing units performs denoising processing on the input content based on the denoising prompt information determined by the mask image and the local edit description text to obtain denoising image features, and uses the denoising image features as input content for the next feature processing unit; the input content of the first feature processing unit includes the noise image, and the denoising processing of the multiple feature processing units is different;

[0016] An image editing feature corresponding to the original image is determined based on the plurality of denoised image features.

[0017] In one embodiment, adjusting the visual impact of the original image on subsequent editing of multiple pixels in the original image based on the text impact to obtain a target visual impact, and determining the overall image features of the original image based on the original image and the target visual impact includes:

[0018] Acquiring a visual impact degree set of the original image, wherein the visual impact degree set includes a visual impact degree of the original image on subsequent editing of each pixel of the original image;

[0019] adjusting multiple visual impact levels in the visual impact level set according to the text impact level of each pixel in the local editing area, and determining a target visual impact level for each pixel in the original image based on the adjusted visual impact level set; the local editing area being the local area marked to be edited in the original image;

[0020] The overall image features of the original image are determined according to the target visual impact degree of each pixel of the original image and the pre-acquired pixel features of each pixel of the original image.

[0021] In one embodiment, adjusting the multiple visual impact levels in the visual impact level set according to the text impact level of each pixel in the local editing area includes:

[0022] In the visual impact degree set, determining respective original visual impact degrees corresponding to pixels in the local editing area;

[0023] For each of the original visual impact levels, determining a downward adjustment range of the original visual impact level according to the text impact level of the corresponding pixel, and determining an adjusted visual impact level according to the downward adjustment range and the original visual impact level; the downward adjustment range is positively correlated with the text impact level;

[0024] The original visual impact levels in the visual impact level set are updated using the adjusted visual impact levels to obtain an adjusted visual impact level set.

[0025] In one embodiment, before obtaining the original image to be partially edited and the partial editing description text, the method further includes:

[0026] Inputting a pre-acquired sample image, a sample mask image corresponding to the sample image, and attribute description text for a local image region in the sample image into the image editing model to be trained; the sample mask image is obtained by performing masking processing on the local image region in the sample image;

[0027] Determining, by the image editing model, respective text influence degrees of a plurality of pixels in the sample image, determining respective sample visual influence degrees of the plurality of pixels in the sample image based on the respective text influence degrees of the plurality of pixels in the sample image, determining, based on the sample image and the sample visual influence degrees, a sample image overall feature corresponding to the sample image, and generating, based on the sample image overall feature, an edited image that performs a partial edit on the sample image;

[0028] According to the difference between the edited image and the sample image, the model parameters of the image editing model are adjusted to obtain a trained image editing model; the image editing model is used to generate the target image based on the input original image and the local editing description text.

[0029] In one embodiment, before inputting the pre-acquired sample image, the sample mask image corresponding to the sample image, and the attribute description text for the local area of ​​the sample image into the image editing model to be trained, the method further includes:

[0030] Acquire a sample image, and segment the sample image using a trained image segmentation model to obtain at least one local image region in the sample image; image contents in the same local image region have matching attributes;

[0031] Masking is performed on each local area of ​​the image to obtain a sample mask image corresponding to the sample image; and the sample image and the segmentation result of the sample image are input into a trained language model to obtain attribute description text output by the language model for each local area of ​​the image.

[0032] According to a second aspect of an embodiment of the present disclosure, there is provided an image editing apparatus, comprising:

[0033] An editing information acquiring unit configured to acquire an original image to be partially edited and a local editing description text;

[0034] a text impact degree determining unit configured to determine a text impact degree of the local edit description text on subsequent edits of a plurality of pixels in the original image;

[0035] a visual impact adjustment unit configured to adjust the visual impact of subsequent editing of the original image on a plurality of pixels in the original image based on the text impact to obtain a target visual impact, and determine an overall image feature of the original image based on the original image and the target visual impact; the overall image feature is used to generate a target image after the original image is partially edited, and the target visual impact is negatively correlated with the text impact;

[0036] The image generating unit is configured to generate the target image according to the overall features of the image.

[0037] According to a third aspect of an embodiment of the present disclosure, a computer device is provided, including:

[0038] processor;

[0039] a memory for storing instructions executable by the processor;

[0040] The processor is configured to execute the instructions to implement any of the above-mentioned image editing methods.

[0041] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of a computer device, the computer device is enabled to execute any of the image editing methods described above.

[0042] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, which includes instructions. When the instructions are executed by a processor of a computer device, the computer device is able to execute the image editing method as described in any one of the above items.

[0043] The technical solutions provided by the embodiments of the present disclosure bring at least the following beneficial effects:

[0044] By adjusting the visual impact of the original image on subsequent editing of multiple pixels in the original image according to the degree of text influence, and making the target visual impact negatively correlated with the text influence, on the one hand, it can reduce the visual dependence on the original image when the text influence is high, and improve the accuracy of text editing. On the other hand, it can increase the attention and use of the visual information of the original image content when the text influence is low, so that the local edited content can be naturally transitioned to the original image content, effectively taking into account both the accuracy of text editing and the naturalness of boundary transitions.

[0045] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0047] Figure 1 It is a flowchart of partial image redrawing in related technology.

[0048] Figure 2 A diagram showing an application environment of an image editing method according to an exemplary embodiment.

[0049] Figure 3 The figure is a flowchart of an image editing method according to an exemplary embodiment.

[0050] Figure 4 The present invention is a flowchart showing a step of determining the impact degree of a text according to an exemplary embodiment.

[0051] Figure 5 is a framework schematic diagram of an image editing model according to an exemplary embodiment.

[0052] Figure 6 is a step flow chart of determining an overall feature of an image according to an exemplary embodiment.

[0053] Figure 7 is a schematic diagram of constructing a triple training sample according to an exemplary embodiment.

[0054] Figure 8 is a flow chart of another image editing method according to an exemplary embodiment.

[0055] Figure 9 is a block diagram of an image editing apparatus according to an exemplary embodiment.

[0056] Figure 10 is a block diagram of a computer device according to an exemplary embodiment. DETAILED DESCRIPTION

[0057] In order for those skilled in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings.

[0058] It should be noted that the terms "first", "second", and the like in the specification and claims of the present disclosure and the above-described drawings are used to distinguish similar objects, and do not necessarily have to describe a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein. The implementation described in the following exemplary embodiments does not represent all the implementations consistent with the present disclosure. Rather, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0059] It should also be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, analyzed data, etc.) involved in the present disclosure are all information and data authorized by the user or authorized by all parties.

[0060] In order for those skilled in the art to better understand the present disclosure, the related art will be introduced first.

[0061] With the development of computer technology and artificial intelligence technology, various images can be easily generated through text instructions. For example, text instructions can be used to limit the editing of a specific area in an image in a specified way and redraw the image within the limited area.

[0062] In some related technologies, based on the original image and text instructions provided by the user, an inversion-type editing method can be used to generate an edited image through image noise reduction. However, the final image obtained by the inversion-type method often has difficulty in accurately aligning the image content with delicate text describing the local part of the image (such as "sparse eyebrows"), resulting in the editing effect deviating from expectations or being too conservative.

[0063] In other related technologies, attempts are made to improve the accuracy of local text editing by local redrawing, for example Figure 1 In the partial redraw framework shown, the model receives a masked editing area (Mask) defined by the user in the original image and the given instruction text, generates the edited content corresponding to the text within the masked editing area, and finally uses the blurred masked editing area (Blurred Mask) to perform boundary transitions to obtain the edited image. However, this approach fails to explicitly model the relationship between the original image details and the editing area, resulting in loss of detail or visual discontinuity at the transition boundary.

[0064] It can be seen that the related art has difficulty in effectively balancing text editing accuracy and the smooth transition of local editing area boundaries. Therefore, the present disclosure provides an image editing method, apparatus, computer device, storage medium, and computer program product to at least address the problem of the related art in balancing text editing accuracy and detail consistency.

[0065] In an exemplary embodiment, the image editing method provided by the present disclosure can be applied to Figure 2 In the application environment shown, a terminal can communicate with a server via a network. The server can have a corresponding data storage system that can store data that the server needs to process, such as an original image to be partially edited and associated text describing the partial edits. In actual applications, the data storage system can be integrated with the server or placed on a cloud or other network server.

[0066] In an optional embodiment, the terminal can obtain the original image to be partially edited and the local editing description text based on the information input by the image editing account, and then transmit the original image and the local editing description text to the server; after obtaining the original image and the local editing description text, the server can determine the degree of influence of the local editing description text on the text of subsequent editing of multiple pixels in the original image, and then adjust the visual influence of the original image on the subsequent editing of multiple pixels in the original image according to the text influence degree to obtain the target visual influence degree, and determine the overall image features of the original image according to the original image and the target visual influence degree, wherein the overall image features are used to generate a target image after the original image is partially edited, and the target visual influence degree is negatively correlated with the text influence degree. Furthermore, the server can generate a corresponding target image according to the overall image features, and then return the target image as a processing result to the terminal.

[0067] It is understood that one or more steps in the image editing method provided herein can be performed by a terminal. In one embodiment, the terminal can completely edit the original image based on the local editing description text. In another embodiment, the terminal and server can cooperate to perform different image editing steps to generate the target image. For example, the server can extract the overall image features and then return the overall image features to the terminal, which then generates the target image based on the overall image features.

[0068] For example, the terminal can be, but is not limited to, various personal computers, laptops, smart phones, tablets, Internet of Things devices and portable wearable devices. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc.; the server can be implemented as an independent server or a server cluster consisting of multiple servers.

[0069] Figure 3 is a flow chart of an image editing method according to an exemplary embodiment. Figure 3 As shown, the image editing method can be applied to a server and includes the following steps.

[0070] In step S310, the original image to be partially edited and the partial editing description text are obtained.

[0071] Among them, the original image can be any image to be locally edited, and the area on the original image to be locally edited can be called the local editing area. In some examples, the original image can be one or more of a face image, a landscape image, an animal image, a plant image, and a medical image.

[0072] The local editing description text can be text that indicates the editing method or editing effect of the local editing area in the original image. For example, the object attributes or image effects that the object in the local editing area is expected to have can be determined, such as "sparse eyebrows", "thick lips" or "clothes with bright colors", and then recorded in text form to obtain the local editing description text.

[0073] In a specific implementation, the terminal can obtain the original image and local edit description text provided by the image editing account, where the image editing account can be any account used when performing image editing. In some optional embodiments, the image editing account can select the current captured real-life image or historical images stored in the terminal as the original image to be locally edited, and provide the associated local edit description text to the terminal via voice input or text input. The terminal can then transmit the obtained original image and local edit description text to the server.

[0074] In step S320, the degree of influence of the local edit description text on the subsequent edit text of multiple pixels in the original image is determined.

[0075] Specifically, local edit description text, which indicates the editing method or editing effect of a local edited area in the original image, often provides different guidance for subsequent editing of pixels in various regions of the original image. For example, taking the local edit description text "Sparse Eyebrows" as an example, since this text indicates that the image edit is a local edit of "eyebrows," areas in the original image whose image semantics are related to eyebrows can pay more attention to the editing effect specified by the local edit description text during subsequent editing. Areas whose image semantics are unrelated to eyebrows (such as the lips or nose) or have a low correlation with eyebrows (such as the area around the eyebrows) can pay less or no attention to the local edit description text during subsequent editing.

[0076] In this regard, after acquiring the original image and the local edit description text, this embodiment can analyze multiple pixels in the original image to determine the degree of textual influence of the local edit description text on the multiple pixels in the original image. The degree of textual influence of the local edit description text on the pixels is also referred to as the text importance score. The degree of textual influence can be understood as the degree of influence of the local edit description text on the editing result. The higher the degree of textual influence, the more the editing method will be determined by the local edit description text during subsequent editing. The lower the degree of textual influence, the smaller the influence of the local edit description text on the editing result, and the textual semantics of the local edit description text can be paid less attention or no attention during subsequent editing.

[0077] In an exemplary embodiment, the degree of text influence of the local editing description text on each pixel in the original image can be determined. For example, the degree of text influence of the local editing description text on pixels in a partial area of ​​the original image can also be determined. Exemplarily, the degree of text influence of only the pixels in the local editing area can be determined, or the partial image area containing the local editing area can be determined, and the degree of text influence of the pixels in the partial image area can be determined.

[0078] In some exemplary embodiments, the degree of text influence can be predicted using a pre-trained model, or pixel-level semantic segmentation can be performed on the original image to determine the category to which the pixels in the original image belong, and then the degree of text influence can be determined based on the semantic association between the category and the local edit description text, where the higher the semantic association, the higher the degree of text influence.

[0079] In step S330, the visual impact of the original image on subsequent editing of multiple pixels in the original image is adjusted according to the degree of text impact to obtain a target visual impact, and the overall image features of the original image are determined based on the original image and the target visual impact; the overall image features are used to generate a target image after local editing of the original image, and the target visual impact is negatively correlated with the text impact.

[0080] In practical applications, when performing local image editing, the original image often contains a wealth of detailed information. This information can be injected into the editing process as auxiliary information to improve the fidelity of the corresponding area. To address this, some related technologies directly fuse the original image when editing the original image based on text instructions. For example, this involves using a cross-attention mechanism to fuse the features of the original image and the text instructions, using the original image as a prompt to guide the model in generating an image of a similar style. However, the inventors have found in practice that this direct injection of the original image leads to an over-reliance on visual conditions when editing the image, while ignoring the text instructions.

[0081] In this regard, in this step, the visual impact of the original image on subsequent editing of multiple pixels in the original image can be adjusted according to the text impact, and the adjusted visual impact is used as the target visual impact.

[0082] The degree of visual impact can be understood as the degree to which the original image's content influences the editing results. The higher the degree of visual impact, the more reliant the original image's information will be on subsequent editing. For example, more of the original image's content can be used as auxiliary information in the image editing process. The lower the degree of visual impact, the less influence the original image's content has on subsequent editing results, allowing for less or no attention to the original image's content during subsequent editing.

[0083] In this embodiment, the target visual impact can be negatively correlated with the text impact. When the text impact is high, it can indicate that the local edit description text editing has a stronger impact on a certain pixel. Therefore, the visual dependence (or visual attention) of the pixel on the original image content can be suppressed, reducing the visual influence of the original image on the pixel, and improving the accuracy of text editing based on the local edit description. When the text impact is low, it can be determined that the local edit description text has a weaker impact on a certain pixel. Therefore, the visual dependence (or visual attention) on the original image content can be enhanced to maintain causal dependence on the original image and avoid unnatural transitions between the boundary of the local edit area and the original image content.

[0084] After obtaining the target visual impact levels of multiple pixels, different levels of attention are applied to the image features of the original image according to the target visual impact levels of the multiple pixels, thereby obtaining the overall image features of the original image, wherein the overall image features can reflect the attention paid to the visual content of each pixel of the original image as a whole. In one embodiment, the pixel features of multiple pixels of the original image can be obtained, and then the overall image features of the original image can be determined based on the product of the target visual impact levels of each of the multiple pixels and the pixel features. Subsequently, the target image after the original image is locally edited can be generated in combination with the overall image features. It can be understood that in this embodiment, on the basis of additionally injecting the original image detail information, the injection intensity of the original image detail information can be adaptively adjusted based on the impact of the local edit description text on multiple pixels in the original image, thereby reasonably balancing the effects of the text condition (i.e., the local edit description text) and the visual condition (i.e., the original image).

[0085] In step S340 , a target image is generated according to the overall image features.

[0086] After obtaining the overall image features, a target image that is a partial edit of the original image can be generated based on the overall image features. In some exemplary embodiments, the overall image features can be input into a trained image generation module that can be used to generate the target image. The image generation module generates and outputs the target image based on the input overall image features.

[0087] In other exemplary embodiments, the overall text features corresponding to the locally edited description text may also be obtained, thereby generating a target image based on the fusion results of the overall text features and the overall image features. For example, the fusion results of the overall text features and the overall image features may be input into an image generation module to generate a target image output by the image generation module. This target image may be obtained by performing a local image edit on the original image based on the locally edited description text. For ease of distinction, the overall text features used for fusion with the overall image features may be referred to as the first overall text features.

[0088] In some exemplary embodiments, the overall feature of the first text can be determined based on a self-attention mechanism or a cross-attention mechanism. For example, the overall feature of the first text can be obtained as follows: :

[0089]

[0090] in, is the query matrix corresponding to the local edit description text determined by the self-attention mechanism or the cross-attention mechanism; The key matrix corresponding to the local edit description text, is the value matrix corresponding to the local edit description text, and d is the feature dimension corresponding to the features in the matrix.

[0091] In the above-mentioned image editing method, an original image to be partially edited and text describing the local edit can be obtained, and the degree of textual influence of the local edit description text on subsequent edits of multiple pixels in the original image can be determined. Then, based on the degree of textual influence, the degree of visual influence of the original image on the subsequent edits of multiple pixels in the original image is adjusted to obtain a target visual influence degree. Based on the original image and the target visual influence degree, the overall image features of the original image are determined, wherein the overall image features are used to generate a target image after the local edit of the original image, and the target visual influence degree is negatively correlated with the textual influence degree. Furthermore, the target image can be generated based on the overall image features. In this embodiment, by adjusting the degree of visual influence of the original image on the subsequent edits of multiple pixels in the original image based on the textual influence degree, and making the target visual influence degree negatively correlated with the textual influence degree, on the one hand, it can reduce the visual dependence on the original image when the textual influence degree is high, thereby improving the accuracy of text editing. On the other hand, it can increase the attention and use of the visual information of the original image content when the textual influence degree is low, so that the local edit content naturally transitions to the original image content, effectively balancing the accuracy of text editing and the naturalness of boundary transition.

[0092] Compared with related technologies, the image editing method of this embodiment not only improves the text editing capability with higher levels of detail in various image local redrawing tasks (such as the local redrawing model of the human face), but also improves the authenticity and fidelity of the editing results. At the same time, the target image obtained by the final editing has significant improvements in style preservation and text alignment, which fully proves that the disclosed method is suitable for various delicate attribute editing.

[0093] In an exemplary embodiment, Figure 4 As shown, in step S320, determining the degree of influence of the local edit description text on the subsequent edited text of multiple pixels in the original image may include the following steps:

[0094] In step S321 , semantic unit text features corresponding to each of the plurality of semantic units in the partial edit description text are obtained, and the overall text features corresponding to the partial edit description text are determined based on the semantic unit text features.

[0095] In a specific implementation, the local edit description text can be segmented according to natural language rules to obtain multiple semantic units, and then the trained text feature extraction module can be used to obtain the semantic unit text features corresponding to each of the multiple semantic units.

[0096] In some exemplary embodiments, the text feature extraction module and the image feature extraction module can be trained by contrastive learning with paired text and image, i.e., contrastive language-image pre-training (CLIP), to obtain trained text feature extraction module and image feature extraction module, wherein the text feature extraction module can be used to obtain semantic unit text features, semantic unit text features As shown below:

[0097]

[0098] in, is a semantic unit, is the number of semantic units, It is the feature dimension of the semantic unit text feature.

[0099] After obtaining the semantic unit text features corresponding to each of the multiple semantic units, the multiple semantic unit text features can be fused to obtain the overall text feature corresponding to the local edit description text. For the sake of distinction, the overall text feature can be referred to as the second overall text feature. In some exemplary embodiments, the multiple semantic unit text features can be pooled to obtain the pooled feature. As the overall feature of the second text.

[0100] In step S322, the image editing features corresponding to the original image are obtained, and the trained text influence degree prediction module determines the text influence degree of each of multiple pixels in the original image based on the feature splicing results of the image editing features and the overall text features; the image editing features are features determined from the edited image during the process of pre-editing the original image according to the local editing description text.

[0101] In this step, the original image can be pre-edited according to the local editing description text, and features of the edited image can be extracted during the pre-editing process, and the extracted features are used as image editing features, wherein the image editing features can reflect the changes in the image before and after image editing. Therefore, by analyzing the image editing features, it can be determined how the local editing description text affects the editing of multiple pixels in the original image.

[0102] On the other hand, a pre-trained text influence prediction module can be obtained, which can also be called a score predictor of the text importance score. In some examples, the text influence prediction module can be supervised and trained using training samples with score labels.

[0103] Furthermore, the image editing features and the overall text features (i.e., the second overall text features) can be spliced ​​together, and the resulting feature splicing result can be input into a text influence prediction module. The module then determines the text influence of each of the multiple pixels in the original image based on the input feature splicing result. In one example, the text influence can be obtained as follows:

[0104]

[0105] Among them, Z is the image editing feature, , is the number of features of the image editing features, is the feature dimension of image editing features, the degree of text influence , feature splicing results , S(⋅) is the text influence prediction module.

[0106] In this embodiment, the text influence degree prediction module determines the text influence degree of multiple pixels in the original image based on the feature splicing results of the image editing features and the overall text features. The image editing features can be determined based on the text content of the local editing description text itself and the actual pre-editing of the image, and the impact of the local editing description text on the editing effects of multiple pixels in the original image can be accurately identified, thereby improving the detection accuracy of the text influence degree.

[0107] In an exemplary embodiment, in step S322, obtaining the image editing features corresponding to the original image may include the following steps:

[0108] The randomly generated noise image and the mask image obtained by performing mask processing on the locally edited region labeled in the original image are input into the trained image generation module. The image generation module includes a plurality of cascaded feature processing units. The image generation module is configured to perform denoising processing on the noise image according to the mask image and the local editing description text, so as to obtain an image in which the locally edited region is edited according to the local editing description text. The denoising prompt information obtained by each feature processing unit according to the mask image and the local editing description text is used to perform denoising processing on the input content, so as to obtain denoised image features. The denoised image features are used as the input content of the next feature processing unit. The input content of the first feature processing unit includes the noise image. The plurality of feature processing units perform different denoising processing. According to the plurality of denoised image features, the image editing features corresponding to the original image are determined.

[0109] In a specific implementation, the image pre-editing can be performed based on an inversion type method. In the inversion type method, a new image can be generated through an image denoising process. In this embodiment, a randomly generated noise image (for example, a randomly generated noise is superimposed on the original image, or a noise image is obtained based on a randomly generated image noise without using the original image) and a mask image obtained by performing mask processing on a locally edited region labeled in the original image can be obtained. In some optional embodiments, the locally edited region in the original image can be labeled by an image editing account. After obtaining the original image carrying the label information, the server can obtain the corresponding mask image. For example, in the mask image of Figure 5 the eye region to be edited can be masked.

[0110] Then, the noise image, the mask image, and the local editing description text can be input into a trained image generation module. The image generation module includes a plurality of cascaded feature processing units. The image generation module can perform denoising processing on the noise image through the plurality of cascaded feature processing units according to the mask image and the local editing description text, so as to obtain an image in which the locally edited region is edited according to the local editing description text.

[0111] In some optional embodiments, the image generation module can be constructed based on a diffusion model. In some examples, as shown in Figure 5 the diffusion model can be included in the image generation module. The mask image and the local editing description text can be directly input into the diffusion model, or can be processed by a reference information injection module and then input into the diffusion model.

[0112] After receiving the noisy image, the first feature processing unit in the image generation module can determine denoising prompt information based on the associated input mask image and the local edit description text. It then denoises the noisy image based on the denoising prompt information and extracts denoised image features. The denoised image features can then be input into the next feature processing unit as input, and denoised again to obtain new denoised image features. Each feature processing unit performs different denoising operations. By sequentially denoising the input content through multiple cascaded feature processing units, an image can be gradually generated that matches the image details in the mask image (i.e., image details outside the mask area) and the local edit description text.

[0113] Then, the influence of the local editing description text on the original image and the pixel editing can be determined based on the obtained multiple denoised image features to obtain the corresponding image editing features. In some embodiments, all denoised image features can be used as image editing features, for example, Figure 5 In the framework shown, multiple feature processing units in the diffusion model cascade can obtain image editing features Z (also called hidden positives). After these image editing features Z are converted from multidimensional data to one-dimensional data (Flatten to Tokens), they can be used as input to the score predictor. Of course, all denoised image features can also be filtered or fused, and the filtered or fused results can be used as image editing features.

[0114] In this embodiment, the input content is denoised by each feature processing unit based on the denoising prompt information determined according to the mask image and the local editing description text to obtain denoised image features, and the image editing features corresponding to the original image are determined based on multiple denoised image features. This can comprehensively consider different aspects of the image information, so that the generated image editing features are more consistent with the overall content and structural logic of the image, and help to more accurately determine the degree of influence of the local editing description text on the image editing in the future.

[0115] In an exemplary embodiment, Figure 6 As shown, in step S330, the visual impact of the original image on the subsequent editing of multiple pixels in the original image is adjusted according to the text impact to obtain a target visual impact, and the overall image features of the original image are determined based on the original image and the target visual impact. The following steps may be included:

[0116] In step S331 , a visual impact degree set of the original image is obtained, where the visual impact degree set includes the visual impact degree of the original image on subsequent editing of each pixel of the original image.

[0117] In a specific implementation, the visual influence degree of each pixel of the original image after subsequent editing of the original image can be obtained, and the plurality of visual influence degrees constitute a visual influence degree set.

[0118] In an optional embodiment, the visual influence degree can be determined through a self-attention mechanism or a cross-attention mechanism. For example, an image feature extraction module trained through contrastive learning can be used to obtain image block features corresponding to each image block of the original image. The plurality of image blocks can be segmented from the image according to a preset interval, and then a query matrix, a key matrix and a value matrix of the original image can be determined according to the image block features. The visual influence degree of each pixel of the original image can be determined according to the following manner:

[0119]

[0120] wherein, is a query matrix corresponding to the original image determined according to the self-attention mechanism or the cross-attention mechanism; is a key matrix corresponding to the original image, and d is a feature dimension corresponding to a feature in the matrix.

[0121] In step S332, the plurality of visual influence degrees in the visual influence degree set are adjusted according to the text influence degrees of each pixel in the local editing region, and the target visual influence degree of each pixel of the original image is determined based on the adjusted visual influence degree set. The local editing region is a local region in the original image that is marked for editing.

[0122] In this step, the visual influence degree can be adjusted according to the text influence degree of each pixel in the local editing region. Specifically, the text influence degree of each pixel in the local editing region can be determined from the plurality of obtained text influence degrees, and then the plurality of visual influence degrees in the visual influence degree set can be adjusted according to the plurality of determined text influence degrees. In some examples, the visual influence degrees corresponding to the same pixel in the visual influence degree set can be adjusted according to the plurality of determined text influence degrees, and in other embodiments, the visual influence degrees of other associated pixels can also be adjusted on this basis, for example, the pixels outside the local editing region and within a preset distance are associated with other pixels, and then they can be adjusted together accordingly.

[0123] Further, each visual influence degree in the adjusted visual influence degree set can be used as the target visual influence degree of each pixel.

[0124] In step S333, the overall image features of the original image are determined based on the target visual impact degree of each pixel of the original image and the pre-acquired pixel features of each pixel of the original image.

[0125] After obtaining the target visual impact of each pixel in the original image, the pixel features of each pixel in the original image can also be obtained. In one example, the pixel features of each pixel in the original image can be obtained based on the value matrix corresponding to the original image. Then, the overall image features of the original image can be determined based on the target visual impact of each pixel and the pixel features. In one example, the overall image features can be obtained as shown below: :

[0126]

[0127] in, is the value matrix of the original image, which can be obtained by linearly transforming the features of multiple image blocks; A collection of adjusted visual impact levels.

[0128] In one example, the overall image features can be and the overall characteristics of the first text The fusion is input into the reference information injection module for processing and then input into the diffusion model. For example, the splicing result of the overall image feature and the overall feature of the first text can be obtained as shown below: :

[0129]

[0130] Then, if Figure 5 As shown, the concatenation result can be input into each layer in the reference information injection module.

[0131] In this embodiment, by adjusting multiple visual impact levels in the visual impact level set according to the text impact level of each pixel in the local editing area, the target visual impact level of each pixel in the original image is obtained. The text impact level of each pixel in the local editing area can be used in a targeted manner to adjust the visual impact level, which helps to improve the accuracy of text editing in the local editing area, while reducing the amount of data processing and improving image editing efficiency.

[0132] In an exemplary embodiment, in step S620, adjusting multiple visual impact levels in the visual impact level set according to the text impact level of each pixel in the local editing area may include the following steps:

[0133] In the visual impact degree set, each original visual impact degree corresponding to the pixels in the local editing area is determined; for each original visual impact degree, the downward adjustment range of the original visual impact degree is determined according to the text impact degree of the corresponding pixel, and the adjusted visual impact degree is determined according to the downward adjustment range and the original visual impact degree; the downward adjustment range is positively correlated with the text impact degree; using the adjusted visual impact degree, the original visual impact degrees in the visual impact degree set are updated to obtain the adjusted visual impact degree set.

[0134] In practical applications, for each pixel in the local editing area, the visual impact degree corresponding to the pixel can be determined in the visual impact degree set as the original visual impact degree. Then, for each original visual impact degree, the downward adjustment amplitude of the visual impact degree can be determined based on the text impact degree of the corresponding pixel, wherein the downward adjustment amplitude of the visual impact degree is positively correlated with the text impact degree, that is, the higher the text impact degree, the greater the decrease in the visual impact degree of the pixel in the local editing area. In one example, the opposite of the text impact degree can be used as the downward adjustment amplitude of the visual image degree. For example, if the text impact degree is x (0 < x < 1), -x can be determined as the downward adjustment amplitude.

[0135] Then, the adjusted visual impact level can be used to update the corresponding original visual impact levels in the visual impact level set to obtain the adjusted visual impact level set. In one example, the adjusted visual impact level set can be determined as follows: :

[0136]

[0137] Wherein, M is a mask, which is used to indicate that the information of the local editing area is retained while the information outside the local editing area is discarded.

[0138] In this embodiment, by determining the downward adjustment range of each original visual impact degree of pixels in the local editing area based on the text impact degree, determining the adjusted visual impact degree based on the downward adjustment range of the visual impact degree and the original visual impact degree, and updating the original visual impact degree in the visual impact degree set, the visual impact degree of each pixel in the local editing area can be accurately adjusted, which helps to ensure the accuracy of controlling the image content in the local editing area through the local editing description text.

[0139] In an exemplary embodiment, before step S310, the following steps may be further included:

[0140] A pre-acquired sample image, a sample mask image corresponding to the sample image, and attribute description text for a local area of ​​the sample image are input into an image editing model to be trained; the image editing model determines the text influence degree of each of multiple pixels in the sample image, and determines the sample visual influence degree of each of multiple pixels in the sample image based on the text influence degree of each of the multiple pixels in the sample image; the overall features of the sample image corresponding to the sample image are determined based on the sample image and the sample visual influence degree, and an edited image for locally editing the sample image is generated based on the overall features of the sample image; the model parameters of the image editing model are adjusted based on the difference between the edited image and the sample image to obtain a trained image editing model; the image editing model is used to generate a target image based on the input original image and the local edit description text.

[0141] The sample mask image is obtained by performing masking processing on a local area of ​​the sample image.

[0142] In a specific implementation, a sample image and a mask image corresponding to the sample image can be obtained. The sample image is any image used to train the image editing model. In one embodiment, the image type of the sample image can be the same as the image type actually edited in the image editing account. For example, if the image editing account edits a face image, the sample image can include a face image. Furthermore, attribute description text for a local area of ​​the sample image can be obtained. The attribute description text can represent the attribute information of the local area of ​​the image in the sample image in text form.

[0143] Furthermore, the matched sample image, sample mask image, and attribute description text can be input into the image editing model to be trained. The image editing model determines the text influence degree of each of the multiple pixels in the sample image. Based on the text influence degree of each of the multiple pixels in the sample image, the visual influence degree of each of the multiple pixels in the sample image, i.e., the sample visual influence degree, is determined. Then, based on the sample image and the sample visual influence degree, the overall image feature corresponding to the sample image, i.e., the overall sample image feature, is determined. Based on the overall sample image feature, an edited image is generated that partially edits the sample image based on the attribute description text. The image editing model can adjust the original visual influence degree of the multiple pixels in the sample image based on the text influence degree, and determine the overall sample image feature corresponding to the sample image based on the adjusted visual influence degree and the sample image. In some embodiments, the process of determining the text influence degree, adjusting the visual influence degree of the sample image based on the text influence degree, and determining the overall sample image feature of the sample image by the image editing model can be the same as the process of determining the text influence degree of the original image, adjusting the visual influence degree of the original image based on the text influence degree, and determining the overall image feature of the original image in one or more of the aforementioned embodiments, and will not be repeated here.

[0144] Then, the model parameters of the image editing model can be adjusted according to the difference between the edited image and the sample image to obtain a trained image editing model. The trained image editing model can generate a target image based on the input original image and the local editing description text. For example, it can be deployed on a server to perform the processing of the aforementioned steps S310 to S340.

[0145] In this embodiment, by using sample images and their corresponding sample mask images and attribute description text to train the image editing model, the image editing model's ability to understand the local semantics of the image can be enhanced. At the same time, based on a large number of training samples, the image editing model can fully learn how to extract the degree of text influence and adjust the degree of visual influence according to the degree of text influence, thereby improving the image editing model's ability to balance text conditions and visual conditions, thereby achieving high-fidelity and high-editing degree local editing of images, while following text instructions, accurately retaining the skin details of the original image, and solving the unnatural problem of boundary areas.

[0146] In an exemplary embodiment, before inputting the pre-acquired sample image, the sample mask image corresponding to the sample image, and the attribute description text for the local area of ​​the sample image into the image editing model to be trained, the following steps may be further included:

[0147] A sample image is obtained and segmented using a trained image segmentation model to obtain at least one local image region in the sample image; the image content in the same local image region has matching attributes; masking is performed on each of the local image regions to obtain a sample mask image corresponding to the sample image; and the sample image and the segmentation result of the sample image are input into a trained language model to obtain attribute description text output by the language model for each local image region.

[0148] In practical applications, a large number of sample images can be obtained and then input into a trained image segmentation model. For each sample image input, the image segmentation model can perform image semantic analysis and segmentation on the sample image, identifying at least one local image region, where each pixel in the same local image region has matching semantic attributes. It is understandable that the same local image region can be continuous or discontinuous. For example, in a facial image, the discontinuous eye region can be treated as the same local image region, or the left and right eyebrows can be divided into two local image regions.

[0149] Then, on the one hand, each local image area in the sample image can be masked separately to obtain the corresponding sample mask image. On the other hand, the sample image and the segmentation result of the sample image can be input into the trained language model, and the language model outputs a fine-grained text description for each local image area to obtain the attribute description text of each local image area.

[0150] For example, taking the sample image generation of face images as an example, Figure 7 As shown, the face image can be input into the image segmentation model, which performs semantic analysis and mask processing and divides the face image into multiple local image regions, such as hair, nose, ears, eyes, lips, etc. For ease of understanding, Figure 7 In the present invention, different colors are masked on multiple local regions of the same sample image at the same time. However, in some embodiments, each local region of the sample image can be masked separately. For example, the lips in the sample image can be masked separately to obtain a sample mask image, and the eyebrows in the sample image can be masked separately to obtain another sample mask image. On the other hand, instructions can be input to the Multimodal Large Language Model (MLLM) to instruct the MLLM to generate fine-grained attribute description text for each local region of the sample image (i.e., the segmentation result of the sample image).

[0151] In this embodiment, the image segmentation model is used to perform image segmentation to obtain the local image regions of the sample image, and the language model outputs the attribute description text of each local image region. The processing capability of the language model can be used to generate fine-grained text-mask pairs, quickly construct a large number of triplet training samples, and effectively improve the image editing model's ability to understand local semantics, helping the model to understand fine-grained image attribute descriptions.

[0152] In order to enable those skilled in the art to better understand the above steps, the embodiment of the present disclosure is illustrated below by using an example, but it should be understood that the embodiment of the present disclosure is not limited thereto.

[0153] like Figure 8 As shown, in this embodiment, the following steps may be included:

[0154] S801, input the pre-acquired sample image, sample mask image and attribute description text into the image editing model to be trained, and the image editing model locally edits the sample image according to the sample image, sample mask image and attribute description text to obtain an edited image.

[0155] Among them, the framework of the image editing model can be as follows Figure 5 shown.

[0156] S802: Adjust the model parameters of the image editing model according to the difference between the edited image and the sample image to obtain a trained image editing model.

[0157] S803: Obtain the original image and local editing description text provided by the image editing account.

[0158] S804, using the image editing model to determine the degree of text influence of the local editing description text on the subsequent editing of multiple pixels in the original image, adjusting the visual influence of the original image on the subsequent editing of multiple pixels in the original image according to the text influence degree, obtaining the target visual influence degree, and determining the overall image features of the original image based on the original image and the target visual influence degree, and generating the target image after the original image is locally edited according to the overall image features and the first overall text features corresponding to the pre-acquired local editing description text.

[0159] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0160] It can be understood that the same / similar parts between the various embodiments of the above method in this specification can be referred to each other, and each embodiment focuses on the differences from other embodiments. For related parts, please refer to the description of other method embodiments.

[0161] Based on the same inventive concept, an embodiment of the present disclosure further provides an image editing device for implementing the above-mentioned image editing method.

[0162] Figure 9 FIG. 1 is a block diagram of an image editing device according to an exemplary embodiment. Figure 9 , the device comprises:

[0163] The editing information acquiring unit 901 is configured to acquire an original image to be partially edited and a local editing description text;

[0164] A text impact degree determining unit 902 is configured to determine a text impact degree of the local edit description text on subsequent edits of multiple pixels in the original image;

[0165] A visual impact adjustment unit 903 is configured to adjust the visual impact of subsequent editing of the original image on multiple pixels in the original image based on the text impact to obtain a target visual impact, and determine an overall image feature of the original image based on the original image and the target visual impact; the overall image feature is used to generate a target image after the original image is partially edited, and the target visual impact is negatively correlated with the text impact.

[0166] The image generating unit 904 is configured to generate the target image according to the overall features of the image.

[0167] In an exemplary embodiment, the text influence degree determination unit 902 is configured to perform:

[0168] Acquire semantic unit text features corresponding to each of the plurality of semantic units in the partial edit description text, and determine overall text features corresponding to the partial edit description text based on the semantic unit text features;

[0169] The image editing features corresponding to the original image are obtained, and a trained text influence degree prediction module is used to determine the text influence degree of each of multiple pixels in the original image based on the feature splicing results of the image editing features and the overall text features; the image editing features are features determined from the edited image during the process of pre-editing the original image according to the local editing description text.

[0170] In an exemplary embodiment, the text influence degree determination unit 902 is configured to perform:

[0171] Obtaining a randomly generated noise image and a mask image obtained by masking the local edited region marked in the original image, and inputting the noise image, the mask image, and the local edit description text into a trained image generation module; the image generation module includes a plurality of cascaded feature processing units, and the image generation module is configured to perform denoising processing on the noise image based on the mask image and the local edit description text to obtain an image in which the local edited region is edited according to the local edit description text;

[0172] Each of the feature processing units performs denoising processing on the input content based on the denoising prompt information determined by the mask image and the local edit description text to obtain denoising image features, and uses the denoising image features as input content for the next feature processing unit; the input content of the first feature processing unit includes the noise image, and the denoising processing of the multiple feature processing units is different;

[0173] An image editing feature corresponding to the original image is determined based on the plurality of denoised image features.

[0174] In an exemplary embodiment, the visual impact adjustment unit 903 is configured to execute:

[0175] Acquiring a visual impact degree set of the original image, wherein the visual impact degree set includes a visual impact degree of the original image on subsequent editing of each pixel of the original image;

[0176] adjusting multiple visual impact levels in the visual impact level set according to the text impact level of each pixel in the local editing area, and determining a target visual impact level for each pixel in the original image based on the adjusted visual impact level set; the local editing area being the local area marked to be edited in the original image;

[0177] The overall image features of the original image are determined according to the target visual impact degree of each pixel of the original image and the pre-acquired pixel features of each pixel of the original image.

[0178] In an exemplary embodiment, the visual impact adjustment unit 903 is configured to execute:

[0179] In the visual impact degree set, determining respective original visual impact degrees corresponding to pixels in the local editing area;

[0180] For each of the original visual impact levels, determining a downward adjustment range of the original visual impact level according to the text impact level of the corresponding pixel, and determining an adjusted visual impact level according to the downward adjustment range and the original visual impact level; the downward adjustment range is positively correlated with the text impact level;

[0181] The original visual impact levels in the visual impact level set are updated using the adjusted visual impact levels to obtain an adjusted visual impact level set.

[0182] In an exemplary embodiment, the apparatus further includes a model training unit, wherein the model training unit is configured to execute:

[0183] Inputting a pre-acquired sample image, a sample mask image corresponding to the sample image, and attribute description text for a local image region in the sample image into the image editing model to be trained; the sample mask image is obtained by performing masking processing on the local image region in the sample image;

[0184] Determining, by the image editing model, respective text influence degrees of a plurality of pixels in the sample image, determining respective sample visual influence degrees of the plurality of pixels in the sample image based on the respective text influence degrees of the plurality of pixels in the sample image, determining, based on the sample image and the sample visual influence degrees, a sample image overall feature corresponding to the sample image, and generating, based on the sample image overall feature, an edited image that performs a partial edit on the sample image;

[0185] According to the difference between the edited image and the sample image, the model parameters of the image editing model are adjusted to obtain a trained image editing model; the image editing model is used to generate the target image based on the input original image and the local editing description text.

[0186] In an exemplary embodiment, before inputting the pre-acquired sample image, the sample mask image corresponding to the sample image, and the attribute description text for the local area of ​​the sample image into the image editing model to be trained, the method further includes:

[0187] Acquire a sample image, and segment the sample image using a trained image segmentation model to obtain at least one local image region in the sample image; image contents in the same local image region have matching attributes;

[0188] Masking is performed on each local area of ​​the image to obtain a sample mask image corresponding to the sample image; and the sample image and the segmentation result of the sample image are input into a trained language model to obtain attribute description text output by the language model for each local area of ​​the image.

[0189] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0190] Each module in the above-mentioned image editing device may be implemented in whole or in part through software, hardware, or a combination thereof. Each module may be embedded in or independent of a processor in a computer device in the form of hardware, or may be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0191] Figure 101 is a block diagram of a computer device 1000 (the computer device may also be referred to as an electronic device) for implementing an image editing method according to an exemplary embodiment. For example, the computer device 1000 may be a server. Figure 10 Computer device 1000 includes a processing component 1020, which further includes one or more processors, and memory resources represented by memory 1022 for storing instructions, such as applications, that can be executed by processing component 1020. The application stored in memory 1022 may include one or more modules, each corresponding to a set of instructions. In addition, processing component 1020 is configured to execute the instructions to perform the above-described method.

[0192] The computer device 1000 may also include a power supply component 1024 configured to perform power management of the computer device 1000, a wired or wireless network interface 1026 configured to connect the computer device 1000 to a network, and an input / output (I / O) interface 1028. The computer device 1000 may operate based on an operating system stored in the memory 1022, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, or the like.

[0193] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as memory 1022 including instructions. The instructions are executable by a processor of the computer device 1000 to perform the above method. The storage medium may be a computer-readable storage medium, such as a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, and the like.

[0194] In an exemplary embodiment, a computer program product is further provided. The computer program product includes instructions, and the instructions can be executed by a processor of the computer device 1000 to implement the above method.

[0195] It should be noted that the above-mentioned devices, computer equipment, computer-readable storage media, computer program products, etc. can also include other implementation methods according to the description of the method embodiments. The specific implementation methods can refer to the description of the relevant method embodiments and will not be described one by one here.

[0196] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the claims.

[0197] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. An image editing method, characterized in that: include: Obtaining the original image to be partially edited and the local editing description text; Determining the extent to which the local edit description text affects subsequent edits of multiple pixels in the original image; adjusting the visual impact of subsequent editing of the original image on a plurality of pixels in the original image based on the text impact to obtain a target visual impact, and determining overall image features of the original image based on the original image and the target visual impact; the overall image features are used to generate a target image after the original image is partially edited, the target visual impact being negatively correlated with the text impact; The target image is generated according to the overall features of the image.

2. The method according to claim 1, characterized in that Determining the degree of influence of the local edit description text on the text of subsequent edits of the plurality of pixels in the original image includes: Acquire semantic unit text features corresponding to each of the plurality of semantic units in the partial edit description text, and determine overall text features corresponding to the partial edit description text based on the semantic unit text features; The image editing features corresponding to the original image are obtained, and a trained text influence degree prediction module is used to determine the text influence degree of each of multiple pixels in the original image based on the feature splicing results of the image editing features and the overall text features; the image editing features are features determined from the edited image during the process of pre-editing the original image according to the local editing description text.

3. The method according to claim 2, characterized in that The obtaining of the image editing feature corresponding to the original image includes: Obtaining a randomly generated noise image and a mask image obtained by masking the local edited region marked in the original image, and inputting the noise image, the mask image, and the local edit description text into a trained image generation module; the image generation module includes a plurality of cascaded feature processing units, and the image generation module is configured to perform denoising processing on the noise image based on the mask image and the local edit description text to obtain an image in which the local edited region is edited according to the local edit description text; Each of the feature processing units performs denoising processing on the input content based on the denoising prompt information determined by the mask image and the local edit description text to obtain denoising image features, and uses the denoising image features as input content for the next feature processing unit; the input content of the first feature processing unit includes the noise image, and the denoising processing of the multiple feature processing units is different; An image editing feature corresponding to the original image is determined based on the plurality of denoised image features.

4. The method according to claim 1, wherein The step of adjusting the visual impact of the original image on subsequent editing of multiple pixels in the original image based on the text impact to obtain a target visual impact, and determining an overall image feature of the original image based on the original image and the target visual impact includes: Acquiring a visual impact degree set of the original image, wherein the visual impact degree set includes a visual impact degree of the original image on subsequent editing of each pixel of the original image; adjusting multiple visual impact levels in the visual impact level set according to the text impact level of each pixel in the local editing area, and determining a target visual impact level for each pixel in the original image based on the adjusted visual impact level set; the local editing area being the local area marked to be edited in the original image; The overall image features of the original image are determined according to the target visual impact degree of each pixel of the original image and the pre-acquired pixel features of each pixel of the original image.

5. The method according to claim 4, characterized in that The adjusting of the multiple visual impact levels in the visual impact level set according to the text impact level of each pixel in the local editing area includes: In the visual impact degree set, determining respective original visual impact degrees corresponding to pixels in the local editing area; For each of the original visual impact levels, determining a downward adjustment range of the original visual impact level according to the text impact level of the corresponding pixel, and determining an adjusted visual impact level according to the downward adjustment range and the original visual impact level; the downward adjustment range is positively correlated with the text impact level; The original visual impact levels in the visual impact level set are updated using the adjusted visual impact levels to obtain an adjusted visual impact level set.

6. The method according to any one of claims 1 to 5, characterized in that Before obtaining the original image to be partially edited and the local editing description text, the method further includes: Inputting a pre-acquired sample image, a sample mask image corresponding to the sample image, and attribute description text for a local image region in the sample image into the image editing model to be trained; the sample mask image is obtained by performing masking processing on the local image region in the sample image; Determining, by the image editing model, respective text influence degrees of a plurality of pixels in the sample image, determining respective sample visual influence degrees of the plurality of pixels in the sample image based on the respective text influence degrees of the plurality of pixels in the sample image, determining, based on the sample image and the sample visual influence degrees, a sample image overall feature corresponding to the sample image, and generating, based on the sample image overall feature, an edited image that performs a partial edit on the sample image; According to the difference between the edited image and the sample image, the model parameters of the image editing model are adjusted to obtain a trained image editing model; the image editing model is used to generate the target image based on the input original image and the local editing description text.

7. The method according to claim 6, characterized in that Before inputting the pre-acquired sample image, the sample mask image corresponding to the sample image, and the attribute description text for the local area of ​​the sample image into the image editing model to be trained, the method further includes: Acquire a sample image, and segment the sample image using a trained image segmentation model to obtain at least one local image region in the sample image; image contents in the same local image region have matching attributes; Masking is performed on each local area of ​​the image to obtain a sample mask image corresponding to the sample image; and the sample image and the segmentation result of the sample image are input into a trained language model to obtain attribute description text output by the language model for each local area of ​​the image.

8. An image editing device, characterized in that: include: An editing information acquiring unit configured to acquire an original image to be partially edited and a local editing description text; a text impact degree determining unit configured to determine a text impact degree of the local edit description text on subsequent edits of a plurality of pixels in the original image; a visual impact adjustment unit configured to adjust the visual impact of subsequent editing of the original image on a plurality of pixels in the original image based on the text impact to obtain a target visual impact, and determine an overall image feature of the original image based on the original image and the target visual impact; the overall image feature is used to generate a target image after the original image is partially edited, and the target visual impact is negatively correlated with the text impact; The image generating unit is configured to generate the target image according to the overall features of the image.

9. A computer device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the image editing method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that When the instructions in the computer-readable storage medium are executed by a processor of a computer device, the computer device is enabled to execute the image editing method according to any one of claims 1 to 7.

11. A computer program product comprising instructions, characterized in that: When the instructions are executed by a processor of a computer device, the computer device is enabled to execute the image editing method according to any one of claims 1 to 7.