Image processing method and device, electronic equipment and storage medium
Through the self-attention network, the association relationships in long-sequence images are extracted and noise prediction is performed, which solves the problem of insufficient consistency of long-sequence visual content, and achieves high-quality and efficient image generation.
Patent Information
- Application Number
- CN202411833254.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-12
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art has shortcomings in maintaining the consistency of long-sequence visual content, resulting in inconsistencies or lack of coherence in the generated content, affecting the content quality and user experience.
By acquiring the to-process noise images, description text and initial image sequences, the association relationships in the image sequence are extracted using the self-attention network, noise prediction and image denoising are performed, and the target image is generated.
Ensure that the generated images maintain high consistency in long sequences, improve the image generation quality and efficiency, and solve the problem of insufficient consistency of visual content in long sequences.
Smart Images

Figure CN119941489A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing, and in particular to an image processing method, device, electronic device and storage medium. Background Art
[0002] With the development of artificial intelligence generated content (AIGC) technology, it has high application value in animation, games and other fields. For example, AI can be used to generate long-sequence visual content. However, the existing technology is insufficient in maintaining the consistency of long-sequence visual content. When processing long-sequence visual content, it is difficult to maintain the consistency of characters, scenes and other content, resulting in inconsistencies or lack of coherence in the generated content, which in turn affects the content quality and user experience. Summary of the invention
[0003] The present application provides an image processing method, device, electronic device and storage medium to at least solve the problem of low consistency between multiple generated images in the related art. The technical solution of the present application is as follows:
[0004] According to a first aspect of an embodiment of the present application, there is provided an image processing method, comprising:
[0005] Acquire a first noise image to be processed, a description text, and an initial image sequence corresponding to the description text;
[0006] Inputting the initial image sequence into a self-attention network for self-attention learning to extract the association relationship between the initial images in the initial image sequence to obtain first image feature information;
[0007] Perform noise prediction based on the first image feature information and the first noise image to be processed to obtain first updated noise information;
[0008] Perform noise prediction based on the first updated noise information to obtain first predicted noise information;
[0009] Image denoising is performed on the first noisy image to be processed based on the first predicted noise information to obtain a target image.
[0010] In an optional embodiment, the initial image sequence includes a first preset number of initial images arranged based on a first preset order, and the initial image sequence is input into a self-attention network for self-attention learning to extract the association relationship between the initial images in the initial image sequence, and the first image feature information is obtained, including:
[0011] Performing noise addition processing on a first preset number of the initial images respectively to obtain a first preset number of initial noise images, each of the initial noise images corresponding to a noise information;
[0012] Inputting a first preset number of the initial noise images into a first image coding network for image coding to obtain image features corresponding to each of the initial noise images;
[0013] Obtaining a first key weight matrix corresponding to the image feature corresponding to each of the initial noise images, and a corresponding first value weight matrix;
[0014] The product of the image feature corresponding to each of the initial noise images and the corresponding first key weight matrix is determined as the first key feature corresponding to each of the initial noise images; the product of the image feature corresponding to each of the initial noise images and the corresponding first value weight matrix is determined as the first value feature corresponding to each of the initial noise images;
[0015] Determine the noise information corresponding to the target noise image as the first query feature; the target noise image is an initial noise image corresponding to the last initial image in the initial image sequence;
[0016] Determine the similarity between the first query feature and the first key feature corresponding to each of the initial noise images, and obtain the first attention data corresponding to each of the initial noise images; the corresponding first attention data is used to characterize the importance of the corresponding first value feature;
[0017] Based on the first attention data corresponding to each of the initial noise images, the corresponding first value features are weighted to obtain the first image feature information.
[0018] In an optional embodiment, before performing noise prediction based on the first updated noise information to obtain first predicted noise information, the method further includes:
[0019] Inputting the initial image, the description text and the first updated noise information into a multimodal cross attention network for multimodal cross attention learning to extract second image feature information corresponding to the initial image and first text feature information corresponding to the description text, and fusing the second image feature information and the first text feature information to obtain fused feature information;
[0020] Correspondingly, performing noise prediction based on the first updated noise information to obtain first predicted noise information includes:
[0021] Noise prediction is performed based on the fused feature information and the first updated noise information to obtain the first predicted noise information.
[0022] In an optional embodiment, the multimodal cross-attention network includes a text cross-attention network and an image cross-attention network, the initial image includes an initial character image and an initial scene image, and the initial image, the description text and the first updated noise information are input into the multimodal cross-attention network for multimodal cross-attention learning to extract the second image feature information corresponding to the initial image and the first text feature information corresponding to the description text include:
[0023] Inputting the initial character image and the initial scene image into a second image coding network for image coding respectively, to obtain character feature information corresponding to the initial character image and scene feature information corresponding to the initial scene image;
[0024] Performing splicing processing on the character feature information and the scene feature information to obtain third image feature information;
[0025] Inputting the third image feature information and the first updated noise information into the text cross attention network for cross attention learning to obtain the second image feature information;
[0026] Inputting the description text into a text encoding network for text encoding to obtain second text feature information;
[0027] The second text feature information and the first updated noise information are input into the image cross-attention network for cross-attention learning to obtain the first text feature information.
[0028] In an optional embodiment, inputting the third image feature information and the first updated noise information into the text cross attention network for cross attention learning to obtain the second image feature information includes:
[0029] Obtaining a second key weight matrix and a corresponding second value weight matrix corresponding to the third image feature information;
[0030] Determine the product of the third image feature information and the corresponding second key weight matrix as the second key feature;
[0031] Determine the product of the third image feature information and the corresponding second value weight matrix as the second value feature;
[0032] determining the first updated noise information as a second query feature;
[0033] Determine the similarity between the second query feature and the second key feature to obtain second attention data; the second attention data is used to characterize the importance of the second value feature;
[0034] The second value feature is weighted based on the second attention data to obtain the second image feature information.
[0035] In an optional embodiment, inputting the second text feature information and the first updated noise information into the second image attention network for cross attention learning to obtain the first text feature information includes:
[0036] Obtaining a third key weight matrix and a corresponding third value weight matrix corresponding to the second text feature information;
[0037] Determine the product of the second text feature information and the corresponding third key weight matrix as the third key feature;
[0038] Determine the product of the second text feature information and the corresponding third value weight matrix as the third value feature;
[0039] determining the first updated noise information as a second query feature;
[0040] Determine the similarity between the second query feature and the third key feature to obtain third attention data; the third attention data is used to characterize the importance of the third value feature;
[0041] The third value feature is weighted based on the third attention data to obtain the first text feature information.
[0042] In an optional embodiment, the first noise image to be processed corresponds to first initial noise information, and performing noise prediction based on the first image feature information and the first noise image to be processed to obtain first updated noise information includes:
[0043] Perform noise prediction based on the first image feature information and the first initial noise information to obtain second predicted noise information;
[0044] The difference information between the initial noise information and the second predicted noise information is determined as the first updated noise information.
[0045] In an optional embodiment, performing image denoising on the first to-be-processed noisy image based on the first predicted noise information to obtain a target image includes:
[0046] Based on the first predicted noise information, the first noise image to be processed is subjected to image denoising to obtain an updated noise image; the first noise image to be processed corresponds to the first initial noise information, the updated noise image corresponds to the second updated noise information, and the second updated noise information is the difference information between the first initial noise information and the first predicted noise information;
[0047] The updated noise image is used again as the first noise image to be processed, and the steps of inputting the initial image sequence into the self-attention network for self-attention learning are repeated until the first noise image to be processed is subjected to image denoising based on the first predicted noise information to obtain a target image, until a preset condition is met and the target image is obtained.
[0048] In an optional embodiment, the method further includes:
[0049] Acquire a description text sequence; the description text sequence includes a second preset number of the description texts arranged in a second preset order, and each of the description texts corresponds to one of the initial image sequences;
[0050] For each of the description texts, the steps of inputting the initial image sequence into the self-attention network for self-attention learning, and performing image denoising on the first to-be-processed noisy image based on the first predicted noise information to obtain a target image are performed to obtain a target image corresponding to each of the description texts;
[0051] Based on the second preset order, the target images corresponding to each of the description texts are sorted to obtain a target image sequence.
[0052] In an optional embodiment, the acquiring the initial image sequence corresponding to the description text includes:
[0053] Acquire a second noise image to be processed;
[0054] Inputting the description text into a text encoding network for text encoding to obtain second text feature information;
[0055] Perform noise prediction based on the second text feature information and the second noise image to be processed to obtain third predicted noise information;
[0056] The second to-be-processed noisy image is subjected to image denoising based on the third predicted noise information to obtain the initial image sequence.
[0057] According to a second aspect of an embodiment of the present application, there is provided an image processing device, the device comprising:
[0058] A first acquisition module is configured to acquire a first noise image to be processed, a description text, and an initial image sequence corresponding to the description text;
[0059] A self-attention learning module is configured to input the initial image sequence into a self-attention network to perform self-attention learning, so as to extract the association relationship between the initial images in the initial image sequence and obtain first image feature information;
[0060] A first noise prediction module is configured to perform noise prediction based on the first image feature information and the first noise image to be processed to obtain first updated noise information;
[0061] A second noise prediction module is configured to perform noise prediction based on the first updated noise information to obtain first predicted noise information;
[0062] The denoising module is configured to perform image denoising on the first to-be-processed noisy image based on the first predicted noise information to obtain a target image.
[0063] According to a third aspect of an embodiment of the present application, there is provided an electronic device for image processing, including:
[0064] processor;
[0065] a memory for storing instructions executable by the processor;
[0066] The processor is configured to execute the instructions to implement the image processing method as described in any of the above embodiments.
[0067] According to a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device executes the image processing method as described in any of the above embodiments.
[0068] According to a fifth aspect of an embodiment of the present application, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the image processing method described in any one of the above embodiments is implemented.
[0069] The technical solution provided by the embodiments of the present application brings at least the following beneficial effects:
[0070] The embodiment of the present application provides an image processing method, device, electronic device and storage medium. The method obtains a first noise image to be processed, a description text, and an initial image sequence corresponding to the description text; inputs the initial image sequence into a self-attention network for self-attention learning to extract the correlation between the initial images in the initial image sequence to obtain first image feature information; performs noise prediction based on the first image feature information and the first noise image to be processed to obtain first updated noise information; performs noise prediction based on the first updated noise information to obtain first predicted noise information; performs image denoising on the first noise image to be processed based on the first predicted noise information to obtain a target image. Thus, by combining text and initial image information for image generation and extracting correlation in long sequence images, it is possible to ensure that the generated image maintains a high degree of consistency in the long sequence, thereby improving the image generation quality and efficiency.
[0071] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] The drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present application, and together with the specification are used to explain the principles of the present application, and do not constitute an improper limitation on the present application.
[0073] Figure 1 The figure is a diagram showing an application environment of an image processing method according to an exemplary embodiment.
[0074] Figure 2 It is a flowchart of an image processing method according to an exemplary embodiment.
[0075] Figure 3 The figure is a schematic diagram of a process of determining first image feature information according to an exemplary embodiment.
[0076] Figure 4 It is a flowchart of a cross-attention learning according to an exemplary embodiment.
[0077] Figure 5 The figure is a schematic diagram of a process of determining feature information of a second image according to an exemplary embodiment.
[0078] Figure 6 The figure is a schematic diagram showing a flow chart of determining first text feature information according to an exemplary embodiment.
[0079] Figure 7 The figure is a flowchart of generating a target image according to an exemplary embodiment.
[0080] Figure 8is a block diagram of an image processing device according to an exemplary embodiment.
[0081] Fig. 9 The invention is a block diagram of an electronic device for image processing according to an exemplary embodiment.
[0082] Fig.10 is a block diagram of another electronic device for image processing according to an exemplary embodiment. DETAILED DESCRIPTION
[0083] In order to enable ordinary persons in the art to better understand the technical solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.
[0084] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the attached claims.
[0085] With the development of artificial intelligence generated content (AIGC) technology, it has become possible to use AI to generate long-sequence comic content. However, existing technologies are insufficient in maintaining the consistency of characters in long-sequence comics, affecting content quality and user experience. When processing long-sequence content, current AI generation technology has difficulty maintaining the consistency of character appearance, scene content, etc., resulting in feature mutations of the same character in different scenes, affecting the coherence and credibility of the content. In addition, traditional attention mechanisms have difficulty effectively capturing long-distance contextual information when processing long-sequence data, resulting in inconsistencies or lack of coherence in the generated content. In addition, insufficient multimodal fusion leads to consistency issues between text description and image generation, and it is easy for images and text to be inconsistent.
[0086] In response to the above problems, this application aims to propose a long sequence image generation method based on long-range dependency self-attention modeling, which can be used in comic creation, animation production, video content generation and other fields. Through the improved attention mechanism, the problem of consistency of long-sequence visual content is effectively solved, the long-range dependency modeling capability is enhanced, and the generated images are ensured to maintain a high degree of consistency in the long sequence, thereby improving the generation quality and efficiency of long-sequence images.
[0087] See also Figure 1 , Figure 101 is an application environment diagram of an image processing method according to an exemplary embodiment. The application environment may include a client 01 and a server 02. The server 02 may communicate with the client 01 in a wired or wireless manner, which is not limited in this application.
[0088] The client 01 can be used to provide image processing services to any user. Specifically, the client 01 can include but is not limited to smart phones, desktop computers, tablet computers, laptop computers, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, smart wearable devices and other types of electronic devices, and can also be software running on the above electronic devices, such as applications. Optionally, the operating system running on the electronic device can include but is not limited to Android, IOS, Linux, Windows, etc.
[0089] The server 02 may provide background services for the client 01. Specifically, the server 02 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services.
[0090] It should be noted that Figure 1 This is only one application environment of the image processing method provided in this application. In practical applications, other application environments may also be included.
[0091] Figure 2 is a flowchart of an image processing method according to an exemplary embodiment. Figure 2 As shown, this method can be used to Figure 1 In the server in, the method may at least include the following steps:
[0092] In step S201, a first noise image to be processed, a description text, and an initial image sequence corresponding to the description text are obtained.
[0093] In a specific embodiment, the first noise image to be processed may be a randomly generated noise image, and the noise image may be denoised by describing the text and the feature information corresponding to the initial image sequence to generate the required image. The description text may include text describing characters, scenes, storylines, etc., and the initial image sequence may include a first preset number of initial images, and the initial images may include initial character images and initial scene images. Specifically, each initial image may be obtained by inputting the description text into the first preset diffusion network (Stable Diffusion) for the Vincent map for image generation, without the need for a high-quality image as a reference for image generation, thereby reducing the need for high-quality reference images. Specifically, the above-mentioned first preset number may be set according to actual application requirements, for example, it may be set to 3, 5, etc.
[0094] Optionally, in the above step S201, the above step of obtaining the initial image sequence corresponding to the description text may include:
[0095] Acquire a second noise image to be processed;
[0096] Inputting the description text into a text encoding network for text encoding to obtain second text feature information;
[0097] Perform noise prediction based on the second text feature information and the second noise image to be processed to obtain third predicted noise information;
[0098] The second to-be-processed noisy image is subjected to image denoising based on the third predicted noise information to obtain an initial image sequence.
[0099] In a specific embodiment, the second noise image to be processed can be a randomly generated noise image, and the text encoding network can be used to extract the features of the text. The specific structure can be set according to the actual application requirements, for example, it can be a text encoder (Text Encoder) in a multimodal pre-trained neural network (Contrastive Language-Image Pre-Training, CLIP). The second text feature information can be used to characterize the character features, scene features, and semantic features of the plot corresponding to the description text.
[0100] In practical applications, the second text feature information and the second noise image to be processed can be input into a preset noise prediction network for noise prediction to obtain the third predicted noise information. Further, the third predicted noise information is removed from the second noise image to be processed, and the noise image without the predicted noise is used as a new second noise image to be processed. The noise prediction and image denoising process are repeated until the preset conditions are met, and the obtained noise image is decoded, and the output image is determined as the above-mentioned initial image. Optionally, the above-mentioned process of determining the initial image can be repeated multiple times to obtain multiple initial images, that is, to obtain an initial image sequence. Specifically, the above-mentioned preset noise prediction network can be a Unet network, and the preset conditions can be set according to actual application requirements, for example, the number of repetitions reaches a preset number, the image without the predicted noise meets the preset requirements, etc.; the preset number and preset requirements can be set according to actual application requirements.
[0101] In step S203, the initial image sequence is input into the self-attention network for self-attention learning to extract the association relationship between the initial images in the initial image sequence and obtain the first image feature information.
[0102] In a specific embodiment, the association relationship may represent the consistency between the initial images in terms of character features, scene style, etc., and the first image feature information may be used to characterize the consistency features between the initial images.
[0103] In an alternative embodiment, Figure 3 is a schematic diagram of a process for determining first image feature information according to an exemplary embodiment. Figure 3 As shown, in the above step S203, the above-mentioned inputting the initial image sequence into the self-attention network for self-attention learning to extract the association relationship between the initial images in the initial image sequence to obtain the first image feature information may include:
[0104] In step S301, a first preset number of initial images are subjected to noise addition processing to obtain a first preset number of initial noise images, each of which corresponds to a piece of noise information.
[0105] In a specific embodiment, the initial image sequence may include a first preset number of initial images arranged based on a first preset order. Optionally, for each initial image, Gaussian noise can be gradually added to the initial image according to a preset noise addition rule and a preset number of steps to obtain a corresponding initial noise image. Specifically, a certain amount of Gaussian noise is added to the initial image to obtain a result image after the first addition of Gaussian noise, and then a certain amount of Gaussian noise is added to the first result image again to obtain a result image after the second addition of Gaussian noise. Repeat the above-mentioned step of adding Gaussian noise for a preset number of steps to obtain a noise image close to Gaussian noise, that is, Gaussian noise is gradually added to the initial image. Specifically, the first preset order and the preset number of steps can be set according to actual application requirements.
[0106] In step S303, a first preset number of initial noise images are input into a first image coding network for image coding to obtain image features corresponding to each initial noise image.
[0107] In a specific embodiment, the first image encoding network may be used to perform feature extraction on the initial noise image to obtain image features corresponding to the initial noise image.
[0108] In step S305, a first key weight matrix corresponding to the image feature corresponding to each initial noise image and a corresponding first value weight matrix are obtained.
[0109] In step S307, the product of the image feature corresponding to each initial noise image and the corresponding first key weight matrix is determined as the first key feature corresponding to each initial noise image.
[0110] In step S309, the product of the image feature corresponding to each initial noise image and the corresponding first value weight matrix is determined as the first value feature corresponding to each initial noise image.
[0111] In step S311, noise information corresponding to the target noise image is determined as a first query feature.
[0112] Specifically, the target noise image may be an initial noise image corresponding to the last initial image in the initial image sequence.
[0113] In step S313, the similarity between the first query feature and the first key feature corresponding to each initial noise image is determined to obtain the first attention data corresponding to each initial noise image.
[0114] Specifically, the above-mentioned corresponding first attention data can be used to characterize the importance of the above-mentioned corresponding first value feature.
[0115] In step S315, based on the first attention data corresponding to each initial noise image, the corresponding first value feature is weighted to obtain first image feature information.
[0116] In a specific embodiment, the first image feature information can be calculated using the following formula:
[0117]
[0118] Among them, Q2 represents the query of the self-attention mechanism, that is, the first query feature mentioned above, K2 represents the key of the self-attention mechanism, that is, the first key feature mentioned above, V2 represents the value of the self-attention mechanism, that is, the first value feature mentioned above, d2 represents the dimension, K2=X2W K2 ,v2=X2W V2 , X2 represents the image features corresponding to each of the above initial noise images, W K2 ,W V2 Represent the above-mentioned first key weight matrix and first value weight matrix respectively.
[0119] Specifically, we can first calculate the image features corresponding to each initial noise image based on the self-attention network and then compare them with W K2 ,W V2 The product between them is used to obtain the first key feature and the first value feature, and the noise information of the initial noise image corresponding to the last initial image is determined as the first query feature. In practical applications, the query feature represents the "interest" of each position in the image feature in other positions. The role of the key feature is to match the query feature in the attention mechanism, and the value feature contains the specific information of each position in the image feature. Then, the similarity between the first query feature and the first key feature can be calculated based on the self-attention network. For example, the product between the first query feature and the first key feature is calculated to obtain Q2K2 T . Calculate the quotient of the similarity and the square root of dimension d2, and get Based on the quotient, the softmax function and V2, the first attention data is calculated, and the attention data is used to characterize the importance of V2. V2 is weighted and summed according to the importance to obtain the first image feature information.
[0120] In the above embodiment, based on the image features of each initial image in the initial image sequence, the corresponding key and value are determined, and the noise information of the initial noise image corresponding to the last initial image in the initial image sequence is determined as a query, and self-attention calculation is performed, that is, the query comes from the last initial image, and the key and value come from the last initial image and other images in the initial image sequence, so as to realize the sharing of keys and values in the sequence, and extract the association relationship in the long sequence of images through the attention calculation across images, and obtain the image feature information that characterizes the consistency between images, so that the subsequently generated images can maintain consistency in the long sequence.
[0121] In step S205, noise prediction is performed based on the first image feature information and the first noise image to be processed to obtain first updated noise information.
[0122] In a specific embodiment, the first noise image to be processed may correspond to the first initial noise information. The first updated noise information may be the noise information corresponding to the noise image obtained after removing the second predicted noise information from the first noise image to be processed.
[0123] Optionally, in the step S205, the step of performing noise prediction based on the first image feature information and the first noise image to be processed to obtain the first updated noise information may include:
[0124] Perform noise prediction based on the first image feature information and the first initial noise information to obtain second predicted noise information;
[0125] The difference information between the initial noise information and the second predicted noise information is determined as the first updated noise information.
[0126] In a specific embodiment, the refinement process of performing noise prediction based on the first image feature information and the first initial noise information to obtain the second predicted noise information can refer to the specific steps of performing noise prediction based on the second text feature information and the second noise image to be processed to obtain the third predicted noise information, which will not be repeated here.
[0127] In the above embodiment, noise prediction is performed based on the first image feature information and the first initial noise information to obtain the second predicted noise information, and the difference information between the initial noise information and the second predicted noise information is determined as the first updated noise information, and then the noise image is subsequently denoised so that the denoising direction is more inclined to consistency between images, so that the subsequently generated images can maintain consistency in a long sequence.
[0128] In step S207, noise prediction is performed based on the fused feature information and the first updated noise information to obtain first predicted noise information.
[0129] Specifically, the fused feature information and the first updated noise information may be input into the preset noise prediction network to perform noise prediction, so as to obtain the first predicted noise information.
[0130] In step S209, image denoising is performed on the first to-be-processed noisy image based on the first predicted noise information to obtain a target image.
[0131] In a specific embodiment, the obtained target image can follow the content of the description text while maintaining semantic consistency between the characters and the scene.
[0132] Optionally, in the above step S209, performing image denoising on the first to-be-processed noisy image based on the first predicted noise information to obtain the target image may include:
[0133] Performing image denoising on the first noise image to be processed based on the first predicted noise information to obtain an updated noise image;
[0134] The updated noise image is used again as the first noise image to be processed, and the initial image sequence is repeatedly input into the self-attention network for self-attention learning, until the first noise image to be processed is subjected to image denoising based on the first predicted noise information to obtain the target image, until the preset conditions are met and the target image is obtained.
[0135] In a specific embodiment, the first noise image to be processed may correspond to the first initial noise information, the updated noise image may correspond to the second updated noise information, and the second updated noise information may be the difference information between the first initial noise information and the first predicted noise information.
[0136] In the above embodiment, the first predicted noise information is removed from the first noise image to be processed, and the obtained noise image is the above updated noise image. Then, the updated noise image is used as the new first noise image to be processed, and the above steps of self-attention learning to image denoising are repeated until the preset conditions are met, the obtained noise image is decoded, and the output image is determined as the target image, which can improve the quality of the generated image. Specifically, the preset conditions can be set according to actual application requirements, for example, the number of repetitions reaches a preset number, the image with the predicted noise removed meets the preset requirements, etc.; the preset number and preset requirements can be set according to actual application requirements.
[0137] Optionally, the noise prediction and image denoising processing can be performed based on a second preset diffusion network. Specifically, the parameters of the second preset diffusion network can reuse the parameters of the first preset diffusion network, thereby reducing the consumption of computing resources in the image generation process.
[0138] In an optional embodiment, before performing noise prediction based on the first updated noise information to obtain the first predicted noise information, the method may further include:
[0139] Inputting the initial image, the description text and the first updated noise information into a multimodal cross attention network for multimodal cross attention learning to extract second image feature information corresponding to the initial image and first text feature information corresponding to the description text, and fusing the second image feature information and the first text feature information to obtain fused feature information;
[0140] Accordingly, performing noise prediction based on the first updated noise information to obtain the first predicted noise information includes:
[0141] Noise prediction is performed based on the fused feature information and the first updated noise information to obtain first predicted noise information.
[0142] In a specific embodiment, the multimodal cross-attention network may include a text cross-attention network and an image cross-attention network. The image cross-attention network may be used to process image features, and the text cross-attention network may be used to process text features. Specifically, an independent cross-attention layer may be set in the Unet network, or a multi-layer text cross-attention network and a multi-layer image cross-attention network may be set. The text cross-attention network and the image cross-attention network process image features and text features respectively. The second image feature information and the first text feature information may be spliced to obtain fused feature information.
[0143] In the above embodiment, through multimodal fusion, text and image inputs can be processed simultaneously to achieve more flexible content generation. Furthermore, by combining the fusion features including text features and image features for noise prediction, the subsequent text and the generated image can be kept consistent, and an image that is more consistent with the text description can be generated, thereby improving the generation quality of long sequence images.
[0144] In an alternative embodiment, Figure 4 is a schematic diagram of a flow chart of cross attention learning according to an exemplary embodiment, such as Figure 4 As shown, the above-mentioned inputting the initial image, the description text and the first updated noise information into the multimodal cross attention network for multimodal cross attention learning to extract the second image feature information corresponding to the initial image and the first text feature information corresponding to the description text may include:
[0145] In step S401, the initial character image and the initial scene image are respectively input into the second image coding network for image coding to obtain character feature information corresponding to the initial character image and scene feature information corresponding to the initial scene image.
[0146] In a specific embodiment, the second image coding network can be used to extract image features, which can be specifically set according to actual application requirements, for example, it can be an image encoder (Vision Encoder) in a multimodal pre-trained neural network (Contrastive Language-Image Pre-Training, CLIP). In practical applications, the character image and the scene image are encoded separately through the image coding network, and the image modality features are extracted, and the details, style and other features of the image can be extracted, so that the subsequently generated images can maintain visual consistency and quality.
[0147] In step S403, the character feature information and the scene feature information are spliced to obtain third image feature information.
[0148] In step S405, the third image feature information and the first updated noise information are input into the text cross-attention network for cross-attention learning to obtain the second image feature information.
[0149] In a specific embodiment, the feature relationship between the character image and the scene image is modeled based on the third image feature information and the first updated noise information based on the image cross-attention network, so that the character image is effectively integrated into the scene image, the character is naturally integrated into the scene, and the semantic consistency between the character and the scene is maintained, so as to obtain the second image feature information.
[0150] In an alternative embodiment, Figure 5 is a schematic diagram of a process for determining feature information of a second image according to an exemplary embodiment. Figure 5 As shown, in the above step S405, the above-mentioned inputting the third image feature information and the first updated noise information into the text cross attention network for cross attention learning to obtain the second image feature information may include:
[0151] In step S501, a second key weight matrix and a corresponding second value weight matrix corresponding to the third image feature information are obtained.
[0152] In step S503, the product of the third image feature information and the corresponding second key weight matrix is determined as the second key feature.
[0153] In step S505, the product of the third image feature information and the corresponding second value weight matrix is determined as the second value feature.
[0154] In step S507, the first updated noise information is determined as the second query feature.
[0155] In step S509, the similarity between the second query feature and the second key feature is determined to obtain second attention data.
[0156] Specifically, the second attention data can be used to characterize the importance of the second value feature.
[0157] In step S511, the second value feature is weighted based on the second attention data to obtain second image feature information.
[0158] In a specific embodiment, the second image feature information can be calculated using the following formula:
[0159]
[0160] Wherein, Q1 represents the second query feature (Query), K1 represents the second key feature (Key), V1 represents the second value feature (Value), d1 represents the dimension, K1=X1W K1 ,V1=X1W V1 , X1 represents the second image feature information, W K1 ,W V1 Respectively represent the above-mentioned second key weight matrix and second value weight matrix.
[0161] Specifically, we can first calculate the second image features based on the cross attention network and then compare them with W K1 ,W V1 The second key feature and the second value feature are obtained by multiplying the first updated noise information to obtain the second query feature. Then, the similarity between the second query feature and the second key feature can be calculated based on the cross attention network. For example, the product between the second query feature and the second key feature is calculated to obtain Q1K1 T . Calculate the quotient of the similarity and the square root of dimension d1, and get Based on the quotient, the softmax function and V1, the second attention data is calculated, and the attention data is used to characterize the importance of V1. According to the importance, V1 is weighted and summed to obtain the second image feature information.
[0162] In the above embodiment, an independent cross-attention layer is set in the Unet network, and the image cross-attention network is used to process image features to ensure that the image modal information is fully processed and understood before fusion, and the character graph is effectively integrated into the scene graph, so that the character is naturally integrated into the scene, maintaining the semantic consistency between the character and the scene.
[0163] In step S407, the description text is input into a text encoding network for text encoding to obtain second text feature information.
[0164] Specifically, the second text feature information can be used to characterize the character features, scene features, and semantic features of the plot corresponding to the description text. In practical applications, the description text is encoded through a text encoding network to obtain text modal features, which not only include the features of the characters and scenes, but also extract the semantic features of the plot, thereby ensuring the subsequent generation of images with narrative coherence.
[0165] In step S409, the second text feature information and the first updated noise information are input into the text cross-attention network for cross-attention learning to obtain the first text feature information.
[0166] In the embodiments of this specification, through the decoupled cross-attention operation, on the one hand, the corresponding Key and Value are determined based on the text features, and the noise is used as the Query for attention calculation. On the other hand, the corresponding Key and Value are determined based on the image features, and the noise is used as the Query for attention calculation, so that the most relevant part of the text description is dynamically focused on during the image generation process, thereby generating an image that is more consistent with the storyline. In addition, in the process of processing the image modal features, the character graph is effectively integrated into the scene graph to ensure the natural integration of the character in the scene, and the image features are integrated with the text features, so that the subsequently generated image can follow the content of the text description and maintain the semantic consistency between the character and the scene.
[0167] In an alternative embodiment, Figure 6 is a schematic diagram of a process for determining first text feature information according to an exemplary embodiment. Figure 6 As shown, in the above step S409, the above-mentioned inputting the second text feature information and the first updated noise information into the text cross attention network for cross attention learning to obtain the first text feature information may include:
[0168] In step S601, a third key weight matrix and a corresponding third value weight matrix corresponding to the second text feature information are obtained.
[0169] In step S603, the product of the second text feature information and the corresponding third key weight matrix is determined as the third key feature.
[0170] In step S605, the product of the second text feature information and the corresponding third value weight matrix is determined as the third value feature.
[0171] In step S607, the first updated noise information is determined as the second query feature.
[0172] In step S609, the similarity between the second query feature and the third key feature is determined to obtain the third attention data.
[0173] Specifically, the third attention data can be used to characterize the importance of the third value feature.
[0174] In step S611, the third value feature is weighted based on the third attention data to obtain the first text feature information.
[0175] Specifically, the above-mentioned refinement step of inputting the second text feature information and the first updated noise information into the text cross-attention network for cross-attention learning to obtain the first text feature information can refer to the above-mentioned specific step of inputting the third image feature information and the first updated noise information into the image cross-attention network for cross-attention learning to obtain the second image feature information, which will not be repeated here.
[0176] In the above embodiment, an independent cross-attention layer is set in Unet, and the text cross-attention network is used to process text features to ensure that the text modal information is fully processed and understood before fusion, and the text features are fused with the character image and the scene image, so that the subsequent noise prediction and denoising based on the fused features can follow the text content and ensure the semantic consistency between the character and the scene. Through multimodal fusion, text and image inputs can be processed simultaneously to achieve more flexible content generation.
[0177] Optionally, the above cross-attention network can use a multi-head attention mechanism to further refine the interactive features of text and images to improve the quality and consistency of the generated results. The multi-head attention mechanism allows the model to perform parallel calculations on multiple focus points, thereby capturing more subtle interactions between text and images. This refined interactive feature can not only improve the level of detail of the generated image, but also enhance the overall coordination and visual effect of the image.
[0178] Optionally, the above method may further include:
[0179] Get the description text sequence;
[0180] For each description text, the steps of inputting the initial image sequence into the self-attention network for self-attention learning, performing image denoising on the first to-be-processed noise image based on the first predicted noise information, and obtaining a target image are performed to obtain the target image corresponding to each description text;
[0181] Based on the second preset order, the target images corresponding to each description text are sorted to obtain a target image sequence.
[0182] In a specific embodiment, the description text sequence may include a second preset number of description texts arranged based on a second preset order, and each description text may correspond to an initial image sequence. The description text sequence may include multiple description texts describing characters, scenes, storylines, etc., such as text descriptions of scripts, long-sequence comics, etc.; the second preset order may be the order of development of the storyline. Specifically, the second preset number may be determined according to actual application requirements.
[0183] In the above embodiment, for each description text in the description text sequence, the above image processing method is executed to obtain a target image corresponding to each description text, and then based on the second preset order, the target image corresponding to each description text is sorted to obtain a sequence of multiple target images arranged in the order of story development. The target image sequence can ensure the consistency between the text and the generated image, as well as the consistency between multiple images in the target image sequence, and ensure the consistency between the characters and scenes in the generated image, thereby improving the coherence of long sequence images and ensuring the generation quality of long sequence images.
[0184] The present application embodiment provides an image processing method for intelligent generation of long sequence images, such as Figure 7As shown, on the one hand, after each initial image in the initial image sequence is subjected to noise processing, the corresponding image features are extracted, and the product of the image features corresponding to each initial noise image and the corresponding first key weight matrix is determined as the first key feature corresponding to each initial noise image, and the product of the image features corresponding to each initial noise image and the corresponding first value weight matrix is determined as the first value feature corresponding to each initial noise image, and the noise information of the initial noise image corresponding to the last initial image in the initial image sequence is determined as the first query feature, and self-attention calculation is performed to extract the correlation relationship between the initial images in the initial image sequence to obtain the first image feature information, and noise prediction is performed based on the first image feature information and the first noise image to be processed to obtain the first updated noise information. On the other hand, character feature information is extracted based on the character image, scene feature information is extracted based on the scene image, and the third image feature information is obtained by splicing, the product of the third image feature information and the corresponding second key weight matrix is determined as the second key feature, the product of the third image feature information and the corresponding second value weight matrix is determined as the second value feature, the first updated noise information is determined as the second query feature, and cross-attention calculation is performed to obtain the second image feature information; second text feature information is extracted based on the description text, the product of the second text feature information and the corresponding third key weight matrix is determined as the third key feature, the product of the second text feature information and the corresponding third value weight matrix is determined as the third value feature, the first updated noise information is determined as the second query feature, and cross-attention calculation is performed to obtain the first text feature information, and the second image feature information and the first text feature information are fused to obtain fused feature information. Further, noise prediction is performed based on the fused feature information and the first updated noise information to obtain the first predicted noise information, and image denoising is performed on the first noise image to be processed based on the first predicted noise information to obtain the target image.
[0185] Therefore, through decoupled cross-attention operations and multimodal fusion, attention calculations are performed based on text features and image features respectively, so that the most relevant parts of the text description can be dynamically focused on during the image generation process, and images that are more consistent with the storyline can be generated. In addition, in the process of processing the image modal features, the character graph is effectively integrated into the scene graph to ensure the natural integration of the character in the scene, and the image features are fused with the text features so that the subsequently generated images can follow the content of the text description and maintain the semantic consistency between the character and the scene. In addition, based on the long-sequence consistency self-attention mechanism, through cross-image attention calculation, the correlation relationship in the long sequence image is extracted, which can ensure that the generated image maintains a high consistency in the long sequence, and improve the image generation quality and efficiency.
[0186] The image processing method proposed in this application can be applied to comic creation, animation production, video content generation and other fields. It can not only improve the efficiency of content creation, but also provide creators with more creative possibilities. For example, in the comic creation process, a coherent picture of characters and scenes can be automatically generated, reducing the time and effort of manual drawing; in animation production, a highly consistent character action sequence can be generated to improve the efficiency and quality of animation production. In addition, it can also be applied in education, entertainment, advertising and other industries. Through this image processing method, educational content can be made more vivid and interactive, entertainment content can be made more diverse, and advertising content can be made more personalized.
[0187] Figure 8 FIG. 1 is a block diagram of an image processing device according to an exemplary embodiment. Figure 8 , the device comprises:
[0188] A first acquisition module 810 is configured to acquire a first noise image to be processed, a description text, and an initial image sequence corresponding to the description text;
[0189] The self-attention learning module 820 is configured to input the initial image sequence into a self-attention network to perform self-attention learning, so as to extract the association relationship between the initial images in the initial image sequence and obtain first image feature information;
[0190] A first noise prediction module 830 is configured to perform noise prediction based on the first image feature information and the first noise image to be processed to obtain first updated noise information;
[0191] A second noise prediction module 840 is configured to perform noise prediction based on the first updated noise information to obtain first predicted noise information;
[0192] The denoising module 850 is configured to perform image denoising on the first to-be-processed noisy image based on the first predicted noise information to obtain a target image.
[0193] In an optional embodiment, the initial image sequence includes a first preset number of the initial images arranged in a first preset order, and the self-attention learning module includes:
[0194] A noise adding unit is configured to perform noise adding processing on a first preset number of the initial images respectively to obtain a first preset number of initial noise images, each of the initial noise images corresponding to a noise information;
[0195] A first image encoding unit is configured to input a first preset number of the initial noise images into a first image encoding network for image encoding, so as to obtain an image feature corresponding to each of the initial noise images;
[0196] A third acquisition unit is configured to acquire a first key weight matrix corresponding to the image feature corresponding to each of the initial noise images, and a corresponding first value weight matrix;
[0197] A first key feature determining unit is configured to perform a multiplication of the image feature corresponding to each of the initial noise images and the corresponding first key weight matrix to determine the first key feature corresponding to each of the initial noise images;
[0198] A first value feature determining unit is configured to perform a multiplication of the image feature corresponding to each of the initial noise images and the corresponding first value weight matrix to determine the first value feature corresponding to each of the initial noise images;
[0199] A first query feature determination unit is configured to determine noise information corresponding to a target noise image as a first query feature; the target noise image is an initial noise image corresponding to the last initial image in the initial image sequence;
[0200] A first attention data determining unit is configured to determine the similarity between the first query feature and the first key feature corresponding to each of the initial noise images, and obtain the first attention data corresponding to each of the initial noise images; the corresponding first attention data is used to characterize the importance of the corresponding first value feature;
[0201] The first image feature information determination unit is configured to perform weighted processing on the corresponding first value feature based on the first attention data corresponding to each of the initial noise images to obtain the first image feature information.
[0202] In an optional embodiment, the device further comprises:
[0203] a cross-attention learning module, configured to input the initial image, the description text and the first updated noise information into a multimodal cross-attention network to perform multimodal cross-attention learning, so as to extract second image feature information corresponding to the initial image and first text feature information corresponding to the description text, and to fuse the second image feature information and the first text feature information to obtain fused feature information;
[0204] Accordingly, the second noise prediction module includes:
[0205] The first predicted noise information determining unit is configured to perform noise prediction based on the fused feature information and the first updated noise information to obtain the first predicted noise information.
[0206] In an optional embodiment, the multimodal cross-attention network includes a text cross-attention network and an image cross-attention network, the initial image includes an initial character image and an initial scene image, and the cross-attention learning module includes:
[0207] A second image encoding unit is configured to input the initial character image and the initial scene image into a second image encoding network for image encoding, respectively, to obtain character feature information corresponding to the initial character image and scene feature information corresponding to the initial scene image;
[0208] a splicing unit configured to perform splicing processing on the character feature information and the scene feature information to obtain third image feature information;
[0209] A first cross-attention learning unit is configured to input the third image feature information and the first updated noise information into the text cross-attention network to perform cross-attention learning to obtain the second image feature information;
[0210] A text encoding unit is configured to input the description text into a text encoding network for text encoding to obtain second text feature information;
[0211] The second cross-attention learning unit is configured to input the second text feature information and the first updated noise information into the image cross-attention network for cross-attention learning to obtain the first text feature information.
[0212] In an optional embodiment, the first cross-attention learning unit comprises:
[0213] A first acquisition unit is configured to acquire a second key weight matrix and a corresponding second value weight matrix corresponding to the third image feature information;
[0214] A second key feature determining unit is configured to determine the product of the third image feature information and the corresponding second key weight matrix as the second key feature;
[0215] A second value feature determining unit is configured to determine the product of the third image feature information and the corresponding second value weight matrix as the second value feature;
[0216] A second query feature determination unit is configured to determine the first updated noise information as a second query feature;
[0217] A second attention data determining unit is configured to determine the similarity between the second query feature and the second key feature to obtain second attention data; the second attention data is used to characterize the importance of the second value feature;
[0218] The second image feature information determination unit is configured to perform weighted processing on the second value feature based on the second attention data to obtain the second image feature information.
[0219] In an optional embodiment, the second cross-attention learning unit comprises:
[0220] A second acquisition unit is configured to execute acquisition of a third key weight matrix and a corresponding third value weight matrix corresponding to the second text feature information;
[0221] A third key feature determining unit is configured to determine the product of the second text feature information and the corresponding third key weight matrix as the third key feature;
[0222] A third value feature determining unit is configured to determine the product of the second text feature information and the corresponding third value weight matrix as the third value feature;
[0223] A second query feature determination unit is configured to determine the first updated noise information as a second query feature;
[0224] A third attention data determining unit is configured to determine the similarity between the second query feature and the third key feature to obtain third attention data; the third attention data is used to characterize the importance of the third value feature;
[0225] The first text feature information determining unit is configured to perform weighted processing on the third value feature based on the third attention data to obtain the first text feature information.
[0226] In an optional embodiment, the first noise image to be processed corresponds to first initial noise information, and the first noise prediction module includes:
[0227] A first noise prediction unit is configured to perform noise prediction based on the first image feature information and the first initial noise information to obtain second predicted noise information;
[0228] The first updated noise information determining unit is configured to determine the difference information between the initial noise information and the second predicted noise information as the first updated noise information.
[0229] In an optional embodiment, the device further comprises:
[0230] A second acquisition module is configured to acquire a description text sequence; the description text sequence includes a second preset number of description texts arranged in a second preset order, and each description text corresponds to one of the initial image sequences;
[0231] An image processing module is configured to execute, for each of the description texts, the steps of inputting the initial image sequence into a self-attention network for self-attention learning, to performing image denoising on the first to-be-processed noisy image based on the first predicted noise information to obtain a target image, and obtain a target image corresponding to each of the description texts;
[0232] The target image sequence determination module is configured to sort the target images corresponding to each of the description texts based on the second preset order to obtain a target image sequence.
[0233] In an optional embodiment, the first acquisition module includes:
[0234] A fourth acquisition unit is configured to acquire a second noise image to be processed;
[0235] A text encoding unit is configured to input the description text into a text encoding network for text encoding to obtain second text feature information;
[0236] A second noise prediction unit is configured to perform noise prediction based on the second text feature information and the second noise image to be processed to obtain third predicted noise information;
[0237] The denoising unit is configured to perform image denoising on the second to-be-processed noisy image based on the third predicted noise information to obtain the initial image sequence.
[0238] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0239] Fig. 9 is a block diagram of an electronic device for image processing according to an exemplary embodiment. The electronic device may be a terminal, and its internal structure diagram may be as shown in FIG. Fig. 9As shown. The electronic device includes a processor, a memory, a network interface, a display screen and an input device connected via a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, an image processing method is implemented. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the electronic device can be a touch layer covering the display screen, or a key, trackball or touchpad provided on the housing of the electronic device, or an external keyboard, touchpad or mouse, etc.
[0240] Fig.10 is a block diagram of another electronic device for image processing according to an exemplary embodiment. The electronic device may be a server, and its internal structure diagram may be as shown in FIG. Fig.10 As shown. The electronic device includes a processor, a memory and a network interface connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, an image processing method is implemented.
[0241] Those skilled in the art will understand that Fig. 9 or Fig.10 The structure shown in the figure is merely a block diagram of a partial structure related to the scheme of the present disclosure, and does not constitute a limitation on the electronic device to which the scheme of the present disclosure is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0242] In an exemplary embodiment, an electronic device for image processing is also provided, including: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the image processing method in the embodiment of the present disclosure.
[0243] In an exemplary embodiment, a computer-readable storage medium is further provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the steps of any image processing method in the above embodiments.
[0244] In an exemplary embodiment, a computer program product is also provided, including a computer program, and when the computer program is executed by a processor, the image processing method provided in any one of the above embodiments is implemented.
[0245] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0246] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present application. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the following claims.
[0247] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. An image processing method, characterized in that: The method comprises: Acquire a first noise image to be processed, a description text, and an initial image sequence corresponding to the description text; Inputting the initial image sequence into a self-attention network for self-attention learning to extract the association relationship between the initial images in the initial image sequence to obtain first image feature information; Perform noise prediction based on the first image feature information and the first noise image to be processed to obtain first updated noise information; Perform noise prediction based on the first updated noise information to obtain first predicted noise information; Image denoising is performed on the first noisy image to be processed based on the first predicted noise information to obtain a target image.
2. The method according to claim 1, characterized in that The initial image sequence includes a first preset number of the initial images arranged in a first preset order, and the initial image sequence is input into a self-attention network for self-attention learning to extract the association relationship between the initial images in the initial image sequence, and the first image feature information is obtained, including: Performing noise addition processing on a first preset number of the initial images respectively to obtain a first preset number of initial noise images, each of the initial noise images corresponding to a noise information; Inputting a first preset number of the initial noise images into a first image coding network for image coding to obtain image features corresponding to each of the initial noise images; Obtaining a first key weight matrix corresponding to the image feature corresponding to each of the initial noise images, and a corresponding first value weight matrix; Determine the product of the image feature corresponding to each of the initial noise images and the corresponding first key weight matrix as the first key feature corresponding to each of the initial noise images; The product of the image feature corresponding to each of the initial noise images and the corresponding first value weight matrix is determined as the first value feature corresponding to each of the initial noise images; Determine the noise information corresponding to the target noise image as the first query feature; the target noise image is an initial noise image corresponding to the last initial image in the initial image sequence; Determine the similarity between the first query feature and the first key feature corresponding to each of the initial noise images, and obtain the first attention data corresponding to each of the initial noise images; the corresponding first attention data is used to characterize the importance of the corresponding first value feature; Based on the first attention data corresponding to each of the initial noise images, the corresponding first value features are weighted to obtain the first image feature information.
3. The method according to claim 1, characterized in that: Before performing noise prediction based on the first updated noise information to obtain first predicted noise information, the method further includes: Inputting the initial image, the description text and the first updated noise information into a multimodal cross attention network for multimodal cross attention learning to extract second image feature information corresponding to the initial image and first text feature information corresponding to the description text, and fusing the second image feature information and the first text feature information to obtain fused feature information; Correspondingly, performing noise prediction based on the first updated noise information to obtain first predicted noise information includes: Noise prediction is performed based on the fused feature information and the first updated noise information to obtain the first predicted noise information.
4. The method according to claim 3, characterized in that The multimodal cross-attention network includes a text cross-attention network and an image cross-attention network, the initial image includes an initial character image and an initial scene image, and the initial image, the description text and the first updated noise information are input into the multimodal cross-attention network for multimodal cross-attention learning to extract the second image feature information corresponding to the initial image and the first text feature information corresponding to the description text include: Inputting the initial character image and the initial scene image into a second image coding network for image coding respectively, to obtain character feature information corresponding to the initial character image and scene feature information corresponding to the initial scene image; Performing splicing processing on the character feature information and the scene feature information to obtain third image feature information; Inputting the third image feature information and the first updated noise information into the text cross attention network for cross attention learning to obtain the second image feature information; Inputting the description text into a text encoding network for text encoding to obtain second text feature information; The second text feature information and the first updated noise information are input into the image cross-attention network for cross-attention learning to obtain the first text feature information.
5. The method according to claim 4, characterized in that The step of inputting the third image feature information and the first updated noise information into the text cross attention network for cross attention learning to obtain the second image feature information comprises: Obtaining a second key weight matrix and a corresponding second value weight matrix corresponding to the third image feature information; Determine the product of the third image feature information and the corresponding second key weight matrix as the second key feature; Determine the product of the third image feature information and the corresponding second value weight matrix as the second value feature; determining the first updated noise information as a second query feature; Determine the similarity between the second query feature and the second key feature to obtain second attention data; the second attention data is used to characterize the importance of the second value feature; The second value feature is weighted based on the second attention data to obtain the second image feature information.
6. The method according to claim 4, characterized in that The step of inputting the second text feature information and the first updated noise information into the image cross attention network for cross attention learning to obtain the first text feature information comprises: Obtaining a third key weight matrix and a corresponding third value weight matrix corresponding to the second text feature information; Determine the product of the second text feature information and the corresponding third key weight matrix as the third key feature; Determine the product of the second text feature information and the corresponding third value weight matrix as the third value feature; determining the first updated noise information as a second query feature; Determine the similarity between the second query feature and the third key feature to obtain third attention data; the third attention data is used to characterize the importance of the third value feature; The third value feature is weighted based on the third attention data to obtain the first text feature information.
7. The method according to claim 1, characterized in that The first noise image to be processed corresponds to first initial noise information, and performing noise prediction based on the first image feature information and the first noise image to be processed to obtain first updated noise information includes: Perform noise prediction based on the first image feature information and the first initial noise information to obtain second predicted noise information; The difference information between the initial noise information and the second predicted noise information is determined as the first updated noise information.
8. The method according to any one of claims 1 to 7, characterized in that: The performing image denoising on the first to-be-processed noisy image based on the first predicted noise information to obtain a target image comprises: Based on the first predicted noise information, the first noise image to be processed is subjected to image denoising to obtain an updated noise image; the first noise image to be processed corresponds to the first initial noise information, the updated noise image corresponds to the second updated noise information, and the second updated noise information is the difference information between the first initial noise information and the first predicted noise information; The updated noise image is used again as the first noise image to be processed, and the steps of inputting the initial image sequence into the self-attention network for self-attention learning are repeated until the first noise image to be processed is subjected to image denoising based on the first predicted noise information to obtain a target image, until a preset condition is met and the target image is obtained.
9. The method according to claim 1, characterized in that: The method further comprises: Acquire a description text sequence; the description text sequence includes a second preset number of the description texts arranged in a second preset order, and each of the description texts corresponds to one of the initial image sequences; For each of the description texts, the steps of inputting the initial image sequence into the self-attention network for self-attention learning, and performing image denoising on the first to-be-processed noisy image based on the first predicted noise information to obtain a target image are performed to obtain a target image corresponding to each of the description texts; Based on the second preset order, the target images corresponding to each of the description texts are sorted to obtain a target image sequence.
10. The method according to claim 1, characterized in that The obtaining of the initial image sequence corresponding to the description text comprises: Acquire a second noise image to be processed; Inputting the description text into a text encoding network for text encoding to obtain second text feature information; Perform noise prediction based on the second text feature information and the second noise image to be processed to obtain third predicted noise information; The second to-be-processed noisy image is subjected to image denoising based on the third predicted noise information to obtain the initial image sequence.
11. An image processing device, characterized in that: The device comprises: A first acquisition module is configured to acquire a first noise image to be processed, a description text, and an initial image sequence corresponding to the description text; A self-attention learning module is configured to input the initial image sequence into a self-attention network to perform self-attention learning, so as to extract the association relationship between the initial images in the initial image sequence and obtain first image feature information; A first noise prediction module is configured to perform noise prediction based on the first image feature information and the first noise image to be processed to obtain first updated noise information; A second noise prediction module is configured to perform noise prediction based on the first updated noise information to obtain first predicted noise information; The denoising module is configured to perform image denoising on the first to-be-processed noisy image based on the first predicted noise information to obtain a target image.
12. An electronic device for image processing, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the image processing method according to any one of claims 1 to 10. 13 . A computer-readable storage medium, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device executes the image processing method according to claim 1 .
Citation Information
Cited By
Image meaning analysis scene consistency evaluation system based on visual model
CN120997650A
Scene consistency evaluation system for image meaning resolution based on visual model
CN120997650B
End-to-end target detection and orientation recognition method, device and equipment
CN122115922A