Method executed by electronic device, electronic device, storage medium and program product
By using the guidance information generated by the first AI network and the second AI network performs super-resolution processing, the problem of excessive model and time-consuming generation of high-resolution images in the prior art is solved, and fast and efficient high-resolution image generation is achieved.
Patent Information
- Application Number
- CN202311585734.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-24
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art is difficult to generate high-resolution images, and the models are usually too large and time-consuming to meet the user's usage needs.
By using the first AI network to generate guidance information, including spatial correlation guidance information and semantic correlation guidance information, the second AI network is guided to perform super-resolution processing and generate high-resolution images.
It realizes rapid and efficient generation of high-resolution images, reducing model size, calculation effort and time-consuming, and improving user experience.
Smart Images

Figure CN120047314A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a method executed by an electronic device, an electronic device, a storage medium, and a program product. Background Art
[0002] With the continuous development and progress of science and technology, it has become possible to use electronic devices to generate pictures beyond traditional imagination, and the generated pictures can be highly consistent with the user's intentions. On this basis, there is a strong user demand for high-quality AI image generation.
[0003] Current image generation technology can generate images with smaller resolutions, such as 512 × 512 pixels. However, users usually want to use electronic devices to obtain higher resolution images, such as 4000 × 3000 pixels, which puts higher requirements on image processing models.
[0004] Currently, the models that can generate high-resolution images are usually too large and time-consuming to meet user needs. Summary of the invention
[0005] The purpose of the embodiments of the present application is to solve the technical problem that the solution model for obtaining high-resolution generated images is usually too large and time-consuming.
[0006] According to one aspect of an embodiment of the present application, a method performed by an electronic device is provided, the method comprising:
[0007] Based on the input information, using a first AI (Artificial Intelligence) network, obtaining a first image and guidance information, where the guidance information includes at least one of spatial correlation guidance information and semantic correlation guidance information;
[0008] Based on the guidance information, a second AI network is used to perform super-resolution processing on the first image to obtain a second image.
[0009] Optionally, the spatial correlation guidance information includes spatial correlation weights between different spatial positions;
[0010] The semantic relevance guidance information includes at least one of a semantic relevance weight between a spatial position and text constraint content, and a semantic relevance weight between different spatial positions.
[0011] Optionally, the first AI network includes at least one first spatial attention module, and the spatial correlation guidance information includes spatial correlation guidance information in at least one first spatial attention module; and / or,
[0012] The semantic relevance guidance information includes semantic relevance guidance information corresponding to at least one word element.
[0013] Optionally, the second AI network includes at least one second spatial attention module and / or at least one second semantic attention module, and based on the guidance information, the second AI network is used to perform super-resolution processing on the first image, including at least one of the following:
[0014] Based on the spatial correlation guidance information, using at least one second spatial attention module, respectively performing spatial attention processing on the input first image features;
[0015] Based on the semantic relevance guidance information, at least one second semantic attention module is used to perform semantic attention processing on the input second image features respectively.
[0016] Optionally, the first AI network includes at least one first semantic attention module, the semantic relevance guidance information includes semantic relevance guidance information corresponding to at least one word in the at least one first semantic attention module, and based on the semantic relevance guidance information, at least one second semantic attention module is used to perform semantic attention processing on the input second image features respectively, including:
[0017] For each word-unit, the semantic relevance guidance information corresponding to the word-unit in at least one first semantic attention module is fused to obtain the semantic relevance guidance information corresponding to the word-unit;
[0018] Based on the semantic relevance guidance information corresponding to at least one word unit, at least one second semantic attention module is used to perform semantic attention processing on the input second image features.
[0019] Optionally, fusing the semantic relevance guidance information corresponding to the word-unit in at least one first semantic attention module includes:
[0020] For each execution of the first AI network, the semantic relevance guidance information corresponding to the word in at least one first semantic attention module is transformed into the same size as the first image and then superimposed to obtain first semantic relevance accumulation guidance information;
[0021] Superimposing the first semantic relevance accumulated guidance information obtained by executing the first AI network each time to obtain second semantic relevance accumulated guidance information;
[0022] The second first semantic relevance accumulated guidance information is normalized.
[0023] Optionally, the second AI network includes at least one second spatial attention module and / or at least one second semantic attention module, and based on the guidance information, the second AI network is used to perform super-resolution processing on the first image, including at least one of the following:
[0024] Based on the spatial correlation guidance information in at least one first spatial attention module, respectively perform spatial attention processing on the first image features corresponding to the corresponding second spatial attention modules;
[0025] The semantic relevance guidance information corresponding to at least one word-unit is fused, and based on the fused semantic relevance guidance information, semantic attention processing is performed on the second image features corresponding to at least one second semantic attention module.
[0026] Optionally, for each of the spatial correlation guidance information in the first spatial attention module, performing spatial attention processing on the first image feature corresponding to the corresponding second spatial attention module based on the spatial correlation guidance information includes:
[0027] Based on the first image feature corresponding to the corresponding second spatial attention module, obtain a fourth image feature and a fifth image feature;
[0028] Based on the spatial correlation guidance information, spatial attention processing is performed on the fourth image feature to obtain a sixth image feature;
[0029] Performing channel attention adjustment on the fifth image feature to obtain the seventh image feature;
[0030] The sixth image feature and the seventh image feature are fused.
[0031] Optionally, performing channel attention adjustment on the fifth image feature includes:
[0032] Compressing the fifth image feature in the spatial dimension to obtain an eighth image feature;
[0033] Based on the eighth image feature, obtaining the channel attention weight of the fifth image feature;
[0034] The channel attention of the fifth image feature is adjusted based on the channel attention weight.
[0035] Optionally, the semantic relevance guidance information corresponding to at least one word-element is fused, including:
[0036] Obtain a weight corresponding to at least one word unit;
[0037] Based on the weight corresponding to at least one word-gram, the semantic relevance guidance information corresponding to at least one word-gram is weightedly fused.
[0038] Optionally, obtaining a weight corresponding to at least one word-element includes:
[0039] Display at least one lemma;
[0040] In response to a first selection operation of a user on a target word-gram among the at least one word-gram, a weight corresponding to the at least one word-gram is determined.
[0041] Optionally, based on the fused semantic relevance guidance information, semantic attention processing is performed on the second image features corresponding to at least one second semantic attention module, including:
[0042] fusing the fused semantic relevance guidance information with a third image feature of at least one scale corresponding to the first image to obtain a global semantic feature of at least one scale;
[0043] Based on the global semantic features of at least one scale, semantic attention processing is performed on the second image features corresponding to the second semantic attention module of the corresponding scale.
[0044] Optionally, the second AI network further includes a correction module, and the method further includes:
[0045] Using the correction module, feature correction is performed in the row direction and / or column direction of the image feature corresponding to the first image.
[0046] Optionally, feature correction is performed in a row direction and / or a column direction of the image feature corresponding to the first image, including:
[0047] Using the dilated convolution, feature correction is performed in the row direction and / or column direction of the image features corresponding to the first image.
[0048] Optionally, using a dilated convolution, feature correction is performed in the row direction and the column direction of the image feature corresponding to the first image, including:
[0049] Determining a ninth image feature and a tenth image feature based on the image feature corresponding to the first image;
[0050] Compressing the ninth image feature in the row direction, and performing at least one dilated convolution in the column direction on the compressed ninth image feature to obtain an image feature corrected in the column direction;
[0051] Compressing the tenth image feature in the column direction, and performing at least one dilated convolution in the row direction on the compressed tenth image feature to obtain an image feature corrected in the row direction;
[0052] The features of each rectified image are fused.
[0053] Optionally, the editing process includes at least one of the following:
[0054] Image completion processing, image expansion processing, text-based image generation processing, image fusion processing, and image style conversion processing.
[0055] Optionally, the first AI network and / or the second AI network is a diffusion network.
[0056] Optionally, based on the input information, before using the first AI network to perform editing processing to obtain the first image and the guidance information, the method further includes:
[0057] Responding to the user's input operation, obtaining the target content text;
[0058] Based on the target content text, determine the corresponding recommended constraint content and display the recommended constraint content;
[0059] In response to a second selection operation of the user on the target constraint content in the recommended constraint content, the target constraint content and the target content text are determined as input information.
[0060] Optionally, the recommended constraint content includes: positive constraint content and / or negative constraint content.
[0061] Optionally, after the first image is obtained by using the first AI network to perform editing based on the input information, the method further includes:
[0062] Displaying the first image;
[0063] Receive super-resolution processing instructions from user feedback;
[0064] Based on the guidance information, a second AI network is used to perform super-resolution processing on the first image, including:
[0065] In response to the super-resolution processing instruction, based on the guidance information, use the second AI network to perform super-resolution processing on the first image;
[0066] The method further includes:
[0067] receiving a first image re-acquisition instruction fed back by a user;
[0068] In response to the instruction to re-acquire the first image, the operation of re-executing the editing process based on the input information using the first AI network to obtain the first image and then displaying the first image again.
[0069] According to another aspect of an embodiment of the present application, another method performed by an electronic device is provided, the method comprising:
[0070] acquiring a third image;
[0071] Performing super-resolution processing on the third image to obtain a fourth image;
[0072] Feature correction is performed in the row direction and / or column direction of the image features corresponding to the fourth image to obtain a fifth image.
[0073] Optionally, performing feature correction in a row direction and / or a column direction of the image feature corresponding to the fourth image includes:
[0074] Using the dilated convolution, feature correction is performed in the row direction and / or column direction of the image features corresponding to the fourth image.
[0075] Optionally, using a dilated convolution, feature correction is performed in a row direction and / or a column direction of image features corresponding to the fourth image, including:
[0076] determining an eleventh image feature and a twelfth image feature based on the image feature corresponding to the fourth image;
[0077] Compressing the eleventh image feature in the row direction, and performing at least one dilated convolution in the column direction on the compressed eleventh image feature to obtain an image feature corrected in the column direction;
[0078] Compressing the twelfth image feature in the column direction, and performing at least one dilated convolution in the row direction on the compressed twelfth image feature to obtain an image feature corrected in the row direction;
[0079] The features of each rectified image are fused.
[0080] According to another aspect of an embodiment of the present application, an electronic device is provided. The electronic device includes: a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to implement the method provided in the embodiment of the present application.
[0081] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method provided in the embodiments of the present application is implemented.
[0082] According to another aspect of an embodiment of the present application, a computer program product is provided. The computer program product includes a computer program. When the computer program is executed by a processor, the method provided in the embodiment of the present application is implemented.
[0083] The method, electronic device, storage medium and program product performed by the electronic device provided in the embodiments of the present application can use the first AI network to perform editing processing based on input information to obtain a first image and guidance information, where the guidance information includes at least one of spatial correlation guidance information and semantic correlation guidance information; based on the guidance information, use the second AI network to perform super-resolution processing on the first image to obtain a second image, that is, the embodiments of the present application save calculations in the second AI network by sharing the guidance information in the first AI network with the second AI network, thereby achieving a fast super-resolution processing process, effectively reducing the model size, calculation amount and time consumption, and improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0084] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in describing the embodiments of the present application.
[0085] Figure 1 A flowchart of a method executed by an electronic device provided in an embodiment of the present application;
[0086] Figure 2 A schematic diagram of a training method for a first AI network and / or a second AI network provided in an embodiment of the present application;
[0087] Figure 3 A schematic diagram of a calculation method of an attention module provided in an embodiment of the present application;
[0088] Figure 4 A schematic diagram of a word element provided in an embodiment of the present application;
[0089] Figure 5 A schematic diagram of an attention module provided in an embodiment of the present application;
[0090] Figure 6 A schematic diagram of a process of accumulating cross-attention weight maps in all stages of all steps provided in an embodiment of the present application;
[0091] Figure 7 A schematic diagram of extracting similar features in adjacent stages provided in an embodiment of the present application;
[0092] Figure 8 A schematic diagram of a self-attention weight sharing process provided in an embodiment of the present application;
[0093] Fig. 9 A schematic diagram of a cross-attention weight sharing process provided in an embodiment of the present application;
[0094] Fig.10 A schematic diagram of a correction module processing flow provided in an embodiment of the present application;
[0095] Fig.11 A schematic diagram of a cascade diffusion model processing solution provided in an embodiment of the present application;
[0096] Fig.12 A schematic diagram of another cascade diffusion model processing solution provided in an embodiment of the present application;
[0097] Fig.13 A schematic diagram of a high-resolution image generation process using a cascade diffusion model provided in an embodiment of the present application;
[0098] Fig.14 A schematic diagram of a usage scenario of the solution provided in an embodiment of the present application;
[0099] Fig.15 A schematic diagram of another usage scenario of the solution provided in the embodiment of the present application;
[0100] Fig.16 A schematic diagram of another method performed by an electronic device provided in an embodiment of the present application;
[0101] Fig.17 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0102] The following description with reference to the accompanying drawings is provided to facilitate a comprehensive understanding of the various embodiments of the present application as defined by the claims and their equivalents. This description includes various specific details to facilitate understanding but should be considered as exemplary only. Therefore, one of ordinary skill in the art will recognize that various changes and modifications can be made to the various embodiments described herein without departing from the scope and spirit of the present application. In addition, for the sake of clarity and conciseness, descriptions of well-known functions and structures may be omitted.
[0103] The terms and expressions used in the following specification and claims are not limited to their dictionary meanings, but are merely used by the inventor to enable a clear and consistent understanding of the present application. Therefore, it should be apparent to those skilled in the art that the following description of various embodiments of the present application is provided for illustration purposes only and not to limit the purpose of the present application as defined in the appended claims and their equivalents.
[0104] It should be understood that the singular forms "a", "an", and "the" may also include plural references unless the context clearly indicates otherwise. Thus, for example, reference to "a component surface" includes reference to one or more such surfaces. When we refer to an element as being "connected" or "coupled" to another element, the one element may be directly connected or coupled to the other element, or the one element and the other element may establish a connection relationship through an intermediate element. In addition, "connected" or "coupled" as used herein may include wireless connection or wireless coupling.
[0105] The term "include" or "may include" refers to the presence of corresponding disclosed functions, operations, or components that can be used in various embodiments of the present application, rather than limiting the presence of one or more additional functions, operations, or features. In addition, the term "include" or "have" may be interpreted as indicating certain characteristics, numbers, steps, operations, constituent elements, components, or combinations thereof, but should not be interpreted as excluding the possibility of the presence of one or more other characteristics, numbers, steps, operations, constituent elements, components, or combinations thereof.
[0106] The term "or" used in various embodiments of the present application includes any of the listed terms and all combinations thereof. For example, "A or B" may include A, may include B, or may include both A and B. When describing multiple (two or more) items, if the relationship between the multiple items is not clearly defined, the multiple items may refer to one, multiple, or all of the multiple items. For example, the description of "parameter A includes A1, A2, A3" may be implemented as parameter A including A1 or A2 or A3, and may also be implemented as parameter A including at least two of the three items A1, A2, and A3.
[0107] Unless defined differently, all terms (including technical terms or scientific terms) used in this application have the same meanings understood by those skilled in the art of this application. Common terms as defined in dictionaries are interpreted as having meanings consistent with the context in the relevant technical field, and should not be interpreted ideally or overly formally, unless clearly defined in this application.
[0108] At least some functions of the device or electronic device provided in the embodiments of the present application can be implemented by an AI model, such as at least one module among multiple modules of the device or electronic device can be implemented by an AI model. Functions associated with AI can be performed by non-volatile memory, volatile memory and processor.
[0109] The processor may include one or more processors. In this case, the one or more processors may be general-purpose processors, such as a central processing unit (CPU), an application processor (AP), etc., or pure graphics processing units, such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an AI-specific processor, such as a neural processing unit (NPU).
[0110] The one or more processors control the processing of input data according to predefined operating rules or artificial intelligence (AI) models stored in non-volatile memory and volatile memory. The predefined operating rules or artificial intelligence models are provided by training or learning.
[0111] Here, providing by learning means obtaining a predefined operating rule or an AI model with desired characteristics by applying a learning algorithm to a plurality of learning data. The learning can be performed in the device or electronic device itself in which the AI according to the embodiment is executed, and / or can be implemented by a separate server / system.
[0112] The AI model may include multiple neural network layers. Each layer has multiple weight values, and each layer performs neural network calculations by calculating between the input data of the layer (such as the calculation results of the previous layer and / or the input data of the AI model) and the multiple weight values of the current layer. Examples of neural networks include, but are not limited to, convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), generative adversarial networks (GANs), and deep Q networks.
[0113] A learning algorithm is a method of using a plurality of learning data to train a predetermined target device (e.g., a robot) to enable, allow, or control the target device to make a determination or prediction. Examples of the learning algorithm include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
[0114] The method provided in this application may involve one or more technical fields such as speech, language, image, video or data intelligence.
[0115] Optionally, when it comes to the field of images or videos, according to the present application, in a method performed in an electronic device, a method for image generation, image super-resolution processing and / or image correction can obtain output data after the target repair area in the image is repaired by using image data as input data of an artificial intelligence model. The artificial intelligence model can be obtained through training. Here, "obtained by training" means training a basic artificial intelligence model with multiple training data through a training algorithm to obtain a predefined operation rule or artificial intelligence model configured to perform the desired feature (or purpose). The method of the present application may relate to the field of visual understanding of artificial intelligence technology, which is a technology for identifying and processing things like human vision, and includes, for example, object recognition, object tracking, image retrieval, human recognition, scene recognition, 3D reconstruction / positioning, or image enhancement.
[0116] The following describes several optional embodiments to illustrate the technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application. It should be noted that the following embodiments can refer to, draw on or combine with each other, and the same terms, similar features and similar implementation steps in different embodiments will not be described repeatedly.
[0117] In an embodiment of the present application, a method executed by an electronic device is provided, such as Figure 1 As shown, the method includes:
[0118] Step S101: Based on input information, use a first AI network to perform editing processing to obtain a first image and guidance information, where the guidance information includes at least one of spatial correlation guidance information and semantic correlation guidance information;
[0119] In the embodiment of the present application, the input information is the basic information for realizing image editing. Optionally, the input information can be text, image, or text plus image, etc., or it can be voice, voice plus image, text plus voice plus image, etc., but not limited to this. That is, the main task of this step is to receive information such as text or text plus image input by the user as input, and realize image editing processing based on the input information. Optionally, the editing processing includes but is not limited to at least one of the following: image completion (inpainting) processing, image expansion (outpainting) processing, text-based image generation processing, image fusion processing, image style conversion processing, etc.
[0120] Optionally, the first AI network can adopt a generative model, such as a GAN (Generative Adversarial Networks) network. Alternatively, the first AI network can adopt a diffusion network model. For a given random noise image, the diffusion process of the diffusion model can remove noise step by step, and this denoising process is performed on the result of the previous denoising. Among them, if the input includes an image, the random noise image can be obtained by adding random noise to the image. If the input does not include an image, the random noise image can be randomly generated according to a predetermined algorithm. In actual applications, the degree of denoising each time can be achieved by setting different training methods according to actual conditions. For example, Figure 2 As shown, the first AI network can be trained in the following manner: obtain the original image for training, and continuously add noise to it until the original image is completely converted into random noise data. The image with noise added each time is retained and used to calculate the loss in the training phase to update the first AI network. For each training, the first AI network only needs to predict the noise that needs to be removed from the current input image. For example, a 70% denoised image is input into the first AI network. The first AI network only needs to predict the denoised image of the current image, and compare it with the image with 20% noise added to calculate the loss (Loss), and update the model parameters. The updated first AI network continues to be used for prediction. Similar prediction processes and model training processes can be deduced by analogy, which will not be repeated here. Alternatively, the type of the first AI network is not limited to this, and it can also be other neural network models. Those skilled in the art can set the type and training method of the first AI network according to actual conditions, and the embodiments of the present application are not limited here.
[0121] For the image editing task based on the first AI network, such as text generation image processing, it is possible to generate ultra-realistic pictures based on the text content input by the user, and these pictures are highly consistent with the meaning of the text. That is, the process is mainly based on the understanding of text and language models to generate images. With the continuous development of large language models and the significant improvement of their text understanding capabilities, the quality of the output first image can be effectively guaranteed.
[0122] In an embodiment of the present application, while using the first AI network to obtain the first image, guidance information in the first AI network can be obtained, wherein the guidance information includes at least one of spatial correlation guidance information and semantic correlation guidance information.
[0123] Optionally, the spatial correlation guidance information includes spatial correlation weights between different spatial positions, for example, may specifically include self-attention weights of the image, but is not limited thereto.
[0124] Optionally, the semantic relevance guidance information includes at least one of a semantic relevance weight between a spatial position and text constraint content, and a semantic relevance weight between different spatial positions, for example, specifically including a cross-attention weight that characterizes the relevance between the text constraint content and each pixel of the image, but is not limited thereto. The text constraint content may come from input information.
[0125] Considering the calculation of guidance information in the AI network, taking the calculation of attention weight (attention map) as an example, it will occupy most of the steps in the attention module, such as Figure 3 As shown in Figure 1, the calculation process of the attention module includes 5 steps, among which the first 4 steps are used to generate attention weights. Specifically, in the first step, the Q (Query) feature, K (Key) feature and V (Value) feature are obtained; in the second step, the K feature is transposed and matrix multiplied with the Q feature to generate a similarity matrix; in the third step, the similarity matrix is normalized, for example, each element is divided by d k is the dimension size of K; in the fourth step, the normalized result is subjected to Softmax processing to obtain the attention weight; in the fifth step, the attention weight is matrix multiplied with the V feature. In order to speed up the image processing process and facilitate deployment on mobile devices, the embodiments of the present application aim to simplify the calculation of spatial correlation guidance information and / or semantic correlation guidance information. Specifically, the repeated calculation of guidance information in the first AI network and the second AI network can be reduced, so the guidance information in the processing process of the first AI network is obtained and shared with the second AI network, thereby eliminating the calculation of guidance information in the second AI network.
[0126] Step S102: Based on the guidance information, use the second AI network to perform super-resolution processing on the first image to obtain a second image.
[0127] In an embodiment of the present application, the first image obtained in step S101 may be an image with a smaller resolution. In step S102, the second AI network receives the small-resolution image obtained by the first AI network, the input information description, and the guidance information in the first AI network as input, and uses this information to convert the small-resolution image into an image with a larger resolution (i.e., super-resolution processing, which may also be referred to as super-resolution processing), so that the generated second image has a higher resolution while maintaining the original details of the first image.
[0128] Optionally, the second AI network can adopt a super-resolution model (also referred to as a super-resolution model), relying on its image restoration and amplification capabilities to perform this step. Optionally, the second AI network can adopt the diffusion process of at least one diffusion network model to achieve high-resolution image generation. In one example, the generation of high-resolution and high-quality images is achieved by cascading two diffusion models. Alternatively, the type of the second AI network is not limited to this, and it can also be other neural network models. Those skilled in the art can set the type and training method of the second AI network according to actual conditions, and the embodiments of the present application are not limited here.
[0129] Through the above method provided by the embodiment of the present application, users can input information such as text, images, or text plus images to generate high-resolution images that are highly consistent with the input intention, which not only broadens the application field of AI image editing technology, but also greatly meets the user's demand for high-quality, high-resolution images. In addition, by sharing the guidance information in the first AI network with the second AI network, the calculation of the guidance information in the high-resolution space (i.e., in the second AI network) is eliminated, thereby achieving a fast super-resolution processing process, effectively reducing the model size, calculation amount and time consumption, and improving the user experience.
[0130] In the embodiment of the present application, an optional implementation is provided for step S101. Specifically, the first AI network includes at least one first spatial attention module, and the spatial correlation guidance information includes spatial correlation guidance information in at least one first spatial attention module. ;
[0131] The spatial correlation guidance information may be represented in the form of a self-attention weight map, but is not limited thereto.
[0132] For the embodiment of the present application, the first AI network may include at least one first spatial attention module, and the spatial attention module may specifically be a self-attention module (Self-attention block), but is not limited thereto. Optionally, the first spatial attention module may adopt the self-attention module of the Transformer model, or may adopt a spatial attention module of other structures, which is not limited in the embodiment of the present application. Among them, the spatial correlation guidance information obtained from multiple first spatial attention modules can be understood as spatial correlation guidance information at different stages, corresponding to different spatial sizes (which can also be understood as different scales).
[0133] Since the spatial correlation guidance information reflects the correlation between image blocks, and the high-resolution image and the low-resolution image have the same semantic features, but differ in spatial dimensions (i.e., length and width dimensions, excluding channel dimensions) and details, the correlation between their image blocks can be shared. Taking the spatial attention module as an example, in the embodiment of the present application, for the self-attention module, the self-attention weights of at least one stage of the relevant resolution generated by the first AI network are used as spatial attention guidance information in the second AI network, thereby achieving efficient self-attention in the second AI network, which greatly saves the amount of computation.
[0134] Optionally, the spatial correlation guidance information of at least one stage obtained (i.e., the spatial correlation guidance information of at least one first spatial attention module) can be stored using a spatial correlation guidance information cache pool (for example, a self-attention weight cache pool) for subsequent acquisition and use by a second AI network.
[0135] In the embodiment of the present application, the semantic relevance guidance information may include semantic relevance guidance information corresponding to at least one word.
[0136] Among them, the semantic relevance guidance information can be represented in the form of a cross-attention weight map, but is not limited to this.
[0137] For the embodiment of the present application, the first AI network may include at least one first semantic attention module, and the semantic attention module may specifically be a cross-attention block. Optionally, the first semantic attention module may adopt the cross-attention module of the Transformer model, or may adopt a semantic attention module of other structures, which is not limited in the embodiment of the present application.
[0138] In the embodiment of the present application, the semantic relevance guidance information can reflect the relationship between different modalities. For example, in the text-to-image generation task, the semantic relevance guidance information reflects the correlation between text and image. Specifically, the semantic relevance guidance information represents the correlation between word units and pixels. A word unit refers to the smallest grammatical unit in a language, such as a character, a word, or a word. As an example, Figure 4As shown, if the input text is "Cat with hatholding a sword in the hand", each word such as "hat", "cat", "sword" can be used as a word unit to obtain the semantic relevance weight (i.e., token-image attention) of each word with respect to any position in the entire image; as another example, if the input text is "Cat with hatholding a sword in the hand", each word such as "Cat", "hat", "sword" can be used as a word unit to obtain the semantic relevance weight of each word with respect to any position in the entire image.
[0139] Optionally, the text constraint content may be obtained based on at least one word-gram.
[0140] Since at least one word (such as text information or image content from user input) is a highly generalized semantic description of the image, it has a strong guiding significance for both image generation and image super-resolution. Therefore, in the embodiment of the present application, taking the semantic attention module using the cross attention module as an example, for the cross attention module, the cross attention weight generated by the first AI network integrates the global correlation between at least one word and each pixel of the image, and can be effectively combined into the subsequent second AI network for its use based on this global semantic guidance information, to achieve semantic guidance of at least one word for image super-resolution.
[0141] It is understandable that the above-mentioned sharing of spatial correlation guidance information and sharing of semantic correlation guidance information can be used separately or in combination. Figure 5 Taking an attention module as an example (the attention modules of other structures are similar), the attention module may include a connected residual (Residual Network, Resnet) module, a self-attention module and a cross-attention module. The first AI network and the second AI network both include multiple attention modules. Since the attention module has a greater impact on the calculation time, whether the self-attention weight in the self-attention module of the first AI network is shared with the second AI network, or the cross-attention weight in the cross-self-attention module of the first AI network is shared with the second AI network, or the self-attention weight in the self-attention module of the first AI network and the cross-attention weight in the cross-self-attention module are shared with the second AI network, the calculation amount of the second AI network can be reduced. In addition, the order of sharing spatial correlation guidance information and semantic correlation guidance information can be in no particular order, and those skilled in the art can set it according to actual conditions, and the embodiments of the present application are not limited here.
[0142] In the embodiment of the present application, the second AI network includes at least one second spatial attention module and / or at least one second semantic attention module, and step S102 may include at least one of the following:
[0143] Step S1021: Based on the spatial correlation guidance information, use at least one second spatial attention module to perform spatial attention processing on the input first image features respectively;
[0144] Step S1022: Based on the semantic relevance guidance information, use at least one second semantic attention module to perform semantic attention processing on the input second image features respectively.
[0145] In combination with the above introduction, it can be seen that the step numbers in step S1021 and step S1022 do not constitute a limitation on the order of the two steps, that is, step S1021 and step S1022 can only execute any one of them, or the execution order of step S1021 and step S1022 can be non-precedence, such as step S1021 can be executed first and then step S1022, or step S1022 can be executed first and then step S1021, or step S1021 and step S1022 can be executed at the same time, etc., and the embodiments of the present application are not limited here.
[0146] In an embodiment of the present application, the first AI network includes at least one first semantic attention module, and the semantic relevance guidance information includes semantic relevance guidance information corresponding to at least one word-gram in the at least one first semantic attention module. Step S1022 may specifically include: for each word-gram, fusing the semantic relevance guidance information corresponding to the word-gram in at least one first semantic attention module to obtain the semantic relevance guidance information corresponding to the word-gram; based on the semantic relevance guidance information corresponding to at least one word-gram, using at least one second semantic attention module to perform semantic attention processing on the input second image features respectively.
[0147] For the embodiments of the present application, the semantic relevance guidance information obtained from different first semantic attention modules can be understood as semantic relevance guidance information at different stages, corresponding to different spatial sizes (which can also be understood as different scales). Among them, each stage can include semantic relevance guidance information corresponding to at least one word. Since semantic features are features of a global nature, the semantic relevance guidance information of all stages in the image editing process of the first AI network is comprehensively considered to improve the robustness and globality of the obtained semantic relevance guidance information.
[0148] Therefore, in an embodiment of the present application, by superimposing the semantic relevance guidance information of each stage of the model in the process of editing the image by the first AI network, it is possible to establish global semantic relevance guidance information for each word corresponding to different positions of the image, that is, to obtain semantic relevance guidance information at the token-to-pixel level.
[0149] Specifically, the semantic relevance guidance information corresponding to the word-gram in at least one first semantic attention module is fused (each word-gram is processed in the same way), which can include: for each execution of the first AI network, the semantic relevance guidance information corresponding to the word-gram in at least one first semantic attention module is transformed into the same size as the first image and then superimposed to obtain first semantic relevance cumulative guidance information; the first semantic relevance cumulative guidance information obtained by each execution of the first AI network is superimposed to obtain second semantic relevance cumulative guidance information; and the second first semantic relevance cumulative guidance information is normalized.
[0150] In other words, in the process of editing the first image using the first AI network, the first AI network may be executed once or multiple times (step). In the embodiment of the present application, for each word, the semantic relevance guidance information of all steps and all stages in the image editing process of the first AI network is comprehensively considered to improve the robustness and globality of the cross-attention weights.
[0151] As an example, Figure 6 As shown, taking the semantic relevance guidance information as a cross-attention weight map as an example, assuming that the user's input information can obtain multiple word units such as Token A, Token B, Token C..., for each word unit, in each step of the image editing process of the first AI network, the cross-attention weight maps of all stages in the model are unified to the same size (T×H×W) as the original image (first image) for superposition, and the cross-attention cumulative weight map of each step is obtained (i.e., the first semantic relevance cumulative guidance information). Then, the cross-attention cumulative weight maps in all steps (all processes) are superimposed in the first AI network during multiple executions, and in order to keep the weight within the range of 0 to 1, the final cumulative result (i.e., the second semantic relevance cumulative guidance information) is divided by the number of steps*stages (i.e., normalized), thereby obtaining the final token-to-pixel level cross-attention weight maps corresponding to multiple word units such as Token A, Token B, Token C..., etc. Among them, the module that performs this process can be called a cross-attention weight fusion module, which can be included in the cross-attention module of the second AI network.
[0152] In the embodiment of the present application, an optional implementation is provided for step S102. Specifically, the second AI network includes at least one second spatial attention module and / or at least one second semantic attention module. Step S102 may include at least one of the following:
[0153] Step S102A: Based on the spatial correlation guidance information in at least one first spatial attention module, perform spatial attention processing on the first image features corresponding to the corresponding second spatial attention module respectively;
[0154] In the embodiment of the present application, the stage of obtaining the spatial correlation guidance information generated by the first AI network (i.e., which first spatial attention modules are shared with the spatial correlation guidance information) and the corresponding stage in the second AI network (i.e., which second spatial attention modules are shared with) are both configurable. For ease of description and understanding, the following is an example of the spatial attention module being a self-attention module.
[0155] As an example, assuming that both the first AI network and the second AI network include 5 self-attention modules, the first self-attention module of the first AI network generates the self-attention weight of the first stage, which corresponds to the first self-attention module in the second AI network. Therefore, the self-attention weight of the first stage can be fused with the first image feature of the first self-attention module in the second AI network, and the other self-attention modules are analogous. That is, the self-attention modules of the same stage of the first AI network and the second AI network can share the self-attention weight.
[0156] Alternatively, considering that adjacent stages can extract features of similar levels, the attention weights will also be similar, e.g. Figure 7 As shown, the self-attention weights can also be shared between the self-attention modules in adjacent stages of the first AI network and the second AI network.
[0157] As an example, assuming that the first AI network includes 7 self-attention modules and the second AI network includes 5 self-attention modules, the 5-stage self-attention weights generated by the 2nd to 6th self-attention modules of the first AI network are shared with the second AI network, and the 2nd self-attention module of the first AI network generates the self-attention weight of the second stage, which corresponds to the 1st self-attention module in the second AI network. Therefore, the self-attention weight of the second stage can be fused with the first image feature of the first self-attention module in the second AI network, and the same applies to other self-attention modules.
[0158] As described above, both the first AI network and the second AI network can be executed multiple times. For the embodiment of the present application, if the number of executions (steps) of the second AI network is the same as the number of executions of the first AI network, the self-attention weights of each step of the first AI network can be shared with the second AI network of the corresponding step, for example, the self-attention weights in the first execution of the first AI network are shared with the first execution of the second AI network, and the self-attention weights in the second execution of the first AI network are shared with the second execution of the second AI network, and so on.
[0159] Alternatively, the number of executions of the second AI network may be less than the number of executions of the first AI network. In this case, the self-attention weights of multiple steps of the first AI network can be averaged and shared with the second AI network of the corresponding number of steps. For example, if the first AI network iterates 50 times to edit and obtain the first image, and the second AI network iterates 10 times to super-resolution to obtain the second image, the self-attention weights obtained by executing the first AI network every 5 times can be averaged and shared with the second AI network for execution once. In practical applications, the number of executions and the corresponding relationship between the first AI network and the second AI network are set according to actual conditions, and the embodiments of the present application are not limited here.
[0160] In actual applications, the number of first spatial attention modules in the first AI network, the number of second spatial attention modules in the second AI network, the spatial correlation guidance information generated by the first AI network for sharing with the second AI network, and the correspondence between the stages and the number of steps can be set according to actual conditions, and the embodiments of the present application are not limited here.
[0161] Optionally, the correspondence between the stages is sequential, for example, the self-attention weights generated in the earlier stages of the first AI network will be shared with the earlier self-attention modules of the second AI network, and the self-attention weights generated in the later stages of the first AI network will be shared with the later self-attention modules of the second AI network. Therefore, in an embodiment of the present application, the self-attention weights generated in the first AI network can be cached in a self-attention weight cache pool, and their order can be recorded for subsequent acquisition and utilization by the second AI network.
[0162] Optionally, due to the flexibility of the stage correspondence, in order to cope with the situation where the shared self-attention weights and the first image features of the corresponding stages may have different resolutions, the self-attention weights generated by the first AI network can be scaled and adapted to the adaptation stage in the second AI network to match the resolution of the first image features of this stage, and then fused with them to adjust their weights on a spatial scale.
[0163] Step S102B: The semantic relevance guidance information corresponding to at least one word-gram is fused, and based on the fused semantic relevance guidance information, semantic attention processing is performed on the second image features corresponding to at least one second semantic attention module.
[0164] In an embodiment of the present application, semantic relevance guidance information corresponding to at least one word unit is fused to obtain fused semantic relevance guidance information, and then the fused semantic relevance guidance information is fused with features in the second AI network (at least one second image feature), thereby embedding the fused semantic relevance guidance information into the features of the second AI network.
[0165] It is understandable that step S102A and step S102B may be used separately or in combination. In addition, the execution order of step S102A and step S102B may be random, and those skilled in the art may set it according to actual conditions, which is not limited in the present embodiment.
[0166] In an embodiment of the present application, an optional implementation manner is provided for step S102A. Specifically, for the spatial correlation guidance information in each first spatial attention module, performing spatial attention processing on the first image feature corresponding to the corresponding second spatial attention module based on the spatial correlation guidance information may include:
[0167] Step S102A1: obtaining a fourth image feature and a fifth image feature based on the first image feature corresponding to the corresponding second spatial attention module;
[0168] In the embodiment of the present application, the first image feature is processed into two copies to perform the subsequent steps S102A2 and S102A3 in parallel to adjust the channel attention, thereby achieving high-quality self-attention in the second AI network while greatly saving the amount of calculation. In practical applications, those skilled in the art can choose a suitable way to process the first image feature into two copies according to actual conditions, and the embodiment of the present application is not limited here.
[0169] As an example, the first image feature of the corresponding stage in the second AI network can be split into two. For example, the first image feature of the corresponding stage in the second AI network can be divided into two in the channel dimension to obtain the fourth image feature and the fifth image feature, thereby reducing the amount of parallel calculation without substantially affecting the calculation effect. Alternatively, the first image feature of the corresponding stage in the second AI network can be copied to obtain the fourth image feature and the fifth image feature (the two image features are the same), thereby making the calculation data more complete.
[0170] Step S102A2: fusing the spatial correlation guidance information with the fourth image feature to obtain a sixth image feature;
[0171] This step can be understood as a spatial attention branch, which uses the spatial correlation guidance information generated by the first spatial attention module corresponding to this stage in the first AI network to perform matrix multiplication with the input fourth image feature to obtain the sixth image feature adjusted by the global spatial attention.
[0172] Step S102A3: performing channel attention adjustment on the fifth image feature to obtain a seventh image feature;
[0173] This step can be understood as a channel attention branch. Among them, different channels in the image features can be understood as information of different frequencies. According to the characteristics of the super-resolution processing task of the second AI network, when enlarging the image, the semantic information of the original low-resolution image (first image) is retained, and more details and texture information are generated, which is manifested in the image as the generation of high-frequency information, that is, the super-resolution processing task pays great attention to high-frequency information. Therefore, through the channel attention mechanism provided in the spatial attention module, different channels can be weighted differently to better help the network learn high-frequency information.
[0174] In an embodiment of the present application, the channel attention weight of the fifth image feature may be preset, or the channel attention weight of the fifth image feature may be calculated based on the fifth image feature. Specifically, the fifth image feature may be compressed in the spatial dimension to obtain the eighth image feature, for example, the fifth image feature is subjected to a pooling operation in the spatial dimension to obtain an eighth image feature whose spatial dimension is 1 and whose number of channels is consistent with that of the fifth image feature; based on the eighth image feature, the channel attention weight of the fifth image feature is obtained, for example, a convolution operation (such as conv 1×1, but not limited thereto) is performed on the eighth image feature to obtain the channel attention weight of the fifth image feature, wherein the feature obtained by the convolution operation is used to characterize the importance of each channel in the fifth image feature. Furthermore, the feature obtained by the convolution operation may be subjected to nonlinear mapping to obtain the channel attention weight of the fifth image feature.
[0175] Furthermore, channel attention adjustment is performed on the fifth image feature based on the channel attention weight. Specifically, the channel attention weight can be multiplied by the original fifth image feature to achieve weighted operation on different channels of the fifth image feature, thereby obtaining the seventh image feature with adjusted channel attention weight.
[0176] Step S102A4: Fusing the sixth image feature and the seventh image feature.
[0177] The way of fusing the sixth image feature and the seventh image feature corresponds to the way of processing the first image feature into two.
[0178] As an example, if the fourth image feature and the fifth image feature are obtained by dividing the first image feature into two parts in the channel dimension, the sixth image feature and the seventh image feature can be spliced in the channel dimension in the original order, so as to ensure that the splicing result is consistent with the first image feature in dimension. Alternatively, if the fourth image feature and the fifth image feature are obtained by copying a copy of the first image feature, the sixth image feature and the seventh image feature can be summed, averaged or weighted summed. Those skilled in the art can select a suitable fusion method according to actual conditions, and the embodiments of the present application are not limited here.
[0179] In other embodiments, the spatial attention branch and the channel attention branch may also be executed in series. For example, the spatial attention branch may be used first, and the output of the spatial attention branch may be processed by the channel attention branch, or the channel attention branch may be used first, and the output of the channel attention branch may be processed by the spatial attention branch, etc. Those skilled in the art may expand the method according to actual conditions, and the embodiments of the present application are not limited thereto.
[0180] Based on at least one of the above embodiments, taking the spatial correlation guidance information as the self-attention weight as an example, Figure 8 As shown, a process example of self-attention weight sharing provided by an embodiment of the present application is given. Specifically, in a self-attention module of an attention module (including a connected residual module, a self-attention module, and a cross-attention module) of a first AI network, the V feature, Q feature, and K feature of the input feature are obtained. The matrix multiplication result of the Q feature and the K feature can obtain a self-attention weight map, and the self-attention weight map and the V feature are multiplied to obtain the output result of the self-attention module. Among them, the self-attention weight map is also saved in the self-attention weight cache pool for the second AI network to obtain and use.
[0181] In the self-attention module of an attention module of the second AI network (including the connected residual module, self-attention module and cross-attention module), the input feature is divided into two parts in the channel dimension, namely, feature F 1 (i.e., the fourth image feature) and feature F 2 (ie the fifth image feature).
[0182] The feature F 1 Send it to the spatial attention branch, obtain the self-attention weight map generated by the self-attention module in the first AI network corresponding to this stage in the self-attention weight cache pool, adjust the scale of this self-attention weight map to be consistent with the input feature in the spatial dimension, and perform feature F 1 Perform matrix multiplication with the scale-adjusted self-attention weight map to obtain the feature F adjusted by the global spatial attention 3(i.e., the sixth image feature). The feature adjustment method may be direct resizing (adjusting the size) or dimension adjustment by convolution.
[0183] The feature F 2 Send it to the channel attention branch, first the feature F 2 Perform pooling operation on the spatial dimension, and get 1 on the spatial dimension, and the channel and F 2 Consistent features, after 1×1 convolution and nonlinear layer, the attention weight on the channel dimension is obtained, and the channel attention weight is combined with the original feature F 2 Multiply them to achieve weighted operations on different channels, thereby obtaining the feature F with adjusted channel attention weights 4 (ie the seventh image feature).
[0184] F 3 and F 4 The output features are generated by splicing in the original order on the channel dimension, thereby ensuring that the output features are consistent with the input features in dimension.
[0185] Among them, the spatial attention branch and the channel attention branch can form a self-attention fusion module. This module can be included in the self-attention module.
[0186] In an embodiment of the present application, the self-attention fusion module can use the self-attention weight map shared by the first AI network to implement global spatial attention adjustment on the features in the second AI network, draw on the features of the associated areas according to their correlation, and assign higher channel attention weights to high-frequency feature channels through the feature attention branch parallel to the spatial attention, thereby saving a lot of computation while generating high-quality high-resolution images.
[0187] In the embodiment of the present application, for step S102B, the semantic relevance guidance information corresponding to at least one word-unit is integrated, which may specifically include:
[0188] Step S102B1: Obtain a weight corresponding to at least one word unit;
[0189] Among them, the dimension of the obtained weight (for the convenience of distinction, it can be called token aggregation weight) is consistent with the number of tokens (words).
[0190] Optionally, the token aggregation weight may be a default weight, which may be set by a person skilled in the art according to actual conditions, and is not limited in the embodiments of the present application.
[0191] Alternatively, the token aggregation weight can also be user-defined. Specifically, at least one word-gram can be displayed by an electronic device; in response to the user's first selection operation on a target word-gram in the at least one word-gram, the weight corresponding to the at least one word-gram is determined. Among them, the display form includes but is not limited to text, voice, etc., and the selection method includes but is not limited to clicking, speaking, etc. That is, in an embodiment of the present application, the user can select the word-gram for which he wants to generate more details based on at least one word-gram corresponding to the original input information. It can be understood that the word-gram selected by the user will be assigned more weight. In actual applications, the distribution ratio of the word-gram selected by the user and the word-gram not selected can be set according to the actual situation, and the embodiment of the present application is not limited here.
[0192] Step S102B2: Based on the weight corresponding to at least one word-gram, weighted fusion is performed on the semantic relevance guidance information corresponding to at least one word-gram;
[0193] For example, taking the case where the semantic relevance guidance information is the cross-attention weight, the cross-attention weight corresponding to at least one word unit and the weight corresponding to at least one word unit can be weighted by matrix multiplication to obtain a fused cross-attention weight.
[0194] In an embodiment of the present application, for step S102B, based on the fused semantic relevance guidance information, semantic attention processing is performed on the second image features corresponding to at least one second semantic attention module, which can specifically include: fusing the fused semantic relevance guidance information with the third image features of at least one scale corresponding to the first image to obtain a global semantic feature of at least one scale; based on the global semantic feature of at least one scale, semantic attention processing is performed on the second image features corresponding to the second semantic attention modules of the corresponding scales.
[0195] Among them, the third image feature refers to the image feature obtained by extracting features from the first image. Optionally, the first image can be sent to a feature extraction network (such as a lightweight convolutional network) including multiple feature extraction modules to extract the third image feature in the first image. In the feature extraction network, the third image features of different scales (which can also be understood as different stages, that is, the outputs of different feature extraction modules, corresponding to different spatial sizes) are fused with the semantic relevance guidance information corresponding to at least one word element to obtain a global semantic feature of at least one scale. For example, taking the example that the semantic relevance guidance information is the cross-attention weight, for the third image features of each stage corresponding to the first image, the third image features are matrix multiplied with the scale-adjusted fused cross-attention weights to obtain the global semantic feature of the stage.
[0196] Furthermore, semantic attention processing is performed on the global semantic features of each scale and the features in the second AI network (the second image features of the corresponding scale), so that the global semantic features are embedded in the features of the second AI network.
[0197] Similarly, the number of first semantic attention modules in the first AI network, the number of second semantic attention modules in the second AI network, the semantic relevance guidance information generated by the first AI network for sharing with the second AI network, and the second semantic attention module in the second AI network that needs to embed global semantic information can be set according to actual conditions, and the embodiments of the present application are not limited here.
[0198] Step S102B3: Fusing the fused semantic relevance guidance information with at least one scale of the third image feature corresponding to the first image respectively.
[0199] Based on at least one of the above embodiments, taking the semantic relevance guidance information as a cross attention weight as an example, Fig. 9 As shown, a process example of cross-attention weight sharing provided by an embodiment of the present application is given. Specifically, after obtaining the token-to-pixel level cross-attention weight maps corresponding to multiple word units, they are weighted by matrix multiplication through a token aggregation weight to obtain a fused cross-attention weight map. The first image is sent to a feature extraction network (such as a convolutional conv network) to extract image features in the first image. The features of different stages (different spatial sizes) in the feature extraction network are matrix multiplied with the scale-adjusted fused cross-attention weight map to obtain the global semantic features of the corresponding stage. The global semantic features of each stage are fused with the features of the same feature space size in the second AI network (i.e., the third image features), so as to embed the global semantic information into the features of the second AI network. Among them, the module that executes this process can be called a cross-attention fusion module, which can be included in the cross-attention module.
[0200] In an embodiment of the present application, since the second AI network and the first AI network have the same semantic meaning, the cross-attention weights can be reused in the second AI network. Among them, by accumulating all the cross-attention weights shared by the first AI network, the cross-attention weights from the word unit to the pixel level can be obtained. The cross-attention fusion module can use the cross-attention weights to embed the shared semantic information into the second AI network through an ultra-lightweight encoder to ensure the consistency between the semantic information of the high-resolution image (second image) and the semantic information of the original low-resolution image (first image). After adjusting the features in the second AI network, the guidance of the super-resolution task by global semantic information can be achieved, so that the second AI network can generate natural and realistic texture details.
[0201] In the embodiment of the present application, the second AI network also includes an injection correction module, and step S102 may also include: using the correction module to perform feature correction in the row direction and / or column direction of the image features corresponding to the first image.
[0202] The first AI network edits the image based on the guidance of input information, where the input information may be a high-level semantic summary that can guide the first AI network to generate semantically correct content, but may not contain rich visual texture information. Therefore, the first AI network may generate details in the detail texture that do not fully match the natural image.
[0203] Among them, for human eyes, regular textures are most obvious if they are misaligned. These details are not obvious in low-resolution images, but due to the super-resolution processing, the misaligned details will be magnified, affecting the user's visual perception.
[0204] Since the human eye is particularly sensitive to the alignment of regular features in the row and column directions, in order to avoid the situation in which the not-so-good detail texture in the low-resolution generated image (i.e., the first image) is magnified during the super-resolution processing, making the originally unobvious misalignment unacceptable to the user, in an embodiment of the present application, the correction module automatically corrects the misaligned texture features by calculating the relationship between the features in the row direction (horizontal direction) and / or column direction (vertical direction), which can significantly improve the visual effect of the large-resolution image.
[0205] Among them, the correction module can be connected before the attention fusion module, and the image features corresponding to the first image at this time refer to the features obtained by directly extracting features from the first image; or, the correction module can also be connected after the attention fusion module, and the image features corresponding to the first image at this time can refer to the features obtained after a certain super-resolution processing is performed on the first image. Alternatively, the correction module can be provided before and after the attention fusion module, and the image features corresponding to the first image include the above two situations, which can be processed separately.
[0206] Specifically, a dilated convolution may be used to perform feature correction in the row direction and / or column direction of the image features corresponding to the first image.
[0207] Among them, compared with ordinary convolution that can only use a small receptive field to obtain the relationship between adjacent square areas, the correction module can achieve strong correlation correction of features in the row direction and / or strong correlation correction of features in the column direction by using hollow convolution with a large receptive field in the row direction and / or column direction respectively, thereby achieving the effect of correcting the misaligned texture generated by the first AI network in the second AI network to repair the misalignment in the sensitive structure of the human eye.
[0208] Optionally, a dilated convolution of at least two dilated indices is cascaded to perform feature correction in a row direction and / or a column direction of image features corresponding to the first image.
[0209] Optionally, at least two dilated convolutions of dilated indices may be connected in series.
[0210] Optionally, the number of atrous convolutions in series can be adjusted according to the feature dimension.
[0211] Optionally, the dilation index corresponding to each dilated convolution can be adjusted according to actual conditions.
[0212] Optionally, the dilation index of the serially connected dilated convolutions can be increasing.
[0213] Optionally, using a dilated convolution, feature correction is performed in the row direction and the column direction of the image feature corresponding to the first image, which may specifically include:
[0214] Step S301: determining a ninth image feature and a tenth image feature based on the image feature corresponding to the first image;
[0215] In the embodiment of the present application, the image features corresponding to the first image are processed into two to achieve feature correction in the row direction and column direction respectively. In practical applications, those skilled in the art can choose a suitable way to process the image features corresponding to the first image into two according to actual conditions, and the embodiment of the present application is not limited here.
[0216] As an example, the image feature corresponding to the first image can be split into two parts, for example, the image feature corresponding to the first image can be divided into two parts in the channel dimension, and the ninth image feature and the tenth image feature can be obtained, so as to reduce the amount of calculation without substantially affecting the calculation effect. Alternatively, the image feature corresponding to the first image can be copied, and the ninth image feature and the tenth image feature (the two image features are the same) can also be obtained, so that the calculation data is more complete.
[0217] Step S302: compressing the ninth image feature in the row direction, and performing at least one dilated convolution in the column direction on the compressed ninth image feature to obtain an image feature corrected in the column direction;
[0218] Optionally, the ninth image feature is globally pooled in the row direction to compress it to obtain a compressed ninth image feature of a dimension of columns*number of channels.
[0219] Furthermore, the compressed ninth image feature is subjected to at least one dilated convolution in the column direction, for example, feature extraction is performed by a dilated convolution in which multiple dilated indices are connected in series, so as to achieve global perception in the column direction and obtain the image feature after correction in the column direction.
[0220] The ninth image feature after the dilated convolution in the column direction can be copied in the row dimension, and the expanded dimension is the same as the original ninth image feature. Further, the ninth image feature after the expanded dimension is matrix multiplied with the original ninth image feature to obtain the image feature after column direction correction.
[0221] Step S303: compressing the tenth image feature in the column direction, and performing at least one dilated convolution in the row direction on the compressed tenth image feature to obtain an image feature corrected in the row direction;
[0222] Optionally, the tenth image feature is globally pooled in the row direction to compress it to obtain a compressed tenth image feature of a row*channel number dimension.
[0223] Furthermore, the compressed tenth image feature is subjected to at least one dilated convolution in the row direction, for example, feature extraction is performed by a dilated convolution of multiple different dilated indices connected in series, so as to achieve global perception in the row direction and obtain the image feature after correction in the row direction.
[0224] The tenth image feature after the dilated convolution in the row direction can be copied in the column dimension, and the expanded dimension is the same as the original tenth image feature. Further, the tenth image feature after the expanded dimension is matrix multiplied with the original tenth image feature to obtain the image feature after correction in the row direction.
[0225] Step S304: Fusing the corrected image features.
[0226] The method of fusing the image features after feature correction and alignment corresponds to the method of processing the image features corresponding to the first image into two.
[0227] As an example, if the ninth image feature and the tenth image feature are obtained by dividing the image feature corresponding to the first image into two parts in the channel dimension, the image features after feature correction and alignment can be spliced in the channel dimension. Alternatively, if the ninth image feature and the tenth image feature are obtained by copying the image feature corresponding to the first image, the image features after feature correction and alignment can be summed, averaged, or weighted summed. Those skilled in the art can select a suitable fusion method according to actual conditions, and the embodiments of the present application are not limited here.
[0228] In other embodiments, the row alignment correction and the column alignment correction may also be performed in series. For example, the row texture features may be corrected first, and then the column texture features of the row alignment result may be corrected; or the column texture features may be corrected first, and then the row texture features of the column alignment result may be corrected, etc. Those skilled in the art may expand the correction according to actual conditions, and the embodiments of the present application are not limited here.
[0229] Based on at least one of the above embodiments, Fig.10 As shown, an example of a processing flow of the correction module provided in an embodiment of the present application is given, which may specifically include:
[0230] Step S10.1: Divide the input features into two equal parts in the channel dimension, one part is sent to the row alignment branch, and the other part is sent to the column alignment branch.
[0231] Step S10.2: On the column alignment branch, firstly, globally pool the input features of the column branch in the row direction to obtain features of the number of columns (H)*channels (C). Then, the features after global pooling are extracted by multiple dilated convolutions connected in series with different dilated indices to achieve global perception in the column direction. Fig.10 As shown in FIG, five dilated convolutions with gradually increasing dilation indices are connected in series, which can cover a maximum range of 13 intervals.
[0232] Step S10.3: The features of the column alignment branch after multiple dilated convolutions are copied in the row dimension, and the expanded dimension is the same as the column branch input feature.
[0233] Step S10.4: Perform matrix multiplication on the expanded dimension features and the column branch input features to obtain features that have been corrected and aligned in the column direction.
[0234] Step S10.5: On the row alignment branch, the operation steps are basically the same as step 10.3, except that global pooling is performed on the row branch input features in the column direction to obtain features of the row (W) * channel (C) number dimension.
[0235] Step S10.6: The features of the row alignment branch after multiple dilated convolutions are copied in the column direction, and the expanded dimension is the same as the row branch input feature.
[0236] Step S10.7: Perform matrix multiplication on the expanded dimension features and the row branch input features to obtain features that are corrected and aligned in the row direction.
[0237] Step S10.8: Concatenate the output features of the two branches in the channel dimension to obtain the final result.
[0238] In the embodiment of the present application, correction is performed in the row direction and / or column direction in the correction module, and correction of misalignment in the row and column directions is automatically achieved through interaction at the feature level, thereby improving the super-resolution effect of regular textures. Under the premise that the human eye is most sensitive to regular textures, the visual perception of the human eye is enhanced, thereby improving the user experience.
[0239] In the embodiment of the present application, a VAE (Variational Auto-Encoder) network can also be used to compress the image into a latent variable space, which is also called a latent space. The image editing process and / or the image super-resolution process are performed in the latent space, which can significantly reduce the dimension of the model, thereby making the image editing process and / or the image super-resolution process faster.
[0240] Based on at least one of the above embodiments, Fig.11 As shown, a high-performance and high-efficiency cascade diffusion model processing solution provided by an embodiment of the present application is provided for generating high-resolution images. Specifically, the solution may include:
[0241] Step S11.1: The user edits the original image, selects the area to be removed (such as "tree crown"), and enters text to guide the model to generate corresponding content.
[0242] Step S11.2: After compressing the image (512×512) to the latent variable space (64×64, configurable) through the VAE network encoder, the text feature representation obtained by encoding the text through the text encoder is combined to guide the generative diffusion model (i.e., the first AI network) to generate a low-resolution image (i.e., the first image), and the low-resolution image is decoded back to the original image form (512×512) through the VAE network decoder.
[0243] Step S11.3: Obtain the self-attention weight map and cross-attention weight map in the generative diffusion model.
[0244] Step S11.4: After resizing the low-resolution image, it is compressed to the latent variable space (256×256) through the VAE network encoder. The self-attention fusion module and cross-attention fusion module in the super-resolution diffusion model (i.e., the second AI model) produce a high-resolution image based on the generated low-resolution image and attention map, and correct it in the row and column directions respectively through the correction module in the super-resolution diffusion model.
[0245] Step S11.5: Decode the high-resolution image back to the original image form (2048×2048) through the VAE network decoder, and output the final high-resolution image (i.e., the second image).
[0246] The embodiment of the present application can save the calculation of high-resolution space and repair the texture dislocation caused by the generated diffusion model by generating an attention sharing mechanism between the diffusion model and the super-resolution diffusion model.
[0247] In addition, due to the different functions of the self-attention module and the cross-attention module, the corresponding self-attention weights and cross-attention weights are shared differently.
[0248] Combined with the introduction of the self-attention fusion module above, we can see that the global spatial self-attention map of adjacent calculations can be used in the super-resolution diffusion model to avoid the huge amount of calculation in the high-resolution space, while also obtaining spatial self-attention. The spatial attention adjustment of features is achieved through scale adjustment. In addition, channel attention adjustment can be achieved through channel weighted branches based on the characteristics of super-resolution tasks, which can better help the network learn high-frequency information, enhance super-resolution performance, and further save calculations.
[0249] Combined with the introduction of the cross-attention fusion module above, it can be seen that the cross-attention weight map accumulated in the generative diffusion model can embed word-unit guidance into super-resolution to ensure semantic consistency and enhance the capabilities of the super-resolution diffusion model.
[0250] Combined with the introduction of the correction module above, it can be seen that the local correction module is based on the large receptive field module of columns and rows, which is used to extract horizontal and vertical relative features, so as to correct the generated texture misalignment, especially for regular textures that humans are sensitive to, such as horizontal or vertical textures, to achieve high-quality super-resolution processing.
[0251] In practical applications, the connection between the self-attention fusion module, the cross-attention fusion module and the correction module in the super-resolution diffusion model is not limited to Fig.11 The structure shown in the figure can also be connected in other ways, such as Fig.12As shown, as long as the super-resolution diffusion model includes the three modules (each of which can be one or more), the above functions can be realized. In addition, these three modules can also be used alone or in combination, and can achieve corresponding functions. Those skilled in the art can use them in combination according to actual conditions, and the embodiments of this application are not limited here.
[0252] Based on at least one of the above embodiments, Fig.13 As shown, a network structure for realizing high-resolution image generation by using a cascade diffusion model provided in an embodiment of the present application is given, wherein the entire structure is based on a U-net (a network containing jump connections), including an encoder, an intermediate part, and a decoder. The encoder extracts features by shrinking the feature size, the intermediate part is a further operation for enriching the features, the decoder expands the spatial size to the image size, and combines the features of the encoder. A residual module (Resnet) block is used to extract features at each stage. The local correction module is located in the shallow layer because it is very lightweight and can repair texture details. The self-attention fusion module and the cross-semantic fusion module are located after the residual module at each stage in turn.
[0253] Specifically, Fig.13 In the figure, the model above is a generative diffusion model, which aims to generate low-resolution images based on input text, images, or text and images. In the image generation process:
[0254] The corresponding stages obtained from the respective attention modules of the generative diffusion model ( Fig.13 The self-attention weight maps of stages 1, 2, 3, 6, and 7 are obtained as examples. They will be cached in the self-attention weight cache pool for subsequent super-resolution diffusion models.
[0255] The corresponding stages obtained from each criss-cross attention module of the generative diffusion model ( Fig.13 The cross-attention weight maps of stages 1 to 7 are obtained as an example, and the cross-attention weight fusion module is used to accumulate all cross-attention weight maps for subsequent super-resolution diffusion model use.
[0256] In addition, the model below is a super-resolution diffusion model, which aims to obtain a high-resolution image based on the super-resolution of a low-resolution image. In the image super-resolution process:
[0257] The self-attention fusion module in the super-resolution diffusion model uses the self-attention map in the self-attention weight cache pool and adapts it to the corresponding stage in the super-resolution model after rescaling ( Fig.13In the example of the second, third, fourth, fifth and sixth stages, the self-attention weight map obtained in the generative diffusion model is an adjacent stage), and the weights are adjusted on a spatial scale by fusing with the features output by the residual module. At the same time, according to the characteristics of the super-resolution task, a dual-branch structure is used to execute the channel attention mechanism in parallel, automatically weighting different channels differently, realizing channel attention adjustment, and better helping the network learn high-frequency information.
[0258] The cross-attention fusion module in the super-resolution diffusion model rescales all the cross-attention accumulated weight maps in the image generation process, establishes the cross-attention weight map for each word corresponding to different positions in the image, and embeds it into the features output by the self-attention module, so as to realize the guidance of semantic information for the features in the super-resolution diffusion model.
[0259] The correction module in the super-resolution diffusion model uses cascaded dilated convolutions in the row and column directions respectively to achieve strong correlation correction of features in the row direction and strong correlation correction of features in the column direction, thereby correcting the misaligned textures generated by the generated diffusion model in the super-resolution diffusion model.
[0260] in, Fig.13 For details not covered in detail, please refer to Fig.11 The description of the above embodiments will not be repeated here.
[0261] It should be noted that in the cascade diffusion model, the number and connection order of the self-attention module, the cross-attention module, and the correction module are all optional. For example, only one correction module can be used in the decoder layer; for another example, the residual unit, the self-attention module, and the cross-attention module are a group of modules, and the decoder and the encoder can use a group of modules respectively; for another example, the self-attention module can be connected after the cross-attention module, etc. Those skilled in the art can select and connect the modules according to the actual situation, and the embodiments of the present application are not limited here.
[0262] Through the combination of these modules, while ensuring the quality of super-resolution generated images, the computational complexity of the attention module based on the diffusion model is reduced, the operating efficiency is optimized, and the misaligned detail information is corrected, achieving fast and high-quality high-resolution image generation.
[0263] In the embodiment of the present application, before step S101, the following steps may be further included:
[0264] Step S001: Responding to the user's input operation, obtaining the target content text;
[0265] In the embodiment of the present application, users can input the content they want to generate an image, and the input method includes but is not limited to text, image, voice, etc. In response to the user's input operation, the target content text corresponding to the user's input content can be obtained. If the input is voice content, the target content text after the voice content is converted into text can be obtained. If the input is an image, the target content text obtained by analyzing the image semantics can be obtained.
[0266] Step S002: Based on the target content text, determine the corresponding recommended constraint content, and display the recommended constraint content;
[0267] In an embodiment of the present application, recommended constraint content determined based on the target content text may be provided to the user, wherein the recommended constraint content is a descriptive prompt for the image, which helps to improve image generation performance.
[0268] Optionally, the recommended constraint content may include positive constraint content, i.e., positive descriptive prompts, such as "high quality", "delicate clothing", "real style", etc., and / or negative constraint content, i.e., negative descriptive prompts, such as "low quality", "oil painting", "blurred", etc.
[0269] Optionally, the text constraint content may be obtained based on the recommendation constraint content and / or word-grams.
[0270] Step S003: In response to the user's second selection operation on the target constraint content in the recommended constraint content, the target constraint content and the target content text are determined as input information.
[0271] In the embodiment of the present application, the user can select one or more of the recommended constraint contents displayed, wherein for the positive descriptive prompts selected by the user, the corresponding features will be enhanced in the subsequent image generation process and image super-resolution process. For the negative descriptive prompts selected by the user, the corresponding features will be subtracted in the subsequent image generation process and image super-resolution process, so as to achieve the purpose of personalized image quality improvement.
[0272] In the embodiment of the present application, after obtaining the first image based on the input information using the first AI network in step S101, the following may be included:
[0273] Step: 401: displaying the first image;
[0274] Step S402A: receiving a super-resolution processing instruction fed back by a user;
[0275] Then step S102 may specifically include: in response to the super-resolution processing instruction, using the second AI network to perform super-resolution processing on the first image based on at least one of the spatial correlation guidance information and the semantic correlation guidance information;
[0276] Furthermore, the method may also include:
[0277] Step S402B: receiving a re-acquisition instruction of the first image from user feedback; in response to the re-acquisition instruction of the first image, re-executing the operation of obtaining the first image based on the input information using the first AI network and displaying the first image again.
[0278] In one example, the user may be asked whether he is satisfied with the first image. If the user chooses to be satisfied, a super-resolution processing instruction is fed back. If the user chooses to be dissatisfied, a re-acquisition instruction of the first image is fed back. In other embodiments, other methods may be used to generate these two instructions, such as displaying "Whether to continue" or directly receiving instructions input by the user, etc., which are not limited in the embodiments of the present application.
[0279] In an embodiment of the present application, if the user is not satisfied with the generated first image, the first AI network can be reused to generate the first image again until the user is satisfied, and the attention weight of the first AI network processing corresponding to the first image that satisfies the user is obtained, and then shared with the second AI network for use, so as to improve the efficiency of generating high-resolution images.
[0280] Based on at least one of the above embodiments, Fig.14 As shown, a usage scenario of the solution provided by the embodiment of the present application is given, and an image can be generated using text, specifically, it can include:
[0281] Step 14.1( Fig.14 Indicated as serial number ①, the other steps are similar and will not be repeated): Get the target content text entered by the user, such as "the cat in the hat holds a sword in his hand".
[0282] Step 14.2: Based on the target content text, determine the corresponding recommended constraint content, including positive descriptive text, such as "high quality", "delicate clothing", "real style", etc., and negative descriptive text, such as "low quality", "oil painting", "blur", etc. The recommended constraint content is displayed to the user; the user selects the provided constraint content; the device obtains the constraint content selected by the user from the displayed recommended constraint content. The constraint content helps to improve the image generation performance.
[0283] Step 14.3: Generate an image based on the constraint content and target content text selected by the user, ask the user whether he is satisfied with the generated result, and proceed to the next step after the user is satisfied with the generated low-resolution image.
[0284] Step 14.4: The user can select words for which he wants to generate more details based on the original input text, such as selecting the words "cat" and "sword" in the text "The cat in the hat holds a sword in his hand".
[0285] Step 14.5: Perform super-resolution processing on the image based on the important words selected by the user, and output the generated high-quality and high-resolution image.
[0286] This allows high-quality, large-resolution images to be quickly generated based on the user's text input.
[0287] Based on at least one of the above embodiments, Fig.15 As shown, another use scenario of the solution provided in the embodiment of the present application is given, which can use images and texts to generate new images, and can be applied to scenarios such as image editing, image expansion, and image restoration. Specifically, it can include:
[0288] Step 15.1( Fig.15 Indicated as serial number ①, other steps are similar and will not be repeated): obtain the user's original image and the target editing area drawn by the user on the original image, and obtain the mask of the target editing area.
[0289] Step 15.2: Ask the user whether there is any desired generated content. If yes, obtain the text content input by the user, such as "cherry blossom tree"; if the user chooses that there is no desired generated content, obtain the text content that the user does not want to generate; if the user chooses that there is no desired generated content (that is, the user has not input), obtain empty text, and randomly generate images according to the background. That is, this step is to determine whether the user has input text content, and obtain it if yes.
[0290] Step 15.3: Based on the target content text (if the user has input text, it can be determined based on the text content; if the user has not input text, it can be determined based on the image content), determine the corresponding recommended constraint content, including positive descriptive text, such as "high quality", "delicate clothing", "real style", etc., and negative descriptive text, such as "low quality", "oil painting", "blur", etc. The recommended constraint content is displayed to the user; the user selects the provided constraint content; and the constraint content selected by the user in the displayed recommended constraint content is obtained. The constraint content helps to improve the image generation performance.
[0291] Step 15.4: Generate a new image based on the constraints selected by the user and the image and text content entered by the user, and ask the user whether he is satisfied with the generated result. If the user is satisfied with the generated low-resolution image, proceed to the next step.
[0292] Step 15.5: The user can select the words for which he wants to generate more details based on the original input text, such as selecting the word "cherry blossom tree" in the text "a cherry blossom tree". If the user does not enter any text, this selection step can be skipped and the subsequent super-resolution processing uses the default aggregation weight.
[0293] Step 15.6: Perform super-resolution processing on the image based on the important words selected by the user or the default aggregation weights, and output the generated high-quality and high-resolution image accordingly.
[0294] After a large number of experiments by the inventors of the present application, it was found that the technical solution provided in the embodiments of the present application can better generate more and more natural detail information and has higher super-resolution quality.
[0295] Second, model size vs. time and comparison with existing models:
[0296]
[0297] It can be seen that the model provided by the embodiment of the present application is greatly reduced in both model size and inference time compared to the model that does not share attention weights, thus meeting the user's needs for high-quality image generation.
[0298] In addition, after adding the correction module, the model provided in the embodiment of the present application can correct the texture of misaligned rows and columns on the super-resolution image, for example, it can make the horizontal lines more straight, greatly improving the visual effect.
[0299] The present application also provides a method executed by an electronic device, such as Fig.16 As shown, the method includes:
[0300] Step S1601: Acquire a third image;
[0301] Among them, the third image can be an image obtained by the user in any way, such as real-time shooting, downloading, or reading from the local, etc.; or the third image can also be an image output by the model during the user's editing process of the image, such as a new image produced based on text and / or images input by the user, etc., and the embodiments of the present application are not limited here.
[0302] Step S1602: performing super-resolution processing on the third image to obtain a fourth image;
[0303] In the embodiment of the present application, super-resolution processing is performed on the third image with a smaller resolution to convert it into an image with a larger resolution, so that the generated fourth image has a higher resolution while maintaining the original details of the third image. Among them, those skilled in the art can select the super-resolution processing method used in this step according to actual conditions, for example, the super-resolution method provided by at least one of the above embodiments can be used, or other super-resolution methods can be used.
[0304] Step S1603: performing feature correction in the row direction and / or column direction of the image features corresponding to the fourth image to obtain a fifth image.
[0305] Since the human eye is particularly sensitive to the alignment of regular features in the row and column directions, in order to avoid the situation in which the not-so-good detail texture in the low-resolution generated image (i.e., the third image) is magnified during the super-resolution processing, making the originally unobvious misalignment unacceptable to the user, in an embodiment of the present application, the misaligned texture features are automatically corrected by calculating the relationship between the features in the row direction (horizontal direction) and / or the column direction (vertical direction), which can significantly improve the visual effect of the large-resolution image.
[0306] Specifically, a dilated convolution may be used to perform feature correction in the row direction and / or column direction of the image features corresponding to the fourth image.
[0307] Among them, compared with ordinary convolution that can only use a small receptive field to obtain the relationship between adjacent square areas, the embodiment of the present application uses a hole convolution with a large receptive field in the row direction and / or the column direction respectively, which can achieve strong correlation correction of features in the row direction and / or strong correlation correction of features in the column direction, thereby achieving the effect of correcting the uneven textures produced by super-resolution processing to repair the misalignment in the sensitive structure of the human eye.
[0308] Optionally, a dilated convolution of at least two dilated indices is cascaded to perform feature correction in a row direction and / or a column direction of image features corresponding to the fourth image.
[0309] Optionally, at least two dilated convolutions of dilated indices may be connected in series.
[0310] Optionally, the number of atrous convolutions in series can be adjusted according to the feature dimension.
[0311] Optionally, the dilation index corresponding to each dilated convolution can be adjusted according to actual conditions.
[0312] Optionally, the dilation index of the serially connected dilated convolutions can be increasing.
[0313] Optionally, using a dilated convolution, feature correction is performed in the row direction and the column direction of the image feature corresponding to the fourth image, which may specifically include:
[0314] Step S16031: determining an eleventh image feature and a twelfth image feature based on the image feature corresponding to the fourth image;
[0315] In the embodiment of the present application, the image features corresponding to the fourth image are processed into two to achieve feature correction in the row direction and the column direction respectively. In practical applications, those skilled in the art can choose a suitable way to process the image features corresponding to the fourth image into two according to actual conditions, and the embodiment of the present application is not limited here.
[0316] As an example, the image feature corresponding to the fourth image can be split into two parts, for example, the image feature corresponding to the fourth image can be divided into two parts in the channel dimension, and the eleventh image feature and the twelfth image feature can be obtained, so as to reduce the amount of calculation without substantially affecting the calculation effect. Alternatively, the image feature corresponding to the fourth image can be copied, and the eleventh image feature and the twelfth image feature (the two image features are the same) can also be obtained, so that the calculation data is more complete.
[0317] Step S16032: compressing the eleventh image feature in the row direction, and performing at least one dilated convolution in the column direction on the compressed eleventh image feature to obtain an image feature corrected in the column direction;
[0318] Optionally, the eleventh image feature is globally pooled in the row direction to compress it, so as to obtain a compressed eleventh image feature of a dimension of columns*number of channels.
[0319] Furthermore, the compressed eleventh image feature is subjected to at least one dilated convolution in the column direction, for example, feature extraction is performed by connecting a plurality of dilated convolutions with different dilated indices in series, so as to achieve global perception in the column direction and obtain the image feature after correction in the column direction.
[0320] The eleventh image feature after the dilated convolution in the column direction can be copied in the row dimension, and the expanded dimension is the same as the original eleventh image feature. Further, the eleventh image feature after the expanded dimension is matrix multiplied with the original eleventh image feature to obtain the image feature after column direction correction.
[0321] Step S16033: compressing the twelfth image feature in the column direction, and performing at least one dilated convolution in the row direction on the compressed twelfth image feature to obtain an image feature corrected in the row direction;
[0322] Optionally, the twelfth image feature is globally pooled in the row direction to compress it to obtain a compressed twelfth image feature of a row*channel number dimension.
[0323] Furthermore, the compressed twelfth image feature is subjected to at least one dilated convolution in the row direction, for example, feature extraction is performed by a plurality of dilated convolutions with different dilated indices connected in series, so as to achieve global perception in the row direction and obtain the image feature after correction in the row direction.
[0324] The twelfth image feature after the dilated convolution in the row direction can be copied in the column dimension, and the expanded dimension is the same as the original twelfth image feature. Further, the twelfth image feature after the expanded dimension is matrix multiplied with the original twelfth image feature to obtain the image feature after correction in the row direction.
[0325] Step S16034: Fuse the corrected image features.
[0326] Among them, the way of fusing the image features after feature correction and alignment corresponds to the way of processing the image features corresponding to the fourth image into two.
[0327] As an example, if the eleventh image feature and the twelfth image feature are obtained by dividing the image feature corresponding to the fourth image into two parts in the channel dimension, the image features after feature correction and alignment can be spliced in the channel dimension. Alternatively, if the eleventh image feature and the twelfth image feature are obtained by copying the image feature corresponding to the fourth image, the image features after feature correction and alignment can be summed, averaged, or weighted summed. Those skilled in the art can select a suitable fusion method according to actual conditions, and the embodiments of the present application are not limited here.
[0328] In other embodiments, the row alignment correction and the column alignment correction may also be performed in series. For example, the row texture features may be corrected first, and then the column texture features of the row alignment result may be corrected; or the column texture features may be corrected first, and then the row texture features of the column alignment result may be corrected, etc. Those skilled in the art may expand the correction according to actual conditions, and the embodiments of the present application are not limited here.
[0329] In the embodiment of the present application, a large receptive field module based on columns and rows is used to correct the row and column directions respectively on the image after super-resolution processing using any super-resolution method. Through the interaction at the feature level, the correction of misalignment in the row and column directions is automatically achieved, and the texture that is not aligned with the rows and columns is corrected. For example, horizontal lines can be generated to be straighter, and the super-resolution effect of regular textures can be improved. Under the premise that the human eye is most sensitive to regular textures, the visual effect is improved, thereby improving the user experience.
[0330] The technical solution provided in the embodiments of the present application can be applied to various electronic devices, including but not limited to mobile terminals, smart terminals, etc., such as smart phones, tablet computers, laptops, smart wearable devices (such as watches, glasses, etc.), smart speakers, vehicle terminals, personal digital assistants, portable multimedia players, navigation devices, etc., but not limited to these. It can be understood by those skilled in the art that, in addition to components specifically used for mobile purposes, the structure according to the embodiments of the present application can also be applied to fixed-type terminals, such as digital televisions, desktop computers, etc.
[0331] The technical solution provided in the embodiments of the present application can also be applied to image generation and super-resolution in a server, such as an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), as well as big data and artificial intelligence platforms.
[0332] Specifically, the technical solution provided in the embodiments of the present application can be applied to image AI editing applications on various electronic devices to improve the speed and performance of generating high-resolution images, thereby producing fascinating image generation results, allowing users to unleash their imagination and create more exquisite images from input text, images, etc.
[0333] An embodiment of the present application also provides an electronic device, which includes a processor and, optionally, may also include a transceiver and / or a memory coupled to the processor, and the processor is configured to execute the steps of the method provided in any optional embodiment of the present application.
[0334] Fig.17 A schematic diagram of the structure of an electronic device applicable to an embodiment of the present invention is shown in FIG. Fig.17 As shown, Fig.17 The electronic device 4000 shown includes: a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, such as through a bus 4002. Optionally, the electronic device 4000 may also include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as data transmission and / or data reception. It should be noted that in actual applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present application. Optionally, the electronic device may be a first network node, a second network node, or a third network node.
[0335] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It may implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of this application. Processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0336] The bus 4002 may include a path to transmit information between the above components. The bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.17 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0337] The memory 4003 may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compressed optical disk, laser disk, optical disk, digital versatile disk, Blu-ray disk, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, without limitation herein.
[0338] The memory 4003 is used to store the computer program for executing the embodiment of the present application, and the execution is controlled by the processor 4001. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the steps shown in the above method embodiment.
[0339] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps and corresponding contents of the aforementioned method embodiment can be implemented.
[0340] The embodiment of the present application also provides a computer program product, including a computer program, which can implement the steps and corresponding contents of the aforementioned method embodiment when executed by a processor.
[0341] The terms "first", "second", "third", "fourth", "1", "2", etc. (if any) in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than that shown or described in the drawings.
[0342] It should be understood that, although each operation step is indicated by arrows in the flowchart of the embodiment of the present application, the implementation order of these steps is not limited to the order indicated by the arrows. Unless clearly stated herein, in some implementation scenarios of the embodiment of the present application, the implementation steps in each flowchart can be performed in other orders according to demand. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on actual implementation scenarios. Some or all of these sub-steps or stages may be executed at the same time, and each sub-step or stage in these sub-steps or stages may also be executed at different times respectively. In different scenarios of execution time, the execution order of these sub-steps or stages may be flexibly configured according to demand, and the embodiment of the present application does not limit this.
[0343] The above text and drawings are provided only as examples to help readers understand the present application. They are not intended to and should not be interpreted as limiting the scope of the present application in any way. Although certain embodiments and examples have been provided, based on the contents disclosed herein, it is obvious to those skilled in the art that the embodiments and examples shown can be changed without departing from the scope of the present application, and other similar implementation means based on the technical ideas of the present application are adopted, which also fall within the protection scope of the embodiments of the present application.
Claims
1. A method performed by an electronic device, It is characterized in that include: Based on the input information, use a first artificial intelligence (AI) network to perform editing processing to obtain a first image and guidance information, wherein the guidance information includes at least one of spatial correlation guidance information and semantic correlation guidance information; Based on the guidance information, a second AI network is used to perform super-resolution processing on the first image to obtain a second image.
2. The method according to claim 1, It is characterized in that The spatial correlation guidance information includes spatial correlation weights between different spatial positions; The semantic relevance guidance information includes at least one of a semantic relevance weight between a spatial position and text constraint content, and a semantic relevance weight between different spatial positions.
3. The method according to claim 1, It is characterized in that The first AI network includes at least one first spatial attention module, and the spatial correlation guidance information includes spatial correlation guidance information in at least one first spatial attention module; and / or, The semantic relevance guidance information includes semantic relevance guidance information corresponding to at least one word unit.
4. The method according to any one of claims 1 to 3, It is characterized in that The second AI network includes at least one second spatial attention module and / or at least one second semantic attention module, and the super-resolution processing of the first image using the second AI network based on the guidance information includes at least one of the following: Based on the spatial correlation guidance information, using the at least one second spatial attention module, respectively perform spatial attention processing on the input first image features; Based on the semantic relevance guidance information, the at least one second semantic attention module is used to perform semantic attention processing on the input second image features respectively.
5. The method according to claim 4, It is characterized in that The first AI network includes at least one first semantic attention module, the semantic relevance guidance information includes semantic relevance guidance information corresponding to at least one word in the at least one first semantic attention module, and based on the semantic relevance guidance information, using the at least one second semantic attention module to perform semantic attention processing on the input second image features respectively, including: For each word-unit, the semantic relevance guidance information corresponding to the word-unit in at least one first semantic attention module is fused to obtain the semantic relevance guidance information corresponding to the word-unit; Based on the semantic relevance guidance information corresponding to at least one word-unit, the at least one second semantic attention module is used to perform semantic attention processing on the input second image features respectively.
6. The method according to claim 5, It is characterized in that The fusing of the semantic relevance guidance information corresponding to the word-unit in at least one first semantic attention module includes: For each execution of the first AI network, the semantic relevance guidance information corresponding to the word in at least one first semantic attention module is transformed into the same size as the first image and then superimposed to obtain first semantic relevance accumulation guidance information; Superimposing the first semantic relevance accumulated guidance information obtained by executing the first AI network each time to obtain second semantic relevance accumulated guidance information; The second first semantic relevance accumulated guidance information is normalized.
7. The method according to claim 3, It is characterized in that The second AI network includes at least one second spatial attention module and / or at least one second semantic attention module, and the super-resolution processing of the first image using the second AI network based on the guidance information includes at least one of the following: Based on the spatial correlation guidance information in at least one first spatial attention module, respectively perform spatial attention processing on the first image features corresponding to the corresponding second spatial attention modules; The semantic relevance guidance information corresponding to at least one word unit is fused, and based on the fused semantic relevance guidance information, semantic attention processing is performed on the second image features corresponding to at least one second semantic attention module.
8. The method according to claim 7, It is characterized in that For each of the spatial correlation guidance information in the first spatial attention module, based on the spatial correlation guidance information, performing spatial attention processing on the first image feature corresponding to the corresponding second spatial attention module, including: Based on the first image feature corresponding to the corresponding second spatial attention module, obtain a fourth image feature and a fifth image feature; Based on the spatial correlation guidance information, performing spatial attention processing on the fourth image feature to obtain a sixth image feature; Performing channel attention adjustment on the fifth image feature to obtain a seventh image feature; The sixth image feature and the seventh image feature are fused.
9. The method according to claim 8, It is characterized in that The performing channel attention adjustment on the fifth image feature comprises: Compressing the fifth image feature in a spatial dimension to obtain an eighth image feature; Based on the eighth image feature, obtaining a channel attention weight of the fifth image feature; Channel attention is adjusted on the fifth image feature based on the channel attention weight.
10. The method according to claim 7, It is characterized in that The fusing of the semantic relevance guidance information corresponding to at least one word-element includes: Obtain a weight corresponding to at least one word unit; Based on the weight corresponding to the at least one word-gram, weighted fusion is performed on the semantic relevance guidance information corresponding to the at least one word-gram.
11. The method according to claim 10, It is characterized in that The obtaining of a weight corresponding to at least one word element includes: displaying the at least one word-gram; In response to a first selection operation of a user on a target word-gram among the at least one word-gram, a weight corresponding to the at least one word-gram is determined.
12. The method according to claim 7, It is characterized in that The step of performing semantic attention processing on the second image features corresponding to at least one second semantic attention module based on the fused semantic relevance guidance information includes: fusing the fused semantic relevance guidance information with a third image feature of at least one scale corresponding to the first image to obtain a global semantic feature of at least one scale; Based on the global semantic features of the at least one scale, semantic attention processing is performed on the second image features corresponding to the second semantic attention modules of the corresponding scales.
13. The method according to any one of claims 1 to 12, It is characterized in that The second AI network further includes a correction module, and the method further includes: Using the correction module, feature correction is performed in the row direction and / or column direction of the image features corresponding to the first image.
14. The method according to claim 13, It is characterized in that The performing feature correction in the row direction and / or column direction of the image feature corresponding to the first image includes: Using dilated convolution, feature correction is performed in the row direction and / or column direction of the image features corresponding to the first image.
15. The method according to claim 14, It is characterized in that The using of the dilated convolution to perform feature correction in the row direction and the column direction of the image feature corresponding to the first image includes: determining a ninth image feature and a tenth image feature based on the image feature corresponding to the first image; Compressing the ninth image feature in the row direction, and performing at least one dilated convolution in the column direction on the compressed ninth image feature to obtain an image feature corrected in the column direction; Compressing the tenth image feature in the column direction, and performing at least one dilated convolution in the row direction on the compressed tenth image feature to obtain an image feature corrected in the row direction; The features of each rectified image are fused.
16. The method according to any one of claims 1 to 15, It is characterized in that The editing process includes at least one of the following: Image completion processing, image expansion processing, text-based image generation processing, image fusion processing, and image style conversion processing.
17. A method performed by an electronic device, It is characterized in that include: acquiring a third image; Performing super-resolution processing on the third image to obtain a fourth image; Feature correction is performed in the row direction and / or column direction of the image features corresponding to the fourth image to obtain a fifth image.
18. An electronic device comprising a memory, a processor and a computer program stored in the memory, It is characterized in that The processor executes the computer program to implement the method according to any one of claims 1 to 17.
19. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 17 is implemented.
20. A computer program product comprising a computer program, It is characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 17 is implemented.