Model training method, image super-resolution processing method and related equipment

Through a model training method, combining coding features and semantic features for fusion attention processing, predicted images are generated and image super-resolution models are trained, which solves the problem that the prior art cannot effectively restore the degraded serious image details, and realizes high-quality image super-resolution processing.

CN119942264APending Publication Date: 2025-05-06BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411997092.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing image super-resolution methods based on diffusion models cannot effectively restore image details that are degraded but semantically important.

Method used

Through a model training method, the sample image and its corresponding high-resolution image are obtained, semantic extraction is performed to obtain high-level and low-level semantic features, combine coding features for fusion attention processing, generate predicted images, and train the model based on high-resolution images to obtain an image super-resolution model.

Benefits of technology

This method can capture semantic information in the image more accurately, effectively restore semantic important details, and generate high-resolution images that are highly similar to the input image and rich in details, improving the quality of image super-resolution processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942264A_ABST
    Figure CN119942264A_ABST
Patent Text Reader

Abstract

The invention provides a model training method, an image super-resolution processing method and related equipment, and relates to the technical field of image processing. The model training method comprises the following steps: acquiring a sample image and a corresponding high-resolution image; performing semantic extraction on the sample image to obtain high-level semantic features and low-level semantic features; performing coding processing on the sample image through the initial model to obtain coding features; performing fusion attention processing on the coding features, the high-level semantic features and the low-level semantic features to obtain output image features; performing decoding processing according to the coding features and the output image features to obtain a prediction image of the sample image; and training the initial model based on the high-resolution image and the prediction image to obtain an image super-resolution model. The image super-resolution model obtained by the method can effectively restore important semantic details of the image super-resolution model, and a high-resolution image which is highly similar to an input image and rich in details is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technology, and in particular to a model training method, an image super-resolution processing method, a model training device, an image super-resolution processing device, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0002] Image super-resolution refers to restoring a high-resolution image from a low-resolution image. Since images inevitably suffer from various forms of degradation during acquisition and transmission, such as low resolution, blur, noise, etc., and given the prevalence of these degradation phenomena in actual situations, image super-resolution has become an important task in the field of image processing.

[0003] In the related art, an image super-resolution method based on a diffusion model is used to perform image super-resolution reconstruction on an image. However, due to the randomness of the diffusion model, it is impossible to restore severely degraded but semantically important image details.

[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention

[0005] The present disclosure provides a model training method, an image super-resolution processing method, a model training device, an image super-resolution processing device, an electronic device, a computer-readable storage medium and a computer program product to overcome the above-mentioned problems or at least partially solve the above-mentioned problems.

[0006] According to one aspect of an embodiment of the present disclosure, a model training method is provided, the method comprising: obtaining a sample image and a high-resolution image corresponding to the sample image; performing semantic extraction on the sample image to obtain high-level semantic features and low-level semantic features of the sample image; encoding the sample image through an initial model to obtain encoding features of the sample image; performing fusion attention processing on the encoding features, the high-level semantic features and the low-level semantic features to obtain output image features of the sample image; performing decoding processing based on the encoding features and the output image features to obtain a predicted image of the sample image; training the initial model based on the high-resolution image and the predicted image to obtain an image super-resolution model.

[0007] In some embodiments of the present disclosure, the initial model includes an image decoder; the decoding processing based on the encoding features and the output image features to obtain the predicted image of the sample image includes: decoding the output image features through the image decoder to obtain the decoding features of the sample image; fusing the encoding features with the decoding features to obtain the predicted image.

[0008] In some embodiments of the present disclosure, the image decoder includes a feature adaptation module, which includes a cross-attention module and a convolution layer; the fusing processing of the encoding feature and the decoding feature to obtain the predicted image includes: fusing the encoding feature and the decoding feature through the cross-attention module to obtain a fused feature; generating a convolution parameter corresponding to the encoding feature through the convolution layer, and convolving the fused feature based on the convolution parameter to obtain the predicted image.

[0009] In some embodiments of the present disclosure, the initial model includes a control network module and a backbone network module, the control network module includes a semantic fusion attention module, and the backbone network module includes a conditional attention module; the fusion attention processing of the encoding features, the high-level semantic features and the low-level semantic features to obtain the output image features of the sample image includes: performing fusion attention processing on the encoding features, the high-level semantic features and the low-level semantic features through the semantic fusion attention module to obtain the fusion attention features of the sample image; performing conditional attention processing on the fusion attention features, the high-level semantic features and noise data through the conditional attention module to obtain the output image features.

[0010] In some embodiments of the present disclosure, the semantic fusion attention module includes a first attention module, a second attention module and a third attention module; the semantic fusion attention module performs fusion attention processing on the coding features, the high-level semantic features and the low-level semantic features to obtain the fusion attention features of the sample image, including: performing cross-attention processing on the coding features and the high-level semantic features through the first attention module to obtain the first cross-attention features of the sample image; performing cross-attention processing on the coding features and the low-level semantic features through the second attention module to obtain the second cross-attention features of the sample image; performing fusion attention processing on the first cross-attention features and the second cross-attention features through the third attention module to obtain the fusion attention features.

[0011] In some embodiments of the present disclosure, the semantic extraction of the sample image to obtain high-level semantic features and low-level semantic features of the sample image includes: performing semantic extraction on the sample image based on a multimodal large language model to obtain high-level semantic hint information and low-level semantic hint information of the sample image; encoding the high-level semantic hint information and the low-level semantic hint information respectively based on a text encoder to obtain the high-level semantic features and the low-level semantic features.

[0012] In some embodiments of the present disclosure, the initial model is trained based on the high-resolution image and the predicted image to obtain an image super-resolution model, including: constructing a model loss function based on the high-resolution image and the predicted image; adjusting the parameters of the image decoder based on the model loss function to obtain an image decoder with adjusted parameters, so as to obtain the image super-resolution model.

[0013] According to another aspect of an embodiment of the present disclosure, there is provided an image super-resolution processing method, the method comprising: acquiring an image to be processed; performing semantic extraction on the image to be processed to obtain high-level semantic features and low-level semantic features of the image to be processed; inputting the image to be processed and the high-level semantic features and low-level semantic features of the image to be processed into an image super-resolution model, and outputting a super-resolution image corresponding to the image to be processed; the image super-resolution model is trained according to the above-mentioned model training method.

[0014] According to another aspect of an embodiment of the present disclosure, a model training device is provided, the device comprising: a sample acquisition module, configured to acquire a sample image and a high-resolution image corresponding to the sample image; a sample semantic extraction module, configured to perform semantic extraction on the sample image to obtain high-level semantic features and low-level semantic features of the sample image; a sample training module, configured to encode the sample image through an initial model to obtain encoding features of the sample image; perform feature extraction on the encoding features, the high-level semantic features and the low-level semantic features to obtain output image features of the sample image; perform decoding based on the encoding features and the output image features to obtain a predicted image of the sample image; and train the initial model based on the high-resolution image and the predicted image to obtain an image super-resolution model.

[0015] According to another aspect of an embodiment of the present disclosure, there is provided an image super-resolution processing device, the device comprising: an image acquisition module, configured to acquire an image to be processed; an image semantic extraction module, configured to perform semantic extraction on the image to be processed, and obtain high-level semantic features and low-level semantic features of the image to be processed; an image processing module, configured to input the image to be processed, the high-level semantic features and the low-level semantic features of the image to be processed into an image super-resolution model, and output a super-resolution image corresponding to the image to be processed; the image super-resolution model is trained according to the above-mentioned model training method.

[0016] According to another aspect of an embodiment of the present disclosure, there is provided an electronic device, comprising: a processor; and a memory for storing processor executable instructions; wherein the processor is configured to execute the executable instructions to implement the above-mentioned model training method, or to implement the above-mentioned image super-resolution processing method.

[0017] According to another aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the above-mentioned model training method, or execute the above-mentioned image super-resolution processing method.

[0018] According to another aspect of an embodiment of the present disclosure, a computer program product is provided, including a computer program, which implements the above-mentioned model training method or the above-mentioned image super-resolution processing method when executed by a processor.

[0019] The model training method provided by the embodiment of the present disclosure has the following advantages: on the one hand, the coding feature is the key information extracted from the sample image. The coding feature is used to guide the decoding process during the model training process, which can more accurately capture the semantic information in the sample image, so that the image super-resolution model can effectively restore the important semantic details of the severely degraded input image, and then generate a high-resolution image that is highly similar to the input image and rich in details; on the other hand, by fusing the coding features, high-level semantic features and low-level semantic features, it is possible to deeply explore the multi-level information in the sample image, enhance the model's sensitivity to image details and semantics, so that the image super-resolution model can better retain and restore the important details and features in the input image when processing the image super-resolution task, and further improve the quality of the generated image.

[0020] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification are used to explain the principles of the present disclosure. Obviously, the accompanying drawings described below are only some embodiments of the present disclosure, and for ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without creative work.

[0022] Figure 1 A schematic diagram showing an exemplary system architecture to which the model training method or image super-resolution processing method of the embodiments of the present disclosure can be applied;

[0023] Figure 2 A flow chart of a model training method according to an embodiment of the present disclosure is shown;

[0024] Figure 3 A model framework diagram of an image super-resolution model according to an embodiment of the present disclosure is shown;

[0025] Figure 4 A module structure diagram of a feature adaptation module of an embodiment of the present disclosure is shown;

[0026] Figure 5 A schematic diagram showing a method of obtaining semantic features based on a multimodal large language model according to an embodiment of the present disclosure is shown;

[0027] Figure 6 A model framework diagram of another image super-resolution model according to an embodiment of the present disclosure is shown;

[0028] Figure 7 A flowchart of an image super-resolution processing method according to an embodiment of the present disclosure is shown;

[0029] Figure 8 A block diagram of a model training device according to an embodiment of the present disclosure is shown;

[0030] Fig. 9 A block diagram of an image super-resolution processing device according to an embodiment of the present disclosure is shown;

[0031] Fig.10 A schematic structural diagram of an electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0032] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be comprehensive and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The same reference numerals in the figures represent the same or similar parts, and thus their repeated description will be omitted.

[0033] The features, structures or characteristics described in the present disclosure may be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure may be practiced while omitting one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, known methods, devices, implementations or operations are not shown or described in detail to avoid blurring the various aspects of the present disclosure.

[0034] The collection, collection, updating, analysis, processing, use, transmission, storage and other aspects of the user personal information involved in this disclosure are in compliance with the relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. Necessary measures are taken for user personal information to prevent illegal access to user personal information data and maintain the security of user personal information, network security and national security.

[0035] The accompanying drawings are only schematic diagrams of the present disclosure, and the same reference numerals in the drawings represent the same or similar parts, so their repeated description will be omitted. Some block diagrams shown in the accompanying drawings do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in at least one hardware module or integrated circuit, or implemented in different networks and / or processor devices and / or microcontroller devices.

[0036] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and steps, nor must they be executed in the order described. For example, some steps can be decomposed, and some steps can be combined or partially combined, so the actual execution order may change according to actual conditions.

[0037] In this specification, the terms "a", "an", "the", "said" and "at least one" are used to indicate the presence of at least one element / component / etc.; the term "plurality" refers to two or more; the terms "comprising", "including" and "having" are used to express an open-ended inclusion and mean that additional elements / components / etc. may exist in addition to the listed elements / components / etc.; the terms "first", "second" and "third" etc. are used only as labels and are not intended to limit the quantity of their objects.

[0038] Figure 1 A schematic diagram of an exemplary system architecture to which the model training method or image super-resolution processing method of an embodiment of the present disclosure can be applied is shown.

[0039] like Figure 1As shown, the system architecture may include a server 101, a network 102 and a terminal device 103. The network 102 is used to provide a medium for a communication link between the terminal device 103 and the server 101. The network 102 may include various connection types, such as wired, wireless communication links or optical fiber cables.

[0040] In an exemplary embodiment, the terminal device 103 for data transmission with the server 101 may include but is not limited to mobile devices such as smart phones, tablet computers, and laptop computers, as well as terminal devices with specific functions or forms such as smart speakers, digital assistants, AR (Augmented Reality) devices, VR (Virtual Reality) devices, and smart wearable devices. Alternatively, the terminal device 103 may also be a personal computer, such as a laptop computer and a desktop computer. Optionally, the operating system running on the electronic device may include but is not limited to Android, IOS, Linux, Windows, etc.

[0041] Server 101 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. In some practical applications, server 101 can also be a server of a network platform, which can be, for example, a trading platform, a live broadcast platform, a social platform, or an audio platform, etc., which is not limited in the embodiments of the present disclosure. Among them, the server can be a single server or a cluster formed by multiple servers, and the present disclosure does not limit the specific architecture of the server.

[0042] In some embodiments of the present disclosure, the process of model training by server 101 may be: obtaining a sample image and a high-resolution image corresponding to the sample image; performing semantic extraction on the sample image to obtain high-level semantic features and low-level semantic features of the sample image; encoding the sample image through the initial model to obtain the encoding features of the sample image; performing fusion attention processing on the encoding features, high-level semantic features and low-level semantic features to obtain output image features of the sample image; performing decoding processing based on the encoding features and the output image features to obtain a predicted image of the sample image; training the initial model based on the high-resolution image and the predicted image to obtain an image super-resolution model.

[0043] In some embodiments of the present disclosure, the process of image super-resolution processing by server 101 may be: obtaining an image to be processed; performing semantic extraction on the image to be processed to obtain high-level semantic features and low-level semantic features of the image to be processed; inputting the image to be processed and the high-level semantic features and low-level semantic features of the image to be processed into an image super-resolution model, and outputting a super-resolution image corresponding to the image to be processed; wherein the image super-resolution model is trained according to the model training method of the above-mentioned embodiment.

[0044] In addition, it should be noted that Figure 1 What is shown is merely one application environment of the model training method or image super-resolution processing method provided by the present disclosure. Figure 1 The number of terminal devices 103, networks 102 and servers 101 is merely illustrative, and any number of terminal devices, networks and servers may be provided according to actual needs.

[0045] Figure 2 A flow chart of a model training method according to an embodiment of the present disclosure is shown. Figure 2 The execution subject of the method provided in the embodiment may be any electronic device, such as Figure 1 The terminal device 101 in the embodiment is also as follows Figure 1 The terminal device 103 in Figure 1 The server 101 and the terminal device 103 in the embodiment jointly implement the model training method, but the present disclosure is not limited thereto. Figure 2 , the model training method provided in the embodiment of the present disclosure includes the following steps.

[0046] Step S210: Acquire a sample image and a high-resolution image corresponding to the sample image.

[0047] In the embodiment of the present disclosure, a training sample set is obtained, which includes a plurality of training sample pairs, each of which includes a sample image and a high-resolution image corresponding to the sample image. The sample image may be an image with a resolution lower than a preset resolution threshold, that is, the sample image may also be called a low-resolution image sample.

[0048] In the disclosed embodiments, datasets such as DIV2K (DIVerse 2Kresolution dataset), DIV8K (DIVerse8Kresolution dataset), Flickr2K (Flickr 2Kresolution dataset), Unsplash2K (a high-resolution image dataset collected from the Unsplash website), and FFHQ (Flickr-Faces-HQ dataset) can be used to synthesize training sample pairs for training using a degradation pipeline based on a model based on ESRGAN (Enhanced Super-Resolution Generative Adversarial Networks).

[0049] Step S220 , performing semantic extraction on the sample image to obtain high-level semantic features and low-level semantic features of the sample image.

[0050] In the embodiments of the present disclosure, high-level semantic features refer to abstract, global information closely related to the overall content, scene, objects and their interrelationships of an image. High-level semantic features reflect the deep meaning and context of an image. For example, high-level semantic features may include object recognition and classification features, scene features, etc.

[0051] Among them, the object recognition and classification features can be used to identify key objects in the image and classify them. For example, it is determined that there is a cat in the image based on the object recognition and classification features. The scene features can be used to determine the type of scene depicted in the image. For example, it is determined that the image scene is a beach based on the scene understanding features. Of course, the high-level semantic features can also include other features, such as spatial position features, which are used for the spatial position relationship between objects in the image, and the embodiments of the present disclosure are not limited to this.

[0052] In the disclosed embodiment, low-level semantic features focus on details and local information in the image, which are directly related to the pixel values ​​and basic visual elements of the image. Exemplarily, low-level semantic features may include color features, texture features, shape features, image quality features, etc.

[0053] Among them, color features are used to extract color information of different areas in the image, for example, a certain area in the image is red. Texture features are used to describe the texture patterns of different areas in the image, for example, a certain area in the image has a rough texture. Shape features are used to identify the shape and outline of objects in the image, for example, a certain object in the image is round. Image quality features are used to evaluate the overall quality of the image, for example, the image is blurred and noisy. Of course, low-level semantic features can also include other features, such as lighting features, which are used to describe the lighting conditions of the image, and the embodiments of the present disclosure are not limited to this.

[0054] In the disclosed embodiment, semantic features are extracted from the sample image to obtain high-level semantic features and low-level semantic features of the sample image. By combining these two features, the initial model can obtain more comprehensive image information, which helps the trained image super-resolution model to maintain the integrity of the overall structure and detail information of the image when generating high-resolution images.

[0055] Step S230, encode the sample image through the initial model to obtain the encoding features of the sample image; perform fusion attention processing on the encoding features, high-level semantic features and low-level semantic features to obtain the output image features of the sample image; perform decoding processing based on the encoding features and the output image features to obtain the predicted image of the sample image.

[0056] In the disclosed embodiment, the initial model is an image super-resolution model pre-constructed based on a diffusion model. The diffusion model is a generative model that realizes image generation by learning the process of gradually adding noise from a clear image until it becomes pure noise (forward process), and the process of gradually denoising from pure noise and restoring it to a clear image (reverse process or generative process).

[0057] Among them, fused attention processing is an advanced deep learning technology that combines the attention mechanism to optimize the feature fusion process. Among them, the attention mechanism is a computational model that simulates the human visual attention allocation process. Fusion attention processing dynamically assigns weights according to the importance of each feature through the attention mechanism. This weight reflects the relative importance of the feature in the current task. These weighted features will be fused together to form a new, more robust feature representation. In the disclosed embodiment, fused attention feature extraction can be a processing operation that uses the network structure in the initial model to perform feature extraction and feature fusion on multiple model input data.

[0058] In the disclosed embodiment, a pre-built initial model is obtained, a sample image, high-level semantic features and low-level semantic features of the sample image are used as model input data of the initial model, the sample image is encoded to obtain the encoded features of the sample image, and then feature extraction is performed on the encoded features, high-level semantic features and low-level semantic features based on the attention mechanism to obtain output image features, and then the output image features are decoded, and the decoding process is based on the encoded features to finally obtain a predicted image of the sample image.

[0059] Step S240: training the initial model based on the high-resolution image and the predicted image to obtain an image super-resolution model.

[0060] In the disclosed embodiment, after obtaining the predicted image, the initial model is trained using the high-resolution image and the predicted image, and after the training is completed, an image super-resolution model is obtained.

[0061] The model training method provided by the embodiment of the present disclosure has the following advantages: on the one hand, the coding feature is the key information extracted from the sample image. The coding feature is used to guide the decoding process during the model training process, which can more accurately capture the semantic information in the sample image, so that the image super-resolution model can effectively restore the important semantic details of the severely degraded input image, and then generate a high-resolution image that is highly similar to the input image and rich in details; on the other hand, by fusing the coding features, high-level semantic features and low-level semantic features, it is possible to deeply explore the multi-level information in the sample image, enhance the model's sensitivity to image details and semantics, so that the image super-resolution model can better retain and restore the important details and features in the input image when processing the image super-resolution task, and further improve the quality of the generated image.

[0062] In some embodiments of the present disclosure, the initial model includes an image decoder; decoding processing is performed according to the encoding features and the output image features to obtain a predicted image of the sample image, including: decoding processing is performed on the output image features through the image decoder to obtain decoding features of the sample image; and the encoding features and the decoding features are fused to obtain a predicted image.

[0063] In the disclosed embodiment, the initial model includes an image decoder. Exemplarily, the image decoder is a variational autoencoder decoder (Decoder of Variational Autoencoder, VAE decoder). In addition, the initial model also includes an image encoder, which encodes the sample image through the image encoder to obtain image sample features of the sample image. Exemplarily, the image encoder is a variational autoencoder (Variational Autoencoder, VAE encoder).

[0064] The image decoder receives as input the output image features, which are obtained by fusing the coding features, high-level semantic features, and low-level semantic features during the forward propagation of the model. The image decoder decodes the output image features to extract the information required to reconstruct the image. After decoding the decoded features of the sample image, the decoded features are fused with the coding features, that is, the features generated by the image decoder are combined with the coding features of the sample image to further refine and optimize the quality of the predicted image. By fusing these two features, more image details and contextual information can be captured, thereby generating a more accurate and realistic predicted image.

[0065] Figure 3FIG. 4 shows a model framework diagram of an image super-resolution model according to an embodiment of the present disclosure. Figure 3 As shown, the model includes an image encoder 310, a diffusion model 320 and an image decoder 330. The sample image is input to the image encoder 310, and the image encoder 310 encodes the sample image to obtain the encoding features of the sample image. The encoding features, the high-level semantic features and the low-level semantic features of the sample image are input to the diffusion model 320. In addition, the noise data is also input to the diffusion model 320, and the diffusion model performs fusion attention processing to obtain the output image features.

[0066] The image encoder 310 includes a plurality of encoding network layers, and the image decoder 330 includes a plurality of decoding network layers, and the encoding network layers correspond to the decoding network layers one by one. For each encoding network layer i of the image encoder 310, the encoding feature G output by the encoding network layer i is i It is sent to the decoding network layer i corresponding to the encoding network layer, and the encoding feature G i The decoded feature F output by the decoding network layer i i Fusion is performed, and a predicted image is finally generated through layer-by-layer decoding and fusion.

[0067] In some embodiments of the present disclosure, an image decoder includes a feature adaptation module, which includes a cross-attention module and a convolution layer; the encoded features and the decoded features are fused to obtain a predicted image, including: fusing the encoded features and the decoded features through the cross-attention module to obtain fused features; generating convolution parameters corresponding to the encoded features through the convolution layer, and convolving the fused features based on the convolution parameters to obtain a predicted image.

[0068] In the disclosed embodiment, a feature adaptation module is provided in each decoding network layer of the image decoder, and the encoding features extracted in the encoding stage and the decoding features gradually constructed in the decoding stage are fused through the feature adaptation module.

[0069] Figure 4 FIG. 2 shows a module structure diagram of a feature adaptation module of an embodiment of the present disclosure. Figure 4 As shown, the feature adaptation module includes a cross attention module and a convolution layer. For the feature adaptation module in the decoding network layer i, the cross attention module of the feature adaptation module receives the encoding feature G from the encoding network layer i i and the decoded features F from the current decoding network layer i As input, cross-attention processing is performed on the input features to obtain the fused features The calculation formula is as follows:

[0070]

[0071] In formula (1), CA represents the function of the cross attention module. In the process of fusion processing through the cross attention model, the encoding feature G i Perform feature extraction, extract key features (K) and value features (V), and decode features F i Perform feature extraction, extract query features (Q), perform normalization based on the extracted query features (Q), key features (K) and value features (V), and the fused features

[0072] The convolutional layer receives the fused features from the cross-attention module As input, generate the encoded feature G i The corresponding convolution parameters, specifically, generate the proportional weight matrix W through the convolution layer i and the bias matrix b i , using convolution parameters to fused features Perform convolution processing, the calculation formula is as follows:

[0073]

[0074] In formula (2), It is the feature after convolution processing by the convolution layer. In the disclosed embodiment, each decoding network layer in the image decoder fuses the encoding features from the corresponding encoding network layer and the decoding features from the current decoding network layer through its feature adaptation module, and finally generates a predicted image through layer-by-layer decoding and fusion.

[0075] In the model training method of the disclosed embodiment, the feature adaptation module of the image decoder realizes the effective fusion of the encoding features and the decoding features, and further processes the fused features using the convolutional layer to finally generate a predicted image, thereby improving the model's ability to capture image details and the accuracy of the overall prediction, thereby being able to generate a more detailed and realistic predicted image.

[0076] In some embodiments of the present disclosure, semantic extraction is performed on a sample image to obtain high-level semantic features and low-level semantic features of the sample image, including: performing semantic extraction on the sample image based on a multimodal large language model to obtain high-level semantic hint information and low-level semantic hint information of the sample image; encoding the high-level semantic hint information and the low-level semantic hint information respectively based on a text encoder to obtain high-level semantic features and low-level semantic features.

[0077] In the disclosed embodiment, a multi-modal large language model (MLLM) is a large language model that can process and integrate multiple modal information such as text, images, audio and video. As the data scale and model capacity increase, MLLM demonstrates excellent capabilities in semantic understanding. Therefore, MLLM is used to parse sample images and extract semantic information from them.

[0078] In the disclosed embodiment, in order to ensure that the output generated by MLLM conforms to human preferences, clear instructions are defined to guide the behavior of the model. For the semantic acquisition of high-level semantic hint information, the multimodal large language model LLaVA (Large Language and Vision Assistant) is selected to generate a descriptive summary, and the instruction is set to "Please provide a descriptive summary of the content of this image". Through this instruction, LLaVA can be stably guided to generate expected high-level semantic hint information including object description, spatial location, scene, etc.

[0079] In the disclosed embodiment, for the semantic acquisition of low-level semantic hint information, it is difficult to directly obtain these details from the image due to the lack of special training of MLLM. Based on this, InternLM-Xcomposer (visual language multimodal large model) fine-tuned on Q-Instruct (human feedback dataset focusing on low-level vision) related to low-level vision is selected as MLLM to extract semantic information, wherein Q-Instruct is a human feedback dataset focusing on low-level vision. After fine-tuning the dataset, InternLM-XComposer can more accurately capture the low-level semantic features of the image. When acquiring low-level semantic hint information, the distribution of image degradation in real-world scenes is analyzed. These degradations are usually caused by factors such as shooting process, encoding conversion compression or transmission, resulting in motion blur, low light conditions, color shift, noise and compression artifacts. Various distortion phenomena, for this case, the instruction "Please describe the quality of this image and evaluate it based on factors such as clarity, color, noise and lighting" is used to generate relevant descriptions of the overall image quality, clarity, noise level, color, etc.

[0080] In the disclosed embodiment, the text encoder can capture and convert the semantic information in the text to make it a useful feature vector in the image generation process. For example, a CLIP (Contrastive Language-Image Pre-training) text encoder can be used.

[0081] Figure 5 FIG. 2 shows a schematic diagram of obtaining semantic features based on a multimodal large language model according to an embodiment of the present disclosure. Figure 5 As shown in the figure, semantic features of sample images are extracted based on MLLM to obtain high-level semantic hint information and low-level semantic hint information of the sample images. Among them, high-level semantic hint information includes image descriptions such as object description, spatial position, scene, etc., and low-level semantic hint information includes descriptions of the overall image quality, clarity, noise level, color, etc. For example, the high-level semantic hint information is "there is a tall white tower with a green roof in the image, the top of the tower is decorated with a cross, surrounded by several trees, and the tower is in front of the sky background", and the low-level semantic hint information is "a photo of a church spire, which looks blurry and out of focus, lack of color and poor focus on the subject resulting in poor quality, taken from a distance, limiting the view of the church spire".

[0082] After obtaining the high-level semantic hint information and the low-level semantic hint information of the sample image, the high-level semantic hint information and the low-level semantic hint information of the sample image are respectively encoded by the CLIP text encoder to obtain the high-level semantic features and the low-level semantic features of the sample image.

[0083] In the model training method of the embodiment of the present disclosure, semantics of different levels are extracted from sample images through MLLM, and the extracted semantic features are encoded to obtain high-level semantic features and low-level semantic features. The high-level semantic features and low-level semantic features are input into the initial model as semantic prior knowledge to guide the training of the initial model, which can improve the semantic understanding ability of the model.

[0084] In some embodiments of the present disclosure, the initial model includes a control network module and a backbone network module, the control network module includes a semantic fusion attention module, and the backbone network module includes a conditional attention module; fusion attention processing is performed on the encoding features, high-level semantic features and low-level semantic features to obtain output image features of the sample image, including: performing fusion attention processing on the encoding features, high-level semantic features and low-level semantic features through the semantic fusion attention module to obtain fusion attention features of the sample image; performing conditional attention processing on the fusion attention features, high-level semantic features and noise data through the conditional attention module to obtain output image features.

[0085] In the disclosed embodiment, the initial model can be constructed based on a controlled diffusion model (Controlled Diffusion Model), wherein the controlled diffusion model includes a control network module and a backbone network module. The control network module acts as a controller in the model and is responsible for providing control information to the backbone network module. The backbone network module is another key component in the model and is responsible for generating output image features based on the control information provided by the control network module. Exemplarily, the control network module is a ControlNet network, and the backbone network module is a U-Net network of a stable model (Stable Diffusion Model), including a U-Net encoder and a U-Net decoder.

[0086] Among them, the control network module includes a semantic fusion attention module, which can be used to process and fuse semantic information from different sources. In the disclosed embodiment, the control network module extracts the encoded features of the sample image, maps them to the latent variable space, and obtains the features of the sample image in the latent variable space. Then, the semantic fusion attention module applies the attention mechanism to the features of the sample image in the latent variable space, the high-level semantic features, and the low-level semantic features, and fuses these features together to generate the fused attention features of the sample image.

[0087] In the disclosed embodiment, the backbone network module includes an encoder and a decoder, and the decoder includes a conditional attention module. The conditional attention module performs conditional attention processing on the fused attention features, high-level semantic features and noise data to obtain output image features, including: inputting the noise data and high-level semantic features into the encoder for feature extraction to obtain the noisy latent variable of the sample image, that is, the noisy feature of the sample image in the latent variable space, which integrates noise and high-level semantic information; inputting the noisy feature of the sample image in the latent variable space into the decoder for decoding processing to obtain decoded features, and then applying the conditional attention mechanism to the decoded features and the fused attention features through the conditional attention module, that is, dynamically adjusting the attention weight of the decoded features according to the information in the fused attention features, and finally obtaining the output image features of the sample image.

[0088] In some embodiments of the present disclosure, a fusion attention module includes a first attention module, a second attention module and a third attention module; the semantic fusion attention module performs fusion attention processing on the coding features, high-level semantic features and low-level semantic features to obtain a fusion attention feature of the sample image, including: performing cross-attention processing on the coding features and high-level semantic features through the first attention module to obtain a first cross-attention feature of the sample image; performing cross-attention processing on the coding features and low-level semantic features through the second attention module to obtain a second cross-attention feature of the sample image; performing fusion attention processing on the first cross-attention feature and the second cross-attention feature through the third attention module to obtain a fusion attention feature.

[0089] In the disclosed embodiment, the fusion attention module includes a first attention module, a second attention module and a third attention module. The first attention module and the second attention module are parallel branches, which are used to perform cross-attention processing on high-level semantic features and low-level semantic features respectively, and then the third attention module is used to perform fusion attention processing on the obtained cross-attention features to obtain fusion attention features.

[0090] Figure 6 FIG. 4 shows a model framework diagram of another image super-resolution model of an embodiment of the present disclosure. Figure 6 As shown, the model includes an image encoder 610, an image decoder 620, a ControlNet network 630, a U-Net encoder 640 and a U-Net decoder 650, the ControlNet network 630 includes a semantic fusion attention module 631, the U-Net decoder 650 includes a conditional attention module 651, and the ControlNet network 630 is copied by the U-Net encoder 640 for initialization.

[0091] The sample image is input to the image encoder 610, and the image encoder 610 encodes the sample image to obtain the encoded features of the sample image. The encoded features of the sample image are input to the ControlNet network 630, and the ControlNet network 630 extracts the encoded features of the sample image and maps them to the latent variable space to obtain the features of the sample image in the latent variable space. The features of the sample image in the latent variable space and the high-level semantic features and low-level semantic features of the sample image are used as inputs to the semantic fusion attention module 631, and the semantic fusion attention module 631 performs fusion attention processing on the input features to obtain the fusion attention features of the sample image.

[0092] like Figure 6As shown, the semantic fusion attention module 631 includes a first attention module 6311, a second attention module 6312 and a third attention module 6313. The features and high-level semantic features of the sample image in the latent variable space are input into the first attention module 6311, and the features of the sample image in the latent variable space are extracted, the query feature (Q) is extracted, the high-level semantic feature is extracted, the key feature (K) and the value feature (V) are extracted, and the extracted query feature (Q), key feature (K) and value feature (V) are normalized to obtain the first cross attention feature. The features and low-level semantic features of the sample image in the latent variable space are input into the second attention module 6312, and the features of the sample image in the latent variable space are extracted, the query feature (Q) is extracted, the low-level semantic feature is extracted, the key feature (K) and the value feature (V) are extracted, and the extracted query feature (Q), key feature (K) and value feature (V) are normalized to obtain the second cross attention feature. After obtaining the first cross-attention feature and the second cross-attention feature, these features are used as the input of the third attention module 6313, and the third attention module 6313 performs fusion attention processing, performs feature extraction on the first cross-attention feature to extract the query feature (Q), performs feature extraction on the second cross-attention feature to extract the key feature (K) and the value feature (V), and performs normalization processing on the extracted query feature (Q), key feature (K) and value feature (V) to obtain the fusion attention feature.

[0093] The noise data and the high-level semantic features of the sample image are input to the U-Net encoder 640 for feature extraction to obtain the noise-added features of the sample image in the latent variable space, which integrates the noise and high-level semantic information. The noise-added features of the sample image in the latent variable space are input to the U-Net decoder 650 for decoding processing to obtain the decoded features, and then the conditional attention mechanism is applied to the decoded features and the fused attention features through the conditional attention module 651 to finally obtain the output image features of the sample image.

[0094] In the model training method of the embodiment of the present disclosure, the fused attention feature of the sample image not only contains the detailed information of the sample image, but also incorporates the abstract representation of high-level and low-level semantics, thereby enhancing the semantic understanding ability of the sample image and helping the model to more accurately capture the key information in the image; and, a conditional attention mechanism is applied to the fused attention feature, that is, the attention weight of the decoded feature can be dynamically adjusted according to the information in the fused attention feature, thereby improving the quality and accuracy of the output image features.

[0095] Furthermore, by performing cross-attention processing on high-level semantic features and low-level semantic features respectively through the first attention module and the second attention module, and then performing fusion attention processing on the obtained cross-attention features, a balance can be achieved between different levels of priors, thereby allowing adaptive selection of semantic features.

[0096] The model training process of the embodiment of the present disclosure may include two stages. The first stage trains the image encoder 610, the ControlNet network 630 (including the semantic fusion attention module 631), the semantic fusion attention module 631 (including the first attention module 6311, the second attention module 6312 and the third attention module 6313) and the conditional attention module 651. The second stage uses the first stage training structure to train the image decoder 620.

[0097] Since real-world images may experience various degradations such as blurring, blocky artifacts, etc., which affect the distortion of high-frequency and low-frequency components (appearing in pixel space and latent variable space), in order to mitigate the impact of degradation and extract robust semantic information from sample images, in the first stage, a degradation-free constraint (DFC) is attached in the pixel space and latent variable space respectively.

[0098] like Figure 6 As shown, a VAE encoder 660 is also included, and the VAE encoder 660 extracts features from the high-resolution image to obtain a downsampling result x of the high-resolution image. hr,i And the downsampling result z in the latent variable space hr,j The sample image is encoded by the image encoder 610, and the pixel space constraint can be applied in the image encoder 610. Exemplarily, in the i-th image encoding layer of the image encoder 610, a single layer convolution is used to map the feature to an image with RGB channels. Application 1 The loss makes The downsampled result x is as close as possible to the high-resolution image hr,i Similarly, the ControlNet network 630 uses a pyramid-like U-Net encoder 640 to encode the semantic features of the latent variables, and the latent variable space constraints can be applied in the ControlNet network 630. The features obtained from the j-th layer of the ControlNet network 630 are mapped to the latent variable space Use L 1 The loss is compared with the down-sampling result z of the high-resolution image in the latent variable space. hr,j Alignment. Combining the constraints in pixel space and variable space, the loss function is:

[0099]

[0100] In formula (3), L DFC Indicates DFC loss.

[0101] During training, the high-resolution images are mapped to the latent variable space z hr , by adding noise ε through a diffusion process following t steps to obtain the noisy latent variable Where t is randomly sampled from the range [1, T], and T is the preset training step. The image super-resolution model depends on the sample image x lr , noisy latent variables High-level semantic featuresc h and low-level semantic features c l To predict the added noise, the optimization goal is:

[0102]

[0103] In formula (4), Can express expectations; ||x|| 2 It can represent the L2 norm of x; ∈ can represent the actual image noise; It can be expressed as the predicted image noise. Finally, the loss function L of the first stage is:

[0104] L=λL DFC +L D (5)

[0105] In Formula (5), λ is the balancing coefficient. During testing, a classifier-free guidance strategy can be adopted, which allows the image super-resolution model to generate better quality images through negative cues without additional training. In each training step, according to the high-level semantic features c h and low-level semantic features c l Make predictions while using high-level negative hints c h,neg Replace c h , using low-level negative hint c l,neg Replace c l , these predictions are fused to obtain the final output:

[0106]

[0107] In formula (6), λ s Is the guiding scale, is the noise-added latent variable of the sample image, Substituting the predicted noise into formula (5), we use “blurred, noisy, unclear, low resolution” as the high-level negative prompt. At the same time, due to the low quality of the sample image, we use “clear, high quality, high resolution, clean” as the low-level negative prompt.

[0108] Through the first stage of training, the trained image encoder 610, ControlNet network 630, semantic fusion attention module 631 (including the first attention module 6311, the second attention module 6312 and the third attention module 6313) and the conditional attention module 651 are obtained, and then the second stage of training is carried out to obtain the trained image decoder 620, and finally the trained image super-resolution model is obtained.

[0109] In some embodiments of the present disclosure, an initial model is trained based on a high-resolution image and a predicted image to obtain an image super-resolution model, including: constructing a model loss function based on the high-resolution image and the predicted image; adjusting parameters of an image decoder based on the model loss function to obtain an image decoder with adjusted parameters to obtain an image super-resolution model.

[0110] In the disclosed embodiment, the initial model may be an image super-resolution model obtained by training in the first stage. The high-resolution image of the sample image is the desired result in the training process. In order to evaluate the difference between the predicted image generated by the model and the high-resolution image, a loss function, such as mean square error (MSE), is defined to quantify the error between the predicted image and the high-resolution image.

[0111] In the disclosed embodiment, the image decoder is trainable. Exemplarily, the image decoder includes multiple image decoding layers, each image decoding layer is provided with a feature adaptation module, and the feature adaptation model of each image decoding layer is trainable. The loss value is calculated using the high-resolution image and the predicted image generated by the model, and then the gradient of the loss function with respect to the model parameters is calculated using the back propagation algorithm, and the parameters of the image decoder are updated using gradient descent. The training process will be iterated multiple times. In each iteration, the model will generate a new predicted image based on the current parameter settings and calculate the corresponding loss value. The loss value is used to adjust the parameters of the image decoder to further reduce the difference between the predicted image and the high-resolution image until the loss value converges to a lower level, or reaches a preset number of training rounds, and obtains an image decoder with adjusted parameters, and finally obtains a trained image super-resolution model.

[0112] It should be noted that the model training process of the embodiment of the present disclosure can also be a stage. During the training process, the model loss function is constructed using the loss function of the first stage and the loss function of the second stage. The image encoder 610, the ControlNet network 630 (including the semantic fusion attention module 631), the semantic fusion attention module 631 (including the first attention module 6311, the second attention module 6312 and the third attention module 6313), the conditional attention module 651 and the image decoder 620 are trained through the constructed model loss function to finally obtain a trained image super-resolution model.

[0113] In the model training method of the embodiment of the present disclosure, a model loss function is constructed based on the high-resolution image and the predicted image, and the parameters of the image decoder are adjusted based on the loss function to obtain an image super-resolution model that can generate high-resolution images. The model can generate high-resolution images and improve the ability to restore high-resolution images from low-resolution images.

[0114] Figure 7 A flowchart of an image super-resolution processing method according to an embodiment of the present disclosure is shown. Figure 7 The execution subject of the method provided in the embodiment may be any electronic device, such as Figure 1 The terminal device 101 in the embodiment is also as follows Figure 1 The terminal device 103 in Figure 1 The server 101 and the terminal device 103 in the embodiment jointly implement the image super-resolution processing method, but the present disclosure is not limited thereto. Figure 7 The image super-resolution processing method provided by the embodiment of the present disclosure includes the following steps.

[0115] Step S710, obtaining an image to be processed.

[0116] In the embodiment of the present disclosure, the image to be processed may be an input image to be subjected to super-resolution processing, wherein the image to be processed may be a low-resolution image input by a user, or other image containing noise.

[0117] Step S720: performing semantic extraction on the image to be processed to obtain high-level semantic features and low-level semantic features of the image to be processed.

[0118] In the disclosed embodiment, after acquiring the image to be processed, semantic extraction is performed on the image to be processed based on MLLM to obtain high-level semantic hint information and low-level semantic hint information of the image to be processed, and then the high-level semantic hint information and low-level semantic hint information of the image to be processed are respectively encoded based on the text encoder to obtain high-level semantic features and low-level semantic features of the image to be processed.

[0119] Step S730, input the image to be processed, the high-level semantic features and the low-level semantic features of the image to be processed into the image super-resolution model, and output the super-resolution image corresponding to the image to be processed; wherein the image super-resolution model is trained according to the model training method of the above embodiment.

[0120] In an embodiment of the present disclosure, a pre-trained image super-resolution model is obtained, which is obtained by training according to the above-mentioned model training method. The training process has been explained in detail above, and the present disclosure will not repeat it again.

[0121] Among them, the image super-resolution model can obtain semantic priors at different levels through MLLM to guide the diffusion model to generate higher-fidelity and more realistic images, which can be called XPSR (Cross-modal Priors for Super-Resolution, cross-modal prior guided image super-resolution model).

[0122] In the disclosed embodiment, the super-resolution image may be a high-resolution image obtained by super-reconstructing the image to be processed. The image to be processed, the high-level semantic features and the low-level semantic features of the image to be processed are used as inputs of the image super-resolution model, and the image super-resolution model performs super-reconstruction processing on the input data, and finally obtains a super-resolution image corresponding to the image to be processed.

[0123] The image super-resolution model includes a trained image encoder, a control network module, a backbone network module and an image decoder. The image super-resolution model is used to perform super-resolution reconstruction on the image to be processed. The process of obtaining a super-resolution image is as follows:

[0124] (1) encoding the image to be processed by an image encoder to obtain encoding features of the image to be processed;

[0125] (2) The control network module extracts the encoded features to obtain the features of the image to be processed in the latent variable space, and the semantic fusion attention module of the control network module performs fusion attention processing on the features of the image to be processed in the latent variable space, the high-level semantic features and the low-level semantic features to obtain the fusion attention features of the image to be processed;

[0126] (3) The encoder in the backbone network module extracts high-level semantic features of the image to be processed, and then uses the extracted features together with the fused attention features of the image to be processed as the input of the decoder in the backbone network module. The decoder in the backbone network module performs decoding processing. During the decoding processing, the conditional attention module of the backbone network module performs conditional attention processing, and finally obtains the output image features of the image to be processed;

[0127] (4) The output image features of the image to be processed are input into the image decoder for decoding, and then fused with the encoded features of the image to be processed output by the image encoder, so as to finally obtain a super-resolution image of the image to be processed.

[0128] By comparing the image super-resolution model (XPSR) of the embodiment of the present disclosure with other models, such as BSRGAN (Bidirectional Super-Resolution Generative Adversarial Network), Real-ESRGAN (Enhanced Super-Resolution Generative Adversarial Network), StableSR (Stable Super Resolution), PASD (Pixel-Aware Stable Diffusion), and SeeSR, it is found that the image obtained by super-reconstruction processing based on the image super-resolution model (XPSR) of the embodiment of the present disclosure is more semantically accurate and detailed than the images obtained by other methods, and can understand the semantic context to generate clear images. In addition, the image super-resolution model (XPSR) based on the embodiment of the present disclosure can significantly improve the resolution of the image. For example, compared with the images obtained by other methods, the images obtained based on the image super-resolution model (XPSR) have more detailed animal hair and clearer textures of leaves and walnuts.

[0129] The image super-resolution processing method provided by the embodiment of the present disclosure uses the coding features of the image encoder to guide the decoding process during the decoding process using the image decoder of the image super-resolution, and can more accurately capture the semantic information in the image to be processed, so that even if there are severely degraded areas in the image to be processed, the important semantic details can be effectively restored, thereby generating a high-resolution image that is highly similar to the image to be processed and rich in details. In addition, by fusing the coding features, high-level semantic features and low-level semantic features of the image to be processed, it is possible to deeply explore the multi-level information in the image to be processed, enhance the sensitivity to image details and semantics, better retain and restore the important details and features in the image to be processed, and further improve the quality of the generated image.

[0130] It can be understood that the same / similar parts between the various embodiments of the above method in this specification can refer to each other, and each embodiment focuses on the differences from other embodiments. For related points, please refer to the description of other method embodiments.

[0131] Figure 8 A block diagram of a model training device according to an embodiment of the present disclosure is shown. Figure 8 The device 800 includes: a sample acquisition module 810, a sample semantic extraction module 820 and a model training module 830.

[0132] Among them, the sample acquisition module 810 is configured to: acquire a sample image and a high-resolution image corresponding to the sample image. The sample semantic extraction module 820 is configured to: perform semantic extraction on the sample image to obtain high-level semantic features and low-level semantic features of the sample image. The sample training module 830 is configured to: encode the sample image through the initial model to obtain the encoding features of the sample image; perform feature extraction on the encoding features, high-level semantic features and low-level semantic features to obtain the output image features of the sample image; perform decoding according to the encoding features and the output image features to obtain the predicted image of the sample image; train the initial model based on the high-resolution image and the predicted image to obtain the image super-resolution model.

[0133] In some embodiments of the present disclosure, the initial model includes an image decoder, and the sample training module 830 is further configured to: decode the output image features through the image decoder to obtain the decoded features of the sample image; and fuse the encoded features with the decoded features to obtain a predicted image.

[0134] In some embodiments of the present disclosure, the image decoder includes a feature adaptation module, the feature adaptation module includes a cross-attention module and a convolution layer, and the sample training module 830 is also configured to: fuse the encoding features and the decoding features through the cross-attention module to obtain fused features; generate convolution parameters corresponding to the encoding features through the convolution layer, and convolve the fused features based on the convolution parameters to obtain a predicted image.

[0135] In some embodiments of the present disclosure, the initial model includes a control network module and a backbone network module, the control network module includes a semantic fusion attention module, the backbone network module includes a conditional attention module, and the sample training module 830 is also configured to: perform fusion attention processing on the encoding features, high-level semantic features and low-level semantic features through the semantic fusion attention module to obtain the fusion attention features of the sample image; perform conditional attention processing on the fusion attention features, high-level semantic features and noise data through the conditional attention module to obtain the output image features.

[0136] In some embodiments of the present disclosure, the semantic fusion attention module includes a first attention module, a second attention module and a third attention module, and the sample training module 830 is also configured to: perform cross-attention processing on the encoding features and the high-level semantic features through the first attention module to obtain the first cross-attention features of the sample image; perform cross-attention processing on the encoding features and the low-level semantic features through the second attention module to obtain the second cross-attention features of the sample image; perform fusion attention processing on the first cross-attention features and the second cross-attention features through the third attention module to obtain the fusion attention features.

[0137] In some embodiments of the present disclosure, the sample semantic extraction module 820 is further configured to: perform semantic extraction on the sample image based on a multimodal large language model to obtain high-level semantic hint information and low-level semantic hint information of the sample image; and encode the high-level semantic hint information and the low-level semantic hint information based on a text encoder to obtain high-level semantic features and low-level semantic features.

[0138] In some embodiments of the present disclosure, the sample training module 830 is further configured to: construct a model loss function based on the high-resolution image and the predicted image; adjust the parameters of the image decoder based on the model loss function to obtain an image decoder with adjusted parameters to obtain an image super-resolution model.

[0139] Fig. 9 FIG. 1 is a block diagram of an image super-resolution processing device according to an embodiment of the present disclosure. Fig. 9 The device 900 includes: an image acquisition module 910, an image semantic extraction module 920 and an image processing module 930.

[0140] The image acquisition module 910 is configured to acquire an image to be processed. The image semantic extraction module 920 is configured to perform semantic extraction on the image to be processed to obtain high-level semantic features and low-level semantic features of the image to be processed. The image processing module 930 is configured to input the image to be processed, the high-level semantic features and the low-level semantic features of the image to be processed into an image super-resolution model, and output a super-resolution image corresponding to the image to be processed; wherein the image super-resolution model is trained according to the model training method of the above embodiment.

[0141] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0142] Fig.10 FIG. 1 is a schematic diagram showing the structure of an electronic device according to an embodiment of the present disclosure. It should be noted that: Fig.10The electronic device 1000 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0143] like Fig.10 As shown, the electronic device 1000 is in the form of a general computing device. The components of the electronic device 1000 may include but are not limited to: the at least one processing unit 1010, the at least one storage unit 1020, and a bus 1030 connecting different system components (including the storage unit 1020 and the processing unit 1010).

[0144] The storage unit stores program codes, which can be executed by the processing unit 1010, so that the processing unit 1010 performs the steps according to various exemplary embodiments of the present invention described in the above “Exemplary Method” section of this specification. For example, the processing unit 1010 can perform the following steps: Figure 2 Follow the steps shown in .

[0145] The storage unit 1020 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 10201 and / or a cache storage unit 10202 , and may further include a read-only storage unit (ROM) 10203 .

[0146] The storage unit 1020 may also include a program / utility 10204 having a set (at least one) of program modules 10205, such program modules 10205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0147] Bus 1030 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0148] The electronic device 1000 may also communicate with one or more external devices 1100 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 1000, and / or communicate with any device that enables the electronic device 1000 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed via an input / output (I / O) interface 1050. Furthermore, the electronic device 1000 may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter 1040. As shown, the network adapter 1040 communicates with other modules of the electronic device 1000 via a bus 1030. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 1000, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0149] In an exemplary embodiment of the present disclosure, a computer-readable storage medium is also provided, on which a program product capable of implementing the above method of the present specification is stored. In some possible implementations, various aspects of the present invention can also be implemented in the form of a program product, which includes a program code, and when the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the above "Exemplary Method" section of the present specification.

[0150] The program product for implementing the above method according to the embodiment of the present invention may adopt a portable compact disk read-only memory (CD-ROM) and include program code, and may be run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto, and in this document, a readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, apparatus, or device.

[0151] The program product may use any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0152] Computer readable signal media may include data signals propagated in baseband or as part of a carrier wave, in which readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Readable signal media may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0153] The program code embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the foregoing.

[0154] Program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, etc., and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).

[0155] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into multiple modules or units to be embodied.

[0156] In addition, although the steps of the method in the present disclosure are described in a specific order in the drawings, this does not require or imply that the steps must be performed in this specific order, or that all the steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps, etc.

[0157] Through the description of the above implementation, it is easy for those skilled in the art to understand that the example implementation described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the implementation of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the implementation of the present disclosure.

[0158] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any modification, use or adaptation of the present disclosure, which follows the general principles of the present disclosure and includes common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are to be considered as exemplary only, and the true scope and spirit of the present disclosure are indicated by the appended claims.

[0159] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A model training method, characterized in that: The method comprises: Acquire a sample image and a high-resolution image corresponding to the sample image; Performing semantic extraction on the sample image to obtain high-level semantic features and low-level semantic features of the sample image; The sample image is encoded by the initial model to obtain the encoding features of the sample image; the encoding features, the high-level semantic features and the low-level semantic features are fused with attention to obtain the output image features of the sample image; decoding is performed according to the encoding features and the output image features to obtain the predicted image of the sample image; The initial model is trained based on the high-resolution image and the predicted image to obtain an image super-resolution model.

2. The method according to claim 1, characterized in that The initial model includes an image decoder; The decoding process is performed according to the coding feature and the output image feature to obtain the predicted image of the sample image, including: The output image features are decoded by the image decoder to obtain the decoded features of the sample image; the encoded features are fused with the decoded features to obtain the predicted image.

3. The method according to claim 2, characterized in that The image decoder includes a feature adaptation module, wherein the feature adaptation module includes a cross-attention module and a convolutional layer; The fusing the encoding feature with the decoding feature to obtain the predicted image includes: The encoding feature and the decoding feature are fused by the cross attention module to obtain a fused feature; The convolution layer generates convolution parameters corresponding to the encoded features, and the fused features are convolved based on the convolution parameters to obtain the predicted image.

4. The method according to claim 1, characterized in that The initial model includes a control network module and a backbone network module, the control network module includes a semantic fusion attention module, and the backbone network module includes a conditional attention module; The step of performing fusion attention processing on the coding feature, the high-level semantic feature and the low-level semantic feature to obtain the output image feature of the sample image includes: Performing fusion attention processing on the encoding feature, the high-level semantic feature and the low-level semantic feature through the semantic fusion attention module to obtain a fusion attention feature of the sample image; The conditional attention module performs conditional attention processing on the fused attention features, the high-level semantic features and the noise data to obtain the output image features.

5. The method according to claim 4, characterized in that The semantic fusion attention module includes a first attention module, a second attention module and a third attention module; The step of performing fusion attention processing on the encoding feature, the high-level semantic feature and the low-level semantic feature through the semantic fusion attention module to obtain the fusion attention feature of the sample image includes: Performing cross-attention processing on the encoding feature and the high-level semantic feature through the first attention module to obtain a first cross-attention feature of the sample image; Performing cross-attention processing on the encoding feature and the low-level semantic feature through the second attention module to obtain a second cross-attention feature of the sample image; The first cross-attention feature and the second cross-attention feature are subjected to fused attention processing by the third attention module to obtain the fused attention feature.

6. The method according to claim 1, characterized in that The step of performing semantic extraction on the sample image to obtain high-level semantic features and low-level semantic features of the sample image includes: Performing semantic extraction on the sample image based on a multimodal large language model to obtain high-level semantic prompt information and low-level semantic prompt information of the sample image; The high-level semantic prompt information and the low-level semantic prompt information are respectively encoded based on a text encoder to obtain the high-level semantic features and the low-level semantic features.

7. The method according to claim 2, characterized in that The step of training the initial model based on the high-resolution image and the predicted image to obtain an image super-resolution model comprises: Constructing a model loss function based on the high-resolution image and the predicted image; The parameters of the image decoder are adjusted based on the model loss function to obtain an image decoder with adjusted parameters, so as to obtain the image super-resolution model.

8. An image super-resolution processing method, characterized in that: The method comprises: Get the image to be processed; Performing semantic extraction on the image to be processed to obtain high-level semantic features and low-level semantic features of the image to be processed; The image to be processed, the high-level semantic features and the low-level semantic features of the image to be processed are input into an image super-resolution model, and a super-resolution image corresponding to the image to be processed is output; the image super-resolution model is trained according to the model training method described in any one of claims 1 to 7.

9. A model training device, characterized in that: The device comprises: A sample acquisition module, configured to acquire a sample image and a high-resolution image corresponding to the sample image; A sample semantic extraction module is configured to perform semantic extraction on the sample image to obtain high-level semantic features and low-level semantic features of the sample image; The sample training module is configured to encode the sample image through the initial model to obtain the encoding features of the sample image; perform feature extraction on the encoding features, the high-level semantic features and the low-level semantic features to obtain the output image features of the sample image; perform decoding according to the encoding features and the output image features to obtain the predicted image of the sample image; and train the initial model based on the high-resolution image and the predicted image to obtain an image super-resolution model.

10. An image super-resolution processing device, characterized in that: The device comprises: An image acquisition module is configured to acquire an image to be processed; An image semantic extraction module is configured to perform semantic extraction on the image to be processed to obtain high-level semantic features and low-level semantic features of the image to be processed; An image processing module is configured to input the image to be processed, the high-level semantic features and the low-level semantic features of the image to be processed into an image super-resolution model, and output a super-resolution image corresponding to the image to be processed; the image super-resolution model is trained according to the model training method described in any one of claims 1 to 7.

11. An electronic device, characterized in that: include: processor; A memory for storing the processor executable instructions; wherein the processor is configured to execute the executable instructions to implement the model training method as described in any one of claims 1 to 7, or to implement the image super-resolution processing method as described in claim 8.

12. A computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enables the electronic device to execute the model training method as described in any one of claims 1 to 7, or execute the image super-resolution processing method as described in claim 8.

13. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the model training method as described in any one of claims 1 to 7 is implemented, or the image super-resolution processing method as described in claim 8 is implemented.