An image processing method, a terminal device and a readable storage medium

By combining a multimodal generative magnification model with a denoising network and a super-resolution magnification network, the image of the terminal device is processed iteratively multiple times, which solves the problem of image detail loss caused by excessive magnification and achieves detail restoration and sharpness improvement of high-resolution images.

CN119255121BActive Publication Date: 2025-12-16HONOR DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410387271.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-30
Publication Date
2025-12-16
Estimated Expiration
2044-03-30

AI Technical Summary

Technical Problem

When the actual magnification expected by the user is too high, the magnified image of the object obtained by the terminal device suffers severe loss of detail and poor image clarity.

Method used

By acquiring the first image, generating textual description information and a noisy image, and inputting them into a multimodal generative amplification model for generative super-resolution processing, the model uses a denoising network and a super-resolution amplification network for multiple iterations to generate a high-resolution image with higher detail restoration.

Benefits of technology

It improves the detail and clarity of the image after super-resolution magnification, enhancing the user's visual experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119255121B_ABST
    Figure CN119255121B_ABST
Patent Text Reader

Abstract

The application provides an image processing method, a terminal device and a readable storage medium, and applies to the terminal technical field. The method comprises the following steps: inputting a low-resolution first image, a noise image and text description information obtained into a multi-modal generative magnification model, so as to output a high-resolution second image conforming to the text description information, and the second image obtained has a high degree of detail restoration and a clear image; and the text description information is generated based on the first image. According to the embodiment of the application, the multi-modal generative magnification model is used for performing generative super-resolution magnification processing on the first image, so that a high-quality second image can be obtained, the second image has a high degree of detail restoration, and the visual experience of a user can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of terminal, and in particular, to an image processing method, a terminal device and a readable storage medium. BACKGROUND

[0002] With the development of terminal devices, more and more photographing modes are set in terminal devices. The terminal device can meet different photographing needs of different users by setting multiple photographing modes.

[0003] At present, when a user uses a terminal device to photograph an object far away from the terminal device, the terminal device can use an actual magnification ratio expected by the user to photograph the object, so as to obtain a magnified object image. However, when the actual magnification ratio expected by the user is too large, the magnified object image obtained by the terminal device will have a serious loss of details and poor definition. SUMMARY

[0004] The present application provides an image processing method, a terminal device and a readable storage medium, which can make the details of the super-resolution magnified image more perfect, improve the definition of the super-resolution magnified image, and further improve the visual experience of the user.

[0005] In a first aspect, the present application provides an image processing method, which comprises:

[0006] obtaining a first image, the resolution of the first image being a first resolution;

[0007] generating text description information of the first image based on the first image; the text description information is used to describe the first image, and the text description information at least includes first information and second information; the first information includes a time corresponding to the first image and / or a place corresponding to the first image; the second information is used to represent a category to which content included in the first image belongs;

[0008] obtaining a noise image; the noise image is used to provide noise distribution characteristics for the first image;

[0009] inputting the first image, the text description information and the noise image into a multi-modal generative magnification model, and outputting a second image; the resolution of the second image is a second resolution, and the second resolution is greater than the first resolution.

[0010] The application can obtain an image to be processed by magnification by acquiring a first image, providing a data basis for subsequent super-resolution magnification processing. Then, based on the obtained first image, text description information of the first image is generated, which can provide a data basis for subsequent generative super-resolution magnification processing. The noise image is obtained, which can provide noise distribution characteristics for the first image. Further, after obtaining the first image, the text description information and the noise image, the first image, the text description information and the noise image can be input into the multi-modal generative magnification model. The multi-modal generative magnification model can perform generative super-resolution magnification processing on the first image based on the input data of the three modalities of the first image, the text description information and the noise image, and can better generate magnified image detail information to obtain a second image with higher detail restoration. That is, after the multi-modal generative magnification model performs generative super-resolution magnification processing on the first image under the guidance and constraint of the first image and the text description information, a higher detail restoration image is obtained. Thus, the multi-modal generative magnification model can perform generative super-resolution magnification processing on a low-resolution (e.g., a first resolution) image (e.g., a first image) and output a high-resolution (e.g., a second resolution) image (e.g., a second image). Moreover, the output high-resolution image has higher detail restoration, higher image clarity and better image quality, and the user's visual experience is better.

[0011] In some possible implementation manners, inputting the first image, the text description information and the noise image into the multi-modal generative magnification model to output the second image includes:

[0012] Inputting the first image, the text description information and the noise image into the multi-modal generative magnification model to perform N iterations to obtain the second image, where N is a positive integer greater than or equal to 1 and less than or equal to a preset iteration threshold.

[0013] In the implementation manners described above, by inputting the data sources of the three modalities of the first image, the text description information and the noise image into the multi-modal generative magnification model and performing N iterations, the detail information in the second image after generative super-resolution magnification processing can be clearer, the quality of the second image after generative super-resolution magnification processing can be higher, and the user's visual experience can be improved.

[0014] In some possible implementation manners, the multi-modal generative magnification model can include a denoising network and a super-resolution magnification network.

[0015] Inputting the first image, the text description information and the noise image into the multi-modal generative magnification model to perform N iterations to obtain the second image includes:

[0016] In the first iteration, the first image, the text description information and the noisy image are input into the denoising network, and the denoising network is used to denoise the noisy image based on the text description information and the first image to obtain a fourth image;

[0017] The fourth image is input into the super-resolution network, and the super-resolution network is used to super-resolution process the fourth image from the denoising network to obtain a fifth image;

[0018] In the Nth iteration, the fifth image, the text description information and the first image are input into the denoising network to obtain a denoised fifth image, and the denoised fifth image is input into the super-resolution network to obtain a super-resolution processed fifth image;

[0019] When N is equal to the preset iteration threshold, the super-resolution processed fifth image is output as the second image.

[0020] In the above implementation, in the first iteration, the denoising network is used to denoise the noisy image based on the text description information and the first image to obtain the fourth image, and the details of the fourth image are clearer than those of the first image. Based on this, the super-resolution network is used to super-resolution process the fourth image to obtain the fifth image with less loss of details. Further, the fifth image, the text description information and the first image are used for the next iteration, and after N iterations, the second image with clearer details and better quality is obtained.

[0021] In some possible implementations, the denoising network includes a first encoder, a first decoder and an attention network.

[0022] In the first iteration, the first image, the text description information and the noisy image are input into the denoising network, and the denoising network is used to denoise the noisy image based on the text description information and the first image to obtain a fourth image, including:

[0023] In the first iteration, the first feature information, the second feature information, the third feature information and the fourth feature information are input into the attention network to obtain fifth feature information; the fifth feature information is used to represent the feature after the attention network fuses multiple features; the first feature information is used to represent the image feature corresponding to the first image; the second feature information is used to represent the text feature corresponding to the text description information; the third feature information is used to represent the noise feature corresponding to the noisy image; and the fourth feature information is used to represent the number feature corresponding to the first iteration;

[0024] The noise image and the fifth feature information are input into the first encoder, and a first encoded feature map is output; the first encoder is configured to perform feature encoding processing on the noise image and the fifth feature information.

[0025] The first encoded feature map is input into the first decoder, and a fourth image is output; the first decoder is configured to perform feature decoding processing on the first encoded feature map.

[0026] In the above implementation manner, the attention network is used to fuse and process multiple types of information, so that the generated detailed information in the first encoded feature map is more accurate, and the detailed information in the fourth image is clearer.

[0027] In some possible implementation manners, the multi-modal generative enlargement model further includes a second encoder, a third encoder, and a fourth encoder.

[0028] Before the first feature information, the second feature information, the third feature information, and the fourth feature information are input into the attention network, the method further includes:

[0029] The first image is input into the second encoder to obtain the first feature information; the second encoder is configured to perform feature encoding processing on the first image.

[0030] The text description information is input into the third encoder to obtain the second feature information; the third encoder is configured to perform feature encoding processing on the text description information.

[0031] The noise image is input into the first encoder to obtain the third feature information.

[0032] The current iteration number is input into the fourth encoder to obtain the fourth feature information; the fourth encoder is configured to perform feature encoding processing on the current iteration number.

[0033] In the above implementation manner, different encoders are used to perform feature encoding processing on different input information to obtain corresponding feature information, which prepares for subsequent feature fusion processing of the attention network.

[0034] In some possible implementation manners, the super-resolution enlargement network includes at least one convolutional layer; the fourth image is input into the super-resolution enlargement network, and the super-resolution enlargement network is used to perform super-resolution enlargement processing on the fourth image from the denoising network to obtain a fifth image, including:

[0035] The fourth image is input into the at least one convolutional layer, and the at least one convolutional layer is used to perform super-resolution enlargement processing on the fourth image based on a preset enlargement ratio to obtain the fifth image.

[0036] In the implementation manner, the at least one convolutional layer is configured to implement the super-resolution magnification processing on the fourth image, so that the loss of the detail information of the fourth image in the super-resolution magnification processing can be avoided, and the second image with clear details and high quality can be obtained.

[0037] In some possible implementation manners, the text description information of the first image is generated based on the first image, and the method comprises the following steps of:

[0038] obtaining first information;

[0039] inputting the first image into a classification model to output second information;

[0040] fusing the first information and the second information to obtain the text description information.

[0041] In the implementation manner, the first information related to the first image is obtained, and the second information of the high-order semantic feature of the first image is obtained through the classification model. Based on this, the text information can be obtained by fusing the first information and the second information. Therefore, the guided condition for the multi-modal generative magnification model to perform the generative super-resolution magnification processing on the first image can be provided, and the accuracy of the detail information in the first image generated by the multi-modal generative magnification model can be improved.

[0042] In some possible implementation manners, the classification model at least comprises an encoder layer and a fully connected layer;

[0043] The inputting the first image into the classification model to output the second information comprises the following steps of:

[0044] inputting the first image into the encoder layer to obtain sixth feature information; the encoder layer is configured to perform feature encoding processing on the first image;

[0045] inputting the sixth feature information into the fully connected layer to obtain the second information; the fully connected layer is configured to perform classification processing on the sixth feature information.

[0046] In the implementation manner, the first image is first subjected to the feature encoding processing, and the sixth feature information obtained by the feature encoding processing is subjected to the classification processing, so that the category to which the content included in the first image belongs can be accurately obtained, and the accuracy of the detail information in the first image generated by the multi-modal generative magnification model can be further improved.

[0047] In some possible implementation manners, before the first image is obtained, the method further comprises the following steps of:

[0048] in response to an operation of starting a camera by a user, displaying a first preview image at a first zoom ratio;

[0049] display a second preview image when the first zoom ratio is adjusted to a second zoom ratio, the second zoom ratio being greater than the first zoom ratio;

[0050] obtaining a first image, comprising:

[0051] in a case where a photographing instruction is received, pre-processing the second preview image to obtain the first image; the pre-processing comprises one or more of the following: denoising processing, automatic white balance processing, demosaicing processing, histogram equalization processing, and image enhancement processing;

[0052] wherein the second image is output, comprising:

[0053] displaying the second image.

[0054] In the above implementation manner, by pre-processing the second preview image after the zoom ratio is adjusted (i.e., the first zoom ratio is adjusted to the second zoom ratio), a first image with enhanced image brightness and contrast can be obtained, thereby providing a good data basis for subsequent processing of the multi-modal generative magnification model.

[0055] In some possible implementation manners, before the second image is displayed, the method further comprises:

[0056] post-processing the second image; the post-processing comprises one or more of the following processing: sharpening processing, compression processing, and format conversion processing.

[0057] In the above implementation manner, by post-processing the second image, the definition of the image after the generative super-resolution magnification processing by the multi-modal generative magnification model can be further enhanced, thereby improving the visual experience of the user.

[0058] In some possible implementation manners, the multi-modal generative magnification model is obtained by using the following training method:

[0059] obtaining a first sample image, a second sample image, text description sample information, and a third sample image; the first sample image and the third sample image are images with the same content and different resolutions, and the resolution of the first sample image is less than that of the third sample image; the second sample image is a noise image used to increase the magnification diversity of the first sample image; the text description sample information is generated based on the first sample image;

[0060] inputting the first sample image, the second sample image, the text description sample information, and the third sample image to the original multi-modal generative magnification model for training until the original multi-modal generative magnification model meets a training iteration condition, and determining the original multi-modal generative magnification model that meets the training iteration condition as the multi-modal generative magnification model.

[0061] In the implementation manner, the multi-modal generative amplification model obtained through multiple iterative training of the original multi-modal generative amplification model with the same network structure as the multi-modal generative amplification model has better effect in generative super-resolution amplification processing of the first image.

[0062] In a second aspect, the present application provides a terminal device, which comprises one or more processors and a memory; the memory is coupled to the one or more processors, and is configured to store computer program codes, the computer program codes comprising computer instructions; the one or more processors invoke the computer instructions to enable the terminal device to perform the image processing method in the first aspect and any possible design of the first aspect.

[0063] In a third aspect, the present application provides a chip system, which is applied to a terminal device; the chip system comprises one or more processors; the one or more processors are configured to invoke computer instructions to enable the terminal device to perform the image processing method in the first aspect and any possible design of the first aspect.

[0064] In the chip system, one chip can be included, or multiple chips can be included; when multiple chips are included in the chip system, the type and number of the chips are not limited.

[0065] In a fourth aspect, the present application provides a computer readable storage medium, which comprises instructions; when the instructions are run on a terminal device, the terminal device is enabled to perform the image processing method in the first aspect and any possible design of the first aspect.

[0066] In a fifth aspect, the present application provides a computer program product; when the computer program product is run on a computer, the computer is enabled to perform the image processing method in the first aspect and any possible design of the first aspect.

[0067] It can be understood that the beneficial effects of the second aspect to the fifth aspect can be referred to the related description in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0068] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0069] Figure 1 The application scenario diagram of the current image processing method provided by an embodiment of the present application is shown in the following figure;

[0070] Figure 2 A schematic diagram of the principle of the image processing method provided by an embodiment of the present application is shown in FIG. 1;

[0071] Figure 3 A schematic diagram of the application scenario of the image processing method provided by an embodiment of the present application is shown in FIG. 2;

[0072] Figure 4 A schematic diagram of the image zooming effect provided by an embodiment of the present application is shown in FIG. 3;

[0073] Figure 5 A schematic diagram of the principle of the image processing method provided by an embodiment of the present application is shown in FIG. 4;

[0074] Figure 6 A flowchart of the image processing method provided by an embodiment of the present application is shown in FIG. 5;

[0075] Figure 7 A flowchart of the image processing method provided by an embodiment of the present application is shown in FIG. 6;

[0076] Figure 8 A schematic diagram of the principle of outputting the first type of information provided by an embodiment of the present application is shown in FIG. 7;

[0077] Figure 9 A schematic diagram of the network architecture of the classification model provided by an embodiment of the present application is shown in FIG. 8;

[0078] Figure 10 A schematic diagram of the principle of the language conversion text module provided by an embodiment of the present application is shown in FIG. 9;

[0079] Figure 11 A schematic diagram of the network architecture of the multi-modal generative zooming model provided by an embodiment of the present application is shown in FIG. 10;

[0080] Figure 12 A schematic diagram of the principle of the diffusion model provided by an embodiment of the present application is shown in FIG. 11;

[0081] Figure 13 A schematic diagram of the application of the multi-modal generative zooming model provided by an embodiment of the present application is shown in FIG. 12;

[0082] Figure 14 A schematic diagram of the network architecture of the multi-modal generative zooming model provided by an embodiment of the present application is shown in FIG. 13;

[0083] Figure 15 A schematic diagram of the training process of the multi-modal generative zooming model provided by an embodiment of the present application is shown in FIG. 14;

[0084] Figure 16 A schematic diagram of the hardware system of the terminal device provided by an embodiment of the present application is shown in FIG. 15;

[0085] Figure 17 Fig. 1 is a schematic diagram of a software system of a terminal device according to an embodiment of the present application. DETAILED DESCRIPTION

[0086] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application. In the description of the embodiments of the present application, unless otherwise specified, " / " represents the meaning of or, for example, A / B can represent A or B; in this document, "and / or" only describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which means that there are three cases of A alone, A and B together, and B alone. In addition, in the description of the embodiments of the present application, "multiple" means two or more than two.

[0087] Hereinafter, the terms "first", "second", "third" are only for description purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first", "second", "third" can explicitly or implicitly include one or more features.

[0088] The image processing method provided by the embodiments of the present application can be applied to the scene of super-resolution magnification processing of images. In some embodiments, the image processing method provided by the embodiments of the present application can be applied to the scene of photographing, video recording, and screen capture, etc. which need to perform super-resolution magnification processing on images.

[0089] The image processing method provided by the embodiments of the present application can be applied to a terminal device. The terminal device can be a terminal device with display hardware and corresponding software support. For example, the terminal device can be a mobile phone, a folding screen, a smart screen, a tablet computer, a wearable terminal device, a vehicle-mounted terminal device, an augmented reality (AR) device, a virtual reality (VR) device, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), a home device, a projector, etc. The embodiments of the present application do not make any limitation on the specific type of the terminal device.

[0090] The specific implementation process of the current image processing method will be described in detail below in combination with Figure 1 and Figure 2 .

[0091] Please refer to Figure 1 , Figure 1 Fig. 1 is a schematic diagram of an application scene of an image processing method.

[0092] AsFigure 1 As shown in the figure, when the user uses the terminal device 100 to shoot the image of the pine tree, in response to the operation of the user adjusting the magnification (for example, the magnification is 60 times), the terminal device 100 can display the preview image after the partial details of the pine tree are magnified. Further, in response to the photographing operation of the user, the terminal device 100 can obtain the magnified pine tree image (for example, the pine tree image with the magnification of 60 times pointed by the arrow in the middle). Figure 1 As shown in the figure, when the user uses the terminal device 100 to shoot the image of the pine tree, in response to the operation of the user adjusting the magnification (for example, the magnification is 60 times), the terminal device 100 can display the preview image after the partial details of the pine tree are magnified. Further, in response to the photographing operation of the user, the terminal device 100 can obtain the magnified pine tree image (for example, the pine tree image with the magnification of 60 times pointed by the arrow in the middle).

[0093] It should be noted that, Figure 1 The structure of the pine tree shown in the figure is a schematic structure, not the structure of the pine tree in actual shooting, and the complete application scenario can be embodied by the schematic pine tree and the terminal device 100.

[0094] It should be noted that, Figure 1 The pine tree image with the magnification of 60 times pointed by the arrow in the middle is a kind of magnification effect diagram obtained by actual shooting, and the effect of the terminal device 100 magnifying the partial details of the pine tree can be embodied by the kind of magnification effect diagram obtained by actual shooting.

[0095] The terminal device 100 can realize the above image processing process through a plurality of software modules. Figure 2 The principle schematic block diagram of the current image processing method is shown. As shown in the figure, Figure 2 The plurality of software modules can include a collection module 10, a preprocessing module 11, a magnification module 12, and a post-processing module 13.

[0096] The terminal device 100 can collect the to-be-shot image through the collection module 10. The to-be-shot image is an original image without any processing.

[0097] Exemplarily, the collection module 10 can be a camera (including a long-focus lens and a short-focus lens, etc.) installed on the terminal device 100. The camera can transmit the content to be shot to the subsequent preprocessing module 11 in the form of a digital image by converting the optical signal into an electrical signal, providing a data basis for subsequent preprocessing.

[0098] After obtaining the to-be-shot image, the collection module 10 can input the to-be-shot image into the preprocessing module 11 for preprocessing (or pre-processing), to obtain the pre-processed to-be-shot image. The pre-processed to-be-shot image is a preview image before the user adjusts the magnification (for example, the image of the pine tree shown in the figure). Figure 1

[0099] ​In the case where the pre-processed to-be-shot image is obtained, in response to an operation of the user adjusting the magnification (for example, the user selects a magnification greater than the current magnification in the terminal device 100), the terminal device 100 can perform magnification processing on part of the details of the pre-processed to-be-shot image by the magnification module 12 to obtain a magnification-processed to-be-shot image. The magnification-processed to-be-shot image is the preview image after the user adjusts the magnification.

[0100] Optionally, in some embodiments, the magnification module 12 can perform super-resolution magnification on the pre-processed to-be-shot image by using an image interpolation algorithm (for example, nearest neighbor interpolation or bilinear interpolation, etc.).

[0101] It should be noted that the magnification module 12 can also perform super-resolution magnification on the pre-processed to-be-shot image by using a generative adversarial algorithm, which is not limited in the embodiments of the present application.

[0102] In the case where the magnification-processed to-be-shot image is obtained, in response to a shooting operation of the user, the terminal device 100 can input the magnification-processed to-be-shot image into the post-processing module 13 for post-processing by the magnification module 12 to obtain a post-processed image. The post-processed image is the image displayed to the user by the terminal device 100 (for example, Figure 1 The magnification of the pine tree image pointed by the arrow in the middle is 60 times.

[0103] However, the details of the image obtained by the above image processing method are lost seriously, and the definition of the image is poor (for example, Figure 1 The magnification of the pine tree image pointed by the arrow in the middle is 60 times.

[0104] Therefore, the embodiments of the present application propose an image processing method, which obtains a first image with low resolution, and obtains a noise image. Then, the text description information of the first image can be generated according to the first image. Further, the obtained first image, text description information and noise image are input into a multi-modal generative magnification model, and the multi-modal generative magnification model can complete the super-resolution magnification processing of the first image based on the information of the three modalities of the first image, text description information and noise image, and output a second image with high resolution. Thus, by inputting information of multiple modalities, the multi-modal generative magnification model can accurately generate the detail information of the first image under the guidance of the text description information, so that the detail information in the second image after super-resolution magnification processing is clearer, and the quality of the second image after super-resolution magnification processing is further improved.

[0105] The following will be described in combination withFigures 3 to 5 The embodiment of the present application provides a kind of specific implementation process of image processing method.

[0106] Please refer to Figure 3 , Figure 3 The application scenario schematic diagram of the image processing method provided by the embodiment of the present application is shown.

[0107] In the case where the user uses the terminal device 100, the terminal device 100 can display the interface 11 as shown in (1) of Figure 3 As shown in (1) of Figure 3 A plurality of application programs are installed in the terminal device 100. The plurality of application programs can include a camera 101, a phonebook, a telephone and a photo album.

[0108] As shown in (1) of Figure 3 In response to the triggering operation (such as single click, multiple clicks or long press operation) of the user on the camera 101 in the interface 11, the terminal device 100 can display the internal interface of the application program of the camera 101.

[0109] As shown in (2) of Figure 3 The interface 12 displayed by the terminal device 100 is an internal interface of the camera 101. The interface 12 can include a plurality of shooting modes of the camera 101. The plurality of shooting modes can include an aperture mode, a night scene mode, a portrait mode, a photographing mode, a video recording mode and a professional mode. The professional mode is set as a mode without adding any algorithm by default.

[0110] It should be noted that the image effects obtained by different shooting modes are different, and in the actual use process, the user can manually switch different shooting modes to obtain different image effects, and the embodiment of the present application is not limited to this.

[0111] Optionally, the image processing method of the embodiment of the present application can be integrated in the camera 101. Alternatively, the image processing method of the embodiment of the present application can be integrated in any one of the plurality of shooting modes. For example, the image processing method of the embodiment of the present application can be integrated in the photographing mode. Considering that the interface of the camera 101 in the terminal device 100 when it is turned on is usually the interface in the photographing mode (for example, the interface 12), therefore, the image processing algorithm of the embodiment of the present application is integrated in the photographing mode, the user only needs to click the control 106, without the need to switch the mode, to obtain the image processed by the image processing algorithm of the embodiment of the present application, which is more convenient for user operation and can increase user experience.

[0112] It should be noted that, in practical applications, those skilled in the art can also integrate the image processing algorithm of this application embodiment into other shooting modes (such as video recording mode) besides the photo shooting mode, according to actual usage needs. This application embodiment does not limit this. For example, when the terminal device 100 integrates the image processing method of this application embodiment into the video recording mode, during the recording process, the image captured during the recording process can be subjected to generative super-resolution magnification processing to obtain an image with higher resolution and clearer details.

[0113] like Figure 3 As shown in Figure (2), the interface 12 may also include a viewfinder 102, controls 103, 104, 105, and 106. In response to the user's triggering operation on different controls in the interface 12, the terminal device 100 can perform different operations.

[0114] Optionally, the viewfinder 102 is used to display the image to be captured (e.g., ...). Figure 3 (2) The image of the pine tree shown in Figure 2) or the video to be recorded, etc. For example, in the photo mode of camera 101, the viewfinder 102 is used to display a preview image of the pine tree to be photographed. In the video recording mode of camera 101, the viewfinder 102 is used to display a video of the pine tree to be recorded.

[0115] It should be noted that, Figure 3 The pine tree in the preview image of the pine tree to be photographed shown in Figure (2) is a schematic structure, intended to fully represent the preview image currently displayed in viewfinder 102. Figure 3 As shown in Figure (2), the interface 12 can also display the optical magnification of the lens in the terminal device 100. The optical magnification of a lens refers to the ability of the lens to magnify the observed object on the imaging sensor or film. The optical magnification of a lens is a fixed value, which is determined by the physical characteristics of the lens and cannot be changed by the user.

[0116] Optionally, the optical magnification of the lens may include the optical magnification of a short-focal-length lens and the optical magnification of a telephoto lens. For example, the optical magnification of the short-focal-length lens is 1, and the optical magnification of the telephoto lens is 3.

[0117] Optionally, control 103 is used to display the optical magnification of the short-focal-length lens in terminal device 100. For example, as shown... Figure 3 As shown in Figure (2), the value 1 can be displayed in control 103.

[0118] Optionally, control 104 is used to display the optical magnification of the telephoto lens in terminal device 100. For example, as shown... Figure 3As shown in (2) of FIG. 1, the control 104 can display the value 3.

[0119] It is considered that the optical magnification of the lens in the terminal device is a fixed value, and therefore, in actual use, the digital zoom magnification of the terminal device 100 needs to be combined to achieve the actual magnification desired by the user. The value of the digital zoom magnification of the terminal device 100 is a value that can be changed by the user through corresponding operation.

[0120] Optionally, the actual magnification desired by the user can be the product of the optical magnification of the lens and the digital zoom magnification of the terminal device 100. When the terminal device 100 detects the actual magnification desired by the user, the terminal device 100 can perform corresponding magnification processing on the image to be captured displayed in the viewfinder 102.

[0121] It should be noted that, in actual application, in order to ensure the quality of the magnified image, the digital zoom magnification of the terminal device 100 is usually set to have an upper limit value (for example, the upper limit value is 100 or the upper limit value is 200) stored in the memory of the terminal device 100. In the case where the quality of the image is not required, the digital zoom magnification of the terminal device 100 can also not be set to have an upper limit value, which is not limited in the embodiments of the present application.

[0122] It should be further noted that, since the actual magnification desired by the user can be the product of the optical magnification of the lens and the digital zoom magnification of the terminal device 100, the actual magnification desired by the user is usually displayed in the terminal device 100, so that the user can directly see the value of the current magnification, and the upper limit value of the digital zoom magnification of the terminal device 100 is usually not displayed in the interface.

[0123] In the case where the actual magnification desired by the user is greater than the optical magnification of the short-focus lens and less than the optical magnification of the long-focus lens, the actual magnification desired by the user is the product of the optical magnification of the short-focus lens and the digital zoom magnification of the terminal device 100.

[0124] For example, the actual magnification desired by the user is 1.5 times, the optical magnification of the short-focus lens is 1 times, that is, the digital zoom magnification of the terminal device 100 is 1.5 times, and at this time, the actual magnification desired by the user is the product of 1 and 1.5.

[0125] Optionally, in some embodiments, in the case where the actual magnification desired by the user is greater than the optical magnification of the short-focus lens and less than the optical magnification of the long-focus lens, in response to the operation of the user adjusting the magnification in the interface 12 (for example, Figure 3The terminal device 100 can display the change process of the actual zoom ratio expected by the user in the control 103, for example, the value displayed in the control 103 gradually increases from 1 to 1.5, and then gradually increases from 1.5 to 2.9 (as shown in (2) of FIG. 1). Figure 3 The terminal device 100 can display the change process of the actual zoom ratio expected by the user in the control 103, for example, the value displayed in the control 103 gradually increases from 1 to 1.5, and then gradually increases from 1.5 to 2.9 (as shown in (2) of FIG. 1).

[0126] It should be noted that the pinch operation can refer to the user performing a sliding operation on the screen with two fingers, so that the relative distance between the two fingers gradually increases or decreases, and finally realizes the zoom processing of the operation object.

[0127] It should be further noted that in response to the user adjusting the zoom ratio in the interface 12, the control 103 can also not display the change process of the actual zoom ratio expected by the user, and the embodiments of the present application are not limited in this regard.

[0128] In response to the user adjusting the zoom ratio in the interface 12, in the case where the actual zoom ratio expected by the user is equal to the optical zoom ratio of the long-focus lens, the terminal device 100 can switch from the short-focus lens to the long-focus lens for shooting. At this time, the actual zoom ratio expected by the user displayed in the terminal device 100 can be switched from the optical zoom ratio of the short-focus lens to the optical zoom ratio of the long-focus lens. For example, Figure 3 The current actual zoom ratio shown in (2) of FIG. 1 is switched from the value 1 displayed in the control 103 to the value 3 displayed in the control 104.

[0129] It should be noted that in addition to the above-mentioned switching of the terminal device 100 from the short-focus lens to the long-focus lens for shooting by the user adjusting the zoom ratio, in response to the user performing a leftward sliding operation or a rightward sliding operation on the display screen of the terminal device 100, the terminal device 100 can also switch from the short-focus lens to the long-focus lens for shooting, and the embodiments of the present application are not limited in this regard.

[0130] It should be further noted that with the user adjusting the zoom ratio, the image to be shot displayed in the viewfinder 102 will also be subjected to corresponding zoom processing. For example, when the user adjusts the zoom ratio to 1.5 times, the image to be shot displayed in the viewfinder 102 will be subjected to 1.5 times zoom processing.

[0131] Optionally, the control 105 is used to display the photographed image (for example, Figure 3 The person in the image shown in (2) of FIG. 1) or the recorded video. In response to the user triggering the control 105, the terminal device 100 can display the photographed image or display the recorded video.

[0132] Optionally, the control 106 is configured to trigger starting photographing or starting recording. In response to the user triggering the control 106, the terminal device 100 can start photographing an image or start recording a video.

[0133] In the case where the actual magnification desired by the user is greater than the optical magnification of the telephoto lens, as shown in (2) of FIG. 1, in response to the user adjusting the magnification in the interface 12 (for example, the user performs a pinch operation of sliding the two fingers on the screen so that the relative distance between the two fingers gradually increases), the terminal device 100 can perform corresponding magnification processing on the image to be photographed displayed in the viewfinder 102. Figure 3

[0134] In the case where the actual magnification desired by the user is greater than the optical magnification of the telephoto lens, the actual magnification desired by the user is the product of the optical magnification of the telephoto lens and the digital zoom magnification of the terminal device 100.

[0135] Exemplarily, as shown in (3) of FIG. 1, the actual magnification desired by the user is 60 times, the optical magnification of the telephoto lens is 3 times, that is, the digital zoom magnification of the terminal device 100 is 20 times, at this time, the actual magnification desired by the user is the product of 3 and 20. Figure 3

[0136] As shown in (3) of FIG. 1, the interface 13 includes a control 107. The control 107 is configured to display the current actual magnification greater than the optical magnification of the telephoto lens. Exemplarily, the control 107 displays the numerical value 60. Figure 3

[0137] It should be noted that the embodiment of the present application takes the control 107 to display the current actual magnification greater than the optical magnification of the telephoto lens as an example for description, and the embodiment of the present application is not limited thereto, for example, the terminal device 100 can also directly display the current actual magnification greater than the optical magnification of the telephoto lens through the control 104.

[0138] As shown in (3) of FIG. 1, the viewfinder 102 of the interface 13 displays the preview image after magnification corresponding to the current actual magnification in the control 107. Exemplarily, the viewfinder 102 of the interface 13 displays the preview image of the pine tree after magnification of 60 times. Figure 3 It should be noted that the pine tree included in the preview image of the pine tree after magnification of 60 times shown in (3) of FIG. 1 is a schematic structure, and the purpose is to fully embody the preview image currently displayed in the viewfinder 102.

[0139] Figure 3 As shown in (3) of FIG. 1, the viewfinder 102 of the interface 13 displays the preview image after magnification corresponding to the current actual magnification in the control 107. Exemplarily, the viewfinder 102 of the interface 13 displays the preview image of the pine tree after magnification of 60 times.

[0140] As shown in (3) of FIG. 1, the viewfinder 102 of the interface 13 displays the preview image after magnification corresponding to the current actual magnification in the control 107. Exemplarily, the viewfinder 102 of the interface 13 displays the preview image of the pine tree after magnification of 60 times. Figure 4 ​​​​As shown in Figure (3), in response to the user's trigger operation on control 106, terminal device 100 can display as shown in Figure (3). Figure 4 The interface 14 is shown in Figure (4). At this time, the control 105 of the interface 14 can display an image of the pine tree magnified 60 times.

[0141] It should be noted that, Figure 4 The pine tree image shown in Figure (4) of the control 105, which is magnified 60 times, is a schematic structure and is intended to fully represent the image currently displayed by the control 105.

[0142] like Figure 4 As shown in Figure (4), in response to a user's triggering operation on control 105, terminal device 100 can display an image of a pine tree magnified 60 times. For example, terminal device 100 can display... Figure 4 The image of the pine tree shown in interface 15 is magnified 60 times.

[0143] like Figure 1 As shown, interface 15 may include controls 108 and 109, etc. Control 108 is used to return the stored image of the pine tree displayed after being magnified 60 times. Control 109 is used to edit the currently displayed image of the pine tree after being magnified 60 times in interface 15 (e.g., cropping the image size).

[0144] It should be noted that, apart from controls 108 and 109, Figure 4 The interface 15 shown may also include share controls, delete controls, etc. In response to user clicks on different controls, the terminal device 100 can perform corresponding processing or display corresponding interfaces. For example, in response to a user click on the delete control, the terminal device 100 can delete the image of a pine tree magnified 60 times from memory. Additionally, Figure 4 The interface 15 shown can also display information such as the time when the current image was captured, but this embodiment does not limit this.

[0145] Figure 3 The image of the pine tree shown in interface 15, magnified 60 times, is compared to... Figure 3 Compared to the pine tree image displayed with a magnification of 60x, the details are more complete, the image is clearer and of better quality, which can enhance the user's visual experience.

[0146] It should be noted that, Figure 3 The image of the pine tree magnified 60 times shown is a magnified effect obtained from an actual photograph. This magnified effect obtained from an actual photograph can demonstrate the effect of the terminal device 100 magnifying some details of the pine tree.

[0147] It should be noted that Figure 3 In the middle of interface 15, the image of the pine tree amplified by 60 times is displayed as an example, and the embodiments of the present application are not limited thereto. For example, the image of the pine tree amplified by 60 times can be displayed on interface 15 (or said to be displayed on the entire interface).

[0148] After the terminal device 100 completes the shooting, the image of the shot scenery 1 amplified by 60 times can be displayed in the control 105 of the interface 14. After the terminal device 100 stores the image of the scenery 1 amplified by 60 times, the image collected by the camera 101 in real time can be displayed in the viewfinder frame 102 of the interface 14 (not shown in (4) of FIG. 1). Figure 5

[0149] It should be understood that each of the above interfaces and the controls in each interface are exemplary and do not constitute a limitation on the interfaces in the embodiments of the present application. In another embodiment of the present application, based on different application scenarios, each interface can include more or fewer controls than the controls in the above examples, and the positions of each control can be adjusted accordingly, which is not limited in the present application. For example, Figure 5 In addition to being displayed on the right side of the viewfinder frame 102, the controls 103 and 104 shown in (2) of FIG. 1 can also be displayed at the bottom of the viewfinder frame 102. For example, Figure 5 As shown in (4) of FIG. 1, after the shooting is completed, the control 107 will not be displayed in the interface.

[0150] It should also be understood that the above descriptions related to the direction (such as right, left, and bottom, etc.) are based on the direction shown in Figure 3 When the direction of the terminal device 100 changes, the direction and position of each control can be automatically adjusted accordingly, and the embodiments of the present application are not limited thereto.

[0151] Based on the above description, in the case where the terminal device includes multiple software modules, the image processing method of the embodiments of the present application can be implemented by multiple software modules.

[0152] Please refer to Figure 3 , Figure 3 A schematic block diagram of the principle of an image processing method provided by an embodiment of the present application is shown.

[0153] As shown in Figure 3 The terminal device 100 can include an acquisition module 301, a preprocessing module 302, a multi-modal generative amplification model 303, a post-processing module 304, and a language conversion module 305.

[0154] When the user adjusts the magnification (for example, as shown in Figure 3 ​before the zooming operation (e.g., the zooming operation shown in (2) of FIG. 1) is performed, the terminal device 100 can acquire a preview image (e.g., the preview image displayed in the viewfinder 102 shown in (1) of FIG. 1) without adjusting the zoom ratio by the acquisition module 301. Figure 4

[0155] In response to the user's operation of adjusting the zoom ratio, the terminal device 100 can acquire a preview image (e.g., the preview image displayed in the viewfinder 102 shown in (3) of FIG. 1) after the zoom ratio is adjusted by the acquisition module 301. Figure 2

[0156] In the case where the preview image after the zoom ratio is adjusted is obtained, in response to the user's photographing operation (e.g., the user's triggering operation on the control 106 shown in (3) of FIG. 1), the acquisition module 301 can input the preview image after the zoom ratio is adjusted into the preprocessing module 302 for preprocessing to obtain a preprocessed image. Figures 3 to 6

[0157] Optionally, in some embodiments, the preprocessing is used to eliminate noise of the preview image after the zoom ratio is adjusted, improve contrast of the preview image, and correct distortion of the preview image, etc. The preprocessing can include one or more of the following: denoising processing, automatic white balance processing, demosaicing processing, image contrast processing, histogram equalization processing, and image enhancement processing.

[0158] By preprocessing the preview image after the zoom ratio is adjusted, the terminal device 100 can improve the quality of the preview image after the zoom ratio is adjusted (e.g., improve the brightness of the preview image after the zoom ratio is adjusted and improve the contrast of the preview image after the zoom ratio is adjusted, etc.), thereby providing a good data basis for subsequent generative super-resolution magnification processing and post-processing.

[0159] Optionally, in some embodiments, in the case where the preprocessed image is obtained, the terminal device 100 can send the preprocessed image to the multi-modal generative magnification model 303 and the language-to-text module 305 respectively by the preprocessing module 302.

[0160] Further, the terminal device 100 can generate text description information corresponding to the preprocessed image based on the preprocessed image by the language-to-text module 305.

[0161] Optionally, in some embodiments, after the text description information corresponding to the preprocessed image is obtained, the terminal device 100 can send the text description information corresponding to the preprocessed image to the multi-modal generative magnification model 303 by the language-to-text module 305.

[0162] ​​​It should be noted that the multi-modal generative magnification model 303 and the language conversion module 305 are independently set in the embodiments of the present application, and the terminal device 100 can also integrate the multi-modal generative magnification model 303 and the language conversion module 305 into one module to implement the image processing method of the embodiments of the present application.

[0163] In addition, the terminal device 100 can also acquire the noise image pre-stored in the terminal device 100 through the acquisition module 306, and send the acquired noise image to the multi-modal generative magnification model 303.

[0164] Optionally, in some embodiments, after obtaining the pre-processed image, the text description information corresponding to the pre-processed image, and the noise image, the terminal device 100 can perform generative super-resolution magnification processing on the pre-processed image based on the input pre-processed image and the text description information corresponding to the pre-processed image and the noise image through the multi-modal generative magnification model 303, to obtain the generative super-resolution magnification processed image. That is, under the guidance and constraint of the text description information and the pre-processed image, the details in the pre-processed image are generated, so that the details of the generative super-resolution magnification processed image are restored to a higher degree, the image is clearer, and the user's visual experience can be improved.

[0165] Optionally, in some embodiments, in the case of obtaining the generative super-resolution magnification processed image, the terminal device 100 can send the generative super-resolution magnification processed image to the post-processing module 304 through the multi-modal generative magnification model 303. Thus, the terminal device 100 can perform post-processing on the generative super-resolution magnification processed image through the post-processing module 304 to obtain an image after post-processing (such as an image displayed in the control 105 shown in (4) of FIG. 1, or an image shown in FIG. 2). Figure 6 Figure 6 By post-processing the generative super-resolution magnification processed image, the clarity of the image can be further improved, thereby further improving the user's visual experience.

[0166] Optionally, in some embodiments, the post-processing is used to further optimize the quality of the generative super-resolution magnification processed image. The terminal device 100 can perform post-processing on the generative super-resolution magnification processed image, so that the generative super-resolution magnification processed image can achieve a specific visual effect and be displayed to the user.

[0167] ​Optionally, the post-processing can include one or more of the following processes: sharpening processing, compression processing, and format conversion processing. The sharpening processing is used to increase the edges and details of the image. The compression processing is used to reduce the storage space occupied by the image without changing the definition of the image. The format conversion processing is used to convert the image into a format such as jpg and store it in the memory of the terminal device 100 to facilitate the user to view the captured image.

[0168] It should be noted that in response to the photographing operation of the user, the terminal device 100 can directly display the image after the generative super-resolution enlargement processing, i.e., the image after the post-processing, without displaying the image output by the pre-processing module 302 or the image output by the multi-modal generative enlargement model 303, thereby making the user experience better.

[0169] As can be seen from the above, in the image processing method in the embodiments of the present application, the pre-processed image, the text description information corresponding to the pre-processed image, and the noise image are input into the multi-modal generative enlargement model 303, which can make the multi-modal generative enlargement model 303 accurately generate the detail information in the pre-processed image under the guidance and constraint of the pre-processed image and the text description information corresponding to the pre-processed image, obtain the image after the generative super-resolution enlargement processing, make the detail information of the image after the generative super-resolution enlargement processing clearer, and improve the quality of the image after the generative super-resolution enlargement processing. Compared with the simple enlargement processing of the enlargement module 12 in the related art Figure 6 , the enlargement effect is better, the image obtained by the terminal device 100 is clearer, and the visual experience of the user can be improved.

[0170] Based on the description of Figure 6 , the specific implementation process of the image processing method provided in the embodiments of the present application will be described in detail below in combination with Figure 5 .

[0171] Please refer to Figure 5 , Figure 3 , which shows the flowchart of the image processing method provided in the embodiments of the present application.

[0172] As shown in Figure 5 , the image processing method in the embodiments of the present application includes the following steps:

[0173] S101, obtaining a first image.

[0174] Optionally, the first image is a low-resolution image to be subjected to super-resolution enlargement. As shown in Figure 3 , the first image is equivalent to Figure 3The pre-processing module 302 outputs the image. Illustratively, in a photographing scenario, a first image after pre-processing can be acquired in response to a photographing operation of a user.

[0175] Optionally, the resolution of the first image is a first resolution. Illustratively, the first resolution can be 512 pixels * 512 pixels.

[0176] Optionally, in some embodiments, the terminal device can display a first preview image at a first zoom ratio in response to an operation of the user starting the camera (such as a trigger operation of the user on the camera 101 as shown in (1) of FIG. 1). By displaying the first preview image, a data basis can be provided for subsequent acquisition of the first image. Figure 3

[0177] Optionally, the first preview image is a raw low-resolution image acquired by the terminal device. As shown in (2) of FIG. 1, the first preview image is equivalent to a preview image acquired by the acquisition module 301 without adjustment of the magnification. Illustratively, the first preview image can be a preview image displayed in the viewfinder 102 as shown in (2) of FIG. 1. Figure 5 Figure 5

[0178] It should be noted that the terminal device can also store the first preview image, and subsequent processing can be performed based on the stored first preview image, which is not limited in the embodiments of the present application.

[0179] Optionally, the first zoom ratio is an actual magnification of the terminal device without adjustment. Illustratively, the first zoom ratio can be an optical magnification of a short-focus lens displayed by the control 103 as shown in (2) of FIG. 1. Figure 3

[0180] It should be noted that the terminal device can also set the first zoom ratio to an optical magnification of a long-focus lens or other magnification (such as a specific numerical value of a digital zoom magnification) according to actual use requirements, which is not limited in the embodiments of the present application.

[0181] When the first zoom ratio is adjusted to a second zoom ratio, the terminal device can display a second preview image. By displaying the second preview image, a data basis can be further provided for subsequent acquisition of the first image.

[0182] Optionally, the second zoom ratio is an actual magnification of the terminal device after adjustment. The second zoom ratio is greater than the first zoom ratio. Illustratively, the second zoom ratio can be a current actual magnification displayed by the control 107 as shown in (3) of FIG. 1. Figure 3

[0183] Optionally, the second preview image is an image after the terminal device adjusts the magnification. As shown in (3) of FIG. 1, the second preview image is equivalent to an image after the acquisition module 301 adjusts the magnification. Illustratively, the second preview image can be a preview image displayed in the viewfinder 102 as shown in (3) of FIG. 1.​​​​​Figure 5 As shown in FIG. 2, the second preview image corresponds to the preview image after the zooming magnification is adjusted by the zooming magnification adjustment module 201. Figure 3 As shown in FIG. 2, the second preview image corresponds to the preview image after the zooming magnification is adjusted by the zooming magnification adjustment module 201. Figure 5 As shown in FIG. 2, the second preview image corresponds to the preview image after the zooming magnification is adjusted by the zooming magnification adjustment module 201.

[0184] In some embodiments, the terminal device can adjust the first zooming magnification to the second zooming magnification in response to the user's operation of adjusting the zooming magnification, such as the pinch operation shown in FIG. 2(2). Figure 5 In some embodiments, the terminal device can adjust the first zooming magnification to the second zooming magnification in response to the user's operation of adjusting the zooming magnification, such as the pinch operation shown in FIG. 2(2).

[0185] In another embodiment, the terminal device can also adjust the first zooming magnification to the second zooming magnification in response to the user's operation of switching the shooting mode.

[0186] It should be noted that in addition to the above two ways, the terminal device can also automatically adjust the first zooming magnification to the second zooming magnification by detecting the distance between the content to be shot (such as scenery, buildings, or people, etc.) and the terminal device, and when it is detected that the distance between the content to be shot and the terminal device exceeds the processing distance of the short-focus lens, the terminal device automatically adjusts the first zooming magnification to the second zooming magnification, which is not limited in the embodiments of the present application.

[0187] Optionally, in some embodiments, the terminal device can pre-process the second preview image (equivalent to the processing performed by the pre-processing module 302 shown in FIG. 2) to obtain the first image when receiving the shooting instruction. Figure 5 By pre-processing the second preview image, a first image with better image brightness and image contrast can be obtained, which can provide a good data basis for subsequent generative super-resolution magnification processing.

[0188] As an example, the terminal device can receive the shooting instruction in response to the user's shooting operation (such as the user's triggering operation on the control 106 shown in FIG. 2(3)). Figure 5 It should be further noted that the terminal device can obtain the first image through various possible implementation manners, which is not limited in the embodiments of the present application. For example, the terminal device can obtain the first image pre-stored in the memory of the terminal device. For another example, the terminal device can obtain the first image by intercepting the image during the video recording process.

[0189] In summary, the terminal device can obtain the first image, which can provide a good data basis for subsequent processing of the multi-modal generative magnification model.

[0190] S102, generating text description information of the first image based on the first image.

[0191] S102, generating text description information of the first image based on the first image.

[0192] Optionally, in some embodiments, the text description information is used to describe the first image. The text description information includes at least first type information and second type information. The first type information can be system information related to the first image obtained by the terminal device. The system information related to the first image refers to attribute information (such as time attribute and place attribute, etc.) other than the content contained in the first image. The second type information can be high-level semantic feature information of the first image (such as color information and texture information of the content contained in the first image, etc.).

[0193] It should be noted that the language type of the text description information is not limited in the embodiments of the present application. The voice type can depend on the user's settings. The text description information can be described in languages including but not limited to Chinese, English, Korean, Russian, and Japanese, etc. For example, the text description information can be "pine trees in Huangshan, Anhui on May 3".

[0194] Optionally, the first type information can include a time corresponding to the first image and / or a place corresponding to the first image. The time corresponding to the first image can be in one or more of the following forms: date (such as 20230503), day of the week (such as Thursday), season (such as spring), and time of day (such as 8:00), etc.

[0195] It should be noted that the time corresponding to the first image can be the shooting time of the first image, the time when the user obtains the first image by screenshot, the time when the user saves the first image after editing, or the current display time of the terminal device 100, etc. The embodiments of the present application are not limited thereto.

[0196] Optionally, the place corresponding to the first image can be in one or more of the following forms: geographic coordinates (such as longitude 122°18′, latitude 36°57′), administrative region (such as Beijing), landmark and building (such as the Forbidden City), electronic fence and position marked on the map, etc. For example, in the shooting scenario, the place corresponding to the first image can be the shooting place.

[0197] It should be noted that the place corresponding to the first image can be determined by the Beidou navigation system or the global positioning system (GPS), and the embodiments of the present application are not limited thereto. For example, a variety of sensors (including electronic compass, accelerometer, and gyroscope, etc.) combined with wireless network signals can determine the place corresponding to the first image.

[0198] Optionally, the second type of information is used to represent a category to which the content included in the first image belongs. Optionally, the category to which the content included in the first image belongs can include plants and trees (such as pine trees, peony flowers, camphor trees, and azalea flowers, etc.), buildings (such as the Forbidden City and the Big Goose Pagoda, etc.), animals (cats and dogs, etc.), figures (such as children and old people, etc.), and vehicles (such as cars, trucks, vans, and bicycles, etc.), etc.

[0199] After S101 is performed, the terminal device can obtain the first image. After the first image is obtained, the terminal device can generate the text description information of the first image based on the system information related to the first image and the high-order semantic feature information of the first image (equivalent to the processing performed by the translation language text module 305 shown in the figure). By generating the text description information of the first image, a data basis can be provided for subsequent processing of the multi-modal generative amplification model. Figure 5

[0200] S103, obtaining a noise image.

[0201] It should be noted that there is no time sequence and order between S102 and S103, that is, S102 and S103 can be executed simultaneously, or can be executed in sequence, and the specific execution sequence can depend on the actual application scenario. For example, S102 is executed first and S103 is executed second. For example, S102 is executed second and S103 is executed first. The embodiments of the present application do not limit this.

[0202] Optionally, in some embodiments, the noise image is used to provide noise distribution characteristics for the first image. The noise image can be an image pre-stored in the terminal device. Exemplarily, the noise image can be a random Gaussian noise image or a Gaussian white noise image. The resolution of the noise image can be 512 pixels * 512 pixels.

[0203] Optionally, in another embodiment, the noise image can also be an image generated by the terminal device according to pre-set software code, which is not limited by the embodiments of the present application.

[0204] Optionally, in some embodiments, the terminal device can obtain the pre-stored noise image from the memory of the terminal device through the obtaining module 306 as shown in the figure. Figure 3 By obtaining the noise image, a data basis can be provided for subsequent processing of the multi-modal generative amplification model, which can ensure the normal operation of the multi-modal generative model, and can increase the diversity of the multi-modal generative amplification model for the generative super-resolution processing of the first image, so that the first image generated by the multi-modal generative amplification model has more diverse detailed information (such as the length and direction of the branches of the pine tree, etc.).

[0205] ​S104, input the first image, the text description information, and the noise image into the multi-modal generative magnification model, and output a second image.

[0206] Optionally, the multi-modal generative magnification model (as shown in Figure 3 ) is a pre-trained model stored in the terminal device. Multi-modal refers to two or more than two different types of imaging technologies or data sources (for example, the three data sources of the first image, the text description information, and the noise image).

[0207] Optionally, in some embodiments, the resolution of the second image is a second resolution, and the second resolution is greater than the first resolution. Exemplarily, the second resolution can be 4096 pixels * 4096 pixels.

[0208] After performing S101 to S103, the terminal device can input the obtained first image, text description information, and noise image into the pre-trained multi-modal generative magnification model for processing, and output the processed second image (equivalent to the processing performed by the multi-modal generative magnification model 303 shown in Figure 4 ). Thus, the terminal device can obtain the second image under the second zoom ratio.

[0209] It should be noted that the multi-modal generative magnification model in the terminal device can use the same noise image to perform generative super-resolution magnification processing on the first image containing different contents, or use different noise images to perform generative super-resolution magnification processing on the first image containing different contents. The embodiments of the present application do not limit this.

[0210] Optionally, in some embodiments, in response to the user's photographing operation, the terminal device can display the second image. Before displaying the second image, the terminal device can also perform post-processing on the second image (equivalent to the processing performed by the post-processing module 304 shown in Figures 7 to 10 ). Thus, the terminal device can obtain the second image after post-processing. The second image after post-processing has a higher degree of detail restoration and better quality, which can improve the user's visual experience.

[0211] Optionally, in some embodiments, the terminal device can store the second image after post-processing. Exemplarily, the second image after post-processing can be the image of the pine tree after magnification by 60 times in the control 105 as shown in (4) of Figure 7 . As shown in (4) of Figure 5 , in response to the user's triggering operation on the control 105 in the interface 14, the terminal device can display the image as shown in Figure 7 , which is convenient for the user to view.

[0212] It should be noted that the image processing method of the embodiment of the present application is a processing method between the time when the camera of the terminal device starts to shoot an image and the time when the shooting of the image is completed. Moreover, there is a certain delay between the time when the camera of the terminal device starts to shoot an image and the time when the shooting of the image is completed, so the first image obtained by the terminal device is only a data basis for the super-resolution enlargement processing of the terminal device, and is not the final image shown to the user. The image finally shown to the user by the terminal device is the second image after post-processing, so as to enhance the experience of the user.

[0213] It should also be noted that the image processing method of the embodiment of the present application can be used in the shooting mode of the camera 101 to obtain a clear second image. When the user needs to compare effects, the shooting mode can be switched to the professional mode for shooting. Since the image processing method of the embodiment of the present application is not used for processing in the professional mode, the obtained image is relatively blurred.

[0214] The image processing method in the embodiment of the present application can obtain the image to be processed by obtaining the first image, and provide a data basis for subsequent super-resolution enlargement processing. Then, based on the obtained first image, the text description information of the first image is generated, which can provide a data basis for subsequent generative super-resolution enlargement processing. The noise image is obtained, which can provide noise distribution characteristics for the first image. Further, after obtaining the first image, the text description information and the noise image, the first image, the text description information and the noise image can be input into the multi-modal generative enlargement model. The multi-modal generative enlargement model can perform generative super-resolution enlargement processing on the first image based on the input data of the three modalities of the first image, the text description information and the noise image, and can better generate the detail information of the enlarged image, and obtain a second image with higher detail restoration degree. That is, after the multi-modal generative enlargement model performs generative super-resolution enlargement processing on the first image under the guidance and constraint of the first image and the text description information, a higher detail restoration degree image is obtained. Thus, the multi-modal generative enlargement model can be used to perform generative super-resolution enlargement processing on a low-resolution (such as a first resolution) image (such as a first image), and output a high-resolution (such as a second resolution) image (such as a second image). Moreover, the high-resolution image has higher detail restoration degree, higher image clarity and better image quality, and the user has better visual experience.

[0215] It should be noted that the multi-modal generative enlargement model in the image processing method of the embodiment of the present application can rely on the terminal device to process the image in real time, or can be implemented independently in the form of a model to realize offline processing of the image. For example, the first image, the text description information and the noise image are input into the independent multi-modal generative model, and the second image is directly output.

[0216] Based on the description of S102, the terminal device of the embodiments of the present application can generate text description information in a plurality of possible implementation manners. Next, a possible implementation manner of generating text description information in the embodiments of the present application is described in detail in combination with Figure 7 . Figure 7 The method can be applied to the language conversion module 305 shown in Figure 5 .

[0217] Please refer to Figure 8 , Figure 8 a flowchart of an image processing method provided by the embodiments of the present application is shown.

[0218] As shown in Figure 9 , the image processing method in the embodiments of the present application includes the following steps:

[0219] S201, obtaining first type information.

[0220] Optionally, in some embodiments, after performing S101, the terminal device 100 can obtain system information related to the first image through the language conversion module 305 as shown in Figure 5 , to obtain the first type information.

[0221] Considering the various forms of the first type information, after obtaining the first type information, the terminal device can perform conversion processing on the first type information through the language conversion module 305, and the form of the first type information after conversion processing is text. By converting the first type information in other forms (such as numbers or symbols, etc.) into first type information in text form, a data basis can be provided for the generation of text description information, so that the generated text description information is more accurate.

[0222] Optionally, in the case where the first type information includes the time corresponding to the first image, and the form of the time corresponding to the first image is not a text form, the terminal device can convert the time corresponding to the first image in other forms into text form time through an index algorithm (such as a hash algorithm or a bitmap index, etc.) in the language conversion module 305.

[0223] Exemplarily, the time corresponding to the first image can be the shooting date of the first image, and the shooting date of the first image is in the form of a number 20230503. Figure 10 a text conversion process of the time corresponding to the first image is shown. As Figure 11As shown, the terminal device matches the obtained digital representation of the shooting date of the first image 20230503 with the date literal description in the index list pre-stored in the terminal device through the language conversion literal module 305. The date literal description in the index list at least includes month and day. That is, the terminal device can match 05 in 20230503 with May in the index list and match 03 in 20230503 with 3 in the index list through the language conversion literal module 305, and obtain that the shooting date of the first image is May 3.

[0224] Optionally, in the case that the first type of information includes the place corresponding to the first image, and the representation of the place corresponding to the first image is not in the literal representation, the terminal device can convert the place corresponding to the first image in other representation into the place in literal representation through a coordinate-to-literal calling function in the language conversion literal module 305.

[0225] Exemplarily, the place corresponding to the first image can be geographical coordinates determined through the global positioning system in the terminal device. The geographical coordinates can be longitude 122°18′ and latitude 36°57′.

[0226] Optionally, the coordinate-to-literal calling function can be:

[0227] String location Address1=Location And Geocoder Util.get Instance(this).get Address By Location(location1);

[0228] Wherein, location Address1 is used to represent the place corresponding to the first image in literal representation; Location And Geocoder Util is a class used to process tasks related to geographical position and address coding; getInstance(this) means returning an instance of the Location And Geocoder Util class; get Address By Location(location1) is used to query and return an address string corresponding to the geographical coordinates of location1.

[0229] Exemplarily, when location1 is longitude 122°18′ and latitude 36°57′, the terminal device can obtain that the place corresponding to the first image is Huangshan Mountain in Anhui through the coordinate-to-literal calling function in the language conversion literal module to convert the longitude and latitude.

[0230] It should be noted that the above is described by taking the index algorithm and the coordinate conversion character calling function integrated in the language conversion character module 305 as an example, and the embodiments of the present application are not limited thereto. For example, the index algorithm and the coordinate conversion character calling function can also be independently set in the terminal device, and the language conversion character module 305 calls the index algorithm and the coordinate conversion character calling function set in the terminal device to realize corresponding processing.

[0231] In summary, the terminal device can obtain the first type of information. By obtaining the first type of information, a data basis can be provided for subsequent generation of text description information.

[0232] S202, input the first image into the classification model, and output the second type of information.

[0233] It should be noted that there is no time sequence and order between S201 and S202. That is, S201 and S202 can be executed simultaneously, or S201 can be executed first and S202 can be executed second, or S201 can be executed second and S202 can be executed first, and the embodiments of the present application do not limit this.

[0234] Optionally, in some embodiments, the language conversion character module 305 includes a pre-trained classification model.

[0235] The terminal device can input the obtained first image into the pre-trained classification model in the language conversion character module, output the second type of information, and thus obtain the category to which the content included in the first image belongs.

[0236] It should be noted that the classification model in the embodiments of the present application can be obtained by training an original classification model on low-resolution sample images without preprocessing and corresponding category annotation results, or can be obtained by training an original classification model on low-resolution sample images with preprocessing and corresponding category annotation results, and the embodiments of the present application do not limit this.

[0237] It should be further noted that the difference between the above two training methods is that the brightness and contrast of the low-resolution sample image with preprocessing are better than those of the low-resolution sample image without preprocessing, and therefore, the classification effect of the classification model obtained by training the original classification model on the low-resolution sample image with preprocessing and the corresponding category annotation result is more accurate.

[0238] Optionally, in some embodiments, as Figure 11As shown, the classification model at least includes an encoder layer and a fully connected layer. The terminal device can input the first image into the encoder layer of the classification model to obtain the sixth feature information. In other words, the terminal device can perform feature encoding processing on the first image through the encoder layer to obtain the sixth feature information. After obtaining the sixth feature information, the terminal device can input the sixth feature information into the fully connected layer to obtain the second type information. In other words, the terminal device can perform classification processing on the sixth feature information through the fully connected layer to obtain the second type information.

[0239] It should be noted that the network structure of the original classification model adopts the same network structure as the network structure of the classification model. That is, the original classification model also includes an encoder layer and a fully connected layer.

[0240] S203, the first type information and the second type information are fused to obtain the text description information.

[0241] Optionally, in some embodiments, after performing S201 and S202, the terminal device can perform fusion processing on the obtained first type information and second type information through the language conversion text module 305 as shown, so as to obtain the text description information, which can prepare for the subsequent processing of the multi-modal generative amplification model 303. Figure 12 As shown, the terminal device can also input the first image and the first type information into the language conversion text module 305 to output the text description information of the first image. For example, when the input first image is a low-resolution pine tree image, the time corresponding to the input first image is May 3, 2023, and the location corresponding to the input first image is longitude 122°18′ and latitude 36°57′, the terminal device outputs the text description information as "pine trees in Huangshan, Anhui on May 3" through the language conversion text module 305.

[0242] Optionally, in some embodiments, as shown Figure 11 , the terminal device can also input the first image and the first type information into the language conversion text module 305 to output the text description information of the first image. For example, when the input first image is a low-resolution pine tree image, the time corresponding to the input first image is May 3, 2023, and the location corresponding to the input first image is longitude 122°18′ and latitude 36°57′, the terminal device outputs the text description information as "pine trees in Huangshan, Anhui on May 3" through the language conversion text module 305.

[0243] Based on the description of S104, the multi-modal generative amplification model in the embodiments of the present application can adopt a variety of network structures. Next, in combination with Figure 12 , a network structure of the multi-modal generative amplification model in the embodiments of the present application will be introduced in detail.

[0244] As shown Figure 13 , the multi-modal generative amplification model in the embodiments of the present application can include a denoising network and a super-resolution amplification network. The terminal device can input the first image, the text description information, and the noise image into the denoising network in the multi-modal generative amplification model for denoising processing, and then input the intermediate image after denoising processing into the super-resolution amplification network in the multi-modal generative amplification model for super-resolution amplification processing to obtain the second image.

[0245] Optionally, in some embodiments, the denoising network can employ any one of the following diffusion models: a stable diffusion (SD) model, a differentiable predictive coding (DDPM) model based on deep learning, a stochastic differential equations (SDE) model based on random differential equations, a score based generative model (SGM), and a diffusion model based on a u-shaped convolutional network (U-NET).

[0246] It should be noted that the above diffusion models are only illustrative examples, and in actual applications, other diffusion models can also be employed by those skilled in the art based on actual needs, and the embodiments of the present application do not limit this.

[0247] As shown in Figure 13 , the process of the diffusion model from XT to X0 is a denoising process, and the process of the diffusion model from X0 to XT is a noise adding process. Wherein, XT represents the original noisy image in the denoising process; XT-1 represents the intermediate denoised image at time step T-1; X1 represents the intermediate denoised image at time step 2; X0 represents the denoised image at time step 1; T-1 represents the total number of time steps, that is, the number of denoising steps required for the diffusion model to go from XT to X0.

[0248] It should be noted that the total number of time steps determines the degree of detail of the denoising process. The degree of denoising at each step is different, and in the denoising process, as the number of time steps increases, the degree of denoising gradually decreases, which can ensure that the quality of the denoised image after denoising is higher.

[0249] It should be further noted that the specific network structure of the diffusion model can be referred to the description in the related art, which is not repeated here.

[0250] Considering that the denoising process of the diffusion model needs to be denoised for a total of T-1 time steps to obtain a denoised image with high quality. Therefore, the multi-modal generative enlargement model in the embodiments of the present application is provided with a preset iteration threshold, and the multi-modal generative enlargement model is iterated multiple times to ensure that a second image with high quality is obtained.

[0251] It should be noted that the preset iteration threshold can be set by those skilled in the art according to actual application needs, and the embodiments of the present application do not limit this. Illustratively, the preset iteration threshold can be set to 1000.

[0252] Based on the description of Figure 13 and Figure 13 , the specific process of the terminal device outputting the second image through the multi-modal generative amplification model in the case of using the diffusion model in the denoising network is introduced in detail below. Figure 13 Optionally, in some embodiments, Figure 13 application schematic diagram of the multi-modal generative amplification model is shown, as shown in Figure 14 the terminal device can input the first image, the text description information and the noise image into the multi-modal generative amplification model to perform N iterations to obtain the second image, where N is a positive integer greater than or equal to 1 and less than or equal to a preset iteration threshold.

[0253] The terminal device can determine the current iteration number (or current iteration step number) through the multi-modal generative amplification model.

[0254] It should be noted that the degrees of denoising processing corresponding to different iteration numbers in the embodiments of the present application are different, so the terminal device can determine the degree of denoising processing in the current iteration number by determining the current iteration number.

[0255] After determining the current iteration number N, the terminal device can determine whether the current iteration number N is equal to 1. When it is determined that the current iteration number N is equal to 1, the terminal device can perform the first iteration.

[0256] When performing the first iteration, as shown in Figure 14 the terminal device can input the first image, the noise image and the text description information into the denoising network, and perform denoising processing on the input noise image based on the input text description information and the input first image through the denoising network to obtain the fourth image.

[0257] After obtaining the fourth image, the terminal device can input the fourth image into the super-resolution amplification network to output the fifth image, and perform super-resolution amplification processing on the fourth image from the denoising network through the super-resolution amplification network to obtain the fifth image. Illustratively, the resolution of the fifth image can be 1024 pixels*1024 pixels.

[0258] In the case where the current iteration number N is equal to 1, as shown in Figure 14 the fifth image is an intermediate image obtained by the multi-modal generative amplification model performing denoising processing once and super-resolution amplification processing once on the noise image under the guidance of the first image and the text description information.

[0259] In a case where the current iteration number N is less than the preset iteration threshold, the Nth iteration is performed. The terminal device can input the fifth image, the text description information, and the first image to the denoising network to obtain a denoising-processed fifth image, and input the denoising-processed fifth image to the super-resolution network to obtain a super-resolution-processed fifth image, and use the super-resolution-processed fifth image, the first image, and the text description information as inputs of the next iteration until the current iteration number N is equal to the preset iteration threshold.

[0260] Exemplarily, in a case where the current iteration number N is 2, the terminal device can perform the 2nd iteration. In the 2nd iteration, the terminal device can input the first image, the fifth image, and the text description information to the denoising network to obtain a denoising-processed fifth image. Further, the terminal device can input the denoising-processed fifth image to the super-resolution network to obtain a super-resolution-processed fifth image. The super-resolution-processed fifth image, the first image, and the text description information are used as inputs of the 3rd iteration.

[0261] It should be noted that in a case where the current iteration number N is less than the preset iteration threshold, each iteration process is similar, and reference can be made to the related description of the 2nd iteration, which will not be repeated here.

[0262] As shown in Figure 14 , in a case where the current iteration number N is equal to the preset iteration threshold, the terminal device can output the super-resolution-processed fifth image as the second image through the multi-modal generative magnification model, and obtain a second image with higher detail restoration degree, thereby improving the clarity of the image processed by the super-resolution.

[0263] The specific implementation process of the image processing method in the embodiment of the present application will be described in detail below in conjunction with Figure 14 .

[0264] As shown in Figure 14 , the multi-modal generative magnification model in the embodiment of the present application can include a second encoder, a third encoder, a fourth encoder, a denoising network, and a super-resolution network. The super-resolution network includes at least one convolutional layer. The denoising network includes a first encoder, a first decoder, and an attention network. The output of the second encoder, the output of the third encoder, and the output of the fourth encoder are used as inputs of the denoising network, and the output of the denoising network is used as the output of the super-resolution network. In the denoising network, data can be exchanged between the first encoder and the attention network, and the output of the first encoder is used as the input of the first decoder. The output of the first decoder is used as the input of the super-resolution network.

[0265] It should be noted that Figure 5The attention network shown in the figure is only illustrative, and in the embodiments of the present application, only the attention network corresponding to the first encoder is shown. In actual applications, those skilled in the art can also set the attention network corresponding to the first decoder according to the use requirements, and the embodiments of the present application do not limit this. In addition, the number and specific network structure of the attention network are not limited in the embodiments of the present application. For example, the number of attention networks can be set according to the number of layers of the first encoder. Alternatively, the attention network can also be set according to the number of layers of the first decoder. The specific network structure of the attention network can refer to the description in the related art, which will not be repeated here.

[0266] In the case where the attention network corresponding to the first decoder is also set in the denoising network, Figures 3 to 14 The first feature information, the second feature information and the fourth feature information in the figure also need to be input into the attention network corresponding to the first decoder for feature fusion processing, and the feature information after the fusion processing is input into the first decoder.

[0267] It should also be noted that, Figure 15 The attention network shown in the figure is used for fusion processing of multiple feature information, and in actual applications, the terminal device can also perform fusion processing of two kinds of feature information through the attention network according to the use requirements of those skilled in the art. For example, the terminal device performs fusion processing of the first feature information and the third feature information through the attention network 1, fusion processing of the second feature information and the third feature information through the attention network 2, and inputs the feature information after the fusion processing of the attention network 1 and the feature information after the fusion processing of the attention network 2 into the first encoder for subsequent processing.

[0268] As Figure 15 shown, the image processing method of the embodiments of the present application specifically includes the following steps:

[0269] S301, input the first image into the second encoder to obtain the first feature information.

[0270] Optionally, the second encoder is used for feature encoding processing of the first image. Illustratively, the second encoder can be an image encoder. The first feature information is used to represent the image features corresponding to the first image.

[0271] S302, input the text description information into the third encoder to obtain the second feature information.

[0272] Optionally, the third encoder is used for feature encoding processing of the text description information. Illustratively, the third encoder can be a text encoder. The second feature information is used to represent the text features corresponding to the text description information.

[0273] It should be noted that the specific network architecture of the third encoder can refer to the specific network architecture of the text encoder described in the related art, which is not described here.

[0274] Optionally, in some embodiments, the third encoder can also be a feature encoder. At this time, the first type of information in any form can be directly input to the feature encoder, and the corresponding feature information is output. That is, when the third encoder is a feature encoder, as shown in the language conversion text module 305 in the terminal device, Figure 15 The language conversion text module 305 in the terminal device is an optional module.

[0275] It should be noted that the specific network architecture of the feature encoder can be described in the related art, which is not described here.

[0276] S303, inputting the noise image to the first encoder to obtain third feature information.

[0277] It should be noted that there is no time sequence and order between S301, S302 and S303. That is, S301, S302 and S303 can be executed at the same time, or can be executed in other order, such as S301 in the middle, S302 in front, and S303 in the back, or S301 in front, S302 in back, S303 in the middle, etc.

[0278] Optionally, the first encoder is used for feature encoding processing of the noise image. Illustratively, the first encoder can be an image encoder. The third feature information is used to represent the noise feature corresponding to the noise image.

[0279] It should be noted that the first encoder and the second encoder can use the same image encoder, or different image encoders, which is not limited by the embodiments of the present application.

[0280] Optionally, the image encoder can include a first encoding layer, a second encoding layer, a third encoding layer and a fourth encoding layer from top to bottom. Illustratively, the first encoding layer can output four feature maps with a size of 64*64. The second encoding layer can output four feature maps with a size of 32*32. The third encoding layer can output four feature maps with a size of 16*16. The fourth encoding layer can output four feature maps with a size of 8*8.

[0281] S304, inputting the current iteration number to the fourth encoder to obtain fourth feature information.

[0282] Optionally, the fourth encoder is configured to perform feature encoding processing on the current iteration number. For example, the fourth encoder can be a step number encoder. The fourth feature information is used to represent the iteration number corresponding to the current iteration (e.g., the first iteration).

[0283] It should be noted that the specific network architecture of the fourth encoder can refer to the specific network architecture of the step number encoder described in the related art, which will not be described here.

[0284] S305, when performing the first iteration, input the first feature information, the second feature information, the third feature information, and the fourth feature information to the attention network to obtain the fifth feature information.

[0285] Optionally, the attention network is configured to perform fusion processing on the features. The fifth feature information is used to represent the features fused by the attention network.

[0286] It should be noted that S303 is an optional step. In some embodiments, the terminal device can directly fuse the first feature information, the second feature information, and the fourth feature information through the attention network, and continue to perform S306 to S309, which can also realize the image processing method of the embodiments of the present application.

[0287] S306, input the noise image and the fifth feature information to the first encoder to output the first encoded feature map.

[0288] Optionally, the first encoder is configured to perform feature encoding processing on the noise image and the fifth feature information.

[0289] By performing S305 and S306, the terminal device can fuse the third feature information corresponding to the noise image together, so that the fusion effect of the attention network is better.

[0290] S307, input the first encoded feature map to the first decoder to output the fourth image.

[0291] Optionally, the first decoder is configured to perform feature decoding processing on the first encoded feature map. For example, the first decoder can be an image decoder.

[0292] Optionally, the image decoder can include a first decoding layer, a second decoding layer, a third decoding layer, and a fourth decoding layer from top to bottom. For example, the first decoding layer can output four feature maps with a size of 8*8. The second decoding layer can output four feature maps with a size of 16*16. The third decoding layer can output four feature maps with a size of 32*32. The fourth decoding layer can output four feature maps with a size of 64*64.

[0293] It should be noted that the number of decoding layers in the image decoder can be the same as or different from the number of encoding layers in the image encoder, and this application does not limit this.

[0294] S308, the fourth image is input into at least one convolutional layer, and the fourth image is super-resolution magnified by at least one convolutional layer based on a preset magnification factor to obtain the fifth image.

[0295] It should be noted that in practical applications, in order to improve image clarity, the super-resolution upsampling network may also include an upsampling layer, and this embodiment does not limit this. The upsampling layer can recover the spatial resolution of the image from the feature map output by the convolutional layer through the pixel shuffle upsampling operation, thereby further improving the image clarity.

[0296] It should also be noted that the preset magnification can be the magnification of a pre-set iteration, or it can be the product of the magnifications of at least two iterations. If the actual magnification of the image to be captured exceeds the preset magnification, the terminal device can use digital zoom to further magnify the image.

[0297] S309, if the current iteration number N is equal to the preset iteration threshold, output the fifth image after super-resolution amplification as the second image.

[0298] It should be noted that the above steps are illustrated using the first iteration as an example. S305 to S308 are the specific implementation steps of the first iteration. When the current iteration number N is less than the preset iteration threshold, the specific implementation process of the Nth iteration is similar to S305 to S308. Please refer to the description of S305 to S308. It will not be repeated here.

[0299] In summary, by inputting three modal data sources—the first image, the noisy image, and the text description information—into a multimodal generative magnification model, the multimodal generative magnification model can perform generative super-resolution magnification processing on the first image under the guidance and constraints of the first image and the text description information, and output a second image with more complete details and higher clarity, thereby improving the user's visual experience.

[0300] The above text combined Figure 14 The specific implementation process of the image processing method according to the embodiments of this application is described in detail below. Figure 16 This paper introduces a possible implementation of the multimodal generative amplification model obtained in the embodiments of this application.

[0301] The multimodal generative amplification model in this embodiment of the application is obtained using the following training method:

[0302] Optionally, in some embodiments, the terminal device can acquire the first sample image, the second sample image, the text description sample information, and the third sample image.

[0303] Optionally, the first sample image and the third sample image are images of the same content but different resolutions, and the resolution of the first sample image is smaller than that of the third sample image. The second sample image is a noise image for increasing the zoom diversity of the first sample image. The text description sample information is generated based on the first sample image.

[0304] It should be noted that the first sample image and the third sample image can be acquired by the terminal device, or acquired by other acquisition devices (such as a single-lens reflex camera), and the embodiments of the present application do not limit this.

[0305] As shown in Figure 17 , the terminal device can input the first sample image, the second sample image, the text description sample information, and the third sample image to the original multi-modal generative zoom model for training until the original multi-modal generative zoom model meets the training iteration condition, and determine the original multi-modal generative zoom model meeting the training iteration condition as the multi-modal generative zoom model.

[0306] Optionally, in some embodiments, the training iteration condition can be a training iteration number threshold or a loss threshold, etc.

[0307] It should be noted that in the actual training process, the person skilled in the art can set the corresponding training iteration condition according to the actual situation to ensure that a multi-modal generative zoom model with better effect is obtained.

[0308] It should also be noted that Figure 16 , the network structure of the original multi-modal generative zoom model can be trained using the network structure of the multi-modal generative zoom model shown in Figure 16 .

[0309] The above describes in detail the specific implementation process of the image processing method in the embodiments of the present application. In actual application, the image processing method of the embodiments of the present application can be implemented by relying on the hardware system and the software system of the terminal device. The hardware system and the software system of a terminal device in the embodiments of the present application are introduced below with reference to Figure 16 and Figure 16 .

[0310] Please refer to Figure 16 , Figure 16 , which shows a hardware system schematic diagram of a terminal device provided by the embodiments of the present application.

[0311] As shown in Figure 16As shown, the terminal device 100 can include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a loudspeaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. Among them, the sensor module 180 can include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, and a bone conduction sensor 180M, etc.

[0312] It should be noted that, Figure 16 The structure shown does not constitute a specific limitation on the hardware system of the terminal device 100. In other embodiments of the present application, the hardware system of the terminal device 100 can include more or fewer components than those shown, or the hardware system of the terminal device 100 can include a combination of some of the components shown, or the hardware system of the terminal device 100 can include sub-components of some of the components shown. For example, Figure 16 The hardware system of the terminal device 100 can include more or fewer components than those shown, or the hardware system of the terminal device 100 can include a combination of some of the components shown, or the hardware system of the terminal device 100 can include sub-components of some of the components shown. For example, Figure 16 The hardware system of the terminal device 100 can include more or fewer components than those shown, or the hardware system of the terminal device 100 can include a combination of some of the components shown, or the hardware system of the terminal device 100 can include sub-components of some of the components shown. For example, Figure 3 The hardware system of the terminal device 100 can include more or fewer components than those shown, or the hardware system of the terminal device 100 can include a combination of some of the components shown, or the hardware system of the terminal device 100 can include sub-components of some of the components shown. For example, Figure 4 The proximity light sensor 180G shown can be optional. Figure 17 The components shown can be implemented in hardware, software, or a combination of software and hardware.

[0313] The processor 110 can include one or more processing units. For example, the processor 110 can include at least one of the following processing units: an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, a neural-network processing unit (NPU). Among them, different processing units can be independent devices, or can be integrated devices.

[0314] The controller can generate operation control signals according to the instruction operation code and the timing signal, and complete the control of fetching and executing instructions.

[0315] The processor 110 can also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory can save instructions or data that the processor 110 has just used or repeatedly uses. If the processor 110 needs to use the instructions or data again, it can directly call from the memory. Avoiding repeated access, reducing the waiting time of the processor 110, thus improving the efficiency of the system.

[0316] Figure 17 The connection relationship between the modules shown is only illustrative and does not constitute a limitation on the connection relationship between the modules of the terminal device 100. Alternatively, the modules of the terminal device 100 can also adopt a combination of various connection modes in the above embodiments.

[0317] The terminal device 100 can realize the display function through the GPU, the display screen 194 and the application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 can include one or more GPUs that execute program instructions to generate or change display information.

[0318] The display screen 194 can be used to display images or videos. The display screen 194 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a mini light-emitting diode (Mini LED), a micro light-emitting diode (Micro LED), a micro OLED, or a quantum dot light emitting diode (QLED). In some embodiments, the terminal device 100 can include 1 or N display screens 194, where N is a positive integer greater than 1. In some embodiments, the unfolded state of the display screen 194 can be any one of the four states shown in FIG. 4. Figure 17

[0319] The terminal device 100 can implement the photographing function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor, etc.

[0320] The ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, the light is transmitted to the camera photosensitive element through the lens, the light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing to convert it into an image visible to the naked eye. The ISP can optimize the noise, brightness, and color of the image through algorithms, and can also optimize the exposure and color temperature of the shooting scene and other parameters. In some embodiments, the ISP can be arranged in the camera 193.

[0321] ​The camera 193 is configured to capture still images or videos. An object projects an optical image through a lens to a photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, which is then transmitted to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into an image signal in a standard format, such as RGB or YUV. In some embodiments, the terminal device 100 can include one or N cameras 193, where N is a positive integer greater than 1.

[0322] The digital signal processor is configured to process digital signals, including digital image signals and other digital signals. For example, when the terminal device 100 selects a frequency point, the digital signal processor is configured to perform Fourier transform on the frequency point energy, etc.

[0323] The terminal device 100 can implement audio functions, such as music playing and recording, through the audio module 170, the speaker 170A, the microphone 170B, the microphone 170C, the earphone interface 170D, and the application processor, etc.

[0324] The touch sensor 180K is also referred to as a touch device. The touch sensor 180K can be disposed on the display screen 194, and the touch sensor 180K and the display screen 194 form a touch screen, which is also referred to as a touch screen. The touch sensor 180K is configured to detect a touch operation acting on or near the touch sensor 180K. The touch sensor 180K can transmit the detected touch operation to the application processor to determine the type of touch event. In other embodiments, the touch sensor 180K can also be disposed on the surface of the terminal device 100 and disposed at a position different from the display screen 194.

[0325] Optionally, in some embodiments, the processor 110 is configured to input the obtained low-resolution first image, the noise image, and the text description information into a multi-modal generative enlargement model, perform generative super-resolution enlargement processing on the low-resolution first image under the guidance of the low-resolution first image and the text description information through the multi-modal generative enlargement model, so as to generate image detail information and obtain a high-resolution second image, and display the second image (such as shown in FIG. 8B) to the user, thereby improving the user's visual experience. Figure 17

[0326] ​The hardware system of the terminal device 100 is described in detail above, and the software system of the terminal device 100 is introduced below. The software system can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservice architecture, or a cloud architecture. Embodiments of the present application take the layered architecture as an example to describe the software system of the terminal device 100.

[0327] Please refer to Figure 5 , Figure 5 A schematic diagram of the software system of a terminal device provided in an embodiment of the present application is shown.

[0328] The software system of the terminal device 100 can be divided into several layers, each layer having a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, as shown in Figure 17 , the software system of the terminal device 100 can be an Android system architecture, which can be divided into four layers from top to bottom: an application layer (APP), an application framework layer, a system runtime library layer, and a kernel layer (linux kernel).

[0329] The application layer can include multiple applications. For example, the multiple applications can be audio applications, settings applications, camera applications, calendar applications, and photo album applications, etc. It can be understood that the multiple applications in the application layer can be third-party applications installed by a user, or system applications.

[0330] As shown in Figure 3 , the camera application can further include a multi-modal generative amplification model and a language translation text module, etc. The multi-modal generative amplification model and the language translation text module can refer to the related descriptions in Figure 4 , which will not be described here. The camera application can further include other modules as shown in Figure 17 , such as a post-processing module, etc. (not shown in Figure 17 ).

[0331] The application framework layer is a framework layer for supporting the running of the multiple applications in the application layer. In some embodiments, the application framework layer can include an Android Archive (AAR) file.

[0332] The system runtime library layer may include graphics hardware, which is used to abstract hardware. Graphics hardware may include external libraries and hardware abstraction layer libraries (HAL libraries).

[0333] The kernel layer is an abstraction layer between hardware and software. For example, the kernel layer may include driver modules such as audio drivers, display drivers, hardware interface drivers (e.g., headphone jack drivers), and sensor drivers.

[0334] The image processing method of this application embodiment can be executed in the application layer. Specifically, the image processing method of this application embodiment can be executed through a camera application in the application layer.

[0335] In some embodiments, the camera application in the terminal device 100, in shooting mode, can perform generative super-resolution magnification processing on the first image under the guidance of the first image and text description information through a multimodal generative magnification model in the camera application, thereby outputting a second image with high detail reproduction, allowing the user to view the second image through controls in the album application or camera application (such as...). Figure 17 (3) The control 105 shown in the figure can be used to view a second image with higher clarity (such as...). ​ (As shown).

[0336] It should be understood that ​ The layered structure shown does not constitute a specific limitation on the software system of terminal device 100. In other embodiments of this application, the software system of terminal device 100 may include more than ​ The layered architecture shown may have more or fewer layers, or each layer of the software system of terminal device 100 may include more than [a certain number of layers]. ​ The embodiments shown may have more or fewer constituent structures, and the present application is not limited to these.

[0337] For example, this application provides a computer-readable storage medium including instructions that, when executed on a terminal device 100, cause the terminal device 100 to implement the methods described in the preceding embodiments.

[0338] For example, this application provides a chip system applied to a terminal device 100; the chip system includes one or more processors; the one or more processors are used to invoke computer instructions to cause the terminal device 100 to perform the methods in the foregoing embodiments.

[0339] Exemplarily, the present application provides a computer program product, which, when running on a computer, causes the terminal device 100 to implement the method in the foregoing embodiments.

[0340] In the foregoing embodiments, all or part of the functions can be implemented by software, hardware, or a combination of software and hardware. When implemented by software, the functions can be implemented in the form of a computer program product in whole or in part. The computer program product includes one or more computer codes or instructions. When the computer program codes or instructions are loaded and executed on a computer, all or part of the processes or functions according to the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer codes or instructions can be stored in a computer-readable storage medium. The computer-readable storage medium can be any available medium accessible by the computer or a data storage device such as a server, data center, etc. that includes one or more available media sets. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.

[0341] A person of ordinary skill in the art can understand that all or part of the processes in the foregoing embodiments can be implemented by a computer program to instruct the relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, the processes of the foregoing embodiments can be included. The foregoing storage medium includes a read only memory (ROM) or a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

Claims

1. An image processing method, characterized by, The method is applied to a terminal device and includes: obtaining a first image, the resolution of the first image being a first resolution; generating text description information of the first image based on the first image; the text description information is used to describe the first image, and the text description information at least includes first type information and second type information; the first type information includes a time corresponding to the first image and / or a shooting location corresponding to the first image; the second type information is used to represent a category to which content included in the first image belongs; the time corresponding to the first image is a shooting time of the first image, a time at which a user obtains the first image by screenshot, a time at which the user saves the first image after editing, or a current display time of the terminal device; obtaining a noise image; the noise image is used to provide noise distribution characteristics for the first image; inputting the first image, the text description information, and the noise image into a multi-modal generative magnification model to perform N iterations to obtain a second image; the resolution of the second image is a second resolution, and the second resolution is greater than the first resolution; the multi-modal generative magnification model includes a denoising network and a super-resolution magnification network, the denoising network is connected with the super-resolution magnification network, and N is a positive integer greater than or equal to 1 and less than or equal to a preset iteration threshold value; the inputting the first image, the text description information, and the noise image into the multi-modal generative magnification model to perform N iterations to obtain a second image includes: in the first iteration, inputting the first image, the text description information, and the noise image into the denoising network, and performing denoising processing on the noise image based on the text description information and the first image through the denoising network to obtain a fourth image; inputting the fourth image into the super-resolution magnification network, and performing super-resolution magnification processing on the fourth image from the denoising network through the super-resolution magnification network to obtain a fifth image; in the Nth iteration, inputting the fifth image, the text description information, and the first image into the denoising network to obtain a fifth image after denoising processing, and inputting the fifth image after denoising processing into the super-resolution magnification network to obtain a fifth image after super-resolution magnification processing; in a case where N is equal to the preset iteration threshold value, outputting the fifth image after super-resolution magnification processing as the second image.

2. The method of claim 1, wherein, the denoising network includes a first encoder, a first decoder, and an attention network; the inputting the first image, the text description information, and the noise image into the denoising network in the first iteration, and performing denoising processing on the noise image based on the text description information and the first image through the denoising network to obtain a fourth image includes: In the first iteration, the first feature information, the second feature information, the third feature information and the fourth feature information are input into the attention network to obtain fifth feature information; the fifth feature information is used to represent the feature after the attention network fuses multiple features; the first feature information is used to represent the image feature corresponding to the first image; the second feature information is used to represent the text feature corresponding to the text description information; the third feature information is used to represent the noise feature corresponding to the noise image; and the fourth feature information is used to represent the number feature corresponding to the first iteration; The noise image and the fifth feature information are input into the first encoder to output a first encoded feature map; the first encoder is used for feature encoding processing on the noise image and the fifth feature information; The first encoded feature map is input into the first decoder to output the fourth image; the first decoder is used for feature decoding processing on the first encoded feature map.

3. The method of claim 2, wherein, The multi-modal generative magnification model further comprises a second encoder, a third encoder and a fourth encoder; Before the first feature information, the second feature information, the third feature information and the fourth feature information are input into the attention network, the method further comprises: The first image is input into the second encoder to obtain the first feature information; the second encoder is used for feature encoding processing on the first image; The text description information is input into the third encoder to obtain the second feature information; the third encoder is used for feature encoding processing on the text description information; The noise image is input into the first encoder to obtain the third feature information; The current iteration number is input into the fourth encoder to obtain the fourth feature information; the fourth encoder is used for feature encoding processing on the current iteration number.

4. The method according to any one of claims 1 to 3, characterized in that, The super-resolution magnification network comprises at least one convolutional layer; The fourth image is input into the super-resolution magnification network, and the super-resolution magnification network is used to perform super-resolution magnification processing on the fourth image from the denoising network to obtain a fifth image, comprising: The fourth image is input into the at least one convolutional layer, and the at least one convolutional layer is used to perform super-resolution magnification processing on the fourth image based on a preset magnification ratio to obtain the fifth image.

5. The method according to any one of claims 1 to 3, characterized in that, The text description information of the first image is generated based on the first image, comprising: Obtaining the first type of information; The first image is input into a classification model to output the second type of information; The first type of information and the second type of information are fused to obtain the text description information.

6. The method of claim 5, wherein, The classification model comprises at least an encoder layer and a fully connected layer; The first image is input into the classification model to output the second type of information, comprising: The first image is input into the encoder layer to obtain sixth feature information; the encoder layer is used for feature encoding processing on the first image; The sixth feature information is input into the full connection layer to obtain the second type of information; and the full connection layer is configured to perform classification processing on the sixth feature information.

7. The method according to any one of claims 1 to 3, characterized in that, Before the first image is acquired, the method further includes: In response to an operation of a user opening the camera, a first preview image at a first zoom ratio is displayed; When the first zoom ratio is adjusted to a second zoom ratio, a second preview image is displayed, the second zoom ratio being greater than the first zoom ratio; The first image is acquired by: In a case where a photographing instruction is received, the second preview image is pre-processed to obtain the first image; the pre-processing includes one or more of the following: denoising processing, automatic white balance processing, demosaicing processing, histogram equalization processing, and image enhancement processing; The second image is displayed. Before the second image is displayed, the method further includes:

8. The method of claim 7, wherein, The second image is post-processed; the post-processing includes one or more of the following processing: sharpening processing, compression processing, and format conversion processing. The multi-modal generative magnification model is obtained by using the following training method:

9. The method according to any one of claims 1 to 3, characterized in that, A first sample image, a second sample image, text description sample information, and a third sample image are acquired; the first sample image and the third sample image are images with the same content and different resolutions, and the resolution of the first sample image is less than that of the third sample image; the second sample image is a noise image used to increase the magnification diversity of the first sample image; the text description sample information is generated based on the first sample image; The first sample image, the second sample image, the text description sample information, and the third sample image are input into an original multi-modal generative magnification model for training until the original multi-modal generative magnification model meets a training iteration condition, and the original multi-modal generative magnification model that meets the training iteration condition is determined as the multi-modal generative magnification model. The terminal device includes one or more processors and a memory; the memory is coupled to the one or more processors, the memory is configured to store computer program code, the computer program code includes computer instructions, and the one or more processors invoke the computer instructions to cause the terminal device to perform the method of any one of claims 1 to 9.

10. A terminal device, comprising: The computer-readable storage medium includes instructions that, when executed on a terminal device, cause the terminal device to perform the method of any one of claims 1 to 9.

11. A computer readable storage medium, characterized in that, The chip system is applied to a terminal device; the chip system includes one or more processors; and the one or more processors are configured to invoke computer instructions to cause the terminal device to perform the method of any one of claims 1 to 9.

12. A chip system, characterized by ​

Citation Information

Patent Citations

  • Shooting method and equipment

    CN113452895A

  • Image processing method and device, electronic equipment and readable storage medium

    CN117422612A