Multi-modal image processing method and device and storage medium
By denoising and modal recognition of multimodal images, and using corresponding amplification and detail enhancement algorithms, the problem of poor image processing effect in the prior art is solved, and high-quality image amplification and detail enhancement is achieved.
Patent Information
- Application Number
- CN202510144368.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art uses unified means when processing multimodal images, resulting in poor image effects after processing and inability to effectively restore image details.
By denoising and modal recognition of the target image, corresponding amplification and detail enhancement algorithms are used according to the image type, and the amplification and detail enhancement process is repeated until the preset standard is reached.
It realizes adaptively enlarging and detail enhancement of multimodal images, ensuring that the enlarged image has high resolution clarity, while retaining the details and reality of the original image to the greatest extent.
Smart Images

Figure CN120070191A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a method, device, and storage medium for processing multimodal images. Background Art
[0002] Multimodal refers to an input that simultaneously contains multiple information modalities, such as images, text, audio, etc., and can adapt to different types of inputs and generate high-quality outputs. Among them, multimodal images refer to images of different information modalities, including photos, wallpapers, or screenshots, etc. Existing technologies all use a unified method to process multimodal images, which results in poor image effects after processing, including the inability to effectively restore image details subsequently, etc. Summary of the Invention
[0003] This application provides a method, device, and storage medium for processing multimodal images, which can adaptively magnify and enhance the details of multimodal images.
[0004] On the one hand, this application provides a method for processing multimodal images, and the method includes:
[0005] Performing denoising processing on the target image to obtain a denoised image;
[0006] Performing modality recognition on the denoised image to obtain the type of the target image;
[0007] According to the type of the target image, magnifying the target image by using an image method algorithm corresponding to the type;
[0008] According to the type of the target image, performing detail enhancement on the magnified target image by using an image detail enhancement algorithm corresponding to the type;
[0009] Evaluating the image after detail enhancement;
[0010] If the image after detail enhancement meets the preset standard, outputting the image after detail enhancement; otherwise, repeating the magnification and / or detail enhancement process until the image after detail enhancement meets the preset standard.
[0011] On the other hand, this application provides a device for processing multimodal images, and the device includes:
[0012] A denoising module, configured to perform denoising processing on the target image to obtain a denoised image;
[0013] A type recognition module, configured to perform modality recognition on the denoised image to obtain the type of the target image;
[0014] An amplification module, configured to amplify the target image by using an image method algorithm corresponding to the type according to the type of the target image;
[0015] An enhancement module, configured to enhance the details of the amplified target image by using an image detail enhancement algorithm corresponding to the type according to the type of the target image;
[0016] An evaluation module, configured to evaluate the image after detail enhancement;
[0017] An output module, configured to output the image after detail enhancement if the image after detail enhancement meets a preset standard; otherwise, repeat the amplification and / or detail enhancement process until the image after detail enhancement meets the preset standard.
[0018] In a third aspect, the present application provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the technical solution of the above-mentioned multi-modal image processing method are implemented.
[0019] In a fourth aspect, the present application provides a storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the technical solution of the above-mentioned multi-modal image processing method are implemented.
[0020] As can be seen from the technical solutions provided by the present application above, after identifying the type of the target image, the target image is amplified by using an image amplification algorithm corresponding to the type according to the type of the target image, and the amplified target image is enhanced in details by using an image detail enhancement algorithm corresponding to the type according to the type of the target image. Compared with traditional image amplification methods that often fail to fully consider the characteristics of different types of images, resulting in distortion in detail reconstruction and realism of the amplified images, the technical solution of the present application can ensure that the amplified image has both high-resolution clarity and maximum retention of the details and realism of the original image through this adaptive amplification and detail enhancement strategy. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0022] Figure 1 is a flowchart of the multi-modal image processing method provided by the embodiment of the present application;
[0023] Figure 2 It is a schematic structural diagram of a processing device for multimodal images provided by an embodiment of the present application;
[0024] Figure 3 It is a schematic structure of an electronic device provided by an embodiment of the present application. Specific embodiments
[0025] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0026] In this specification, adjectives such as first and second can only be used to distinguish one element or action from another element or action, and do not necessarily require or imply any actual such relationship or order. Where circumstances permit, reference to an element or component or step (etc.) should not be construed as being limited to only one of the elements, components, or steps, but can be one or more of the elements, components, or steps, etc.
[0027] In this specification, for ease of description, the dimensions of the various parts shown in the drawings are not drawn in actual proportional relationship.
[0028] Multimodal refers to an input that simultaneously contains multiple information modalities, such as images, text, audio, etc., and can adapt to different types of inputs and generate high-quality outputs. Among them, multimodal images refer to images of different information modalities, including photos, wallpapers, or screenshots, etc. When dealing with multimodal images in the prior art, the types or characteristics of images of different modalities are not considered, and all use the same means for processing. For example, the same magnification algorithm is used to magnify images of different modalities, which results in poor image effects after processing, including the inability to effectively restore image details subsequently, etc.
[0029] In view of the above problems in the prior art, the present application proposes a method for processing multimodal images, and its flowchart is as shown in the appendix Figure 1 shown, mainly including steps S101 to S106, which are described in detail as follows:
[0030] Step S101: Denoise the target image to obtain a denoised image.
[0031] For an image, its high-frequency part represents the regions with rich details and rapid changes in the image. The information in the high-frequency part usually includes details, textures, and edges in the image, etc. The changes in these parts are relatively drastic, manifested as rapid changes in pixel values, especially sharp edges, fine textures, and noise in the image. High-frequency information is usually related to the details, fine textures, edges, and noise in the image. Noise usually appears as irregular pixel changes in the image, often being small and rapid changes, so they are concentrated in the high-frequency region of the image. Image noise is usually a random signal or artifact caused by various factors during image acquisition, transmission, or processing. It will affect the quality of the image. Common types of noise include Gaussian noise, salt-and-pepper noise, Poisson noise, and quantization noise, etc. Since noise tends to be concentrated in the high-frequency part of the image, if the target image (i.e., the image to be processed) is not denoised before magnification and detail enhancement, the noise will be magnified simultaneously during the magnification process, resulting in the noise in the image becoming more obvious. The finally reconstructed high-resolution image not only has unclear details but may also have obvious noise, making the image look unnatural. Therefore, the target image can be denoised first to obtain a denoised image. As for the technical means of denoising, it can be achieved by means such as Gaussian filtering, non-local means filtering, or a deep learning denoising network, etc.
[0032] Step S102: Perform modal recognition on the denoised image to obtain the type of the target image.
[0033] Performing modal recognition on the denoised image is essentially to extract and then classify the features of the denoised image to determine the type of the target image. Considering that different images have different features and different feature extraction schemes have their own advantages and disadvantages, in the embodiments of the present application, the advantages of various different feature extraction schemes can be combined to fuse the features extracted by different feature extraction schemes. Specifically, as an embodiment of the present application, performing modal recognition on the denoised image to obtain the type of the target image can be achieved through steps S1021 to S1023, which are described in detail as follows:
[0034] Step S1021: Use a deep convolutional neural network DCNN, a scale-invariant feature transform SIFT model, and a histogram of oriented gradients HOG model to extract the features F DCNN 、F SIFT and F HOG of the denoised image respectively.
[0035] Deep convolutional neural networks (DCNNs) can automatically learn high-level features from images (such as semantic information, complex textures, and shapes), and can generally capture more complex image structures. This is the advantage of using DCNNs to extract features from noisy images. Algorithms based on Scale-Invariant Feature Transform (SIFT) are invariant under image scale, rotation, and affine transformation, and can extract local invariant features; they perform excellently in extracting local information such as textures, corners, and edges, and have significant advantages especially in tasks such as object recognition and image matching. As for the Histogram of Oriented Gradients (HOG) model, it is very suitable for describing the shape information of objects (such as human detection, vehicle detection, etc.). By dividing the image into small regions and calculating the histogram of gradient directions within the regions, local edge features of the image can be captured. It has strong robustness to illumination and pose changes. Embodiments of this application can utilize the advantages of the above feature extraction schemes, and use at least any two of the deep convolutional neural network DCNN, the Scale-Invariant Feature Transform SIFT model, and the Histogram of Oriented Gradients HOG model to extract the features F of the denoised image DCNN , F SIFT and F HOG , that is, use DCNN to extract the feature F of the denoised image DCNN , use the SIFT model to extract the feature F of the denoised image SIFT , use the HOG model to extract the feature F of the denoised image HOG , select a combination of any two of these three schemes or use these three schemes to extract the features F of the denoised image DCNN , F SIFT and F HOG .
[0036] Step S1022: Fuse at least any two of the features F DCNN , F SIFT and F HOG of the denoised image to obtain the fused feature of the denoised image.
[0037] Specifically, fusing at least any two of the features F DCNN , F SIFT and F HOG of the denoised image to obtain the fused feature of the denoised image can be: determining the fusion weights corresponding to the features F DCNN , F SIFT and F HOG of the denoised image; according to the fusion weights w DCNN , F SIFT and F HOG corresponding to the features F DCNN , w SIFT and w HOG of the denoised image, for the features F of the denoised imageDCNN , F SIFT and F HOG are weighted averaged to obtain the fused feature F of the denoised image fuse , that is, F fuse = w DCNN *F DCNN + w SIFT *F SIFT + w HOG *F HOG . In the above embodiment, a scheme for determining the feature F DCNN , F SIFT and F HOG of the denoised image and the corresponding fusion weights w DCNN , w SIFT and w HOG is to first define a target variable Y; then, construct a linear regression model to fit the relationship between the fused feature F fuse and the target variable; finally, use the least squares method to optimize the model parameters to obtain the fusion weights w DCNN , w SIFT and w HOG .
[0038] Another scheme for determining the feature F DCNN , F SIFT and F HOG of the denoised image and the corresponding fusion weights w DCNN , w SIFT and w HOG is: define the initial values of the fusion weights w DCNN , w SIFT and w HOG ; construct a loss function with the fusion weights w DCNN , w SIFT and w HOG as parameters; use the gradient descent method to optimize the fusion weights w DCNN , w SIFT and w HOG to minimize the loss function and obtain the values of the fusion weights w DCNN , w SIFT and w HOG corresponding to the minimum of the loss function.
[0039] Step S1023: Classify the denoised image based on the fused feature of the denoised image to obtain the type of the target image.
[0040] After the fused features of the denoised image are input into the trained classifier, the trained classifier outputs the probability that the denoised image belongs to a certain type according to the fused features. For example, the probability that the image belongs to an image with a smooth area exceeding a preset threshold area, an image with a texture feature exceeding a preset threshold complexity, or an image with an edge contrast exceeding a preset threshold contrast, etc., so as to obtain the type of the target image. It should be noted that the above-mentioned trained classifier can be obtained by training a model based on algorithms such as support vector machines, random forests, or convolutional neural networks. A method for training the above classifier may include steps S1 to S4, which are described as follows:
[0041] Step S1: Determine the loss function of the classifier to be trained.
[0042] In model training, the loss function is generally used to measure the difference between the predicted category of the model and the true label, maximize the margin and minimize the classification error. Cross-entropy loss or mean squared error, etc. can all be used as the loss function.
[0043] Step S2: Input the feature vector or the feature vector and label into the classifier to be trained to train the classifier.
[0044] The feature vector or the feature vector and label can be a pre-prepared training set, and the parameters of the classifier to be trained (such as the hyperplane of the SVM or the weights of the convolutional neural network) are updated through these training sets. This continuously updates the parameters of the classifier to be trained through an optimization algorithm (such as gradient descent) until the loss function converges.
[0045] Step S3: Tune the hyperparameters of the classifier to be trained.
[0046] During the process of training the classifier to be trained, some key hyperparameters of the classifier to be trained (such as the learning rate, regularization parameter, and number of training epochs, etc.) can be adjusted, and the optimal hyperparameters of the classifier are selected through cross-validation or the validation set.
[0047] Step S4: Evaluate the classifier with the optimal hyperparameters.
[0048] Evaluating the classifier includes two links: verification and improvement. Among them, the verification of the classifier refers to evaluating the performance of the classifier on the validation set using accuracy (that is, the proportion of correctly classified samples in the total samples), precision, recall rate, and confusion matrix, etc. If the performance of the classifier is not good, then adjust the feature extraction method, change the classifier, add more training data, or adopt a more complex model.
[0049] Step S103: According to the type of the target image, use an image magnification algorithm corresponding to the type of the target image to magnify the target image.
[0050] In the embodiments of the present application, according to the type of the target image, magnifying the target image by using an image magnification algorithm corresponding to the type of the target image may be: according to the texture feature, smoothness, and edge contrast of the target image, magnifying the target image by using an image method algorithm corresponding to the texture feature, smoothness, and / or edge contrast. Specifically, if the target image belongs to an image with a smooth area exceeding a preset threshold area, the bilinear interpolation or bicubic interpolation algorithm is used to magnify the target image; if the target image belongs to an image with a texture feature exceeding a preset threshold complexity, the super-resolution algorithm based on a deep convolutional neural network is used to magnify the target image; if the target image belongs to an image with an edge contrast exceeding a preset threshold contrast, the edge-preserving algorithm is used to magnify the target image, and so on. It should be noted that if the target image has multiple characteristics, a regional processing scheme can be adopted, that is, for different regions in the target image, the corresponding image magnification algorithm is used for processing. For example, if the texture feature of some regions in the target image exceeds the preset threshold complexity and the edge contrast of some regions exceeds the preset threshold contrast, one solution is to first identify the regions in the target image where the texture feature exceeds the preset threshold complexity and the regions where the edge contrast exceeds the preset threshold contrast; then, for the regions where the texture feature exceeds the preset threshold complexity, the super-resolution algorithm based on a deep convolutional neural network is used to magnify them, and for the regions where the edge contrast exceeds the preset threshold contrast, the edge-preserving algorithm is used to magnify them.
[0051] Step S104: According to the type of the target image, perform detail enhancement on the magnified target image by using an image detail enhancement algorithm corresponding to the type of the target image.
[0052] Different from the prior art that uses the same detail enhancement scheme for each type of image in a multi-modal image, resulting in the image becoming distorted or unnatural, in the embodiments of the present application, according to the type of the target image, an image detail enhancement algorithm corresponding to the type of the target image is used to perform detail enhancement on the magnified target image. Specifically, if the target image belongs to an image with a smooth area exceeding a preset threshold area, contrast stretching is used to perform detail enhancement on the magnified target image; if the target image belongs to an image with a texture feature exceeding a preset threshold complexity, detail enhancement is performed on the magnified target image based on the Laplacian pyramid algorithm; if the target image belongs to an image with an edge contrast exceeding a preset threshold contrast, detail enhancement is performed on the magnified target image based on the edge enhancement algorithm.
[0053] Step S105: Evaluate the image after detail enhancement.
[0054] Specifically, evaluate the image after detail enhancement, where evaluating the image after detail enhancement includes evaluating the image after detail enhancement using any one or a combination of the structural similarity index, peak signal-to-noise ratio, mean square error, and visual information fidelity. It should be noted that since each index has its advantages and disadvantages in evaluating image quality, the above-mentioned indexes for evaluating the image after detail enhancement can be used alone or several indexes can be combined according to the type and characteristics of the target image, etc. For example, although the mean square error is simple and easy to evaluate image quality, it cannot well reflect the human eye's perception of image quality; another example is that although the structural similarity index can take into account the structural information of the image and is more in line with the human eye's perception characteristics than the mean square error and peak signal-to-noise ratio, it may not be sensitive enough to image noise or small-scale detail changes, and so on. In this case, a single index may not be able to comprehensively reflect image quality, especially when there is noise or distortion in the image. Therefore, a multi-index comprehensive evaluation method is usually adopted. For example, the structural similarity index focuses on structural information, and the peak signal-to-noise ratio focuses on the overall error. Therefore, the structural similarity index and the peak signal-to-noise ratio can complement each other.
[0055] Step S106: If the image after detail enhancement meets the preset standard, output the image after detail enhancement; otherwise, repeat the magnification and / or detail enhancement process until the image after detail enhancement meets the preset standard.
[0056] If the image after detail enhancement meets the preset standard, directly output the image after detail enhancement; otherwise, repeat the magnification and / or detail enhancement process, that is, repeat the process of step S103 and step S104, until the image after detail enhancement meets the preset standard. The above-mentioned preset standard is usually a preset threshold. For example, for the evaluation index of the structural similarity index, a structural similarity index threshold can be preset. If the structural similarity index of the image after detail enhancement exceeds this threshold, output the image after detail enhancement; otherwise, repeat the magnification and / or detail enhancement process until the image after detail enhancement reaches or exceeds this threshold. In addition, as mentioned above, if multiple indexes are used to comprehensively evaluate the image after detail enhancement, a normalized weight can be set for each evaluation index according to the characteristics of the target image or the user's attention. An index with a higher user attention is set with a higher weight, and vice versa, a lower weight is set. For example, assume that the structural similarity index and the peak signal-to-noise ratio are used to comprehensively evaluate the image after detail enhancement, and the user pays more attention to the structural information. Then, a higher weight (for example, 0.6) is assigned to the structural similarity index during evaluation, and a lower weight (for example, 0.4) is assigned to the peak signal-to-noise ratio index. That is, compare the value of 0.6*SSIM + 0.4*PSNR with the preset threshold to evaluate whether the image after detail enhancement meets the preset standard, where SSIM represents the structural similarity index and PSNR represents the peak signal-to-noise ratio.
[0057] From the above attached Figure 1 As can be seen from the processing method of the multi-modal image in the above example, after identifying the type of the target image, according to the type of the target image, an image method corresponding to the type is used to magnify the target image, and according to the type of the target image, an image detail enhancement algorithm corresponding to the type is used to enhance the details of the magnified target image. Compared with the traditional image magnification method, which often fails to fully consider the characteristics of different types of images, resulting in distortion in detail reconstruction and realism of the magnified image, the technical solution of this application can ensure that the magnified image has both high-resolution clarity and can retain the details and realism of the original image to the greatest extent through this adaptive magnification and detail enhancement strategy.
[0058] Please refer to the attached Figure 2 , which is a processing device for multi-modal images provided by an embodiment of this application. The device may include a denoising module 201, a type recognition module 202, a magnification module 203, an enhancement module 204, an evaluation module 205, and an output module 206, which are described in detail as follows:
[0059] The denoising module 201 is used to perform denoising processing on the target image to obtain a denoised image;
[0060] The type recognition module 202 is used to perform modal recognition on the denoised image to obtain the type of the target image;
[0061] The magnification module 203 is used to magnify the target image according to the type of the target image by using an image magnification algorithm corresponding to the type of the target image;
[0062] The enhancement module 204 is used to enhance the details of the magnified target image according to the type of the target image by using an image detail enhancement algorithm corresponding to the type of the target image;
[0063] The evaluation module 205 is used to evaluate the image after detail enhancement;
[0064] The output module 206 is used to output the image after detail enhancement if the image after detail enhancement meets the preset standard, otherwise, repeat the magnification and / or detail enhancement process until the image after detail enhancement meets the preset standard.
[0065] From the above attached Figure 2As can be seen from the processing device for the exemplary multimodal image, after identifying the type of the target image, according to the type of the target image, an image magnification algorithm corresponding to the type is adopted to magnify the target image, and according to the type of the target image, an image detail enhancement algorithm corresponding to the type is adopted to enhance the details of the magnified target image. Compared with the traditional image magnification method that often fails to fully consider the characteristics of different types of images, resulting in distortion in terms of detail reconstruction and realism in the magnified image, the technical solution of this application can ensure that the magnified image has both the clarity of high resolution and can retain the details and realism of the original image to the greatest extent through this adaptive magnification and detail enhancement strategy.
[0066] Figure 3 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 3 shown, the electronic device 3 in this embodiment mainly includes: a processor 30, a memory 31, and a computer program 32 stored in the memory 31 and executable on the processor 30, such as a program for the processing method of multimodal images. When the processor 30 executes the computer program 32, the steps in the embodiments of the above-mentioned multimodal image processing method are implemented, such as Figure 1 the steps S101 to S106 shown. Alternatively, when the processor 30 executes the computer program 32, the functions of each module / unit in the above-mentioned device embodiments are implemented, such as Figure 2 the functions of the denoising module 201, the type recognition module 202, the magnification module 203, the enhancement module 204, the evaluation module 205, and the output module 206 shown.
[0067] Exemplarily, the computer program 32 for the method of processing multi-modal images mainly includes: denoising the target image to obtain a denoised image; performing modal recognition on the denoised image to obtain the type of the target image; according to the type of the target image, using an image magnification algorithm corresponding to the type of the target image to magnify the target image; according to the type of the target image, using an image detail enhancement algorithm corresponding to the type of the target image to enhance the details of the magnified target image; evaluating the image after detail enhancement; if the image after detail enhancement meets the preset standard, outputting the image after detail enhancement, otherwise, repeating the magnification and / or detail enhancement process until the image after detail enhancement meets the preset standard. The computer program 32 can be divided into one or more modules / units, and one or more modules / units are stored in the memory 31 and executed by the processor 30 to complete this application. One or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 32 in the electronic device 3. For example, the computer program 32 can be divided into the functions of a denoising module 201, a type recognition module 202, a magnification module 203, an enhancement module 204, an evaluation module 205, and an output module 206 (modules in the virtual device). The specific functions of each module are as follows: The denoising module 201 is used to denoise the target image to obtain a denoised image; the type recognition module 202 is used to perform modal recognition on the denoised image to obtain the type of the target image; the magnification module 203 is used to, according to the type of the target image, use an image magnification algorithm corresponding to the type of the target image to magnify the target image; the enhancement module 204 is used to, according to the type of the target image, use an image detail enhancement algorithm corresponding to the type of the target image to enhance the details of the magnified target image; the evaluation module 205 is used to evaluate the image after detail enhancement; the output module 206 is used to, if the image after detail enhancement meets the preset standard, output the image after detail enhancement, otherwise, repeating the magnification and / or detail enhancement process until the image after detail enhancement meets the preset standard.
[0068] The electronic device 3 may include but is not limited to the processor 30 and the memory 31. Those skilled in the art can understand that Figure 3 merely examples of the electronic device 3, which do not constitute a limitation on the electronic device 3, may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the electronic device may also include input / output devices, network access devices, buses, etc.
[0069] The so-called processor 30 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0070] The memory 31 may be an internal storage unit of the electronic device 3, such as the hard disk or memory of the electronic device 3. The memory 31 may also be an external storage device of the electronic device 3, such as a plug-in hard disk equipped on the electronic device 3, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 31 may also include both the internal storage unit of the electronic device 3 and the external storage device. The memory 31 is used to store computer programs and other programs and data required by the electronic device. The memory 31 may also be used to temporarily store data that has been output or is to be output.
[0071] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above device can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0072] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0073] Those of ordinary skill in the art will realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0074] In the embodiments provided in this application, it should be understood that the disclosed apparatus / devices and methods can be implemented in other ways. For example, the apparatus / device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.
[0075] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0076] In addition, the functional units in each embodiment of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0077] When the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of this application, it can also be completed by instructing relevant hardware through a computer program. The computer program for the multi-modal image processing method can be stored in a storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented, that is, denoising the target image to obtain a denoised image; performing modal recognition on the denoised image to obtain the type of the target image; according to the type of the target image, using an image magnification algorithm corresponding to the type of the target image to magnify the target image; according to the type of the target image, using an image detail enhancement algorithm corresponding to the type of the target image to enhance the details of the magnified target image; evaluating the image after detail enhancement; if the image after detail enhancement meets the preset standard, output the image after detail enhancement, otherwise, repeat the magnification and / or detail enhancement process until the image after detail enhancement meets the preset standard. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The storage medium can include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the storage medium does not include electrical carrier signals and telecommunication signals.
[0078] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit it; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of this application, and should all be included in the protection scope of this application. The above-mentioned specific implementation manners have further elaborated on the purpose, technical solutions and beneficial effects of this application. It should be understood that the above-mentioned are only the specific implementation manners of this application, and are not used to limit the protection scope of this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this application should all be included in the protection scope of this invention.
Claims
1. A method for processing a multimodal image, characterized in that: The method comprises: Perform denoising on the target image to obtain a denoised image; Performing modality recognition on the denoised image to obtain the type of the target image; According to the type of the target image, the target image is enlarged by using an image enlargement algorithm corresponding to the type; According to the type of the target image, using an image detail enhancement algorithm corresponding to the type to perform detail enhancement on the enlarged target image; Evaluate detail-enhanced images; If the detail-enhanced image meets the preset standard, the detail-enhanced image is output; otherwise, the enlargement and / or detail enhancement process is repeated until the detail-enhanced image meets the preset standard.
2. The method for processing multimodal images according to claim 1, characterized in that: The performing modality recognition on the denoised image to obtain the type of the target image includes: The features F of the denoised image are extracted using a deep convolutional neural network DCNN, a scale-invariant feature transform SIFT model, and a histogram of oriented gradients HOG model. DCNN 、F SIFT and F HOG ; The feature F of the denoised image DCNN 、F SIFT and F HOG fusing at least any two of the above to obtain a fusion feature of the denoised image; The denoised image is classified based on the fusion features of the denoised image to obtain the type of the target image.
3. The method for processing multimodal images according to claim 2, characterized in that: The feature F of the denoised image DCNN 、F SIFT and F HOG At least any two of the above are fused to obtain the fusion features of the denoised image, including: Determine the feature F of the denoised image DCNN 、F SIFT and F HOG The corresponding fusion weight; According to the feature F of the denoised image DCNN 、F SIFT and F HOG The corresponding fusion weight w DCNN 、w SIFT and w HOG , the feature F of the denoised image DCNN 、F SIFT and F HOG A weighted average is performed to obtain the fusion features of the denoised image.
4. The method for processing multimodal images according to claim 1, characterized in that: The step of enlarging the target image according to the type of the target image by using an image enlarging algorithm corresponding to the type includes: According to the texture features, smoothness and edge contrast of the target image, the target image is enlarged by using an image enlargement algorithm corresponding to the texture features, smoothness and / or edge contrast.
5. The method for processing multimodal images according to claim 4, characterized in that: The step of magnifying the target image according to the texture features, smoothness and edge contrast of the target image and using an image magnification algorithm corresponding to the texture features, smoothness and / or edge contrast comprises: If the target image is an image containing a smooth area exceeding a preset threshold area, a bilinear interpolation algorithm or a bicubic interpolation algorithm is used to enlarge the target image; If the target image is an image whose texture features exceed a preset threshold complexity, a super-resolution algorithm based on a deep convolutional neural network is used to enlarge the target image; If the target image is an image whose edge contrast exceeds a preset threshold contrast, an edge-preserving algorithm is used to enlarge the target image.
6. The method for processing a multimodal image according to any one of claims 1 to 5, characterized in that: According to the type of the target image, using an image detail enhancement algorithm corresponding to the type to perform detail enhancement on the enlarged target image includes: If the target image is an image containing a smooth area exceeding a preset threshold area, contrast stretching is used to enhance the details of the enlarged target image; If the target image is an image whose texture features exceed a preset threshold complexity, performing detail enhancement on the enlarged target image based on a Laplacian pyramid algorithm; If the target image is an image whose edge contrast exceeds a preset threshold contrast, detail enhancement is performed on the amplified target image based on an edge enhancement algorithm.
7. The multimodal image processing method according to claim 1, characterized in that: The evaluating the detail-enhanced image includes evaluating the detail-enhanced image using any one or a combination of several of a structural similarity index, a peak signal-to-noise ratio, a mean square error, and visual information fidelity.
8. A multimodal image processing device, characterized in that: The device comprises: A denoising module is used to perform denoising on the target image to obtain a denoised image; A type recognition module, used to perform modality recognition on the denoised image to obtain the type of the target image; an enlargement module, configured to enlarge the target image according to the type of the target image by using an image enlargement algorithm corresponding to the type; An enhancement module, configured to enhance the details of the amplified target image by using an image detail enhancement algorithm corresponding to the type of the target image; An evaluation module, used to evaluate the image after detail enhancement; The output module is used to output the detail-enhanced image if the detail-enhanced image meets the preset standard, and otherwise, repeat the enlargement and / or detail enhancement process until the detail-enhanced image meets the preset standard.
9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.