Image correction and recovery method and system based on image-text multi-modal large model

CN121616476BActive Publication Date: 2026-09-22JIANGSU HAOHAN INFORMATION TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511882932.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-15
Publication Date
2026-09-22
Estimated Expiration
2045-12-15

AI Technical Summary

Benefits of technology

通过遍历目标监测区域进行多维图像采集,获得多维图像数据集;构建退化图像图文多模态大模型,通过所述退化图像图文多模态大模型对所述多维图像数据集进行退化分析,提取退化图像特征与退化文本特征;将所述退化图像特征与所述退化文本特征进行融合,将融合结果同步至图像恢复模型进行多层级交互重建,生成图像初始校正恢复数据;基于所述图像初始校正恢复数据进行损失优化,生成多退化图像校正恢复数据。也就是说,通过引入退化图像图文多模态大模型提取退化特征并与图像恢复网络进行多层级交互重建,无需区分退化类型即可实现多退化图像高质量校正恢复,提高了图像恢复精度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121616476B_ABST
    Figure CN121616476B_ABST
Patent Text Reader

Abstract

The application provides an image correction and recovery method and system based on a picture-text multi-modal large model, relates to the technical field of image correction, and the method comprises the following steps: performing multi-dimensional image acquisition by traversing a target monitoring area to obtain a multi-dimensional image dataset; constructing a degraded image picture-text multi-modal large model to perform degradation analysis on the multi-dimensional image dataset; fusing degraded image features and degraded text features, and synchronizing the fusion result to an image recovery model for multi-level interactive reconstruction; and performing loss optimization based on initial image correction and recovery data to generate multi-degraded image correction and recovery data. The application solves the technical problem that, in the prior art, due to the lack of perception and differentiation ability of the image recovery model for degradation types, the image recovery quality is poor in a multi-source degradation scene, and the degradation features are extracted by introducing the degraded image picture-text multi-modal large model and the image recovery network is subjected to multi-level interactive reconstruction, thereby improving the image recovery precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image correction technology, specifically to an image correction and restoration method and system based on a large multimodal model of images and text. Background Technology

[0002] In practical applications, images are often affected by various degradation factors during acquisition and transmission, including blurring, noise, compression artifacts, uneven illumination, and weather occlusion. However, most existing image restoration methods are based on a single degradation model, such as designing separate models for tasks like dehazing, deraining, desnowing, and low-light enhancement. Accurately identifying the degradation type before image restoration and selecting the appropriate restoration model for different degradation tasks not only increases the complexity of use but also significantly reduces the image restoration effect when multiple degradations are superimposed or the degradation factors are unclear. This makes it difficult to meet the high-quality image restoration requirements in complex real-world environments, thus affecting the overall image restoration quality.

[0003] In summary, existing technologies suffer from poor image restoration quality in multi-source degradation scenarios due to the lack of perception and differentiation capabilities of image restoration models for degradation types. Summary of the Invention

[0004] The purpose of this application is to provide an image correction and restoration method and system based on a large image-text multimodal model, in order to solve the technical problem in the prior art that the image restoration model lacks the ability to perceive and distinguish the type of degradation, resulting in poor image restoration quality in multi-source degradation scenarios.

[0005] To achieve the above objectives, this application provides an image correction and restoration method and system based on a large multimodal model of images and text.

[0006] Firstly, this application provides an image correction and restoration method based on a large-scale image-text multimodal model. This method is implemented through an image correction and restoration system based on a large-scale image-text multimodal model. The method includes: traversing the target monitoring area to acquire multidimensional images and obtain a multidimensional image dataset; constructing a degraded image-text multimodal model; performing degradation analysis on the multidimensional image dataset using the degraded image-text multimodal model to extract degraded image features and degraded text features; fusing the degraded image features and the degraded text features; synchronizing the fusion result to an image restoration model for multi-level interactive reconstruction to generate initial image correction and restoration data; and performing loss optimization based on the initial image correction and restoration data to generate multi-degraded image correction and restoration data.

[0007] Optionally, environmental analysis is performed based on the target monitoring area to determine multiple environmental conditions; multiple sets of image acquisition devices in the target monitoring area are activated according to the multiple environmental conditions to perform multidimensional image acquisition of the target monitoring area, obtaining multiple sets of original environmental images; the multiple sets of original environmental images are traversed for quality screening to generate image screening results; text is combined based on the image screening results to construct multiple image-text pairs; the multiple image-text pairs are added to the multidimensional image dataset.

[0008] Optionally, the multidimensional image dataset is used to train the image-text multimodal large model end-to-end according to the model network architecture to construct a degraded image-text multimodal large model; the multidimensional image dataset is synchronized to the degraded image-text multimodal large model for multi-scale analysis to determine multi-scale features; the multi-scale features are fused according to a feature pyramid network, and global average pooling is performed based on the feature fusion result to generate image feature vectors; degradation perception is performed based on the image feature vectors to extract degraded image features; the multidimensional image dataset is synchronized to the degraded image-text multimodal large model for hidden analysis to determine hidden features; attention pooling is performed on the hidden features to generate keyword features, and the keyword features are projected according to the image feature vectors to generate text feature vectors; degradation perception is performed based on the text feature vectors to extract degraded text features.

[0009] Optionally, a dual-encoder structure is adopted, which includes a visual encoder and a text encoder; image features are extracted from the multidimensional image dataset based on the visual encoder; text features are extracted from the multidimensional image dataset based on the text encoder; semantic analysis is performed on the image features and the text features to construct a shared semantic space; the image features and the text features are mapped to the shared semantic space for spatial alignment to construct the model network architecture.

[0010] Optionally, the image features and text features are mapped to the shared semantic space for similarity analysis to construct a similarity matrix; based on the similarity matrix, the image to text is compared to construct a first contrastive learning loss parameter; based on the similarity matrix, the text to image is compared to construct a second contrastive learning loss parameter; the first contrastive learning loss parameter and the second contrastive learning loss parameter are combined and aligned to construct the model network architecture.

[0011] Optionally, a large image-text multimodal model is introduced, and the multiple image-text pairs are synchronized to the large image-text multimodal model in batches according to the model network architecture; forward propagation calculation is performed through the large image-text multimodal model to obtain initial image features and initial text features; backpropagation is performed based on the initial image features and the initial text features to generate backpropagation information; the visual encoder and the text encoder are updated based on the backpropagation information, and the large image-text multimodal model is trained end-to-end according to the update results to construct the degraded image-text multimodal model.

[0012] Optionally, the degraded image features and the degraded text features are fused in a multi-mode manner to generate degraded enhancement features; the degraded enhancement features are then synchronized to the image restoration model for multi-level feature interaction: S1: The multi-level feature interaction includes an encoding stage and a decoding stage; S2: Based on the encoding stage, the degraded image features and the degraded text features are subjected to cross-attention interaction to extract degraded semantic features; S3: Based on the decoding stage, the degraded enhancement features are subjected to skip connections, and upsampling is performed based on the connection results to obtain an upsampled dataset; S4: Image restoration is performed based on the upsampled dataset to generate image detail information, and continuous interaction is performed between the image detail information and the degraded enhancement features to generate the initial image correction and restoration data.

[0013] Optionally, the image detail information is mapped to the image space for multi-scale parallel analysis to generate a residual image; the residual image is added element-wise to the multi-dimensional image dataset to generate an initial restoration result; the initial restoration result is cropped by pixel values ​​to determine the pixel restoration result; a progressive interaction is performed between the pixel restoration result and the degradation enhancement features to generate pixel restoration detail interaction content; multi-scale detail reconstruction is performed according to the pixel restoration detail interaction content to generate the initial image correction and restoration data.

[0014] Optionally, a comprehensive quality assessment is performed on the initial image correction and restoration data to generate an image quality assessment result; multi-scale loss analysis is performed based on the image quality assessment result to generate a multi-scale loss score; a multi-task joint optimization objective is set, and dynamic contribution calculation is performed based on the multi-scale loss score according to the multi-task joint optimization objective to obtain multiple contribution degrees; the initial image correction and restoration data is dynamically balanced according to the multiple contribution degrees to construct a weight adjustment strategy; the weight adjustment strategy is executed to perform multiple rounds of optimization correction on the initial image correction and restoration data to generate an optimized correction result; the optimized correction result is verified, and when the verification is successful, the multi-degraded image correction and restoration data is generated.

[0015] Secondly, this application also provides an image correction and restoration system based on a large image-text multimodal model, used to execute the image correction and restoration method based on a large image-text multimodal model as described in the first aspect. The image correction and restoration system based on a large image-text multimodal model includes: a multidimensional image acquisition module for traversing the target monitoring area to acquire multidimensional images and obtain a multidimensional image dataset; a degradation analysis module for constructing a degraded image-text multimodal model, performing degradation analysis on the multidimensional image dataset through the degraded image-text multimodal model, and extracting degraded image features and degraded text features; a multi-level interactive reconstruction module for fusing the degraded image features and the degraded text features, synchronizing the fusion result to an image restoration model for multi-level interactive reconstruction, and generating initial image correction and restoration data; and a loss optimization module for performing loss optimization based on the initial image correction and restoration data to generate multi-degraded image correction and restoration data.

[0016] One or more technical solutions provided in this application have at least the following technical effects or advantages: A multidimensional image dataset is obtained by traversing the target monitoring area to acquire multidimensional images. A degraded image-text multimodal model is constructed, and degradation analysis is performed on the multidimensional image dataset using this model to extract degraded image features and degraded text features. The degraded image features and degraded text features are then fused, and the fusion result is synchronized to an image restoration model for multi-level interactive reconstruction to generate initial image correction and restoration data. Loss optimization is performed based on the initial image correction and restoration data to generate multi-degraded image correction and restoration data. In other words, by introducing a degraded image-text multimodal model to extract degradation features and performing multi-level interactive reconstruction with an image restoration network, high-quality correction and restoration of multi-degraded images can be achieved without distinguishing between degradation types, thus improving image restoration accuracy.

[0017] The above description is merely an overview of the technical solution of this application. To better understand the technical means of this application and to facilitate its implementation according to the description, and to make the above and other objects, features, and advantages of this application more apparent, specific embodiments of this application are described below. It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent through the following description. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating the image correction and restoration method based on a large multimodal model of images and text in this application.

[0020] Figure 2 This is a schematic diagram of the image correction and restoration system based on a large multimodal model of images and text, as described in this application.

[0021] Figure labeling: 11 Multidimensional image acquisition module, 12 Degradation analysis module, 13 Multi-level interactive reconstruction module, 14 Loss optimization module. Detailed Implementation

[0022] This application provides an image correction and restoration method and system based on a large-scale image-text multimodal model, which solves the technical problem in existing technologies where the lack of perception and differentiation of degradation types in image restoration models leads to poor image restoration quality in multi-source degradation scenarios. By introducing a large-scale image-text multimodal model to extract degradation features and performing multi-level interactive reconstruction with the image restoration network, high-quality correction and restoration of multi-degraded images can be achieved without distinguishing degradation types, thus improving image restoration accuracy.

[0023] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. It should be understood that this application is not limited to the exemplary embodiments described herein. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application. It should also be noted that, for ease of description, only the parts related to this application are shown in the accompanying drawings, not all of them.

[0024] Example 1, please refer to the appendix. Figure 1 This application provides an image correction and restoration method based on a large image-text multimodal model. The method is applied to an image correction and restoration system based on a large image-text multimodal model, and specifically includes the following steps: Multidimensional image acquisition is performed by traversing the target monitoring area to obtain a multidimensional image dataset.

[0025] Furthermore, this application also includes the following steps: performing environmental analysis based on the target monitoring area to determine multiple environmental conditions; activating multiple sets of image acquisition devices in the target monitoring area according to the multiple environmental conditions to perform multidimensional image acquisition of the target monitoring area, obtaining multiple sets of original environmental images; traversing the multiple sets of original environmental images for quality screening to generate image screening results; combining text based on the image screening results to construct multiple image-text pairs; and adding the multiple image-text pairs to the multidimensional image dataset.

[0026] Specifically, environmental analysis is conducted on the target monitoring area. By reviewing historical meteorological data and environmental sensor networks, environmental factors that significantly impact image quality are identified and categorized. The target monitoring area is the specific spatial range where image acquisition and analysis are required, such as a city intersection, a section of highway, a parking lot, or a field. Multiple environmental conditions affect image quality, including specific natural or artificial lighting conditions such as weather conditions (sunny, foggy, rainy, snowy, hazy, dusty), lighting conditions (daytime, nighttime, low light, backlight, strong light), and seasonal conditions (different times of day, spring, summer, autumn, and winter).

[0027] Multiple sets of image acquisition devices in the target monitoring area are activated based on various environmental conditions. Capture or record screenshots from different angles, times, and lighting conditions when each environmental condition occurs, thereby obtaining multiple sets of raw environmental images. The image acquisition devices are camera groups deployed within the target monitoring area, including visible light cameras of different models and focal lengths, as well as potentially existing infrared cameras, used to capture images from different angles and under different conditions.

[0028] The process iterates through multiple sets of original environmental images for quality screening. It calculates the sharpness and information entropy of each image and sets thresholds to automatically filter out completely black / white images caused by focus failure or signal interruption. This is supplemented by manual sampling to generate the final image screening results, eliminating low-quality images. The final image screening results consist of high-quality images retained after the quality screening process, along with their corresponding quality labels, indicating whether to keep or discard them.

[0029] Based on the filtered images, multiple image-text pairs are constructed by combining text descriptions. Each filtered image is labeled with a text description indicating its content and degradation type. These image-text pairs are then added to a multidimensional image dataset for subsequent model training. The text descriptions include terms such as clear, low, haze, small rain, and heavy rain, where clear corresponds to a clear image, low to a low-light image, small rain to a light rain image, and heavy rain to a heavy rain image.

[0030] By collecting, filtering, and semantically enhancing image data under different environmental conditions, the resulting training dataset not only contains rich multi-scene image information but also has clear degradation semantic descriptions, thereby significantly improving the recognition ability and restoration accuracy of subsequent image restoration models for multi-source degradation scenes.

[0031] A degraded image-text multimodal model is constructed, and degradation analysis is performed on the multidimensional image dataset using the degraded image-text multimodal model to extract degraded image features and degraded text features.

[0032] Furthermore, this application also includes the following steps: using the multidimensional image dataset to perform end-to-end training on the image-text multimodal large model according to the model network architecture, constructing a degraded image-text multimodal large model; synchronizing the multidimensional image dataset to the degraded image-text multimodal large model for multi-scale analysis to determine multi-scale features; fusing the multi-scale features according to a feature pyramid network, and performing global average pooling based on the feature fusion result to generate image feature vectors; performing degradation perception based on the image feature vectors to extract degraded image features; synchronizing the multidimensional image dataset to the degraded image-text multimodal large model for hidden analysis to determine hidden features; performing attention pooling on the hidden features to generate keyword features, and projecting the keyword features according to the image feature vectors to generate text feature vectors; performing degradation perception based on the text feature vectors to extract degraded text features.

[0033] Specifically, using a multidimensional image dataset and following the constructed model network architecture, the large-scale image-text multimodal model was trained end-to-end using the Adam optimizer with a learning rate set to 5e. -5 With a batch size of 128, the model was trained for 20 epochs on a dataset containing 268,000 image-text pairs until the contrastive loss converged, thus constructing a stable degraded image-text multimodal large model. End-to-end training means that the entire process from input data to output results is completed through a unified network during model training. The model can learn directly from the data without the need for manually designing intermediate processing steps.

[0034] For image paths, the multidimensional image dataset is synchronized to a degraded image-text multimodal large model for multi-scale analysis. Utilizing the four-stage output of its visual encoder, four different scale multi-scale features are determined, with dimensions of 56*56*256, 28*28*512, 14*14*1024, and 7*7*2048, respectively. These multi-scale features are fused using a feature pyramid network, generating an enhanced set of multi-scale features through top-down and lateral connections. Global average pooling is then applied to the top-level feature map after fusion, generating a 2048-dimensional image feature vector. The feature pyramid network is a network structure for effectively fusing multi-scale features. Through top-down paths and lateral connections, it fuses deep, strong semantic features with shallow, high-resolution features, thereby generating feature maps with both strong semantics and high resolution at each scale. Global average pooling is a pooling operation that averages all pixel values ​​in each channel of the feature map, compressing it into a single value. For a C-channel feature map, global average pooling outputs a C-dimensional vector, reducing dimensionality and enhancing the feature's robustness to spatial location. Degradation perception based on image feature vectors involves extracting features related to image degradation, such as blur, noise, and low light, from the image feature vectors to obtain degraded image features.

[0035] For text paths, the multidimensional image dataset is synchronized to a degraded image-text multimodal large model for latent analysis. A text encoder processes the text to determine the latent features output by the last Transformer layer—the implicit semantics inferred from the context, not directly from the image. Attention pooling is applied to the latent features, calculating the attention weight for each word, and the weighted sum is used to generate key word features, which are vector representations of the core semantics of the text. These key word features are projected onto the image feature vector and passed through a fully connected layer (768→512) to generate a 512-dimensional text feature vector. Degradation perception is then performed based on this text feature vector to extract semantically clear degraded text features.

[0036] Through multi-scale analysis and global average pooling, the extracted degraded image features contain both detailed and global semantic information. Through attention pooling, the text features can focus on the core words describing the degradation and filter out the interference of irrelevant words.

[0037] Furthermore, this application also includes the following steps: employing a dual encoder structure, wherein the dual encoder structure includes a visual encoder and a text encoder; extracting image features from the multidimensional image dataset based on the visual encoder; extracting text features from the multidimensional image dataset based on the text encoder; performing semantic analysis on the image features and the text features to construct a shared semantic space; mapping the image features and the text features to the shared semantic space for spatial alignment to construct the model network architecture.

[0038] Furthermore, this application also includes the following steps: mapping the image features and the text features to the shared semantic space for similarity analysis to construct a similarity matrix; comparing the image to text based on the similarity matrix to construct a first contrastive learning loss parameter; comparing the text to the image based on the similarity matrix to construct a second contrastive learning loss parameter; and combining and aligning the first contrastive learning loss parameter and the second contrastive learning loss parameter to construct the model network architecture.

[0039] Specifically, a dual-encoder structure is employed, comprising a visual encoder and a text encoder. One encoder processes image data, while the other processes text data. Together, they extract features from both the image and the text. The visual encoder transforms the input image from pixel space to a high-dimensional feature vector space, capturing image features; the text encoder transforms the input text from a sequence of words to a high-dimensional feature vector space, capturing text features.

[0040] A visual encoder is used to extract image features from a multidimensional image dataset. A convolutional neural network is used to extract semantic information from the images, identifying key elements such as object boundaries, colors, and textures. A text encoder is then used to process the text descriptions corresponding to the images, transforming the text information into feature vectors.

[0041] Image and text features extracted from the visual encoder and text encoder are mapped to the same shared semantic space, ensuring they have the same dimension. This shared semantic space is a common high-dimensional vector space into which image and text features are projected. In this shared semantic space, semantically similar images and texts have geometrically close feature vectors, while semantically unrelated ones are far apart.

[0042] Image and text features of a batch are mapped to a shared semantic space, similarity analysis is performed, and a similarity matrix is ​​constructed. The cosine similarity between all possible image-text pairs is calculated to form the similarity matrix. For example, assuming a batch has N image-text pairs, this matrix is ​​an N×N square matrix. The element Si,j in the matrix represents the cosine similarity score between the i-th image feature and the j-th text feature in the shared semantic space. ,in Represents the features of the i-th image. This represents the feature of the j-th text. Indicates the temperature coefficient. and This represents an image and its corresponding text description.

[0043] Based on the similarity matrix, image-to-text comparison is performed to construct the first contrastive learning loss parameters. The image-to-text loss is... Where N represents the number of samples. For the i-th image in the batch, it forms a positive sample with the corresponding i-th text, and the similarity Si,i should be as high as possible; while with the other N-1 negative sample texts in the batch, i.e., texts j≠i, the similarity Si,j should be as low as possible. The first contrastive learning loss parameter is used to measure the model's ability to find the correct description of a given image from a batch of texts, ensuring that the similarity between matching image and text pairs is maximized, while the similarity between unmatched text and image pairs is minimized.

[0044] A second contrastive learning loss parameter is constructed based on the similarity matrix used to compare text and images. The loss for text-to-image comparison is... For the j-th text in a batch, it forms a positive sample with its corresponding j-th image, and the similarity = Sj,j should be high; while with the other N-1 negative sample images in the batch, i.e., images i≠j, the similarity Si,j should be low. The second contrastive learning loss parameter is used to measure the model's ability to find the correctly matched image from a batch of images for a given text description, ensuring that similar texts have the highest similarity with their corresponding images and the lowest similarity with other images. For example, a batch contains N=32 image-text pairs, including images I1 (thick fog), I2 (heavy rain), ..., I32 (clear), and their corresponding texts T1: thick fog, T2: heavy rain, ..., T32: clear. 32 image features and 32 text features are calculated, resulting in a 32x32 similarity matrix. Ideally, the values ​​on the diagonal of the matrix, i.e., positive samples, should be high. Positive sample similarity S1,1=0.92, i.e., image I1 thick fog and text T1 thick fog; negative sample similarity S1,2=0.15, i.e., image I1 thick fog and text T2 heavy rain; negative sample similarity S1,3=0.08, i.e., image I1 thick fog and text T3.

[0045] The first and second contrastive learning loss parameters are combined and aligned to form the overall loss function. This involves calculating the arithmetic mean of the two to ensure a balanced optimization of the model's retrieval capabilities in both the image and text directions, thus constructing the model's network architecture. Through a dual-encoder structure and contrastive learning loss, a shared semantic space is built, achieving precise association and alignment between visual degradation patterns and textual semantic descriptions. This enables the model not only to understand the content of the image but also to identify quality issues present within it.

[0046] Furthermore, this application also includes the following steps: introducing a large image-text multimodal model, synchronizing the multiple image-text pairs to the large image-text multimodal model in batches according to the model network architecture; performing forward propagation calculation through the large image-text multimodal model to obtain initial image features and initial text features; performing backpropagation based on the initial image features and the initial text features to generate backpropagation information; updating the visual encoder and the text encoder based on the backpropagation information, and performing end-to-end training on the large image-text multimodal model according to the update results to construct the degraded image-text multimodal model.

[0047] Specifically, a large-scale image-text multimodal model is introduced. Multiple image-text pairs are synchronized to the large-scale model in batches according to the model's network architecture. This involves inputting the image and its corresponding text into the visual encoder and text encoder, respectively. Forward propagation computation is performed through the large-scale image-text multimodal model, meaning data flows from the model's input layer to the output layer. The input data undergoes layer-by-layer computation to ultimately obtain the model's predicted output. The image passes through the visual encoder to output initial image features, and the text passes through the text encoder to output initial text features.

[0048] Based on the first and second contrastive learning loss parameters of the model network architecture, backpropagation is performed on the initial image and text features. The gradient of the loss with respect to each parameter in the model is calculated using the chain rule, generating backpropagation information, i.e., the gradient tensor. Backpropagation calculates the gradient of the loss function with respect to each model parameter from the output layer back to the input layer based on the difference between the output calculated in the forward propagation and the true label, i.e., the loss. The backpropagation information, the calculated gradients, indicates the direction and magnitude in which each model parameter should be adjusted to reduce the loss.

[0049] Based on the backpropagation information, the visual encoder and text encoder are updated, that is, millions or even hundreds of millions of parameters of the visual encoder and text encoder are updated through the optimizer. This process is repeated tens of thousands of times, and the large-scale image-text multimodal model is trained end-to-end based on the parameter update results until the model loss converges and its performance is stable on downstream tasks, finally constructing a large-scale degraded image-text multimodal model.

[0050] By minimizing the contrast loss, the visual encoder and text encoder are forced to adjust their parameters so that the feature vectors of matched image-text pairs are close to each other in the shared semantic space, while those of mismatched pairs are far apart. The resulting degraded image-text multimodal large model becomes the degradation perceptron of the entire image restoration system, generating high-quality, semantically relevant degradation features for any input image or text prompt.

[0051] The degraded image features and the degraded text features are fused together, and the fusion result is synchronized to the image restoration model for multi-level interactive reconstruction to generate initial image correction and restoration data.

[0052] Furthermore, this application also includes the following steps: performing multi-mode fusion of the degraded image features and the degraded text features to generate degraded enhancement features; synchronizing the degraded enhancement features to the image restoration model for multi-level feature interaction: S1: The multi-level feature interaction includes an encoding stage and a decoding stage; S2: Based on the encoding stage, performing cross-attention interaction between the degraded image features and the degraded text features to extract degraded semantic features; S3: Based on the decoding stage, performing skip connections on the degraded enhancement features, and upsampling based on the connection results to obtain an upsampled dataset; S4: Performing image restoration based on the upsampled dataset to generate image detail information, and continuously interacting with the degraded enhancement features based on the image detail information to generate the initial image correction and restoration data.

[0053] Furthermore, this application also includes the following steps: mapping the image detail information to the image space for multi-scale parallel analysis to generate a residual image; adding the residual image to the multi-dimensional image dataset element-wise to generate an initial restoration result; cropping the pixel values ​​of the initial restoration result to determine the pixel restoration result; performing progressive interaction between the pixel restoration result and the degradation enhancement features to generate pixel restoration detail interaction content; and performing multi-scale detail reconstruction according to the pixel restoration detail interaction content to generate the initial image correction and restoration data.

[0054] Specifically, the degradation image features obtained from degradation analysis are fused with degradation text features using a multi-modal method to generate degradation enhancement features. These features include both specific visual degradation patterns from the image and abstract degradation semantic information from the text. The degradation enhancement features are then synchronized to the image restoration model for multi-level feature interaction. Multi-level feature interaction refers to the repeated exchange and integration of information between the degradation enhancement features and the image's internal features at multiple different levels of the image restoration model.

[0055] Multi-level feature interaction comprises encoding and decoding stages. The encoding stage, the first half of the image restoration network, typically uses a series of downsampling operations to gradually increase the receptive field, extracting and compressing deep semantic features of the image. The decoding stage, the second half, typically uses a series of upsampling operations to gradually restore the spatial resolution of the image and reconstruct details. In the encoding stage, for each layer's image features, it is treated as Q, and the degradation enhancement features are treated as K and V, undergoing cross-attention interaction. This allows the network to focus on degradation-related regions when extracting image features, thereby extracting degradation semantic features containing the degradation context. Cross-attention interaction allows features from one sequence (Query) to query features from another sequence (Key, Value) and extract relevant information. In the decoding stage, skip connections are made to the degradation enhancement features, and upsampling is performed based on the connection results to gradually obtain higher-resolution feature maps, i.e., the upsampled dataset. Skip connections directly pass the feature map of a layer in the encoding stage to the corresponding layer in the decoding stage, helping to solve the gradient vanishing problem and passing the detailed information captured by the encoder to the decoder, thus better reconstructing high-frequency details. Image restoration is performed based on the upsampled dataset, progressively generating image detail information such as sharp edges and textures. Degradation enhancement features interact with the generated image detail information to ensure that the restored details move in the direction of removing specific degradations, rather than blindly reconstructing them. Finally, the decoder's last layer outputs the initial corrected image restoration data.

[0056] An image restoration model is a neural network used to reconstruct a clear image from a degraded image. It typically employs an encoder-decoder structure, and its training process is as follows: Load a pre-trained large-scale multimodal model of degraded images. Extract small batches of data from the sample set. ,in Indicates a degraded image. express The corresponding text description, express The corresponding clear image, This indicates the batch size, that is, the number of images sent to the network at one time, and the number of recovered images output by the network. By using a large multimodal model of degraded images, we can obtain the features of degraded images. and text description features Random selection or As output Using convolutional layers Mapped to , Indicates the number of channels. and Let represent the width and height of the image, respectively. Calculate cross-attention, with the input being... and a: Calculation ,in, This is the weight matrix. For bias vectors, Presentation layer normalization, output ,Will Divided by channel and b: Calculate attention ,in c represents the softmax function; c: for the output Processed using a multilayer perceptron, the multilayer perceptron output... Final output ;d: for the output Downsampling was performed to obtain .Will and As input, perform cross-attention calculation, execute steps a to d, and obtain... .Will and As input, perform cross-attention calculation again, execute steps a to d, and obtain... Repeat this process N times to obtain a set of features. .Will and As input, steps a through c are performed, and then the output is upsampled to obtain... .Will and As input, steps a through c are performed, and then the output is upsampled to obtain... Repeat the process N times to obtain the output. Using convolutional layers to Mapped to Output Constructing the loss function , where N represents the number of samples. The gradient descent algorithm is used to update the image restoration model.

[0057] Image detail information is mapped to the image space, its channel count is increased to 3 (RGB channels), and multi-scale parallel analysis is performed to generate a residual image. The image space is a standard RGB or grayscale pixel space, where each pixel is represented by a specific numerical value, such as 0-255. Multi-scale parallel analysis uses multiple convolutional kernels with different receptive fields to process features simultaneously, capturing contextual information across different ranges, from local texture to larger-scale structure. In image restoration, a very effective strategy is to have the network learn the difference between the sharp image and the degraded image, i.e., the residual. The residual image typically contains degraded components that need to be removed, such as noise, rain streaks, or fog effects, or details that need to be enhanced.

[0058] The residual image is element-wise added to the corresponding degraded input image in the multidimensional image dataset, using the formula I_initial = I_input + R, where R is the residual image and I_input is the corresponding degraded input image in the multidimensional image dataset, thus generating the initial restoration result I_initial. Pixel value clipping is then applied to the initial restoration result, restricting all pixel values ​​to the range [0, 1.0] or [0, 255] to determine the pixel restoration result. Pixel value clipping forcibly constrains pixel values ​​outside the legal range to this range to prevent overflow artifacts.

[0059] A progressive interaction is performed based on the pixel restoration results and degradation enhancement features. The pixel restoration results are fed back into a lightweight convolutional network, while the degradation enhancement features are injected into this network through a cross-attention mechanism, guiding it to focus on repairing areas with remaining flaws and generating pixel restoration detail interaction content, i.e., refined features. Progressive interaction refers to interacting the initial restoration results with guidance information, i.e., with degradation enhancement features, for a second or even multiple rounds of fine-tuning to gradually refine the restoration effect. Pixel restoration detail interaction content is a new feature generated during the progressive interaction process, containing image regions and details that need further adjustment based on the previous round's results and degradation guidance information.

[0060] Multi-scale detail reconstruction is performed based on pixel-level detail recovery interaction content. This involves reconstructing and fusing the final image details at multiple scales to ensure consistency from subtle textures to overall structure, resulting in initial corrected image restoration data. Through residual learning and progressive interaction, the results are fine-tuned using the initially restored image and invariant degradation priors. This effectively eliminates minor imperfections or unnaturalness that may have been introduced in the first round of restoration, ensuring the authenticity of details.

[0061] Through multi-level feature interactions, particularly cross-attention in the encoding stage, degradation information is deeply integrated into the image feature extraction process, enabling the image restoration model to understand the nature of degradation and its impact on the scene. Through continuous interaction and skip connections in the decoding stage, while restoring image details, the generation of these details is always constrained by the degradation semantics, thus avoiding discrepancies between the restoration results and the scene semantics, while preserving rich textures.

[0062] Loss optimization is performed based on the initial image correction and restoration data to generate multi-degraded image correction and restoration data.

[0063] Furthermore, this application also includes the following steps: performing a comprehensive quality assessment on the initial image correction and restoration data to generate an image quality assessment result; performing multi-scale loss analysis based on the image quality assessment result to generate a multi-scale loss score; setting a multi-task joint optimization objective, and performing dynamic contribution calculation based on the multi-scale loss score according to the multi-task joint optimization objective to obtain multiple contribution degrees; dynamically balancing the initial image correction and restoration data according to the multiple contribution degrees to construct a weight adjustment strategy; executing the weight adjustment strategy to perform multiple rounds of optimization correction on the initial image correction and restoration data to generate an optimized correction result; verifying the optimized correction result, and generating the multi-degraded image correction and restoration data when the verification is successful.

[0064] Specifically, a comprehensive quality assessment is performed on the initial corrected and restored image data. The quality of the restored image is quantitatively analyzed from different dimensions to obtain image quality assessment results, typically including scores for multiple quality dimensions. Multi-scale loss analysis is then performed on the image quality assessment results, calculating loss functions at different spatial scales of the image to evaluate the restoration effect of the image restoration model at different levels. Based on the image quality assessment results, multi-scale loss analysis is performed, calculating losses at three scales, such as the original scale, half-scale, and quarter-scale, generating multi-scale loss scores. For example, if the average PSNR of the initial corrected and restored image data on the SIDD validation set is 38.5 dB, SSIM is 0.95, but LPIPS is 0.085, it indicates that there is room for improvement in perceptual quality. Multi-scale loss analysis on a single image example yields a PSNR of 39.1, an SSIM of 0.952, and LPIPS of 0.089. PSNR is Peak Signal-to-Noise Ratio, SSIM is the Structural Similarity Index, and LPIPS is the learned perceptual image patch similarity.

[0065] A multi-task joint optimization objective is set, which simultaneously optimizes multiple potentially conflicting objectives, such as a training objective function that considers multiple recovery sub-objectives like fidelity, perceptual quality, and texture naturalness. Multi-scale loss scores are dynamically calculated based on the contribution of the multi-task joint optimization objective, obtaining multiple contribution values ​​based on the ratio of each loss term's current value to its moving average.

[0066] A total loss function is constructed based on multiple contribution levels, thereby establishing an adaptive weight adjustment strategy. This strategy is then applied to the initial image correction and restoration data through multiple rounds of optimization, generating optimized correction results. The optimized correction results are then validated; if the validation passes, the final multi-degradation image correction and restoration data is obtained.

[0067] For example, tests were conducted on a dataset, using PSNR and SSIM as evaluation metrics. The test results are shown in Table 1. Higher values ​​indicate better image restoration. Peak signal-to-noise ratio (PSNR) is an objective metric for measuring image reconstruction quality. A higher value indicates smaller pixel-level errors between the restored image and the original clear image, resulting in better quality. The unit is decibels (dB). Structural similarity is another objective metric for measuring image quality, more in line with human visual perception. It assesses the degree to which an image retains structural information. A higher value, closer to 1, indicates better preservation of the image's structure. Restormer is a high-performance image restoration model based on the Transformer architecture. By using a multi-head attention mechanism and a gated feedforward network, it effectively models the long-range dependencies of images, achieving leading performance on multiple restoration tasks. OKNet is a dedicated image restoration network that focuses on model efficiency or knowledge distillation. In the table, its performance is slightly lower than Restormer. AirNet is an earlier model or one focused on specific weather conditions, such as defogging models. Its general restoration capabilities may not be as good as the aforementioned models. PromptIR is a pioneering method that introduced the concept of prompts to the field of image restoration for the first time. By inputting a prompt representing the type of degradation, such as rain, it guides the model to perform targeted restoration, which is an important step in achieving adaptive restoration. TransWeather is a Transformer model specifically designed to eliminate various weather degradations, such as rain, snow, and fog, aiming to solve multiple weather problems with a unified model.

[0068] Table 1. Performance of each model on a specific dataset

[0069] Through comprehensive quality assessment and multi-task joint optimization, combined with dynamic contribution calculation and weight adjustment strategies, the system automatically identifies the weak links in the current restoration results and dynamically adjusts the training focus, so that the final multi-degraded image correction and restoration data achieves an optimal balance in various quality indicators, which is both clear and realistic, as well as natural and comfortable, thus truly meeting the needs for high-quality image restoration in real-world complex environments.

[0070] In summary, the image correction and restoration method based on a large multimodal model of images and text provided in this application has the following technical effects: A multidimensional image dataset is obtained by traversing the target monitoring area to acquire multidimensional images. A degraded image-text multimodal model is constructed, and degradation analysis is performed on the multidimensional image dataset using this model to extract degraded image features and degraded text features. The degraded image features and degraded text features are then fused, and the fusion result is synchronized to an image restoration model for multi-level interactive reconstruction to generate initial image correction and restoration data. Loss optimization is performed based on the initial image correction and restoration data to generate multi-degraded image correction and restoration data. In other words, by introducing a degraded image-text multimodal model to extract degradation features and performing multi-level interactive reconstruction with an image restoration network, high-quality correction and restoration of multi-degraded images can be achieved without distinguishing between degradation types, thus improving image restoration accuracy.

[0071] Example 2: Based on the same inventive concept as the image correction and restoration method based on a large image-text multimodal model in Example 1, this application also provides an image correction and restoration system based on a large image-text multimodal model. Please refer to the appendix. Figure 2 The image correction and restoration system based on a large image-text multimodal model includes: The multidimensional image acquisition module 11 is used to traverse the target monitoring area to acquire multidimensional images and obtain a multidimensional image dataset; the degradation analysis module 12 is used to construct a large-scale multimodal model of degraded images and text, and to perform degradation analysis on the multidimensional image dataset through the large-scale multimodal model of degraded images and text, extracting degraded image features and degraded text features; the multi-level interactive reconstruction module 13 is used to fuse the degraded image features and the degraded text features, and synchronize the fusion result to the image restoration model for multi-level interactive reconstruction to generate initial image correction and restoration data; the loss optimization module 14 is used to perform loss optimization based on the initial image correction and restoration data to generate multi-degraded image correction and restoration data.

[0072] Furthermore, the multidimensional image acquisition module 11 in the image correction and restoration system based on a large image-text multimodal model is also used for: performing environmental analysis based on the target monitoring area to determine multiple environmental conditions; activating multiple sets of image acquisition devices in the target monitoring area according to the multiple environmental conditions to perform multidimensional image acquisition of the target monitoring area, obtaining multiple sets of original environmental images; traversing the multiple sets of original environmental images for quality screening to generate image screening results; combining text based on the image screening results to construct multiple image-text pairs; and adding the multiple image-text pairs to the multidimensional image dataset.

[0073] Furthermore, the degradation analysis module 12 in the image correction and restoration system based on a multi-modal image-text model is also used for: using the multi-dimensional image dataset to perform end-to-end training on the multi-modal image-text model according to the model network architecture to construct a degraded image-text multi-modal model; synchronizing the multi-dimensional image dataset to the degraded image-text multi-modal model for multi-scale analysis to determine multi-scale features; fusing the multi-scale features according to a feature pyramid network, and performing global average pooling based on the feature fusion result to generate image feature vectors; performing degradation perception based on the image feature vectors to extract degraded image features; synchronizing the multi-dimensional image dataset to the degraded image-text multi-modal model for hidden analysis to determine hidden features; performing attention pooling on the hidden features to generate keyword features, and projecting the keyword features according to the image feature vectors to generate text feature vectors; and performing degradation perception based on the text feature vectors to extract degraded text features.

[0074] Furthermore, the degradation analysis module 12 in the image correction and restoration system based on a large image-text multimodal model is also used to: employ a dual encoder structure, which includes a visual encoder and a text encoder; extract image features from the multidimensional image dataset based on the visual encoder; extract text features from the multidimensional image dataset based on the text encoder; perform semantic analysis on the image features and the text features to construct a shared semantic space; and map the image features and the text features to the shared semantic space for spatial alignment to construct the model network architecture.

[0075] Furthermore, the degradation analysis module 12 in the image correction and restoration system based on a multimodal large model of images and text is also used for: mapping the image features and the text features to the shared semantic space for similarity analysis to construct a similarity matrix; comparing images to text based on the similarity matrix to construct a first contrastive learning loss parameter; comparing text to images based on the similarity matrix to construct a second contrastive learning loss parameter; and combining and aligning the first contrastive learning loss parameter and the second contrastive learning loss parameter to construct the model network architecture.

[0076] Furthermore, the degradation analysis module 12 in the image correction and restoration system based on the image-text multimodal large model is also used for: introducing the image-text multimodal large model, synchronizing the multiple image-text pairs to the image-text multimodal large model in batches according to the model network architecture; performing forward propagation calculation through the image-text multimodal large model to obtain initial image features and initial text features; performing backpropagation based on the initial image features and the initial text features to generate backpropagation information; updating the visual encoder and the text encoder based on the backpropagation information, and performing end-to-end training on the image-text multimodal large model according to the update results to construct the degraded image-text multimodal large model.

[0077] Furthermore, the multi-level interactive reconstruction module 13 in the image correction and restoration system based on a multi-modal large model of image and text is also used for: multi-modal fusion of the degraded image features and the degraded text features to generate degraded enhancement features; synchronizing the degraded enhancement features to the image restoration model for multi-level feature interaction: S1: the multi-level feature interaction includes an encoding stage and a decoding stage; S2: based on the encoding stage, the degraded image features and the degraded text features are subjected to cross-attention interaction to extract degraded semantic features; S3: based on the decoding stage, the degraded enhancement features are subjected to skip connections, and upsampling is performed according to the connection results to obtain an upsampled dataset; S4: image restoration is performed based on the upsampled dataset to generate image detail information, and continuous interaction is performed between the image detail information and the degraded enhancement features to generate the initial image correction and restoration data.

[0078] Furthermore, the multi-level interactive reconstruction module 13 in the image correction and restoration system based on a multi-modal large model is also used to: map the image detail information to the image space for multi-scale parallel analysis to generate a residual image; add the residual image to the multi-dimensional image dataset element by element to generate an initial restoration result; crop the pixel values ​​of the initial restoration result to determine the pixel restoration result; perform progressive interaction between the pixel restoration result and the degradation enhancement features to generate pixel restoration detail interaction content; and perform multi-scale detail reconstruction according to the pixel restoration detail interaction content to generate the initial image correction and restoration data.

[0079] Furthermore, the loss optimization module 14 in the image correction and restoration system based on a multimodal large model of images and text is also used for: performing a comprehensive quality assessment on the initial image correction and restoration data to generate an image quality assessment result; performing multi-scale loss analysis based on the image quality assessment result to generate a multi-scale loss score; setting a multi-task joint optimization objective, and performing dynamic contribution calculation based on the multi-scale loss score according to the multi-task joint optimization objective to obtain multiple contribution degrees; dynamically balancing the initial image correction and restoration data according to the multiple contribution degrees to construct a weight adjustment strategy; executing the weight adjustment strategy to perform multiple rounds of optimization correction on the initial image correction and restoration data to generate an optimized correction result; verifying the optimized correction result, and generating the multi-degraded image correction and restoration data when the verification is successful.

[0080] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The image correction and restoration method and specific examples based on the image-text multimodal large model in the foregoing embodiment 1 are also applicable to the image correction and restoration system based on the image-text multimodal large model in this embodiment. Through the foregoing detailed description of the image correction and restoration method based on the image-text multimodal large model, those skilled in the art can clearly understand the image correction and restoration system based on the image-text multimodal large model in this embodiment. Therefore, for the sake of brevity, it will not be described in detail here.

[0081] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0082] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of this application and its equivalents, this application also intends to include such modifications and variations.

Claims

1. An image correction and restoration method based on a large multimodal model of images and text, characterized in that, include: Multidimensional image acquisition is performed by traversing the target monitoring area to obtain a multidimensional image dataset; A degraded image-text multimodal large model is constructed, and degradation analysis is performed on the multidimensional image dataset using the degraded image-text multimodal large model to extract degraded image features and degraded text features; The degraded image features and the degraded text features are fused, and the fusion result is synchronized to the image restoration model for multi-level interactive reconstruction to generate initial image correction and restoration data, including: The degraded image features and the degraded text features are fused in a multi-modal manner to generate degradation enhancement features; The degradation enhancement features are synchronized to the image restoration model for multi-level feature interaction: S1: The multi-level feature interaction includes an encoding stage and a decoding stage; S2: Based on the encoding stage, the degraded image features and the degraded text features are subjected to cross-attention interaction to extract degraded semantic features; S3: Based on the decoding stage, skip connections are made to the degraded enhancement features, and upsampling is performed according to the connection results to obtain an upsampled dataset; S4: Based on the upsampled dataset, perform image restoration to generate image detail information, and continuously interact with the degradation enhancement features according to the image detail information to generate the initial image correction and restoration data; Loss optimization is performed based on the initial image correction and restoration data to generate multi-degraded image correction and restoration data.

2. The image correction and restoration method based on a large multimodal image-text model as described in claim 1, characterized in that, To obtain a multidimensional image dataset, the method involves traversing the target monitoring area to perform multidimensional image acquisition. Environmental analysis is conducted based on the target monitoring area to determine multiple environmental conditions; According to the aforementioned multiple environmental conditions, multiple sets of image acquisition devices in the target monitoring area are activated to perform multi-dimensional image acquisition of the target monitoring area, thereby obtaining multiple sets of original environmental images. The multiple sets of original environmental images are traversed for quality screening to generate image screening results. Based on the image filtering results, text is combined to construct multiple image-text pairs; Add the multiple image-text pairs to the multidimensional image dataset.

3. The image correction and restoration method based on a large multimodal model of images and text as described in claim 2, characterized in that, A degraded image-text multimodal model is constructed. This model is then used to perform degradation analysis on the multidimensional image dataset, extracting degraded image features and degraded text features. The method includes: The multidimensional image dataset is used to train the large image-text multimodal model end-to-end according to the model network architecture to construct a large image-text multimodal model for degraded images. The multidimensional image dataset is synchronized to the degraded image image-text multimodal large model for multi-scale analysis to determine multi-scale features; The multi-scale features are fused according to the feature pyramid network, and global average pooling is performed based on the feature fusion result to generate image feature vectors. Degradation perception is performed based on the image feature vectors, and degraded image features are extracted; The multidimensional image dataset is synchronized to the degraded image image-text multimodal large model for hidden feature analysis to determine hidden features; Attention pooling is performed on the hidden features to generate keyword features. The keyword features are then projected onto the image feature vector to generate a text feature vector. Degradation is detected based on the text feature vector, and the degraded text features are extracted.

4. The image correction and restoration method based on a large multimodal image-text model as described in claim 3, characterized in that, The process and methods for constructing the model network architecture include: A dual encoder structure is adopted, which includes a visual encoder and a text encoder; Based on the visual encoder, image features are extracted from the multidimensional image dataset; The text encoder is used to extract text features from the multidimensional image dataset. Semantic analysis is performed on the image features and the text features to construct a shared semantic space; The image features and text features are mapped to the shared semantic space for spatial alignment, thereby constructing the model network architecture.

5. The image correction and restoration method based on a large multimodal model of images and text as described in claim 4, characterized in that, The method involves mapping the image features and text features to the shared semantic space for spatial alignment, and constructing the model network architecture. The image features and text features are mapped to the shared semantic space for similarity analysis to construct a similarity matrix; Based on the similarity matrix, the image is compared to the text, and a first contrastive learning loss parameter is constructed; Based on the similarity matrix, text-to-image comparison is performed to construct a second contrastive learning loss parameter; The first contrastive learning loss parameter and the second contrastive learning loss parameter are combined and aligned to construct the model network architecture.

6. The image correction and restoration method based on a large multimodal model of images and text as described in claim 4, characterized in that, The degraded image multimodal model is constructed by end-to-end training of the image-text multimodal model according to the model network architecture using the multidimensional image dataset. The method includes: A large image-text multimodal model is introduced, and the multiple image-text pairs are synchronized to the large image-text multimodal model in batches according to the model's network architecture; The initial image features and initial text features are obtained by performing forward propagation calculations using the aforementioned multimodal image-text model. Backpropagation is performed based on the initial image features and the initial text features to generate backpropagation information; The visual encoder and the text encoder are updated based on the backpropagation information, and the image-text multimodal large model is trained end-to-end according to the update results to construct the degraded image-text multimodal large model.

7. The image correction and restoration method based on a large multimodal image-text model as described in claim 1, characterized in that, The method involves continuously interacting with the image detail information and the degradation enhancement features to generate the initial image correction and restoration data, including: The image detail information is mapped to the image space for multi-scale parallel analysis to generate a residual image; The residual image is added element-wise to the multidimensional image dataset to generate an initial recovery result; The initial recovery result is cropped by pixel values ​​to determine the pixel recovery result; Based on the pixel restoration results and the degradation enhancement features, a progressive interaction is performed to generate pixel restoration detail interaction content; Multi-scale detail reconstruction is performed based on the pixel-level detail recovery interaction content to generate the initial image correction and restoration data.

8. The image correction and restoration method based on a large multimodal model of images and text as described in claim 1, characterized in that, Loss optimization is performed based on the initial image correction and restoration data to generate multi-degradation image correction and restoration data. The method includes: A comprehensive quality assessment is performed on the initial image correction and restoration data to generate image quality assessment results; Based on the image quality assessment results, multi-scale loss analysis is performed to generate a multi-scale loss score. A multi-task joint optimization objective is set, and dynamic contribution calculation is performed based on the multi-scale loss score according to the multi-task joint optimization objective to obtain multiple contribution degrees. The initial image correction and restoration data are dynamically balanced according to the multiple contribution levels to construct a weight adjustment strategy; The weight adjustment strategy is executed to perform multiple rounds of optimization correction on the initial image correction and restoration data, generating optimized correction results; The optimized correction results are verified, and when the verification is successful, the multi-degraded image correction and restoration data is generated.

9. An image correction and restoration system based on a large multimodal model of images and text, characterized in that, The steps for implementing the image correction and restoration method based on a large image-text multimodal model according to any one of claims 1 to 8, wherein the image correction and restoration system based on a large image-text multimodal model comprises: The multidimensional image acquisition module is used to traverse the target monitoring area to acquire multidimensional images and obtain a multidimensional image dataset. The degradation analysis module is used to construct a large multimodal model of degraded images and text, and to perform degradation analysis on the multidimensional image dataset through the large multimodal model of degraded images and text to extract degraded image features and degraded text features; The multi-level interactive reconstruction module is used to fuse the degraded image features with the degraded text features, synchronize the fusion result to the image restoration model for multi-level interactive reconstruction, and generate initial image correction and restoration data. The loss optimization module is used to perform loss optimization based on the initial image correction and restoration data to generate multi-degraded image correction and restoration data.

Citation Information

Patent Citations

  • Degenerated text image recovery method

    CN113222865A

  • Melanoma lesion area segmentation method based on CLIP multi-mode fusion network

    CN120807924A