A universal image restoration system, method and device
By designing a multi-scale feature extraction module and feature enhancement and fusion module, combined with an adaptive multi-modal prompt generation module, the problems of poor generalization capabilities of model and insufficient multi-scale feature extraction and fusion capabilities in the existing technology are solved, and high-quality image restoration effect is achieved.
Patent Information
- Application Number
- CN202411583957.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-11-07
AI Technical Summary
When the prior art deals with multiple image degradation problems, the model has poor generalization ability and is difficult to deal with multiple degradation problems at the same time. The ability to extract and fusion multiple scale features is insufficient, resulting in limited detail recovery effect of the model at different scales.
A general image restoration system is designed, including a multi-scale feature extraction module (MSFE), feature enhancement and fusion module (IEFM), and an adaptive multi-modal prompt generation module. Through multiple depth convolution heads, the self-attention module and a hybrid multi-scale feedforward neural network are transposed to achieve multi-scale extraction and fusion of features, and dynamically adjust the prompt content through the adaptive multi-modal prompt generation module.
It improves the generalization ability and robustness of the system in different tasks, can achieve excellent restoration effects in complex environments, and significantly improves the quality and detail retention capabilities of image restoration.
Smart Images

Figure CN119359580B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a universal image restoration system, method and device. Background Art
[0002] At present, with the continuous advancement of digital image processing technology, image restoration tasks have become increasingly important in many application scenarios, especially when dealing with various image degradation problems. Common image degradation phenomena include blur, noise, low resolution, and uneven brightness, which seriously affect the quality and visual effects of the image.
[0003] Traditional image restoration methods are usually independently designed for a single type of degraded image, resulting in poor generalization of the model and difficulty in dealing with multiple degradation problems at the same time; moreover, the insufficient ability to extract and fuse multi-scale features leads to limited detail restoration effect of the model at different scales; and the lack of an effective multimodal information fusion mechanism limits the further improvement of image detail restoration effect. Summary of the invention
[0004] The object of the present invention is to provide a universal image restoration system, method and device to solve at least one of the above-mentioned technical problems existing in the prior art.
[0005] In a first aspect, in order to solve the above technical problems, the present invention provides a general image restoration system, comprising a first convolutional layer, a first multi-scale feature extraction module (MSFE), a second multi-scale feature extraction module, a third multi-scale feature extraction module, a fourth multi-scale feature extraction module, a fifth multi-scale feature extraction module, a sixth multi-scale feature extraction module, a seventh multi-scale feature extraction module, a first splicing unit, a second splicing unit, a third splicing unit, a first information prompt block, a second information prompt block, a third information prompt block, a first feature enhancement and fusion module (IEFM), a second feature enhancement and fusion module, a third feature enhancement and fusion module, a second convolutional layer and a first addition unit;
[0006] The input end of the first convolutional layer is used to receive the original image of the restored image; the output end of the first convolutional layer is connected to the input end of the first multi-scale feature extraction module;
[0007] The first output end of the first multi-scale feature extraction module is connected to the input end of the second multi-scale feature extraction module; the first output end of the second multi-scale feature extraction module is connected to the input end of the third multi-scale feature extraction module; the output end of the third multi-scale feature extraction module is connected to the input end of the fourth multi-scale feature extraction module; the output end of the fourth multi-scale feature extraction module is connected to the input end of the first information prompt block; the output end of the first information prompt block is connected to the first input end of the first splicing unit; the second input end of the first splicing unit is connected to the second output end of the third multi-scale feature extraction module; the output end of the first splicing unit is connected to the input end of the first feature enhancement and fusion module; the output end of the first feature enhancement and fusion module is connected to the input end of the fifth multi-scale feature extraction module; the output end of the fifth multi-scale feature extraction module is connected to the input end of the second information prompt block; the output end of the second information prompt block is connected to the first input end of the second splicing unit; the second input end of the second splicing unit is connected to the second The second output end of the multi-scale feature extraction module; the output end of the second splicing unit is connected to the input end of the second feature enhancement and fusion module; the output end of the second feature enhancement and fusion module is connected to the input end of the sixth multi-scale feature extraction module; the output end of the sixth multi-scale feature extraction module is connected to the input end of the third information prompt block; the output end of the third information prompt block is connected to the first input end of the third splicing unit; the second input end of the third splicing unit is connected to the second output end of the first multi-scale feature extraction module; the output end of the third splicing unit is connected to the input end of the third feature enhancement and fusion module; the output end of the third feature enhancement and fusion module is connected to the input end of the seventh multi-scale feature extraction module; the output end of the seventh multi-scale feature extraction module is connected to the input end of the second convolutional layer; the output end of the second convolutional layer is connected to the first input end of the first adding unit; the second input end of the first adding unit is connected to the input end of the first convolutional layer; the output end of the first adding unit is used to output a restored image.
[0008] In a feasible implementation, the multi-scale feature extraction module (MSFE) specifically includes: a first normalization layer, a multi-depth convolutional head transposed self-attention module (MDTA), a second normalization layer, and a hybrid multi-scale feedforward neural network (MSFN);
[0009] The input end of the first normalization layer is connected to the input end of the multi-scale feature extraction module; the output end of the first normalization layer is connected to the input end of the multi-depth convolution head transposed self-attention module; the output end of the multi-depth convolution head transposed self-attention module and the input end of the multi-scale feature extraction module are commonly connected to the input end of the second normalization layer; the output end of the second normalization layer is connected to the input end of the hybrid multi-scale feedforward neural network; the output end of the hybrid multi-scale feedforward neural network serves as the output end of the multi-scale feature extraction module;
[0010] Through the above architecture, using the self-attention mechanism of MDTA in the channel dimension, we can flexibly select important features, and use MSFN to effectively fuse features of different scales, so that we can retain the fine-grained details of the image and extract high-level semantic information. Through dynamic adjustment and optimization, we not only improve the generalization ability of this system in different tasks, but also enhance the robustness of this system, so that it can achieve excellent restoration effects in complex environments.
[0011] In a feasible implementation, the multiple depth convolution head transposed self-attention module (MDTA) specifically includes a third normalization layer, a third convolution layer, a fourth convolution layer, a fifth convolution layer, a sixth convolution layer, a seventh convolution layer, an eighth convolution layer, a ninth convolution layer, a first reshaping layer, a second reshaping layer, a third reshaping layer, a fourth reshaping layer, a first multiplication unit, a second multiplication unit and a second addition unit;
[0012] The input end of the third normalization layer is connected to the input end of the transposed self-attention module of the multiple depth convolution head, and is used to perform linear normalization processing on the previous sequence data; the output end of the third normalization layer is respectively connected to the input end of the third convolution layer, the input end of the fifth convolution layer, and the input end of the seventh convolution layer;
[0013] The output end of the third convolutional layer is connected to the input end of the fourth convolutional layer for calculating the query ; The output end of the fourth convolutional layer is connected to the input end of the first reshaping layer; the output end of the first reshaping layer is connected to the first input end of the first multiplication unit;
[0014] The output end of the fifth convolutional layer is connected to the input end of the sixth convolutional layer to calculate the key ; The output end of the sixth convolutional layer is connected to the input end of the second reshaping layer; the output end of the second reshaping layer is connected to the second input end of the first multiplication unit; the output end of the first multiplication unit is connected to the first input end of the second multiplication unit for matrix multiplication;
[0015] The output end of the seventh convolutional layer is connected to the input end of the eighth convolutional layer to calculate the value ; The output end of the eighth convolutional layer is connected to the input end of the third reshaping layer; the output end of the third reshaping layer is connected to the second input end of the second multiplication unit; the output end of the second multiplication unit is connected to the input end of the fourth reshaping layer for matrix multiplication; the output end of the fourth reshaping layer is connected to the input end of the ninth convolutional layer; the output end of the ninth convolutional layer is connected to the first input end of the second addition unit for element-wise addition; the second input end of the second addition unit is connected to the input end of the third normalization layer; the output end of the second addition unit is connected to the output end of the transposed self-attention module of the multiple depth convolution head for outputting data in the next order;
[0016] Through the above architecture, important features can be flexibly selected in the channel dimension.
[0017] In a feasible implementation manner, the first information prompt block specifically includes: a first adaptive multimodal prompt generation module (AMPG) and a first prompt interaction module (PIM);
[0018] The input end of the first adaptive multimodal prompt generation module is connected to the input end of the first information prompt block; the output end of the first adaptive multimodal prompt generation module is connected to the input end of the first prompt interaction module; the input end of the first prompt interaction module is also connected to the input end of the first information prompt block; the output end of the first prompt interaction module is connected to the output end of the first information prompt block; in this way, the detailed information before and after processing by the adaptive multimodal prompt generation module can be taken into account.
[0019] In a feasible implementation manner, the first adaptive multimodal prompt generation module (AMPG) specifically includes a visual information prompt branch, a text information prompt branch, a third adding unit and a tenth convolutional layer;
[0020] The input end of the visual information prompting branch is used to receive the visual information of the previous sequence; the output end of the visual information prompting branch is connected to the first input end of the third adding unit;
[0021] The input end of the text information prompt branch is used to receive a preset task label (for example, denoising, defogging, and rain removal); the output end of the text information prompt branch is connected to the second input end of the third adding unit;
[0022] The output end of the third adding unit is connected to the input end of the tenth convolutional layer; the output end of the tenth convolutional layer serves as the output end of the adaptive multimodal prompt generation module;
[0023] Through the above architecture, visual information and semantic information can be deeply integrated, improving the system's poor convergence and information redundancy problems and improving the efficiency of information use.
[0024] In a feasible implementation, the visual information prompt branch specifically includes: a first global average pooling layer (GAP), an eleventh convolution layer, a first sigmod function layer, a first linear combination unit, a second global average pooling layer, a twelfth convolution layer, a first activation function layer, a thirteenth convolution layer, a first global maximum pooling layer (GMP), a fourteenth convolution layer, a second activation function layer, a fifteenth convolution layer, a fourth addition unit, a second sigmod function layer and a third multiplication unit;
[0025] The input end of the first global average pooling layer is used to receive the visual information of the previous order; the output end of the first global average pooling layer is connected to the input end of the eleventh convolutional layer; the output end of the eleventh convolutional layer is connected to the input end of the first sigmod function layer; the output end of the first sigmod function layer is connected to the first input end of the first linear combination unit; the second input end of the first linear combination unit is used to receive (learnable) parameters; the output end of the first linear combination unit is respectively connected to the input end of the second global average pooling layer and the input end of the first global maximum pooling layer;
[0026] The output end of the second global average pooling layer is connected to the input end of the twelfth convolutional layer; the output end of the twelfth convolutional layer is connected to the input end of the first activation function layer; the output end of the first activation function layer is connected to the input end of the thirteenth convolutional layer; the output end of the thirteenth convolutional layer is connected to the first input end of the fourth adding unit;
[0027] The output end of the first global maximum pooling layer is connected to the input end of the fourteenth convolutional layer; the output end of the fourteenth convolutional layer is connected to the input end of the second activation function layer; the output end of the second activation function layer is connected to the input end of the fifteenth convolutional layer; the output end of the fifteenth convolutional layer is connected to the second input end of the fourth adding unit;
[0028] The output end of the fourth adding unit is connected to the input end of the second sigmod function layer; the output end of the second sigmod function layer is connected to the first input end of the third multiplying unit; the second input end of the third multiplying unit is also connected to the output end of the first linear combination unit for element-by-element multiplication processing; the output end of the third multiplying unit is used as the output end of the visual information prompting branch;
[0029] Through the above architecture, adaptive adjustments can be made to different visual information, so that the weight ratio of each module can be dynamically adjusted according to the characteristics of different tasks. The parameter configuration of the model can also be flexibly optimized according to the diversity and complexity of the input information.
[0030] In a feasible implementation, the text information prompt branch specifically includes: a CLIP model, a sixteenth convolution layer, a fourth normalization layer, a third activation function layer, a seventeenth convolution layer, a fourth activation function layer, an eighteenth convolution layer, a third global average pooling layer, a second global maximum pooling layer, a fourth splicing unit, a nineteenth convolution layer, a third sigmod function layer and a fourth multiplication unit;
[0031] The input end of the CLIP model (Contrastive Language-Image Pre-Training, a deep learning model that can be pre-trained from unlabeled image and text data by self-supervised contrast learning, so that the model can understand the semantic connection between images and texts) is used to receive a preset task label; the output end of the CLIP model is connected to the input end of the sixteenth convolutional layer; the output end of the sixteenth convolutional layer is connected to the input end of the fourth normalization layer; the output end of the fourth normalization layer is connected to the input end of the third activation function layer; the output end of the third activation function layer is connected to the input end of the seventeenth convolutional layer; the output end of the seventeenth convolutional layer is connected to the input end of the fourth activation function layer; the output end of the fourth activation function layer is connected to the input end of the eighteenth convolutional layer; the output end of the eighteenth convolutional layer is respectively connected to the input end of the third global average pooling layer and the input end of the second global maximum pooling layer;
[0032] The output end of the third global average pooling layer is connected to the first input end of the fourth splicing unit; the output end of the second global maximum pooling layer is connected to the second input end of the fourth splicing unit for dimensional splicing; the output end of the fourth splicing unit is connected to the input end of the nineteenth convolutional layer; the output end of the nineteenth convolutional layer is connected to the input end of the third sigmod function layer; the output end of the third sigmod function layer is connected to the first input end of the fourth multiplication unit; the second input end of the fourth multiplication unit is connected to the output end of the eighteenth convolutional layer for element-wise multiplication; the output end of the fourth multiplication unit is used as the output end of the text information prompt branch;
[0033] Through the above architecture, specific task labels (prompt texts) can be introduced for different restoration tasks, so that the system can perform targeted restoration for different types of image distortions.
[0034] In a feasible implementation manner, the first prompt interaction module (PIM) specifically includes: a fifth splicing unit, an eighth multi-scale feature extraction module, a twentieth convolutional layer, and a twenty-first convolutional layer;
[0035] The fifth splicing unit is used to splice the feature information extracted by the fourth multi-scale feature extraction module and the output result of the first adaptive multi-mode prompt generation module by dimension, and input them to the input end of the eighth multi-scale feature extraction module; the output end of the eighth multi-scale feature extraction module is connected to the input end of the twentieth convolutional layer; the output end of the 20th convolutional layer is connected to the input end of the twenty-first convolutional layer; the output end of the twenty-first convolutional layer is used as the output end of the first information prompt block.
[0036] Based on the architecture of the first information prompt block, the structures of the second and third information prompt blocks can be obtained by analogy.
[0037] In a feasible implementation, the hybrid multi-scale feedforward neural network (MSFN) specifically includes: a fifth normalization layer, a twenty-second convolution layer, a twenty-third convolution layer, a fifth activation function layer, a sixth splicing unit, a twenty-fourth convolution layer, a sixth activation function layer, a twenty-fifth convolution layer, a twenty-sixth convolution layer, a seventh activation function layer, a seventh splicing unit, a twenty-seventh convolution layer, an eighth activation function layer, an eighth splicing unit, a twenty-eighth convolution layer and a fifth addition unit;
[0038] The input end of the fifth normalization layer is used as the input end of the hybrid multi-scale feedforward neural network; the output end of the fifth normalization layer is connected to the input end of the twenty-second convolutional layer and the input end of the twenty-fifth convolutional layer respectively;
[0039] The output end of the 22nd convolutional layer is connected to the input end of the 24th convolutional layer; the output end of the 24th convolutional layer is connected to the input end of the fifth activation function layer; the output end of the fifth activation function layer is connected to the first input end of the sixth splicing unit;
[0040] The output end of the 25th convolutional layer is connected to the input end of the 26th convolutional layer; the output end of the 26th convolutional layer is connected to the input end of the seventh activation function; the output end of the seventh activation function is connected to the first input end of the seventh splicing unit;
[0041] The second input end of the sixth splicing unit is connected to the output end of the seventh activation function for dimensional splicing; the output end of the sixth splicing unit is connected to the input end of the twenty-fourth convolutional layer; the output end of the twenty-fourth convolutional layer is connected to the input end of the sixth activation function layer; the output end of the sixth activation function layer is connected to the first input end of the eighth splicing unit;
[0042] The second input end of the seventh splicing unit is also connected to the output end of the fifth activation function; the output end of the seventh splicing unit is connected to the input end of the twenty-seventh convolutional layer; the output end of the twenty-seventh convolutional layer is connected to the input end of the eighth activation function layer; the output end of the eighth activation function layer is connected to the first input end of the eighth splicing unit;
[0043] The output end of the eighth splicing unit is connected to the input end of the twenty-eighth convolutional layer; the output end of the twenty-eighth convolutional layer is connected to the first input end of the fifth adding unit; the second input end of the fifth adding unit is connected to the input end of the fifth normalization layer; the output end of the fifth adding unit serves as the output end of the hybrid multi-scale feedforward neural network.
[0044] In a feasible implementation manner, each feature enhancement and fusion module (IEFM) specifically includes a first branch, a second branch, a third branch and a sixth adding unit;
[0045] The first branch includes a twenty-ninth convolutional layer;
[0046] The second branch includes a 30th convolutional layer, a 31st convolutional layer, a 9th activation function layer, a 32nd convolutional layer, a 4th sigmod function layer and a 5th multiplication unit connected in sequence; the 5th multiplication unit is also provided with a second input end connected to the output end of the 29th convolutional layer for performing element-by-element multiplication calculation;
[0047] The third branch includes a fourth global average pooling layer, a thirty-third convolutional layer, a tenth activation function layer, a thirty-fourth convolutional layer, a fifth sigmod function layer and a sixth multiplication unit connected in sequence; the sixth multiplication unit is also provided with a second input end connected to the second output end of the fifth multiplication unit for performing element-by-element multiplication calculation;
[0048] The first input end of the sixth adding unit is connected to the first output end of the fifth multiplying unit; the second input end of the sixth adding unit is connected to the output end of the sixth multiplying unit; the output end of the sixth adding unit serves as the output end of the feature enhancement and fusion module;
[0049] Through the above architecture, the contextual information of the spatial dimension and the complex nonlinear relationship between features can be captured, which avoids the loss of multi-scale information and improves the quality of image restoration.
[0050] In the second aspect, based on the same inventive concept, the present application also provides a general image restoration method, which specifically includes:
[0051] Step 1: Degrade the image , input the convolution layer, perform preprocessing operations, and obtain the preprocessed image ;
[0052] Step 2: Preprocess the image Perform multi-scale feature extraction, including:
[0053] Through linear normalization, normalization is performed; then through multiple deep convolution heads, the self-attention mechanism is transposed, and convolution and deep convolution operations are performed to obtain the query ,key and value ; respectively for query ,key and value Perform a reshape operation; then query and key , perform matrix multiplication and activation function (softmax) calculation to obtain the attention feature map ; Attention feature map With value , perform matrix multiplication and reshape operation to obtain the first feature map ; For the first feature map After convolution, the preprocessed image Perform element-wise addition to obtain the second feature map ;
[0054] For the second feature map Linear normalization is performed to obtain shallow feature information; then different convolution operations are performed on the shallow feature information through a hybrid multi-scale feedforward network (MSFN) to obtain the third feature map And the fourth characteristic diagram ; For the third feature map And the fourth characteristic diagram , respectively perform activation function (Relu) and splicing by dimension to obtain the fifth feature map And the sixth characteristic diagram ;
[0055] For the fifth feature map Perform convolution and activation function (Relu) operations to obtain the seventh feature map ;
[0056] For the sixth characteristic graph Perform convolution and activation function (Relu) operations to obtain the eighth feature map ;
[0057] The seventh characteristic map And the eighth characteristic diagram , after splicing by dimension, convolution is performed to obtain the ninth feature map ;
[0058] The ninth characteristic map With the second feature map , add the elements together to get the tenth feature map ; Downsample the tenth feature map to obtain the eleventh feature map ;
[0059] Step 3: Iterate step 2 three times to obtain the twelfth, thirteenth and fourteenth feature maps. ;
[0060] Step 4: For the fourteenth feature map , extracting feature information that integrates visual information and text information, including:
[0061] Perform global average pooling, convolution and activation function (softmax) operations in sequence to obtain the first feature weight ; Weight the first feature with (learnable) parameters , perform element-wise multiplication to obtain visual cue feature information ; Visual cue feature information First, global average pooling and global maximum pooling are performed, and then convolution operations are performed respectively, and then element-wise addition is performed, and then the visual information prompt weight is calculated by the softmax function. ; Weight the visual information Feature information with visual cues , multiply by elements to get the parameters of visual cue feature information ;
[0062] The preset task labels are passed through the pre-trained CLIP model, convolution operation and activation function (GELU) to obtain text prompt feature information ; After performing global average pooling and global maximum pooling respectively, splicing is performed by dimension; and then the semantic information prompt weight is obtained through convolution operation and softmax function calculation. ; Weight the semantic information Feature information with text prompts , multiply by elements to get the parameters of text prompt feature information ;
[0063] The parameters of the visual cue feature information Parameters with text prompt feature information , add them by dimension and perform convolution operation to get the final generated parameters ;
[0064] The parameters With the fourteenth characteristic figure , prompt interaction: perform dimension splicing and multi-scale feature extraction; then obtain the first feature information through convolution and upsampling operations ;
[0065] Step 5: The first feature information With the thirteenth characteristic figure , splicing by dimension to obtain the fifteenth feature map; then performing image enhancement and fusion operations, specifically including:
[0066] Perform convolution operation on the fifteenth feature map to obtain the second feature information ;
[0067] For the fifteenth feature map, the first dynamic weight is obtained through convolution and activation function (softmax) calculation ;
[0068] Then perform global average pooling, convolution and activation function (softmax) calculation on the fifteenth feature map to obtain the second dynamic weight ;
[0069] The second characteristic information , first dynamic weight and the second dynamic weight , perform fusion, and obtain the first fusion result , expressed by the formula:
[0070] ;
[0071] Step 6: The first fusion result , through multi-scale feature extraction, decoding operation is performed to obtain the second fusion result ;
[0072] Step 7: The second fusion result , refer to step 4 to obtain the third feature information; combine the third feature information with the twelfth feature map, refer to steps 5-6 to obtain a third fusion result;
[0073] The third fusion result is referred to step 4 to obtain the fourth feature information; the fourth feature information and the eleventh feature map are referred to steps 5-6 to obtain the fourth fusion result ;
[0074] The fourth fusion result , after multiple multi-scale feature extractions, a convolution operation is performed to obtain the fifth fusion result;
[0075] Combine the fifth fusion result with the degraded image Add elements by element to get the restored image .
[0076] On the third aspect, based on the same inventive concept, the present application also provides a universal image restoration device, including a processor, a memory and a bus, wherein the memory stores instructions and data read by the processor, and the processor is used to call the instructions and data in the memory to implement the universal image restoration system as described above, and the bus connects the functional components for transmitting information.
[0077] By adopting the above technical solution, the present invention has the following beneficial effects:
[0078] The present invention provides a universal image restoration system, method and device, which can restore various image degradation problems and have the characteristics of good versatility (All-in-One), good adaptability and strong generalization ability. The multi-scale feature extraction module of the present solution can capture the details and structural information of the image at different scales, effectively solving the key information loss problem caused by the single scale in the background technology, thereby greatly improving the quality of image restoration and the detail retention ability. The adaptive multi-mode prompt generation module of the present solution can more accurately identify and process different image degradation problems by combining visual and semantic prompt information and dynamically adjusting the prompt content, thus overcoming the defect of single information in the background technology. The image enhancement and fusion module of the present solution can effectively reduce information loss, significantly enhance the restoration effect, and improve the detail integrity and visual quality of the restored image. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings in the following description are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0080] Figure 1 A general image restoration system architecture diagram provided by an embodiment of the present invention;
[0081] Figure 2 The architecture diagram of the multi-depth convolution head transposed self-attention module provided by the embodiment of the present invention;
[0082] Figure 3 An architecture diagram of an adaptive multimodal prompt generation module provided in an embodiment of the present invention;
[0083] Figure 4 A diagram of the prompt interaction module architecture provided by an embodiment of the present invention;
[0084] Figure 5 A diagram of the architecture of a hybrid multi-scale feedforward neural network provided in an embodiment of the present invention;
[0085] Figure 6 A diagram of the architecture of a feature enhancement and fusion module provided in an embodiment of the present invention;
[0086] Figure 7 A comparison diagram of rain removal effects provided by an embodiment of the present invention: Figure a is a degraded image containing traces of white raindrops; Figure b is a restored image after removing the white raindrops;
[0087] Figure 8 A comparison diagram of the defogging effect provided by an embodiment of the present invention: Figure a is a degraded image containing traces of white fog; Figure b is a restored image after removing the white fog. DETAILED DESCRIPTION
[0088] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0089] In the description of the present invention, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only, and cannot be understood as indicating or implying relative importance.
[0090] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0091] The present invention is further explained below in conjunction with specific implementation modes.
[0092] It should also be noted that the following specific embodiments or specific implementations are a series of optimized settings listed in the present invention to further explain the specific content of the invention, and these settings can be used in combination or in association with each other.
[0093] Embodiment 1:
[0094] like Figure 1 As shown, a general image restoration system provided by the present embodiment includes a first convolution layer (the convolution kernel size is 3*3), a first multi-scale feature extraction module (MSFE), a second multi-scale feature extraction module, a third multi-scale feature extraction module, a fourth multi-scale feature extraction module, a fifth multi-scale feature extraction module, a sixth multi-scale feature extraction module, a seventh multi-scale feature extraction module (the operation is repeated twice to better restore the detail information in the image and improve the system's restoration effect on various types of degradation), a first splicing unit, a second splicing unit, a third splicing unit, a first information prompt block, a second information prompt block, a third information prompt block, a first feature enhancement and fusion module (IEFM), a second feature enhancement and fusion module, a third feature enhancement and fusion module, a second convolution layer (size is 3*3) and a first adding unit;
[0095] The input end of the first convolutional layer is used to receive the original image of the restored image; the output end of the first convolutional layer is connected to the input end of the first multi-scale feature extraction module;
[0096] The first output end of the first multi-scale feature extraction module is connected to the input end of the second multi-scale feature extraction module; the first output end of the second multi-scale feature extraction module is connected to the input end of the third multi-scale feature extraction module; the output end of the third multi-scale feature extraction module is connected to the input end of the fourth multi-scale feature extraction module; the output end of the fourth multi-scale feature extraction module is connected to the input end of the first information prompt block; the output end of the first information prompt block is connected to the first input end of the first splicing unit; the second input end of the first splicing unit is connected to the second output end of the third multi-scale feature extraction module; the output end of the first splicing unit is connected to the input end of the first feature enhancement and fusion module; the output end of the first feature enhancement and fusion module is connected to the input end of the fifth multi-scale feature extraction module; the output end of the fifth multi-scale feature extraction module is connected to the input end of the second information prompt block; the output end of the second information prompt block is connected to the first input end of the second splicing unit; the second input end of the second splicing unit is connected to the second The second output end of the multi-scale feature extraction module; the output end of the second splicing unit is connected to the input end of the second feature enhancement and fusion module; the output end of the second feature enhancement and fusion module is connected to the input end of the sixth multi-scale feature extraction module; the output end of the sixth multi-scale feature extraction module is connected to the input end of the third information prompt block; the output end of the third information prompt block is connected to the first input end of the third splicing unit; the second input end of the third splicing unit is connected to the second output end of the first multi-scale feature extraction module; the output end of the third splicing unit is connected to the input end of the third feature enhancement and fusion module; the output end of the third feature enhancement and fusion module is connected to the input end of the seventh multi-scale feature extraction module; the output end of the seventh multi-scale feature extraction module is connected to the input end of the second convolutional layer; the output end of the second convolutional layer is connected to the first input end of the first adding unit; the second input end of the first adding unit is connected to the input end of the first convolutional layer; the output end of the first adding unit is used to output a restored image.
[0097] Furthermore, the multi-scale feature extraction module (MSFE) specifically includes: a first normalization layer, a multiple deep convolutional head transposed self-attention module (MDTA), a second normalization layer and a hybrid multi-scale feedforward neural network (MSFN);
[0098] The input end of the first normalization layer is connected to the input end of the multi-scale feature extraction module; the output end of the first normalization layer is connected to the input end of the multi-depth convolution head transposed self-attention module; the output end of the multi-depth convolution head transposed self-attention module and the input end of the multi-scale feature extraction module are commonly connected to the input end of the second normalization layer; the output end of the second normalization layer is connected to the input end of the hybrid multi-scale feedforward neural network; the output end of the hybrid multi-scale feedforward neural network serves as the output end of the multi-scale feature extraction module;
[0099] Through the above architecture, using the self-attention mechanism of MDTA in the channel dimension, we can flexibly select important features, and use MSFN to effectively fuse features of different scales, so that we can retain the fine-grained details of the image and extract high-level semantic information. Through dynamic adjustment and optimization, we not only improve the generalization ability of this system in different tasks, but also enhance the robustness of this system, so that it can achieve excellent restoration effects in complex environments.
[0100] Furthermore, if Figure 2 As shown, the multiple depth convolution head transposed self-attention module (MDTA) specifically includes a third normalization layer, a third convolution layer (size is 1*1), a fourth convolution layer (size is 3*3), a fifth convolution layer (size is 1*1), a sixth convolution layer (size is 3*3), a seventh convolution layer (size is 1*1), an eighth convolution layer (size is 3*3), a ninth convolution layer (size is 1*1), a first reshaping layer, a second reshaping layer, a third reshaping layer, a fourth reshaping layer, a first multiplication unit, a second multiplication unit and a second addition unit;
[0101] The input end of the third normalization layer is connected to the input end of the transposed self-attention module of the multiple depth convolution head, and is used to perform linear normalization processing on the previous sequence data; the output end of the third normalization layer is respectively connected to the input end of the third convolution layer, the input end of the fifth convolution layer, and the input end of the seventh convolution layer;
[0102] The output end of the third convolutional layer is connected to the input end of the fourth convolutional layer for calculating the query ; The output end of the fourth convolutional layer is connected to the input end of the first reshaping layer; the output end of the first reshaping layer is connected to the first input end of the first multiplication unit;
[0103] The output end of the fifth convolutional layer is connected to the input end of the sixth convolutional layer to calculate the key ; The output end of the sixth convolutional layer is connected to the input end of the second reshaping layer; the output end of the second reshaping layer is connected to the second input end of the first multiplication unit; the output end of the first multiplication unit is connected to the first input end of the second multiplication unit for matrix multiplication;
[0104] The output end of the seventh convolutional layer is connected to the input end of the eighth convolutional layer to calculate the value ; The output end of the eighth convolutional layer is connected to the input end of the third reshaping layer; the output end of the third reshaping layer is connected to the second input end of the second multiplication unit; the output end of the second multiplication unit is connected to the input end of the fourth reshaping layer for matrix multiplication; the output end of the fourth reshaping layer is connected to the input end of the ninth convolutional layer; the output end of the ninth convolutional layer is connected to the first input end of the second addition unit for element-wise addition; the second input end of the second addition unit is connected to the input end of the third normalization layer; the output end of the second addition unit is connected to the output end of the transposed self-attention module of the multiple depth convolution head for outputting data in the next order;
[0105] Through the above architecture, important features can be flexibly selected in the channel dimension.
[0106] Further, the first information prompt block specifically includes: a first adaptive multimodal prompt generation module (AMPG) and a first prompt interaction module (PIM);
[0107] The input end of the first adaptive multimodal prompt generation module is connected to the input end of the first information prompt block; the output end of the first adaptive multimodal prompt generation module is connected to the input end of the first prompt interaction module; the input end of the first prompt interaction module is also connected to the input end of the first information prompt block; the output end of the first prompt interaction module is connected to the output end of the first information prompt block; in this way, the detailed information before and after processing by the adaptive multimodal prompt generation module can be taken into account.
[0108] Furthermore, if Figure 3 As shown, the first adaptive multimodal prompt generation module (AMPG) specifically includes a visual information prompt branch, a text information prompt branch, a third adding unit and a tenth convolutional layer (size is 3*3);
[0109] The input end of the visual information prompting branch is used to receive the visual information of the previous sequence; the output end of the visual information prompting branch is connected to the first input end of the third adding unit;
[0110] The input end of the text information prompt branch is used to receive a preset task label (for example, denoising, defogging, and rain removal); the output end of the text information prompt branch is connected to the second input end of the third adding unit;
[0111] The output end of the third adding unit is connected to the input end of the tenth convolutional layer; the output end of the tenth convolutional layer serves as the output end of the adaptive multimodal prompt generation module;
[0112] Through the above architecture, visual information and semantic information can be deeply integrated, improving the system's poor convergence and information redundancy problems and improving the efficiency of information use.
[0113] Furthermore, the visual information prompt branch specifically includes: a first global average pooling layer (GAP), an eleventh convolutional layer (size is 1*1), a first sigmod function layer, a first linear combination unit, a second global average pooling layer, a twelfth convolutional layer (size is 1*1), a first activation function layer, a thirteenth convolutional layer (size is 1*1), a first global maximum pooling layer (GMP), a fourteenth convolutional layer (size is 1*1), a second activation function layer, a fifteenth convolutional layer (size is 1*1), a fourth addition unit, a second sigmod function layer and a third multiplication unit;
[0114] The input end of the first global average pooling layer is used to receive the visual information of the previous order; the output end of the first global average pooling layer is connected to the input end of the eleventh convolutional layer; the output end of the eleventh convolutional layer is connected to the input end of the first sigmod function layer; the output end of the first sigmod function layer is connected to the first input end of the first linear combination unit; the second input end of the first linear combination unit is used to receive learnable parameters; the output end of the first linear combination unit is respectively connected to the input end of the second global average pooling layer and the input end of the first global maximum pooling layer;
[0115] The output end of the second global average pooling layer is connected to the input end of the twelfth convolutional layer; the output end of the twelfth convolutional layer is connected to the input end of the first activation function layer; the output end of the first activation function layer is connected to the input end of the thirteenth convolutional layer; the output end of the thirteenth convolutional layer is connected to the first input end of the fourth adding unit;
[0116] The output end of the first global maximum pooling layer is connected to the input end of the fourteenth convolutional layer; the output end of the fourteenth convolutional layer is connected to the input end of the second activation function layer; the output end of the second activation function layer is connected to the input end of the fifteenth convolutional layer; the output end of the fifteenth convolutional layer is connected to the second input end of the fourth adding unit;
[0117] The output end of the fourth adding unit is connected to the input end of the second sigmod function layer; the output end of the second sigmod function layer is connected to the first input end of the third multiplying unit; the second input end of the third multiplying unit is also connected to the output end of the first linear combination unit for element-by-element multiplication processing; the output end of the third multiplying unit is used as the output end of the visual information prompting branch;
[0118] Through the above architecture, adaptive adjustments can be made to different visual information, so that the weight ratio of each module can be dynamically adjusted according to the characteristics of different tasks. The parameter configuration of the model can also be flexibly optimized according to the diversity and complexity of the input information.
[0119] Furthermore, the visual information prompt generation process of the visual information prompt branch can be expressed by the formula:
[0120] ;
[0121] in, Represents input visual information; Indicates Parameters of visual information; Represents the amount of visual information in each set of input; Represents visual cue feature information.
[0122] Furthermore, the text information prompt branch specifically includes: a CLIP model, a sixteenth convolutional layer (size is 1*1), a fourth normalization layer, a third activation function layer, a seventeenth convolutional layer (size is 1*1), a fourth activation function layer, an eighteenth convolutional layer (size is 1*1), a third global average pooling layer, a second global maximum pooling layer, a fourth splicing unit, a nineteenth convolutional layer (size is 7*7), a third sigmod function layer and a fourth multiplication unit;
[0123] The input end of the CLIP model is used to receive a preset task label; the output end of the CLIP model is connected to the input end of the sixteenth convolutional layer; the output end of the sixteenth convolutional layer is connected to the input end of the fourth normalization layer; the output end of the fourth normalization layer is connected to the input end of the third activation function layer; the output end of the third activation function layer is connected to the input end of the seventeenth convolutional layer; the output end of the seventeenth convolutional layer is connected to the input end of the fourth activation function layer; the output end of the fourth activation function layer is connected to the input end of the eighteenth convolutional layer; the output end of the eighteenth convolutional layer is respectively connected to the input end of the third global average pooling layer and the input end of the second global maximum pooling layer;
[0124] The output end of the third global average pooling layer is connected to the first input end of the fourth splicing unit; the output end of the second global maximum pooling layer is connected to the second input end of the fourth splicing unit for dimensional splicing; the output end of the fourth splicing unit is connected to the input end of the nineteenth convolutional layer; the output end of the nineteenth convolutional layer is connected to the input end of the third sigmod function layer; the output end of the third sigmod function layer is connected to the first input end of the fourth multiplication unit; the second input end of the fourth multiplication unit is connected to the output end of the eighteenth convolutional layer for element-wise multiplication; the output end of the fourth multiplication unit is used as the output end of the text information prompt branch;
[0125] Through the above architecture, specific task labels (prompt texts) can be introduced for different restoration tasks, so that the system can perform targeted restoration for different types of image distortions.
[0126] Furthermore, the text information prompt generation process of the text information prompt branch can be expressed by the formula:
[0127] ;
[0128] in, Indicates that the information of each layer is extracted before Gao Heqian Wide information; Indicates the preset task tags, including ;in, Indicates the number of task labels for each set of input; Represents text prompt feature information and belongs to ;in, Indicates the number of image channels; this allows accurate extraction of text feature information through the CLIP model and reshaping clipping operations.
[0129] Furthermore, the specific calculation process of the first adaptive multimodal prompt generation module (AMPG) can be expressed by the formula:
[0130] ;
[0131] ;
[0132] ;
[0133] in, Indicates the dimension-wise splicing operation; Indicates that two 1*1 convolutions were performed consecutively; Parameters representing text prompt feature information; Parameters representing feature information of visual cues; Represents the final generated parameters, which are used to guide the model to dynamically adjust for different tasks.
[0134] Furthermore, if Figure 4 As shown, the first prompt interaction module (PIM) specifically includes: a fifth splicing unit, an eighth multi-scale feature extraction module, a twentieth convolutional layer (size is 3*3) and a twenty-first convolutional layer (size is 1*1);
[0135] The fifth splicing unit is used to splice the feature information extracted by the fourth multi-scale feature extraction module and the output result of the first adaptive multi-mode prompt generation module by dimension, and input them to the input end of the eighth multi-scale feature extraction module; the output end of the eighth multi-scale feature extraction module is connected to the input end of the twentieth convolutional layer; the output end of the 20th convolutional layer is connected to the input end of the twenty-first convolutional layer; the output end of the twenty-first convolutional layer is used as the output end of the first information prompt block.
[0136] Based on the architecture of the first information prompt block, the structures of the second and third information prompt blocks can be obtained by analogy.
[0137] Furthermore, if Figure 5 As shown, the hybrid multi-scale feedforward neural network (MSFN) specifically includes: a fifth normalization layer, a twenty-second convolutional layer (size is 1*1), a twenty-third convolutional layer (size is 3*3), a fifth activation function layer, a sixth splicing unit, a twenty-fourth convolutional layer (size is 3*3), a sixth activation function layer, a twenty-fifth convolutional layer (size is 1*1), a twenty-sixth convolutional layer (size is 7*7), a seventh activation function layer, a seventh splicing unit, a twenty-seventh convolutional layer (size is 7*7), an eighth activation function layer, an eighth splicing unit, a twenty-eighth convolutional layer (size is 1*1) and a fifth addition unit;
[0138] The input end of the fifth normalization layer is used as the input end of the hybrid multi-scale feedforward neural network; the output end of the fifth normalization layer is connected to the input end of the twenty-second convolutional layer and the input end of the twenty-fifth convolutional layer respectively;
[0139] The output end of the 22nd convolutional layer is connected to the input end of the 24th convolutional layer; the output end of the 24th convolutional layer is connected to the input end of the fifth activation function layer; the output end of the fifth activation function layer is connected to the first input end of the sixth splicing unit;
[0140] The output end of the 25th convolutional layer is connected to the input end of the 26th convolutional layer; the output end of the 26th convolutional layer is connected to the input end of the seventh activation function; the output end of the seventh activation function is connected to the first input end of the seventh splicing unit;
[0141] The second input end of the sixth splicing unit is connected to the output end of the seventh activation function for dimensional splicing; the output end of the sixth splicing unit is connected to the input end of the twenty-fourth convolutional layer; the output end of the twenty-fourth convolutional layer is connected to the input end of the sixth activation function layer; the output end of the sixth activation function layer is connected to the first input end of the eighth splicing unit;
[0142] The second input end of the seventh splicing unit is also connected to the output end of the fifth activation function; the output end of the seventh splicing unit is connected to the input end of the twenty-seventh convolutional layer; the output end of the twenty-seventh convolutional layer is connected to the input end of the eighth activation function layer; the output end of the eighth activation function layer is connected to the first input end of the eighth splicing unit;
[0143] The output end of the eighth splicing unit is connected to the input end of the twenty-eighth convolutional layer; the output end of the twenty-eighth convolutional layer is connected to the first input end of the fifth adding unit; the second input end of the fifth adding unit is connected to the input end of the fifth normalization layer; the output end of the fifth adding unit serves as the output end of the hybrid multi-scale feedforward neural network.
[0144] Furthermore, if Figure 6 As shown, each feature enhancement and fusion module (IEFM) specifically includes a first branch, a second branch, a third branch and a sixth adding unit;
[0145] The first branch includes a twenty-ninth convolutional layer (size is 1*1);
[0146] The second branch includes a 30th convolutional layer (size is 3*3), a 31st convolutional layer (size is 3*3), a 9th activation function layer, a 32nd convolutional layer (size is 3*3), a 4th sigmod function layer and a 5th multiplication unit connected in sequence; the 5th multiplication unit is also provided with a second input end connected to the output end of the 29th convolutional layer for performing element-by-element multiplication calculation;
[0147] The third branch includes a fourth global average pooling layer, a thirty-third convolutional layer (size is 1*1), a tenth activation function layer, a thirty-fourth convolutional layer (size is 1*1), a fifth sigmod function layer and a sixth multiplication unit connected in sequence; the sixth multiplication unit is also provided with a second input end connected to the second output end of the fifth multiplication unit for performing element-by-element multiplication calculation;
[0148] The first input end of the sixth adding unit is connected to the first output end of the fifth multiplying unit; the second input end of the sixth adding unit is connected to the output end of the sixth multiplying unit; the output end of the sixth adding unit serves as the output end of the feature enhancement and fusion module;
[0149] Through the above architecture, the contextual information of the spatial dimension and the complex nonlinear relationship between features can be captured, which avoids the loss of multi-scale information and improves the quality of image restoration.
[0150] Furthermore, the calculation process of the feature enhancement and fusion module (IEFM) can be expressed by the formula:
[0151] ;
[0152] ;
[0153] in, Indicates the output result of the previous sequence; Represents feature information; Represents the output result of the feature enhancement and fusion module; this can better learn and fuse the detailed information of the image and better restore the fine structure and texture details in the image.
[0154] Embodiment 2:
[0155] This embodiment provides a general image restoration method based on the first embodiment, which specifically includes:
[0156] Step 1: Degrade the image (Size is ), input convolution layer (size is 3*3), perform preprocessing operation, and obtain preprocessed image ;
[0157] Step 2: Preprocess the image Perform multi-scale feature extraction, including:
[0158] Through linear normalization, normalization is performed; then through the multi-depth convolution head transposition self-attention mechanism, convolution (3 times, size is 1*1) and depth convolution (size is 1*1) operations are performed to obtain the query ,key and value ; respectively for query ,key and value Perform a reshape operation ( The size is ; The size is ; The size is ); and then query and key , perform matrix multiplication and activation function (softmax) calculation to obtain the attention feature map (Size is ); Attention feature map With value , perform matrix multiplication and reshape operation to obtain the first feature map (Size is ); for the first feature map After convolution (size is 1*1), it is then combined with the preprocessed image Perform element-wise addition to obtain the second feature map (Size is );
[0159] For the second feature map Linear normalization is performed to obtain shallow feature information; then different convolution operations (sizes are 3*3 and 7*7) are performed on the shallow feature information through a hybrid multi-scale feedforward network (MSFN) to obtain the third feature map And the fourth characteristic diagram ; For the third feature map And the fourth characteristic diagram , respectively perform activation function (Relu) and splicing by dimension to obtain the fifth feature map And the sixth characteristic diagram ;
[0160] For the fifth feature map Perform convolution (size is 3*3) and activation function (Relu) operations to obtain the seventh feature map ;
[0161] For the sixth characteristic graph Perform convolution (size is 7*7) and activation function (Relu) operations to obtain the eighth feature map ;
[0162] The seventh characteristic map And the eighth characteristic diagram , after splicing by dimension, convolution is performed (size is 1*1) to obtain the ninth feature map ;
[0163] The ninth characteristic map With the second feature map , add the elements together to get the tenth feature map ; Downsample the tenth feature map to obtain the eleventh feature map (Size is );
[0164] Step 3: Iterate step 2 three times to obtain the twelfth, thirteenth and fourteenth feature maps. (Size is );
[0165] Step 4: For the fourteenth feature map , extracting feature information that integrates visual information and text information, including:
[0166] Perform global average pooling, convolution (size is 1*1) and activation function (softmax) operations in sequence to obtain the first feature weight ; Weight the first feature with (learnable) parameters , perform element-wise multiplication to obtain visual cue feature information ; Visual cue feature information First, global average pooling and global maximum pooling are performed, and then (twice) convolution (size is 1*1) operations are performed and then element-wise addition is performed. Finally, the visual information cue weight is calculated by the softmax function. ; Weight the visual information Feature information with visual cues , multiply by elements to get the parameters of visual cue feature information ;
[0167] The preset task label is passed through the pre-trained CLIP model, convolution (size is 1*1) operation and activation function (GELU) to obtain the text prompt feature information ; After performing global average pooling and global maximum pooling respectively, splicing is performed by dimension; and then the semantic information prompt weight is obtained through convolution (size is 7*7) operation and softmax function calculation. ; Weight the semantic information Feature information with text prompts , multiply by elements to get the parameters of text prompt feature information ;
[0168] The parameters of the visual cue feature information Parameters with text prompt feature information , add by dimension and perform convolution (size is 3*3) to get the final generated parameters ;
[0169] The parameters With the fourteenth characteristic figure , Prompt Interaction Module (PIM): performs dimension splicing and multi-scale feature extraction; then obtains the first feature information through (twice) convolution (size is 3*3 and 1*1) and upsampling operation (Size is );
[0170] Step 5: The first feature information With the thirteenth characteristic figure , splicing by dimension to obtain the fifteenth feature map; then performing image enhancement and fusion operations, specifically including:
[0171] Perform convolution (size 1*1) on the fifteenth feature map to obtain the second feature information ;
[0172] For the fifteenth feature map, the first dynamic weight is obtained through (multiple) convolutions (size is 3*3) and activation function (softmax) calculations. ;
[0173] Then, for the fifteenth feature map, perform global average pooling, (multiple) convolutions (size is 1*1) and activation function (softmax) calculations to obtain the second dynamic weight. ;
[0174] The second characteristic information , first dynamic weight and the second dynamic weight , perform fusion, and obtain the first fusion result , expressed by the formula:
[0175] ;
[0176] Step 6: The first fusion result , through multi-scale feature extraction, decoding operation is performed to obtain the second fusion result ;
[0177] Step 7: The second fusion result , refer to step 4 to obtain the third feature information; combine the third feature information with the twelfth feature map, refer to steps 5-6 to obtain a third fusion result;
[0178] The third fusion result is referred to step 4 to obtain the fourth feature information; the fourth feature information and the eleventh feature map are referred to steps 5-6 to obtain the fourth fusion result (Size is );
[0179] The fourth fusion result , after performing multiple (for example, twice) multi-scale feature extraction, a convolution operation (size is 3*3) is performed to obtain the fifth fusion result;
[0180] Combine the fifth fusion result with the degraded image Add elements by element to get the restored image .
[0181] Through the method of this embodiment, an effective rain removal effect can be achieved, such as Figure 7 As shown:
[0182] Among them, Figure a is a degraded image containing traces of white raindrops; Figure b is a restored image after removing the white raindrops.
[0183] Through the method of this embodiment, an effective defogging effect can be achieved, such as Figure 8 As shown:
[0184] Among them, Figure a is a degraded image containing traces of white fog; Figure b is a restored image after removing the white fog.
[0185] Embodiment three:
[0186] This embodiment provides a universal image restoration device, including a processor, a memory and a bus, wherein the memory stores instructions and data read by the processor, the processor is used to call the instructions and data in the memory to implement the universal image restoration system as described above, and the bus connects the functional components to transmit information.
[0187] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A general image restoration system, characterized in that: It includes a first convolution layer, a first multi-scale feature extraction module, a second multi-scale feature extraction module, a third multi-scale feature extraction module, a fourth multi-scale feature extraction module, a fifth multi-scale feature extraction module, a sixth multi-scale feature extraction module, a seventh multi-scale feature extraction module, a first splicing unit, a second splicing unit, a third splicing unit, a first information prompt block, a second information prompt block, a third information prompt block, a first feature enhancement and fusion module, a second feature enhancement and fusion module, a third feature enhancement and fusion module, a second convolution layer and a first addition unit; The input end of the first convolutional layer is used to receive the original image of the restored image; the output end of the first convolutional layer is connected to the input end of the first multi-scale feature extraction module; The first output end of the first multi-scale feature extraction module is connected to the input end of the second multi-scale feature extraction module; the first output end of the second multi-scale feature extraction module is connected to the input end of the third multi-scale feature extraction module; the output end of the third multi-scale feature extraction module is connected to the input end of the fourth multi-scale feature extraction module; the output end of the fourth multi-scale feature extraction module is connected to the input end of the first information prompt block; the output end of the first information prompt block is connected to the first input end of the first splicing unit; the second input end of the first splicing unit is connected to the second output end of the third multi-scale feature extraction module; the output end of the first splicing unit is connected to the input end of the first feature enhancement and fusion module; the output end of the first feature enhancement and fusion module is connected to the input end of the fifth multi-scale feature extraction module; the output end of the fifth multi-scale feature extraction module is connected to the input end of the second information prompt block; the output end of the second information prompt block is connected to the first input end of the second splicing unit; the second input end of the second splicing unit is connected to the second The second output end of the multi-scale feature extraction module; the output end of the second splicing unit is connected to the input end of the second feature enhancement and fusion module; the output end of the second feature enhancement and fusion module is connected to the input end of the sixth multi-scale feature extraction module; the output end of the sixth multi-scale feature extraction module is connected to the input end of the third information prompt block; the output end of the third information prompt block is connected to the first input end of the third splicing unit; the second input end of the third splicing unit is connected to the second output end of the first multi-scale feature extraction module; the output end of the third splicing unit is connected to the input end of the third feature enhancement and fusion module; the output end of the third feature enhancement and fusion module is connected to the input end of the seventh multi-scale feature extraction module; the output end of the seventh multi-scale feature extraction module is connected to the input end of the second convolutional layer; the output end of the second convolutional layer is connected to the first input end of the first adding unit; the second input end of the first adding unit is connected to the input end of the first convolutional layer; the output end of the first adding unit is used to output a restored image.
2. The system according to claim 1, characterized in that The multi-scale feature extraction module specifically includes: a first normalization layer, a multi-depth convolution head transposed self-attention module, a second normalization layer and a hybrid multi-scale feedforward neural network; The input end of the first normalization layer is connected to the input end of the multi-scale feature extraction module; the output end of the first normalization layer is connected to the input end of the multiple deep convolution head transposed self-attention module; the output end of the multiple deep convolution head transposed self-attention module and the input end of the multi-scale feature extraction module are commonly connected to the input end of the second normalization layer; the output end of the second normalization layer is connected to the input end of the hybrid multi-scale feedforward neural network; the output end of the hybrid multi-scale feedforward neural network serves as the output end of the multi-scale feature extraction module.
3. The system according to claim 2, characterized in that The multiple deep convolution head transposed self-attention module specifically includes a third normalization layer, a third convolution layer, a fourth convolution layer, a fifth convolution layer, a sixth convolution layer, a seventh convolution layer, an eighth convolution layer, a ninth convolution layer, a first reshaping layer, a second reshaping layer, a third reshaping layer, a fourth reshaping layer, a first multiplication unit, a second multiplication unit and a second addition unit; The input end of the third normalization layer is connected to the input end of the transposed self-attention module of the multiple depth convolution head, and is used to perform linear normalization processing on the previous sequence data; the output end of the third normalization layer is respectively connected to the input end of the third convolution layer, the input end of the fifth convolution layer, and the input end of the seventh convolution layer; The output end of the third convolutional layer is connected to the input end of the fourth convolutional layer for calculating the query ; The output end of the fourth convolutional layer is connected to the input end of the first reshaping layer; the output end of the first reshaping layer is connected to the first input end of the first multiplication unit; The output end of the fifth convolutional layer is connected to the input end of the sixth convolutional layer to calculate the key ; The output end of the sixth convolutional layer is connected to the input end of the second reshaping layer; the output end of the second reshaping layer is connected to the second input end of the first multiplication unit; the output end of the first multiplication unit is connected to the first input end of the second multiplication unit for matrix multiplication; The output end of the seventh convolutional layer is connected to the input end of the eighth convolutional layer to calculate the value ; The output end of the eighth convolutional layer is connected to the input end of the third reshaping layer; the output end of the third reshaping layer is connected to the second input end of the second multiplication unit; the output end of the second multiplication unit is connected to the input end of the fourth reshaping layer for matrix multiplication; The output end of the fourth reshaping layer is connected to the input end of the ninth convolutional layer; the output end of the ninth convolutional layer is connected to the first input end of the second adding unit for element-wise addition; The second input end of the second adding unit is connected to the input end of the third normalization layer; the output end of the second adding unit is connected to the output end of the transposed self-attention module of the multiple depth convolution head, for outputting data to the next order.
4. The system according to claim 3, characterized in that The first information prompt block specifically includes: a first adaptive multimodal prompt generation module and a first prompt interaction module; The input end of the first adaptive multimodal prompt generation module is connected to the input end of the first information prompt block; the output end of the first adaptive multimodal prompt generation module is connected to the input end of the first prompt interaction module; the input end of the first prompt interaction module is also connected to the input end of the first information prompt block; the output end of the first prompt interaction module is connected to the output end of the first information prompt block.
5. The system according to claim 4, characterized in that The first adaptive multimodal prompt generation module specifically includes a visual information prompt branch, a text information prompt branch, a third adding unit and a tenth convolutional layer; The input end of the visual information prompting branch is used to receive the visual information of the previous sequence; the output end of the visual information prompting branch is connected to the first input end of the third adding unit; The input end of the text information prompt branch is used to receive a preset task label; the output end of the text information prompt branch is connected to the second input end of the third adding unit; The output end of the third adding unit is connected to the input end of the tenth convolutional layer; The output end of the tenth convolutional layer is used as the output end of the adaptive multimodal prompt generation module; The visual information prompt branch specifically includes: a first global average pooling layer, an eleventh convolution layer, a first sigmoid function layer, a first linear combination unit, a second global average pooling layer, a twelfth convolution layer, a first activation function layer, a thirteenth convolution layer, a first global maximum pooling layer, a fourteenth convolution layer, a second activation function layer, a fifteenth convolution layer, a fourth addition unit, a second sigmoid function layer and a third multiplication unit; The input end of the first global average pooling layer is used to receive the visual information of the previous order; the output end of the first global average pooling layer is connected to the input end of the eleventh convolutional layer; the output end of the eleventh convolutional layer is connected to the input end of the first sigmod function layer; the output end of the first sigmod function layer is connected to the first input end of the first linear combination unit; the second input end of the first linear combination unit is used to receive parameters; the output end of the first linear combination unit is respectively connected to the input end of the second global average pooling layer and the input end of the first global maximum pooling layer; The output end of the second global average pooling layer is connected to the input end of the twelfth convolutional layer; the output end of the twelfth convolutional layer is connected to the input end of the first activation function layer; the output end of the first activation function layer is connected to the input end of the thirteenth convolutional layer; the output end of the thirteenth convolutional layer is connected to the first input end of the fourth adding unit; The output end of the first global maximum pooling layer is connected to the input end of the fourteenth convolutional layer; the output end of the fourteenth convolutional layer is connected to the input end of the second activation function layer; the output end of the second activation function layer is connected to the input end of the fifteenth convolutional layer; the output end of the fifteenth convolutional layer is connected to the second input end of the fourth adding unit; The output end of the fourth adding unit is connected to the input end of the second sigmod function layer; the output end of the second sigmod function layer is connected to the first input end of the third multiplying unit; the second input end of the third multiplying unit is also connected to the output end of the first linear combination unit for element-by-element multiplication processing; the output end of the third multiplying unit is used as the output end of the visual information prompting branch; The text information prompt branch specifically includes: a CLIP model, a sixteenth convolution layer, a fourth normalization layer, a third activation function layer, a seventeenth convolution layer, a fourth activation function layer, an eighteenth convolution layer, a third global average pooling layer, a second global maximum pooling layer, a fourth splicing unit, a nineteenth convolution layer, a third sigmod function layer and a fourth multiplication unit; The input end of the CLIP model is used to receive a preset task label; the output end of the CLIP model is connected to the input end of the sixteenth convolutional layer; the output end of the sixteenth convolutional layer is connected to the input end of the fourth normalization layer; the output end of the fourth normalization layer is connected to the input end of the third activation function layer; the output end of the third activation function layer is connected to the input end of the seventeenth convolutional layer; the output end of the seventeenth convolutional layer is connected to the input end of the fourth activation function layer; the output end of the fourth activation function layer is connected to the input end of the eighteenth convolutional layer; the output end of the eighteenth convolutional layer is respectively connected to the input end of the third global average pooling layer and the input end of the second global maximum pooling layer; The output end of the third global average pooling layer is connected to the first input end of the fourth splicing unit; the output end of the second global maximum pooling layer is connected to the second input end of the fourth splicing unit for dimensional splicing; the output end of the fourth splicing unit is connected to the input end of the nineteenth convolutional layer; the output end of the nineteenth convolutional layer is connected to the input end of the third sigmod function layer; the output end of the third sigmod function layer is connected to the first input end of the fourth multiplication unit; the second input end of the fourth multiplication unit is connected to the output end of the eighteenth convolutional layer for element-wise multiplication; the output end of the fourth multiplication unit serves as the output end of the text information prompt branch.
6. The system according to claim 5, characterized in that The first prompt interaction module specifically includes: a fifth splicing unit, an eighth multi-scale feature extraction module, a twentieth convolutional layer and a twenty-first convolutional layer; The fifth splicing unit is used to splice the feature information extracted by the fourth multi-scale feature extraction module and the output result of the first adaptive multi-mode prompt generation module by dimension, and input them to the input end of the eighth multi-scale feature extraction module; the output end of the eighth multi-scale feature extraction module is connected to the input end of the twentieth convolutional layer; the output end of the 20th convolutional layer is connected to the input end of the twenty-first convolutional layer; the output end of the twenty-first convolutional layer is used as the output end of the first information prompt block.
7. The system according to claim 6, characterized in that The hybrid multi-scale feedforward neural network specifically includes: a fifth normalization layer, a twenty-second convolution layer, a twenty-third convolution layer, a fifth activation function layer, a sixth splicing unit, a twenty-fourth convolution layer, a sixth activation function layer, a twenty-fifth convolution layer, a twenty-sixth convolution layer, a seventh activation function layer, a seventh splicing unit, a twenty-seventh convolution layer, an eighth activation function layer, an eighth splicing unit, a twenty-eighth convolution layer and a fifth adding unit; The input end of the fifth normalization layer is used as the input end of the hybrid multi-scale feedforward neural network; the output end of the fifth normalization layer is connected to the input end of the twenty-second convolutional layer and the input end of the twenty-fifth convolutional layer respectively; The output end of the 22nd convolutional layer is connected to the input end of the 24th convolutional layer; the output end of the 24th convolutional layer is connected to the input end of the fifth activation function layer; the output end of the fifth activation function layer is connected to the first input end of the sixth splicing unit; The output end of the 25th convolutional layer is connected to the input end of the 26th convolutional layer; the output end of the 26th convolutional layer is connected to the input end of the seventh activation function; the output end of the seventh activation function is connected to the first input end of the seventh splicing unit; The second input end of the sixth splicing unit is connected to the output end of the seventh activation function for dimensional splicing; the output end of the sixth splicing unit is connected to the input end of the twenty-fourth convolutional layer; the output end of the twenty-fourth convolutional layer is connected to the input end of the sixth activation function layer; the output end of the sixth activation function layer is connected to the first input end of the eighth splicing unit; The second input end of the seventh splicing unit is also connected to the output end of the fifth activation function; the output end of the seventh splicing unit is connected to the input end of the twenty-seventh convolutional layer; the output end of the twenty-seventh convolutional layer is connected to the input end of the eighth activation function layer; the output end of the eighth activation function layer is connected to the first input end of the eighth splicing unit; The output end of the eighth splicing unit is connected to the input end of the twenty-eighth convolutional layer; the output end of the twenty-eighth convolutional layer is connected to the first input end of the fifth adding unit; the second input end of the fifth adding unit is connected to the input end of the fifth normalization layer; the output end of the fifth adding unit serves as the output end of the hybrid multi-scale feedforward neural network.
8. The system according to claim 7, characterized in that Each feature enhancement and fusion module specifically includes a first branch, a second branch, a third branch and a sixth adding unit; The first branch includes a twenty-ninth convolutional layer; The second branch includes a 30th convolutional layer, a 31st convolutional layer, a 9th activation function layer, a 32nd convolutional layer, a 4th sigmod function layer and a 5th multiplication unit connected in sequence; the 5th multiplication unit is also provided with a second input end connected to the output end of the 29th convolutional layer for performing element-by-element multiplication calculation; The third branch includes a fourth global average pooling layer, a thirty-third convolutional layer, a tenth activation function layer, a thirty-fourth convolutional layer, a fifth sigmod function layer and a sixth multiplication unit connected in sequence; the sixth multiplication unit is also provided with a second input end connected to the second output end of the fifth multiplication unit for performing element-by-element multiplication calculation; The first input end of the sixth adding unit is connected to the first output end of the fifth multiplying unit; the second input end of the sixth adding unit is connected to the output end of the sixth multiplying unit; the output end of the sixth adding unit serves as the output end of the feature enhancement and fusion module.
9. A general image restoration method, characterized in that: Specifically include: Step 1: Degrade the image , input the convolution layer, perform preprocessing operations, and obtain the preprocessed image ; Step 2: Preprocess the image Perform multi-scale feature extraction, including: Through linear normalization, normalization is performed; then through multiple deep convolution heads, the self-attention mechanism is transposed, and convolution and deep convolution operations are performed to obtain the query ,key and value ; respectively for query ,key and value Perform a reshape operation; then query and key , perform matrix multiplication and activation function calculation to obtain the attention feature map ; Attention feature map With value , perform matrix multiplication and reshape operation to obtain the first feature map ; For the first feature map After convolution, the preprocessed image Perform element-wise addition to obtain the second feature map ; For the second feature map Linear normalization is performed to obtain shallow feature information; then different convolution operations are performed on the shallow feature information through a hybrid multi-scale feedforward network to obtain the third feature map And the fourth characteristic diagram ; For the third feature map And the fourth characteristic diagram , respectively perform activation functions and splicing by dimension to obtain the fifth feature map And the sixth characteristic diagram ; For the fifth feature map Perform convolution and activation function operations to obtain the seventh feature map ; For the sixth characteristic graph Perform convolution and activation function operations to obtain the eighth feature map ; The seventh characteristic map And the eighth characteristic diagram , after splicing by dimension, convolution is performed to obtain the ninth feature map ; The ninth characteristic map With the second feature map , add element by element to get the tenth feature map ; Downsample the tenth feature map to obtain the eleventh feature map ; Step 3: Iterate step 2 three times to obtain the twelfth, thirteenth and fourteenth feature maps. ; Step 4: For the fourteenth feature map , extracting feature information that integrates visual information and text information, including: Perform global average pooling, convolution and activation function operations in sequence to obtain the first feature weight ; Weight the first feature With parameters , perform element-wise multiplication to obtain visual cue feature information ; Visual cue feature information First, global average pooling and global maximum pooling are performed, and then convolution operations are performed respectively, and then element-wise addition is performed, and then the visual information prompt weight is calculated by the softmax function. ; Weight the visual information Feature information with visual cues , multiply by elements to get the parameters of visual cue feature information ; The preset task labels are used through the pre-trained CLIP model, convolution operation and activation function to obtain text prompt feature information. ; After performing global average pooling and global maximum pooling respectively, splicing is performed by dimension; and then the semantic information prompt weight is obtained through convolution operation and softmax function calculation. ; Give semantic information hint weight Feature information with text prompts , multiply by elements to get the parameters of text prompt feature information ; Parameters that set visual cue feature information Parameters with text prompt feature information , add them by dimension and perform convolution operation to get the final generated parameters ; The parameters With the fourteenth characteristic figure , prompt interaction: perform dimension splicing and multi-scale feature extraction; then obtain the first feature information through convolution and upsampling operations ; Step 5: The first feature information With the thirteenth characteristic figure , splicing by dimension to obtain the fifteenth feature map; then performing image enhancement and fusion operations, specifically including: Perform convolution operation on the fifteenth feature map to obtain the second feature information ; For the fifteenth feature map, the first dynamic weight is obtained through convolution and activation function calculation ; Then, perform global average pooling, convolution and activation function calculation on the fifteenth feature map to obtain the second dynamic weight ; The second characteristic information , first dynamic weight and the second dynamic weight , perform fusion, and obtain the first fusion result , expressed by the formula: ; Step 6: The first fusion result , through multi-scale feature extraction, decoding operation is performed to obtain the second fusion result ; Step 7: The second fusion result , refer to step 4 to obtain the third feature information; combine the third feature information with the twelfth feature map, refer to steps 5-6 to obtain a third fusion result; The third fusion result is referred to step 4 to obtain the fourth feature information; the fourth feature information and the eleventh feature map are referred to steps 5-6 to obtain the fourth fusion result ; The fourth fusion result , after multiple multi-scale feature extractions, a convolution operation is performed to obtain the fifth fusion result; Combine the fifth fusion result with the degraded image Add elements by element to get the restored image .
10. A general image restoration device, characterized in that: It includes a processor, a memory and a bus, wherein the memory stores instructions and data read by the processor, and the processor is used to call the instructions and data in the memory to implement the universal image restoration system as described in any one of claims 1-8, and the bus connects the functional components for transmitting information.
Citation Information
Patent Citations
Fast image segmentation method based on multi-scale Transform
CN117173205A
Image restoration method based on double-prompt guidance Transform
CN118396858A