Whiteboard image enhancement method and device based on multi-scale prompt learning
Through the method based on multi-scale prompt learning, the problem of degradation of whiteboard images in complex scenarios is solved, the image details and quality are improved, and the image enhancement effect is significantly improved.
Patent Information
- Application Number
- CN202510053523.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to effectively solve the problem of degradation of whiteboard images in complex scenarios, especially the lack of feature learning of different types of degradation and degrees of degradation, resulting in poor image enhancement effect.
Using a multi-scale prompt learning method, high-quality whiteboard images are reconstructed through downsampling feature extraction, convolutional attention enhancement and multi-scale prompt enhancement operations.
Effectively capture multi-scale information in the image, suppress noise and background interference, improve the details and quality of whiteboard images, and significantly improve the image enhancement effect.
Smart Images

Figure CN120013781A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing, and in particular to a whiteboard image enhancement method and device based on multi-scale prompt learning. Background Art
[0002] With the development of 5G communications, remote classrooms are becoming more and more widely used. In remote classrooms, whiteboards are usually an indispensable medium for information transmission. In the above process, the camera first captures the whiteboard, and then transmits the captured real-time image to the class through the network. However, in real scenes, the whiteboard image captured by the camera is affected by a variety of complex factors, resulting in image quality degradation. These complex factors often include: noise, brightness, shadows, blur, and highlights.
[0003] In the prior art, natural image enhancement technology based on deep learning has made breakthrough progress, but most of the above image enhancement methods are only applicable to image degradation problems in specific single scenarios. At present, there is still a lack of solutions to the problem of whiteboard image degradation in complex scenarios. The difficulty of whiteboard image degradation lies in the different types and degrees of degradation. The existing deep learning network cannot effectively learn the characteristics of different types and degrees of degradation, resulting in poor whiteboard image enhancement effect.
[0004] Therefore, how to improve the image enhancement effect in scenarios with various types of whiteboard image degradation and varying degrees of degradation has become a technical problem that needs to be solved at present. Summary of the invention
[0005] The present application provides a whiteboard image enhancement method and device based on multi-scale prompt learning to solve the technical problem of how to improve the image enhancement effect in scenarios with various types of whiteboard image degradation and different degrees of degradation.
[0006] In order to solve the above technical problems, the embodiment of the present application provides a whiteboard image enhancement method based on multi-scale prompt learning, including:
[0007] Extracting first features of different scales from the first image to be enhanced by performing a first preset number of downsampling feature extraction operations;
[0008] Through the convolution attention enhancement operation, the second feature output by the last downsampling operation is enhanced to obtain the third feature;
[0009] Combining the first features at each scale, enhancing the third feature by performing the third feature prompt enhancement operation for the first preset number of times, to obtain a fourth feature;
[0010] According to the fourth feature, the first image is reconstructed to obtain a second image with enhanced image quality.
[0011] Compared with the prior art, the embodiments of the present application have the following beneficial effects: through a first preset number of downsampling feature extraction operations, first features with different scales are extracted from the first image to be enhanced, thereby effectively capturing multi-scale information in the image, and then enhancing the expression ability of the image at all levels; further through the convolution attention enhancement operation, in the process of image feature extraction, the second feature is enhanced so that the noise and irrelevant information in the image are suppressed, and the details and quality of the whiteboard image are improved; further combined with the first features of each scale, the third feature is targetedly prompted and enhanced, thereby effectively learning various degradation types and degradation degrees of the whiteboard image at each scale, improving the effect of enhancing the handwriting content in the whiteboard image, and effectively solving the degradation problem of the whiteboard image.
[0012] In some embodiments of the first aspect of the present application, extracting first features of different scales from the first image to be enhanced by performing a first preset number of downsampling feature extraction operations includes:
[0013] Inputting the first image into a first convolutional layer to extract a fifth feature;
[0014] After the fifth feature is input into the first convolution module, a downsampling operation is performed to complete the downsampling feature extraction operation once;
[0015] The downsampling feature extraction operation is repeatedly performed a first preset number of times, and each time the downsampling feature extraction operation is performed, the feature output by the first convolution module is used as the first feature.
[0016] Compared with the prior art, the above embodiment has the following beneficial effects: by extracting features through multiple downsampling, feature information of different scales is captured from the low-quality first image, which helps to enable the subsequently generated prompt features to effectively handle different degradation types at different scales, thereby improving the image enhancement effect.
[0017] In some embodiments of the first aspect of the present application, the step of enhancing the second feature output by the last downsampling operation through a convolutional attention enhancement operation to obtain a third feature includes:
[0018] The output feature of the downsampling operation during the last downsampling feature extraction operation is used as the second feature;
[0019] After inputting the second feature into the second convolutional layer, the sixth feature is extracted through the first activation layer;
[0020] The sixth feature is sequentially and repeatedly passed through the third convolutional layer and the first spatial attention layer for a second preset number of times to extract the seventh feature;
[0021] After the seventh feature is input into the fourth convolutional layer, the third feature is extracted through the second activation layer.
[0022] Compared with the prior art, the above embodiment has the following beneficial effects: the second feature obtained through multiple downsampling feature extraction contains a large number of high-level semantic features in the first image, and the above-mentioned high-level semantic features are further processed to capture the spatial semantic relationship in the second feature, thereby obtaining the sixth feature, and by introducing the spatial attention mechanism, the network enhances specific areas, such as text or important graphic areas, according to the spatial semantic information in the sixth feature, and obtains the seventh feature, thereby achieving the purpose of suppressing degradation and enhancing handwriting.
[0023] In some embodiments of the first aspect of the present application, reconstructing the first image based on the fourth feature to obtain a second image with enhanced image quality includes: splicing the fourth feature and the fifth feature and inputting them into a fifth convolutional layer to obtain the second image.
[0024] Compared with the prior art, the above embodiment has the following beneficial effects: the fourth feature is an image feature obtained through a series of feature extraction, convolution and enhancement operations, which contains rich spatial and detail information, and the fifth feature is the original feature extracted from the first image, which contains the basic content of the original image. By splicing the fourth feature and the fifth feature, the detail information extracted from the image and the original basic information are combined, so that the enhanced image maintains the enhanced details without losing the basic information of the original image.
[0025] In some embodiments of the first aspect of the present application, the combining the first features at each scale, enhancing the third feature through the first preset number of third feature prompt enhancement operations to obtain the fourth feature, includes:
[0026] Extracting a first hint feature of a corresponding scale from the third feature currently input, and inputting the first hint feature and the first feature of the corresponding scale into the second convolution module at the same time after upsampling, so as to complete a single hint enhancement operation of the third feature;
[0027] The third feature prompt enhancement operation is repeatedly performed a first preset number of times, and the feature output by the second convolution module during the last third feature prompt enhancement operation is used as the fourth feature.
[0028] Compared with the prior art, the above embodiment has the following beneficial effects: by extracting the first prompt feature of the corresponding scale from the third feature, it is possible to suppress degradation interference from features of different scales during subsequent feature extraction, and restore the extracted features to a higher resolution through an upsampling operation to enhance image details, and by repeatedly performing the above steps, it is possible to gradually remove the degraded parts in the whiteboard image, such as noise, blur, highlight areas, etc., through multi-scale information, and ultimately achieve the purpose of optimizing the whiteboard image enhancement effect.
[0029] In some embodiments of the first aspect of the present application, extracting the first prompt feature of the corresponding scale from the third feature currently input includes:
[0030] After the third feature passes through the first downsampling layer, the sixth convolutional layer and the third activation layer in sequence, the eighth feature is obtained;
[0031] Multiplying the eighth feature by a first prompt parameter of a scale corresponding to the third feature to obtain a ninth feature;
[0032] The ninth feature is sequentially passed through the seventh convolution layer, the first upsampling layer, and the eighth convolution layer to obtain the tenth feature;
[0033] After concatenating the tenth feature with the third feature, the concatenated feature is input into a ninth convolutional layer to obtain the first prompt feature.
[0034] Compared with the prior art, the above embodiment has the following beneficial effects: first, through a series of feature extraction operations, a refined eighth feature is extracted from the third feature, and the eighth feature is multiplied by the first prompt parameter, so as to strengthen the area and feature information that need to be strengthened at this scale of the eighth feature and suppress degradation interference, obtain the ninth feature, improve the image enhancement effect, further restore the resolution of the ninth feature, and splice the tenth feature strengthened by the first prompt parameter with the original third feature to obtain a spliced feature.
[0035] In a second aspect, an embodiment of the present application further provides a whiteboard image enhancement device based on multi-scale prompt learning, comprising: a first feature extraction module, a third feature extraction module, a fourth feature extraction module and an image quality enhancement module;
[0036] The first feature extraction module is used to extract first features of different scales from the first image to be enhanced by performing a first preset number of downsampling feature extraction operations;
[0037] The third feature extraction module is used to enhance the second feature output by the last downsampling operation through a convolution attention enhancement operation to obtain a third feature;
[0038] The fourth feature extraction module combines the first features at each scale and enhances the third feature through the first preset number of third feature prompt enhancement operations to obtain a fourth feature;
[0039] The image quality enhancement module is used to reconstruct the first image according to the fourth feature to obtain a second image with enhanced image quality.
[0040] In some embodiments of the second aspect of the present application, extracting first features of different scales from the first image to be enhanced by performing a first preset number of downsampling feature extraction operations includes:
[0041] Inputting the first image into a first convolutional layer to extract a fifth feature;
[0042] After the fifth feature is input into the first convolution module, a downsampling operation is performed to complete the downsampling feature extraction operation once;
[0043] The downsampling feature extraction operation is repeatedly performed a first preset number of times, and each time the downsampling feature extraction operation is performed, the feature output by the first convolution module is used as the first feature.
[0044] In some embodiments of the second aspect of the present application, the combining the first features at each scale, enhancing the third feature through the first preset number of third feature prompt enhancement operations to obtain the fourth feature, includes:
[0045] Extracting a first hint feature of a corresponding scale from the third feature currently input, and inputting the first hint feature and the first feature of the corresponding scale into the second convolution module at the same time after upsampling, so as to complete a single hint enhancement operation of the third feature;
[0046] The third feature prompt enhancement operation is repeatedly performed a first preset number of times, and the feature output by the second convolution module during the last third feature prompt enhancement operation is used as the fourth feature.
[0047] In some embodiments of the second aspect of the present application, extracting the first prompt feature of the corresponding scale from the third feature currently input includes:
[0048] After the third feature passes through the first downsampling layer, the sixth convolutional layer and the third activation layer in sequence, the eighth feature is obtained;
[0049] Multiplying the eighth feature by a first prompt parameter of a scale corresponding to the third feature to obtain a ninth feature;
[0050] The ninth feature is sequentially passed through the seventh convolution layer, the first upsampling layer, and the eighth convolution layer to obtain the tenth feature;
[0051] After concatenating the tenth feature with the third feature, the concatenated feature is input into a ninth convolutional layer to obtain the first prompt feature. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 A schematic diagram of a process of a whiteboard image enhancement method based on multi-scale prompt learning provided in some embodiments of the present application;
[0053] Figure 2 A schematic diagram of the structure of a whiteboard image enhancement network based on multi-scale prompt learning provided in some embodiments of the present application;
[0054] Figure 3 A schematic diagram of the structure of a convolution module provided in some embodiments of the present application;
[0055] Figure 4 A schematic diagram of a network structure of a convolutional attention module provided in some embodiments of the present application;
[0056] Figure 5 A schematic diagram of a spatial attention layer network structure provided for some embodiments of the present application;
[0057] Figure 6 A schematic diagram of a network structure of a prompt generation module provided in some embodiments of the present application;
[0058] Figure 7 This is a schematic structural diagram of a whiteboard image enhancement device based on multi-scale prompt learning provided in some embodiments of the present application. DETAILED DESCRIPTION
[0059] In the prior art, natural image enhancement technology based on deep learning has made breakthrough progress, but most of the above image enhancement methods are only applicable to image degradation problems in specific single scenarios. At present, there is still a lack of solutions to the problem of whiteboard image degradation in complex scenarios. The difficulty of whiteboard image degradation lies in the different types and degrees of degradation. The existing deep learning network cannot effectively learn the characteristics of different types and degrees of degradation, resulting in poor whiteboard image enhancement effect.
[0060] In order to solve the above technical problems, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0061] Embodiment 1
[0062] Please refer to Figure 1 , a whiteboard image enhancement method based on multi-scale prompt learning provided by an embodiment of the present application, including S10 to S40; reference Figure 2 , is a whiteboard image enhancement network based on multi-scale prompt learning provided in some embodiments of the present application. Next, combined with Figure 2 The network structure shown above introduces Figure 1 S10 to S40 shown, wherein S10 to S40 are specifically:
[0063] S10: extracting first features of different scales from the first image to be enhanced through a first preset number of downsampling feature extraction operations.
[0064] Further, in some embodiments of the present application, extracting first features of different scales from the first image to be enhanced by performing a first preset number of downsampling feature extraction operations includes:
[0065] Inputting the first image into a first convolutional layer to extract a fifth feature;
[0066] After the fifth feature is input into the first convolution module, a downsampling operation is performed to complete the downsampling feature extraction operation once;
[0067] The downsampling feature extraction operation is repeatedly performed a first preset number of times, and each time the downsampling feature extraction operation is performed, the feature output by the first convolution module is used as the first feature.
[0068] Preferably, reference Figure 2 In some embodiments of the present application, after the first image is input to the first convolution layer to extract the fifth feature, it passes through the first convolution module and the downsampling operation three times in sequence, and each time it passes through the first convolution module, the corresponding first feature is output to the second convolution module at the same scale.
[0069] The present application captures feature information of different scales from a low-quality first image through multiple downsampling feature extraction, which helps to enable the subsequently generated prompt features to effectively handle different degradation types at different scales, thereby improving the image enhancement effect.
[0070] Preferably, reference Figure 3 , which is a schematic diagram of the structure of a convolution module provided in some embodiments of the present application. The first convolution module described in the present application can be used Figure 3The network structure shown in the figure includes two convolutional layers and two ReLU activation function layers, in which the fifth feature is first input to the tenth convolutional layer, then through the ReLU activation function, further input to the eleventh convolutional layer again, and finally through the ReLU activation function again to output the first feature. Through multiple convolution operations and activation functions, the feature information of the first image at different scales is extracted to prepare for the subsequent whiteboard image enhancement at the same scale.
[0071] S20: Enhance the second feature output by the last downsampling operation through a convolutional attention enhancement operation to obtain a third feature.
[0072] Further, in some embodiments of the present application, the step of enhancing the second feature output by the last downsampling operation through the convolution attention enhancement operation to obtain the third feature includes:
[0073] The output feature of the downsampling operation during the last downsampling feature extraction operation is used as the second feature;
[0074] After inputting the second feature into the second convolutional layer, the sixth feature is extracted through the first activation layer;
[0075] The sixth feature is sequentially and repeatedly passed through the third convolutional layer and the first spatial attention layer for a second preset number of times to extract the seventh feature;
[0076] After the seventh feature is input into the fourth convolutional layer, the third feature is extracted through the second activation layer.
[0077] Preferably, in some embodiments of the present application, reference Figure 2 ,After the second feature is obtained through the last downsampling operation, the second feature is input to the convolutional attention module, the convolutional attention enhancement operation is performed, and the third feature is output through the convolutional attention module.
[0078] Preferably, reference Figure 4 , which is a schematic diagram of the network structure of a convolutional attention module provided in some embodiments of the present application, wherein the second feature passes through the second convolution layer, and then passes through the first activation layer with the activation function being the ReLU activation function to obtain the seventh feature; the seventh feature is further repeatedly input into the third convolution layer and the first spatial attention layer twice to obtain the third feature; finally, the third feature is input into the fourth convolution layer and the second activation layer with the activation function being the ReLU activation function to obtain the third feature.
[0079] The second feature obtained after multiple downsampling feature extraction contains a large number of high-level semantic features in the first image. The above high-level semantic features are further processed to capture the spatial semantic relationship in the second feature, thereby obtaining the sixth feature. By introducing the spatial attention mechanism, the network enhances specific areas, such as text or important graphic areas, according to the spatial semantic information in the sixth feature, and obtains the seventh feature, thereby achieving the purpose of suppressing degradation and enhancing handwriting.
[0080] Preferably, in some embodiments of the present application, the first spatial attention layer in the present application can adopt Figure 5 The network structure of the spatial attention layer shown in the figure includes: four convolutional layers, a maximum pooling layer, an addition layer, a multiplication layer, a second upsampling layer and a Sigmoid layer. When the seventh feature passes through the first spatial attention layer, it firstly extracts features through the twelfth convolutional layer, and then successively passes through the thirteenth convolutional layer, the maximum pooling layer, the fourteenth convolutional layer, and the second upsampling layer to obtain deeper spatial features. The features extracted from the twelfth convolutional layer are added to the deeper features output by the second upsampling layer for residual learning, and then the spatial features are output after passing through the fifteenth convolutional layer and the Sigmoid layer, and the spatial features are multiplied with the seventh features to enhance the important areas.
[0081] S30: In combination with the first features at each scale, the third feature is enhanced by performing the third feature prompt enhancement operation for the first preset number of times to obtain a fourth feature.
[0082] Further, in some embodiments of the present application, the combining the first features at each scale, enhancing the third feature through the first preset number of third feature prompt enhancement operations, and obtaining the fourth feature includes:
[0083] Extracting a first hint feature of a corresponding scale from the third feature currently input, and inputting the first hint feature and the first feature of the corresponding scale into the second convolution module at the same time after upsampling, so as to complete a single hint enhancement operation of the third feature;
[0084] The third feature prompt enhancement operation is repeatedly performed a first preset number of times, and the feature output by the second convolution module during the last third feature prompt enhancement operation is used as the fourth feature.
[0085] By extracting the first prompt feature of the corresponding scale from the third feature, it is possible to select appropriate prompt information from features of different scales during subsequent feature extraction, and restore the extracted features to a higher resolution through upsampling operations, enhance image details, and remove degradation. By repeatedly performing the above steps, it is possible to gradually remove the degraded parts in the whiteboard image, such as noise, blur, highlight areas, etc., through multi-scale information, and ultimately achieve the purpose of optimizing the whiteboard image enhancement effect.
[0086] Preferably, reference Figure 2 In some embodiments of the present application, after the convolutional attention module outputs the fourth feature, the fourth feature is further input into the third feature hint enhancement operation consisting of the hint generation module, the upsampling operation, and the second convolution module, and the third feature hint enhancement operation is repeated three times, wherein the hint feature generation module is used to extract the first hint feature of the corresponding scale from the third feature. At the same time, each time the second convolution module is passed, the first feature input by the first convolution module at the same scale is received to obtain the fourth feature output by the last second convolution module.
[0087] Furthermore, in some embodiments of the present application, extracting a first prompt feature of a corresponding scale from the third feature currently input includes:
[0088] After the third feature passes through the first downsampling layer, the sixth convolutional layer and the third activation layer in sequence, the eighth feature is obtained;
[0089] Multiplying the eighth feature by a first prompt parameter of a scale corresponding to the third feature to obtain a ninth feature;
[0090] The ninth feature is sequentially passed through the seventh convolution layer, the first upsampling layer, and the eighth convolution layer to obtain the tenth feature;
[0091] After concatenating the tenth feature with the third feature, the concatenated feature is input into a ninth convolutional layer to obtain the first prompt feature.
[0092] Preferably, in some embodiments of the present application, Figure 2 The prompt generation module shown can be used Figure 6The network structure of the prompt generation module shown in the figure includes: four convolutional layers, a downsampling layer, an upsampling layer, a Softmax layer, a multiplication operation, and a feature concatenation operation, and also includes a first prompt parameter, which is a learnable parameter learned during the network training process and is updated following the network training. The size of the prompt parameter is consistent with the size of the input third feature. The prompt generation modules of different scales have different first prompt parameters, thereby realizing multi-scale dynamic adaptive generation of the first prompt feature. Therefore, the prompt features generated at multiple scales can effectively learn and characterize the degradation types and degradation degrees of various inputs at different scales.
[0093] Among them, when the third feature passes through the prompt generation module, it first passes through the first downsampling layer, the sixth convolution layer, and the third activation layer with the Softmax function as the activation function, and then outputs the eighth feature, multiplies the eighth feature with the first prompt parameter at this scale, and obtains the ninth feature. The ninth feature is further passed through the seventh convolution layer, the first upsampling layer, and the eighth convolution layer in sequence to restore the resolution of the ninth feature, thereby obtaining the tenth feature, and finally the tenth feature is spliced with the third feature to obtain the first prompt feature. The prompt generation module first extracts the refined eighth feature from the third feature through a series of feature extraction operations, multiplies the eighth feature with the first prompt parameter that has been pre-trained, thereby strengthening the area and feature information that need to be emphasized in the eighth feature at this scale and suppressing degradation interference, obtaining the ninth feature, improving the effect of image enhancement, further restoring the resolution of the ninth feature, and splicing the tenth feature strengthened by the first prompt parameter with the original third feature to obtain the spliced feature.
[0094] S40: Reconstruct the first image according to the fourth feature to obtain a second image with enhanced image quality.
[0095] Furthermore, in some embodiments of the present application, reconstructing the first image according to the fourth feature to obtain a second image with enhanced image quality includes: splicing the fourth feature and the fifth feature and inputting them into a fifth convolutional layer to obtain the second image.
[0096] Preferably, reference Figure 2In some embodiments of the present application, the fourth feature and the fifth feature are spliced in the channel dimension and input into the fifth convolution layer for reconstruction to obtain a high-definition whiteboard image. Among them, since the fourth feature is an image feature obtained through a series of feature extraction, convolution and enhancement operations, it contains rich spatial and detail information, and the fifth feature is the original feature extracted from the first image, which contains the basic content of the original image and can effectively reflect the basic information of the image. By splicing the fourth feature and the fifth feature, the detail information extracted from the image and the original basic information are combined, so that the enhanced image not only maintains the enhanced details, but also does not lose the basic information of the original image.
[0097] In summary, the whiteboard image enhancement method based on multi-scale prompt learning provided by the embodiment of the present application has the following beneficial effects: through a first preset number of downsampling feature extraction operations, first features with different scales are extracted from the first image to be enhanced, thereby effectively capturing multi-scale information in the image, and then enhancing the expression ability of the image at each scale; further through the convolution attention enhancement operation, in the process of image feature extraction, the second feature is enhanced so that the noise and background interference in the image are suppressed, and the details and quality of the whiteboard image are improved; further combined with the first features of each scale, the third feature is targetedly prompted and enhanced, so as to effectively learn various degradation types and degradation degrees of the whiteboard image at each scale, improve the effect of enhancing the handwriting content in the whiteboard image, and effectively solve the degradation problem of the whiteboard image.
[0098] Embodiment 2
[0099] refer to Figure 7 , a whiteboard image enhancement device based on multi-scale prompt learning provided in an embodiment of the present application, includes: a first feature extraction module 11, a third feature extraction module 12, a fourth feature extraction module 13 and an image quality enhancement module 14.
[0100] Furthermore, in some embodiments of the present application, the first feature extraction module 11 is used to extract first features of different scales from the first image to be enhanced through a first preset number of downsampling feature extraction operations; the third feature extraction module 12 is used to enhance the second feature output by the last downsampling operation through a convolution attention enhancement operation to obtain a third feature; the fourth feature extraction module 13 combines the first features of each scale and enhances the third feature through the first preset number of third feature prompt enhancement operations to obtain a fourth feature; the image quality enhancement module 14 is used to reconstruct the first image based on the fourth feature to obtain a second image with enhanced image quality.
[0101] Furthermore, in some embodiments of the present application, the first features of different scales are extracted from the first image to be enhanced through a first preset number of downsampling feature extraction operations, including: inputting the first image into a first convolutional layer to extract a fifth feature; after inputting the fifth feature into a first convolutional module, performing a downsampling operation to complete the downsampling feature extraction operation once; repeatedly performing the downsampling feature extraction operation a first preset number of times, and using the feature output by the first convolutional module as the first feature each time the downsampling feature extraction operation is performed.
[0102] Furthermore, in some embodiments of the present application, the convolutional attention enhancement operation is used to enhance the second feature output by the last downsampling operation to obtain the third feature, including: taking the output feature of the downsampling operation as the second feature during the last downsampling feature extraction operation; after inputting the second feature into the second convolutional layer, extracting the sixth feature through the first activation layer; sequentially and repeatedly passing the sixth feature through the third convolutional layer and the first spatial attention layer for a second preset number of times to extract the seventh feature; after inputting the seventh feature into the fourth convolutional layer, extracting the third feature through the second activation layer.
[0103] Furthermore, in some embodiments of the present application, reconstructing the first image according to the fourth feature to obtain a second image with enhanced image quality includes: splicing the fourth feature and the fifth feature and inputting them into a fifth convolutional layer to obtain the second image.
[0104] Furthermore, in some embodiments of the present application, the first features at each scale are combined, and the third features are enhanced by the first preset number of third feature prompt enhancement operations to obtain a fourth feature, including: extracting the first prompt feature of the corresponding scale from the third feature currently input, and inputting the first prompt feature into the second convolution module simultaneously with the first feature of the corresponding scale after an upsampling operation to complete a single third feature prompt enhancement operation; repeatedly performing the third feature prompt enhancement operation for the first preset number of times, and using the feature output by the second convolution module during the last third feature prompt enhancement operation as the fourth feature.
[0105] Furthermore, in some embodiments of the present application, extracting a first hint feature of a corresponding scale from a third feature of the current input includes: passing the third feature through a first downsampling layer, a sixth convolutional layer, and a third activation layer in sequence to obtain an eighth feature; multiplying the eighth feature by a first hint parameter of a scale corresponding to the third feature to obtain a ninth feature; passing the ninth feature through a seventh convolutional layer, a first upsampling layer, and an eighth convolutional layer in sequence to obtain a tenth feature; concatenating the tenth feature with the third feature and inputting the resultant feature into a ninth convolutional layer to obtain the first feature.
[0106] It can be understood that the above-mentioned device item embodiment corresponds to the method item embodiment of the present invention. The whiteboard image enhancement device based on multi-scale prompt learning provided by the embodiment of the present invention can implement any method item embodiment of the present invention, that is, the whiteboard image enhancement method based on multi-scale prompt learning provided by Example 1.
[0107] In summary, the embodiment of the present application provides a whiteboard image enhancement device based on multi-scale prompt learning, which has the following beneficial effects: through a first preset number of downsampling feature extraction operations, first features with different scales are extracted from the first image to be enhanced, thereby effectively capturing multi-scale information in the image, and then enhancing the expression ability of the image at each scale; further through the convolution attention enhancement operation, in the process of image feature extraction, the second feature is enhanced so that the noise and degradation interference information in the image are suppressed, and the details and accuracy of the whiteboard image quality are improved; further combined with the first feature of each scale, the third feature is targetedly prompted and enhanced, so as to effectively learn various degradation types and degradation degrees of the whiteboard image at each scale, improve the effect of enhancing the handwriting content in the whiteboard image, and effectively solve the degradation problem of the whiteboard image.
[0108] Embodiment 3
[0109] Based on the above-mentioned embodiment of the whiteboard image enhancement method based on multi-scale prompt learning, another embodiment of the present application provides a whiteboard image enhancement terminal device based on multi-scale prompt learning. The whiteboard image enhancement terminal device based on multi-scale prompt learning includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the whiteboard image enhancement method based on multi-scale prompt learning of any embodiment of the present application is implemented.
[0110] Exemplarily, in this embodiment, the computer program may be divided into one or more modules, which are stored in the memory and executed by the processor to complete the present application. The one or more modules may be a series of computer program instruction segments capable of completing specific functions, which are used to describe the execution process of the computer program in the whiteboard image enhancement device based on multi-scale prompt learning.
[0111] The whiteboard image enhancement device based on multi-scale prompt learning can be a computing device such as a desktop computer, a notebook, a handheld computer, a cloud server, etc. The whiteboard image enhancement terminal device based on multi-scale prompt learning can include, but is not limited to, a processor and a memory.
[0112] The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc. The processor is the control center of the whiteboard image enhancement device based on multi-scale prompt learning, and uses various interfaces and lines to connect the various parts of the whiteboard image enhancement device based on multi-scale prompt learning. The memory may be used to store the computer program and / or module, and the processor implements various functions of the whiteboard image enhancement device based on multi-scale prompt learning by running or executing the computer program and / or module stored in the memory, and calling the data stored in the memory. The memory may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function, etc.; the data storage area may store data created according to the use of the mobile phone, etc. In addition, the memory may include high-speed random access memory and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0113] Embodiment 4
[0114] Based on the above-mentioned embodiment of the whiteboard image enhancement method based on multi-scale prompt learning, another embodiment of the present application provides a storage medium, wherein the storage medium includes a stored computer program, wherein when the computer program is running, the device where the storage medium is located is controlled to execute the whiteboard image enhancement method based on multi-scale prompt learning of any embodiment of the present application.
[0115] In this embodiment, the storage medium is a computer-readable storage medium, and the computer program includes computer program code, which may be in source code form, object code form, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0116] The specific embodiments described above further describe the purpose, technical solutions and beneficial effects of the present application in detail. It should be understood that the above description is only a specific embodiment of the present application and is not intended to limit the scope of protection of the present application. It is particularly pointed out that for those skilled in the art, any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of protection of the present application.
Claims
1. A whiteboard image enhancement method based on multi-scale cue learning, characterized in that: include: Extracting first features of different scales from the first image to be enhanced by performing a first preset number of downsampling feature extraction operations; Through the convolution attention enhancement operation, the second feature output by the last downsampling operation is enhanced to obtain the third feature; Combining the first features at each scale, enhancing the third feature by performing the third feature prompt enhancement operation for the first preset number of times, to obtain a fourth feature; According to the fourth feature, the first image is reconstructed to obtain a second image with enhanced image quality.
2. A whiteboard image enhancement method based on multi-scale prompt learning as claimed in claim 1, characterized in that: The step of extracting first features of different scales from the first image to be enhanced by performing a first preset number of downsampling feature extraction operations includes: Inputting the first image into a first convolutional layer to extract a fifth feature; After the fifth feature is input into the first convolution module, a downsampling operation is performed to complete the downsampling feature extraction operation once; The downsampling feature extraction operation is repeatedly performed a first preset number of times, and each time the downsampling feature extraction operation is performed, the feature output by the first convolution module is used as the first feature.
3. A whiteboard image enhancement method based on multi-scale prompt learning as claimed in claim 2, characterized in that: The step of enhancing the second feature output by the last downsampling operation through the convolution attention enhancement operation to obtain the third feature includes: The output feature of the downsampling operation during the last downsampling feature extraction operation is used as the second feature; After inputting the second feature into the second convolutional layer, the sixth feature is extracted through the first activation layer; The sixth feature is sequentially and repeatedly passed through the third convolutional layer and the first spatial attention layer for a second preset number of times to extract the seventh feature; After the seventh feature is input into the fourth convolutional layer, the third feature is extracted through the second activation layer.
4. A whiteboard image enhancement method based on multi-scale prompt learning as claimed in claim 2, characterized in that: Reconstructing the first image according to the fourth feature to obtain a second image with enhanced image quality includes: splicing the fourth feature and the fifth feature and inputting them into a fifth convolutional layer to obtain the second image.
5. The whiteboard image enhancement method based on multi-scale prompt learning as claimed in claim 1, characterized in that: The step of combining the first features at each scale and enhancing the third feature by performing the third feature prompt enhancement operation for the first preset number of times to obtain a fourth feature includes: Extracting a first hint feature of a corresponding scale from the third feature currently input, and inputting the first hint feature and the first feature of the corresponding scale into the second convolution module at the same time after upsampling, so as to complete a single hint enhancement operation of the third feature; The third feature prompt enhancement operation is repeatedly performed a first preset number of times, and the feature output by the second convolution module during the last third feature prompt enhancement operation is used as the fourth feature.
6. A whiteboard image enhancement method based on multi-scale prompt learning as claimed in claim 5, characterized in that: The step of extracting a first prompt feature of a corresponding scale from a third feature currently input includes: After the third feature passes through the first downsampling layer, the sixth convolutional layer and the third activation layer in sequence, the eighth feature is obtained; Multiplying the eighth feature by a first prompt parameter of a scale corresponding to the third feature to obtain a ninth feature; The ninth feature is sequentially passed through the seventh convolution layer, the first upsampling layer, and the eighth convolution layer to obtain the tenth feature; After concatenating the tenth feature with the third feature, the concatenated feature is input into a ninth convolutional layer to obtain the first prompt feature.
7. A whiteboard image enhancement device based on multi-scale prompt learning, characterized in that: include: A first feature extraction module, a third feature extraction module, a fourth feature extraction module and an image quality enhancement module; The first feature extraction module is used to extract first features of different scales from the first image to be enhanced by performing a first preset number of downsampling feature extraction operations; The third feature extraction module is used to enhance the second feature output by the last downsampling operation through a convolution attention enhancement operation to obtain a third feature; The fourth feature extraction module combines the first features at each scale and enhances the third feature through the first preset number of third feature prompt enhancement operations to obtain a fourth feature; The image quality enhancement module is used to reconstruct the first image according to the fourth feature to obtain a second image with enhanced image quality.
8. The whiteboard image enhancement device based on multi-scale prompt learning as claimed in claim 7, characterized in that: The step of extracting first features of different scales from the first image to be enhanced by performing a first preset number of downsampling feature extraction operations includes: Inputting the first image into a first convolutional layer to extract a fifth feature; After the fifth feature is input into the first convolution module, a downsampling operation is performed to complete the downsampling feature extraction operation once; The downsampling feature extraction operation is repeatedly performed a first preset number of times, and each time the downsampling feature extraction operation is performed, the feature output by the first convolution module is used as the first feature.
9. The whiteboard image enhancement device based on multi-scale prompt learning as claimed in claim 7, characterized in that: The step of combining the first features at each scale and enhancing the third feature by performing the third feature prompt enhancement operation for the first preset number of times to obtain a fourth feature includes: Extracting a first hint feature of a corresponding scale from the third feature currently input, and inputting the first hint feature and the first feature of the corresponding scale into the second convolution module at the same time after upsampling, so as to complete a single hint enhancement operation of the third feature; The third feature prompt enhancement operation is repeatedly performed a first preset number of times, and the feature output by the second convolution module during the last third feature prompt enhancement operation is used as the fourth feature.
10. The whiteboard image enhancement device based on multi-scale prompt learning as claimed in claim 9, characterized in that: The step of extracting a first prompt feature of a corresponding scale from a third feature currently input includes: After the third feature passes through the first downsampling layer, the sixth convolutional layer and the third activation layer in sequence, the eighth feature is obtained; Multiplying the eighth feature by a first prompt parameter of a scale corresponding to the third feature to obtain a ninth feature; The ninth feature is sequentially passed through the seventh convolution layer, the first upsampling layer, and the eighth convolution layer to obtain the tenth feature; After concatenating the tenth feature with the third feature, the concatenated feature is input into a ninth convolutional layer to obtain the first prompt feature.