Image processing method and system and medium

By preprocessing the image to be repaired, and using gated convolution processing and multi-scale ReLU linear attention processing, the problem that image repair methods in the prior art ignore local details and global context information is difficult to balance, achieving high-quality image repair effects.

CN119991516AInactive Publication Date: 2025-05-13INSPUR GENERSOFT CO LTD

Patent Information

Application Number
CN202510465829.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing image repair methods based on deep learning are prone to ignore local details of the image, resulting in the loss of key feature information, making it difficult to accurately restore complex local texture structures in the image, and it is difficult to find a balance between global context information and local feature expression, resulting in blurred image and artifacts after repair.

Method used

By preprocessing the image to be repaired, its size is consistent with the training set, and using gated convolution processing and multi-scale ReLU linear attention processing, dynamically control information flow and extract multi-scale features to preserve important features and suppress invalid information.

Benefits of technology

It improves the detail clarity and authenticity of the repaired image, reduces the occurrence of artifacts, enhances the accurate recovery ability of complex local texture structures, and improves the fusion between the repaired image and the original image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991516A_ABST
    Figure CN119991516A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, and discloses an image processing method and system and a medium, and the method comprises the steps: carrying out the preprocessing of a to-be-restored image, so as to enable the size of the to-be-restored image to be consistent with the size of a training set corresponding to an image restoration model; and repairing the preprocessed to-be-repaired image through the image repairing model, wherein the repairing comprises gating convolution processing and multi-scale ReLU linear attention processing. According to the method, the overall quality of the restored image is remarkably improved, the visual effect of the restored image is closer to that of an original image, and details are more exquisite and vivid.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to an image processing method, system and medium. Background Art

[0002] With the development of computer technology, digital images have gradually become the main medium for recording and preserving information. However, during the transmission and storage process, digital images will inevitably suffer from quality degradation problems such as pixel loss. Therefore, image restoration is an important research topic in the field of computer vision and image processing, aiming to restore the original visual quality and information of damaged, missing or degraded images. Currently, image restoration is mainly used in the following fields: (1) In the field of cultural relics protection and restoration, many historical relics and art paintings have been damaged due to the passage of time, environmental impacts or human factors. Through image restoration technology, experts can virtually restore cultural relics and foresee the effects of restoration to determine the best restoration plan. (2) In the social media and entertainment industries, users have increasing demands for image quality. People want to share clear and beautiful photos, but there may be unwanted obstructions or distracting objects in the shooting scene. Image restoration technology can help users easily remove unwanted areas in photos, making them more visually attractive and meeting the needs of social sharing.

[0003] (3) Old photos carry precious memories and historical records, but due to the passage of time, environmental influences or improper storage conditions, these photos often become damaged by fading, creases, stains, etc. Image restoration technology uses advanced algorithms to automatically identify these defects, fill in the missing image information, and inject new life into precious memories.

[0004] In the process of implementing the present invention, the inventors found that the prior art has the following technical problems: Existing deep learning-based image restoration methods tend to ignore local details of the image, resulting in the loss of key feature information and difficulty in accurately restoring the complex local texture structure in the image, making the restored image blurry in details and prone to artifacts. In addition, in the process of image restoration, the deep learning-based graphics restoration model has difficulty finding a balance between global context information and local feature expression, resulting in a low degree of integration between the restored image and the original image, and a poor restoration effect. Summary of the invention

[0005] In view of the above problems, the present disclosure provides an image processing method, system and medium that overcome the above problems or at least partially solve the above problems. The technical solution is as follows: In order to achieve the above object, in a first aspect, the present invention provides an image processing method, comprising: Preprocessing the image to be repaired so that the size of the image to be repaired is consistent with the size of the training set corresponding to the image repair model; The preprocessed image to be repaired is repaired by the image repair model, and the repair includes gated convolution processing and multi-scale ReLU linear attention processing.

[0006] Preferably, the image to be repaired is preprocessed, specifically: Performing data augmentation processing on the image to be restored that is scaled to a specified size; Generate a degraded image based on the image to be restored after data augmentation processing and the mask data; The data dimension is expanded based on the degraded image and the mask data.

[0007] Preferably, the gated convolution processing specifically includes: Performing parallel convolution operations on the features in the image to be repaired; Selecting features to be enhanced and features to be suppressed from the features based on a preset activation function and a preset Hadamard product; The features to be enhanced and the features to be suppressed are respectively subjected to convolution processing.

[0008] Preferably, the multi-scale ReLU linear attention processing specifically includes: Performing layer normalization processing on the features in the image to be repaired; The features after layer normalization processing are equally divided into three processing branches according to the channel dimension of the features; The first two processing branches of the three processing branches are subjected to feature extraction according to different convolution kernels, and the last processing branch of the three processing branches is subjected to identity mapping processing; The features corresponding to the branches after feature extraction and mapping processing are superimposed respectively.

[0009] Preferably, the features in the image to be repaired are subjected to layer normalization processing, specifically: Perform statistics on the data in each channel of the input feature to determine its mean and variance; Performing standardization processing on the input features based on the mean and the variance so that the feature data distribution in each channel satisfies a preset mean and variance range; After the normalization process, the normalized feature data is linearly transformed through learnable scaling parameters and translation parameters to restore or adjust the distribution range of the feature representation.

[0010] Preferably, the image restoration model is a network composed of a module corresponding to the gated convolution processing and a module corresponding to the multi-scale ReLU linear attention processing; Each layer in the downsampling and upsampling process of the image restoration model consists of gated convolution modules.

[0011] Preferably, before preprocessing the image to be repaired, the method further includes: Acquire a data set for training the image restoration model; Training the data set according to a preset number of training iterations, and expanding training samples to the data set by using data enhancement technology during the training process; The types of data enhancement technology include at least: random flipping, rotation, brightness adjustment, contrast adjustment and hue adjustment.

[0012] Preferably, after the preprocessed image to be repaired is repaired by the image repair model, the method further includes: The modified image to be repaired and the data samples in the data set are subjected to a loss function to generate a loss value; The image restoration model is iteratively optimized using the loss value.

[0013] In a second aspect, the present invention provides an image processing system, comprising: A preprocessing module, used for preprocessing the image to be repaired so that the size of the image to be repaired is consistent with the size of the training set corresponding to the image repair model; A restoration module is used to restore the preprocessed image to be restored through the image restoration model, wherein the restoration includes gated convolution processing and multi-scale ReLU linear attention processing.

[0014] In a third aspect, the present invention provides a computer-readable storage medium storing a computer program, which, when executed, implements any of the methods described above.

[0015] The embodiment of the present application discloses an image processing method, system and medium. The method ensures that the data input to the image restoration model has a unified standard by adjusting the size of the image to be restored to be consistent with the training set, and ensures that the model can more accurately apply the knowledge it has learned when processing the image to be restored, avoiding information loss or deformation caused by inconsistent sizes, thereby improving the quality of the restored image; using gated convolution processing to extract features and control information flow of the image to be restored, it can dynamically and selectively retain important features and suppress invalid or redundant information, effectively solving the problem that the existing technology easily ignores local details, making the restored image clearer and more realistic in details, reducing the appearance of artifacts, and greatly improving the ability to accurately restore complex local texture structures; using multi-scale ReLU linear attention processing, not only reduces the computational complexity, but also enhances the ability to capture long-distance dependencies by efficiently obtaining the global receptive field. At the same time, by combining multi-scale feature learning, it can effectively capture global dependencies while paying attention to fine-grained local information, thereby achieving a good balance between global context information and local feature expression, improving the fusion degree of the restored image and the original image, and improving the restoration effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings: Figure 1 A flowchart of an image processing method provided by an embodiment of the present invention; Figure 2 A schematic diagram of the structure of an image processing system provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0017] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art. It should be noted that, in the absence of conflict, the embodiments of the present invention and the features described in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0018] like Figure 1 As shown, in some embodiments of the present application, this embodiment provides an image processing method, specifically, the method includes the following steps: Step S101: preprocess the image to be repaired so that the size of the image to be repaired is consistent with the size of the training set corresponding to the image repair model.

[0019] As mentioned above, the image to be repaired is preprocessed to ensure that its size is consistent with the image size in the dataset used to train the image repair model. This process is crucial because it ensures that the images input to the model have consistent size standards, thereby avoiding information loss or deformation due to size differences. Specifically, this step includes adjusting the size of the image to be repaired (such as scaling it to 256×256 pixels) so that it can match the data format used when training the model. In addition, it may also include operations such as rotating and mirroring the image to expand the dataset and increase the robustness and generalization ability of the model.

[0020] For example, suppose we have an old photo to be restored, whose original size is 1024×768 pixels, and has problems such as fading and stains. According to the method of the present invention, in step S101, it is first necessary to scale this old photo to the same size as the model training set, such as 256×256 pixels. Then, a series of data augmentation operations can be performed on the reduced image, such as random flipping, rotating 90 degrees or horizontal flipping, etc., which can not only expand the number of training samples, but also enhance the model's ability to recognize and repair images from different perspectives. In this way, even if the original image is taken under a variety of different conditions, it can be ensured that the model can effectively repair it.

[0021] It should be noted that in a specific implementation scenario, a color correction scheme can be adopted on the basis of the above scheme, that is, before or after resizing, the color information of the image can be corrected to compensate for the color deviation caused by a long time or other reasons, which helps to improve the realism and visual quality of the repaired image; a noise removal scheme can be adopted, that is, in order to more accurately restore the image details, a denoising algorithm can be applied to the image before resizing, which can not only improve the overall clarity of the image, but also reduce the artifact phenomenon that may occur in the subsequent repair process; an adaptive resolution adjustment scheme is adopted, that is, in addition to the fixed size adjustment, a resolution adjustment mechanism can be used to dynamically adjust the resolution of the image to be repaired according to the specific situation. For example, for particularly important but severely damaged areas, a higher resolution method can be used to process them to ensure that the key parts are repaired to the best effect; an intelligent cropping and filling scheme is adopted, that is, for images with mismatched proportions or containing a large amount of irrelevant background information, non-essential parts can be removed by intelligent cropping, and the missing areas can be filled with a specific algorithm, so that the final resized image is more focused on the core content and the repair efficiency and quality are improved. The above optional schemes all belong to the protection scope of this application.

[0022] Step S102: repairing the preprocessed image to be repaired by using the image repair model, wherein the repairing includes gated convolution processing and multi-scale ReLU linear attention processing.

[0023] As mentioned above, this step records the process of using a specific image restoration model to restore the preprocessed image to be restored, where the restoration process includes gated convolution processing and multi-scale ReLU linear attention processing. Gated convolution processing is mainly used to control the flow of information, and reduces the impact of invalid pixels by selectively enhancing or suppressing features, thereby improving the model's expressiveness in complex data. The multi-scale ReLU linear attention mechanism aims to efficiently obtain the global receptive field while combining feature extraction methods at different scales to enhance the model's ability to capture local to global dependencies. This combination not only improves computational efficiency, but also ensures powerful feature extraction capabilities and nonlinear modeling capabilities, which helps to generate higher quality restored images.

[0024] For example, suppose there is an old photo that has been resized and contains severe fading and crease problems. In step S102, the photo is first input into the trained GaLAformer image restoration model. In this process, the gated convolution process identifies and enhances areas that require special attention (such as creases) while suppressing irrelevant background noise. Next, the multi-scale ReLU linear attention process analyzes the entire image, identifies key features from local details to overall structure, and uses these features to fill in the missing parts and restore the original visual quality of the image. For example, for the subtle texture of a person's face in the photo, multi-scale feature learning ensures that even the smallest wrinkles can be accurately restored without losing important details.

[0025] It should be noted that in a specific implementation scenario, an adaptive gated adjustment scheme can be adopted on the basis of the above scheme, that is, the activation function parameters in the gated convolution are dynamically adjusted according to the specific situation of the input image, so that the model can more flexibly cope with different types or degrees of image damage; a deep context fusion scheme is adopted, that is, on the basis of the multi-scale ReLU linear attention mechanism, a deeper context information fusion strategy is introduced, such as cross-layer connection or attention weighted average, to further enhance the model's ability to capture long-distance dependencies; a mixed loss optimization scheme is adopted, that is, in addition to the basic reconstruction loss, more types of loss functions such as perceptual loss and style loss can also be considered to be added to facilitate a more comprehensive evaluation and optimization of the repair effect. This method can make the repaired image not only close to the original image at the pixel level, but also as consistent as possible in visual perception and artistic style; a real-time feedback and iterative repair scheme is adopted, that is, based on a feedback repair mechanism, allowing users to provide feedback based on the preliminary repair results, and then automatically adjusting the model parameters based on these feedbacks and performing iterative repairs, which will greatly improve the user experience and ensure that the final output meets the user's personalized needs. The above optional schemes all belong to the scope of protection of this application.

[0026] In some embodiments of the present application, in order to effectively prepare the images to be repaired so that they are more suitable for training and testing image repair models, thereby improving the quality and efficiency of repair, the images to be repaired are preprocessed, specifically: Performing data augmentation processing on the image to be restored that is scaled to a specified size; Generate a degraded image based on the image to be restored after data augmentation processing and the mask data; The data dimension is expanded based on the degraded image and the mask data.

[0027] As mentioned above, first, the image to be repaired is resized to a specific size (for example, 256×256 pixels) to match the standard size used when training the model. This step is to ensure that all images input into the model have a consistent size, thereby avoiding information loss or deformation due to inconsistent sizes. Next, data augmentation is performed on these resized images. Data augmentation is a technique that creates more training samples by applying a series of transformations, such as random rotation, flipping (horizontally or vertically), adjusting brightness, contrast, etc. Doing so not only increases the diversity of the training set, but also enhances the robustness and generalization ability of the model, enabling it to make more accurate predictions in different situations.

[0028] After completing data augmentation, the next step is to generate degraded images, which are images that simulate various damage conditions that may occur in reality (such as scratches, stains, fading, etc.). Here, mask data is combined with the image to be repaired after data augmentation to generate this degradation effect. The mask data is essentially a binary matrix that defines which areas in the image should be considered "damaged" or "missing". By applying this mask to the original image, degraded images with different types of damage or missing parts can be simulated.

[0029] The last step involves expanding the data dimension in order to accommodate the requirements of subsequent model inputs and provide more information for training. Specifically, this process involves superimposing the degraded image with the corresponding mask data to form a new multi-channel input. For example, if the original image is a single-channel grayscale image, then after adding the mask data, the data dimension that is finally input to the model may be expanded to four channels (i.e., the original image + three additional mask channels). Doing so not only increases the amount of information for each input sample, but also helps the model better understand the damage pattern and its location in the image, thereby improving the repair accuracy.

[0030] In some embodiments of the present application, in order to effectively control the information flow, the image restoration model is more accurate and efficient when facing complex image damage. The gated convolution processing specifically includes: Performing parallel convolution operations on the features in the image to be repaired; Selecting features to be enhanced and features to be suppressed from the features based on a preset activation function and a preset Hadamard product; The features to be enhanced and the features to be suppressed are respectively subjected to convolution processing.

[0031] As mentioned above, at this stage, the image to be repaired is first fed into one or more convolutional layers. These convolutional layers can automatically learn and extract various features in the image (such as edges, textures, etc.). Unlike traditional convolution operations, parallel processing is used here, which means that different parts of the image can be processed by multiple convolution kernels at the same time, thereby improving efficiency and the diversity of feature extraction. Parallel convolution operations not only speed up the calculation process, but also capture richer feature information from different angles, which is crucial for subsequent feature screening.

[0032] The feature map obtained after the convolution operation will further pass through an activation function (such as the Sigmoid function), which generates a weight between 0 and 1 for each eigenvalue. This weight determines the importance of the corresponding feature. Then, the weight matrix generated by the activation function is combined with the original feature map using the preset Hadamard product (i.e., element-by-element multiplication). This can effectively enhance or suppress each position in the feature map according to the weight. Specifically, the features corresponding to the positions with weights close to 1 are considered important and should be enhanced; while the positions with weights close to 0 are considered unimportant or even harmful and should be suppressed.

[0033] After feature screening is completed, the next step is to perform independent convolution processing on the selected features to be enhanced and suppressed. This processing method allows the model to adopt different strategies for different types of features: for features to be enhanced, their representation ability can be further enhanced through appropriate convolution operations; while for features to be suppressed, more conservative convolution settings may be applied to reduce their impact. This differentiated processing helps to better preserve and restore the key details of the image while avoiding interference from irrelevant or harmful information, ultimately achieving the purpose of improving the restoration effect.

[0034] In some embodiments of the present application, in order to improve the model's ability to capture complex image features, while enhancing its ability to understand and integrate information of different scales, the multi-scale ReLU linear attention processing specifically includes: Performing layer normalization processing on the features in the image to be repaired; The features after layer normalization processing are equally divided into three processing branches according to the channel dimension of the features; The first two processing branches of the three processing branches are subjected to feature extraction according to different convolution kernels, and the last processing branch of the three processing branches is subjected to identity mapping processing; The features corresponding to the branches after feature extraction and mapping processing are superimposed respectively.

[0035] As mentioned above, at this stage, the feature map obtained after the image to be repaired is first subjected to layer normalization after convolution and other operations. Layer normalization stabilizes the training process and accelerates model convergence by standardizing all pixel values ​​in each feature map (adjusting their mean and variance). Specifically, the mean and variance of the data in each channel of the input feature are calculated, and the input features are standardized based on these statistics so that the distribution of feature data in each channel meets the preset mean and variance range. Afterwards, the standardized feature data is linearly transformed through learnable scaling parameters and translation parameters to restore or adjust the distribution range of the feature representation.

[0036] After layer normalization, the feature map is evenly split into three independent branches along the channel dimension. This segmentation allows the model to analyze image features from different angles and scales, thereby enhancing the model's ability to capture multi-level information. Each branch can focus on extracting a specific type of feature, which helps improve the comprehensiveness and accuracy of the final restoration effect.

[0037] For the first two branches, different sizes of convolution kernels (such as 3×3 and 5×5) are used for feature extraction. This ensures that the model can capture detailed information at different scales: smaller convolution kernels are more suitable for capturing local details, while larger convolution kernels help to identify structural information in a larger range. The third branch is directly processed by identity mapping, that is, no transformation or filtering is performed on the data of this branch, and the original features are kept unchanged. This method retains the basic structural information of the input features and prevents important details from being lost during complex processing.

[0038] The last step is to reassemble the features corresponding to each branch after different processing (including feature extraction and identity mapping). Usually, this step is achieved by simple element addition or splicing. The superimposed feature map combines rich information at different scales, including both fine local details and broad global context, thus providing strong support for subsequent restoration work.

[0039] In some embodiments of the present application, in order to effectively solve the problem of inconsistent data distribution between different batches or channels, improve the efficiency and stability of model training, and maintain the necessary flexibility to meet the needs of various application scenarios, the features in the image to be repaired are subjected to layer normalization processing, specifically: Perform statistics on the data in each channel of the input feature to determine its mean and variance; Performing standardization processing on the input features based on the mean and the variance so that the feature data distribution in each channel satisfies a preset mean and variance range; After the normalization process, the normalized feature data is linearly transformed through learnable scaling parameters and translation parameters to restore or adjust the distribution range of the feature representation.

[0040] As mentioned above, in this step, we first need to calculate statistics for each channel in the feature map to be processed (that is, the data obtained after operations such as convolution). Specifically, for all pixel values ​​in each channel, calculate their average (mean) and variance. The "channel" here refers to a dimension of the feature map. For example, there are three channels (red, green, and blue) in an RGB image, and there may be more channels in more complex neural network layers. By calculating these statistics, we can understand the overall distribution of the data in the current channel, providing a basis for subsequent standardization processing.

[0041] Next, the feature data in each channel is normalized using the mean and variance calculated in the previous step. The goal of normalization is to adjust the distribution of the data in each channel so that it has the same mean and variance, usually set to a standard normal distribution with a mean of 0 and a variance of 1. This process can be achieved using the following formula: ; in is the original input data, is the mean, is the variance, is a small constant that prevents the denominator from being zero. This is done to eliminate scale differences between batches or channels, thereby improving the stability of model training.

[0042] After standardization, although the data distribution becomes more consistent, sometimes this strict standardization may affect the performance of the model because some tasks or model architectures may require inputs in a specific range or form. Therefore, in this step, two learnable parameters are introduced: the scaling parameter γ and the translation parameter β. By applying a linear transformation to the standardized data , the distribution range of data can be flexibly adjusted according to actual needs. These parameters are automatically optimized during the training process, allowing the model to adaptively adjust the form of feature representation according to specific task requirements, retaining the benefits of standardization while avoiding information loss caused by forced standardization.

[0043] In some embodiments of the present application, in order to not only accurately restore image quality in complex image damage situations, but also effectively manage and optimize information flow to ensure detail preservation and enhancement during the conversion process from low resolution to high resolution. The image restoration model is a network composed of a module corresponding to the gated convolution process and a module corresponding to the multi-scale ReLU linear attention process; Each layer in the downsampling and upsampling process of the image restoration model consists of gated convolution modules.

[0044] As mentioned above, the core of the image restoration model is composed of two main types of modules: gated convolution modules and multi-scale ReLU linear attention modules. These two modules are responsible for different functions and work together to achieve efficient image restoration. The gated convolution module is mainly used to control the flow of information. It reduces the impact of invalid pixels by selectively enhancing or suppressing specific parts of the feature map, which helps to improve the model's expressiveness in complex data and can more accurately restore the key details of the image; the multi-scale ReLU linear attention module aims to efficiently obtain the global receptive field while combining feature extraction methods of different scales to enhance the model's ability to capture local to global dependencies, which not only reduces the computational complexity, but also improves the model's ability to capture long-distance dependencies.

[0045] In image processing networks, downsampling (usually used to reduce spatial resolution but increase the number of channels) and upsampling (on the contrary, used to restore spatial resolution) are common operation steps. These operations are essential for resizing images, extracting high-level features, and ultimately generating high-resolution restoration results. In the model proposed in the present invention, a gated convolution module is used in each step of both downsampling and upsampling. The advantages include: the gating mechanism allows the model to dynamically adjust its processing method according to the input data, thereby more effectively extracting shallow spatial features during downsampling while reducing the interference of invalid pixels; during upsampling, the gated convolution module helps the model better restore detail information because it can enhance or suppress specific features as needed to ensure that key information is not lost; by consistently using gated convolution modules in the downsampling and upsampling paths of the entire network, it can ensure that the information flow is optimally managed, thereby improving the expressiveness and robustness of the overall model.

[0046] In some embodiments of the present application, in order to ensure that the finally trained model has good generalization performance, the image restoration task can be effectively performed in practical applications. Before preprocessing the image to be restored, the following steps are also included: Acquire a data set for training the image restoration model; Training the data set according to a preset number of training iterations, and expanding training samples to the data set by using data enhancement technology during the training process; The types of data enhancement technology include at least: random flipping, rotation, brightness adjustment, contrast adjustment and hue adjustment.

[0047] As mentioned above, before starting to train the image restoration model, you first need to collect or obtain a suitable training dataset. The choice of dataset is crucial because it directly affects the learning effect and generalization ability of the model. This dataset contains a large number of image samples to be restored and their corresponding complete (undamaged) versions, so that the model can learn how to restore from the damaged state to the original state; covering various types of image damage, such as scratches, stains, fading, etc., to ensure that the model can work effectively in different scenarios.

[0048] Once the training dataset is determined, the next step is to use this data to train the image restoration model according to the preset number of training iterations. The number of training iterations refers to the number of times the entire dataset is traversed by the model during the entire training process. In each iteration, the model starts with forward propagation to calculate the predicted value, and then adjusts its internal parameters (i.e. weights) based on the difference between the actual value and the predicted value. This process is usually implemented through the backpropagation algorithm and the loss function is used to quantify the error size. The preset number of training iterations can be determined based on experimental results or experience, with the goal of allowing the model to fully learn the patterns in the data without overfitting.

[0049] To further improve the robustness and generalization of the model, data augmentation techniques are applied to the existing dataset during training. This approach generates new training samples by applying a series of transformations to existing images, thereby increasing the diversity and size of the training set. Data augmentation techniques include at least the following types: random flipping, which flips the image horizontally or vertically to increase the number of training samples that are directional invariant; rotation, which randomly rotates the image by a certain angle so that the model learns to recognize features in different directions; brightness adjustment, which changes the overall brightness of the image to simulate different lighting conditions; contrast adjustment, which changes the contrast between different areas in the image to help the model adapt to different visual intensities; and hue adjustment, which fine-tunes the color distribution of the image so that the model can handle images with large color variations.

[0050] In some embodiments of the present application, in order to ensure that as the training progresses, the model can gradually learn how to better restore the quality of the damaged image. After the pre-processed image to be repaired is repaired by the image repair model, it also includes: The modified image to be repaired and the data samples in the data set are subjected to a loss function to generate a loss value; The image restoration model is iteratively optimized using the loss value.

[0051] As mentioned above, after completing an image restoration operation, the restored image processed by the model is obtained. At this point, it is necessary to compare this restoration result with the corresponding complete (undamaged) version of the image in the dataset. The loss function is used to quantify this difference or error. Commonly used loss functions include but are not limited to L1 loss, L2 loss, perceptual loss, style loss, and adversarial loss. Each loss function has its specific purpose. For example, L1 loss is usually used to measure pixel-level differences, while perceptual loss focuses on higher-level feature similarities. By inputting the restored image and the real image into the selected loss function, one or more loss values ​​can be generated. These loss values ​​reflect the gap between the current output of the model and the ideal output. The lower the value, the better the model performance.

[0052] After obtaining the loss value, the next step is to use this information to adjust the parameters within the model in order to reduce the error in future predictions. This step is called the iterative optimization process of the model. Optimization is usually achieved through the back-propagation algorithm. During this process, the gradient calculated based on the loss value will be back-propagated along the network layer and the weight parameters of each layer will be updated accordingly. Iterative optimization is a cyclical process until a predetermined stopping condition is reached, such as reaching the maximum number of training rounds, the loss value no longer decreases significantly, or a certain performance indicator is met. In each iteration, the model will try to find a better set of parameter settings so that the next run can produce a restoration result that is closer to the real image. This process helps to continuously improve the accuracy and stability of the model, enabling it to perform well in various situations.

[0053] Compared with the prior art, the embodiment of the present application discloses an image processing method, which ensures that the data input to the image restoration model has a unified standard by adjusting the size of the image to be restored to be consistent with the training set, and ensures that the model can more accurately apply its learned knowledge when processing the image to be restored, avoiding information loss or deformation caused by inconsistent sizes, thereby improving the quality of the restored image; using gated convolution processing to extract features and control information flow of the image to be restored, it can dynamically and selectively retain important features and suppress invalid or redundant information, effectively solving the problem that the prior art easily ignores local details, making the restored image clearer and more realistic in details, reducing the appearance of artifacts, and greatly improving the ability to accurately restore complex local texture structures; using multi-scale ReLU linear attention processing, not only the computational complexity is reduced, but also the ability to capture long-distance dependencies is enhanced by efficiently obtaining the global receptive field. At the same time, by combining multi-scale feature learning, it is possible to effectively capture global dependencies while paying attention to fine-grained local information, thereby achieving a good balance between global context information and local feature expression, improving the fusion degree of the restored image with the original image, and improving the restoration effect. In summary, through the above means, this method not only overcomes the problems existing in existing deep learning-based image restoration methods, such as loss of key feature information, difficulty in accurately restoring complex local texture structures, and difficulty in finding a balance between global contextual information and local feature expression, but also significantly improves the overall quality of the restored image, making it closer to the original image in visual effect and more refined and realistic in detail expression.

[0054] Based on the same inventive concept as the above method, the present application embodiment also proposes an image processing system, such as Figure 2 FIG. 1 is a schematic diagram of the structure of an image processing system, the system comprising: A preprocessing module, used for preprocessing the image to be repaired so that the size of the image to be repaired is consistent with the size of the training set corresponding to the image repair model; A restoration module is used to restore the preprocessed image to be restored through the image restoration model, wherein the restoration includes gated convolution processing and multi-scale ReLU linear attention processing.

[0055] Based on the same inventive concept as the above method, an embodiment of the present application further proposes a computer-readable storage medium on which a computer program is stored, and when the program is executed by a processor, the steps of any of the above-mentioned embodiments of the method are implemented. Among them, computer-readable storage media may include, but are not limited to, any type of disk, including floppy disks, optical disks, DVDs (Digital Video Discs), CD-ROMs (Compact Disc Read-Only Memory), microdrives and magneto-optical disks, ROMs (Read-Only Memory), RAMs (Random Access Memory), EPROMs (Erasable Programmable Read-Only Memory), EEPROMs (Electrically Erasable Programmable Read Only Memory), DRAMs (Dynamic Random Access Memory), VRAMs (Video Random Access Memory), flash memory devices, magnetic or optical cards, nanosystems (including molecular memory ICs), or any type of media or device suitable for storing instructions and / or data.

[0056] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0057] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1A device that provides the functions specified in a block or multiple blocks.

[0058] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0059] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0060] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. An image processing method, characterized in that: include: Preprocessing the image to be repaired so that the size of the image to be repaired is consistent with the size of the training set corresponding to the image repair model; The preprocessed image to be repaired is repaired by the image repair model, and the repair includes gated convolution processing and multi-scale ReLU linear attention processing.

2. The method according to claim 1, characterized in that The image to be repaired is preprocessed as follows: Performing data augmentation processing on the image to be restored that is scaled to a specified size; Generate a degraded image based on the image to be restored after data augmentation processing and the mask data; The data dimension is expanded based on the degraded image and the mask data.

3. The method according to claim 1, characterized in that The gated convolution processing specifically includes: Performing parallel convolution operations on the features in the image to be repaired; Selecting features to be enhanced and features to be suppressed from the features based on a preset activation function and a preset Hadamard product; The features to be enhanced and the features to be suppressed are respectively subjected to convolution processing.

4. The method according to claim 1, characterized in that The multi-scale ReLU linear attention processing specifically includes: Performing layer normalization processing on the features in the image to be repaired; The features after layer normalization processing are equally divided into three processing branches according to the channel dimension of the features; The first two processing branches of the three processing branches are subjected to feature extraction according to different convolution kernels, and the last processing branch of the three processing branches is subjected to identity mapping processing; The features corresponding to the branches after feature extraction and mapping processing are superimposed respectively.

5. The method according to claim 4, characterized in that The features in the image to be repaired are subjected to layer normalization processing, specifically: Perform statistics on the data in each channel of the input feature to determine its mean and variance; Performing standardization processing on the input features based on the mean and the variance so that the feature data distribution in each channel satisfies a preset mean and variance range; After the normalization process, the normalized feature data is linearly transformed through learnable scaling parameters and translation parameters to restore or adjust the distribution range of the feature representation.

6. The method according to any one of claims 1 to 5, characterized in that: The image restoration model is a network composed of a module corresponding to the gated convolution processing and a module corresponding to the multi-scale ReLU linear attention processing; Each layer in the downsampling and upsampling process of the image restoration model consists of gated convolution modules.

7. The method according to claim 6, characterized in that Before preprocessing the image to be repaired, it also includes: Acquire a data set for training the image restoration model; Training the data set according to a preset number of training iterations, and expanding training samples to the data set by using data enhancement technology during the training process; The types of data enhancement technology include at least: random flipping, rotation, brightness adjustment, contrast adjustment and hue adjustment.

8. The method according to claim 7, characterized in that After the preprocessed image to be repaired is repaired by the image repair model, the method further includes: The modified image to be repaired and the data samples in the data set are subjected to a loss function to generate a loss value; The image restoration model is iteratively optimized using the loss value.

9. An image processing system, characterized in that: include: A preprocessing module, used for preprocessing the image to be repaired so that the size of the image to be repaired is consistent with the size of the training set corresponding to the image repair model; A restoration module is used to restore the preprocessed image to be restored through the image restoration model, wherein the restoration includes gated convolution processing and multi-scale ReLU linear attention processing.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Smoke detection method and system based on lightweight neural network, medium and equipment

    CN116434002A

  • Image restoration system based on gated convolution and gated attention mechanism

    CN116523779A

  • Text feature-introduced image restoration method

    CN117314778A

  • Shielding handwritten medical record Chinese character image restoration method based on gating convolution and SCPAM attention module

    CN117455813A

  • Image restoration method and device based on generative adversarial network

    CN117495724A

Cited By

  • Lightweight image restoration method and device based on multi-scale features, and medium

    CN120374463A

  • Lightweight image restoration method, device and medium based on multi-scale features

    CN120374463B