Image processing model, server, image processing method, and storage medium

Through the features extraction, quality factor prediction and dynamic modulation in the image processing model, the problem of degradation in the quality of compressed images is solved, and efficient and accurate image recovery and quality improvement are achieved.

CN120034655APending Publication Date: 2025-05-23GUANGZHOU XIAOPENG MOTORS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510193961.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Due to bandwidth limitations and data storage pressure, images captured by the camera usually require image compression, resulting in image information loss and compression artifacts, affecting image quality and accuracy of subsequent applications.

Method used

An image processing model is provided, including a feature extractor, a quality factor predictor, a flexible controller and an image recoverer. The model efficiently and accurately recovers compressed images through a closed-loop process of feature extraction, quality factor prediction, dynamic modulation and adaptive recovery.

Benefits of technology

It realizes efficient recovery of compressed images, removes artifacts, improves image clarity and detailed information, enhances the generalization ability and scene adaptability of the image processing model, and avoids the tedious process of training individual models for different compression quality factors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034655A_ABST
    Figure CN120034655A_ABST
Patent Text Reader

Abstract

The invention discloses an image processing model, a server, an image processing method and a computer readable storage medium. The image processing model comprises a feature extractor, a quality factor predictor, a flexible controller and an image restorer. The feature extractor is configured to determine image feature information of the compressed image according to the acquired compressed image. The quality factor predictor is configured to decouple the image feature information and determine a compression quality factor. The flexible controller is configured to determine a quality factor modulation parameter based on the compressed quality factor. And the image restorer is configured to perform restoration processing on the compressed image according to the image feature information and the quality factor modulation parameter to generate a target image. Thus, when multi-source images with different compression qualities are processed, efficient and accurate recovery of the compressed images is realized through a closed-loop process of feature decoupling-dynamic modulation-adaptive recovery.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing, and more specifically, to an image processing model, a server, an image processing method, and a computer-readable storage medium. Background Art

[0002] In the related art, due to bandwidth limitations and data storage pressure, cameras usually need to compress and transmit images. However, images often lose part of the image information during the compression process, and compression artifacts may appear. In this way, the quality of the images captured by the camera will be reduced, affecting subsequent applications. Summary of the invention

[0003] The present application provides an image processing model, a server, an image processing method and a computer-readable storage medium.

[0004] The embodiment of the present application provides an image processing model, the image processing model comprising a feature extractor, a quality factor predictor, a flexible controller and an image restorer;

[0005] The feature extractor is configured to determine image feature information of the compressed image based on the acquired compressed image;

[0006] The quality factor predictor is configured to perform decoupling processing on the image feature information to determine a compression quality factor;

[0007] The flexible controller is configured to determine a quality factor modulation parameter based on the compression quality factor;

[0008] The image restorer is configured to perform restoration processing on the compressed image according to the image feature information and the quality factor modulation parameter to generate a target image.

[0009] Thus, the image processing model provided by the present application includes a feature extractor, a quality factor predictor, a flexible controller and an image restorer. Among them, the feature extractor can determine the image feature information of the compressed image based on the acquired compressed image. The quality factor predictor can decouple the image feature information and determine the compression quality factor. The flexible controller can determine the quality factor modulation parameter based on the compression quality factor. The image restorer can restore the compressed image based on the image feature information and the quality factor modulation parameter to generate a target image. In this way, when processing multi-source images with different compression qualities, the image processing model realizes efficient and accurate restoration of the compressed image through a closed-loop process of feature decoupling-dynamic modulation-adaptive restoration. In addition, the compression quality factor is used to guide the restoration process, so that the image processing model has stronger generalization ability and scene adaptability, avoiding the cumbersome process of training separate image processing models for different compression quality factors, and improving development efficiency and deployment flexibility.

[0010] In some implementations, the feature extractor is configured to:

[0011] The compressed image information is extracted and processed in stages to determine the image feature information, wherein the extraction process includes at least one sub-extraction processing stage, each of the sub-extraction processing stages is performed based on a corresponding first preset number of channels, a current sub-extraction processing stage is performed based on sub-image feature information determined in a previous sub-extraction processing stage, and the image feature information is sub-image feature information determined in the last sub-extraction processing stage.

[0012] In this way, the feature extractor can extract and process the compressed image information in stages to determine the image feature information, wherein the extraction process includes at least one sub-extraction process stage, each sub-extraction process stage is performed based on the corresponding first preset number of channels, the current sub-extraction process stage is performed based on the sub-image feature information determined in the previous sub-extraction process stage, and the image feature information is the sub-image feature information determined in the last sub-extraction process stage. In this way, gradually extracting features at different levels can better capture the detailed information of the image and enhance the robustness of feature extraction. In addition, the feature information extracted in stages can be conveniently used in subsequent image restoration processing to improve overall processing efficiency.

[0013] In some embodiments, the feature extractor is configured to:

[0014] Based on a first preset number of convolution kernels and the first preset number of channels corresponding to the current sub-extraction processing stage, a convolution operation is performed on the input compressed image information to determine a convolution output result, wherein the input compressed image information is sub-image feature information outputted in the previous sub-extraction processing stage;

[0015] Performing a normalization operation on the convolution output result to determine a normalized result;

[0016] Based on a preset activation function, converting the normalized result to determine an activation result;

[0017] Based on a second preset number of convolution kernels corresponding to the current sub-extraction processing stage, performing channel alignment processing on the activation result to determine temporary current sub-image feature information;

[0018] Based on a third preset number of convolution kernels and a preset step-size convolution operation corresponding to the current sub-extraction processing stage, down-sampling processing is performed on the temporary current sub-image feature information to determine the current sub-image feature information.

[0019] In this way, the feature extractor can perform a convolution operation on the input compressed image information based on the first preset number of convolution kernels and the first preset number of channels corresponding to the current sub-extraction processing stage, and determine the convolution output result, where the input compressed image information is the sub-image feature information outputted by the previous sub-extraction processing stage. Next, the feature extractor can perform a normalization operation on the convolution output result to determine the normalization result. Then, the feature extractor can perform a conversion process on the normalization result based on a preset activation function to determine the activation result. Subsequently, the feature extractor can perform a channel alignment process on the activation result based on the second preset number of convolution kernels corresponding to the current sub-extraction processing stage to determine the temporary current sub-image feature information. Finally, the feature extractor can perform a down-sampling process on the temporary current sub-image feature information based on the third preset number of convolution kernels and the preset step size convolution operation corresponding to the current sub-extraction processing stage to determine the current sub-image feature information. In this way, through operations such as convolution, normalization, activation and down-sampling, the feature extractor can effectively extract deep features from the compressed image, and provide information support for subsequent image restoration and compression quality factor prediction.

[0020] In certain embodiments, the flexible controller is configured to:

[0021] Based on a preset embedding matrix, determining a correlation relationship of the compression quality factors according to the compression quality factors;

[0022] The quality factor modulation parameter is determined according to the association relationship.

[0023] In this way, the flexible controller can determine the correlation relationship of the compression quality factors according to the compression quality factors based on the preset embedding matrix. Then, the flexible controller can determine the quality factor modulation parameter according to the correlation relationship. In this way, the flexible controller can obtain the correlation relationship between different compression quality factors, and determine the quality factor modulation parameter according to the correlation relationship, thereby enhancing the adaptability to images with various compression degrees based on the quality factor modulation parameter and improving the generalization ability.

[0024] In some embodiments, the image restorer is configured to:

[0025] According to the image feature information and the quality factor modulation parameters, the compressed image is restored in stages to generate the target image, wherein the restoration process includes at least one sub-restoration processing stage, each of the sub-restoration processing stages is performed based on a corresponding second preset number of channels, the current sub-restoration processing stage is performed based on a sub-restoration image determined in a previous sub-restoration processing stage, and the target image is a sub-restoration image generated in the last sub-restoration processing stage.

[0026] In this way, the image restorer can restore the compressed image in stages according to the image feature information and the quality factor modulation parameters to generate a target image, wherein the restoration process includes at least one sub-restoration process stage, each sub-restoration process stage is performed based on the corresponding second preset channel number, the current sub-restoration process stage is performed based on the sub-restoration image determined in the previous sub-restoration process stage, and the target image is the sub-restoration image generated in the last sub-restoration process stage. In this way, through the staged restoration process, the image restorer can effectively remove compression artifacts, improve image clarity, and restore image detail information to generate high-quality images.

[0027] In some embodiments, the number of sub-restoration processing stages of the restoration process is the same as the number of sub-extraction processing stages of the extraction process.

[0028] In this way, the number of sub-restoration processing stages of the restoration process is the same as the number of sub-extraction processing stages of the extraction process. In this way, the corresponding relationship between the number of sub-restoration processing stages of the restoration process and the number of sub-extraction processing stages of the extraction process ensures the flow of information in the feature extraction and image restoration process. The sub-restoration processing stage can make full use of deep feature information, thereby better restoring image detail information and improving image restoration effect.

[0029] In some embodiments, the image restorer is configured to:

[0030] Performing upsampling processing on the image feature information to generate an upsampling processing result;

[0031] Performing fusion processing on the upsampling processing results to generate a fusion processing result;

[0032] According to the quality factor modulation parameter, an affine transformation is performed on the fusion processing result to generate a current sub-restored image.

[0033] In this way, the image restorer is configured to perform upsampling processing on the image feature information to generate an upsampling processing result. Next, the image restorer is configured to perform fusion processing on the upsampling processing result to generate a fusion processing result. Finally, the image restorer is configured to perform affine transformation processing on the fusion processing result according to the quality factor modulation parameter to generate the current sub-restored image. In this way, through upsampling, fusion processing and affine transformation processing, the resolution and quality of the image are gradually restored, and the restoration strategy is dynamically adjusted according to the compression degree to enhance the generalization ability of the image processing model.

[0034] In certain embodiments, the quality factor predictor employs a multilayer perceptron.

[0035] Thus, the quality factor predictor adopts a multilayer perceptron. In this way, the multilayer perceptron can learn the complex mapping relationship between input and output, thereby achieving high-precision prediction of the compression quality factor. In addition, the structure of the multilayer perceptron can be flexibly adjusted to better adapt to different task requirements.

[0036] In some embodiments, the quality factor predictor is optimized based on a first loss function, wherein the first loss function is Where N represents the batch size of training images, represents the predicted compression quality factor for each training image, Represents the actual compression quality factor of each training image.

[0037] In this way, the quality factor predictor is optimized based on the first loss function, which is Where N represents the batch size of training images, represents the predicted compression quality factor for each training image, Represents the actual compression quality factor of each training image. In this way, the first loss function can effectively evaluate the error between the predicted result and the true value, and optimize the image processing model parameters through the back propagation algorithm, so as to predict the compression quality factor with high accuracy.

[0038] In some embodiments, the image processing model is optimized using a comprehensive loss function, wherein the comprehensive loss function is Loss all =Loss res +λLoss QF , where Loss QFis the first loss function, Loss res is the second loss function, λ is the parameter for weight balance between the first loss function and the second loss function, and the second loss function is Where N represents the batch size of training images, represents the restored image, Represents the original high-resolution image.

[0039] In this way, the image processing model is optimized using a comprehensive loss function, which is Loss all =Loss res +λLoss QF , where Loss QF is the first loss function, Loss res is the second loss function, λ is the parameter for weight balance between the first loss function and the second loss function, and the second loss function is Where N represents the batch size of training images, represents the restored image, represents the original high-resolution image. In this way, the second loss function can effectively evaluate the error between the restored image and the original high-resolution image, and optimize the image processing model parameters through the back propagation algorithm, thereby improving the image restoration effect. By adjusting the value of λ, the two loss functions can be balanced, and the impact of the comprehensive loss function on the optimization of the image processing model can be determined, so as to better adapt to different task requirements.

[0040] An embodiment of the present application provides a server, on which the above-mentioned image processing model is deployed.

[0041] The present application provides an image processing method, which is based on the above-mentioned image processing model and includes:

[0042] Determining image feature information of the compressed image according to the acquired compressed image;

[0043] Decoupling the image feature information to determine a compression quality factor;

[0044] Determining a quality factor modulation parameter according to the compression quality factor;

[0045] The compressed image is restored according to the image feature information and the quality factor modulation parameter to generate a target image.

[0046] In this way, the image processing model determines the image feature information of the compressed image based on the acquired compressed image. Next, the image processing model decouples the image feature information and determines the compression quality factor. Then, the image processing model determines the quality factor modulation parameter based on the compression quality factor. Finally, the image processing model restores the compressed image based on the image feature information and the quality factor modulation parameter to generate the target image. In this way, based on the above-mentioned image processing model, the resolution and detail information of the image can be improved, making the image clearer and more realistic.

[0047] An embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned image processing method are implemented.

[0048] Additional aspects and advantages of the embodiments of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:

[0050] Figure 1 It is a structural schematic diagram of an image processing model in an embodiment of the present application;

[0051] Figure 2 It is a schematic diagram of the architecture of the image processing model of the implementation mode of the present application;

[0052] Figure 3 It is a flowchart of the image processing method according to an embodiment of the present application. DETAILED DESCRIPTION

[0053] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals represent the same or similar elements or elements having the same or similar functions from beginning to end. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the embodiments of the present application, and cannot be understood as limiting the embodiments of the present application.

[0054] In smart cockpit applications, cameras can be used to implement multiple functions such as driver monitoring, passenger behavior analysis, gesture recognition, etc. However, due to bandwidth limitations and data storage pressure, images captured by cameras usually require image compression algorithms for compression transmission, such as the JPEG (Joint Photographic Experts Group) image compression algorithm. Although these image compression algorithms can effectively reduce the size of image files, they also bring negative effects such as image information loss and compression artifacts. These negative effects will significantly reduce the quality of images captured by the camera, thereby affecting the accuracy and reliability of smart cockpit applications such as user status detection, fatigue monitoring, and gesture recognition.

[0055] Taking the JPEG compression algorithm as an example, PEG is a lossy compression algorithm that achieves compression by removing redundant information from the image. This compression process will lose some image information, resulting in a decrease in image quality, and compression artifacts such as block effects, blurring, and color distortion may appear. In user status detection applications, the system needs to recognize features such as the user's facial expressions and head posture, and compression artifacts will affect the recognition accuracy of these features, resulting in the system being unable to accurately judge the user's fatigue level, distraction, etc. In gesture recognition applications, the system needs to recognize the gestures of passengers, and compression artifacts will affect the recognition accuracy of gestures, resulting in the system being unable to accurately judge the passenger's intentions, such as whether the air conditioning temperature needs to be adjusted or music needs to be played.

[0056] Block effect refers to the fact that during the image compression process, obvious boundaries may appear between adjacent pixel blocks, forming a blocky structure, which affects the smoothness of the image.

[0057] Blur means that the JPEG compression algorithm will reduce the resolution and detail information of the image, causing the image to appear blurry and affecting the clarity of the image.

[0058] Color distortion refers to the fact that the JPEG compression algorithm may change the saturation and brightness of certain colors in the image, causing the image to appear distorted.

[0059] Based on the above questions, please refer to Figure 1 An embodiment of the present application provides an image processing model, which includes a feature extractor, a quality factor predictor, a flexible controller and an image restorer.

[0060] Among them, the feature extractor can determine the image feature information of the compressed image based on the acquired compressed image.

[0061] The quality factor predictor can decouple image feature information and determine the compression quality factor.

[0062] The flexible controller is capable of determining a quality factor modulation parameter according to the compression quality factor.

[0063] The image restorer can restore the compressed image according to the image feature information and the quality factor modulation parameters to generate the target image.

[0064] Specifically, see Figure 2 The image processing model is an end-to-end image processing model based on deep convolutional neural networks (DCNN). Through the collaborative work of four modules: feature extractor, quality factor predictor, flexible controller and image restorer, it realizes the removal of artifacts and blur restoration of the acquired compressed image to generate high-quality images. In some embodiments, the deep convolutional neural network can use other advanced deep learning frameworks, such as vision transformer (ViT), and a hybrid architecture of deep convolutional neural network and vision transformer, which is not limited here. The unique design and adjustability of the image processing model enable it to adapt to different compression conditions and generate high-quality images, providing important support for smart cockpit applications.

[0065] Deep Convolutional Neural Networks (DCNN) is a type of image processing model with a deep network structure that can extract abstract features to achieve functions such as image recognition, image classification, and target detection.

[0066] Please refer to Figure 2 , the feature extractor is a module in the image processing model used to extract the features of the input compressed image. Through a series of operations such as convolution and pooling, the compressed image is extracted and processed to determine the feature representation with higher-level semantic information, that is, the image feature information, to provide data information support for subsequent tasks. In other words, the feature extractor converts raw data such as pixel values ​​into features with higher-level semantic information, such as edges, textures, shapes, etc. These features can represent the semantic content of the image, such as faces, vehicles, and roads. In the feature extractor, deep convolutional network architectures such as ResNet, UNet, or Xception can be selected for feature extraction, but it is necessary to ensure that the structure of the extractor corresponds to the network architecture in the image restorer in order to realize jump connections. The realization of feature reuse and information flow through jump connections can alleviate the gradient vanishing problem of the image processing model, accelerate the convergence of the image processing model, and prevent network degradation.

[0067] Compressed images refer to a digital representation of an image, that is, the original image is processed through a certain compression algorithm to reduce the file size and make it more suitable for storage and transmission.

[0068] Image feature information refers to the feature information extracted from the image, which contains the feature information of different levels and scales of the image. Each layer of feature map corresponds to a different level of abstraction of the image. For example, a low-level feature map may contain basic features such as edges and textures, while a high-level feature map may contain advanced semantic features such as objects and scenes.

[0069] The compression quality factor QF is a parameter used in the image compression algorithm to control the degree of compression and image quality, that is, to adjust the quantization table during the compression process, and the quantization table determines which image information can be discarded during the compression process. The compression quality factor QF is an integer between 0 and 100, where 0 represents the maximum compression (lowest quality) and 100 represents the minimum compression (highest quality). A higher QF value means a smaller value in the quantization table, so more image details are retained, but the corresponding file size will also be larger. A lower QF value means a larger value in the quantization table, so more image details are discarded, but the file size will be smaller. The quality factor predictor can predict the compression quality factor QF of the image based on the image feature information output by the feature extractor, that is, convert the image feature information into an understandable QF value, thereby helping the image restorer to better restore the image. For example, when the image compression degree is high, the QF value is large, and the image restorer will pay more attention to artifact removal; when the compression degree is low, the QF value is small, and the image restorer will pay more attention to detail retention.

[0070] Decoupling refers to separating the high-order features of the quality factor that can predict the compression quality factor from other image feature information, so that the image processing model can flexibly process images with different compression levels. After decoupling, the image processing model no longer needs to train a separate image processing model for each specific QF value, but can flexibly process images with different compression levels.

[0071] Flexible controllers usually use neural network structures such as multilayer perceptrons (MLPs), receive deep features output by feature extractors as input, predict the compression quality factor of the image based on the received deep features, and output a series of quality factor modulation parameters (α, β). These quality factor modulation parameter pairs are applied to different scales of the image restorer to adjust the restoration strategy. The flexible controller adjusts the strength of the image restorer for artifact removal and detail retention according to the size of the QF value. When the QF value is high, the flexible controller outputs smaller α and β values, so that the image restorer retains more image details and reduces artifact removal. When the QF value is low, the flexible controller outputs larger α and β values, so that the image restorer removes more image artifacts and reduces detail retention.

[0072] Multilayer Perceptron refers to one of the most basic structures in neural networks, which consists of multiple layers of neurons, each layer is fully connected to the next layer. MLP contains at least one hidden layer, and the intermediate layers except the input layer and the output layer are called hidden layers.

[0073] The quality factor modulation parameters refer to the parameter pair (α, β) output by the flexible controller, which are applied to different scales of the image restorer to adjust the restoration strategy. The α parameter controls the strength of the image restorer for artifact removal, and the β parameter controls the strength of the image restorer for detail preservation.

[0074] The image restorer refers to a deep convolutional neural network, which receives other image feature information output by the feature extractor except the high-order features of the quality factor and the quality factor modulation parameter pair as input, and outputs a restored clear image.

[0075] Restoration processing refers to restoring the image to its original high-quality state by removing compression artifacts in the compressed image and enhancing image details.

[0076] It should be noted that the image processing model provided in the embodiment of the present application is used to restore the image after compression processing by the compression algorithm to a high-quality clear image. The compression algorithm includes the JPEG compression algorithm, the PNG-8 compression algorithm, the WebP compression algorithm and the TIFF compression algorithm. In some embodiments, the compressed image can be expanded to a multi-frame image sequence, and the information of the time dimension is used to further improve the artifact repair and detail restoration effects. In addition, other auxiliary data (such as sensor data or environmental parameters) can also be introduced to improve the adaptability of the image processing model to different scenes through multimodal learning.

[0077] First, the feature extractor extracts the deep feature information of the image, that is, the image feature information, from the input compressed image. Then, the quality factor predictor decouples the extracted image feature information, decouples the information related to QF, and predicts the QF value of the image, that is, the compression quality factor. Then, the flexible controller generates a series of modulation parameters (α, β) based on the predicted QF value and applies them to the image restorer, thereby dynamically adjusting the restoration strategy according to the QF value. Finally, the image restorer uses the image feature information extracted by the feature extractor and the quality factor modulation parameters to reconstruct the compressed low-quality image into a clear high-quality image.

[0078] In summary, the image processing model provided by the present application includes a feature extractor, a quality factor predictor, a flexible controller and an image restorer. Among them, the feature extractor can determine the image feature information of the compressed image based on the acquired compressed image. The quality factor predictor can decouple the image feature information and determine the compression quality factor. The flexible controller can determine the quality factor modulation parameter based on the compression quality factor. The image restorer can restore the compressed image based on the image feature information and the quality factor modulation parameter to generate a target image. In this way, when processing multi-source images with different compression qualities, the image processing model realizes efficient and accurate restoration of the compressed image through a closed-loop process of feature decoupling-dynamic modulation-adaptive restoration. In addition, the compression quality factor is used to guide the restoration process, so that the image processing model has stronger generalization ability and scene adaptability, avoiding the cumbersome process of training separate image processing models for different compression quality factors, and improving development efficiency and deployment flexibility.

[0079] In some embodiments, the feature extractor can perform extraction processing on compressed image information in stages to determine image feature information, wherein the extraction processing includes at least one sub-extraction processing stage, each sub-extraction processing stage is performed based on the corresponding first preset number of channels, the current sub-extraction processing stage is performed based on the sub-image feature information determined in the previous sub-extraction processing stage, and the image feature information is the sub-image feature information determined in the last sub-extraction processing stage.

[0080] Specifically, the sub-extraction processing stage refers to multiple extraction processing stages in which the feature extractor extracts features from the compressed image. Each stage extracts features at different levels of the image. Each sub-extraction processing stage is implemented by a residual block, and each residual block includes multiple convolutional layers, ReLU activation functions, and batch normalization operations.

[0081] The convolution operation can be used to extract local features of an image by sliding a small convolution kernel (filter) on the image and calculating the dot product between the convolution kernel and the local area of ​​the image to extract features such as edges, textures, and colors of the image.

[0082] The ReLU (Rectified Linear Unit) activation function can output negative values ​​as 0 and keep positive values ​​unchanged, thereby avoiding the gradient vanishing problem, and can increase the nonlinearity of the image processing model and improve the generalization ability of the image processing model.

[0083] Batch normalization can speed up the training of image processing models and improve the generalization ability of image processing models. By normalizing the input data of each batch, the mean of the data is adjusted to 0 and the variance is adjusted to 1. In this way, batch normalization can reduce the sensitivity of the image processing model to initialization and improve the stability of the image processing model.

[0084] Residual blocks can be used to build deep networks. Residual blocks can effectively solve the gradient vanishing problem in deep networks by introducing residual connections, that is, using residual connections to directly connect the input to the output so that the gradient can be directly passed to the deep network.

[0085] Residual blocks, batch normalization, ReLU activation, and convolution operations work together to help the feature extractor effectively learn the complex features of the image. Figure 2 In the example given in the implementation mode of the present application, the extraction process of the feature extractor is divided into four sub-extraction processing stages. In some implementation modes, the number of sub-extraction processing stages can be determined according to specific task requirements and is not limited here.

[0086] The first preset number of channels refers to the number of channels that are preset. The number of channels refers to the depth of feature extraction, which is represented by the number of feature maps contained in the feature map in some embodiments. As the extraction process progresses, the number of channels usually increases to extract more complex and higher-level features. Figure 2 In the examples given in the implementation of the present application, the first preset number of channels in each sub-extraction processing stage is 64, 128, 256 and 512. In some implementations, the number of channels in each sub-extraction processing stage can be determined according to specific task requirements and is not limited here.

[0087] Each sub-extraction processing stage is performed based on the sub-image feature information determined in the previous sub-stage. This can help the feature extractor gradually learn and understand the relationship between image features and perform feature fusion. In addition, extraction based on the feature information of the previous sub-stage can avoid repeated extraction of the same features and improve the efficiency of feature extraction.

[0088] The image feature information is the sub-image feature information determined in the last sub-extraction processing stage. The feature information extracted in the last sub-extraction processing stage includes the highest-level features of the image, such as global semantic information, image style, etc.

[0089] It should be noted that in addition to building feature extractors with deep convolutional network architectures such as ResNet, UNet, or Xception, other types of feature extraction networks, such as DenseNet or EfficientNet, can also be used to improve the efficiency and quality of feature extraction. In addition, the feature extractor can also introduce a multi-scale feature fusion module to more comprehensively capture image details at different resolutions. Alternatively, an attention mechanism (such as SE module or CBAM) can be added to the feature extraction stage to more accurately focus on artifacts and blurred areas.

[0090] In this way, the feature extractor can extract and process the compressed image information in stages to determine the image feature information, wherein the extraction process includes at least one sub-extraction process stage, each sub-extraction process stage is performed based on the corresponding first preset number of channels, the current sub-extraction process stage is performed based on the sub-image feature information determined in the previous sub-extraction process stage, and the image feature information is the sub-image feature information determined in the last sub-extraction process stage. In this way, gradually extracting features at different levels can better capture the detailed information of the image and enhance the robustness of feature extraction. In addition, the feature information extracted in stages can be conveniently used in subsequent image restoration processing to improve overall processing efficiency.

[0091] In some embodiments, the feature extractor can perform a convolution operation on the input compressed image information based on a first preset number of convolution kernels and a first preset number of channels corresponding to the current sub-extraction processing stage, and determine the convolution output result, where the input compressed image information is the sub-image feature information output by the previous sub-extraction processing stage.

[0092] Next, the feature extractor can perform a normalization operation on the convolution output result to determine the normalized result.

[0093] Then, the feature extractor can transform the normalized result based on a preset activation function to determine the activation result.

[0094] Subsequently, the feature extractor can perform channel alignment processing on the activation results based on a second preset number of convolution kernels corresponding to the current sub-extraction processing stage to determine temporary current sub-image feature information.

[0095] Finally, the feature extractor can downsample the temporary current sub-image feature information based on the third preset number of convolution kernels and the preset step-size convolution operation corresponding to the current sub-extraction processing stage to determine the current sub-image feature information.

[0096] Specifically, the processing flow of each sub-extraction processing stage is as follows: First, a convolution operation is performed on the sub-image feature information output by the previous sub-extraction processing stage using a first preset number of convolution kernels and a first preset number of channels to extract local features of the image. In the first sub-extraction processing stage, the input compressed image information is the compressed image obtained by the feature extractor.

[0097] Next, the convolution output is normalized to eliminate the dimensional differences between different features and improve the convergence and generalization capabilities of the feature processor. Dimensional differences refer to the differences in numerical values ​​and ranges between different features. For example, in image features, the pixel values ​​of different color channels may be at different levels. For example, the pixel values ​​of the red channel may be in the range of 0-255, while the pixel values ​​of the blue channel may be in the range of 0-150. This dimensional difference can cause problems such as gradient vanishing or gradient exploding during the training of the image processing model, affecting the performance of the feature extractor.

[0098] Then, the normalized result is transformed nonlinearly using a preset activation function to enhance the nonlinear expression capability of the image processing model. The nonlinear transformation of the activation function refers to mapping the output of the linear function to the output space of the nonlinear function, which enables the feature extractor to learn more complex feature representations.

[0099] Subsequently, a second preset number of convolution kernels is used to perform channel alignment on the activation results to ensure that the number of feature map channels at different stages is consistent, which is convenient for subsequent operations.

[0100] Finally, a third preset number of convolution kernels and a stride convolution operation are used to downsample the temporary current sub-image feature information, reduce the feature map resolution, reduce the amount of calculation, and capture more abstract features of the image.

[0101] It should be noted that in some implementations, after the channel alignment process and the downsampling process, it is also necessary to perform a normalization operation and a nonlinear conversion using a preset activation function.

[0102] Please refer to Figure 2 In the example given in the implementation mode of the present application, for the input compressed image, the feature extractor first uses 64 64-channel convolution kernels to perform a convolution operation on the image, and then performs batch normalization and ReLU activation to obtain the feature map of the first stage (i.e., sub-image feature information). Next, 128 128-channel convolution kernels are used to perform a convolution operation on the feature map of the first stage, and batch normalization and ReLU activation are performed to obtain the feature map of the second stage. And so on, until the feature map of the fourth stage is obtained. The feature map of each stage undergoes channel alignment and downsampling operations, and finally obtains deep features for image restoration and compression quality factor prediction.

[0103] Each sub-extraction processing stage uses two 3×3 convolution layers (the first preset number of convolution kernels) in the residual block to extract local features of the image. The 3×3 convolution layer can capture local information of the image and generate feature maps of multiple channels. Next, the convolution output results are normalized, such as using batch normalization to eliminate the scale differences between different feature maps. Then, a preset activation function, such as ReLU, is used to perform a nonlinear transformation on the normalized results to enhance the expressive power of the image processing model. Subsequently, since the number of output channels of the two 3×3 convolution layers may be different, in order to merge them, 1×1 convolution (the second preset number of convolution kernels) is required for channel alignment. 1×1 convolution can convert feature maps of different channels into the same number of channels by learning a weight matrix, thereby achieving channel alignment. Finally, in order to reduce the size of the feature map, a 2×2 stride convolution (the third preset number of convolution kernels and the preset stride convolution) can be used for downsampling. A stride of 2 means that the convolution kernel moves 2 pixels each time, thereby halving the size of the feature map.

[0104] In this way, the feature extractor can perform a convolution operation on the input compressed image information based on the first preset number of convolution kernels and the first preset number of channels corresponding to the current sub-extraction processing stage, and determine the convolution output result, where the input compressed image information is the sub-image feature information outputted by the previous sub-extraction processing stage. Next, the feature extractor can perform a normalization operation on the convolution output result to determine the normalization result. Then, the feature extractor can perform a conversion process on the normalization result based on a preset activation function to determine the activation result. Subsequently, the feature extractor can perform a channel alignment process on the activation result based on the second preset number of convolution kernels corresponding to the current sub-extraction processing stage to determine the temporary current sub-image feature information. Finally, the feature extractor can perform a down-sampling process on the temporary current sub-image feature information based on the third preset number of convolution kernels and the preset step size convolution operation corresponding to the current sub-extraction processing stage to determine the current sub-image feature information. In this way, through operations such as convolution, normalization, activation and down-sampling, the feature extractor can effectively extract deep features from the compressed image, and provide information support for subsequent image restoration and compression quality factor prediction.

[0105] In some embodiments, the flexible controller can determine the correlation relationship of the compression quality factors according to the compression quality factors based on a preset embedding matrix.

[0106] Then, the flexible controller can determine the quality factor modulation parameters according to the association relationship.

[0107] Specifically, the preset embedding matrix can be used in the image processing model to map the compression quality factors to a low-dimensional continuous vector space so that the flexible controller can learn the semantic relationship and contextual information between the quality factors. The number of rows in the preset embedding matrix is ​​equal to the number of compression quality factors, and the number of columns is equal to the dimension of the embedding vector. Each compression quality factor corresponds to a row in the matrix, and its value determines the position of the compression quality factor in the embedding space.

[0108] The correlation relationship of compression quality factors refers to the mutual influence and interaction between different compression quality factors on the image restoration task. In some embodiments, the correlation relationship of compression quality factors can be determined by calculating the feature similarity (such as cosine similarity) of images with different quality factors to analyze the influence of quality factors on image features.

[0109] The quality factor modulation parameters refer to the parameter pair generated by the flexible controller in the image processing model, denoted as α parameter and β parameter, which can be used to dynamically adjust the processing strategy of the image restorer according to the compression quality factor of the image. The α and β parameters control the affine transformation of the quality factor attention block in the image restorer, respectively. The α parameter controls the feature map scaling, and the β parameter controls the feature map translation. By adjusting the α and β parameters, the degree of artifact removal and detail preservation can be balanced. For example, increasing the α parameter will enhance the detail recovery effect, but may result in insufficient artifact removal. Reducing the α parameter will enhance the artifact removal effect, but may result in detail loss.

[0110] The flexible controller uses a preset embedding matrix to map the compression quality factor into a low-dimensional continuous vector space. The dimension of the embedding matrix can be adjusted according to actual conditions.

[0111] Then, the flexible controller distinguishes different compression quality factors in the vector space through learning and establishes the correlation between the compression quality factors.

[0112] Finally, the flexible controller determines the quality factor modulation parameters based on the learned associations. These quality factor modulation parameters are used to control the quality factor attention block in the image restorer to achieve a balance between artifact removal and detail preservation.

[0113] It should be noted that the flexible controller can also use dynamic weight generation networks (Dynamic Weight Networks) to replace the current modulation parameter generation module and directly adjust the restoration strategy by learning specific task weights. In addition, the flexible controller can also introduce more complex attention mechanisms (such as dynamic convolution or dynamic kernel selection) to adaptively adjust the restoration process according to the input image quality.

[0114] In this way, the flexible controller can determine the correlation relationship of the compression quality factors according to the compression quality factors based on the preset embedding matrix. Then, the flexible controller can determine the quality factor modulation parameter according to the correlation relationship. In this way, the flexible controller can obtain the correlation relationship between different compression quality factors, and determine the quality factor modulation parameter according to the correlation relationship, thereby enhancing the adaptability to images with various compression degrees based on the quality factor modulation parameter and improving the generalization ability.

[0115] In certain embodiments, the image restorer is capable of:

[0116] According to the image feature information and the quality factor modulation parameters, the compressed image is restored in stages to generate a target image, wherein the restoration process includes at least one sub-restoration processing stage, each sub-restoration processing stage is performed based on the corresponding second preset number of channels, the current sub-restoration processing stage is performed based on the sub-restoration image determined in the previous sub-restoration processing stage, and the target image is the sub-restoration image generated in the last sub-restoration processing stage.

[0117] Specifically, the restoration process refers to the process of removing artifacts and blurring from a compressed image to ultimately generate a clear image. In the implementation of the present application, the entire restoration process is divided into multiple stages (i.e., sub-restoration stages), each of which is responsible for processing a specific part or specific feature of the image. It should be noted that the division of the restoration process is related to the division of the extraction process in the feature extractor.

[0118] The second preset number of channels refers to the number of feature channels used in each sub-restoration processing stage, and is related to the number of channels used in each stage of the feature extractor extraction processing.

[0119] The sub-restored image refers to the intermediate image generated in each sub-restore processing stage, which serves as the input of the next stage processing.

[0120] The target image refers to the final image generated by the last sub-restoration processing stage, that is, the clear image after restoration.

[0121] The image restoration process is divided into multiple stages to achieve fine control of the image restoration effect and improve the efficiency of the model. Then, the deep features of the image (i.e., image feature information) are used to identify artifacts and blurred areas in the image and perform targeted repairs. The image restoration strategy is dynamically adjusted according to the degree of image compression to achieve a balance between artifact removal and detail retention.

[0122] Through the iteration of multiple sub-restoration processing stages, the clarity and details of the image are gradually improved, and finally a high-quality target image is generated.

[0123] In this way, the image restorer can restore the compressed image in stages according to the image feature information and the quality factor modulation parameters to generate a target image, wherein the restoration process includes at least one sub-restoration process stage, each sub-restoration process stage is performed based on the corresponding second preset channel number, the current sub-restoration process stage is performed based on the sub-restoration image determined in the previous sub-restoration process stage, and the target image is the sub-restoration image generated in the last sub-restoration process stage. In this way, through the staged restoration process, the image restorer can effectively remove compression artifacts, improve image clarity, and restore image detail information to generate high-quality images.

[0124] In some embodiments, the number of sub-restoration processing stages of the restoration process is the same as the number of sub-extraction processing stages of the extraction process.

[0125] Specifically, by setting the same number of sub-restoration processing stages and sub-extraction processing stages, feature reuse can be achieved in the image restorer. Each sub-restoration processing stage can obtain a feature map (i.e., image feature information) from the corresponding sub-extraction processing stage for further processing and reconstruction.

[0126] The output of each sub-extraction processing stage in the feature extractor is connected to the corresponding sub-restoration processing stage in the image restorer through jump connections, ensuring smooth flow of information between different stages, avoiding the gradient vanishing problem, and accelerating model convergence.

[0127] In this way, the number of sub-restoration processing stages of the restoration process is the same as the number of sub-extraction processing stages of the extraction process. In this way, the corresponding relationship between the number of sub-restoration processing stages of the restoration process and the number of sub-extraction processing stages of the extraction process ensures the flow of information in the feature extraction and image restoration process. The sub-restoration processing stage can make full use of deep feature information, thereby better restoring image detail information and improving image restoration effect.

[0128] In certain embodiments, the image restorer is capable of:

[0129] Perform upsampling processing on the image feature information to generate an upsampling processing result;

[0130] Performing fusion processing on the upsampling processing results to generate a fusion processing result;

[0131] According to the quality factor modulation parameter, the fusion processing result is affine transformed to generate the current sub-restored image.

[0132] Specifically, upsampling refers to an operation used to increase the resolution of an image to make it look clearer, by adding new pixels to the image through interpolation methods, thereby increasing the size of the image. Upsampling can be implemented by nearest neighbor interpolation, bilinear interpolation, and bicubic interpolation. Nearest neighbor interpolation refers to setting the new pixel to the color value of the pixel closest to it. Bilinear interpolation refers to setting the new pixel to the average color value of the four surrounding pixels. Bicubic interpolation refers to the use of more complex interpolation methods, which can produce smoother images.

[0133] Fusion processing refers to fusing the feature maps generated during the upsampling process with the feature maps of the corresponding stage in the feature extractor to obtain a more comprehensive and richer feature representation, thereby generating a clear target image.

[0134] Affine transformation processing refers to scaling, rotating, translating or shearing a geometric figure in an image in some way to obtain a new geometric figure. The affine transformation formula is as follows:

[0135]

[0136] Among them, F out and F in represents the feature map before and after the affine transformation, ⊙ represents element multiplication, represents element-by-element addition. α parameter and β parameter are parameter pairs generated by the flexible controller in the image processing model.

[0137] Please refer to Figure 2 In the example of the implementation method of this application, the image restorer is divided into four stages to gradually restore the image resolution and fuse multi-level features: the first sub-restoration processing stage, initial feature fusion, uses jump connections to obtain shallow features (such as edges, textures, etc.) from the first stage of the feature extractor. Then, the shallow features are combined with the feature map of the current stage of the restorer to retain the details. Finally, a low-resolution feature map with preliminary restoration is obtained.

[0138] In the second sub-restoration stage, upsampling and QF attention blocks are applied, and the feature map output from the first sub-restoration stage is upsampled to a higher resolution using bilinear interpolation. Then, skip connections are used to obtain the features (such as local structure, etc.) of the second stage of the feature extractor. Four QF attention blocks are applied, and the feature map is affine transformed using the parameter pair (α, β). Finally, a medium-resolution feature map with reduced artifacts and enhanced details is output.

[0139] In the third sub-restoration processing stage, further refinement and resolution improvement are performed, and upsampling is continued to a higher resolution. Then, skip connections are used to obtain the features of the third stage of the feature extractor (such as global semantic information). Four QF attention blocks are applied again, and the restoration strategy is adjusted dynamically according to the QF. Finally, a high-resolution feature map with clearer details is output.

[0140] In the fourth sub-restoration processing stage, the final restoration and output are upsampled to the original image resolution. Then, the jump connection is used to obtain the deep features (such as overall scene information) of the fourth stage of the feature extractor. And 4 QF attention blocks are applied again to further optimize the feature map. Then, the feature map is converted to RGB channels through 1×1 convolution to generate the final repaired image. Finally, the target image with artifacts eliminated and details preserved is output.

[0141] Among them, the jump connection refers to combining the shallow features (details) of the feature extractor with the deep features (semantics) of the restorer to avoid information loss.

[0142] The QF attention block refers to a special attention mechanism module that dynamically adjusts the attention on the feature map according to the quality factor modulation parameters of the image, so as to better remove artifacts and restore details.

[0143] In this way, the image restorer can perform upsampling processing on the image feature information to generate an upsampling processing result. Then, the image restorer can perform fusion processing on the upsampling processing result to generate a fusion processing result. Finally, the image restorer can perform affine transformation processing on the fusion processing result according to the quality factor modulation parameter to generate the current sub-restored image. In this way, through upsampling, fusion processing and affine transformation processing, the resolution and quality of the image are gradually restored, and the restoration strategy is dynamically adjusted according to the compression degree to enhance the generalization ability of the image processing model.

[0144] In certain embodiments, the quality factor predictor employs a multi-layer perceptron.

[0145] Specifically, the quality factor predictor adopts a multilayer perceptron (MLP) as its core network structure. MLP is a classic neural network structure consisting of multiple fully connected layers, each layer contains multiple neurons, and maps input data to output through an activation function.

[0146] The use of multi-layer perceptrons can effectively extract deep features output by the feature extractor and convert them into feature representations required for compression quality factor prediction. Through multiple layers of nonlinear transformations, MLP can capture the complex relationship between features, thereby more accurately predicting the compression quality factor of the image. In addition, the parameters of MLP can be adjusted through training to adapt to different task requirements. For example, in the quality factor predictor, the prediction accuracy and training speed can be controlled by adjusting the structure of MLP. In addition, MLP can be easily expanded to more complex network structures, such as increasing the number of hidden layers or using different activation functions. This provides greater flexibility for the model and can better adapt to different application scenarios.

[0147] It should be noted that, in some embodiments, the flexible controller also uses a multilayer perceptron.

[0148] In this way, the multilayer perceptron can learn the complex mapping relationship between input and output, thereby achieving high-precision prediction of compression quality factors. In addition, the structure of the multilayer perceptron can be flexibly adjusted to better adapt to different task requirements.

[0149] In some embodiments, the quality factor predictor is optimized based on a first loss function, where the first loss function is Where N represents the batch size of training images, represents the predicted compression quality factor for each training image, Represents the actual compression quality factor of each training image.

[0150] Specifically, the batch size of training images refers to the number of image samples that are input into the image processing model for learning during one training process. The batch size can determine the amount of memory consumed when training the image processing model and affect the computational efficiency and performance of the image processing model.

[0151] The first loss function is used to measure the difference between the model's predicted image compression quality factor and the actual quality factor. The first loss function can be L1 loss (mean absolute error, MAE) or L2 loss (mean square error, MSE).

[0152] In this way, the first loss function can effectively evaluate the error between the predicted result and the true value, and optimize the image processing model parameters through the back propagation algorithm, so that the compression quality factor can be predicted with high precision.

[0153] In some embodiments, the image processing model is optimized using a comprehensive loss function, where the comprehensive loss function is Loss all =Loss res +λLoss QF , where Loss QFis the first loss function, Loss res is the second loss function, λ is the parameter for weight balance between the first loss function and the second loss function, and the second loss function is Where N represents the batch size of training images, represents the restored image, Represents the original high-resolution image.

[0154] Specifically, the second loss function is used to measure the difference between the restored image and the original high-resolution image. This comprehensive loss function combines the image restoration loss (Loss_res) and the quality factor prediction loss (Loss_QF) to ensure that the restored image is close to the original high-resolution image at the pixel level. The QF prediction enhances the model's adaptability to different compression conditions. The weight parameter λ is also used to allow the optimization focus to be adjusted according to task requirements.

[0155] In this way, the second loss function can effectively evaluate the error between the restored image and the original high-resolution image, and optimize the image processing model parameters through the back propagation algorithm, thereby improving the image restoration effect. By adjusting the value of λ, the two loss functions can be balanced and the impact of the comprehensive loss function on the optimization of the image processing model can be determined, so as to better adapt to different task requirements.

[0156] An embodiment of the present application provides a server, on which the above-mentioned image processing model is deployed.

[0157] Specifically, the image processing model deployed on the server will process the input compressed images, remove compression artifacts, and improve image clarity, thereby providing high-quality image data for the smart cockpit.

[0158] See also Figure 3 The present application provides an image processing method, the method comprising:

[0159] 01: Determine image feature information of the compressed image according to the obtained compressed image;

[0160] 02: Decouple the image feature information and determine the compression quality factor;

[0161] 03: Determine the quality factor modulation parameter according to the compression quality factor;

[0162] 04: According to the image feature information and quality factor modulation parameters, the compressed image is restored to generate the target image.

[0163] The image processing method of the embodiment of the present application can be implemented by the image processing device of the embodiment of the present application. Specifically, the image processing device includes a determination module and a generation module. The determination module is used to determine the image feature information of the compressed image according to the acquired compressed image. The determination module is also used to decouple the image feature information and determine the compression quality factor. The determination module is also used to determine the quality factor modulation parameter according to the compression quality factor. The generation module is used to restore the compressed image according to the image feature information and the quality factor modulation parameter to generate a target image.

[0164] Specifically, the compressed image is efficiently processed through the image processing model, and finally a clear and detail-rich target image is generated, providing strong support for image analysis and decision-making in the intelligent cockpit.

[0165] In summary, based on the above image processing model, the resolution and detail information of the image can be improved, making the image clearer and more realistic.

[0166] The present application also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the vehicle control method described above are implemented.

[0167] It is understood that a computer program includes computer program code. The computer program code may be in source code form, object code form, executable file or some intermediate form. Computer readable storage media may include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), and software distribution medium.

[0168] In the description of this specification, the descriptions with reference to the terms "specifically", "further", "particularly", "understandably", etc. are intended to mean that the specific features, structures, materials or characteristics described in conjunction with the embodiments or examples are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms are not intended to refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they are contradictory.

[0169] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code that includes one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.

[0170] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.

Claims

1. An image processing model, characterized in that: The image processing model includes a feature extractor, a quality factor predictor, a flexible controller and an image restorer; The feature extractor is configured to determine image feature information of the compressed image based on the acquired compressed image; The quality factor predictor is configured to perform decoupling processing on the image feature information to determine a compression quality factor; The flexible controller is configured to determine a quality factor modulation parameter based on the compression quality factor; The image restorer is configured to perform restoration processing on the compressed image according to the image feature information and the quality factor modulation parameter to generate a target image.

2. The image processing model according to claim 1, characterized in that The feature extractor is configured as: The compressed image information is extracted and processed in stages to determine the image feature information, wherein the extraction process includes at least one sub-extraction processing stage, each of the sub-extraction processing stages is performed based on a corresponding first preset number of channels, a current sub-extraction processing stage is performed based on sub-image feature information determined in a previous sub-extraction processing stage, and the image feature information is sub-image feature information determined in the last sub-extraction processing stage.

3. The image processing model according to claim 2, characterized in that: The feature extractor is configured as: Based on a first preset number of convolution kernels and the first preset number of channels corresponding to the current sub-extraction processing stage, a convolution operation is performed on the input compressed image information to determine a convolution output result, wherein the input compressed image information is sub-image feature information outputted in the previous sub-extraction processing stage; Performing a normalization operation on the convolution output result to determine a normalized result; Based on a preset activation function, converting the normalized result to determine an activation result; Based on a second preset number of convolution kernels corresponding to the current sub-extraction processing stage, performing channel alignment processing on the activation result to determine temporary current sub-image feature information; Based on a third preset number of convolution kernels and a preset step-size convolution operation corresponding to the current sub-extraction processing stage, down-sampling processing is performed on the temporary current sub-image feature information to determine the current sub-image feature information.

4. The image processing model according to claim 1, characterized in that The flexible controller is configured to: Based on a preset embedding matrix, determining a correlation relationship of the compression quality factors according to the compression quality factors; The quality factor modulation parameter is determined according to the association relationship.

5. The image processing model according to claim 2, characterized in that: The image restorer is configured as follows: According to the image feature information and the quality factor modulation parameters, the compressed image is restored in stages to generate the target image, wherein the restoration process includes at least one sub-restoration processing stage, each of the sub-restoration processing stages is performed based on a corresponding second preset number of channels, the current sub-restoration processing stage is performed based on a sub-restoration image determined in a previous sub-restoration processing stage, and the target image is a sub-restoration image generated in the last sub-restoration processing stage.

6. The image processing model according to claim 5, characterized in that The number of the sub-restoration processing stages of the restoration process is the same as the number of the sub-extraction processing stages of the extraction process.

7. The image processing model according to claim 5, characterized in that The image restorer is configured as follows: Performing upsampling processing on the image feature information to generate an upsampling processing result; Performing fusion processing on the upsampling processing results to generate a fusion processing result; According to the quality factor modulation parameter, an affine transformation is performed on the fusion processing result to generate a current sub-restored image.

8. The image processing model according to claim 1, characterized in that: The quality factor predictor adopts a multi-layer perceptron.

9. The image processing model according to claim 1, characterized in that: The quality factor predictor is optimized based on a first loss function, wherein the first loss function is Where N represents the batch size of training images, represents the predicted compression quality factor for each training image, Represents the actual compression quality factor of each training image.

10. The image processing model according to claim 9, characterized in that: The image processing model is optimized using a comprehensive loss function, which is Loss all =Loss res +λLoss Qf , Among them, Loss QF is the first loss function, Loss res is the second loss function, λ is the parameter for weight balance between the first loss function and the second loss function, and the second loss function is Where N represents the batch size of training images, represents the restored image, Represents the original high-resolution image.

11. A server, characterized in that: The server is deployed with an image processing model as described in any one of claims 1-10.

12. An image processing method, characterized in that: The method is based on the image processing model according to any one of claims 1 to 10, and the method comprises: Determining image feature information of the compressed image according to the acquired compressed image; Decoupling the image feature information to determine a compression quality factor; Determining a quality factor modulation parameter according to the compression quality factor; The compressed image is restored according to the image feature information and the quality factor modulation parameter to generate a target image.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to claim 12 are implemented.