Image Processing Method, Image Processing Apparatus, and Computer-Readable Storage Medium

Through deep learning and adaptive wavelet transformation, local features and wavelet features are extracted, fusion processing is integrated to generate the degree of skin grinding, solving the problem of loss of details of traditional methods and unnatural treatment, and achieving high-quality facial grinding effect.

CN116977206BActive Publication Date: 2025-06-27ZHEJIANG HUACHUANG VISION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310803082.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-30
Publication Date
2025-06-27
Estimated Expiration
2043-06-30

AI Technical Summary

Technical Problem

Traditional facial skin treatment methods can easily lead to loss of image details or unnatural processing effects.

Method used

A facial skin grinding treatment method based on deep learning and adaptive wavelet transformation is adopted. This method extracts the local features of the image, determines the optimal wavelet function, extracts the wavelet features, and fuses the local features and the wavelet features to generate the degree of skin wear of each pixel point, and performs adaptive processing.

Benefits of technology

Effectively extract and process the detailed features of face images, achieve high-quality facial skin grinding effect, avoid details loss and unnatural problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116977206B_ABST
    Figure CN116977206B_ABST
Patent Text Reader

Abstract

The present application provides an image processing method, an image processing apparatus, and a computer-readable storage medium. The image processing method includes: obtaining an image to be processed; extracting local features of the image to be processed; determining an optimal wavelet function based on the local features; extracting wavelet features of the image to be processed according to the optimal wavelet function; fusing the local features and the wavelet features to obtain fused features; using the fused features to generate the skin smoothing degree of each pixel point of the image to be processed; and processing the image to be processed according to the skin smoothing degree of each pixel point to obtain a face skin-smoothed image. Through the above manner, the image processing apparatus effectively extracts and processes the detail features of a face image by virtue of the excellent extraction ability of texture features through adaptive wavelet transform, and achieves a high-quality face skin smoothing effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and particularly to an image processing method, an image processing device, and a computer-readable storage medium. Background Art

[0002] Face skin smoothing processing is an important technology in the field of image processing and is widely used in scenarios such as live streaming, short videos, and advertisements. Its main purpose is to smooth the human skin while retaining image details and improve the beauty of the portrait. Traditional face skin smoothing processing methods are mainly based on image filtering technology, but these methods often easily lead to the loss of image details or unnatural processing effects. Summary of the Invention

[0003] This application provides an image processing method, an image processing device, and a computer-readable storage medium.

[0004] This application provides an image processing method, and the image processing method includes:

[0005] Obtain an image to be processed;

[0006] Extract local features of the image to be processed;

[0007] Determine an optimal wavelet function based on the local features;

[0008] Extract wavelet features of the image to be processed according to the optimal wavelet function;

[0009] Fuse the local features and the wavelet features to obtain fused features;

[0010] Generate the skin smoothing degree of each pixel point of the image to be processed by using the fused features;

[0011] Process the image to be processed according to the skin smoothing degree of each pixel point to obtain a face skin smoothing image.

[0012] Among them, the process of processing the image to be processed according to the skin smoothing degree of each pixel point to obtain a face skin smoothing image includes:

[0013] Generate the skin smoothing weight of each pixel point according to the skin smoothing degree of each pixel point, where the skin smoothing weight includes an original skin smoothing weight and a neighboring skin smoothing weight;

[0014] Perform weighted processing on the original pixel value of each pixel point and the original skin smoothing weight to obtain an original skin smoothing pixel value;

[0015] Perform weighted processing on the neighboring pixel value of each pixel point and the neighboring skin smoothing weight to obtain a neighboring skin smoothing pixel value;

[0016] Add the original skin-smoothing pixel value and the adjacent skin-smoothing pixel value to obtain the face skin-smoothing pixel value of each pixel point;

[0017] Combine the face skin-smoothing pixel values of all pixel points to generate the face skin-smoothing image.

[0018] Among them, the combining the face skin-smoothing pixel values of all pixel points to generate the face skin-smoothing image includes:

[0019] Combine the face skin-smoothing pixel values of all pixel points to generate a face skin-smoothing weighted image;

[0020] Fuse the face skin-smoothing weighted image and the image to be processed to generate the final face skin-smoothing image.

[0021] Among them, the fusing the local feature and the wavelet feature to obtain a fused feature includes:

[0022] Use a convolutional layer to extract high-level features from the wavelet feature;

[0023] Fuse the local feature, the wavelet feature, and the high-level feature to obtain the fused feature.

[0024] Among them, the fusing the local feature and the wavelet feature to obtain a fused feature includes:

[0025] Use upsampling or downsampling to extract multi-scale local features of the local feature, and fuse the multi-scale local features with the wavelet feature to obtain the fused feature;

[0026] Alternatively, use upsampling or downsampling to extract multi-scale wavelet features of the wavelet feature, and fuse the multi-scale wavelet features with the local feature to obtain the fused feature.

[0027] Among them, after extracting the wavelet feature of the image to be processed according to the optimal wavelet function, the image processing method further includes:

[0028] Input the local feature and the wavelet feature into a pre-trained deep learning model;

[0029] Obtain the predicted enhancement degree of each pixel point output by the deep learning model;

[0030] Use the predicted enhancement degree of each pixel point to weight and increase the pixel value of each pixel point to obtain the texture enhancement pixel value of each pixel point.

[0031] Among them, the determining the optimal wavelet function based on the local feature includes:

[0032] Obtain a number of candidate wavelet functions according to the local features;

[0033] Obtain the wavelet function representation of each pixel point in the local features under each candidate wavelet function;

[0034] Take the similarity between the wavelet function representation and the local features as the similarity value of the candidate wavelet function;

[0035] Determine the candidate wavelet function with the highest similarity value as the optimal wavelet function.

[0036] Among them, determining the optimal wavelet function based on the local features includes:

[0037] Perform skin detection on the image to be processed, and divide the image to be processed into a skin area and a non-skin area;

[0038] Determine the first local features of the skin area and the second local features of the non-skin area based on the local features;

[0039] Determine the first optimal wavelet function based on the first local features;

[0040] Determine the second optimal wavelet function based on the second local features.

[0041] Among them, after determining the optimal wavelet function based on the local features, the image processing method further includes:

[0042] Obtain the texture information of the local features;

[0043] Determine the wavelet parameters of the optimal wavelet function according to the texture information;

[0044] Among them, the wavelet parameters include wavelet scale and / or wavelet direction.

[0045] Among them, extracting the local features of the image to be processed includes:

[0046] Process the image to be processed through an edge detection operator to obtain an edge intensity map;

[0047] Process the image to be processed through a corner detection algorithm to obtain a corner response map;

[0048] Fuse the edge intensity map and the corner response map to generate a comprehensive feature map;

[0049] Extract the local features of the image to be processed from the comprehensive feature map.

[0050] The present application also provides an image processing apparatus, which includes a processor and a memory. Program data is stored in the memory, and the processor is configured to execute the program data to implement the image processing method as described above.

[0051] The present application also provides a computer-readable storage medium, which is used to store program data. When the program data is executed by a processor, it is configured to implement the above-mentioned image processing method.

[0052] The beneficial effects of the present application are as follows: The image processing apparatus acquires an image to be processed; extracts local features of the image to be processed; determines an optimal wavelet function based on the local features; extracts wavelet features of the image to be processed according to the optimal wavelet function; fuses the local features and the wavelet features to obtain fused features; generates the skin smoothing degree of each pixel of the image to be processed by using the fused features; and processes the image to be processed according to the skin smoothing degree of each pixel to obtain a face skin-smoothing image. In the above manner, the image processing apparatus effectively extracts and processes the detailed features of the face image through the excellent ability of adaptive wavelet transform to extract texture features, and achieves a high-quality face skin-smoothing effect. Description of the Drawings

[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. Among them:

[0054] Figure 1 is a schematic flowchart of an embodiment of the image processing method provided by the present application;

[0055] Figure 2 is Figure 1 a specific schematic flowchart of step S14 of the image processing method shown;

[0056] Figure 3 is a schematic structural diagram of an embodiment of the image processing apparatus provided by the present application;

[0057] Figure 4 is a schematic structural diagram of an embodiment of the computer-readable storage medium provided by the present application. Detailed Embodiments

[0058] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0059] To solve the problems of the prior art, the present application proposes a face skin smoothing processing method based on deep learning and adaptive wavelet transform. This method first uses a deep learning model to extract local features of an image, and then adaptively adjusts the feature extraction parameters according to the image content. Next, an attention mechanism is introduced to assign weights to different features, and a more accurate feature extraction is achieved through a feature fusion strategy optimization. In addition, this method also uses adaptive wavelet transform technology to select an appropriate wavelet function according to the local features of the image, so as to achieve a more refined frequency domain decomposition. Finally, the low-frequency components are smoothed, the high-frequency components are retained to maintain the image details, and the processed frequency domain information is recombined to obtain the skin-smoothed image.

[0060] Specifically, please refer to Figure 1 , Figure 1 which is a schematic flowchart of an embodiment of the image processing method provided by the present application.

[0061] Among them, the image processing method of the present application is applied to an image processing device. Among them, the image processing device of the present application can be a server, or a system in which a server and a terminal device cooperate with each other. Correspondingly, each part included in the image processing device, such as each unit, subunit, module, and submodule, can be all set in the server, or can be respectively set in the server and the terminal device.

[0062] Furthermore, the above server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server. When the server is software, it can be implemented as multiple software or software modules, such as software or software modules used to provide a distributed server, or can be implemented as a single software or software module, which is not specifically limited here. In some possible implementation manners, the image processing method of the embodiments of the present application can be implemented by a processor calling computer-readable instructions stored in a memory.

[0063] Specifically, as Figure 1 shown, the image processing method of the embodiments of the present application specifically includes the following steps:

[0064] Step S11: Obtain an image to be processed.

[0065] The image to be processed in this application is an image including a target face, which can be a video frame intercepted from a playing video, a real-time video frame in a live video, or a monitoring image in a monitoring device.

[0066] Furthermore, after obtaining the original image to be processed, the image processing device can also perform image preprocessing before processing. The specific content is as follows:

[0067] The image preprocessing in this application includes, but is not limited to: operations such as image scaling, image normalization, and operations to reduce the impact of noise, such as bilateral filtering, non-local means filtering, etc. These operations ensure that the input images have the same size and similar characteristics, facilitating subsequent processing.

[0068] Specifically, the image scaling in this application refers to adjusting the image size to a predetermined dimension. This is to ensure that all input images have the same size, thus facilitating subsequent processing. The scaling operation is usually implemented using interpolation methods, such as bilinear interpolation, bicubic interpolation, etc. Among these methods, bilinear interpolation is the most commonly used because it can achieve a good balance between computational efficiency and image quality.

[0069] The normalization operation in this application adjusts the pixel values in the image to a standard range, such as [0, 1] or [-1, 1]. This is to eliminate the scale differences in the image data, so that images from different sources have similar characteristics during subsequent processing. Normalization can be achieved by subtracting the minimum value from each pixel value in the image and dividing by the difference between the maximum value and the minimum value.

[0070] During the face skin smoothing process, the noise in the image will affect the skin smoothing effect. Therefore, this application can also perform noise reduction processing on the image to be processed. The noise reduction methods include, but are not limited to: bilateral filtering, non-local means filtering, etc.

[0071] Bilateral filtering is a non-linear filtering method that combines information in the spatial domain and the pixel value domain to smooth the image. Bilateral filtering can preserve image edges and detail information while eliminating noise. This method can be implemented using image processing libraries (such as OpenCV, scikit-image, etc.).

[0072] Non-local means filtering is a noise reduction method based on the similarity between pixels. It replaces the value of each pixel with the weighted average of similar pixels in its neighborhood. This method has good denoising performance, especially in terms of preserving image structure and detail information. Non-local means filtering can be implemented using image processing libraries (such as OpenCV, scikit-image, etc.).

[0073] The image processing device can also detect and align the faces in the image to be processed, and locate and correct the face images through face detection and alignment algorithms. This can ensure that the positions and angles of the faces in the image are consistent, providing a good foundation for subsequent processing.

[0074] The face detection and alignment method provided in this application, that is, locating and correcting the face image through face detection and alignment algorithms for subsequent processing.

[0075] Specifically, face detection is the process of locating the face area in the image. This can be achieved by using deep learning models such as MTCNN, SSD, etc., or traditional face detection algorithms such as Haar cascade classifiers. These methods can identify the faces in the image and return the coordinates of their bounding boxes.

[0076] Face key point detection: After detecting the face area, it is necessary to detect the face key points (such as eyes, nose, mouth, etc.) for face alignment. Face key point detection can be achieved by deep learning models such as Dlib's 68-point detector, Face Alignment Network, etc., or traditional methods such as Active Shape Models, Active Appearance Models, etc.

[0077] Face alignment is to perform operations such as rotation, scaling, and translation on the detected face area to make the face have a standard pose and size. This can ensure that the subsequent skin smoothing process is more stable and efficient. The implementation process of face alignment is as follows:

[0078] i. Rotation: According to the detected key points, such as the eye positions, calculate the tilt angle of the face. Then rotate the image by the corresponding angle to make the face horizontal.

[0079] ii. Scaling: Based on the key points, such as the distance between the eyes and the mouth, calculate the size of the face. Then scale the face area to a predetermined size for subsequent processing.

[0080] iii. Translation: Translate the face area to the center position of the image to ensure that the face is located at a suitable position in the image.

[0081] iv. Cropping: After alignment, crop the face area for subsequent skin smoothing processing. The cropping operation can be achieved according to the key points, such as eyes, nose, mouth, etc., and the predetermined bounding box size. This can ensure that the cropped image only contains the face area, excluding the background and other irrelevant information.

[0082] After completing the above preprocessing and face detection and alignment operations, the processed face image can be fed into the skin smoothing algorithm, that is, the subsequent image processing steps are completed, such as skin smoothing based on frequency domain analysis, skin smoothing based on image filtering, etc. These algorithms will perform beauty processing on the face image, reduce the blemishes and wrinkles on the skin surface, and improve the beauty of the face image.

[0083] Step S12: Extract the local features of the image to be processed.

[0084] In the embodiment of the present application, the image processing device extracts local features by methods such as edge detection and corner detection. It should be noted that the image processing device can use both edge detection and corner detection to extract the local features of the image to be processed, or only use edge detection, or corner detection to extract the local features of the image to be processed, and no specific limitation is made here. The local features of the image to be processed help to describe the detailed information of the face, such as eyes, nose and mouth, etc.

[0085] Specifically, the image processing device applies an edge detection operator, such as Sobel or Canny, to process the image to obtain an edge intensity map. The edge intensity map reflects the information of the object boundaries in the image and helps to extract the contours and details of the face.

[0086] The image processing device applies a corner detection algorithm, such as Harris or FAST, to process the image to obtain a corner response map. The corner response map reflects the distribution of local features in the image and helps to extract the key points of the face.

[0087] The image processing device fuses the edge intensity map and the corner response map to form a comprehensive feature map. Feature fusion can adopt simple addition, weighted summation or more complex methods, such as attention mechanism, etc. The comprehensive feature map contains the contour, detail and key point information of the face, which helps the subsequent feature analysis and application.

[0088] Among them, edge detection is a technique for identifying the boundaries of objects in an image. Edge detection operators identify edges by calculating the gradient or second derivative of the pixel values in the image. Commonly used edge detection operators include:

[0089] Sobel operator: The Sobel operator detects edges by calculating the first-order gradient of the pixel values in the image. The Sobel operator applies convolution kernels in the horizontal and vertical directions respectively, and then adds the gradients in the two directions to obtain the edge intensity map.

[0090] Canny Operator: The Canny operator is a multi-stage edge detection method. First, the image is Gaussian filtered to reduce noise, and then the Sobel operator is used to calculate the gradient magnitude and direction. Next, non-maximum suppression (NMS) is applied to eliminate non-edge pixels, and finally, a double-threshold method is used to separate strong edge pixels from weak edge pixels to form the final edge map.

[0091] Among them, corner detection is a technique used to identify points with obvious local features in an image. Corners are usually located at the intersections of object boundaries and have high gradient changes. Commonly used corner detection algorithms include:

[0092] Harris Corner Detection: The Harris corner detection algorithm detects corners by calculating the local gradient changes of each pixel in the image. First, the Sobel operator is used to calculate the gradient magnitude and direction of the image. Then, the local structure matrix of each pixel is calculated, and pixels with significant changes are found through eigenvalue analysis. Finally, corners are selected through non-maximum suppression (NMS) and thresholding operations.

[0093] FAST Corner Detection: FAST (Features from Accelerated Segment Test) is a fast corner detection method based on the relative relationship of pixels. The FAST algorithm first binarizes the image, and then sets a circular region around the pixel to determine whether the pixel is a corner by comparing the pixel values in the region with the central pixel value. The FAST algorithm has high computational speed and real-time performance.

[0094] Furthermore, the image processing device can also perform local feature analysis, that is, after extracting local features, these features need to be analyzed to provide a basis for subsequent skin smoothing processing. Local feature analysis includes methods such as feature matching, feature clustering, and feature classification, aiming to find pixel regions with similar features and the relationships between these regions. Specifically, it includes the following processes:

[0095] Feature Matching: By calculating the similarity between features, the most matching other features are found for each local feature. Feature matching methods include Euclidean distance, cosine similarity, etc. Feature matching helps to identify local regions with similar properties and provides a reference for subsequent skin smoothing processing.

[0096] Feature Clustering: Cluster analysis is performed on local features, and pixels with similar features are divided into the same category. Commonly used feature clustering algorithms include K-means, spectral clustering, etc. Feature clustering helps to further refine the similarities and differences of local features and improve the accuracy of skin smoothing processing.

[0097] Feature Classification: By training supervised learning models such as support vector machines and neural networks, local features are classified into different categories, such as eyes, nose, mouth, etc. Feature classification can be trained using a pre-annotated face dataset to obtain a classification model with high recognition ability for local features. Feature classification helps to accurately identify each part of the face and provides a more accurate reference for subsequent skin smoothing processing.

[0098] Application of Local Features: After local feature extraction and analysis are completed, these features can be applied to subsequent operations such as skin smoothing processing. The specific application methods can be selected and designed according to requirements and data characteristics. For example:

[0099] Skin Smoothing Processing Based on Local Features: According to the extracted local features, local skin smoothing processing is performed on the face image. This can be achieved by using the local features as a mask to restrict the area where the skin smoothing filter is applied. In this way, the skin smoothing processing can act more precisely on specific local areas and avoid affecting the details of other parts.

[0100] Enhancement of Beauty Effect Based on Local Features: According to the extracted local features, the beauty effect of the face image is enhanced. For example, according to the features of key parts such as eyes, nose, and mouth, local brightness, contrast, and color adjustments can be made to these areas to improve the overall beauty effect.

[0101] Addition of Special Effects Based on Local Features: According to the extracted local features, special effects are added to the face image. For example, a flash effect can be added around the eyes, or a lipstick effect can be added to the mouth, etc. These special effects can be accurately added according to the position and shape of the local features, making the final effect more realistic and natural.

[0102] In summary, the image processing device extracts local features through methods such as edge detection and corner detection, which helps to describe the detailed information of the face, such as eyes, nose, and mouth. Then, through steps such as preprocessing, local feature extraction, local feature analysis, and application, these features can be used for operations such as skin smoothing processing, enhancement of beauty effect, and addition of special effects, improving the quality and beauty of the face image.

[0103] Step S13: Determine the optimal wavelet function based on local features.

[0104] In the embodiment of the present application, the image processing device realizes adaptive wavelet transform for local features, that is, dynamically adjusts the parameters and functions of wavelet transform according to the local characteristics and texture information of the image, so as to better extract and analyze image features. The specific content includes but is not limited to: adaptive parameter adjustment, dynamic adjustment of wavelet functions, selection of the best wavelet function, and regional adaptive wavelet transform.

[0105] Among them, adaptive parameter adjustment means adjusting the parameters of wavelet transform according to the local characteristics of the image to better extract texture features. Dynamically adjusting the wavelet function means selecting an appropriate wavelet function according to the texture characteristics of the image region. Optimal wavelet function selection means comparing the performance of different wavelet functions and selecting the best wavelet function for transformation. Region adaptive wavelet transform means using different wavelet functions for transformation for different regions of the image to better capture the texture information within the region.

[0106] Specifically, the adaptive wavelet transform of the present application specifically includes the following steps:

[0107] Local characteristic analysis: Analyze the local characteristics of the image, such as texture, edges, corners, etc. These characteristics help to determine appropriate wavelet transform parameters and functions. Local characteristic analysis can be achieved by methods such as sliding windows and image block partitioning.

[0108] Adaptive parameter adjustment: Adjust the parameters of wavelet transform according to the local characteristics of the image, such as scale, direction, etc. For example, for regions with rich texture information, smaller scales and more directions can be selected to better capture texture features. Adaptive parameter adjustment can be achieved by optimization methods such as genetic algorithms and simulated annealing.

[0109] Dynamically adjusting the wavelet function: Select an appropriate wavelet function according to the texture characteristics of the image region. Different wavelet functions have different characteristics, such as orthogonality, compact support, symmetry, etc. By dynamically adjusting the wavelet function, image features can be better extracted according to local texture characteristics. Dynamically adjusting the wavelet function can be achieved by learning methods (such as deep learning, support vector machines, etc.) or empirical rules.

[0110] Optimal wavelet function selection: Compare the performance of different wavelet functions and select the best wavelet function for transformation. This can be achieved by evaluation metrics (such as reconstruction error, signal-to-noise ratio, etc.) or classification accuracy on the training set and other methods.

[0111] Region adaptive wavelet transform: Use different wavelet functions for transformation for different regions of the image to better capture the texture information within the region. This can be achieved by methods such as image block partitioning and sliding windows. In each region, apply the wavelet function that has been adaptively parameter-adjusted and dynamically adjusted for transformation, and summarize the transformation results into a comprehensive feature representation.

[0112] Feature extraction and fusion: Extract features from the results of region-adaptive wavelet transform and fuse these features with other types of features, such as edges, corners, etc. Feature fusion methods can be simple feature vector concatenation, weighted summation, or using more complex fusion strategies, such as using an attention mechanism to adaptively adjust the weights between features. In addition, a multi-scale fusion strategy can be utilized to fuse wavelet features at different scales with other features to capture multi-scale information.

[0113] Feature dimensionality reduction: Since adaptive wavelet transform may generate a large number of features, it is necessary to perform dimensionality reduction on the features to reduce computational complexity and avoid overfitting. Commonly used dimensionality reduction methods include principal component analysis (PCA), linear discriminant analysis (LDA), etc. These methods can project the original high-dimensional features into a low-dimensional space while maintaining the main structure and information of the data.

[0114] Classification and regression: Input the dimensionality-reduced features into a classifier or regressor for processing the target task. For example, for an object detection task, the features can be input into a classifier such as a support vector machine (SVM) or a random forest (RF) for identifying the target region. For a regression task, such as pixel prediction, the features can be input into a regressor such as a neural network, support vector regression (SVR), etc. for predicting pixel values.

[0115] The adaptive wavelet transform of this application can dynamically adjust the parameters and functions of the wavelet transform according to the local characteristics and texture information of the image to better extract and analyze image features. Through the above implementation steps, an image processing system based on adaptive wavelet transform can be constructed to effectively process various complex image tasks, such as object detection, texture analysis, etc.

[0116] Next, the specific implementation processes of each part of the adaptive wavelet transform will be introduced separately:

[0117] Calculate the similarity of wavelet functions: To select an appropriate wavelet function, it is necessary to calculate the similarity between different wavelet functions and the image feature space. The following methods can be used to calculate the similarity:

[0118] For each wavelet function, the image processing device calculates its representation in the feature space. This can be achieved by applying the wavelet function to each pixel in the feature space and calculating its response. Then, calculate the distance between the wavelet function representation and the feature space. Metric methods such as Euclidean distance and cosine similarity can be used.

[0119] The following are the specific calculation processes and algorithms:

[0120] 1. Prepare the wavelet function library: First, create a wavelet function library that contains various types of wavelet functions, such as Haar, Daubechies, Symlet, etc. These wavelet functions can be one-dimensional or two-dimensional, depending on the dimensionality of the feature space of the image.

[0121] 2. Calculate the wavelet response in the feature space: For each wavelet function, apply it to each pixel in the feature space. During this process, convolution operations can be used to convolve the wavelet function with the feature space to calculate the wavelet response at each pixel position. The resulting response matrix represents the response of each pixel in the feature space under the current wavelet function.

[0122] 3. Calculate the similarity metric: For each wavelet function, calculate the distance between its representation and the feature space. Multiple similarity metric methods can be used, such as Euclidean distance, cosine similarity, Pearson correlation coefficient, etc. Store these distance or similarity values in a matrix or list for subsequent selection of the best wavelet function.

[0123] It should be noted that the data input for calculating the wavelet function similarity comes from the previous feature selection step. The result after feature selection will be used as the input for calculating the similarity between different wavelet functions and the image. Output of the calculation of wavelet function similarity: The calculated similarity matrix or list is used as the output and passed to the next step, namely, selecting the best wavelet function. At this stage, this application will use the similarity values to determine the wavelet function that is most similar to the image feature space.

[0124] Specifically, the input of wavelet similarity calculation is the output of local feature extraction, and its output will be used to dynamically adjust the wavelet function.

[0125] During this process, the local feature extraction step converts the original image into a set of features that describe the local information of the image, such as edges, corners, etc. These features can be regarded as a high-level representation of the original image, and they can more effectively capture the structural and texture information of the image.

[0126] These features are used as the input for calculating the wavelet function similarity. Specifically, for each wavelet function, apply it to these features and calculate the response of each feature under the wavelet function. These responses can be regarded as the representation of the features under the wavelet function, and they capture the similarity between the features and the wavelet function.

[0127] Furthermore, the image processing device calculates the distance or similarity between each wavelet function representation and the features. These distance or similarity values represent the matching degree of each wavelet function with the features. Therefore, the output of this step is a list of similarity values equal to the number of wavelet functions.

[0128] This list is used as the input for dynamically adjusting the wavelet function. Specifically, the wavelet function that best matches the features is selected according to this list. This wavelet function will be used in subsequent wavelet transform steps to extract finer texture features.

[0129] In summary, calculating the similarity of wavelet functions in this application can effectively connect the local feature extraction step and the dynamic wavelet function adjustment step. Through this step, the most suitable wavelet function for the current features can be selected, thereby improving the effect and accuracy of wavelet transform. The whole process starts from the input features, calculates the responses of each pixel under different wavelet functions, and then calculates the distance or similarity between different wavelet function representations and the feature space. Finally, a similarity matrix or list is output for subsequent selection of the best wavelet function. Throughout the process, the data flow starts from feature input, goes through wavelet response calculation and similarity measurement, and finally outputs a similarity matrix or list.

[0130] Select the best wavelet function: According to the calculated similarity, select the wavelet function that is most similar to the image feature space as the best wavelet function. Optimization algorithms such as greedy search and simulated annealing can be used to implement the search for the best wavelet function.

[0131] Dynamically adjust the wavelet function: In practical applications, the wavelet function can be dynamically adjusted according to the changes in the image. For example, when dealing with the task of face skin retouching, different wavelet functions can be selected for different regions according to the feature differences between the skin region and the non-skin region. Specific implementation steps:

[0132] 1. Perform skin detection on the input image and divide the image into skin regions and non-skin regions.

[0133] 2. Calculate the local features of the skin region and the non-skin region respectively.

[0134] 3. According to the local features, select the best wavelet function for the skin region and the non-skin region respectively. The similarity between the feature vectors of each pixel in the two regions and different wavelet functions can be calculated respectively, and then the wavelet function with the highest similarity can be selected.

[0135] Dynamically adjusting the wavelet function enables the selection of appropriate wavelet functions according to the characteristics of different regions. The following is the detailed process and algorithm:

[0136] 1. Skin detection: Use color models such as HSV, YCbCr, etc. to convert the input image into a specific color space. Set appropriate thresholds and perform thresholding on the color space to divide the image into skin regions and non-skin regions. Morphological operations such as erosion and dilation can be used to further optimize the skin detection results.

[0137] 2. Local feature calculation: Calculate local features for the skin region and non-skin region respectively. Local features can include information such as texture, edges, color, etc. Feature extraction methods such as SIFT, SURF, HOG, etc. can be used, or the deep learning model mentioned above can be used to extract local features.

[0138] 3. Optimal wavelet function selection: According to the results of calculating the wavelet function similarity previously, select the optimal wavelet function for the skin region and non-skin region respectively. The similarity between the feature vectors of each pixel in the two regions and different wavelet functions can be calculated respectively, and then the wavelet function with the highest similarity is selected.

[0139] It should be noted that the input data for this stage comes from the following parts: (1) The input image, used for skin detection; (2) The wavelet function similarity matrix or list calculated previously; (3) The local features of the skin region and non-skin region. Data output: The optimal wavelet functions selected for the skin region and non-skin region respectively are used as the output and passed to the next step, namely regional adaptive wavelet transform.

[0140] The data flow of dynamically adjusting the wavelet function in this application is specifically as follows:

[0141] 1. The input image is first used for skin detection to obtain the division of the skin region and non-skin region.

[0142] 2. Using the division results of the skin region and non-skin region, calculate the local features of the two regions respectively.

[0143] 3. According to the calculated local features, referring to the wavelet function similarity matrix or list calculated previously, select the optimal wavelet function for the two regions respectively.

[0144] 4. The selected optimal wavelet function is used as the output and passed to the next step, namely regional adaptive wavelet transform.

[0145] Specifically, the process of dynamically adjusting the wavelet function includes skin detection, local feature calculation, and optimal wavelet function selection. First, perform skin detection on the input image to divide the image into the skin region and non-skin region. Then, calculate the local features of these two regions respectively. Finally, according to the local features and the wavelet function similarity matrix or list calculated previously, select the optimal wavelet function for the skin region and non-skin region respectively. The selected optimal wavelet function will be used as the output and passed to the next step, namely regional adaptive wavelet transform.

[0146] Relationship between dynamically adjusting wavelet functions and the previous step: When performing local feature extraction, this application obtains the feature information of the input image in the skin area and non-skin area, including texture, edges, color, etc. This local feature information will be used in the step of dynamically adjusting wavelet functions to help this application select the most suitable wavelet function for each area. Therefore, the result of local feature extraction will serve as an important input for the step of dynamically adjusting wavelet functions.

[0147] Relationship between dynamically adjusting wavelet functions and the next step: In the step of dynamically adjusting wavelet functions, this application will select the most suitable wavelet functions for the skin area and non-skin area respectively and perform wavelet transform. The result of the wavelet transform, that is, the image processed by the wavelet function, will be used as the input for the high-level feature extraction step. In the high-level feature extraction step, this application will use methods such as deep learning models to further extract the high-level features of the image, and these high-level features can better reflect the content and structural information of the image. Therefore, the result of dynamically adjusting wavelet functions will directly affect the effect of high-level feature extraction.

[0148] Generally speaking, dynamically adjusting wavelet functions is a key link connecting the two steps of local feature extraction and high-level feature extraction. It selects the most suitable wavelet function according to local features, processes the image, and then passes the processed image to the high-level feature extraction step, providing a better input for subsequent image processing tasks.

[0149] Region-adaptive wavelet transform: Apply the selected optimal wavelet function to the skin area and non-skin area of the input image respectively for wavelet transform. This can select appropriate wavelet functions according to the characteristics of different regions, thereby achieving a more refined frequency-domain decomposition.

[0150] The following are the detailed processes and algorithms, as well as the flow direction of the data stream:

[0151] 1. Input data: The input data includes the original image, skin detection results, that is, the skin area and non-skin area, and the optimal wavelet function selected in the previous step. These data can be obtained through the output results of the previous stages, or imported from external data sources as needed.

[0152] 2. Region division: Divide the original image into the skin area and non-skin area according to the skin detection results. This can be achieved through image segmentation algorithms, such as thresholding method, region growing method, etc. The segmented regions can be represented by a binary mask, where the positions with pixel value 1 correspond to the skin area, and the positions with pixel value 0 correspond to the non-skin area.

[0153] 3. Regional wavelet transform: Apply the selected optimal wavelet function to the skin region and non-skin region respectively for wavelet transform. First, create an empty image with the same size as the original image for each region. Then, copy the skin region and non-skin region of the original image to these two empty images according to the binary mask. Next, perform wavelet transform on these two images. Existing wavelet transform libraries such as the PyWavelets library in Python can be used for calculation.

[0154] 4. Frequency domain decomposition: Wavelet transform decomposes an image into multiple frequency sub-bands, including the low-frequency component, i.e., the smooth part, and the high-frequency component, i.e., the detail part. Generally, the low-frequency component is used to represent the general structure of the image, while the high-frequency component is used to represent the detail information of the image. Frequency domain decomposition can help this application distinguish skin texture and blemishes, thus achieving more refined skin retouching.

[0155] 5. Data output: Output the decomposed low-frequency component and high-frequency component for use in the next step, i.e., face skin retouching. The output data includes the low-frequency component, high-frequency component, and binary mask of the skin region and non-skin region.

[0156] 6. In the face skin retouching stage, the image processing device needs to operate according to the output result of the regional adaptive wavelet transform. Specifically, the image processing device will smooth the low-frequency component of the skin region to eliminate skin blemishes and fine lines; at the same time, retain the high-frequency component to maintain the detail information of the image. The non-skin region does not need to be retouched and can retain the original low-frequency and high-frequency components.

[0157] Step S14: Extract the wavelet features of the image to be processed according to the optimal wavelet function.

[0158] In the embodiment of this application, after the image processing device determines the optimal wavelet function and its wavelet parameters through the adaptive wavelet transform in step S13, it extracts the wavelet features of the image to be processed according to the optimal wavelet function.

[0159] Furthermore, the image processing device can also perform texture enhancement prediction and texture enhancement of the image according to the local features and wavelet features of the image to be processed.

[0160] For details, please refer to Figure 2 , Figure 2 is Figure 1 the specific process schematic diagram of step S14 of the image processing method shown.

[0161] Specifically, as Figure 2 shown, the image processing method of the embodiment of this application specifically includes the following steps:

[0162] Step S141: Input the local features and wavelet features into a pre-trained deep learning model.

[0163] In the embodiment of the present application, the image processing device constructs a deep learning model, which can specifically be composed of a convolutional neural network, to predict the enhancement degree of each pixel point. The input of the model is the local features and wavelet features of the image, and the output is the enhancement degree of each pixel point.

[0164] Step S142: Obtain the predicted enhancement degree of each pixel point output by the deep learning model.

[0165] Step S143: Use the predicted enhancement degree of each pixel point to increase the pixel value of each pixel point by weighting, and obtain the texture-enhanced pixel value of each pixel point.

[0166] In the embodiment of the present application, the image processing device performs texture enhancement on the original image according to the enhancement degree of each pixel point predicted by the deep learning model. Specifically, the predicted enhancement degree is used as the weight to perform weighted adjustment on the pixel values of the original image. For example, for pixel points with a higher predicted enhancement degree, increase their pixel values to make their textures more prominent; while for pixel points with a lower predicted enhancement degree, keep their pixel values unchanged.

[0167] Furthermore, after texture enhancement, the image processing device needs to perform some post-processing on the enhanced image to eliminate possible noise or color blocks and improve the visual effect of the image. The post-processing can include operations such as filtering, denoising, and sharpening.

[0168] It should be noted that during this process, it is necessary to pay attention to maintaining the integrity and coordination of the image. Specifically, one cannot only focus on individual pixel points while ignoring the surrounding pixel points. If the texture of a single pixel point is over-enhanced, it may cause obvious noise or color blocks in the image, destroying the overall effect of the image. Therefore, while enhancing the texture, it is necessary to ensure the overall coordination of the image.

[0169] Step S15: Fuse the local features and wavelet features to obtain fused features.

[0170] In the embodiment of the present application, the image processing device fuses the local features and wavelet features to obtain fused features, which are used to input into the model in subsequent steps for skin smoothing processing of the image face.

[0171] Furthermore, the image processing device can further extract high-level features from the wavelet features, which are helpful for describing more complex and abstract information of the human face. The image processing device uses convolutional layers, such as standard convolutional layers, deformable convolutional layers, adaptive local pattern convolutional layers, etc., to extract high-level features from the data after adaptive wavelet transform.

[0172] Specifically, the image processing device inputs the data after adaptive wavelet transform into a deep convolutional neural network, which can include multiple convolutional layers, activation functions, and pooling layers. The convolutional layer is used to extract high-level features from the data after wavelet transform, and the activation function can introduce non-linear transformation into the network. The pooling layer helps to reduce the dimension of the data, thereby reducing the computational complexity.

[0173] Under the action of multiple convolutional layers, the network will gradually extract more and more abstract features from the input data, which are helpful for describing the complex information of the human face. In this process, the image processing device can use various types of convolutional layers, such as standard convolutional layers, deformable convolutional layers, and adaptive local pattern convolutional layers, etc., to enhance the expression ability of the network.

[0174] After the process of high-level feature extraction is completed, the image processing device can fuse these features with the local features extracted previously, and then continue with subsequent processing, such as global average pooling, fully connected layers, etc. Through these steps, the image processing device can obtain a model that can effectively perform face skin smoothing processing.

[0175] The feature fusion of this application integrates local features, wavelet features, and high-level features. Feature fusion can use methods such as simple feature concatenation, weighted summation, etc., so that the model can utilize various types of features simultaneously, thereby improving the model performance.

[0176] Specifically, feature fusion combines features of different levels and sources to form a comprehensive feature representation. In face skin smoothing processing, feature fusion can include combining local features, such as edges, textures, etc., with global features, such as facial contours, face key points, etc. Feature fusion methods can be simple feature vector concatenation, weighted summation, or using more complex fusion strategies, such as using an attention mechanism to adaptively adjust the weights between features.

[0177] In addition, the image processing device can also utilize a multi-scale fusion strategy to fuse features of different scales with other features to capture multi-scale information. Specifically, the image processing device can extract multi-scale local features of the local features by using upsampling or downsampling, fuse the multi-scale local features with the wavelet features to obtain the fused features; or extract multi-scale wavelet features of the wavelet features by using upsampling or downsampling, and fuse the multi-scale wavelet features with the local features to obtain the fused features.

[0178] After feature fusion, the image processing device obtains a feature representation containing rich information, providing strong support for subsequent skin smoothing processing.

[0179] Furthermore, in order to reduce the dimension of the features for subsequent processing, the image processing device can use global average pooling (GAP) to reduce the dimension of the fused features. Based on the features with reduced dimensions, the image processing device can also use a fully connected layer to further process and non-linearly transform the features to enhance the expressive ability of the model.

[0180] Specifically, global average pooling (GAP) is a method for feature dimension reduction and spatial compression. It calculates the average value of all elements in each feature map to obtain a feature vector. In face skin smoothing processing, global average pooling can compress the feature representation into a low-dimensional feature vector, which helps reduce the computational amount and memory occupancy while retaining key information. The implementation of GAP can use the corresponding functions or modules provided by deep learning frameworks such as TensorFlow and PyTorch.

[0181] The fully connected layer (FC) is a basic structure in a neural network that establishes dense connections between input features and output features. In face skin smoothing processing, the fully connected layer can be used to map the feature vector after global average pooling to a new feature space for subsequent processing. The implementation of the fully connected layer can use the corresponding functions or modules provided by deep learning frameworks (such as TensorFlow and PyTorch). The fully connected layer is usually combined with activation functions (such as ReLU and sigmoid) to introduce non-linear characteristics. The parameters (such as weights and biases) of the fully connected layer can be optimized through the gradient descent algorithm during the training process.

[0182] Step S16: Generate the skin smoothing degree of each pixel point of the image to be processed by using the fused features.

[0183] In the embodiments of the present application, the image processing device can design a corresponding output layer according to the requirements of specific tasks. For the skin smoothing processing task, the output layer can be a fully connected layer with a single neuron, and its activation function can be selected according to the value range of the task, such as a linear activation function, a sigmoid function, etc.

[0184] Specifically, 1. The output layer is the last layer of the neural network and is responsible for generating the final skin smoothing result. In face skin smoothing processing, the output layer can be a regression layer or a convolutional layer, which is used to predict the skin smoothing degree of each pixel point based on the previous feature representation. The following is the detailed implementation process of the output layer:

[0185] Regression layer: The regression layer can map the output of the fully connected layer to a continuous value range (such as [0,1]), representing the skin smoothing degree of each pixel point. The implementation of the regression layer can use the corresponding functions or modules provided by deep learning frameworks (such as TensorFlow, PyTorch, etc.). Usually, the activation function of the regression layer is sigmoid or tanh to ensure that the output value is within the specified range.

[0186] Convolutional layer: The convolutional layer can map the local information in the feature representation to an output image with the same size as the input image, representing the skin smoothing degree of each pixel point. The implementation of the convolutional layer can use the corresponding functions or modules provided by deep learning frameworks (such as TensorFlow, PyTorch, etc.). Usually, the activation function of the convolutional layer is sigmoid or tanh to ensure that the output value is within the specified range.

[0187] Step S17: Process the image to be processed according to the skin smoothing degree of each pixel point to obtain a face skin smoothing image.

[0188] In the embodiments of the present application, the image processing device uses the skin smoothing result output by the model, combines the texture features and color information in the original image, and performs skin smoothing processing on the face. Pixel-level fusion techniques, such as weight fusion, Poisson fusion, etc., can be used to generate the skin-smoothed face image.

[0189] Specifically, the image processing device performs skin smoothing processing on the input image according to the prediction result of the output layer. The specific method is to apply the predicted skin smoothing degree to each pixel point of the input image, thereby obtaining the skin-smoothed image. The implementation of generating the skin smoothing result can be performed using an image processing library, such as OpenCV, scikit-image, etc.

[0190] Furthermore, the image processing device performs an inverse preprocessing operation on the skin-smoothed face image, such as restoring the image size to the original size, restoring the original pixel value range, etc. This step ensures that the skin-smoothed image has the same size and pixel value range as the original image.

[0191] The image processing device fuses the retouched face image with the original image to generate the final retouched effect image. Image fusion algorithms such as alpha blending, Poisson fusion, etc. can be used to ensure that the retouched face blends naturally with the original background. Finally, the image processing device outputs the synthesized retouched effect image as the final output.

[0192] Through the above steps, the image processing device can apply the retouching result of the model to the actual face image to generate a retouched effect image with natural and smooth skin texture.

[0193] Furthermore, in order to improve the retouching effect, the image processing device also needs to perform some post-processing operations, such as sharpening and contrast enhancement on the retouched image. The post-processing operations can be selected and implemented according to actual needs and effect requirements.

[0194] In the embodiments of this application, steps such as feature fusion, global average pooling, fully connected layer, and output layer in the face retouching method are responsible for extracting key features from the input image and generating the retouching result. Through feature fusion, features from different sources are combined together; global average pooling compresses the feature representation into a low-dimensional feature vector; the fully connected layer maps the feature vector to a new feature space; the output layer predicts the retouching degree of each pixel point. These steps together constitute an efficient and accurate face retouching process.

[0195] In a specific implementation, when the data of the output layer is generated, that is, when the image processing device obtains a prediction image with the same size as the input image, each pixel value on this prediction image represents the retouching degree of the corresponding pixel on the original image.

[0196] The following steps are to apply this prediction result to the original face image to generate the retouched image:

[0197] Face image retouching: The image processing device achieves this by applying the predicted retouching degree to each pixel point of the original image. A common method is to use the predicted retouching degree as a weight to perform weighted averaging on the original image to smooth the image. For example, if the predicted retouching degree is 0.8, then the new pixel value is 0.8 times the original pixel value plus 0.2 times the neighboring pixel value. In this way, areas with a high retouching degree, such as skin wrinkles and blemishes, will be smoothed, while areas with a low retouching degree, such as the edges of eyes and lips, will remain clear.

[0198] Inverse preprocessing: After the skin smoothing process is completed, the image processing device performs inverse preprocessing operations on the image to restore the image size to its original size and recover the original pixel value range, etc. This step is necessary because preprocessing operations such as scaling and normalization were performed on the image before inputting it into the neural network, and now these operations need to be reversed to ensure that the skin-smoothed image has the same size and pixel value range as the original image.

[0199] Synthesize the final result: The image processing device fuses the skin-smoothed face image with the original image to generate the final skin smoothing effect image. This step usually involves image fusion algorithms such as transparency fusion, Poisson fusion, etc. Specifically, the face region in the original image can be extracted first, then the skin-smoothed face image is fused into this region, and finally the fused face region replaces the face region in the original image to obtain the final skin smoothing effect image.

[0200] Through the above steps, the prediction results of the neural network model can be used to perform skin smoothing processing on the face image and generate a skin smoothing effect image.

[0201] It should be noted that the actual implementation may vary depending on the specific requirements of the application and the characteristics of the data. For example, some special steps such as color balance and brightness adjustment may need to be added to the skin smoothing process to improve the naturalness and realism of the skin smoothing effect. Different neural network structures and training strategies may also need to be used to adapt to different skin smoothing tasks and performance requirements.

[0202] Generally speaking, a good skin smoothing effect for faces requires a balance between smoothing the skin texture and maintaining image details. Excessive skin smoothing may cause the image to lose its naturalness, while insufficient skin smoothing may not effectively remove skin blemishes and wrinkles. Therefore, choosing an appropriate skin smoothing degree prediction model, optimizing the image processing strategy, and performing fine post-processing are all the keys to achieving a high-quality skin smoothing effect for faces.

[0203] At the same time, this application can also take into account the limitations of computing resources and processing speed. For example, for real-time or near-real-time face skin smoothing applications such as video chats and live broadcasts, this application needs to select methods and technologies that can complete the processing within a limited time. In this case, the neural network model may need to be simplified or optimized to improve its running speed. Additionally, the processing speed can also be increased by using more efficient image processing and computing technologies such as GPU acceleration and parallel computing.

[0204] The image processing method of this application can achieve the following effects:

[0205] Multi-level feature extraction: Through deep learning and wavelet transform, feature extraction is performed on face images from multiple levels. This includes the detailed features extracted from wavelet transform and the high-level semantic features extracted through deep networks. This multi-level feature extraction strategy can more comprehensively express the texture and details of the skin, helping to improve the quality and naturalness of the skin smoothing effect.

[0206] End-to-end processing flow: Integrate all key steps, including wavelet transform, feature extraction, feature fusion, skin smoothing processing, etc., into a unified deep learning framework to form an end-to-end processing flow. This end-to-end design can avoid the complex process of manually designing features and adjusting parameters, and can also directly learn effective feature representations from the original face images, significantly improving the processing efficiency and effect.

[0207] Adaptive processing strategy: The parameters of wavelet transform and skin smoothing processing can be dynamically adjusted according to the specific situation of the face image. For example, the scale and direction of wavelet transform can be dynamically adjusted according to the complexity and variability of skin texture; the intensity and style of the skin smoothing effect can also be adjusted according to the individual differences and characteristics of the face. This adaptive processing strategy can enable the method of this application to obtain excellent skin smoothing effects in various different scenarios and conditions.

[0208] Consider the visual perception and aesthetic habits of the human eye: When designing and adjusting the skin smoothing effect, the visual perception and aesthetic habits of the human eye are fully considered. This application will try to maintain the naturalness of skin texture and avoid the "wax figure" effect of the face caused by excessive skin smoothing; this application will also consider the sensitivity of the human eye to color and brightness and adjust the color and brightness of the skin smoothing effect.

[0209] Powerful scalability and adaptability: Based on open deep learning frameworks and standard wavelet transforms, it can be easily combined with other technologies and methods. For example, this application can introduce new deep learning models to improve the effect of feature extraction and fusion; this application can also combine other image processing technologies (such as HDR, super-resolution, etc.) to further improve the quality and details of the skin smoothing effect. This powerful scalability and adaptability enable the method of this application to adapt to various different application scenarios and requirements, helping to promote the progress and development of face skin smoothing technology.

[0210] Provide a more natural and refined skin smoothing effect: Through the combination of deep learning and adaptive wavelet transform, while retaining the skin texture details, it can effectively smooth the color and brightness of the skin, providing a more natural and refined skin smoothing effect. This meticulous processing method can not only improve the aesthetics of face images, but also avoid the "wax figure" effect of the face caused by excessive skin smoothing.

[0211] Strong adaptability: Regardless of different face types, lighting conditions, skin texture complexities, etc., it can flexibly adjust the parameters of wavelet transform and skin smoothing processing in an adaptive manner to achieve the optimal skin smoothing effect.

[0212] High real-time performance: Due to the adoption of an end-to-end deep learning framework, all processing steps are completed in one forward propagation. Therefore, compared with traditional multi-step processing methods, the processing speed can be greatly improved to meet the requirements of real-time skin smoothing.

[0213] High training and inference efficiency: Based on the adaptive parameter optimization of deep learning, the model training and inference processes are more efficient. Since the model can directly learn skin smoothing parameters from the data without manual setting, the efficiency of model training and inference is greatly improved.

[0214] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order and does not constitute any limitation to the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0215] To implement the image processing method of the above embodiment, the present application also proposes an image processing device. For details, please refer to Figure 3 , Figure 3 which is a schematic structural diagram of an embodiment of the image processing device provided by the present application.

[0216] The image processing device 300 of the embodiment of the present application includes a memory 31 and a processor 32. Among them, the memory 31 and the processor 32 are coupled.

[0217] The memory 31 is used to store program data, and the processor 32 is used to execute the program data to implement the image processing method described in the above embodiment.

[0218] In this embodiment, the processor 32 can also be called a CPU (Central Processing Unit, central processing unit). The processor 32 may be an integrated circuit chip with signal processing capabilities. The processor 32 may also be a general-purpose processor, a digital signal processor (DSP, Digital Signal Process), an application-specific integrated circuit (ASIC, ApplicationSpecific Integrated Circuit), a field-programmable gate array (FPGA, Field Programmable GateArray), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor 32 may also be any conventional processor, etc.

[0219] To implement the image processing method of the above embodiments, the present application further provides a computer-readable storage medium, such as Figure 4 As shown, the computer-readable storage medium 400 is used to store program data 41. When the program data 41 is executed by a processor, it is used to implement the image processing method as described in the above embodiments.

[0220] The present application further provides a computer program product. Among them, the above computer program product includes a computer program, and the above computer program can be operated to cause a computer to execute the image processing method as described in the embodiments of the present application. This computer program product can be a software installation package.

[0221] When the image processing method described in the above embodiments of the present application exists in the form of a software function unit and is sold or used as an independent product, it can be stored in a device, such as a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present invention. And the aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, etc., which can store program codes.

[0222] The above are only the embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. An image processing method, characterized in that, The described image processing method includes: Obtain the image to be processed; Extract the local features of the image to be processed; Determine the optimal wavelet function based on the local features; Extract the wavelet features of the image to be processed according to the optimal wavelet function; Fuse the local features and the wavelet features to obtain fused features; Use the fused features to generate the skin smoothing degree of each pixel point of the image to be processed; Process the image to be processed according to the skin smoothing degree of each pixel point to obtain a face skin-smoothed image.

2. The image processing method according to claim 1, wherein The process of processing the image to be processed according to the skin smoothing degree of each pixel point to obtain a face skin-smoothed image includes: Generate the skin smoothing weight of each pixel point according to the skin smoothing degree of each pixel point, wherein the skin smoothing weight includes an original skin smoothing weight and a neighboring skin smoothing weight; Perform weighted processing on the original pixel value of each pixel point and the original skin smoothing weight to obtain an original skin-smoothed pixel value; Perform weighted processing on the neighboring pixel value of each pixel point and the neighboring skin smoothing weight to obtain a neighboring skin-smoothed pixel value; Add the original skin-smoothed pixel value and the neighboring skin-smoothed pixel value to obtain the face skin-smoothed pixel value of each pixel point; Combine the face skin-smoothed pixel values of all pixel points to generate the face skin-smoothed image.

3. The image processing method according to claim 2, wherein The process of combining the face skin-smoothed pixel values of all pixel points to generate the face skin-smoothed image includes: Combine the face skin-smoothed pixel values of all pixel points to generate a face skin-smoothed weighted image; Fuse the face skin-smoothed weighted image and the image to be processed to generate a final face skin-smoothed image.

4. The image processing method according to claim 1, wherein The process of fusing the local features and the wavelet features to obtain fused features includes: Extract high-level features from the wavelet features using a convolutional layer; Fuse the local features, the wavelet features, and the high-level features to obtain the fused features.

5. The image processing method according to claim 1 or 4, wherein The process of fusing the local features and the wavelet features to obtain fused features includes: Extract multi-scale local features of the local features using upsampling or downsampling, and fuse the multi-scale local features with the wavelet features to obtain the fused features; Or, extract multi-scale wavelet features of the wavelet features using upsampling or downsampling, and fuse the multi-scale wavelet features with the local features to obtain the fused features.

6. The image processing method according to claim 1, wherein After extracting the wavelet features of the image to be processed according to the optimal wavelet function, the image processing method further includes: Input the local features and the wavelet features into a pre-trained deep learning model; Obtain the predicted enhancement degree of each pixel point output by the deep learning model; The pixel value of each pixel point is weighted and increased by using the prediction enhancement degree of each pixel point, so as to obtain the texture enhanced pixel value of each pixel point.

7. The image processing method according to claim 1, wherein the determining the optimal wavelet function based on the local feature includes: obtaining a plurality of candidate wavelet functions according to the local feature; obtaining the wavelet function representation of each pixel point in the local feature under each candidate wavelet function; using the similarity between the wavelet function representation and the local feature as the similarity value of the candidate wavelet function; determining the candidate wavelet function with the highest similarity value as the optimal wavelet function.

8. The image processing method according to claim 1 or 7, wherein the determining the optimal wavelet function based on the local feature includes: performing skin detection on the image to be processed, and dividing the image to be processed into a skin area and a non-skin area; determining a first local feature of the skin area and a second local feature of the non-skin area based on the local feature; determining a first optimal wavelet function based on the first local feature; determining a second optimal wavelet function based on the second local feature.

9. The image processing method according to claim 1 or 7, wherein after determining the optimal wavelet function based on the local feature, the image processing method further includes: obtaining the texture information of the local feature; determining the wavelet parameters of the optimal wavelet function according to the texture information; wherein the wavelet parameters include a wavelet scale and / or a wavelet direction.

10. The image processing method according to claim 1, wherein the extracting the local feature of the image to be processed includes: processing the image to be processed by an edge detection operator to obtain an edge intensity map; processing the image to be processed by a corner detection algorithm to obtain a corner response map; fusing the edge intensity map and the corner response map to generate a comprehensive feature map; extracting the local feature of the image to be processed from the comprehensive feature map.

11. An image processing apparatus, characterized in that, The image processing device includes a processor and a memory, wherein program data is stored in the memory, and the processor is configured to execute the program data to implement the image processing method according to any one of claims 1-10.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used for storing program data, and when the program data is executed by a processor, it is used to implement the image processing method according to any one of claims 1-10.

Citation Information

Patent Citations

  • Dynamic illumination face image quality enhancement method based on multi-scale attention mechanism

    CN115880225A

  • Image processing method and device

    CN116012270A