Method for tea image bit and spatial resolution joint reconstruction based on frequency domain feature decoupling

The method of jointly reconstructing tea images by decoupling frequency domain features and bit and spatial resolution, and using shared and non-shared weight modules for feature extraction and frequency band optimization, solves the problems of insufficient collaborative optimization and low efficiency in existing tea image reconstruction techniques, and achieves efficient image enhancement.

CN121746182BActive Publication Date: 2026-05-01SICHUAN AGRI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SICHUAN AGRI UNIV
Filing Date
2026-02-25
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing tea leaf image reconstruction algorithms suffer from insufficient collaborative optimization, low efficiency, and unsatisfactory reconstruction quality, especially in the case of artifacts and detail loss caused by frequency domain energy coupling conflicts.

Method used

A method for joint reconstruction of tea images based on frequency domain feature decoupling and bit and spatial resolution is adopted. Common features across degradation types are extracted by a shared weight module, and frequency band optimization is performed by a non-shared weight module. Combined with a frequency domain fusion module, adaptive weighted fusion is performed to construct a hybrid network architecture.

Benefits of technology

It effectively solves the problems of missing frequency features and task conflicts, and achieves efficient enhancement of tea images in complex degradation scenarios, improving image quality and computational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121746182B_ABST
    Figure CN121746182B_ABST
Patent Text Reader

Abstract

The present application relates to the field of tea image reconstruction, and particularly relates to a tea image bit and spatial resolution joint reconstruction method based on frequency domain feature decoupling. The scheme comprises the following steps: obtaining an original low-bit low spatial resolution tea image, and performing pretreatment; constructing a tea image reconstruction neural network, the tea image reconstruction neural network comprising a shared weight module, a non-shared weight module and a frequency domain fusion module; training the tea image reconstruction neural network using the pretreated low-bit low spatial resolution tea image; after the training is completed, freezing all parameters of the shared weight module, the non-shared weight module and the frequency domain fusion module, inputting a low spatial resolution low-bit tea image to be reconstructed into the trained neural network, and outputting a final reconstructed image; evaluating the tea image reconstruction quality through a peak signal-to-noise ratio and a structural similarity index, and comprehensively evaluating the reconstruction effect of the tea reconstruction neural network in combination with a subjective visual effect. The present application is suitable for tea image reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of tea image reconstruction, and specifically to a method for joint reconstruction of tea image bits and spatial resolution based on frequency domain feature decoupling. Background Technology

[0002] In the agricultural field, joint enhancement of bit depth and spatial resolution in tea images has significant research and application value. Increasing the bit depth of an image can effectively repair color gradation faults in low bit depth environments, restoring the continuous color transitions and complex texture features of tea leaves; while improving spatial resolution helps to achieve high-precision reconstruction of subtle phenotypes in tea leaves (such as leaf margin serrations and hair distribution), which is crucial for understanding tea growth mechanisms, genetic characteristics, and environmental responses. In terms of pest and disease detection and yield calculation, joint enhancement technology can help detect signs of pests and diseases earlier and more accurately, and precisely estimate plant morphological characteristics.

[0003] However, existing mainstream algorithms suffer from significant architectural flaws. Split-based methods employ a cascaded processing framework, treating super-resolution reconstruction and bit augmentation as independent tasks, sequentially performing spatial detail restoration and bit artifact suppression through a serial network. For example, the ZP-Bic algorithm extends bit resolution to 16 bits with zero-padding and uses bicubic interpolation for spatial resolution reconstruction; the J-DNN algorithm uses a cascaded network architecture, with two SRNet-structured sub-networks handling bit augmentation and spatial super-resolution tasks respectively. While these methods retain some task specificity, they neglect the energy coupling effect of different degradation types in the frequency domain. The sharpening of high-frequency textures by the super-resolution network exacerbates the visibility of low-frequency bit artifacts, while the smoothing operation of the subsequent bit augmentation module leads to a secondary loss of high-frequency details. More seriously, the step-by-step training strategy causes gradient propagation conflicts in the mid-frequency region, which corresponds to key image phenotypic features; the gradient adversarial loss results in structural jagged artifacts in the reconstructed image. Furthermore, the split architecture requires maintaining two independent sets of network parameters, leading to wasted computational resources and a surge in deployment costs. The essence of this gradient adversarial approach lies in the fact that the pseudo-contours generated by low-bit degradation are low-frequency interference, while the blurring caused by low spatial resolution is a loss of high-frequency information. In tea leaf images, traditional cascaded inpainting often amplifies low-frequency bit noise while enhancing high-frequency details (such as fine hairs), and erases subtle textures while smoothing bit artifacts. Therefore, this architecture solves the energy coupling conflict between different degradation types in the frequency domain from a physical perspective through frequency domain decoupling. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method for joint reconstruction of tea image bit and spatial resolution based on frequency domain feature decoupling, which solves the problems of insufficient collaborative optimization, low efficiency and unsatisfactory reconstruction quality in the prior art.

[0005] This invention achieves the above objectives by adopting the following technical solution: This invention provides a method for joint reconstruction of tea image bit and spatial resolution based on frequency domain feature decoupling, comprising:

[0006] S1. Obtain the original low-bit, low-spatial-resolution tea leaf image and perform preprocessing.

[0007] S2. Construct a neural network for reconstructing tea leaf images;

[0008] S3. Train the tea image reconstruction neural network using the preprocessed low-bit, low-spatial-resolution tea images.

[0009] Phase 1, Independent Training:

[0010] The shared weight module is used to extract common features across degradation types, providing basic feature support for subsequent frequency band optimization, and outputting a common feature map containing global structural information of tea images;

[0011] The training process for the shared weight module includes:

[0012] Input a preprocessed low spatial resolution, low bit tea image patch, and extract common features through grouped convolution, pointwise convolution, and SCConv structure of the shared weight module. Use the shared weight loss function to optimize parameters to ensure that the extracted common features accurately capture the global structure of the tea.

[0013] The non-shared weight module is used to perform frequency band optimization for the degradation problem of tea images in different frequency bands, and solves high-frequency spatial degradation, low-frequency bit artifacts and mid-frequency coupling interference respectively, and outputs high-frequency, mid-frequency and low-frequency optimized feature maps.

[0014] The training process for non-shared weight modules includes:

[0015] The common features extracted from the input are used to generate corresponding high-frequency features, mid-frequency features, and low-frequency features through the high-frequency branch, mid-frequency branch, and low-frequency branch of the non-shared weight module, respectively.

[0016] The parameters of each branch are optimized by using a non-shared weight loss function in different frequency segments. The high-frequency branch focuses on leaf edge detail recovery, the low-frequency branch focuses on pseudo-contour suppression, and the mid-frequency branch focuses on lateral vein structure preservation.

[0017] The second stage involves end-to-end fine-tuning of the volumetric neural network.

[0018] After the shared and non-shared modules are pre-trained independently, the parameters of the shared and non-shared modules are fixed, and only the frequency domain fusion module is trained. During the pre-training of the overall neural network, the frequency domain fusion module receives the output of the non-shared weight module, generates fusion weights and outputs them. The fusion weights are optimized through the first hybrid loss function to ensure that the overall neural network is adapted to the double degradation repair of tea images.

[0019] During fine-tuning, a second hybrid loss function, which includes L1 pixel loss and frequency domain loss, is used to optimize the parameters of the frequency domain fusion module. L1 pixel loss is used to ensure pixel-level accuracy of tea color and repair color level breaks caused by low bit depth. Frequency domain loss, on the other hand, forces the model to learn more realistic high-frequency phenotypes by constraining the image spectrum distribution, thereby suppressing structural jagged artifacts in the reconstruction process.

[0020] After training is complete, freeze all parameters of the shared weight module, the non-shared weight module, and the frequency domain fusion module;

[0021] S4. Input the low spatial resolution, low bit image of tea leaves to be reconstructed into the trained neural network and output the final reconstructed image.

[0022] S5. The reconstruction quality of tea images is evaluated by peak signal-to-noise ratio and structural similarity index, and the reconstruction effect of the tea reconstruction neural network is comprehensively evaluated by combining subjective visual effects, including the sharpness of tea leaf edges and the degree of lateral vein restoration.

[0023] Furthermore, in step S1, the preprocessing includes scale normalization, pixel value normalization, and adaptive data augmentation;

[0024] The scaling includes: using reflection fill to adjust the size of the tea leaf image to an integer multiple of 128×128, and then cropping it into a 128×128 non-overlapping image block to ensure consistent image size and adapt to network input requirements; for images in the testing phase, only reflection fill is performed to an integer multiple of 128×128, without cropping to preserve the complete tea leaf shape.

[0025] The pixel value normalization includes: mapping the pixel values ​​of the tea image from the original range to the [0,1] interval;

[0026] The adaptive data augmentation includes:

[0027] Horizontal flip: 50% probability, to avoid deviation of the left and right texture of tea leaves;

[0028] Random rotation: angle range -15°~15°, step size 1°, simulating field shooting angle;

[0029] Brightness perturbation: ±15%, simulating changes in light intensity between morning and evening to adapt to tea images in different field scenes.

[0030] Furthermore, in step S2, the tea image reconstruction neural network includes a shared weight module, a non-shared weight module, and a frequency domain fusion module;

[0031] The shared weight module consists of three parts: hybrid convolutional layers, SCConv attention submodule, and residual connections.

[0032] The hybrid convolutional layer adopts an alternating structure of grouped convolution and pointwise convolution, with a total of 3 recurrent units and the number of channels being 32, 64 and 128 respectively. The first grouped convolutional layer uses a 3×3 convolutional kernel with a stride of 1 and padding of 1. The number of groups is equal to half the number of input channels, which is used for channel decoupling. The pointwise convolutional layer uses a 1×1 convolutional kernel with a stride of 1. The number of channels is consistent with the output of the grouped convolution, which is used to reconstruct channel correlation. The non-linear expression is introduced through the GELU activation function to adapt to the complex texture features of tea images.

[0033] The SCConv attention submodule includes a spatial attention branch and a channel attention branch. The spatial attention branch extracts local spatial information of the tea image through 3×3 convolution, including leaf edge and lateral vein position information, and generates spatial weights. The channel attention branch performs global average pooling and global max pooling on the feature map, and outputs channel weights after concatenation through 2 fully connected layers. Finally, the attention weight is equal to the product of the spatial weight and the channel weight, and dynamically allocates the processing weights of high-frequency texture and low-frequency color.

[0034] The output of each hybrid convolutional layer and the SCConv module is added to the input feature map via residual connections.

[0035] The non-shared weight module includes a high-frequency branch, a mid-frequency branch, and a low-frequency branch. The high-frequency branch uses 4-directional gradient-sensitive convolution to recover high-frequency details of tea leaves, including leaf edge serrations and tea bud hairs. It also uses a dynamic scaling factor to adaptively adjust edge sharpness and outputs a high-frequency optimized feature map.

[0036] The low-frequency branch is based on chroma-constrained residual blocks and contains two 3×3 convolutions with a chroma-aware layer inserted in the middle. The chroma-aware layer converts the feature map into the YCrCb space, calculates the pixel similarity between the Cr and Cb channels, applies smoothing constraints to regions with similarity greater than a set threshold, suppresses color level breaks caused by bit degradation, and outputs a low-frequency optimized feature map.

[0037] The mid-frequency branch includes a complex-domain wavelet convolutional layer and an MFDblock. The complex-domain wavelet convolutional layer uses the Daubechies-4 wavelet basis to decompose the feature map into real and imaginary parts. The real part corresponds to the intensity of the side vein texture, and the imaginary part corresponds to the texture direction. The features are fused into real features through 1×1 convolution. The local features and mid-frequency features within a set range are extracted by the 3×3 and 5×5 parallel convolutions of the MFDblock, respectively. The CAB is combined to enhance the channel where the side vein is located, and the mid-frequency optimized feature map is output.

[0038] The frequency domain fusion module consists of three parts: a frequency domain energy calculation layer, a weight generation layer, and a feature fusion layer.

[0039] The frequency domain energy calculation layer receives the high-frequency optimized feature map, mid-frequency optimized feature map, and low-frequency optimized feature map output by the non-shared module, performs fast Fourier transform on each to obtain frequency domain features, and calculates the energy of each frequency band.

[0040] The weight generation layer generates the corresponding fusion weights by normalizing the frequency domain energy of each frequency band using the Softmax function;

[0041] The feature fusion layer performs a weighted summation of the three frequency band feature maps according to their weights, and outputs the final reconstructed high bit-value, high spatial resolution tea image.

[0042] The beneficial effects of this invention are as follows:

[0043] This invention constructs a hybrid serial-parallel network architecture with shared and non-shared weights, dynamically extracts cross-degradation common features through a spatial-channel collaborative attention mechanism, and utilizes a three-band decoupled processor to optimize high-frequency, mid-frequency, and low-frequency features by frequency band, effectively solving the problems of missing mid-frequency features and task conflicts in existing methods.

[0044] This invention achieves adaptive weighted fusion of features across three frequency bands through a non-competitive frequency domain fusion gating network. It maintains cross-scale consistency in serial hierarchical feature transfer and eliminates task conflicts in parallel frequency division processing, thereby achieving efficient enhancement of tea images in complex degradation scenarios. Attached Figure Description

[0045] Figure 1 This is a flowchart of a method for joint reconstruction of tea image bit and spatial resolution based on frequency domain feature decoupling provided in an embodiment of the present invention. Detailed Implementation

[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0047] This invention provides a method for joint reconstruction of tea image bit and spatial resolution based on frequency domain feature decoupling, such as... Figure 1 As shown, it specifically includes:

[0048] S1. Obtain the original low-bit, low-spatial-resolution tea leaf image and perform preprocessing.

[0049] Preprocessing includes scale normalization, pixel value normalization, and adaptive data augmentation.

[0050] Scale normalization includes: using reflection fill to adjust the size of the tea leaf image to an integer multiple of 128×128 (e.g., if the original image is 200×200, fill it to 256×256), and then crop it into a 128×128 non-overlapping image block to ensure consistent image size and adapt to network input requirements; for images in the testing phase, only reflection fill is performed to an integer multiple of 128×128, without cropping to preserve the complete shape of the tea leaf.

[0051] Pixel value normalization includes mapping the pixel values ​​of the tea image from their original range (e.g., 0~15 for 4-bit, 0~255 for 8-bit) to the [0,1] interval, using the following formula:

[0052] ;

[0053] in, Represents the original pixel value. , These represent the maximum and minimum pixel values ​​for a single image, respectively, to meet the stability requirements of deep learning model training.

[0054] Adaptive data augmentation includes:

[0055] Horizontal flip: 50% probability, to avoid deviation of the left and right texture of tea leaves;

[0056] Random rotation: angle range -15°~15°, step size 1°, simulating field shooting angle;

[0057] Brightness perturbation: ±15%, simulating changes in light intensity between morning and evening to ensure the model can generalize and adapt to tea images in different field scenes.

[0058] S2. Construct a neural network for reconstructing tea leaf images;

[0059] The neural network for reconstructing tea images includes a shared weight module, a non-shared weight module, and a frequency domain fusion module;

[0060] The shared weight module is used to extract common features across degradation types (bit degradation, spatial resolution degradation), providing basic feature support for subsequent frequency band optimization, and outputting a common feature map containing global structural information of tea images.

[0061] The shared weight module consists of three parts: hybrid convolutional layers, SCConv attention submodule, and residual connections.

[0062] The hybrid convolutional layer adopts an alternating structure of grouped convolution and pointwise convolution, with a total of 3 recurrent units and the number of channels being 32, 64 and 128 respectively. The first grouped convolutional layer uses a 3×3 convolutional kernel with a stride of 1 and padding of 1. The number of groups is equal to half the number of input channels, which is used for channel decoupling. The pointwise convolutional layer uses a 1×1 convolutional kernel with a stride of 1. The number of channels is consistent with the output of the grouped convolution, which is used to reconstruct channel correlation. The nonlinear expression is introduced through the GELU activation function to adapt to the complex texture features of tea images.

[0063] The SCConv attention submodule includes a spatial attention branch and a channel attention branch. The spatial attention branch extracts local spatial information of the tea image through 3×3 convolution, including leaf edge and lateral vein position information, and generates spatial weights. The channel attention branch performs global average pooling and global max pooling on the feature map, and outputs channel weights after concatenation through 2 fully connected layers. Finally, the attention weight is equal to the product of the spatial weight and the channel weight, and dynamically allocates the processing weights of high-frequency texture and low-frequency color.

[0064] The output of each hybrid convolutional layer and the SCConv module is added to the input feature map via residual connections to avoid gradient vanishing and preserve the original feature information.

[0065] The non-shared weight module is used to perform frequency band optimization for the degradation problem of tea images in different frequency bands, and solves high-frequency spatial degradation, low-frequency bit artifacts and mid-frequency coupling interference respectively, and outputs optimized feature maps for high frequency, mid frequency and low frequency.

[0066] The non-shared weight module includes high-frequency branches, mid-frequency branches, and low-frequency branches;

[0067] The high-frequency branch uses 4-directional gradient (0°, 45°, 90°, 135°) sensitive convolution to recover high-frequency details of tea leaves, including leaf edge serrations and tea bud hairs. It also uses a dynamic scaling factor to adaptively adjust edge sharpness and outputs a high-frequency optimized feature map.

[0068] Dynamic scaling factor , This indicates the output of the shared module, α∈[0.2,0.8].

[0069] Where b represents the bias term, which works together with the weight matrix W on the shared feature Fshare. The result is mapped to the interval [0, 1] by the activation function σ, thereby realizing the dynamic weighted adjustment of features in different frequency bands.

[0070] The low-frequency branch is based on chroma-constrained residual blocks and contains two 3×3 convolutions with a chroma-sensing layer inserted in the middle. The chroma-sensing layer converts the feature map into the YCrCb space, calculates the pixel similarity between the Cr and Cb channels, applies smoothing constraints to regions with similarity greater than a set threshold (such as the smooth region of a tea leaf), suppresses color level breaks caused by bit degradation, and outputs a low-frequency optimized feature map.

[0071] The mid-frequency branch includes a complex-domain wavelet convolutional layer and an MFDblock. The complex-domain wavelet convolutional layer uses the Daubechies-4 wavelet basis to decompose the feature map into real and imaginary parts. The real part corresponds to the intensity of the side vein texture, and the imaginary part corresponds to the texture direction. The features are fused into real features through 1×1 convolution. The local features and mid-frequency features within a set range are extracted by the 3×3 and 5×5 parallel convolutions of the MFDblock, respectively. The CAB is combined to enhance the channel where the side vein is located, and the mid-frequency optimized feature map is output.

[0072] The frequency domain fusion module learns adaptive fusion weights based on frequency domain energy to weightedly fuse high-frequency, mid-frequency, and low-frequency optimized feature maps from the non-shared weight module, thereby obtaining a final high-quality tea image. The frequency domain fusion module only performs frequency domain feature fusion operations and does not involve additional degradation repair mapping. The fusion weights are adaptively learned by the network and are guaranteed to be non-negative and sum to 1.

[0073] The frequency domain fusion module consists of three parts: a frequency domain energy calculation layer, a weight generation layer, and a feature fusion layer.

[0074] The frequency domain energy calculation layer receives the high-frequency optimized feature map, mid-frequency optimized feature map, and low-frequency optimized feature map output by the non-shared module, performs Fast Fourier Transform on each to obtain the frequency domain features, and calculates the energy of each frequency band. The calculation formula is as follows:

[0075] ;

[0076] in, Indicates the energy of each frequency band. This represents the optimized feature map for each frequency band. Indicates frequency band, , Indicates high frequency band, Indicates mid-frequency band, This indicates the mid-frequency band, where H and W represent the feature map height and width, respectively.

[0077] The weight generation layer generates the corresponding fusion weights for each frequency band by normalizing the frequency domain energy of each band using the Softmax function, as follows:

[0078] ;

[0079] in, Indicates the temperature coefficient;

[0080] The above formula is used to calculate... , , , respectively representing high-frequency weight, mid-frequency weight, and low-frequency weight;

[0081] The feature fusion layer performs a weighted summation of the three frequency band feature maps according to their weights, and outputs the final reconstructed high-bit-value, high-spatial-resolution tea image, as follows:

[0082] ;

[0083] in, This represents the final reconstructed high-bit, high-spatial-resolution image of tea leaves. Represents a high-frequency optimized feature map. This represents the mid-frequency optimized feature map. This represents a low-frequency optimized feature map.

[0084] The purpose of the loss function is to optimize the model's performance by minimizing the difference between the network's output image and the real high-quality tea image, ensuring that each module works together to meet the dual degradation repair needs of tea images.

[0085] Loss function of shared weight module for:

[0086] ;

[0087] In the formula, H represents the height of the feature map, W represents the width of the feature map, and C represents the number of channels in the feature map.

[0088] By comparing the common feature maps output by the shared modules Common feature maps corresponding to real high-quality images To address the differences, mean squared error (MSE) is used to calculate the error and optimize the shared module, ensuring the accuracy of common feature extraction.

[0089] The loss function for the non-shared weight module is:

[0090] ;

[0091] in:

[0092] ;

[0093] ;

[0094] in This represents the complex domain wavelet transform, used to project mid-frequency features onto a specific space to calculate the difference between the reconstructed value and the true value.

[0095] ;

[0096] Optimize each branch of the non-shared module by using frequency band loss.

[0097] The loss function of the frequency domain fusion module is:

[0098] ;

[0099] When training the frequency domain fusion module, minimize the final reconstructed image. Compared to real high bit-value, high spatial resolution tea images By addressing the MSE differences, the fusion weights are optimized to make the output image closer to the quality of a real image.

[0100] S3. Train the tea image reconstruction neural network using the preprocessed low-bit, low-spatial-resolution tea images.

[0101] The first stage involves independent pre-training of the shared and non-shared weight modules: In this stage, the dual degradation repair task is broken down into two sub-tasks: common feature extraction and frequency band optimization. The shared and non-shared modules are trained separately.

[0102] The shared weight module is used to extract common features across degradation types, providing basic feature support for subsequent frequency band optimization, and outputting a common feature map containing global structural information of tea images.

[0103] The shared module training process includes:

[0104] The input is a preprocessed low spatial resolution, low bit depth tea leaf image patch. Common features are extracted using a shared module's "group convolution + pointwise convolution + SCConv" structure, and a shared weight loss function is applied. Parameters were optimized to ensure that the extracted common features accurately captured the global structure of the tea leaves (such as leaf outline and canopy distribution). At this stage, L1 pixel loss was also employed to initially guarantee pixel-level restoration of the tea leaf color and outline. The optimizer used was AdamW with an initial learning rate of 2e-4 and 100,000 training iterations.

[0105] The non-shared weight module is used to perform frequency band optimization for the degradation problem of tea images in different frequency bands, and solves high-frequency spatial degradation, low-frequency bit artifacts and mid-frequency coupling interference respectively, and outputs high-frequency, mid-frequency and low-frequency optimized feature maps.

[0106] The training process for non-shared modules includes:

[0107] The common features extracted from the input are used to generate corresponding high-frequency features, mid-frequency features, and low-frequency features through the high-frequency branch, mid-frequency branch, and low-frequency branch of the non-shared module, respectively.

[0108] Using non-shared weight loss function The parameters of each branch are optimized by frequency band. The high-frequency branch focuses on leaf edge detail recovery, the low-frequency branch focuses on pseudo-contour suppression, and the mid-frequency branch focuses on lateral vein structure preservation. The optimizer used is AdamW with an initial learning rate of 2e-4 and 80,000 training iterations.

[0109] The second stage involves end-to-end fine-tuning of the entire neural network:

[0110] After the shared and non-shared modules are pre-trained independently, the parameters of the shared and non-shared modules are fixed, and only the frequency domain fusion module is trained. The main goal of this stage is to learn adaptive fusion weights based on frequency domain energy to weight and fuse the high-frequency, mid-frequency, and low-frequency optimized feature maps from the non-shared weight module to obtain the final high-quality tea image.

[0111] During overall training, the input is a low spatial resolution, low bit depth tea leaf image I. LR-LBD The parameters of the shared weight module and the non-shared weight module are fixed, and only the frequency domain fusion module is trained. The frequency domain fusion module receives the three-band feature maps output by the non-shared module, generates fusion weights, and outputs I. final The fusion weights are optimized by using the first hybrid loss function to ensure that the overall network is adapted to the double degradation repair of tea images.

[0112] The first mixed loss function is:

[0113] ;

[0114] During training, the optimizer uniformly adopted AdamW (weight decay = 1e-4, β1 = 0.9, β2 = 0.999); the learning rate adopted a "linear warm-up + cosine annealing" strategy: it increased from 1e-6 to 2e-4 in the first 100,000 iterations, and then decayed to 0.5 times the current learning rate every 50,000 iterations, with a minimum of 1e-6; the training batch size was 16, and the total number of iterations was 500,000.

[0115] During fine-tuning, the second hybrid loss function L is used. total2 = L1+ λL fft Optimize the parameters of the frequency domain fusion module. L1 pixel loss is used to ensure pixel-level accuracy of tea color and repair color banding caused by low bit depth. Frequency domain loss L fft By constraining the image spectral distribution, the model is forced to learn more realistic high-frequency phenotypes, fundamentally suppressing structural jagged artifacts in the reconstruction process. The parameters of the frequency domain fusion module are optimized using a hybrid loss function. The optimization strategy adopts AdamW combined with cosine annealing scheduling (the first 100,000 warm-ups, and then the learning rate decays to 0.5 times every 50,000 times), and the training is carried out for 320,000 times.

[0116] After training is complete, freeze all parameters of the shared weight module, the non-shared weight module, and the frequency domain fusion module;

[0117] S4. Input the low spatial resolution, low bit image of tea leaves to be reconstructed into the trained neural network and output the final reconstructed image.

[0118] The goal of the inference stage is to generate high-quality tea images in real time, adapting to mobile monitoring scenarios in the field. To meet the computing power limitations of edge devices, this architecture introduces SCConv spatial channel reconstruction convolution. By identifying and suppressing spatial and channel redundancy in the feature map, it significantly reduces memory consumption and computational cost while maintaining high-precision reconstruction, ensuring the real-time inference performance of the algorithm on portable tea garden monitoring terminals.

[0119] Input image: Input a low spatial resolution, low bit depth image of tea leaves to be reconstructed. LR-LBD (e.g., 4-bit+8x downsampled tea bud images and canopy images taken in the field). These images do not need to be cropped. They are simply filled to an integer multiple of 128×128 by reflection, and then the pixel values ​​are normalized (mapped to [0,1]) before being input into the neural network.

[0120] Network model inference: The core of the shared weight module adopts SCConv (Spatial Channel Reconstruction Convolution), utilizing its Spatial Reconstruction Unit (SRU) to separate spatial feature redundancy, and in conjunction with the Channel Reconstruction Unit (CRU) to reduce channel interference across degenerate tasks. This design can extract the global contour of tea leaves while suppressing noise propagation caused by low bits, providing clean baseline features for subsequent frequency division optimization. The shared module receives I... LR-LBD Output a common feature map F containing the global structure of tea leaves. share The non-shared module receives F share High-frequency optimized feature F is output through high-frequency branching. high (Focusing on reconstructing the morphology of tea bud tips, leaf margin serrations, and distribution of white down), low-frequency optimized feature F is output through low-frequency branching. low (Eliminates color banding on leaf surfaces caused by 4-bit sampling, restoring continuous color transitions such as light green and emerald green), outputs intermediate frequency optimized feature F through the intermediate frequency branch. mid (Texture compensation is performed to enhance the glossiness and lateral vein patterns of tea leaves); the frequency domain fusion module receives feature maps from three frequency bands, calculates frequency domain energy, generates adaptive fusion weights, and outputs the final reconstructed image I. final The physical significance of introducing the mid-frequency branch lies in resolving the gradient conflict between high-frequency sharpening and low-frequency smoothing. Since the side veins of tea leaves belong to the mid-to-high frequency transition region, this branch, through adaptive residual learning, specifically compensates for texture details that are misidentified as noise in super-resolution reconstruction or mistakenly erased in bit enhancement, ensuring I... finalThe frequency domain energy distribution of the image is continuous and natural. The frequency domain fusion module integrates a lightweight energy sensing module, which calculates F using global average pooling (GAP) and an adaptive activation function. high F mid F low The characteristic activation intensity of each channel. The system automatically generates a dynamic scaling factor based on the degree of input degradation. , , By using weighted fusion to balance the edge sharpness and color smoothness of the image, it avoids the graying or over-sharpening phenomena caused by traditional fixed weights.

[0121] Image post-processing: For I final Perform pixel value denormalization (map from [0,1] back to 0~65535, adapting to 16-bit output), crop the reflection fill area before inference, and restore it to the original size of the input image; perform CLAHE enhancement on the denormalized image (only on the Y channel of YCrCb space, contrast limit = 2.0, grid size 8×8) to enhance the distinction between tea leaves and background (branches, soil).

[0122] Impact of joint reconstruction on downstream tasks: The enhanced images processed by this architecture have significant benefits for intelligent tea processing: In the tea bud recognition task, the recall rate of small target tea buds is significantly improved due to the restoration of high-frequency details; In the automatic grading task of tea tenderness, the grading accuracy (mAP) can achieve a significant gain compared with the original low-quality image because the bit depth enhancement restores the subtle color and texture differences, providing high-quality data support for the precise management of tea trees.

[0123] The inference process is end-to-end, requiring only one forward propagation to complete all processing steps from low-quality input to high-quality output. The model does not require retraining during deployment, and can quickly generate output images from any low spatial resolution, low bit tea image.

[0124] S5. The reconstruction quality of tea images is evaluated by peak signal-to-noise ratio and structural similarity index, and the reconstruction effect of the tea reconstruction neural network is comprehensively evaluated by combining subjective visual effects, including the sharpness of tea leaf edges and the degree of lateral vein restoration.

[0125] Peak signal-to-noise ratio The calculation method is as follows:

[0126] ;

[0127] Among them, MAX I 2 This represents the square of the maximum possible value of a pixel in a tea leaf image. For an 8-bit image, the pixel value ranges from 0 to 255. I=255, the numerator is 255 2 For 16-bit images, MAX I =65535, the numerator is 65535 2 The molecule represents the maximum theoretical intensity of the tea image signal. MSE represents the mean squared error, which is the average squared difference between each pixel in the real high-quality tea image and the reconstructed image, expressed as:

[0128] ;

[0129] Peak signal-to-noise ratio (PSNR) reflects the error level of tea image reconstruction. A higher value indicates that the pixel difference between the reconstructed image and the real image is smaller, and the image quality is better. For the Tea-Degrade-3K test set (including Tea_Leaf, Tea_Shoot, and Tea_Canopy sub-datasets), the present invention can achieve a PSNR of 36.7dB on the Tea_Leaf sub-dataset, which is 10.2dB higher than the traditional ZP-Bic algorithm (26.5dB).

[0130] The Structural Similarity Index (SSIM) is calculated using the following formula:

[0131] ;

[0132] Where x is the original high-quality tea image, and y is the reconstructed tea image; μ x and μ y σ represents the pixel mean of image x and y, respectively; x 2 and σ y 2 σ represents the pixel variance of image x and y, respectively; xy Represents the pixel covariance of image x and y; C1 = (K1 × MAX) I ) 2 C2=(K2×MAX) I ) 2 Small constants (K1=0.01, K2=0.03) are introduced to avoid the denominator being zero.

[0133] The structural similarity index evaluates the consistency of structural information in tea images. The closer the value is to 1, the more consistent the reconstructed image is with the real image in terms of structural features such as texture and contour, and the closer it is to the quality perceived by the human eye. The present invention achieves an SSIM of 0.963 on the Tea_Shoot subset dataset, which is 3.3% higher than the DCAFusion algorithm (0.930).

[0134] This invention comprehensively evaluates the model's ability to repair double degradation of tea images by quantitative analysis of peak signal-to-noise ratio and structural similarity index, combined with subjective visual assessment of leaf edge serration clarity, lateral vein branch integrity, and leaf pseudo-contour suppression effect, and verifies the model's practicality in complex field scenarios.

[0135] The above description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein and should not be construed as excluding other embodiments. It can be used in various other combinations, modifications, and environments, and can be altered within the scope of the concept described herein through the above teachings or related technologies or knowledge. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention should be within the protection scope of the appended claims.

Claims

1. A method for joint reconstruction of tea image bit and spatial resolution based on frequency domain feature decoupling, characterized in that, include: S1. Obtain the original low-bit, low-spatial-resolution tea leaf image and perform preprocessing. S2. Construct a neural network for reconstructing tea images, wherein the neural network for reconstructing tea images includes a shared weight module, a non-shared weight module, and a frequency domain fusion module; The shared weight module consists of three parts: hybrid convolutional layers, SCConv attention submodule, and residual connections. The hybrid convolutional layer adopts an alternating structure of grouped convolution and pointwise convolution, with a total of 3 recurrent units and the number of channels being 32, 64 and 128 respectively. The first grouped convolutional layer uses a 3×3 convolutional kernel with a stride of 1 and padding of 1. The number of groups is equal to half the number of input channels, which is used for channel decoupling. The pointwise convolutional layer uses a 1×1 convolutional kernel with a stride of 1. The number of channels is consistent with the output of the grouped convolution, which is used to reconstruct channel correlation. The non-linear expression is introduced through the GELU activation function to adapt to the complex texture features of tea images. The SCConv attention submodule includes a spatial attention branch and a channel attention branch. The spatial attention branch extracts local spatial information of the tea image through 3×3 convolution, including leaf edge and lateral vein position information, and generates spatial weights. The channel attention branch performs global average pooling and global max pooling on the feature map, and outputs channel weights after concatenation through 2 fully connected layers. Finally, the attention weight is equal to the product of the spatial weight and the channel weight, and dynamically allocates the processing weights of high-frequency texture and low-frequency color. The output of each hybrid convolutional layer and the SCConv module is added to the input feature map via residual connections; S3. Train the tea image reconstruction neural network using the preprocessed low-bit, low-spatial-resolution tea images. Phase 1, Independent Training: The shared weight module is used to extract common features across degradation types, providing basic feature support for subsequent frequency band optimization, and outputting a common feature map containing global structural information of tea images; The training process for the shared weight module includes: Input a preprocessed low spatial resolution, low bit tea image patch, and extract common features through grouped convolution, pointwise convolution, and SCConv structure of the shared weight module. Use the shared weight loss function to optimize parameters to ensure that the extracted common features accurately capture the global structure of the tea. The non-shared weight module is used to perform frequency band optimization for the degradation problem of tea images in different frequency bands, and solves high-frequency spatial degradation, low-frequency bit artifacts and mid-frequency coupling interference respectively, and outputs high-frequency, mid-frequency and low-frequency optimized feature maps. The training process for non-shared weight modules includes: The common features extracted from the input are used to generate corresponding high-frequency features, mid-frequency features, and low-frequency features through the high-frequency branch, mid-frequency branch, and low-frequency branch of the non-shared weight module, respectively. The parameters of each branch are optimized by using a non-shared weight loss function in different frequency segments. The high-frequency branch focuses on leaf edge detail recovery, the low-frequency branch focuses on pseudo-contour suppression, and the mid-frequency branch focuses on lateral vein structure preservation. The second stage involves end-to-end fine-tuning of the volumetric neural network. After the shared and non-shared modules are pre-trained independently, the parameters of the shared and non-shared modules are fixed, and only the frequency domain fusion module is trained. During the pre-training of the overall neural network, the frequency domain fusion module receives the output of the non-shared weight module, generates fusion weights and outputs them. The fusion weights are optimized through the first hybrid loss function to ensure that the overall neural network is adapted to the double degradation repair of tea images. During fine-tuning, a second hybrid loss function, which includes L1 pixel loss and frequency domain loss, is used to optimize the parameters of the frequency domain fusion module. L1 pixel loss is used to ensure pixel-level accuracy of tea color and repair color level breaks caused by low bit depth. Frequency domain loss, on the other hand, forces the model to learn more realistic high-frequency phenotypes by constraining the image spectrum distribution, thereby suppressing structural jagged artifacts in the reconstruction process. After training is complete, freeze all parameters of the shared weight module, the non-shared weight module, and the frequency domain fusion module; S4. Input the low spatial resolution, low bit image of tea leaves to be reconstructed into the trained neural network and output the final reconstructed image. S5. The reconstruction quality of tea images is evaluated by peak signal-to-noise ratio and structural similarity index, and the reconstruction effect of the tea reconstruction neural network is comprehensively evaluated by combining subjective visual effects, including the sharpness of tea leaf edges and the degree of lateral vein restoration.

2. The method for joint reconstruction of tea image bit and spatial resolution based on frequency domain feature decoupling according to claim 1, characterized in that, In step S1, preprocessing includes scale normalization, pixel value normalization, and adaptive data augmentation; The scaling includes: using reflection fill to adjust the size of the tea leaf image to an integer multiple of 128×128, and then cropping it into a 128×128 non-overlapping image block to ensure consistent image size and adapt to network input requirements; for images in the testing phase, only reflection fill is performed to an integer multiple of 128×128, without cropping to preserve the complete tea leaf shape. The pixel value normalization includes: mapping the pixel values ​​of the tea image from the original range to the [0,1] interval; The adaptive data augmentation includes: Horizontal flip: 50% probability, to avoid deviation of the left and right texture of tea leaves; Random rotation: angle range -15°~15°, step size 1°, simulating field shooting angle; Brightness perturbation: ±15%, simulating changes in light intensity between morning and evening to adapt to tea images in different field scenes.

3. The method for joint reconstruction of tea image bit and spatial resolution based on frequency domain feature decoupling according to claim 1, characterized in that, The non-shared weight module includes a high-frequency branch, a mid-frequency branch, and a low-frequency branch. The high-frequency branch uses 4-directional gradient-sensitive convolution to recover high-frequency details of tea leaves, including leaf edge serrations and tea bud hairs. It also uses a dynamic scaling factor to adaptively adjust edge sharpness and outputs a high-frequency optimized feature map. The low-frequency branch is based on chroma-constrained residual blocks and contains two 3×3 convolutions with a chroma-aware layer inserted in the middle. The chroma-aware layer converts the feature map into the YCrCb space, calculates the pixel similarity between the Cr and Cb channels, applies smoothing constraints to regions with similarity greater than a set threshold, suppresses color level breaks caused by bit degradation, and outputs a low-frequency optimized feature map. The mid-frequency branch includes a complex-domain wavelet convolutional layer and an MFDblock. The complex-domain wavelet convolutional layer uses the Daubechies-4 wavelet basis to decompose the feature map into real and imaginary parts. The real part corresponds to the intensity of the side vein texture, and the imaginary part corresponds to the texture direction. The features are fused into real features through 1×1 convolution. The local features and mid-frequency features within a set range are extracted by the 3×3 and 5×5 parallel convolutions of the MFDblock, respectively. The CAB is combined to enhance the channel where the side vein is located, and the mid-frequency optimized feature map is output.

4. The method for joint reconstruction of tea image bit and spatial resolution based on frequency domain feature decoupling according to claim 3, characterized in that, The frequency domain fusion module consists of three parts: a frequency domain energy calculation layer, a weight generation layer, and a feature fusion layer. The frequency domain energy calculation layer receives the high-frequency optimized feature map, mid-frequency optimized feature map, and low-frequency optimized feature map output by the non-shared module, performs fast Fourier transform on each to obtain frequency domain features, and calculates the energy of each frequency band. The weight generation layer generates the corresponding fusion weights by normalizing the frequency domain energy of each frequency band using the Softmax function; The feature fusion layer performs a weighted summation of the three frequency band feature maps according to their weights, and outputs the final reconstructed high bit-value, high spatial resolution tea image.

Citation Information

Patent Citations

  • Tea leaf picking point detection method based on light field camera

    CN115205842A

  • Pig image BDE reconstruction system and method based on differential image rate filtering

    CN119991857A