Sugarcane tail counting method based on wavelet attention fusion mechanism

Through the wavelet attention fusion mechanism, the color similarity interference, fixed frequency domain decomposition defects and occlusion problems in sugarcane tail counting are solved, and efficient and accurate counting of sugarcane tails is achieved.

CN120339202APending Publication Date: 2025-07-18GUANGXI YUEGUI GUANGYE HOLDINGS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510373903.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing convolutional neural network model has problems such as color similarity interference, fixed frequency domain decomposition defects, occlusion and small target miss detection and loss function design in sugarcane tail counting, resulting in insufficient counting accuracy.

Method used

Using a method based on the wavelet attention fusion mechanism, the sucrose tail feature decomposition, enhancement and counting model optimization is performed by learning wavelet transform, space-channel attention weights and multitasking loss functions.

Benefits of technology

It improves the feature distinction of color-like targets, counting robustness in occlusion scenarios and high-frequency details retention capabilities, reduces false detection and missed detection, and improves counting accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339202A_ABST
    Figure CN120339202A_ABST
Patent Text Reader

Abstract

The invention relates to the field of image processing, in particular to a wavelet attention fusion mechanism-based sugarcane tail counting method, which comprises the following steps of: acquiring an original feature map of a sugarcane tail through a camera; processing the original feature map through wavelet rolling weaving to obtain a sub-band feature map; performing weighted feature enhancement on the sub-band feature map through the attention weight to obtain a weighted sub-band feature map; processing the original feature map through a space attention weight and a channel attention weight to obtain a weighted feature map; performing up-sampling on the final feature map through a decoder, generating a density map, and obtaining the number of the sugarcane tails according to the density map; collecting image data of sugarcane tails and constructing a data set; and constructing an initial counting model, and performing model training through the data set to obtain an optimized counting model. According to the method, the feature discrimination degree and the high-frequency detail retention capability of targets with similar colors and the counting robustness in a shielding scene can be synchronously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and in particular to a sugarcane tail counting method based on a wavelet attention fusion mechanism. Background Art

[0002] In the sugar-making process, sugarcane tail counting is mainly used to optimize raw material processing and quality control. The tails of sugarcane usually have a lower sugar content and more fibers, which may affect the sugar-making efficiency and the quality of the finished product. With the development of modern agricultural intelligence and industrial quality inspection automation, the demand for dense target counting technology is increasing in scenarios such as sugarcane tail statistics and cell counting. Traditional counting methods based on convolutional neural networks (CNNs) have significant limitations in dealing with targets with similar colors (such as both sugarcane tails and the main bodies being green), dense occlusion, and small-sized targets:

[0003] Color similarity interference: Existing CNN models rely on RGB space features and are difficult to distinguish targets with subtle color differences (such as sugarcane tails and stalks), resulting in feature confusion and misdetection;

[0004] Fixed frequency domain decomposition defect: Traditional discrete wavelet transform (DWT) uses predefined filters (such as Haar), which cannot adapt to the data characteristics and have insufficient ability to extract edge features of complex textures;

[0005] Occlusion and missed detection of small targets: In dense overlapping scenarios, existing spatial attention mechanisms are difficult to model the local details of occluded targets, and the high-frequency information of small targets is easily lost due to downsampling operations;

[0006] Single loss function design: Mainstream methods only optimize the error of the density map in the spatial domain and ignore the frequency domain gradient alignment constraint, resulting in blurred edges in the prediction results and cumulative counting deviations.

[0007] Existing improvement schemes such as frequency domain enhancement networks mostly use fixed wavelet bases and lack an end-to-end learnable mechanism; while attention models mostly focus on the spatial or channel dimensions and are not deeply integrated with wavelet band decomposition, making it difficult to achieve a balance between noise suppression and key feature enhancement. Summary of the Invention

[0008] To solve the above problems, the present invention provides a sugarcane tail counting method based on a wavelet attention fusion mechanism, which can simultaneously improve the feature discrimination of targets with similar colors, the ability to retain high-frequency details, and the counting robustness in occlusion scenarios.

[0009] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0010] A sugarcane tail counting method based on a wavelet attention fusion mechanism, comprising the following steps:

[0011] S1. Take pictures of the sugarcane tails on the conveyor belt through a camera to obtain the original feature map of the sugarcane tails;

[0012] S2. Obtain the original feature map in step S1, and process the original feature map through wavelet convolution to decompose the original feature map into a low-frequency sub-band and a high-frequency sub-band, and obtain the sub-band feature map;

[0013] S3. Obtain the sub-band feature maps in step S2, generate attention weights for each sub-band feature map, and perform weighted feature enhancement on the sub-band feature maps through the attention weights to obtain the weighted sub-band feature maps;

[0014] S4. Obtain the original feature map in step S1, and process the original feature map through spatial attention weights and channel attention weights to obtain the weighted feature map;

[0015] S5. Concatenate the weighted sub-band feature maps and the weighted feature maps along the channel dimension, and obtain the final feature map through convolutional fusion and residual connection;

[0016] S6. Obtain the final feature map in step S5, and perform upsampling on the final feature map through a decoder to restore the original size of the final feature map, and generate a density map. Extract the center points of the targets in the density map to obtain the number of sugarcane tails in the density map;

[0017] S7. Collect the image data of the sugarcane tails and generate the corresponding density maps as labels to construct a data set;

[0018] S8. Construct an initial counting model through steps S2 - S6, and perform model training through the data set to obtain an optimized counting model.

[0019] Further, in step S1, in step S1, the original image includes an RGB image size, a scene description text, and physical parameters. The RGB image size is H×W×3, where H is the height of the image; W is the width of the image; the scene description text is a scene description, and the scene description text is used to provide semantic prior information; the physical parameters include the notch size, the conveyor belt speed, and the maximum stacking density.

[0020] Further, in step S2, the learnable wavelet transform decomposes the original feature map into four sub-bands through a convolutional kernel:

[0021] X wavelet =DWT(X)=X*W DWT Formula (1)

[0022] Where X waveletFor the sub-band feature map, which includes a low-frequency sub-band, a horizontal high-frequency sub-band, a vertical high-frequency sub-band, and a diagonal high-frequency sub-band; DWT is the wavelet transform; X is the original feature map; W DWT is the convolution kernel.

[0023] Further, in step S2, the gradient magnitude is calculated for each high-frequency sub-band to obtain a gradient alignment term, and a gradient alignment constraint is introduced in the wavelet domain through the gradient alignment term:

[0024]

[0025]

[0026] Among them, is the gradient magnitude; is the horizontal direction gradient; is the vertical direction gradient; L wavelet is the gradient alignment term; is the gradient of the predicted density map in the k-th sub-band; is the gradient of the true density map in the k-th sub-band.

[0027] Further, in step S3, a band attention mechanism is set. The band attention mechanism generates attention weights for each sub-band feature map through global average pooling and a two-layer fully connected network:

[0028]

[0029] Among them, is the attention weight; σ(·) is the Sigmoid function; f fc (·) is the two-layer fully connected network; GAP(·) is the global pooling; is the sub-band feature map;

[0030] The weighted sub-band feature map is reconstructed to the original resolution through the learnable inverse wavelet transform, then there is:

[0031] X' wavelet = IDWT(X wavelet ) = X wavelet *W IDWT Formula (5)

[0032] Among them, IDWT is the learnable inverse wavelet transform; X' wavelet is the sub-band feature map at the original resolution; W IDWT ∈R 2 ×2×C×4C is the trainable transposed convolution kernel to restore the feature space size to H×W.

[0033] Furthermore, in step S4, the channel attention branch generates channel attention weights through two layers of fully connected networks:

[0034] A channel = σ(f fc (GAP(X))) Equation (6)

[0035] where A channel ∈ R C×1×1 is the channel attention weight; X is the original feature map; f fc (·) is two layers of fully connected networks; GAP(·) is global pooling;

[0036] The spatial attention branch generates spatial attention weights through a 3×3 convolution:

[0037] A spatial = σ(f conv (X)) Equation (7)

[0038] where A spatial ∈ R 1×H×W is the spatial attention weight; σ(·) is the Sigmoid function; f conv is the convolution function;

[0039] Processing the original feature map according to the spatial attention weight and the channel attention weight, then we have:

[0040] X att = X · A channel · A spatial Equation (8)

[0041] where X att ∈ R C×H×W is the weighted feature map.

[0042] Furthermore, in step S5, after the weighted sub-band feature map and the weighted feature map are concatenated, they are fused through a 1×1 convolution to obtain a fused map:

[0043] X fused = f conv1x1 (Concat(X' wavelet , X att )) Equation (9)

[0044] where X fused is the fused map, f conv1x1 is the 1×1 convolution; Concat is concatenation along the channel dimension;

[0045] The fused map obtains the final feature map through a residual connection:

[0046] Y = X + X fusedFormula (10)

[0047] Wherein, Y is the final feature map; X is the original feature map; X fused is the fusion map.

[0048] Further, in step S8, the training method of the initial counting model is: processing the data set through steps S2 - S6, and performing loss calculation, and optimizing the initial counting model according to the loss calculation to obtain an optimized counting model.

[0049] Further, the loss calculation includes the loss in the spatial domain, the loss of gradient alignment in the wavelet domain, and the total loss:

[0050] The loss in the spatial domain is:

[0051]

[0052] Wherein, D pred is the predicted density map; D gt is the true density map;

[0053] The loss of gradient alignment in the wavelet domain is:

[0054]

[0055] Wherein, is the gradient of the predicted density map in the k - th sub - band; is the gradient of the true density map in the k - th sub - band;

[0056] The total loss combines the loss in the spatial domain and the loss of gradient alignment in the wavelet domain, and performs backpropagation and parameter update to optimize the model. The total loss is:

[0057]

[0058] Wherein, L WGDL is the total loss function; LH is the horizontal high - frequency sub - band; HL is the vertical high - frequency sub - band, HH is the diagonal high - frequency sub - band; λ is the balance coefficient.

[0059] The beneficial effects of the present invention are:

[0060] By performing dynamic matching of the target frequency domain characteristics through learnable wavelet transform, it is possible to effectively distinguish the subtle color difference features between sugarcane stalks and sugarcane tails, reduce the false detection rate caused by color confusion, and accurately distinguish targets with similar colors; by processing the original feature map with spatial attention weights and channel attention weights, based on a spatial-channel dual-path fusion network, combined with a high-frequency feature enhancement mechanism for occluded regions, the ability to capture local details of densely overlapping sugarcane tails is improved, and the problem of missed detection is reduced; a band attention mechanism is used to adaptively adjust the multi-band feature weights to improve the target boundary localization accuracy; a band attention mechanism is used to adaptively adjust the multi-band feature weights to improve the target boundary localization accuracy; through a multi-task joint loss function of the loss in the spatial domain, the gradient alignment loss in the wavelet domain, and the total loss, it is possible to simultaneously constrain the frequency domain reconstruction quality and optimize the counting accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 is a schematic structural diagram of a sugarcane tail counting method based on a wavelet attention fusion mechanism according to a preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0062] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0063] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0064] Please refer to Figure 1 , a sugarcane tail counting method based on a wavelet attention fusion mechanism according to a preferred embodiment of the present invention, includes the following steps:

[0065] S1. Shoot the sugarcane tails on the conveyor belt through a camera to obtain the original feature map of the sugarcane tails.

[0066] In step S1, in step S1, the original image includes the RGB image size, scene description text, and physical parameters. The RGB image size is H×W×3, where H is the height of the image; W is the width of the image; the scene description text is the scene description, and the scene description text is used to provide semantic prior information; the physical parameters include the notch size, conveyor belt speed, and maximum stacking density.

[0067] S2. Obtain the original feature map of step S1, and process the original feature map through wavelet convolution to decompose the original feature map into a low-frequency sub-band and high-frequency sub-bands, and obtain sub-band feature maps.

[0068] In step S2, the learnable wavelet transform decomposes the original feature map into four sub-bands through a convolution kernel:

[0069] X wavelet = DWT(X) = X * W DWT Formula (1)

[0070] where X wavelet is the sub-band feature map, which includes a low-frequency sub-band, a horizontal high-frequency sub-band, a vertical high-frequency sub-band, and a diagonal high-frequency sub-band; DWT is the wavelet transform; X is the original feature map; W DWT is the convolution kernel.

[0071] In step S2, the gradient magnitude is calculated for each high-frequency sub-band to obtain a gradient alignment term, and a gradient alignment constraint is introduced in the wavelet domain through the gradient alignment term:

[0072]

[0073] where is the gradient magnitude; is the horizontal direction gradient; is the vertical direction gradient; L wavelet is the gradient alignment term; is the gradient of the predicted density map in the k-th sub-band; is the gradient of the ground truth density map in the k-th sub-band.

[0074] In this embodiment, the low-frequency sub-band retains the main structure of the cane tail, and the high-frequency sub-bands encode the edge texture details, so as to be able to effectively distinguish targets with similar colors. Moreover, based on the frequency band decomposition and gradient alignment optimization of the Haar wavelet, this embodiment can capture the cane tail edge texture features through the high-frequency sub-bands (LH / HL / HH) to solve the problem of similar colors. The wavelet domain gradient alignment constraint is used to enhance the robustness of the model to occluded regions.

[0075] S3. Obtain the sub-band feature maps of step S2, generate attention weights for each sub-band feature map, and perform weighted feature enhancement on the sub-band feature maps through the attention weights to obtain weighted sub-band feature maps.

[0076] In step S3, a frequency band attention mechanism is set up. The frequency band attention mechanism generates attention weights for each sub-band feature map through global average pooling and a two-layer fully connected network:

[0077]

[0078] Among them, is the attention weight; σ(·) is the Sigmoid function; f fc (·) is a two-layer fully connected network; GAP(·) is global pooling; is the sub-band feature map;

[0079] The weighted sub-band feature map is reconstructed to the original resolution through the learnable inverse wavelet transform, then there is:

[0080] X' wavelet = IDWT(X wavelet ) = X wavelet *W IDWT Formula (5)

[0081] Among them, IDWT is the learnable inverse wavelet transform; X' wavelet is the sub-band feature map at the original resolution; W IDWT ∈R 2 ×2×C×4C is a trainable transposed convolution kernel to restore the feature space size to H×W.

[0082] In this embodiment, by setting the frequency band attention mechanism, the effective frequency band can be dynamically enhanced.

[0083] S4. Obtain the original feature map in step S1, and process the original feature map through the spatial attention weight and the channel attention weight to obtain a weighted feature map.

[0084] In step S4, the channel attention branch generates the channel attention weight through a two-layer fully connected network:

[0085] A channel = σ(f fc (GAP(X))) Formula (6)

[0086] Among them, A channel ∈R C×1×1 is the channel attention weight; X is the original feature map; f fc (·) is a two-layer fully connected network; GAP(·) is global pooling;

[0087] The spatial attention branch generates the spatial attention weight through a 3×3 convolution:

[0088] A spatial = σ(f conv (X)) Formula (7)

[0089] Among them, A spatial ∈R 1×H×W is the spatial attention weight; σ(·) is the Sigmoid function; fconv Convolution function;

[0090] Processing the original feature map according to the spatial attention weight and the channel attention weight, we have:

[0091] X att = X · A channel · A spatial Formula (8)

[0092] Wherein, X att ∈ R C×H×W is the weighted feature map.

[0093] S5. Concatenate the weighted sub-band feature map and the weighted feature map along the channel dimension, and obtain the final feature map through residual connection after convolution fusion.

[0094] In step S5, after the weighted sub-band feature map and the weighted feature map are concatenated, they are fused through a 1×1 convolution to obtain a fusion map:

[0095] X fused = f conv1x1 (Concat(X' wavelet , X att )) Formula (9)

[0096] Wherein, X fused is the fusion map, f conv1x1 is a 1×1 convolution; Concat is concatenation along the channel dimension;

[0097] The fusion map obtains the final feature map through residual connection:

[0098] Y = X + X fused Formula (10)

[0099] Wherein, Y is the final feature map; X is the original feature map; X fused is the fusion map.

[0100] In this embodiment, steps S2 - S5 can significantly improve the model's ability to capture high-frequency edge features and the background noise suppression effect by fusing the learnable wavelet transform and the dual attention mechanism. The learnable wavelet transform and the dual attention mechanism are composed of two parallel branches: the wavelet path realizes multi-band feature decomposition and reconstruction, and the spatial-channel attention path models global context information. The two work together to provide a robust feature representation for dense target counting.

[0101] In this embodiment,

[0102] The Haar wavelet is the simplest orthogonal wavelet, and its basis function consists of two piecewise constant functions:

[0103] Scaling Function:

[0104] Wavelet Function:

[0105] In the two - dimensional case, the Haar wavelet generates four filters through the tensor product of the scaling function and the wavelet function:

[0106] Low - pass filter (LL):

[0107] Horizontal high - pass filter (LH):

[0108] Vertical high - pass filter (HL):

[0109] Diagonal high - pass filter (HH):

[0110] Two - dimensional Haar wavelet decomposition:

[0111] For the input image The two - dimensional Haar wavelet decomposition generates four sub - bands through filtering and downsampling operations in the row and column directions:

[0112] Low - frequency sub - band (LL): X LL = L·I·L T Retains the low - frequency approximate information (main structure) of the image.

[0113] Horizontal high - frequency sub - band (LH): Captures edges and textures in the horizontal direction (such as the longitudinal texture of the sugarcane tail).

[0114] Vertical high - frequency sub - band (HL): X HL = H v ·I·L T Captures edges and textures in the vertical direction.

[0115] Diagonal high - frequency sub - band (HH): Captures edges and textures in the diagonal direction.

[0116] The size of each sub - band is 1 / 2 of the original image.

[0117] In the loss function, the gradient alignment term L wavelet Is achieved by minimizing the gradient difference between the predicted density map and the true density map in the wavelet domain:

[0118]

[0119] Where: To predict the gradient of the density map in the k-th sub-band; To be the gradient of the ground truth density map in the k-th sub-band.

[0120] S6. Obtain the final feature map of step S5, and upsample the final feature map through a decoder to restore the original size of the final feature map, and generate a density map. Extract the center points of the objects in the density map to obtain the number of sugarcane tails in the density map.

[0121] In step S6, the density map is generated through a series of convolutional operations, and thresholding or other post-processing methods can be applied to extract the center points of the objects, thereby obtaining the count of sugarcane tails.

[0122] S7. Collect the image data of sugarcane tails and generate the corresponding density map as a label to construct a dataset.

[0123] S8. Construct an initial counting model through steps S2 - S6, and train the model with the dataset to obtain an optimized counting model.

[0124] In step S8, the training method of the initial counting model is: process the dataset through steps S2 - S6, and perform loss calculation. Optimize the initial counting model according to the loss calculation to obtain an optimized counting model.

[0125] The loss calculation includes the loss in the spatial domain, the loss of gradient alignment in the wavelet domain, and the total loss:

[0126] The loss in the spatial domain is:

[0127]

[0128] where D pred is the predicted density map; D gt is the ground truth density map;

[0129] The loss of gradient alignment in the wavelet domain is:

[0130]

[0131] where is the gradient of the predicted density map in the k-th sub-band; is the gradient of the ground truth density map in the k-th sub-band;

[0132] The total loss combines the loss in the spatial domain and the loss of gradient alignment in the wavelet domain, and performs backpropagation and parameter update to optimize the model. The total loss is:

[0133]

[0134] where L WGDLis the total loss function; LH is the horizontal high-frequency subband; HL is the vertical high-frequency subband, and HH is the diagonal high-frequency subband; λ is the balance coefficient, defaulting to 0.3, which is used to control the weight of the gradient alignment term.

[0135] is the spatial domain MSE loss; is the wavelet domain gradient alignment loss.

[0136] In this embodiment, the adaptive decomposition of the target frequency band is realized through the learnable wavelet transform, the key region feature expression is strengthened by combining the frequency band attention mechanism, and the local details and global semantic information of the occluded target are captured by using the spatial-channel dual-path fusion network. Through the multi-task joint optimization strategy, the feature discrimination degree of color-similar targets, the high-frequency detail retention ability, and the counting robustness in the occlusion scenario are improved synchronously.

Claims

1. A sugarcane tail counting method based on a wavelet attention fusion mechanism, characterized in that, It includes the following steps: S1. Shoot the cane tails on the conveyor belt through a camera to obtain the original feature map of the cane tails; S2. Obtain the original feature map in step S1, and process the original feature map through wavelet convolution to decompose the original feature map into a low-frequency sub-band and a high-frequency sub-band, and obtain the sub-band feature map; S3. Obtain the sub-band feature map in step S2, generate attention weights for each sub-band feature map, and perform weighted feature enhancement on the sub-band feature map through the attention weights to obtain the weighted sub-band feature map; S4. Obtain the original feature map in step S1, and process the original feature map through spatial attention weights and channel attention weights to obtain the weighted feature map; S5. Concatenate the weighted sub-band feature map and the weighted feature map along the channel dimension, and obtain the final feature map through convolution fusion and residual connection; S6. Obtain the final feature map in step S5, and perform upsampling on the final feature map through a decoder to restore the original size of the final feature map, and generate a density map. Extract the center points of the targets from the density map to obtain the number of cane tails in the density map; S7. Collect the image data of the cane tails and generate the corresponding density map as a label to construct a dataset; S8. Construct an initial counting model through steps S2 - S6, and perform model training through the dataset to obtain an optimized counting model.

2. The sugarcane tail counting method based on the wavelet attention fusion mechanism according to claim 1, wherein: In step S1, the original image includes an RGB image size, a scene description text, and physical parameters. The RGB image size is H×W×3, where H is the height of the image; W is the width of the image; the scene description text is a scene description, and the scene description text is used to provide semantic prior information; the physical parameters include notch size, conveyor belt speed, and maximum stacking density.

3. A sugarcane tail counting method based on a wavelet attention fusion mechanism according to claim 1, characterized in that: In step S2, the learnable wavelet transform decomposes the original feature map into four sub-bands through a convolution kernel: X wavelet = DWT(X) = X * W DWT Equation (1) Among them, X wavelet is a sub-band feature map, which includes a low-frequency sub-band, a horizontal high-frequency sub-band, a vertical high-frequency sub-band, and a diagonal high-frequency sub-band; DWT is a wavelet transform; X is the original feature map; W DWT is a convolution kernel.

4. A sugarcane tail counting method based on a wavelet attention fusion mechanism according to claim 3, characterized in that: In step S2, calculate the gradient amplitude for each high-frequency sub-band to obtain the gradient alignment term, and introduce the gradient alignment constraint in the wavelet domain through the gradient alignment term: Among them, is the gradient magnitude; is the horizontal direction gradient; is the vertical direction gradient; L wavelet is the gradient alignment term; is the gradient of the predicted density map in the k-th sub-band; is the gradient of the ground truth density map in the k-th sub-band.

5. The sugarcane tail counting method based on the wavelet attention fusion mechanism according to claim 3, wherein: In step S3, set the frequency band attention mechanism. The frequency band attention mechanism generates attention weights for each sub-band feature map through global average pooling and a two-layer fully connected network: Among them, is the attention weight; σ(·) is the Sigmoid function; f fc (·) is a two-layer fully connected network; GAP(·) is global pooling; is the subband feature map; The weighted sub-band feature map is reconstructed to the original resolution through the learnable inverse wavelet transform, so there is: X' wavelet = IDWT(X wavelet ) = X wavelet * W IDWT Equation (5) Among them, IDWT is the learnable inverse wavelet transform; X' wavelet is the sub-band feature map of the original resolution; W IDWT ∈R 2×2×C×4C is the trainable transposed convolutional kernel to restore the feature space size to H×W.

6. The sugarcane tail counting method based on the wavelet attention fusion mechanism according to claim 5, wherein: In step S4, the channel attention branch generates channel attention weights through a two-layer fully connected network: Among them, A channel ∈R C×1×1 is the channel attention weight; X is the original feature map; f fc (·) is a two-layer fully connected network; GAP(·) is global pooling; The spatial attention branch generates spatial attention weights through a 3×3 convolution: A spatial = σ(f conv (X)) Equation (7) Among them, A spatial ∈R 1×H×W is the spatial attention weight; σ(·) is the Sigmoid function; f conv is the convolution function; Process the original feature map according to the spatial attention weights and channel attention weights, so there is: X att = X · A channel · A spatial Formula (8) Among them, X att ∈R C×H×W is the weighted feature map.

7. A sugarcane tail counting method based on a wavelet attention fusion mechanism according to claim 6, characterized in that: In step S5, after the weighted sub-band feature map and the weighted feature map are concatenated, they are fused through a 1×1 convolution to obtain a fusion map: X fused = f conv1x1 (Concat(X' wavelet , X att )) Formula (9) Among them, X fused is a fusion graph, f conv1x1 is a 1×1 convolution; Concat is concatenation along the channel dimension; The fusion map obtains the final feature map through residual connection: Y = X + X fused Formula (10) Among them, Y is the final feature map; X is the original feature map; X fused is the fusion map.

8. A sugarcane tail counting method based on a wavelet attention fusion mechanism according to claim 6, characterized in that: In step S8, the training method of the initial counting model is as follows: the data set is processed through steps S2 - S6, and loss calculation is performed. The initial counting model is optimized according to the loss calculation to obtain an optimized counting model.

9. The sugarcane tail counting method based on the wavelet attention fusion mechanism according to claim 8, wherein: The loss calculation includes the loss in the spatial domain, the loss of gradient alignment in the wavelet domain, and the total loss: The loss in the spatial domain is: Among them, D pred is the predicted density map; D gt is the true density map; The loss of gradient alignment in the wavelet domain is: Among them, is the gradient of the predicted density map in the k-th sub-band; is the gradient of the true density map in the k-th sub-band; The total loss combines the loss in the spatial domain and the loss of gradient alignment in the wavelet domain, and performs backpropagation and parameter update to optimize the model. The total loss is: Among them, L WGDL is the total loss function; LH is the horizontal high-frequency sub-band; HL is the vertical high-frequency sub-band, and HH is the diagonal high-frequency sub-band; λ is the balance coefficient.