Deep learning cotton growth condition accurate identification method based on ground observation image
By introducing multi-semantic space and channel attention mechanisms, new upsampling and downsampling modules and improved detection heads into the target recognition model, the problem of limited recognition capabilities of existing models in complex environments is solved, and high-precision recognition of the cotton growth stage is achieved, which meets the needs of intelligent agriculture.
Patent Information
- Application Number
- CN202510007848.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-06
AI Technical Summary
When facing a complex cotton growth environment, the existing target recognition model lacks detection accuracy and cannot effectively identify obscured objects and small objects. In complex environments, the recognition ability is limited, making it difficult to meet the accurate identification needs of different growth stages of cotton.
By introducing multi-semantic spatial and channel attention mechanisms into the backbone network, combining a brand new upsampling and downsampling module and improved detection head, a deep learning cotton growth accuracy method for ground observation images is constructed. This method improves the accuracy and efficiency of the model's identification of cotton growth stage by reducing the loss of feature details and obtaining more useful features.
It significantly improves the accuracy and efficiency of the model's identification of cotton growth stage, meets the real-time and accuracy requirements in the large-scale agricultural production environment, and provides technical support for intelligent agriculture.
Smart Images

Figure CN119942325A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision and intelligent agricultural management, and specifically relates to a method for accurately identifying cotton growth conditions by deep learning of ground observation images. Background Art
[0002] With the development of intelligent agriculture, accurate crop growth monitoring is crucial to improving agricultural production efficiency and crop management. In the cotton planting process, timely and accurate identification of cotton growth can help farmers reasonably arrange field management measures, such as fertilization, irrigation, and pest control, thereby improving yield and quality. Traditional cotton growth monitoring mostly relies on manual observation, which is time-consuming, labor-intensive, and highly subjective, and it is difficult to meet the needs of large-scale cotton field management.
[0003] In recent years, with the rapid development of deep learning and computer vision technology, crop growth monitoring technology based on image processing has gradually become an effective means. When facing the complex cotton growth environment, the existing target recognition model has problems such as insufficient detection accuracy and limited ability to recognize small targets, and cannot fully meet the needs of accurate recognition of cotton at different growth stages. Summary of the invention
[0004] The purpose of the present invention is to provide a method for accurately identifying cotton growth conditions by deep learning of ground observation images in response to the specific needs of the cotton growth stage. By introducing multi-semantic space and channel attention mechanism, new upsampling and downsampling modules and improved detection head, the model's recognition accuracy and efficiency for cotton growth stages are effectively improved, meeting the real-time and accuracy requirements in large-scale agricultural production environments, and providing technical support for the development of smart agriculture.
[0005] Aiming at the problems faced by the target recognition model, such as insufficient recognition accuracy of occluded objects and small objects, and weak recognition ability in complex scenes, the present invention constructs a method for accurately identifying cotton growth conditions based on deep learning of ground observation images by introducing multi-semantic space and channel attention mechanism, new upsampling and downsampling modules, and improved detection head, from the two dimensions of reducing the loss of feature details in model training and obtaining more useful features. The algorithm's feature extraction ability for occluded objects and small objects is improved, and at the same time, multiple semantic information is obtained to solve the problems of low recognition accuracy of occluded objects and small objects and limited recognition ability in complex environments.
[0006] The main process of the method includes: first, in the backbone network (backbone), the multi-semantic space and channel attention module (C2f-MSCA) is used to divide the cotton feature map into two groups of sub-features on average, and the spatial attention mechanism based on depthwise separable 1D convolution and the channel attention mechanism based on global average pooling and global maximum pooling are respectively used to calculate the multi-semantic space and channel attention mechanism features of the cotton growth stage, and the two groups of sub-features are spliced together to obtain the feature output of the multi-semantic space and channel attention mechanism. Subsequently, the Dysample module is used in the feature fusion network (neck) to shuffle the feature map, and then the offset is generated by linear projection, and the generated offset is added to the original ground observation image to generate a sampling set. Using the generated sampling set, the grid_sample function is used to perform bilinear interpolation on the input features to achieve upsampling; using the AWD module, the high and low frequency subbands of the features are separated by filtering using Haar wavelet low-pass filter and Haar wavelet high-pass filter in the horizontal direction, and then the high and low frequency subbands are separated in the vertical direction, and the high and low frequency subbands are added according to the weights to obtain the down-sampled output. Finally, a high-frequency information enhancement mechanism is introduced in the detection head. Different convolution operations are performed on the high and low-frequency sub-bands separated by wavelet transform, and the high-frequency information is enhanced. The processed high and low-frequency sub-bands are reconstructed into feature maps with the same size as the input using inverse wavelet transform, and the decoupling head is used to accurately identify the cotton growth conditions based on the feature maps.
[0007] To achieve the above purpose, the technical solution of the present invention is: a method for accurately identifying cotton growth conditions by deep learning of ground observation images, comprising:
[0008] Adding multi-semantic space and channel attention mechanism to the backbone network, we can obtain the characteristics of multi-semantic space and channel attention mechanism in cotton growth stage.
[0009] Use the Dysample upsampling module and AWD downsampling module in the feature fusion network to achieve feature fusion of different scales;
[0010] Construct a detection head that introduces a high-frequency information enhancement mechanism to perform high-precision identification of cotton growth conditions.
[0011] In one embodiment of the present invention, the specific implementation method of adding a multi-semantic space and channel attention mechanism to the backbone network to obtain the characteristics of the multi-semantic space and channel attention mechanism of the cotton growth stage is: using the multi-semantic space and channel attention module C2f-MSCA in the backbone network backbone, dividing the cotton feature map into two groups of sub-features on average, and using a spatial attention mechanism based on deeply separable 1D convolution and a channel attention mechanism based on global average pooling and global maximum pooling respectively, calculating the multi-semantic space and channel attention mechanism characteristics of the cotton growth stage, splicing the two groups of sub-features together, and obtaining the feature output of the multi-semantic space and channel attention mechanism.
[0012] In one embodiment of the present invention, the specific implementation method of using the Dysample upsampling module and the AWD downsampling module in the feature fusion network to achieve feature fusion of different scales is as follows: a multi-scale feature fusion method is used in the feature fusion network neck to fuse feature maps from different convolutional layers in the channel dimension; the Dysample module is used to perform channel shuffling on the feature map, and then a linear projection is used to generate an offset, and the generated offset is added to the initial ground observation image to generate a sampling set; using the generated sampling set, the grid_sample function is used to perform bilinear interpolation on the input features to achieve upsampling; using the AWD module, a Haar wavelet low-pass filter and a Haar wavelet high-pass filter are used to separate the high and low frequency sub-bands of the features in the horizontal direction, and then four sub-bands are separated in the vertical direction, and the four sub-bands are added according to the weights to obtain the down-sampled output.
[0013] In one embodiment of the present invention, the specific implementation method of constructing a detection head that introduces a high-frequency information enhancement mechanism to perform high-precision cotton growth condition identification is as follows: a high-frequency information enhancement mechanism is introduced in the detection head, different convolution operations are performed on the high and low frequency sub-bands separated by the wavelet transform, and the high-frequency information is enhanced, the processed high and low frequency sub-bands are reconstructed into a feature map consistent with the input size using an inverse wavelet transform, and the decoupling head is used to accurately identify the cotton growth condition on the feature map.
[0014] In one embodiment of the present invention, the multi-semantic space and channel attention mechanism are added to the backbone network to obtain the characteristics of the multi-semantic space and channel attention mechanism of the cotton growth stage, which specifically includes the following steps:
[0015] Step S11: Divide the initial feature map with dimensions C×H×W into two groups with dimensions Sub-feature map, C represents the total number of channels, H and W represent the height and width of the feature map respectively;
[0016] Step S12: Sub-feature graph X 1 Decomposing along the height and width dimensions, we create two dimensions: and One-way feature map of W: and Then each unidirectional feature map is divided into 4 channels The independent sub-feature graphs are used, and the depth-separable 1D convolution with kernel sizes of 3, 5, 7, and 9 is used to obtain different semantic features and reduce the number of model parameters; the results after convolution are reconstructed to obtain a dimension of The vertical attention feature map and dimension are W's horizontal attention feature map, and finally the vertical attention feature map and the horizontal attention feature map and the sub-feature map X 1 Multiplying together gives the dimension The spatial feature map X S , the calculation of the spatial feature map is as follows:
[0017]
[0018]
[0019] X S =Atten H ×Atten W ×X 1
[0020] In the formula, represents a one-dimensional depth-wise separable convolution, k i represents the size of the kernel, Represents the result in the height dimension, represents the result on the width dimension, σ(·) represents the Sigmoid function, GN represents group normalization, and X S Represents the spatial feature map, Atten H Represents the vertical attention feature map, Atten W Horizontal attention feature map;
[0021] Step S13: Sub-feature graph X 2 After the maximum pooling layer and the average pooling layer, two sets of feature maps are obtained, and the dimensions are Add the results of the two sets of feature maps after passing through the weight sharing network, and the dimension is Channel attention feature map, and finally combine the channel attention feature map with the sub-feature map X 2 Multiplying together gives the dimension The channel feature map is calculated as follows:
[0022] Atten C =σ(MLP(AvgPool(X 2)+MLP(MaxPool(X 2 )))
[0023] X C =Atten C ×X 2
[0024] Where, X C Represents the channel feature map, Atten C represents the channel attention feature map, MLP represents the weight sharing network, AvgPool represents the average pooling layer, and MaxPool represents the maximum pooling layer;
[0025] Step S14: X S and X C The concatenation is performed along the channel dimension. The concatenation is calculated as follows:
[0026] X=Concat(X S ,X C )
[0027] Where Concat represents the cat function in PyTorch (an open source Python machine learning library).
[0028] In one embodiment of the present invention, in step S11, the initial feature map with dimensions C×H×W is evenly divided into two groups with dimensions The sub-feature map is calculated as follows:
[0029]
[0030] Where, X i represents the i-th sub-feature map, where i∈[1,2].
[0031] In one embodiment of the present invention, the Dysample up-sampling module and the AWD down-sampling module are used in the feature fusion network to achieve feature fusion of different scales, which specifically includes the following steps:
[0032] Step S21: perform channel shuffling on the feature map of dimension C×H×W, where C represents the total number of channels, H and W represent the height and width of the feature map, respectively, and group the feature map into groups of 4 channels: The 4-channel features 4×H×W are re-divided into 2H×2W blocks, and these divided blocks are rearranged in the spatial dimension to adjust their size to
[0033] Step S22, use linear projection to project the feature map to 8×2H×2W to generate an offset, multiply the offset by 0.25 to ensure that the offset just meets the theoretical boundary condition between overlap and non-overlap, and add the generated offset to the initial grid to generate a sampling set with a dimension of 8×2H×2W;
[0034] Step S23, the filtering calculation formula in the horizontal direction is as follows:
[0035] f L (x,y)=∑ k f(x,k)L[y-2k]
[0036] f H (x,y)=∑ k f(x,k)H[y-2k]
[0037] where f L is the result of low-pass filtering, f H is the result of high-pass filtering, f represents the input image, x represents the row index of the image, y represents the column index of the image, k is the index in the convolution operation, L is the low-pass filter, and H is the high-pass filter;
[0038] The low-pass filter L and high-pass filter H are defined as:
[0039]
[0040] Step S24, the calculation formula for filtering in the vertical direction is as follows:
[0041] LL(x,y)=∑ k f L (k,y)L[x-2k]
[0042] LH(x,y)=∑ k f L (k,y)H[x-2k]
[0043] HL(x,y)=∑ k f H (k,y)L[x-2k]
[0044] HH(x,y)=∑ k f H (k,y)H[x-2k]
[0045] Among them, LL represents the approximate progeny, LH represents the horizontal high-frequency progeny, HL represents the vertical high-frequency progeny, and HH represents the diagonal high-frequency progeny;
[0046] Step S25, the calculation formula for adding the downsampling results according to the weights is as follows:
[0047] X=W1 ×LL+W 2 ×LH+W 3 ×HL+W 4 ×HH
[0048] Where W i Represents the weight, i=1,2,3,4, and X represents the result after addition.
[0049] In one embodiment of the present invention, the construction of a detection head that introduces a high-frequency information enhancement mechanism to perform high-precision cotton growth condition identification specifically includes the following steps:
[0050] Step S31, the high frequency enhancement module uses the calculation method of step S24 and step S25 to obtain the characteristic low frequency, horizontal high frequency, vertical high frequency and diagonal high frequency;
[0051] Step S32: using the characteristics of horizontal, vertical and diagonal details, perform 1×3 horizontal convolution, 3×1 vertical convolution and 1×3 and 3×1 horizontal and vertical combined convolution on the three generations respectively, and perform 3×3 ordinary convolution on the low-frequency generation;
[0052] Step S33, multiply the three high-frequency sub-generations by an enhancement factor of 1.5;
[0053] Step S34: Set each dimension to The sub-bands are upsampled to obtain four sub-bands with dimensions of C×H×W. The upsampled sub-bands are low-pass and high-pass filtered respectively, and the results obtained after filtering all sub-bands are added together. The calculation formula is as follows:
[0054] f LL (x,y)=∑ m ∑ n LL(m,n)·L(x-2m)·L(y-2n)
[0055] f LH (x,y)=∑ m ∑ n LH(m,n)·L(x-2m)·H(y-2n)
[0056] f HL (x,y)=∑ m ∑ n HL(m,n)·H(x-2m)·L(y-2n)
[0057] f HH (x,y)=∑ m ∑ n HH(m,n)·H(x-2m)·L(y-2n)
[0058] The overall reconstruction formula is:
[0059] f reconstructed (m,n)=f LL (m,n)+f LH (m,n)+f HL (m,n)+f HH (m,n)
[0060] In the formula, f LL (m,n) is the low frequency part, f LH (m,n),f HL (m,n),f HH (m,n) is the high frequency part, x is the row index of the image, y is the column index of the image, m and n are the indexes of the convolution operation, and f reconstructed (m,n) is the result after inverse wavelet transform reconstruction.
[0061] The present invention also provides a deep learning cotton growth condition accurate identification system based on ground observation images, comprising a memory, a processor, and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, the method steps described above can be implemented.
[0062] The present invention also provides a computer-readable storage medium, on which computer program instructions that can be executed by a processor are stored. When the processor executes the computer program instructions, the method steps described above can be implemented.
[0063] Compared with the prior art, the present invention has the following beneficial effects: BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 is a schematic diagram of the network model structure in an embodiment of the present invention;
[0065] Figure 2 This is a c2f-MSCA structure diagram in an embodiment of the present invention.
[0066] Figure 3 This is a Dysample structure diagram in an embodiment of the present invention.
[0067] Figure 4 This is the AWD structure diagram in the embodiment of the present invention.
[0068] Figure 5 This is a structural diagram of a high frequency enhancement module in an embodiment of the present invention.
[0069] Figure 6 It is a cotton growth stage recognition diagram based on YOLOv8 in an embodiment of the present invention;
[0070] Figure 7 This is a cotton growth condition identification diagram of the method in the embodiment of the present invention. DETAILED DESCRIPTION
[0071] The technical solution of the present invention is described in detail below in conjunction with the accompanying drawings.
[0072] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present application belongs.
[0073] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.
[0074] The present invention provides a method for accurately identifying cotton growth conditions by deep learning of ground observation images, comprising:
[0075] Adding multi-semantic space and channel attention mechanism to the backbone network, we can obtain the characteristics of multi-semantic space and channel attention mechanism in cotton growth stage.
[0076] Use the Dysample upsampling module and AWD downsampling module in the feature fusion network to achieve feature fusion of different scales;
[0077] Construct a detection head that introduces a high-frequency information enhancement mechanism to perform high-precision identification of cotton growth conditions.
[0078] The following is the specific implementation process of the present invention.
[0079] In order to verify the effectiveness of the method, a series of experiments were carried out using the cotton growth stage dataset shown in Table 1. The differences in recognition effects between the YOLOv8 algorithm and the method of the present invention were compared, and the recall rate (R), precision rate (P) and mean average precision (mAP) were used to evaluate the recognition results of the algorithm, and the parameter amount (Params) and floating point operations per second (FLOPS) were used to evaluate the recognition speed of the algorithm.
[0080] Table 1
[0081]
[0082] like Figures 1 to 5 As shown, this embodiment provides a method for accurately identifying cotton growth conditions by deep learning of ground observation images, including the following steps:
[0083] Step S1: Model training
[0084] The acquired cotton ground observation images were cropped and normalized; the cotton location information was drawn on the processed images using the roboflow website; the processed images and their cotton location information were input into the recognition network for weight parameter training, and the model with the best performance during the training process was selected as the final model.
[0085] Step S2: Model evaluation
[0086] The final model is evaluated using the test set data to ensure that the model can make accurate recognition in the real world.
[0087] Step S3: Cotton growth stage identification
[0088] The final model is used to recognize images collected by ground-based cameras and output growth stage information.
[0089] As shown in Table 2, the evaluation results of the YOLOv8 algorithm and this example method for recognition on the cotton growth stage dataset are shown. The evaluation indicators mAP, P, and R are used to measure the quality of the recognition results. The closer the value is to 1, the better the recognition result is. The evaluation indicators Params and FLOPS are used to measure the recognition speed. The smaller the value is, the faster the recognition speed is.
[0090] Table 2
[0091]
[0092] By analyzing the results in Table 2, it can be observed that this example method has a significant improvement in recognition results and recognition speed compared with the original model.
[0093] like Figure 6 and Figure 7 As shown in the figure, the cotton data of an oasis farm in Korla, Xinjiang is identified. It can be seen that the algorithm in this example has significantly improved the number of identifications and the recognition accuracy.
[0094] The present invention also provides a deep learning cotton growth condition accurate identification system based on ground observation images, comprising a memory, a processor, and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, the method steps described above can be implemented.
[0095] The present invention also provides a computer-readable storage medium, on which computer program instructions that can be executed by a processor are stored. When the processor executes the computer program instructions, the method steps described above can be implemented.
[0096] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.
[0097] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0098] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0099] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0100] The above is only a preferred embodiment of the present invention, and does not limit the present invention in other forms. Any technician familiar with the profession may use the above disclosed technical content to change or modify it into an equivalent embodiment with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the technical solution of the present invention still belongs to the protection scope of the technical solution of the present invention.
Claims
1. A method for accurately identifying cotton growth conditions by deep learning of ground observation images, characterized in that: include: Adding multi-semantic space and channel attention mechanism to the backbone network, we can obtain the characteristics of multi-semantic space and channel attention mechanism in cotton growth stage. Use the Dysample upsampling module and AWD downsampling module in the feature fusion network to achieve feature fusion of different scales; Construct a detection head that introduces a high-frequency information enhancement mechanism to perform high-precision identification of cotton growth conditions.
2. The method for accurately identifying cotton growth conditions by deep learning of ground observation images according to claim 1, characterized in that: The specific implementation method of adding the multi-semantic space and channel attention mechanism to the backbone network to obtain the characteristics of the multi-semantic space and channel attention mechanism in the cotton growth stage is: using the multi-semantic space and channel attention module C2f-MSCA in the backbone network backbone, dividing the cotton feature map into two groups of sub-features on average, and respectively using the spatial attention mechanism based on deep separable 1D convolution and the channel attention mechanism based on global average pooling and global maximum pooling to calculate the multi-semantic space and channel attention mechanism characteristics of the cotton growth stage, splicing the two groups of sub-features together, and obtaining the feature output of the multi-semantic space and channel attention mechanism.
3. The method for accurately identifying cotton growth conditions by deep learning of ground observation images according to claim 1, characterized in that: The specific implementation method of using the Dysample up-sampling module and the AWD down-sampling module in the feature fusion network to realize the fusion of features of different scales is as follows: a multi-scale feature fusion method is used in the feature fusion network neck to fuse feature maps from different convolutional layers in the channel dimension; the Dysample module is used to perform channel shuffling on the feature map, and then a linear projection is used to generate an offset, and the generated offset is added to the initial ground observation image to generate a sampling set; using the generated sampling set, the grid_sample function is used to perform bilinear interpolation on the input features to achieve upsampling; using the AWD module, a Haar wavelet low-pass filter and a Haar wavelet high-pass filter are used to separate the high and low frequency sub-bands of the features in the horizontal direction, and then four sub-bands are separated in the vertical direction, and the four sub-bands are added according to the weights to obtain the down-sampled output.
4. The method for accurately identifying cotton growth conditions by deep learning of ground observation images according to claim 3 is characterized in that: The specific implementation method of constructing a detection head that introduces a high-frequency information enhancement mechanism to perform high-precision cotton growth condition identification is as follows: a high-frequency information enhancement mechanism is introduced into the detection head, different convolution operations are performed on the high and low frequency sub-bands separated by wavelet transform, and the high-frequency information is enhanced, the processed high and low frequency sub-bands are reconstructed into a feature map with the same size as the input using an inverse wavelet transform, and the decoupling head is used to accurately identify the cotton growth condition on the feature map.
5. The method for accurately identifying cotton growth conditions by deep learning of ground observation images according to claim 1 or 2, characterized in that: The multi-semantic space and channel attention mechanism are added to the backbone network to obtain the characteristics of the multi-semantic space and channel attention mechanism in the cotton growth stage, which specifically includes the following steps: Step S11: Divide the initial feature map with dimensions C×H×W into two groups with dimensions Sub-feature map, C represents the total number of channels, H and W represent the height and width of the feature map respectively; Step S12: Sub-feature graph X 1 Decomposing along the height and width dimensions, we create two dimensions: and One-way feature map of : and Then each unidirectional feature map is divided into 4 channels The independent sub-feature graphs are used, and the depth-separable 1D convolution with kernel sizes of 3, 5, 7, and 9 is used to obtain different semantic features and reduce the number of model parameters; the results after convolution are reconstructed to obtain a dimension of The vertical attention feature map and dimension are W's horizontal attention feature map, and finally the vertical attention feature map and the horizontal attention feature map and the sub-feature map X 1 Multiplying together gives the dimension The spatial feature map X S , the calculation of the spatial feature map is as follows: In the formula, represents a one-dimensional depth-wise separable convolution, k i represents the size of the kernel, Represents the result in the height dimension, represents the result on the width dimension, σ(·) represents the Sigmoid function, GN represents group normalization, and X S Represents the spatial feature map, Atten H Represents the vertical attention feature map, Atten W Horizontal attention feature map; Step S13: Sub-feature graph X 2 After the maximum pooling layer and the average pooling layer, two sets of feature maps are obtained, and the dimensions are Add the results of the two sets of feature maps after passing through the weight sharing network, and the dimension is Channel attention feature map, and finally combine the channel attention feature map with the sub-feature map X 2 Multiplying together gives the dimension The channel feature map is calculated as follows: Atten C =σ(MLP(AvgPool(X 2 )+MLP(MaxPool(X 2 ))) X C =Atten C ×X 2 Where, X C Represents the channel feature map, Atten C represents the channel attention feature map, MLP represents the weight sharing network, AvgPool represents the average pooling layer, and MaxPool represents the maximum pooling layer; Step S14: X S and X C The concatenation is performed along the channel dimension. The concatenation is calculated as follows: X=Concat(X S ,X C ) Where Concat represents the cat function in the open source Python machine learning library PyTorch.
6. The method for accurately identifying cotton growth conditions by deep learning of ground observation images according to claim 5, characterized in that: In step S11, the initial feature map with dimensions C×H×W is evenly divided into two groups with dimensions The sub-feature map is calculated as follows: Where, X i represents the i-th sub-feature map, where i∈[1,2].
7. The method for accurately identifying cotton growth conditions by deep learning of ground observation images according to claim 1 or 3, characterized in that: The Dysample up-sampling module and the AWD down-sampling module are used in the feature fusion network to achieve feature fusion of different scales, which specifically includes the following steps: Step S21: perform channel shuffling on the feature map of dimension C×H×W, where C represents the total number of channels, H and W represent the height and width of the feature map, respectively, and group the feature map into groups of 4 channels: The 4-channel features 4×H×W are re-divided into 2H×2W blocks, and these divided blocks are rearranged in the spatial dimension to adjust their size to Step S22, linearly project the feature map to 8×2H×2W to generate an offset, multiply the offset by 0.25 to ensure that the offset just meets the theoretical boundary condition between overlap and non-overlap, and add the generated offset to the initial grid to generate a sampling set with a dimension of 8×2H×2W; Step S23, the filtering calculation formula in the horizontal direction is as follows: f L (x,y)=∑ k f(x,k)L[y-2k] f H (x,y)=∑ k f(x,k)H[y-2k] where f L is the result of low-pass filtering, f H is the result of high-pass filtering, f represents the input image, x represents the row index of the image, y represents the column index of the image, k is the index in the convolution operation, L is the low-pass filter, and H is the high-pass filter; The low-pass filter L and high-pass filter H are defined as: Step S24, the calculation formula for filtering in the vertical direction is as follows: LL(x,y)=∑ k f L (k,y)L[x-2k] LH(x,y)=∑ k F L (k,y)H[x-2k] <h2 style=";text-align:left;direction:ltr">HL(x,y)=∑<h2 style=";text-align:left;direction:ltr"> k <h2 style=";text-align:left;direction:ltr"> f<h2 style=";text-align:left;direction:ltr"> H <h2 style=";text-align:left;direction:ltr"> (k,y)L[x-2k] HH(x,y)=∑ k f H (k,y)H[x-2k] Among them, LL represents the approximate progeny, LH represents the horizontal high-frequency progeny, HL represents the vertical high-frequency progeny, and HH represents the diagonal high-frequency progeny; Step S25, the calculation formula for adding the downsampling results according to the weights is as follows: X=W1×LL+W2×LH+W3×HL+W4×HH Where W i Represents the weight, i=1,2,3,4, and X represents the result after addition.
8. The method for accurately identifying cotton growth conditions by deep learning of ground observation images according to claim 7, characterized in that: The construction of the detection head introducing the high-frequency information enhancement mechanism to perform high-precision cotton growth condition identification specifically includes the following steps: Step S31, the high frequency enhancement module uses the calculation method of step S24 and step S25 to obtain the characteristic low frequency, horizontal high frequency, vertical high frequency and diagonal high frequency; Step S32: using the characteristics of horizontal, vertical and diagonal details, perform 1×3 horizontal convolution, 3×1 vertical convolution and 1×3 and 3×1 horizontal and vertical combined convolution on the three generations respectively, and perform 3×3 ordinary convolution on the low-frequency generation; Step S33, multiply the three high-frequency sub-generations by an enhancement factor of 1.5; Step S34: Set each dimension to The sub-bands are upsampled to obtain four sub-bands with dimensions of C×H×W, and low-pass and high-pass filtering operations are performed on the upsampled sub-bands respectively, and the results obtained after filtering all sub-bands are added; The calculation formula is as follows: f LL (x,y)=∑ m ∑ n LL(m,n)·L(x-2m)·L(y-2n) f LH (x,y)=∑ m ∑ n LH(m,n)·L(x-2m)·H(y-2n) <h2 style=";text-align:left;direction:ltr">f<h2 style=";text-align:left;direction:ltr"> HL <h2 style=";text-align:left;direction:ltr"> (x,y)=∑<h2 style=";text-align:left;direction:ltr"> m <h2 style=";text-align:left;direction:ltr"> ∑<h2 style=";text-align:left;direction:ltr"> n <h2 style=";text-align:left;direction:ltr"> HL(m,n)·H(x-2m)·L(y-2n) f HH (x,y)=∑ m ∑ n HH(m,n)·H(x-2m)·L(y-2n) The overall reconstruction formula is: f reconstructed (m,n)=f LL (m,n)+f LH (m,n)+f HL (m,n)+f HH (m,n) In the formula, f LL (m,n) is the low frequency part, f LH (m,n),f HL (m,n),f HH (m,n) is the high frequency part, x is the row index of the image, y is the column index of the image, m and n are the indexes of the convolution operation, and f reconstructed (m,n) is the result after inverse wavelet transform reconstruction.
9. A deep learning cotton growth condition accurate recognition system based on ground observation images, characterized in that: The method comprises a memory, a processor and computer program instructions stored in the memory and executable by the processor. When the processor executes the computer program instructions, the method steps as claimed in any one of claims 1 to 8 can be implemented.
10. A computer-readable storage medium storing computer program instructions that can be executed by a processor, wherein when the processor executes the computer program instructions, the method steps according to any one of claims 1 to 8 can be implemented.