A Remote Sensing Identification and Extraction Method and System for White Lotus Seedpods Based on the Optimized YOLOX Network

By introducing the CA attention mechanism and ASFF feature fusion layer into the YOLOX network, the YOLOX target detection network is optimized, and the error detection problem in low recognition accuracy and complex environments in lotus pod recognition are solved, achieving higher recognition accuracy and operation efficiency.

CN118736402BActive Publication Date: 2025-07-01江西德都食品科技有限公司
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410716546.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-04
Publication Date
2025-07-01
Estimated Expiration
2044-06-04

AI Technical Summary

Technical Problem

The prior art has problems such as low recognition accuracy, complex training, and insufficient recognition of small targets in lotus pod recognition, especially in complex environments, which are prone to false detection and missed detection.

Method used

By adding CA attention mechanism and ASFF feature fusion layer to the YOLOX network, the YOLOX target detection network is optimized to improve the detection sensitivity and identification accuracy of small pod targets.

Benefits of technology

It significantly improves the accuracy and recall of lotus pod recognition, improves the robustness and generalization capabilities of the model, and speeds up the inference speed and optimizes memory footprint.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118736402B_ABST
    Figure CN118736402B_ABST
Patent Text Reader

Abstract

The present invention proposes a remote sensing identification and extraction method and system for white lotus seedpods based on an optimized YOLOX network. Among them, the method includes: collecting seedpod images, performing data augmentation and label annotation on them to obtain a seedpod image dataset; dividing the seedpod image dataset into a training set and a test set; constructing a YOLOX seedpod detection network model; by adding a CA attention mechanism and an ASFF feature fusion layer to the YOLOX seedpod detection network model, an optimized YOLOX object detection network is obtained; using the training set to train the optimized YOLOX object detection network; applying the test set to test and evaluate the trained optimized YOLOX object detection network to obtain the seedpod image detection result. The solution proposed by the present invention can significantly enhance the detection sensitivity and recognition accuracy of the network for small seedpod targets by integrating the CA attention mechanism and the ASFF feature fusion module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of remote sensing identification, and particularly relates to a method and system for remote sensing identification and extraction of white lotus seedpods based on an optimized YOLOX network. Background Art

[0002] The lotus seedpod is a special aquatic economic crop, which is characterized by strong production regionality, wide planting area, long life cycle, rich nutritional value and high economic value. It has a large market demand, broad prospects and considerable economic benefits. However, at present, the monitoring of the planting distribution, area and yield of lotus seedpods mainly relies on reported statistics and field surveys, which are not only time-consuming and laborious, but also difficult to comprehensively reflect its planting area, distribution information and dynamic changes. The unmanned aerial vehicle (UAV) remote sensing technology has a series of advantages such as a short data update cycle and an economical and effective monitoring method. Applying it to the rapid monitoring of the distribution and dynamic changes of Chinese herbal medicine planting can provide a decision-making basis for relevant departments to guide the reasonable planting of lotus seedpods. Establishing an efficient UAV image lotus seedpod recognition model is the primary task for realizing the remote sensing monitoring of lotus seedpods. In recent years, due to the strong ability of deep learning to extract high-dimensional features of targets, it has been widely applied to related problems.

[0003] At present, in the automated harvesting operations or fruit maturity prediction of solanaceous fruits such as strawberries, kiwifruits, apples, grapes and cucumbers, target detection models are often used to identify the fruits. However, solanaceous fruits hang directly below the branches and usually show a regular natural downward trend, so their recognition is relatively easy to achieve.

[0004] The growth of lotus seedpods is different from that of the above-mentioned solanaceous fruits. The heights, orientations and bending degrees of lotus seedpods during development are different, and the maturity periods are also inconsistent. Under the influence of wind, there is also an irregular swinging phenomenon of lotus seedpods. Therefore, when identifying lotus seedpods from a top-down perspective, the growth posture needs to be considered. Moreover, the maturity of lotus seedpods is a gradual process, and the characteristics of lotus seedpods at different maturity stages are not significantly different: fresh lotus seedpods are light green, lotus leaves are emerald green, lotus flower colors are reddish or white, the color difference between lotus seedpods and lotus leaves is not obvious, and they are similar in shape and significantly different in color from lotus flowers; at the fully mature stage, lotus seedpods are bluish green, the lotus seeds slightly protrude from the surface of the lotus seedpods, and the posture of the lotus seedpods droops; at the withered and mature stage, lotus seedpods are brownish, the lotus seeds are black, and the posture is upright. In addition, since the lotus seedpod is an aquatic plant growing in a complex lotus pond environment, there are a large number of obstacles such as lotus leaves and lotus flowers, and the phenomena of overlap and occlusion are serious, and false detection and missed detection are very likely to occur. In summary, the identification of lotus seedpods in UAV images needs to consider many factors to achieve efficient and accurate identification of lotus seedpods.

[0005] Although some current studies have used the PCNN model, YOLO V2 model, or YOLO V5 model to achieve the recognition of lotus pods, the training of the PCNN model is complex and its accuracy is low. The YOLO V2 model has insufficient recognition accuracy for small targets, and the YOLO V5 model mainly improves the training speed and deployment of the model, without sufficient improvement in recognition accuracy. For the target of lotus pod recognition, targeted improvements are still needed. Summary of the Invention

[0006] To solve the above technical problems, the present invention proposes a technical solution for a remote sensing recognition and extraction method of white lotus pods based on an optimized YOLOX network to solve the above technical problems.

[0007] The first aspect of the present invention discloses a remote sensing recognition and extraction method of white lotus pods based on an optimized YOLOX network, and the method includes:

[0008] Step S1: Collect lotus pod images, perform data augmentation and label annotation on them to obtain a lotus pod image dataset; divide the lotus pod image dataset into a training set and a test set;

[0009] Step S2: Construct a YOLOX lotus pod detection network model by sequentially connecting an Input layer, a backbone network with CSPDarkNet53 as the basic architecture, a feature enhancement network adopting a composite structure scheme of FPN+PAN, and an output prediction layer;

[0010] Step S3: Add a CA attention mechanism and an ASFF feature fusion layer to the YOLOX lotus pod detection network model to obtain an optimized YOLOX target detection network;

[0011] Step S4: Use the training set to train the optimized YOLOX target detection network;

[0012] Step S5: Apply the test set to test and evaluate the trained optimized YOLOX target detection network to obtain the lotus pod image detection result.

[0013] According to the method of the first aspect of the present invention, in the step S3, the adding a CA attention mechanism and an ASFF feature fusion layer to the YOLOX lotus pod detection network model to obtain an optimized YOLOX target detection network includes:

[0014] Add a CA attention mechanism between the backbone network and the feature enhancement network, and add an ASFF feature fusion layer between the feature enhancement network and the output prediction layer to obtain an optimized YOLOX target detection network.

[0015] According to the method of the first aspect of the present invention, in the step S3, adding the CA attention mechanism between the backbone network and the feature enhancement network includes:

[0016] Attach CA attention modules to the output feature maps of the second, third, and fourth CSP modules of the backbone network respectively to obtain attention mechanism feature maps; input the attention mechanism feature maps into the feature enhancement network.

[0017] According to the method of the first aspect of the present invention, in the step S3, attaching CA attention modules to the output feature maps respectively to obtain attention mechanism feature maps includes:

[0018] Input the output feature maps into a residual module to obtain residual features;

[0019] Perform average pooling in the X direction and average pooling in the Y direction on the residual features to obtain an X-direction pooled feature map and a Y-direction pooled feature map;

[0020] Concatenate the X-direction pooled feature map and the Y-direction pooled feature map, and obtain a first convolutional feature map through a convolutional layer;

[0021] Perform batch normalization and non-linear transformation on the first convolutional feature map to obtain a non-linear feature map;

[0022] Apply convolutional operations to the non-linear feature map in the horizontal direction and the vertical direction respectively to obtain a second convolutional feature map and a third convolutional feature map;

[0023] Apply the Sigmoid function to the second convolutional feature map and the third convolutional feature map respectively to obtain a first function feature map and a second function feature map;

[0024] Multiply the first function feature map, the second function feature map and the residual features to obtain an attention mechanism feature map.

[0025] The second aspect of the present invention discloses a white lotus seedpod remote sensing identification and extraction system based on an optimized YOLOX network, and the system includes:

[0026] A first processing module, configured to collect lotus seedpod images, perform data augmentation and label annotation on them to obtain a lotus seedpod image dataset; divide the lotus seedpod image dataset into a training set and a test set;

[0027] A second processing module, configured to construct a YOLOX lotus seedpod detection network model by sequentially connecting an Input layer, a backbone network with CSPDarkNet53 as the basic architecture, a feature enhancement network adopting a composite structure scheme of FPN+PAN, and an output prediction layer;

[0028] The third processing module is configured to obtain an optimized YOLOX object detection network by adding a CA attention mechanism and an ASFF feature fusion layer to the YOLOX lotus pod detection network model;

[0029] The fourth processing module is configured to train the optimized YOLOX object detection network by referring to the training set;

[0030] The fifth processing module is configured to test and evaluate the trained optimized YOLOX object detection network using the test set to obtain the lotus pod image detection result.

[0031] For the system according to the second aspect of the present invention, the third processing module is specifically configured to, the obtaining an optimized YOLOX object detection network by adding a CA attention mechanism and an ASFF feature fusion layer to the YOLOX lotus pod detection network model includes:

[0032] An optimized YOLOX object detection network is obtained by adding a CA attention mechanism between the backbone network and the feature enhancement network and adding an ASFF feature fusion layer between the feature enhancement network and the output prediction layer.

[0033] For the system according to the second aspect of the present invention, the third processing module is specifically configured to, the adding a CA attention mechanism between the backbone network and the feature enhancement network includes:

[0034] CA attention modules are respectively attached to the output feature maps of the second, third, and fourth CSP modules of the backbone network to obtain attention mechanism feature maps; the attention mechanism feature maps are input into the feature enhancement network.

[0035] For the system according to the second aspect of the present invention, the third processing module is specifically configured to, the respectively attaching CA attention modules to the output feature maps to obtain attention mechanism feature maps includes:

[0036] The output feature map is input into a residual module to obtain a residual feature;

[0037] Average pooling in the X direction and average pooling in the Y direction are performed on the residual feature to obtain an X-direction pooled feature map and a Y-direction pooled feature map;

[0038] The X-direction pooled feature map and the Y-direction pooled feature map are concatenated and a first convolutional feature map is obtained through a convolutional layer;

[0039] Batch normalization and non-linear transformation are performed on the first convolutional feature map to obtain a non-linear feature map;

[0040] Apply convolution operations to the non-linear feature map in the horizontal and vertical directions respectively to obtain a second convolutional feature map and a third convolutional feature map;

[0041] Apply the Sigmoid function to the second convolutional feature map and the third convolutional feature map respectively to obtain a first function feature map and a second function feature map;

[0042] Multiply the first function feature map, the second function feature map and the residual feature to obtain an attention mechanism feature map.

[0043] A third aspect of the present invention discloses an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the steps in a method for remote sensing identification and extraction of white lotus seedpods based on an optimized YOLOX network according to any one of the first aspects of the present disclosure are implemented.

[0044] A fourth aspect of the present invention discloses a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the steps in a method for remote sensing identification and extraction of white lotus seedpods based on an optimized YOLOX network according to any one of the first aspects of the present disclosure are implemented.

[0045] In summary, the solution proposed by the present invention can significantly enhance the detection sensitivity and recognition accuracy of the network for small lotus seedpod targets by incorporating the CA (Channel Attention) attention mechanism and the ASFF (Adaptive Spatial Feature Fusion) feature fusion module. Description of the Drawings

[0046] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0047] Figure 1 It is a flowchart of a method for remote sensing identification and extraction of white lotus seedpods based on an optimized YOLOX network according to an embodiment of the present invention;

[0048] Figure 2 It is a network structure diagram of the basic YOLOX according to an embodiment of the present invention;

[0049] Figure 3 It is a network structure diagram of the optimized YOLOX according to an embodiment of the present invention;

[0050] Figure 4It is a network structure diagram of the CA attention mechanism according to an embodiment of the present invention;

[0051] Figure 5 It is a line graph of the change of key indicators during the model training process according to an embodiment of the present invention;

[0052] Figure 6 It is a detection result diagram according to an embodiment of the present invention;

[0053] Figure 7 It is a structure diagram of a white lotus and lotus pod remote sensing identification and extraction system based on an optimized YOLOX network according to an embodiment of the present invention;

[0054] Figure 8 It is a structure diagram of an electronic device according to an embodiment of the present invention. Detailed implementation manners

[0055] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0056] The first aspect of the present invention discloses a method for remote sensing identification and extraction of white lotus and lotus pods based on an optimized YOLOX network. Figure 1 It is a flowchart of a method for remote sensing identification and extraction of white lotus and lotus pods based on an optimized YOLOX network according to an embodiment of the present invention, as Figure 1 shown, the method includes:

[0057] Step S1: Collect lotus pod images, perform data augmentation and label annotation on them to obtain a lotus pod image dataset; divide the lotus pod image dataset into a training set and a test set;

[0058] Step S2: Construct a YOLOX lotus pod detection network model by sequentially connecting an Input layer, a backbone network with CSPDarkNet53 as the basic architecture, a feature enhancement network adopting a composite structure scheme of FPN+PAN, and an output prediction layer;

[0059] Step S3: Add a CA attention mechanism and an ASFF feature fusion layer to the YOLOX lotus pod detection network model to obtain an optimized YOLOX object detection network;

[0060] Step S4: Use the training set to train the optimized YOLOX object detection network;

[0061] Step S5: Use the test set to test and evaluate the trained optimized YOLOX object detection network to obtain the detection results of lotus pod images.

[0062] In step S1, collect lotus pod images, perform data augmentation and label annotation on them to obtain a lotus pod image dataset; divide the lotus pod image dataset into a training set and a test set.

[0063] Specifically, adopt a drone acquisition strategy, use the DJI Mavic 3M drone to perform aerial photography tasks over the lotus fields, and successfully capture 300 real-scene images of lotus pods covering various angles, lighting conditions, and growth stages. To further improve the model's generalization ability and cope with the diversity and complexity in practical applications, the dataset expansion in this paper uses a random combination of four core data augmentation methods: random brightness adjustment (including increase and decrease), and horizontal and vertical flipping of images. The original 300 images are expanded to 1500; use the open-source annotation tool LabelImg to finely annotate the enhanced lotus pod images, draw accurate rectangular boxes around each lotus pod instance, and store the annotated data in a way that conforms to the YOLO format. Finally, the construction of a dedicated dataset for lotus pod detection is completed. Finally, on the premise of ensuring that the dataset is sufficient and balanced, according to the industry's common practice, the entire labeled lotus pod detection dataset is divided into a training set and a test set in a ratio of 9:1. Among them, 90% of the samples are used for model training to learn the feature patterns of identifying lotus pods, and the remaining 10% of the samples are used as an independent test set to objectively evaluate the model's performance on unknown data, thereby effectively guiding the model optimization and iteration process.

[0064] In step S2, construct the YOLOX lotus pod detection network model by sequentially connecting the Input layer, the backbone network with CSPDarkNet53 as the basic architecture, the feature enhancement network adopting the composite structure scheme of FPN+PAN, and the output prediction layer.

[0065] Specifically, as Figure 2As shown, the overall architecture of the YOLOX lotus pod detection network model is divided into four key components, which are connected in an orderly manner to jointly achieve the object detection function. First is the Input module, which is responsible for receiving the image data to be detected. The size of the received image is adjustable and can be flexibly set according to different task scenarios and performance requirements. At this stage, the system will pre-adjust the size and standardize the input image to meet the input specification requirements of the subsequent network model. Then comes the Backbone level. The optimized version of CSPDarkNet53 is selected as the core network of the basic architecture. This part integrates various sub-structure units such as Focus, CSPNet, and SPPBottleNet, aiming to extract the key feature expressions of the image. Further is the Neck level design, which adopts the composite structure scheme of FPN+PAN. The FPN component undertakes the task of gradually delivering powerful features rich in semantic information from the high level downwards. At the same time, the PAN structure adds a reversely constructed feature pyramid on the basis of FPN, mainly used to transmit significant feature signals helpful for precise positioning from the bottom upwards. Finally is the Predict layer, which uses the Anchor Free method and no longer relies on predefined anchor boxes. It directly predicts the bounding box coordinates and class probabilities of the target through regression. Among them, the Focus module cuts and splices the channel of the original image in the spatial dimension, halves the width and height of the image, and quadruples the number of channels at the same time. This operation not only increases the richness of features but also reduces the computational amount of the subsequent layers. The CBS module consists of a convolutional layer (Conv), a batch normalization layer (BatchNorm), and an activation function (Silu). It can improve the feature expression ability and the robustness of the model through normalization and non-linear activation. The CSP (Cross Stage Partial Network) module improves the model's expression ability through the cross-connection of partial features. It consists of multiple convolutional layers and residual connections. It can divide the input feature map into two parts, one part passes through the convolutional layer and residual connections, and the other part is directly connected. Finally, the two parts of the feature maps are spliced. The SPP module extracts features with different receptive fields through multi-scale pooling operations, enhances the multi-scale information of the feature map, and makes the model more robust when dealing with targets of different sizes. The YOLOHead module is responsible for generating the final object detection results. Through multi-layer convolutional operations, it generates the position information, confidence, and class probabilities of each prediction box. Finally, the optimal detection results are selected through NMS.

[0066] Although YOLOX performs excellently in all aspects, there are still some problems in identifying small lotus pod targets in complex scenarios. For example, there are phenomena such as occlusion in the distribution of lotus pods in the complex environment of the lotus field, resulting in many misdetections and missed detections. At the same time, the small scale of the lotus pod and the complex background of the lotus field also affect the detection speed and accuracy. To address the above problems, this embodiment proposes an optimized YOLOX lotus pod target detection model, which further obtains a higher running speed while ensuring the accuracy rate.

[0067] In step S3, an optimized YOLOX target detection network is obtained by adding a CA attention mechanism and an ASFF feature fusion layer to the YOLOX lotus pod detection network model.

[0068] In some embodiments, in step S3, as Figure 3 shown, the obtaining of the optimized YOLOX target detection network by adding a CA attention mechanism and an ASFF feature fusion layer to the YOLOX lotus pod detection network model includes:

[0069] An optimized YOLOX target detection network is obtained by adding a CA attention mechanism between the backbone network and the feature enhancement network and adding an ASFF feature fusion layer between the feature enhancement network and the output prediction layer.

[0070] The adding of the CA attention mechanism between the backbone network and the feature enhancement network includes:

[0071] To achieve attention-guided feature enhancement, CA attention modules are respectively attached to the output feature maps of the second, third, and fourth CSP modules of the backbone network to obtain attention mechanism feature maps, enabling the network to dynamically readjust and allocate the importance of different channels during the process of gradually extracting abstract features; the attention mechanism feature maps are input into the feature enhancement network.

[0072] As Figure 4 shown, the obtaining of the attention mechanism feature maps by respectively attaching CA attention modules to the output feature maps includes:

[0073] The output feature maps are input into a residual module to obtain residual features;

[0074] Average pooling in the X direction and average pooling in the Y direction are performed on the residual features to obtain an X-direction pooled feature map of C×H×1 and a Y-direction pooled feature map of C×1×W;

[0075] The X-direction pooled feature map and the Y-direction pooled feature map are concatenated and passed through a convolutional layer to obtain a first convolutional feature map with a dimension of C / r×1×(W + H);

[0076] Batch normalization and non-linear transformation are performed on the first convolutional feature map to obtain a non-linear feature map, with the dimension remaining unchanged;

[0077] Convolution operations are applied to the non-linear feature map in the horizontal and vertical directions respectively to obtain a second convolutional feature map of C×H×1 and a third convolutional feature map of C×H×1;

[0078] The Sigmoid function is applied to the second convolutional feature map and the third convolutional feature map respectively to obtain a first function feature map of C×1×W and a second function feature map of C×1×W;

[0079] The first function feature map, the second function feature map and the residual feature are multiplied to obtain an attention mechanism feature map.

[0080] In step S4, the optimized YOLOX object detection network is trained by referring to the training set.

[0081] Specifically, the training set is input into the optimized YOLOX lotus pod object detection network model for training; and the model training parameters are set as follows: the number of iterations = 2000 times, the batch size = 32, the optimizer is SGD, the initial learning rate = 0.01, and the momentum = 0.937. Figure 5 To optimize the change trend of the key indicators of the YOLOX lotus pod detection model during training, it can be seen that the model tends to converge after more than 2000 iterations, indicating that the model has been able to extract the features of the lotus pod from the input data and can accurately predict the position of the lotus pod, but the performance of the model still needs to be further evaluated.

[0082] In step S5, the trained optimized YOLOX object detection network is tested and evaluated using the test set to obtain the lotus pod image detection results.

[0083] Specifically, accuracy (Precision), recall (Recall), average precision (Average Precision) and F1 score (F1-Score); accuracy (P) represents the proportion of samples that are truly lotus pods among the samples predicted by the model as lotus pods; recall (R) represents the proportion of all actual lotus pods correctly found by the model. Average precision (AP) is used to measure the average value of the area under the precision and recall curves at various IoU thresholds. The F1 score is the harmonic mean of precision and recall, which comprehensively reflects the balance performance of the model in terms of precision and recall ability. The specific calculation formulas are as follows:

[0084]

[0085] Where P(R) is the precision of the recall measurement, N is the number of target detections; TP represents the number of correctly predicted lotus pods by the model, FP represents the number of images where non-lotus pod regions are identified as lotus pods, and FN represents the number of samples where lotus pods are identified as non-lotus pod regions.

[0086] Finally, the test set is input into the trained original YOLOX lotus pod detection network model and the optimized YOLOX lotus pod detection network model. Figure 6 The test results using the optimized YOLOX model show that the optimized YOLOX model can well identify and locate lotus pods. Table 1 shows the results of each evaluation index during the training of the optimized YOLOX lotus pod detection network model.

[0087] Table 1

[0088] P / % R / % AP / % F1 YOLOX 82.42 81.93 81.90 0.82 Optimized YOLOX 83.54 83.54 84.72 0.84

[0089] Verified, the optimized YOLOX lotus pod detection network model shows a significant performance improvement. In the test set, the recognition accuracy of lotus pod targets climbs to 83.54%, and at the same time, the recall rate of lotus pod targets also reaches the same value of 83.54%, and the corresponding F1 score is 0.84. Compared with the original model, this optimization has achieved a steady increase in various key indicators, specifically: the accuracy has increased by 1.12 percentage points, the recall rate has increased by 1.6 percentage points, and the F1 score has increased by 2.82 percentage points. In addition, it is worth noting that the number of model parameters and the computational load have both been effectively reduced, which reveals that the present invention can successfully balance detection speed while maintaining high detection accuracy, achieving a good balance between accuracy and efficiency.

[0090] In summary, to enhance the detection ability of small target lotus pods: integrating the CA attention mechanism in the YOLOX network can dynamically adjust the importance of each channel of the feature map. Especially for the detection of small target lotus pods, by focusing on the key channel features related to target detection, it helps the model capture weak but important target signals, thereby improving the detection accuracy and recall rate of small-sized lotus pods. The ASFF feature fusion module can flexibly fuse feature information from different levels, especially having good adaptability to cross-scale targets. For lotus pod targets of different sizes, through adaptive spatial feature fusion, it can make up for the deficiencies of a single-scale feature map in small target detection.

[0091] Improving the robustness and generalization ability of feature representation: The CA attention mechanism can effectively suppress irrelevant background noise, highlight the potential lotus pod target features, enhance the discriminability of the model for lotus pod features, and maintain stable detection performance even in complex backgrounds or under large lighting changes. The ASFF module enables the model to make full use of the advantages of multi-scale features when dealing with various complex scenarios and lotus pods in different poses by fusing multi-level features, thereby improving the robustness and generalization ability of the model in different situations.

[0092] Speeding up the inference speed and optimizing memory occupancy: The CA attention mechanism performs lightweight operations to selectively enhance only the features in the channel dimension, avoiding excessive computational overhead, which helps the YOLOX network maintain a relatively fast running speed while ensuring detection performance. Compared with traditional methods such as FPN, the ASFF feature fusion can enhance the effectiveness of features while reducing redundant calculations through a more efficient feature integration method, which is particularly crucial for scenarios such as real-time transmission on embedded devices or drones, enabling the construction of a lotus pod detection system with fast response and low resource consumption.

[0093] The second aspect of the present invention discloses a remote sensing identification and extraction system for white lotus pods based on an optimized YOLOX network. Figure 7 It is a structural diagram of a remote sensing identification and extraction system for white lotus pods based on an optimized YOLOX network according to an embodiment of the present invention; as Figure 7 shown, the system 100 includes:

[0094] A first processing module 101, configured to collect lotus pod images, perform data augmentation and label annotation on them to obtain a lotus pod image dataset; and divide the lotus pod image dataset into a training set and a test set;

[0095] A second processing module 102, configured to construct a YOLOX lotus pod detection network model by sequentially connecting an Input layer, a backbone network with CSPDarkNet53 as the basic architecture, a feature enhancement network adopting a composite structure scheme of FPN+PAN, and an output prediction layer;

[0096] A third processing module 103, configured to obtain an optimized YOLOX target detection network by adding a CA attention mechanism and an ASFF feature fusion layer to the YOLOX lotus pod detection network model;

[0097] A fourth processing module 104, configured to train the optimized YOLOX target detection network by referring to the training set;

[0098] A fifth processing module 105, configured to test and evaluate the trained optimized YOLOX target detection network using the test set to obtain lotus pod image detection results.

[0099] For the system according to the second aspect of the present invention, the first processing module 101 is specifically configured as follows. To further improve the generalization ability of the model and cope with the diversity and complexity in practical applications, in this paper, the dataset expansion adopts a random combination of four core data augmentation methods, namely random brightness adjustment (including increasing and decreasing) and horizontal and vertical flipping of images, to expand the original 300 images to 1500 images. The open-source annotation tool LabelImg is used to finely annotate the enhanced lotus pod images, draw accurate rectangular frames around each lotus pod instance, and store the annotated data in a manner conforming to the YOLO format, finally completing the construction of the dataset dedicated to lotus pod detection. Finally, on the premise of ensuring that the dataset is sufficient and balanced, in accordance with the common practice in the industry, the entire annotated lotus pod detection dataset is divided into a training set and a test set in a ratio of 9:1. Among them, 90% of the samples are used for model training to learn the feature patterns for identifying lotus pods, and the remaining 10% of the samples are used as an independent test set to objectively evaluate the performance of the model on unknown data, thereby effectively guiding the model optimization and iteration process.

[0100] For the system according to the second aspect of the present invention, the second processing module 102 is specifically configured as, such as Figure 2As shown, the overall architecture of the YOLOX lotus pod detection network model is divided into four key components, which are connected in an orderly manner to jointly achieve the object detection function. First is the Input module, which is responsible for receiving the image data to be detected. The size of the received image is adjustable and can be flexibly set according to different task scenarios and performance requirements. At this stage, the system will pre-adjust the size and standardize the input image to meet the input specification requirements of the subsequent network model. Then comes the Backbone level, where an optimized version of CSPDarkNet53 is selected as the core network of the basic architecture. This part integrates various sub-structure units such as Focus, CSPNet, and SPPBottleNet, aiming to extract the key feature expressions of the image. Further is the Neck level design, which adopts the composite structure scheme of FPN+PAN. The FPN component is responsible for gradually delivering powerful features rich in semantic information from the high level downwards. At the same time, the PAN structure adds a reversely constructed feature pyramid on the basis of FPN, mainly used to transmit significant feature signals that contribute to precise positioning from the bottom upwards. Finally is the Predict layer, which uses the Anchor Free method and no longer relies on predefined anchor boxes. It directly predicts the bounding box coordinates and class probabilities of the target through regression. Among them, the Focus module cuts and splices the spatial dimensions of the original image, halves the width and height of the image, and quadruples the number of channels at the same time. This operation not only increases the richness of features but also reduces the computational amount of the subsequent layers. The CBS module consists of a convolutional layer (Conv), a batch normalization layer (BatchNorm), and an activation function (Silu). It can improve the feature expression ability and the robustness of the model through normalization and non-linear activation. The CSP (Cross Stage Partial Network) module improves the expression ability of the model through the cross-connection of partial features. It consists of multiple convolutional layers and residual connections. It can divide the input feature map into two parts, one part passes through the convolutional layer and residual connections, and the other part is directly connected. Finally, the two parts of the feature maps are spliced. The SPP module extracts features with different receptive fields through multi-scale pooling operations, enhances the multi-scale information of the feature map, and makes the model more robust when dealing with targets of different sizes. The YOLOHead module is responsible for generating the final object detection results. Through multiple convolutional operations, it generates the position information, confidence, and class probabilities of each prediction box. Finally, the optimal detection results are selected through NMS.

[0101] Although YOLOX performs excellently in all aspects, there are still some problems in recognizing small lotus pod targets in complex scenarios. For example, there are phenomena such as occlusion in the distribution of lotus pods in the complex environment of the lotus field, resulting in many false detections and missed detections. At the same time, the small scale of the lotus pod and the complex background of the lotus field also affect the detection speed and accuracy. To address the above problems, this embodiment proposes an optimized YOLOX lotus pod target detection model, which further obtains a higher running speed while ensuring the accuracy rate.

[0102] According to the system of the second aspect of the present invention, the third processing module 103 is specifically configured as, as Figure 3 shown, the optimized YOLOX target detection network obtained by adding a CA attention mechanism and an ASFF feature fusion layer to the YOLOX lotus pod detection network model includes:

[0103] An optimized YOLOX target detection network is obtained by adding a CA attention mechanism between the backbone network and the feature enhancement network and adding an ASFF feature fusion layer between the feature enhancement network and the output prediction layer.

[0104] The adding of the CA attention mechanism between the backbone network and the feature enhancement network includes:

[0105] To achieve attention-guided feature enhancement, CA attention modules are respectively attached to the output feature maps of the second, third, and fourth CSP modules of the backbone network to obtain attention mechanism feature maps, enabling the network to dynamically readjust and allocate the importance of different channels during the process of gradually extracting abstract features; the attention mechanism feature maps are input into the feature enhancement network.

[0106] As Figure 4 shown, the attaching of the CA attention modules to the output feature maps respectively to obtain the attention mechanism feature maps includes:

[0107] The output feature maps are input into a residual module to obtain residual features;

[0108] Average pooling in the X direction and average pooling in the Y direction are performed on the residual features to obtain an X-direction pooled feature map of C×H×1 and a Y-direction pooled feature map of C×1×W;

[0109] The X-direction pooled feature map and the Y-direction pooled feature map are concatenated and passed through a convolutional layer to obtain a first convolutional feature map with a dimension of C / r×1×(W + H);

[0110] Batch normalization and non-linear transformation are performed on the first convolutional feature map to obtain a non-linear feature map with the dimension remaining unchanged;

[0111] Apply convolution operations to the non - linear feature map in the horizontal and vertical directions respectively to obtain a second convolution feature map of C×H×1 and a third convolution feature map of C×H×1;

[0112] Apply the Sigmoid function to the second convolution feature map and the third convolution feature map respectively to obtain a first function feature map of C×1×W and a second function feature map of C×1×W;

[0113] Multiply the first function feature map, the second function feature map and the residual feature to obtain an attention mechanism feature map.

[0114] According to the system of the second aspect of the present invention, the fourth processing module 104 is specifically configured to input a training set into the optimized YOLOX lotus pod target detection network model for training; and the model training parameters are set as follows: the number of iterations = 2000 times, the batch size = 32, the optimizer is SGD, the initial learning rate = 0.01, and the momentum = 0.937. Figure 5 To optimize the trend of key indicators during the training process of the YOLOX lotus pod detection model, it can be seen that the model tends to converge after more than 2000 iterations, indicating that the model can already extract the features of the lotus pod from the input data and can accurately predict the position of the lotus pod, but the performance of the model still needs to be further evaluated.

[0115] According to the system of the second aspect of the present invention, the fifth processing module 105 is specifically configured to calculate the accuracy (Precision), recall rate (Recall), average precision (Average Precision) and F1 - score (F1 - Score); the accuracy (P) represents the proportion of samples that are truly lotus pods among the samples predicted by the model as lotus pods; the recall rate (R) represents the proportion of all actually existing lotus pods correctly found by the model. The average precision (AP) is used to measure the average value of the area under the precision - recall curve at each IoU threshold. The F1 - score is the harmonic mean of the precision and recall rates, which comprehensively reflects the balanced performance of the model in terms of precision and recall ability. The specific calculation formulas are as follows:

[0116]

[0117] In the formula, P(R) is the precision measured by recall, N is the number of object detections; TP represents the number of lotus pods correctly predicted by the model, FP represents the number of samples that identify non - lotus pod regions as lotus pods, and FN represents the number of samples that identify lotus pods as non - lotus pod regions.

[0118] Finally, input the test set into the trained original YOLOX lotus pod detection network model and the optimized YOLOX lotus pod detection network model. Figure 6The test results of using the optimized YOLOX model show that the optimized YOLOX model can well identify and locate the lotus pods. Table 1 shows the results of various evaluation metrics of the optimized YOLOX lotus pod detection network model during the training process.

[0119] After verification, the optimized YOLOX lotus pod detection network model has shown significant performance improvement. In the test set, the recognition accuracy of the lotus pod target has climbed to 83.54%, and at the same time, the recall rate of the lotus pod target has also reached the same value of 83.54%, and the corresponding F1 score is 0.84. Compared with the original model, this optimization has achieved a steady increase in various key indicators. Specifically, the accuracy has increased by 1.12 percentage points, the recall rate has increased by 1.6 percentage points, and the F1 score has increased by 2.82 percentage points. In addition, it is worth noting that the number of model parameters and the computational load have been effectively reduced, which reveals that the present invention can successfully balance the detection speed while maintaining a high detection accuracy, achieving a good balance between accuracy and efficiency.

[0120] The third aspect of the present invention discloses an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, it implements the steps in a method for remote sensing recognition and extraction of white lotus pods based on an optimized YOLOX network according to any one of the first aspects disclosed in the present invention.

[0121] Figure 8 For the structural diagram of an electronic device according to an embodiment of the present invention, as Figure 8 shown, the electronic device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the electronic device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be achieved through WIFI, a carrier network, near field communication (NFC), or other technologies. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the outer shell of the electronic device, or an external keyboard, a touchpad, or a mouse, etc.

[0122] Those skilled in the art can understand, Figure 8The structure shown is only a structural diagram of the part related to the technical solution of the present disclosure, and does not constitute a limitation on the electronic device to which the solution of this application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.

[0123] The fourth aspect of the present invention discloses a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the steps in a method for remote sensing identification and extraction of white lotus seedpods based on an optimized YOLOX network according to any one of the first aspects disclosed in the present invention are implemented.

[0124] Please note that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification. The above embodiments only represent several implementation manners of this application, and their descriptions are relatively specific and detailed, but they should not be construed as a limitation on the scope of the invention patent. It should be pointed out that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of the patent of this application should be subject to the appended claims.

Claims

1. A method for remote sensing identification and extraction of white lotus pods based on optimized YOLOX network, characterized in that: The method comprises: Step S1, collecting lotus pod images and performing data enhancement and labeling on them to obtain a lotus pod image dataset; dividing the lotus pod image dataset into a training set and a test set; Step S2, constructing the YOLOX lotus detection network model by orderly connecting the input layer with CSPDarkNet53 as the backbone network of the basic architecture, the feature enhancement network adopting the composite structure solution of FPN+PAN, and the output prediction layer; Step S3, by adding CA attention mechanism and ASFF feature fusion layer to the YOLOX lotus detection network model, an optimized YOLOX object detection network is obtained; Step S4, referencing the training set to train the optimized YOLOX object detection network; Step S5, using the test set to test and evaluate the trained optimized YOLOX object detection network to obtain a lotus pod image detection result; In step S3, the optimized YOLOX target detection network obtained by adding CA attention mechanism and ASFF feature fusion layer to the YOLOX lotus detection network model includes: By adding a CA attention mechanism between the backbone network and the feature enhancement network, and adding an ASFF feature fusion layer between the feature enhancement network and the output prediction layer, an optimized YOLOX object detection network is obtained; the ASFF feature fusion layer includes three modules, ASFF1, ASFF2 and ASFF3, and adaptive fusion of multi-scale features is achieved through the three modules, wherein the three features output by the feature enhancement network are used as inputs of each of the three modules; In the step S3, adding a CA attention mechanism between the backbone network and the feature enhancement network includes: The CA attention modules are respectively attached to the output feature maps of the second, third and fourth CSP modules of the backbone network to obtain attention mechanism feature maps; and the attention mechanism feature maps are input into the feature enhancement network.

2. The method for remote sensing identification and extraction of white lotus pods based on optimized YOLOX network according to claim 1, characterized in that: In step S3, the CA attention modules are respectively attached to the output feature maps, and the attention mechanism feature maps obtained include: Inputting the output feature map into a residual module to obtain residual features; Performing average pooling in the X direction and average pooling in the Y direction on the residual features to obtain an X-direction pooling feature map and a Y-direction pooling feature map; The X-direction pooling feature map and the Y-direction pooling feature map are concatenated, and a first convolution feature map is obtained through a convolution layer; Perform batch normalization and nonlinear transformation on the first convolution feature map to obtain a nonlinear feature map; Applying convolution operations to the nonlinear feature map in the horizontal direction and the vertical direction respectively to obtain a second convolution feature map and a third convolution feature map; Applying a Sigmoid function to the second convolution feature map and the third convolution feature map respectively to obtain a first function feature map and a second function feature map; The first function feature map, the second function feature map and the residual feature are multiplied to obtain an attention mechanism feature map.

3. A white lotus seed pod remote sensing identification and extraction system based on optimized YOLOX network, characterized in that: The system comprises: The first processing module is configured to collect lotus pod images and perform data enhancement and label annotation on them to obtain a lotus pod image dataset; and divide the lotus pod image dataset into a training set and a test set; The second processing module is configured to construct a YOLOX lotus detection network model by orderly connecting the input layer, the backbone network with CSPDarkNet53 as the basic architecture, the feature enhancement network adopting the composite structure scheme of FPN+PAN, and the output prediction layer; The third processing module is configured to obtain an optimized YOLOX object detection network by adding a CA attention mechanism and an ASFF feature fusion layer to the YOLOX lotus detection network model; The ASFF feature fusion layer includes three modules, ASFF1, ASFF2 and ASFF3, through which adaptive fusion of multi-scale features is realized, wherein the three features output by the feature enhancement network are used as inputs of each of the three modules; A fourth processing module is configured to train the optimized YOLOX object detection network by referencing the training set; A fifth processing module is configured to apply the test set to test and evaluate the trained optimized YOLOX object detection network to obtain a lotus pod image detection result; The optimized YOLOX target detection network obtained by adding CA attention mechanism and ASFF feature fusion layer to the YOLOX lotus detection network model includes: By adding a CA attention mechanism between the backbone network and the feature enhancement network, and adding an ASFF feature fusion layer between the feature enhancement network and the output prediction layer, an optimized YOLOX object detection network is obtained; The adding of CA attention mechanism between the backbone network and the feature enhancement network includes: The CA attention modules are respectively attached to the output feature maps of the second, third and fourth CSP modules of the backbone network to obtain attention mechanism feature maps; and the attention mechanism feature maps are input into the feature enhancement network.

4. The white lotus seed pod remote sensing identification and extraction system based on optimized YOLOX network according to claim 3 is characterized in that: The CA attention modules are respectively added to the output feature maps, and the obtained attention mechanism feature maps include: Inputting the output feature map into a residual module to obtain residual features; Performing average pooling in the X direction and average pooling in the Y direction on the residual features to obtain an X-direction pooling feature map and a Y-direction pooling feature map; The X-direction pooling feature map and the Y-direction pooling feature map are concatenated, and a first convolution feature map is obtained through a convolution layer; Perform batch normalization and nonlinear transformation on the first convolution feature map to obtain a nonlinear feature map; Applying convolution operations to the nonlinear feature map in the horizontal direction and the vertical direction respectively to obtain a second convolution feature map and a third convolution feature map; Applying a Sigmoid function to the second convolution feature map and the third convolution feature map respectively to obtain a first function feature map and a second function feature map; The first function feature map, the second function feature map and the residual feature are multiplied to obtain an attention mechanism feature map.

5. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps of the white lotus and lotus seed pod remote sensing identification and extraction method based on optimized YOLOX network described in any one of claims 1 to 2 are implemented.

6. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the steps of the white lotus and lotus seed pod remote sensing identification and extraction method based on the optimized YOLOX network described in any one of claims 1 to 2 are implemented.

Citation Information

Patent Citations

  • Remote sensing image target detection method based on improved lightweight YOLOv4

    CN116363529A

  • Remote sensing image dotted independent room detection method and system based on YOLOX network

    CN116630807A

  • Lotus seedpod lightweight target detection method based on YOLOv5

    CN117788997A