A method, device, equipment and medium for testing the fruit of sesame plants

Through the improved LEHP-DETR network model, the identification difficulty caused by irregular shape, multiple branches and occlusion in the flax fruit test species was solved, and the accurate identification and detection of flax fruit was achieved.

CN118823749BActive Publication Date: 2025-05-06GANSU AGRI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410801512.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-20
Publication Date
2025-05-06
Estimated Expiration
2044-06-20

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify the fruits of flax plants due to their irregular outline size and shape, many branches, random growth positions, and different cracking levels, resulting in different degrees of fruit ripening.

Method used

The improved LEHP-DETR network model is adopted. This model improves the RT-DETR network model by introducing the RepNCSPELAN4 module, ADown module, ContextAggregation module and TFE module, as well as the newly designed HWD-ADown module, HiLo-AIFi module and DSSFF module, to form a lightweight, efficient and accurate small object detection model.

Benefits of technology

It significantly improves the accuracy and recall of small target detection, can clearly identify low-contrast, small targets and high occlusion fruits, and comprehensively capture different frequency characteristics in the image, reducing the difficulty of flax plant fruit examination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118823749B_ABST
    Figure CN118823749B_ABST
Patent Text Reader

Abstract

The invention discloses a method, device, equipment and medium for testing the fruit of sesame plants, and relates to the technical field of target recognition. The invention introduces a RepNCSPELAN4 module, an ADown module, a Context Aggregation module and a TFE module, and designs three modules, namely, HWD-ADown, HiLo-AIFi and DSSFF, to improve and fuse the three parts of BackBone, Efficient Hybrid Encoder and Head of the RT-DETR basic model, and proposes a new LEHP-DETR network model. The model can effectively distinguish different frequency features in an image, thereby successfully capturing local and edge detail features and global structural features of the image, and not only significantly improves the accuracy when identifying small targets with low contrast, high occlusion and unclear features; moreover, it realizes lightweight and improves the inference speed, which is more conducive to the promotion and application in actual production. Then, the sesame plant data set is used to train and test the model, and the results show that it can accurately identify the fruits of sesame plants, that is, it is biased towards the recognition of small fruit targets in sesame plant images, and can greatly improve the efficiency of testing the seeds of sesame plants.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of target recognition, and in particular to a method, device, equipment and medium for testing the fruit of sesame plants. Background Art

[0002] Sesame is rich in essential fatty acids and nutrients, and has been favored by more and more consumers. It is recognized by international nutritionists and physiologists as one of the important functional crops. At the same time, due to its strong fiber toughness, it is an important industrial and textile fabric. The innovation of germplasm resources and the cultivation of new and excellent varieties are the fundamental way out for the healthy and sustainable development of the sesame industry, and rapid and accurate seed testing is one of the most critical basic tasks for the development of the sesame industry. At present, sesame seed testing is basically completely dependent on manual processing, which is inefficient, has a high error rate, and is subjective. In addition, with the acceleration of urbanization, labor costs continue to increase, which seriously restricts the development of the sesame industry. How to use computer technology to achieve convenient, efficient and accurate seed testing has become the focus of global sesame breeding experts.

[0003] In recent years, the development of computer vision and deep learning technology has provided new ideas for solving this problem. In particular, target detection technology has become an important research direction in the agricultural field. Target detection technology provides a basis for visual analysis by locating and identifying objects in images. Some researchers have proposed a target detection model based on YOLO, but due to its large number of parameters and floating-point operations, the timeliness of the model is low; and due to the complex and changeable background of crop growth environment, the model is prone to false detection or missed detection, which leads to precision loss. All these limit the promotion and application of such models in actual agricultural production.

[0004] In summary, the application of small target detection technology in the agricultural field still faces a series of problems, such as difficulty in feature selection, loss of feature information, and false detection and missed detection caused by differences in feature responses at different scales. Sesame fruit has irregular outlines, sizes and shapes due to the different number of ovaries and their respective development levels. The number of branches is large and the growth position is random, which will cause large-area occlusion of the fruit. The degree of maturity of the fruit is different, resulting in different degrees of cracking. These phenomena increase the difficulty of sesame plant fruit testing, that is, sesame fruit (capsule) small target detection, making the various problems existing in small target detection more prominent in sesame plant testing. Summary of the invention

[0005] The embodiments of the present invention provide a method, device, equipment and medium for testing the fruit of sesame plants, which can solve the problems in the prior art that sesame plants generally have irregular capsules in size and shape, a large number of branches, random growth positions, large-area occlusion of the fruit, and different degrees of fruit cracking, which increase the difficulty of testing the fruit of sesame plants, make various problems existing in small target detection more difficult in sesame plant testing, and cannot accurately identify the small target of the capsule in sesame plant testing.

[0006] The embodiment of the present invention provides a method for testing the fruit of sesame plants, comprising the following steps:

[0007] Obtain an image dataset of sesame plants;

[0008] Construct an RT-DETR network model, introduce the RepNCSPELAN4 module, ADown module and ContextAggregation module into the BackBone part of the RT-DETR network model, merge the ADown module and the Haar waveletdownsampling module into the HWD-ADown module and then introduce it into the BackBone part, and merge the RepNCSPELAN4 module, ADown module, ContextAggregation module and HWD-ADown module to replace the original ResNet structure in the BackBone part of the RT-DETR network model and merge them into the BackBone part to form an improved RAHC-BackBone part;

[0009] In the Efficient Hybrid Encoder part of the RT-DETR network model, the Multi-headAttention module in the AIFI module is replaced with the HiLoAttention module, the HiLoAttention module and the AIFI module are merged into the HiLo-AIFi module and introduced into the Efficient Hybrid Encoder part, the SSFF module and the DySample module are merged into the DSSFF module and introduced into the Efficient Hybrid Encoder part, and the TFE module is introduced and merged with the HiLo-AIFi module and the DSSFF module into the improved HTD-Efficient Hybrid Encoder part;

[0010] The P2 layer detection head is introduced into the Head part of the RT-DETR network model, and the P2 layer detection head and the Head part are merged into the improved P2-Head part;

[0011] The improved RAHC-BackBone part, the improved HTD-Efficient Hybrid Encoder part and the improved P2-Head part are integrated to obtain the LEHP-DETR network model that is improved on the RT-DETR network model; and the LEHP-DETR network model is trained using the sesame plant image dataset to obtain the trained LEHP-DETR network model;

[0012] The image of sesame plants is input into the trained LEHP-DETR network model; through the improved RAHC-BackBone part, the feature map with significant edge information and detail information is obtained; through the improved HTD-EfficientHybrid Encoder part, the feature map with significant contrast between background and target fruit is obtained; through the improved P2-Head part, the fruit species of sesame plants are obtained.

[0013] Preferably, the step of acquiring an image dataset of sesame plants comprises:

[0014] Use an industrial camera to take multiple images of sesame plants;

[0015] The Labelmg tool was used to annotate sesame capsules in multiple sesame plant images and to construct an image dataset of sesame plants.

[0016] Preferably, the RepNCSPELAN4 module comprises:

[0017] The RepNCSPELAN4 module includes a RepNCSP module and an ELAN module;

[0018] The RepNCSP module is used to extract features, and the ELAN module is used to enhance the recognition of targets.

[0019] Preferably, the HiLo-AIFi module includes:

[0020] Replace the Multi-headAttention module in the AIFI module with the HiLoAttention module, and merge the HiLoAttention module and the AIFI module to form the HiLo-AIFi module;

[0021] The Multi-headAttention module relies on the self-attention mechanism to generate an attention distribution by calculating the similarity between each element in the input sequence and all other elements;

[0022] The HiLoAttention module divides attention into two parts: a low-frequency part and a high-frequency part. The low-frequency part is used to capture global structural information, and the high-frequency part is used to extract local detail information.

[0023] Preferably, the TFE module comprises:

[0024] The TFE module obtains three tensors L, M and S, takes the height and width of the M tensor as the target size, adjusts L to the same size as M through adaptive pooling operation, and performs maximum pooling and average pooling before adding them;

[0025] Use interpolation to adjust S to the same size as M, and concatenate L, M, and S into a tensor according to the channel dimension, i.e., dim=1; the formula for obtaining the output feature map of the TFE module is as follows:

[0026] F TFE =Concat(F 1 ,F m ,F s )

[0027] Among them: F TFE F represents the feature map output by the TFE module; 1 、F m and F s Representing feature maps of large, medium, and small sizes, respectively;

[0028] F TFE By F 1 、F m and F s Splicing to obtain; F TFE and F m The resolution is the same, and the number of channels is F m Three times.

[0029] Preferably, the P2 layer detection head has a resolution of 160×160 pixels and is downsampled twice in the backbone network.

[0030] Preferably, a ContextAggregation module is inserted between the RAHC-BackBone part and the HTD-EfficientHybridEncoder part of the LEHP-DETR network model;

[0031] The ContextAggregation module is used to enable the high-level feature map to retain the detail information in the low-level feature map.

[0032] The embodiment of the present invention also provides a sesame plant fruit testing device, comprising:

[0033] An image module, used to obtain an image dataset of sesame plants;

[0034] Model module, build RT-DETR network model, introduce RepNCSPELAN4 module, ADown module and ContextAggregation module into BackBone part of RT-DETR network model, merge ADown module and Haarwavelet downsampling module into HWD-ADown module and then introduce into BackBone part, merge RepNCSPELAN4 module, ADown module, Context Aggregation module and HWD-ADown module and then replace the original ResNet structure in BackBone part of RT-DETR network model and merge with BackBone part to form improved RAHC-BackBone part;

[0035] In the Efficient Hybrid Encoder part of the RT-DETR network model, the Multi-headAttention module in the AIFI module is replaced with the HiLoAttention module, the HiLoAttention module and the AIFI module are merged into the HiLo-AIFi module and introduced into the Efficient Hybrid Encoder part, the SSFF module and the DySample module are merged into the DSSFF module and introduced into the Efficient Hybrid Encoder part, and the TFE module is introduced and merged with the HiLo-AIFi module and the DSSFF module into the improved HTD-Efficient Hybrid Encoder part;

[0036] The P2 layer detection head is introduced into the Head part of the RT-DETR network model, and the P2 layer detection head and the Head part are merged into the improved P2-Head part;

[0037] The improved RAHC-BackBone part, the improved HTD-Efficient Hybrid Encoder part and the improved P2-Head part are integrated to obtain the LEHP-DETR network model that is improved on the RT-DETR network model; and the LEHP-DETR network model is trained using the sesame plant image dataset to obtain the trained LEHP-DETR network model;

[0038] The recognition module inputs the image of sesame plants into the trained LEHP-DETR network model; through the improved RAHC-BackBone part, the feature map with significant edge information and detail information is obtained; through the improved HTD-Efficient Hybrid Encoder part, the feature map with significant contrast between background and target fruit is obtained; through the improved P2-Head part, the fruit species of sesame plants are obtained.

[0039] An embodiment of the present invention further provides an electronic device, including a memory and a processor;

[0040] The memory is used to store computer programs;

[0041] The processor is used to implement the steps of the above-mentioned method for testing the fruit of sesame plants when executing the computer program stored in the memory.

[0042] An embodiment of the present invention further provides a computer-readable storage medium for storing a computer program, which, when executed by a processor, implements the steps of the above-mentioned method for testing the fruit of sesame plants.

[0043] The embodiment of the present invention provides a method, device, equipment and medium for testing the fruit of sesame plants. Compared with the prior art, the beneficial effects thereof are as follows:

[0044] The present invention introduces a RepNCSPELAN4 module, an ADown module, a ContextAggregation module and a TFE module, as well as a newly designed HWD-ADown module, a HiLo-AIFi module and a DSSFF module, respectively, to replace and fuse a BackBone part, an Efficient Hybrid Encoder part and a Head part in an RT-DETR network model to form an improved LEHP-DETR network model, so that the model can clearly identify small targets with low contrast, small targets with high occlusion and small targets with inconspicuous fruits, and can comprehensively capture different frequency features in an image, so that the recognition accuracy of small targets is greatly improved, and then the LEHP-DETR network model is trained by using an image data set of sesame plants, so that the trained LEHP-DETR network model can be biased towards the recognition of small targets in sesame plant images, that is, it can easily identify the fruits of sesame plants with different development degrees, irregular fruit outlines, large numbers of branches, mutually occluded positions and different cracking degrees, thereby greatly reducing the difficulty of testing the fruits of sesame plants. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1A schematic diagram of the overall process of a sesame plant fruit testing method, device, equipment and medium provided by an embodiment of the present invention;

[0046] Figure 2 A schematic diagram of the LEHP-DETR network structure of a sesame plant fruit testing method, device, equipment and medium provided in an embodiment of the present invention;

[0047] Figure 3 A schematic diagram of the HWD-ADown structure of a sesame plant fruit testing method, device, equipment and medium provided in an embodiment of the present invention;

[0048] Figure 4 A schematic diagram of the HiLo-AIFI structure of a sesame plant fruit testing method, device, equipment and medium provided in an embodiment of the present invention;

[0049] Figure 5 A schematic diagram of the DSSF module structure of a sesame plant fruit testing method, device, equipment and medium provided in an embodiment of the present invention;

[0050] Figure 6 A schematic diagram of the TFE module structure of a sesame plant fruit testing method, device, equipment and medium provided in an embodiment of the present invention;

[0051] Figure 7 A schematic diagram of a sesame data set collection process for a sesame plant fruit seed testing method, device, equipment and medium provided by an embodiment of the present invention;

[0052] Figure 8 A schematic diagram of a comparison of the practical applications of YOLOv5m, YOLOv8m, YOLOv9c, RT-DETR-R18 and LHEP-DETR for a method, device, equipment and medium for testing the fruit of a sesame plant provided by an embodiment of the present invention;

[0053] Fig. 9 A schematic diagram of partial details of the practical application of the LHEP-DETR model of a sesame plant fruit testing method, device, equipment and medium provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0054] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below in conjunction with the accompanying drawings. In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without violating the connotation of the present invention, so the present invention is not limited by the specific embodiments disclosed below.

[0055] like Figure 1 As shown, an embodiment of the present invention provides a method for testing the fruit of sesame plants, and proposes an improved model LEHP-DETR based on RT-DETR; the model significantly improves the accuracy and recall rate of small target detection, reduces the size of the model, the number of parameters and the number of floating-point operations, and improves the reasoning speed of the model by introducing the RepNCSPELAN4 module, the ADown module, the ContextAggregation module and the TFE module, as well as the newly designed HWD-ADown module, the HiLo-AIFi module and the DSSFF module.

[0056] Specifically:

[0057] LEHP-DETR structure:

[0058] The new target detection model LEHP-DETR proposed in this paper is a further improvement of the R18 model based on RT-DETR. The structure of the model is as follows: Figure 2 As shown, it mainly consists of three main parts: RAHC-BackBone, HTD-Efficient Hybrid Encoder and P2-Head.

[0059] RAHC-BackBone improves the traditional backbone network by introducing the RepNCSPELAN4 module, ADown module and the designed DSSFF module, thereby improving the feature expression capability and reducing the computational complexity; this improvement not only achieves the lightweight of the model, but also improves the reasoning speed, especially in small target detection tasks.

[0060] HTD-Efficient Hybrid Encoder designs the HiLo-AIFI module and the DSSFF module, and further introduces the TFE module to achieve an overall improvement in model performance; the HiLo-AIFI module better captures the local details and global structure of the image by distinguishing between high-frequency and low-frequency information, thereby enhancing the generalization ability of the model; the DSSFF module enriches the detail information and improves the multi-scale feature extraction capability of the model by integrating feature maps of different scales; the TFE module accurately captures the detail information of small targets by combining features of different sizes in the spatial dimension.

[0061] In addition, the ContextAggregation module is introduced between AHC-BackBone and HTD-Efficient Hybrid Encoder, aiming to enhance the model’s understanding and utilization of contextual information, while optimizing the feature transmission process to ensure that high-level feature maps can better retain the detail information in low-level feature maps; in the P2-Head part, the P2 detection head is innovatively adopted to make full use of the rich details and edge information in the low-level feature maps, significantly improving the detection performance of small targets; the LEHP-DETR model achieves a balance between real-time performance and detection accuracy, and through its innovative structural design, it shows significant performance advantages in the sesame fruit small target detection task.

[0062] RAHC-BackBone structure:

[0063] In the RAHC-BackBone structure, the present invention replaces the traditional ResNet structure with the RepNCSPELAN4 module, the ADown module and the HWD-ADown module; the purpose of this change is to optimize the computational complexity and parameter quantity of the model, thereby making the model more lightweight and efficient. The details of the structure are as follows Figure 3; Although the ResNet architecture enhances the depth and expression ability of the model through residual connections, its high number of parameters and computational complexity as well as single-scale feature extraction limit its application in small target detection; In order to overcome these limitations, the present invention adopts the RepNCSPELAN4 module, which integrates the RepNCSP module and the ELAN module. The RepNCSP module is used to extract features, while the ELAN module is used to enhance the detection capability of the target, thereby improving the detection accuracy; In addition, the design of these modules is relatively lightweight and has low computational complexity, which is suitable for processing a large number of images; The ADown module is mainly used for feature extraction and downsampling operations. , the ADown module can retain details and important features, which is crucial for feature fusion and feature transfer; specifically, the ADown module is a convolution block composed of a series of convolution layers and pooling layers. Its function is to gradually reduce the size of the feature map and increase the number of channels through multiple convolution and pooling operations, so as to better extract the features of the target; the design of the HWD-ADown module combines the ADown module and the Haarwaveletdownsampling module; the construction of this module has significant advantages in retaining image details and edge information, which is crucial for object positioning and recognition; in the HWD-ADown module, the present invention uses the Haar wavelet downsampling module to replace the 3x3 convolution in the ADown module; in order to reduce the computational complexity and parameter amount of the model, the present invention only replaces the last ADown module in BackBone with the HWD-ADown module.

[0064] HTD-EfficientHybrid Encoder structure:

[0065] In the HTD-EfficientHybrid Encoder structure, the present invention designs the HiLo-AIFI module and the DSSFF module, and introduces the TFE module; Figure 4As shown in the figure, in the HiLo-AIFI module, the present invention innovatively replaces the traditional Multi-headAttention module in AIFI with the HiLoAttention module; the Multi-headAttention module mainly relies on the self-attention mechanism, and generates the attention distribution by calculating the similarity between each element in the input sequence and all other elements; however, for the small target detection task, the small target usually occupies only a small area in the image, so the effective information contained is very limited, which makes it difficult for Multi-headAttention to effectively capture the characteristics of the small target; in contrast, HiLo Attention divides attention into two parts: low frequency (Low Frequency) and high frequency (High Frequency); the low frequency part focuses on capturing global structural information, while the high frequency part focuses on extracting local detail information; this mechanism of separating and processing different frequency information enables HiLoAttention to more effectively capture and utilize contextual information in the image, especially when processing small target detection tasks, it can more accurately process the relationship between small targets and complex backgrounds; at the same time, the module can more comprehensively capture different frequency features in the image, which enables the model to have better generalization ability for different types of images and scenes.

[0066] In order to meet the processing requirements of features of different scales, more comprehensively capture and fuse target information of different scales, obtain richer context information, and improve the accuracy and robustness of detection, the present invention designs a DSSFF module, such as Figure 5 As shown in the figure, this module combines the SSFF module with the DySample module, and replaces the Upsample module in the SSFF module with the DySample module. The traditional Upsample module increases the size of the feature map through interpolation or copying operations. Although it is simple to calculate and fast, it has limitations in detail preservation and edge preservation. In contrast, the DySample module can better preserve the edge detail information of the image, especially when processing small target detection, it can capture and locate small targets more accurately. In addition, DySample can also dynamically adjust the size of the feature map according to the needs of different tasks, thereby improving the performance of the model. Figure 2 In the HTD-Efficient Hybrid Encoder part shown, the DSSFF module fuses the output features extracted from RAHC-BackBone; by effectively fusing these feature maps, sesame fruits with different spatial scales, sizes and shapes can be captured; in DSSFF, different feature maps are normalized to the same size and superimposed together through dynamic upsampling as the input of three-dimensional convolution, thereby combining multi-scale features.

[0067] In addition, the present invention also introduces a TFE module; the structure is as follows Figure 6 As shown in the figure; Since different feature layers of the RAHC-Backbone structure have different sizes, the traditional feature pyramid network (FPN) fusion mechanism only upsamples the small-sized feature map and adds it to the previous layer of features, which may lead to ignoring the rich detail information in the larger-sized feature layer; In order to capture the detailed information of sesame fruits and enhance the detection ability of dense small targets, the TFE module splices three different-sized feature maps of Large (L), Medium (M), and Sampll (S) in the spatial dimension; The basic principle of the TFE module obtains three tensors L, M, and S, takes the height and width of the M tensor as the target size, adjusts L to the same size as M through adaptive pooling operations, and performs maximum pooling and average pooling before adding; Use interpolation operations to adjust S to the same size as M; Splice L, M, and S into a tensor according to the channel dimension (dim=1); The formula is as follows:

[0068] F TFE =Concat(F 1 ,F m ,F s )

[0069] Among them: F TFE F represents the feature map output by the TFE module; 1 、F m and F s Represents the feature maps of large, medium and small sizes respectively; F TFE By F 1 、F m and F s Splicing to obtain; F TFE and F m The resolution is the same, and the number of channels is F m Three times.

[0070] In order to more effectively integrate features from different modules and improve the model's ability to utilize target context information in images, the present invention introduces a ContextAggregation module; as a general building block for multi-head context aggregation, this module aims to achieve effective integration of different context information by effectively aggregating spatial information; this module can not only capture long-range interactions like Transformer, but also explore inductive biases like local convolution, thereby achieving faster convergence speed; the ContextAggregation module adopts an efficient context aggregation mechanism, so that information between different layers can be directly transmitted, which not only improves computational efficiency but also reduces the consumption of computing resources; this design takes into account both global and local information, which is helpful for understanding the overall structure and context of the image, but also can capture details and specific features, providing a richer and more effective feature representation for subsequent target detection or recognition tasks.

[0071] P2-Head structure:

[0072] In view of the challenges in the sesame fruit target detection task, the present invention focuses on the occlusion problem caused by the small size and dense distribution of the target; after multiple downsampling, the feature information of such targets is often lost, and even with a P3 layer detection head with a higher resolution, it is difficult to effectively detect; in order to overcome this problem, the present invention introduces a detection head based on the P2 layer features in the Head part, and its structure is as follows Figure 2 As shown in the P2-Head in the figure; the P2 layer detection head has a resolution of 160×160 pixels, which means that only two downsampling operations are performed in the backbone network, so richer underlying feature information can be retained; by adding the P2 detection head to the model, the present invention significantly improves the detection accuracy of small targets; the P2 detection head can capture the features of small targets more accurately, thereby enhancing the generalization ability of the model and making it more reliable in practical applications.

[0073] Specific experiments:

[0074] The present invention creates a dataset for sesame plants, which covers nearly 500 sesame varieties and a total of 2157 images are collected. Figure 7 As shown in the figure; through data enhancement, it was expanded to 6539 images; the fruits in the sesame plant images were labeled using the LabelImg tool and divided into training set, validation set and test set in a ratio of 7:1:2.

[0075] The evaluation indicators include: the present invention uses model size (Model Size), number of parameters (Parameters), number of floating-point operations (FLOPs), precision (Precision), recall rate (Recall), mean average precision (mAP) and frame rate (FPS) as evaluation indicators.

[0076] The ablation experiment results are as follows:

[0077] In order to evaluate the impact of different modules in the RAHC-BackBone model on the model performance, the present invention conducted an ablation experiment; the experimental results show that while reducing the Model Size, Parameters and FLOPs, the Precision, Recall and FPS are improved. Therefore, the present invention can conclude that the RA, HA and CA modules realize the lightweight of the RAHC-BackBone model, improve the accuracy of the model, and also improve the inference speed of the model.

[0078] In order to further explore the specific impact of HiLo-AIFi and TDS-CCFM modules on the HTD-Efficient Hybrid Encoder model, the present invention continued to carry out a series of ablation experiments; the experimental results show that Precision, Recall, mAP50:95 and FPS are significantly improved. Although the FPS is slightly reduced, it can still meet the real-time requirements.

[0079] In order to comprehensively evaluate the impact of RAHC-BackBone (RAHC), HTD-Efficient Hybrid Encoder (HTD) and P2-Head (P2) on the performance of the LEHP-DETR model proposed in the present invention, the present invention implemented a comprehensive ablation experiment. The experimental results reveal the contribution of each component to improving the performance of the LEHP-DETR model, and show the performance trade-offs under different component combinations, so as to promote and apply it in actual production.

[0080] The comparative experimental results are as follows:

[0081] In order to fully verify the efficiency and accuracy of the LEHP-DETR model proposed in the present invention in the sesame plant fruit target detection task, the present invention compares it with the current advanced target detection algorithms, such as Figure 8 As shown in the figure; the models involved in the comparison include the basic models RT-DETR-R18, YOLOv5m, YOLOv8m and YOLOv9c; in order to ensure the fairness and effectiveness of the comparison, the experimental platforms and data sets of all five models are kept consistent; the comparative experimental results of different models show that the new model LHEP-DETR proposed in this invention has achieved significant improvements in various indicators.

[0082] Through comparative experiments, it is proved that the new model proposed in this invention achieves excellent target detection performance while maintaining light weight and time efficiency.

[0083] Performance comparison on independent test samples:

[0084] In order to evaluate the performance of the model in practical applications, the present invention carefully selected 10 external images that were not included in the original dataset to simulate actual application scenarios. Figure 8 As shown in the figure, during the target detection process, the detection details of interest in the image are selected with white boxes and partially enlarged to more clearly demonstrate the detection capability of the model, as shown in the figure. Fig. 9 shown; through Figure 8 From the intuitive display, it can be clearly observed that the YOLOv5m, YOLOv8m and YOLOv9c models have obvious missed detection phenomena in the sesame plant fruit small target detection task, and the overall detection effect is poor, which makes it difficult to meet the basic target detection needs; in contrast, the new model LEHP-DETR proposed in the present invention shows superior performance in this task, almost achieving complete detection of all targets, and missed detection and false detection are extremely rare; this practical comparison verifies that the new model proposed in the present invention has reliability and stability for actual application scenarios; therefore, LEHP-DETR provides a new and effective solution for lightweight, efficient and accurate small target detection tasks in practical applications.

[0085] In the present invention, an improved model LEHP-DETR based on RT-DETR is proposed to address the challenge of small targets in sesame plant fruit detection; through a series of innovative improvements, the present invention significantly improves the Precision and Recall of small target detection; specifically, the RepNCSPELAN4 module and the ADown module are introduced in the BackBone part, which effectively enhances the feature representation capability and reduces the feature map resolution, thereby reducing the computational complexity; the HWD-ADown module is designed, which combines ADown and Haarwavelet downsampling to improve the accuracy and robustness of target detection; in order to better aggregate contextual information, the ContextAggregation module is introduced; in the Efficient Hybrid Encoder part, the HiLo-AIFi module is designed, which captures local details and global structural information in the image by separating high-frequency and low-frequency information, thereby improving the detection accuracy; the DSSFF module is also designed, which combines the Scale Sequence FeatureFusionModule and DySample to enhance the multi-scale information extraction capability of the model; in addition, the TripleFeature Encoding is introduced Module (TFE) splices features of different sizes in the spatial dimension to further improve the accuracy of target detection; in the Head part, the P2 detection head is introduced to improve the model's detection accuracy for small targets.

[0086] The embodiment of the present invention also provides a sesame plant fruit testing device, comprising:

[0087] The image module is used to obtain the image dataset of sesame plants.

[0088] Model module, build the RT-DETR network model, introduce the RepNCSPELAN4 module, ADown module and ContextAggregation module into the BackBone part of the RT-DETR network model, and merge the ADown module and the Haarwavelet downsampling module into the HWD-ADown module and then introduce it into the BackBone part, and merge the RepNCSPELAN4 module, ADown module, Context Aggregation module and HWD-ADown module to replace the original ResNet structure in the BackBone part of the RT-DETR network model and merge them with the BackBone part to form an improved RAHC-BackBone part.

[0089] In the Efficient Hybrid Encoder part of the RT-DETR network model, the Multi-headAttention module in the AIFI module is replaced with the HiLoAttention module, the HiLoAttention module and the AIFI module are merged into the HiLo-AIFi module and introduced into the Efficient Hybrid Encoder part, the SSFF module and the DySample module are merged into the DSSFF module and introduced into the Efficient Hybrid Encoder part, and the TFE module is introduced and merged with the HiLo-AIFi module and the DSSFF module into the improved HTD-Efficient Hybrid Encoder part.

[0090] The P2 layer detection head is introduced into the Head part of the RT-DETR network model, and the P2 layer detection head and the Head part are merged into the improved P2-Head part.

[0091] The improved RAHC-BackBone part, the improved HTD-Efficient Hybrid Encoder part and the improved P2-Head part are fused to obtain the LEHP-DETR network model that is improved on the RT-DETR network model; and the LEHP-DETR network model is trained using the sesame plant image dataset to obtain the trained LEHP-DETR network model.

[0092] The recognition module inputs the image of sesame plants into the trained LEHP-DETR network model; through the improved RAHC-BackBone part, the feature map with significant edge information and detail information is obtained; through the improved HTD-Efficient Hybrid Encoder part, the feature map with significant contrast between background and target fruit is obtained; through the improved P2-Head part, the fruit species of sesame plants are obtained.

[0093] An embodiment of the present invention further provides an electronic device, including a memory and a processor.

[0094] The memory is used to store computer programs.

[0095] When the processor is used to execute the computer program stored in the memory, the steps of the above-mentioned method for testing the fruit of sesame plants are implemented.

[0096] An embodiment of the present invention further provides a computer-readable storage medium for storing a computer program, which, when executed by a processor, implements the steps of the above-mentioned method for testing the fruit of sesame plants.

[0097] The above-mentioned embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.

Claims

1. A method for testing the fruit of sesame plants, characterized in that: The following steps are involved: Obtain an image dataset of sesame plants; Construct an RT-DETR network model, introduce the RepNCSPELAN4 module, ADown module and Context Aggregation module into the BackBone part of the RT-DETR network model, merge the ADown module and the Haar waveletdownsampling module into the HWD-ADown module and then introduce it into the BackBone part, and merge the RepNCSPELAN4 module, ADown module, Context Aggregation module and HWD-ADown module to replace the original ResNet structure in the BackBone part of the RT-DETR network model and merge them into the BackBone part to form an improved RAHC-BackBone part; In the Efficient Hybrid Encoder part of the RT-DETR network model, the Multi-headAttention module in the AIFI module is replaced with the HiLo Attention module, the HiLo Attention module and the AIFI module are merged into the HiLo-AIFi module and introduced into the Efficient Hybrid Encoder part, the SSFF module and the DySample module are merged into the DSSFF module and introduced into the Efficient Hybrid Encoder part, and the TFE module is introduced and merged with the HiLo-AIFi module and the DSSFF module into the improved HTD-Efficient Hybrid Encoder part; the TFE module is a TripleFeature Encoding module; The P2 layer detection head is introduced into the Head part of the RT-DETR network model, and the P2 layer detection head and the Head part are merged into the improved P2-Head part; The improved RAHC-BackBone part, the improved HTD-Efficient Hybrid Encoder part and the improved P2-Head part are integrated to obtain the LEHP-DETR network model which is an improvement on the RT-DETR network model. and using a sesame plant image dataset to train the LEHP-DETR network model to obtain a trained LEHP-DETR network model; The image of sesame plants is input into the trained LEHP-DETR network model; through the improved RAHC-BackBone part, the feature map with significant edge information and detail information is obtained; through the improved HTD-EfficientHybrid Encoder part, the feature map with significant contrast between background and target fruit is obtained; through the improved P2-Head part, the fruit species of sesame plants are obtained.

2. The method for testing the fruit of sesame plants according to claim 1, characterized in that: The step of obtaining an image dataset of sesame plants comprises: Use an industrial camera to take multiple images of sesame plants; The Labelmg tool was used to annotate sesame capsules in multiple sesame plant images and to construct an image dataset of sesame plants.

3. The method for testing the fruit of sesame plants according to claim 1, characterized in that: The RepNCSPELAN4 module includes: The RepNCSPELAN4 module includes a RepNCSP module and an ELAN module; The RepNCSP module is used to extract features, and the ELAN module is used to enhance the recognition of targets.

4. The method for testing the fruit of sesame plants according to claim 1, characterized in that: The HiLo-AIFi module includes: Replace the Multi-head Attention module in the AIFI module with the HiLo Attention module, and merge the HiLoAttention module and the AIFI module to form the HiLo-AIFi module; The Multi-head Attention module relies on the self-attention mechanism to generate an attention distribution by calculating the similarity between each element in the input sequence and all other elements; The HiLo Attention module divides attention into two parts: a low-frequency part and a high-frequency part. The low-frequency part is used to capture global structural information, and the high-frequency part is used to extract local detail information.

5. The method for testing the fruit of sesame plants according to claim 1, characterized in that: The TFE module comprises: The TFE module is constructed by taking three tensors , and ,Will The height and width of the tensor are used as the target size, and the adaptive pooling operation is used to convert Adjust to The same size, and then add them after maximum pooling and average pooling; Use interpolation to Adjust to The same size and , and According to the channel dimension, dim=1 is concatenated into a tensor; the formula for obtaining the output feature map of the TFE module is as follows: ; in: The feature map representing the output of the TFE module; , and Representing feature maps of large, medium, and small sizes, respectively; Depend on , and Splicing obtained; and The resolution is the same, and the number of channels is Three times.

6. The method for testing the fruit of sesame plants according to claim 1, characterized in that: The P2 layer detection head has a resolution of 160×160 pixels and is downsampled twice in the backbone network.

7. The method for testing the fruit of sesame plants according to claim 1, characterized in that: A ContextAggregation module is inserted between the RAHC-BackBone part and the HTD-Efficient Hybrid Encoder part of the LEHP-DETR network model; The ContextAggregation module is used to enable the high-level feature map to retain the detail information in the low-level feature map.

8. A device for testing the fruit of sesame plants, characterized in that: include: An image module, used to obtain an image dataset of sesame plants; Model module, build RT-DETR network model, introduce RepNCSPELAN4 module, ADown module and Context Aggregation module into BackBone part of RT-DETR network model, merge ADown module and Haarwavelet downsampling module into HWD-ADown module and then introduce into BackBone part, merge RepNCSPELAN4 module, ADown module, Context Aggregation module and HWD-ADown module and then replace the original ResNet structure in BackBone part of RT-DETR network model and merge with BackBone part to form improved RAHC-BackBone part; In the Efficient Hybrid Encoder part of the RT-DETR network model, the Multi-headAttention module in the AIFI module is replaced with the HiLo Attention module, the HiLo Attention module and the AIFI module are merged into the HiLo-AIFi module and introduced into the Efficient Hybrid Encoder part, the SSFF module and the DySample module are merged into the DSSFF module and introduced into the Efficient Hybrid Encoder part, and the TFE module is introduced and merged with the HiLo-AIFi module and the DSSFF module into the improved HTD-Efficient Hybrid Encoder part; the TFE module is a TripleFeature Encoding module; The P2 layer detection head is introduced into the Head part of the RT-DETR network model, and the P2 layer detection head and the Head part are merged into the improved P2-Head part; The improved RAHC-BackBone part, the improved HTD-Efficient Hybrid Encoder part and the improved P2-Head part are integrated to obtain the LEHP-DETR network model which is an improvement on the RT-DETR network model. and using a sesame plant image dataset to train the LEHP-DETR network model to obtain a trained LEHP-DETR network model; The recognition module inputs the image of sesame plants into the trained LEHP-DETR network model; through the improved RAHC-BackBone part, the feature map with significant edge information and detail information is obtained; through the improved HTD-Efficient Hybrid Encoder part, the feature map with significant contrast between background and target fruit is obtained; through the improved P2-Head part, the fruit species of sesame plants are obtained.

9. An electronic device, characterized in that: include: Memory and processor; The memory is used to store computer programs; The processor is used to implement the steps of the sesame plant fruit testing method as described in any one of claims 1 to 7 when executing the computer program stored in the memory.

10. A computer-readable storage medium, characterized in that: Used to store a computer program, which, when executed by a processor, implements the steps of a method for testing the fruit of a sesame plant as described in any one of claims 1 to 7.