Handheld crop character detection system

By designing a handheld crop trait detection system, using a retractable camera rod, a 360-degree rotating camera and an improved YOLOv8 object detection algorithm, the problems of poor portability, strong subjectivity and insufficient accuracy in the prior art are solved, and efficient, accurate and consistent plant trait detection is achieved.

CN120101847APending Publication Date: 2025-06-06YUNNAN AGRICULTURAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510176801.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the prior art, plant trait detection equipment has poor portability and relies on subjective judgment of manual experience, and there are problems of environmental interference and insufficient accuracy.

Method used

A handheld crop trait detection system is designed, including a retractable camera rod, a 360-degree rotary camera and an embedded control panel, combined with the improved YOLOv8 object detection algorithm and EMA attention mechanism to achieve automated detection and image recognition.

Benefits of technology

The detection of plant traits with strong portability, high objectivity and strong resistance to environmental interference has been achieved, which significantly improves the accuracy and consistency of the detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120101847A_ABST
    Figure CN120101847A_ABST
Patent Text Reader

Abstract

The invention discloses a handheld crop character detection system, and relates to the technical field of agriculture. The system comprises a waistband, a camera rod, a platform, a camera, a bracket, a display, a protective shell, a power supply and a control panel, a ubuntu operating system is embedded in the control panel, and an improved YOLOv8 target detection algorithm is arranged on the operating system; the improvement is that a Faster Net network is utilized to replace a backbone network in an original model; adding an EMA attention mechanism in the neck network; plant character pictures are collected through a camera to form a data set, and crop character analysis is carried out on the data set through an improved YOLOv8 target detection algorithm. The method is convenient to wear and carry, and compared with a traditional subjective crop character judgment basis, the method has the advantages of being accurate and better in robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of agriculture, and in particular to a handheld crop trait detection system. Background Art

[0002] At present, the detection method of plant traits mainly relies on the experience and subjective judgment of professionals. Although there are some devices for detecting plant traits on the market, these devices are generally large and fixed and cannot be carried around. Inspectors often need to pick the plant parts to be tested, or move the whole plant to the equipment for analysis. This method is not only troublesome, but also causes damage to the picked parts, resulting in deviations between the test results and the actual values. For example, when testing leaf length or cluster extraction, the impact of the picking and processing process may affect the accuracy of the final result. In addition, traditional manual testing is limited by differences in personnel experience, and the judgment of the same trait may vary from person to person, resulting in inconsistent results. Human factors may interfere with the test results, especially when dealing with subtle quality differences, large errors are prone to occur, affecting the accuracy of the results.

[0003] Problems with existing technologies 1. Poor portability. Existing equipment is generally large and fixed. Inspectors cannot carry the equipment for on-site inspections and must pick parts of the plants. This is not only cumbersome, but may also damage the plants, resulting in deviations between the results and the actual situation. 2. Strong subjectivity. Traditional manual inspection methods are highly dependent on the experience and subjective judgment of inspectors. Different people may have different judgments on the same traits, resulting in inaccurate and unreliable results. 3. Environmental interference. When collecting images, due to the influence of factors such as lighting, angle, and environment, it may not be possible to obtain ideal image quality, thereby affecting the accuracy of the test results. 4. Insufficient accuracy. Traditional methods are difficult to accurately detect subtle quality differences, especially some difficult-to-observe plant traits, such as the color of the tip of the male spike leaf glumes, the color and length of the ear, etc., which can easily lead to missed detections or misjudgments. Summary of the invention

[0004] In order to solve the problems of poor portability of the above-mentioned existing equipment, strong subjectivity of traditional manual detection and insufficient interference accuracy, the present invention provides a handheld crop trait detection system.

[0005] To implement the above technical solution, the details are as follows:

[0006] A handheld plant trait detection system comprises: a waist fixing belt 1, on which a protective shell 7, a camera rod 2 and a bracket 5 are arranged;

[0007] One end of the camera rod 2 is hinged to the waist fixing belt 1 and can swing up and down; the other end is hinged to the platform 3, and the platform 3 can swing up and down 180 degrees; a camera 4 is arranged on the platform 3, and the camera can rotate 360 ​​degrees;

[0008] The camera pole 2 is composed of a telescopic pole, which can be extended in four sections, and the total extension length can reach four meters;

[0009] The camera rod 2 and the camera 4 can be fixed by screw threads;

[0010] A power supply 8 and a control panel 9 are arranged in the protective shell 7;

[0011] An operating screen 6 is provided on the bracket 5;

[0012] The power supply 8, the control panel 9, the camera 4 and the operation screen 6 are electrically connected;

[0013] Power supply 8 provides stable power support for the entire system;

[0014] The operation screen 6 presents the test results in real time;

[0015] Control panel 9 uses a control panel embedded with the Ubuntu operating system;

[0016] During use, the user adjusts the camera rod 2 and the platform 3 through the remote control 10, including: extension, rotation and tilt, to ensure the accuracy of detection and avoid data errors caused by improper position;

[0017] The phenotypic image of the crop is collected by the camera 4, and the collected image is input into the operation screen 6 for the phenotypic image recognition of the crop;

[0018] The operation screen 6 displays the acquisition system and the analysis system;

[0019] The operations performed by the system include:

[0020] S1, manually labeling the plant trait pictures collected by the camera 4, performing enhancement processing on the plant trait pictures and dividing the data sets;

[0021] The manual labeling method is: using the labelimg software embedded in the control panel 9 to manually label the collected plant trait pictures, wherein the labeling type is the crop appearance characteristic level;

[0022] The enhancement process is to randomly rotate, translate, scale, crop, and flip the collected images to simulate different environmental changes in actual scenes;

[0023] The data set is divided into training set, test set and validation set in a ratio of 8:1:1;

[0024] S2, analyzing crop traits of the dataset using the improved YOLOv8 target detection algorithm in control panel 9;

[0025] Improvements to the YOLOv8 target detection algorithm include: using the FasterNet network to replace the backbone network module in the original model;

[0026] Add EMA attention mechanism in the neck network module to improve the feature extraction ability of the model;

[0027] The plant trait pictures collected by camera 4 are sequentially passed through the backbone network, neck network and head network in the improved YOLOv8 target detection algorithm;

[0028] The steps of trait analysis are as follows:

[0029] S2.1, performing image feature segmentation operation on the collected plant trait pictures;

[0030] S2.2, process the image after the feature segmentation operation, the steps are as follows:

[0031] S2.2.1, the image after the feature segmentation operation is increased in nonlinear expression ability through the basic stage 1;

[0032] Foundation Phase 1 consists of three parts:

[0033] (1) In the basic stage, partial convolution is used to perform 3×3 convolution calculations on some channels of the image, and the remaining channels remain unchanged and are passed to the next layer, which can reduce a lot of calculations. The expression is as follows:

[0034] 0 = r × C 2 ×H×W

[0035] In the formula, C is the number of channels, H is the image height, W is the image width, and r is the partial calculation ratio;

[0036] (2) Use point-by-point convolution to perform a 1×1 convolution operation on the image obtained in (1) to exchange information from different channels;

[0037] (3) Normalize the output image to reduce the internal variable offset, and then process it through the activation function to increase the nonlinear expression ability of the image;

[0038] S2.2.2, the image after basic stage 1 is downsampled through the image feature fusion 1 operation;

[0039] Image feature fusion 1 is C 1 Channel, downsample to get 1 / 8 size image;

[0040] S2.2.3, the image after the image feature fusion 1 operation goes through the basic stage 2;

[0041] The steps of image processing in basic stage 2 are similar to those in basic stage 1, except that the image processed in basic stage 2 is downsampled to 1 / 8 size after image feature fusion 1;

[0042] S2.2.4, the image after basic stage 2 is downsampled through the image feature fusion 2 operation;

[0043] Image feature fusion 2 is C 2 Channel, downsample to get 1 / 16 size image;

[0044] S2.2.5, the image after the image feature fusion 2 operation goes through the basic stage 3;

[0045] The steps of image processing in basic stage 3 are similar to those in basic stage 1 and basic stage 2, except that the image processed in basic stage 3 is downsampled to 1 / 16 size after image feature fusion 1;

[0046] S2.2.6, the image after basic stage 3 is downsampled through the image feature fusion 3 operation;

[0047] Image feature fusion 3 is C 3 Channel, downsample to get 1 / 32 size image;

[0048] S2.2.7, perform feature fusion pyramid operation on image features fused with 1 to 3 images;

[0049] The processed image is fused with 1 to 3 features through the feature fusion pyramid for cross-layer fusion, so that the result can focus on both detail features and global features.

[0050] S2.3. Input the processed feature segmentation image into the neck network to obtain the final output feature map. The steps are as follows:

[0051] S2.3.1, the image after the feature fusion pyramid operation is upsampled and input into the joint module with the image obtained in the basic stage 3;

[0052] The upsampling operation uses the nearest neighbor interpolation method to fill the missing data positions by copying the nearest pixels of the original data and increase the image space size;

[0053] The joint module concatenates the two sets of images according to the channel dimension. After the joint, the spatial size of the image remains unchanged and the number of channels is added;

[0054] S2.3.2, the result of S2.3.1 is input into the fusion layer, and after upsampling operation, it is input into the joint module with the image obtained in the basic stage 2;

[0055] The fusion layer includes 1×1 dimensionality reduction convolution, bottleneck layer, joint module and 1×1 fusion channel convolution. After the image enters, it is convolved for dimensionality reduction and then extracted by 3×3 convolution of the bottleneck layer. The bottleneck layer processing results are spliced ​​by the joint module, and finally the feature channel is fused by 1×1 convolution.

[0056] S2.3.3, input the result of S2.3.2 into the fusion layer, and output the detection result through the detection head 1 after the EMA attention mechanism;

[0057] The EMA attention mechanism is as follows: after the feature input, it will be decomposed into sub-features, and the feature maps of the sub-features will be pooled along the height and width directions respectively. Then, the two one-dimensional vectors are concatenated and transformed using a 1×1 convolution to obtain a two-dimensional matrix. The two-dimensional matrix is ​​activated by the sigmoid function, and the output is multiplied element-by-element with the original feature map to obtain a weighted feature map. At the same time, the original feature map is transformed using a 3×3 convolution to obtain another transformed feature map. Next, the weighted and transformed feature maps are pooled, and the softmax function is used to calculate two attention weight matrices. At the same time, the two attention weight matrices are fused using matrix multiplication, and the result is adjusted to the same shape as the original feature map. Finally, the fused attention weights are activated using the sigmoid function and multiplied with the original grouped feature map to obtain the final output feature map.

[0058] The detection head discriminates the input features. After receiving the data, the detection head will have two branches. One branch is used for target recognition, which consists of two 3×3 convolutions and one 1×1 convolution. This branch calculates the number of feature channels and obtains the border loss function, including two parts CIoU and DFL, and finally outputs the border position and confidence of the detected target. The other branch is used for classification detection, which consists of two 3×3 convolutions and one 1×1 convolution. It calculates the number of categories NC, measures it with the BCE binary cross entropy loss function, and finally outputs the category.

[0059] S2.3.4, input the result of S2.3.2 into the fusion layer, after passing through the EMA attention mechanism, it is input into the joint module together with the result of S2.3.1 of the input fusion layer through the convolution operation, and after passing through the fusion layer, the detection result is output from the detection head 2 through the EMA attention mechanism;

[0060] S2.3.5. Input the result of S2.3.2 into the fusion layer, and after passing through the EMA attention mechanism, it is input into the joint module together with the result of S2.3.1 of the input fusion layer through the convolution operation. After passing through the fusion layer, it is input into the joint module together with the feature fusion pyramid operation through the EMA attention mechanism. After passing through the fusion layer, it is output from the detection head 3 through the EMA attention mechanism.

[0061] After the system obtains the final test information, it compares the test results with the trait analysis standards to obtain the specific analysis results of the traits. Then, the system stores the source image information and analysis results in text form in the established MySQL database for subsequent use.

[0062] Beneficial effects of the present invention:

[0063] Compared with the existing technology, the present invention is more portable; compared with the existing equipment, the present invention avoids the tedious steps of picking parts of plants in the traditional method, reduces damage to plants, and avoids the result deviation caused by improper operation. For example, when collecting plant height information, the traditional method requires cutting the plants or transplanting them to a designated collection room for collection, while the present device can be worn on the body and directly collect plants in the field.

[0064] Compared with the prior art, the present invention is more objective; the traditional manual detection method is highly dependent on the experience and subjective judgment of the detection personnel, and different personnel may have different judgments on the same trait, thus affecting the accuracy and reliability of the detection results. The present invention uses deep learning to make rational judgments on the traits, ensuring that the results are accurate and consistent each time, and ensuring that the detection results are more objective. Taking the color analysis of plant leaves as an example, two pictures with similar colors are difficult to distinguish with the naked eye, but the detection model can generate a color histogram through its tiny color feature changes, and the two photos can be distinguished by comparing the histograms.

[0065] Compared with the prior art, the present invention has strong resistance to environmental interference. During the image acquisition process, factors such as illumination, angle and environment may affect the image quality and thus the detection accuracy. The present invention has strong environmental adaptability and can effectively overcome these external interferences to ensure the image quality and accuracy of the detection results. The realization of this function mainly relies on the EMA attention mechanism in the analysis model. The attention mechanism is used to focus the model extraction on the key position of target extraction, reduce the weight of the part affected by the environment, and thus improve the environmental adaptability.

[0066] Compared with the prior art, the present invention has high analysis accuracy: traditional methods are difficult to achieve high accuracy when detecting subtle quality differences, especially for some difficult-to-observe plant traits, such as the color of the tip of the male spike leaf glumes, which are prone to missed detection or misjudgment. The present invention can accurately capture these subtle differences and significantly improve the accuracy of detection. By using slice-assisted detection, the image is divided into small pieces for separate detection during detection, which can accurately detect details that are ignored during overall detection and improve the analysis accuracy of small targets. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 It is a schematic diagram of the overall structure of the present invention;

[0068] Figure 2 It is a schematic diagram of each module of the system of the present invention;

[0069] Figure 3 It is the flow chart of the system;

[0070] Figure 4 This is the improved yolov8 structure diagram of the present invention;

[0071] Figure 5 Schematic diagram for improving the model effect, where part a is the input image; part b is the confidence of yolov8 detection; part c is the heat map of yolov8 detection; part d is the confidence of the detection of the present invention; part e is the heat map of the detection of the present invention;

[0072] In the accompanying drawings: 1. belt; 2. camera pole; 3. platform; 4. camera; 5. bracket; 6. display; 7. protective case; 8. power supply; 9. control panel. DETAILED DESCRIPTION

[0073] The present invention is further described in detail below in conjunction with specific embodiments.

[0074] like Figure 1 As shown, a handheld crop trait detection system comprises: a waist fixing belt 1, on which a protective shell 7, a camera rod 2 and a bracket 5 are arranged;

[0075] One end of the camera rod 2 is hinged to the waist fixing belt 1 and can swing up and down; the other end is hinged to the platform 3, and the platform 3 can swing up and down 180 degrees; a camera 4 is arranged on the platform 3, and the camera can rotate 360 ​​degrees;

[0076] The camera pole 2 is composed of a telescopic pole, which can be extended in four sections, and the total extension length can reach four meters;

[0077] The camera rod 2 and the camera 4 can be fixed by screw threads;

[0078] Camera 4 can be a depth camera or an RGB camera;

[0079] A power supply 8 and a control panel 9 are arranged in the protective shell 7;

[0080] An operating screen 6 is provided on the bracket 5;

[0081] The power supply 8, the control panel 9, the camera 4 and the operation screen 6 are electrically connected;

[0082] Power supply 8 provides stable power support for the entire system, ensuring that each part operates efficiently during the working process;

[0083] The operation screen 6 presents the test results in real time, making it convenient for users to view and make necessary adjustments;

[0084] Control panel 9 uses a control panel embedded with the Ubuntu operating system;

[0085] like Figure 2 As shown, during use, the user adjusts the camera rod 2 and the platform 3 through the remote control 10, including: extension, rotation and tilt, to ensure the accuracy of detection and avoid data errors caused by improper position;

[0086] The phenotypic image of the crop is collected by the camera 4, and the collected image is input into the operation screen 6 for the phenotypic image recognition of the crop;

[0087] The operation screen 6 displays the acquisition system and the analysis system, such as Figure 3 As shown;

[0088] The collection system includes: a rice trait collection module, a corn trait collection module and a data management module; after creating a new project, select the traits to be collected, including: a rice trait collection module, a corn trait collection module, and the corn trait collection module or the rice trait collection module records the traits related to corn or rice. After selecting the traits to be collected, take a photo through a camera 4. If the photo taken is not the trait part, the camera rod 2 and the platform 3 need to be adjusted through a remote control 10. If it is a correct photo, the correct photo is saved in a control panel 9 through the data management module. Exiting the project can exit to the initial interface of the crop trait collection and analysis system so as to select the trait collection system and create a new project and trait analysis system;

[0089] The analysis system includes: a rice trait analysis module, a corn trait analysis module, an analysis module, a result display module and a data management module. The corn trait analysis module or the rice trait analysis module stores the traits of corn or rice, and the corresponding collected photos and the traits to be analyzed can be analyzed. The result display module displays the trait pictures to be analyzed and the results of the display analysis are saved in the control panel 9 through the data management module. The exit project can exit to the initial interface of the crop trait collection and analysis system to select the trait collection system and the trait analysis system;

[0090] The operations performed by the system include:

[0091] S1, manually labeling the plant trait pictures collected by the camera 4, performing enhancement processing on the plant trait pictures and dividing the data sets;

[0092] The manual labeling method is: using the labelimg software embedded in the control panel 9 to manually label the collected plant trait pictures, wherein the labeling type is the crop appearance characteristic level;

[0093] The enhancement process is to randomly rotate, translate, scale, crop, and flip the collected images to simulate different environmental changes in actual scenes;

[0094] For example, randomly rotating the image helps the model adapt to changes in the target in different directions. By randomly cropping the image and retaining the target area, the model can learn the subtle information of the target. The color and brightness adjustment of the image simulates the changes under different lighting conditions, so that the model can maintain good recognition capabilities in different environments. Adding noise processing can help the model maintain high accuracy when facing imperfect images in real environments. Perspective transformation can simulate image changes under different perspectives, so that the model can cope with changes in the angle of the object. Through these enhancements, diverse samples can be generated during the training process, thereby improving the robustness and accuracy of the model;

[0095] The data set is divided into training set, test set and validation set in a ratio of 8:1:1;

[0096] In this embodiment, the processed data set contains a total of 2648 images.

[0097] S2, analyzing crop traits of the dataset using the improved YOLOv8 target detection algorithm in control panel 9;

[0098] like Figure 4 As shown in the figure, the improvements to the YOLOv8 target detection algorithm include: replacing the backbone network module in the original model with the Faster Net network; adding the EMA attention mechanism to the backbone network module and the detection head module to improve the feature extraction capability of the model;

[0099] The plant trait pictures collected by camera 4 are sequentially passed through the backbone network, neck network and head network in the improved YOLOv8 target detection algorithm;

[0100] The steps of trait analysis are as follows:

[0101] S2.1, performing image feature segmentation operation on the collected plant trait pictures;

[0102] The plant trait image collected by camera 4 is an RGB image of 224×224×3. After feature segmentation, the image size is 56×56×64. The feature segmentation is: using a convolution with a convolution kernel of 3×3 to reduce the spatial size of the input data and expand the number of channels of the input data;

[0103] S2.2, process the image after the feature segmentation operation, the steps are as follows:

[0104] S2.2.1, the image after the feature segmentation operation is increased in nonlinear expression ability through the basic stage 1;

[0105] Foundation Phase 1 consists of three parts:

[0106] (1) In the basic stage, partial convolution is used to perform 3×3 convolution calculations on some channels of the image, and the remaining channels remain unchanged and are passed to the next layer, which can reduce a lot of calculations. The expression is as follows:

[0107] 0 = r × C 2 ×H×W

[0108] Wherein, C is the number of channels, H is the image height, W is the image width, and r is the partial calculation ratio (1 / 4 in this embodiment);

[0109] (2) Use point-by-point convolution to perform a 1×1 convolution operation on the image obtained in (1) to exchange information from different channels;

[0110] (3) Normalize the output image to reduce the internal variable offset, and then process it through the activation function to increase the nonlinear expression ability of the image;

[0111] S2.2.2, the image after basic stage 1 is downsampled through the image feature fusion 1 operation;

[0112] Image feature fusion 1 is C 1 Channel, downsample to get 1 / 8 size image;

[0113] S2.2.3, the image after the image feature fusion 1 operation goes through the basic stage 2;

[0114] The steps of image processing in basic stage 2 are similar to those in basic stage 1, except that the image processed in basic stage 2 is downsampled to 1 / 8 size after image feature fusion 1;

[0115] S2.2.4, the image after basic stage 2 is downsampled through the image feature fusion 2 operation;

[0116] Image feature fusion 2 is C 2 Channel, downsample to get 1 / 16 size image;

[0117] S2.2.5, the image after the image feature fusion 2 operation goes through the basic stage 3;

[0118] The steps of image processing in basic stage 3 are similar to those in basic stage 1 and basic stage 2, except that the image processed in basic stage 3 is downsampled to 1 / 16 size after image feature fusion 1;

[0119] S2.2.6, the image after basic stage 3 is downsampled through the image feature fusion 3 operation;

[0120] Image feature fusion 3 is C3 Channel, downsample to get 1 / 32 size image;

[0121] S2.2.7, perform feature fusion pyramid operation on image features fused with 1 to 3 images;

[0122] The processed image is fused with 1 to 3 features through the feature fusion pyramid for cross-layer fusion, so that the result can focus on both detail features and global features.

[0123] S2.3. Input the processed feature segmentation image into the neck network to obtain the final output feature map. The steps are as follows:

[0124] S2.3.1, the image after the feature fusion pyramid operation is upsampled and input into the joint module with the image obtained in the basic stage 3;

[0125] The upsampling operation uses the nearest neighbor interpolation method to fill the missing data positions by copying the nearest pixels of the original data and increase the image space size;

[0126] The joint module concatenates the two sets of images according to the channel dimension. After the joint, the spatial size of the image remains unchanged and the number of channels is added;

[0127] S2.3.2, the result of S2.3.1 is input into the fusion layer, and after upsampling operation, it is input into the joint module with the image obtained in the basic stage 2;

[0128] The fusion layer includes 1×1 dimensionality reduction convolution, bottleneck layer, joint module and 1×1 fusion channel convolution. After the image enters, it is convolved for dimensionality reduction and then extracted by 3×3 convolution of the bottleneck layer. The bottleneck layer processing results are spliced ​​by the joint module, and finally the feature channel is fused by 1×1 convolution.

[0129] S2.3.3, input the result of S2.3.2 into the fusion layer, and output the detection result through the detection head 1 after the EMA attention mechanism;

[0130] The EMA attention mechanism is as follows: after the feature input, it will be decomposed into sub-features, and the feature maps of the sub-features will be pooled along the height and width directions respectively. Then, the two one-dimensional vectors are concatenated and transformed using a 1×1 convolution to obtain a two-dimensional matrix. The two-dimensional matrix is ​​activated by the sigmoid function, and the output is multiplied element-by-element with the original feature map to obtain a weighted feature map. At the same time, the original feature map is transformed using a 3×3 convolution to obtain another transformed feature map. Next, the weighted and transformed feature maps are pooled, and the softmax function is used to calculate two attention weight matrices. At the same time, the two attention weight matrices are fused using matrix multiplication, and the result is adjusted to the same shape as the original feature map. Finally, the fused attention weights are activated using the sigmoid function and multiplied with the original grouped feature map to obtain the final output feature map.

[0131] The detection head discriminates the input features. After receiving the data, the detection head will have two branches. One branch is used for target recognition, which consists of two 3×3 convolutions and one 1×1 convolution. This branch calculates the number of feature channels and obtains the border loss function, including two parts CIoU and DFL, and finally outputs the border position and confidence of the detected target. The other branch is used for classification detection, which consists of two 3×3 convolutions and one 1×1 convolution. It calculates the number of categories NC, measures it with the BCE binary cross entropy loss function, and finally outputs the category.

[0132] S2.3.4, input the result of S2.3.2 into the fusion layer, after passing through the EMA attention mechanism, it is input into the joint module together with the result of S2.3.1 of the input fusion layer through the convolution operation, and after passing through the fusion layer, the detection result is output from the detection head 2 through the EMA attention mechanism;

[0133] S2.3.5. Input the result of S2.3.2 into the fusion layer, and after passing through the EMA attention mechanism, it is input into the joint module together with the result of S2.3.1 of the input fusion layer through the convolution operation. After passing through the fusion layer, it is input into the joint module together with the feature fusion pyramid operation through the EMA attention mechanism. After passing through the fusion layer, it is output from the detection head 3 through the EMA attention mechanism.

[0134] After the system obtains the final test information, it compares the test results with the trait analysis standards to obtain the specific analysis results of the trait. Then, the system stores the source image information and analysis results in text form in the established MySQL database for subsequent use.

[0135] like Figure 5As shown, part a is the input image, part b is the confidence of yolov8 detection, and part d is the confidence of the detection of the present invention; it can be seen that under the same detection conditions, the confidence of the detection of the present invention is higher, and the properties shown in part e are more complete and have better effects compared with the heat map of yolov8 detection as shown in part c.

[0136] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit the scope of protection of the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the essence and scope of the technical solution of the present invention.

Claims

1. A handheld crop trait detection system, characterized in that: include: A waist fixing belt (1) is provided with a protective shell (7), a camera rod (2) and a bracket (5); One end of the camera rod (2) is hinged to the waist fixing belt (1) and can swing up and down; the other end is hinged to the platform (3) and the platform (3) can swing up and down; a camera (4) is arranged on the platform (3) and the camera (4) can rotate; A power supply (8) and a control panel (9) are arranged in the protective shell (7); An operating screen (6) is provided on the bracket (5); The power supply (8), the control panel (9) and the operation screen (6) are electrically connected; The operation screen (6) displays a collection system and an analysis system; The operations performed by the system include: S1, manually labeling the plant trait images collected by the camera (4), performing enhancement processing on the plant trait images, and dividing the data sets; S2. Analyze crop traits of the data set using the improved YOLOv8 target detection algorithm in the control panel (9) to obtain crop trait detection results.

2. A handheld crop trait detection system according to claim 1, characterized in that: The improved YOLOv8 target detection algorithm includes: using the Faster Net network to replace the backbone network module in the original model; adding an EMA attention mechanism in the neck network module.

3. A handheld crop trait detection system according to claim 1, characterized in that: The steps of analyzing the crop traits of the data set using the improved YOLOv8 target detection algorithm in the control board (9) to obtain the crop trait detection results are as follows: S2.1, performing image feature segmentation operation on the collected plant trait pictures; S2.2, processing the image after the feature segmentation operation; S2.

3. Input the processed feature segmentation image into the neck network to obtain the final output feature map.

4. A handheld crop trait detection system according to claim 3, characterized in that: The steps of processing the image after the feature segmentation operation are as follows: S2.2.1, the image after the feature segmentation operation is increased in nonlinear expression ability through the basic stage 1; Foundation Phase 1 consists of three parts: (1) In the basic stage, partial convolution is used to perform 3×3 convolution calculation on some channels of the image. The expression is as follows: 0=r×C 2 ×H×W In the formula, C is the number of channels, H is the image height, W is the image width, and r is the partial calculation ratio; (2) Perform a 1×1 convolution operation on the image obtained in (1) using point-by-point convolution; (3) Normalize the output image; S2.2.2, the image after basic stage 1 is downsampled through the image feature fusion 1 operation; Image feature fusion 1 is C1 channel, and down-sampled to get 1 / 8 size image; S2.2.3, the image after the image feature fusion 1 operation goes through the basic stage 2; The steps of image processing in basic stage 2 are similar to those in basic stage 1, except that the image processed in basic stage 2 is downsampled to 1 / 8 size after image feature fusion 1; S2.2.4, the image after basic stage 2 is downsampled through the image feature fusion 2 operation; Image feature fusion 2 is C2 channel, and down-sampled to obtain 1 / 16 size image; S2.2.5, the image after the image feature fusion 2 operation goes through the basic stage 3; The steps of image processing in basic stage 3 are similar to those in basic stage 1 and basic stage 2, except that the image processed in basic stage 3 is downsampled to 1 / 16 size after image feature fusion 1; S2.2.6, the image after basic stage 3 is downsampled through the image feature fusion 3 operation; Image feature fusion 3 is C3 channel, and down-sampled to get 1 / 32 size image; S2.2.

7. Fuse the image features of 1 to 3 images to perform feature fusion pyramid operation.

5. The handheld crop trait detection system according to claim 3, characterized in that: The steps of inputting the processed feature segmentation image into the neck network to obtain the final output feature map are as follows: S2.3.1, the image after the feature fusion pyramid operation is upsampled and input into the joint module with the image obtained in the basic stage 3; S2.3.2, the result of S2.3.1 is input into the fusion layer, and after upsampling operation, it is input into the joint module with the image obtained in the basic stage 2; S2.3.3, input the result of S2.3.2 into the fusion layer, and output the detection result through the detection head 1 after the EMA attention mechanism; S2.3.4, input the result of S2.3.2 into the fusion layer, after passing through the EMA attention mechanism, it is input into the joint module together with the result of S2.3.1 of the input fusion layer through the convolution operation, and after passing through the fusion layer, the detection result is output from the detection head 2 through the EMA attention mechanism; S2.3.

5. Input the result of S2.3.2 into the fusion layer, and after passing through the EMA attention mechanism, it is input into the joint module together with the result of S2.3.1 of the input fusion layer through the convolution operation. After passing through the fusion layer, it is input into the joint module together with the feature fusion pyramid operation through the EMA attention mechanism. After passing through the fusion layer, it is output from the detection head 3 through the EMA attention mechanism.