Crop row navigation line identification method and system based on deep learning
By constructing an improved YOLOX-Tiny crop detection model and third-order polynomial curve fitting technology, the problem of low accuracy and inability to deal with bending rows of traditional crop row recognition methods is solved, and high-precision crop row detection and navigation line fitting are achieved.
Patent Information
- Application Number
- CN202510658328.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-21
AI Technical Summary
Traditional crop row recognition methods are based on image processing and geometric features, and are susceptible to changes in light and uneven crop growth state, resulting in low recognition accuracy and ineffective processing of curved or nonlinear crop rows.
Using a deep learning-based method, an improved YOLOX-Tiny crop detection model was constructed, and multi-scale features were extracted through backbone network, NAS-FPN and Head modules, and combined with the feature points of the third-order polynomial curve to fit the cooperative physical rows, the slope and intercept of the navigation line were determined.
It improves the accuracy of crop row detection and the accuracy of navigation line fitting, can effectively deal with the problem of bending crop rows, and enhances the adaptability and real-time nature of the system.
Smart Images

Figure CN120182941A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision, and particularly to a method and system for identifying crop row navigation lines based on deep learning. Background Art
[0002] The method for identifying crop row navigation lines based on deep learning is a method that uses deep learning algorithms, especially technologies such as convolutional neural networks (CNNs), to automatically identify and track the paths of crop rows in farmland. This method is usually used in agricultural automation equipment, such as driverless agricultural machinery or robots, to improve the efficiency of agricultural production.
[0003] With the development of technology, artificial intelligence technology has advanced by leaps and bounds, promoting the development of all walks of life towards unmanned and intelligent directions. To achieve intelligent agriculture, the accuracy and real-time performance of field automatic navigation technology are extremely important. Crop rows, as the iconic reference objects in the farmland environment, can guide the movement and operation of agricultural robots. An effective seedling row recognition algorithm can greatly improve the accuracy of agricultural robot navigation and operation, and can provide spatial constraints for subsequent precise fertilization, precise weeding, and precise control of intelligent agricultural machinery, thereby achieving the beneficial effects of saving fertilizers and pesticides, increasing the yield per mu, effectively reducing the labor intensity, reducing environmental pollution, reducing production costs, and improving resource utilization rate and the yield and quality of agricultural products.
[0004] However, traditional crop row recognition methods are usually based on image processing and geometric features, and are easily affected by changes in illumination and uneven crop growth states, resulting in low accuracy of crop row recognition. Under different illumination conditions and seasonal changes, the detection effect will significantly decline. Most traditional navigation line fitting methods use linear regression or Hough transform, and can only fit straight or approximately straight crop rows, and cannot effectively process curved or non-linear crop rows, making it difficult to meet the actual production and operation requirements. Summary of the Invention
[0005] In order to solve the technical problems that traditional crop row recognition methods are usually based on image processing and geometric features, are easily affected by changes in illumination and uneven crop growth states, resulting in low accuracy of crop row recognition, and under different illumination conditions and seasonal changes, the detection effect will significantly decline, and most traditional navigation line fitting methods use linear regression or Hough transform, and can only fit straight or approximately straight crop rows, and cannot effectively process curved or non-linear crop rows, making it difficult to meet the actual production and operation requirements, the present invention provides a method and system for identifying crop row navigation lines based on deep learning.
[0006] The technical solutions provided by the embodiments of the present invention are as follows:
[0007] First aspect:
[0008] A method for identifying crop row navigation lines based on deep learning provided by an embodiment of the present invention includes:
[0009] S1: Obtain crop image data; S2: Construct a crop detection model based on improved YOLOX-Tiny, where the crop detection model includes a backbone network module, a NAS-FPN module, and a Head module connected in sequence; S3: Input the crop image data into the crop detection model for training; S4: Obtain real-time crop image data; S5: Input the real-time crop image data into the trained crop detection model, and output the detection box data of the crop, where the detection box data includes the coordinates of the upper left corner and the lower right corner of the detection box; S6: Calculate the center point coordinates of the detection box according to the detection box data, and determine the feature points of the crop row; S7: Fit the feature points with a third-order polynomial curve to determine the curve equation of the crop row; S8: Based on the curve equation, determine the navigation line slope and the navigation line intercept for tracking the crop row; S9: Determine the position and direction of the navigation line according to the navigation line slope and the navigation line intercept.
[0010] Second aspect:
[0011] A crop row navigation line identification system based on deep learning provided by an embodiment of the present invention includes:
[0012] A processor; a memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the method for identifying crop row navigation lines based on deep learning as described in the first aspect is implemented.
[0013] Third aspect:
[0014] A computer-readable storage medium provided by an embodiment of the present invention, on which a computer program is stored, and when the program is executed by a processor, the method for identifying crop row navigation lines based on deep learning as described in the first aspect is implemented.
[0015] The beneficial effects brought by the technical solution provided by the embodiment of the present invention at least include:
[0016] In an embodiment of the present invention, by constructing a crop detection model based on the improved YOLOX-Tiny, the global feature capture capability of the crop row recognition model in complex backgrounds is enhanced, so that the model can better adapt to different environmental conditions and effectively improve the detection accuracy of crop rows. Targeted training is performed through the model to ensure the high accuracy of the trained model in actual detection, and the detection frame of the crop row can be accurately output to ensure accurate identification of the crop position and range. By fitting the feature points of the crop row through a third-order polynomial, the nonlinear structure of the crop row can be more accurately described, and the problem of curved crop rows can be effectively handled, thereby improving the fitting accuracy and actual adaptability of the navigation line. By determining the slope and intercept of the navigation line through a curve equation, the navigation line of the crop row can be accurately positioned, thereby improving navigation accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0018] Figure 1 A schematic flow chart of a crop row navigation line recognition method based on deep learning provided by an embodiment of the present invention;
[0019] Figure 2 A schematic diagram of a crop detection model structure provided by an embodiment of the present invention;
[0020] Figure 3 A schematic diagram of a crop row navigation line fitting process provided by an embodiment of the present invention;
[0021] Figure 4 A structural schematic diagram of a crop row navigation line recognition system based on deep learning provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0022] The technical solution of the present invention is described below in conjunction with the accompanying drawings.
[0023] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "example" is intended to present the concept in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.
[0024] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, their intended meanings are the same. "of", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, their intended meanings are the same.
[0025] In the embodiments of the present invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, their intended meanings are the same.
[0026] To make the technical problems, technical solutions, and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0027] Refer to the attached Figure 1 , which shows a schematic flowchart of a crop row navigation line recognition method provided by an embodiment of the present invention based on deep learning.
[0028] The embodiments of the present invention provide a crop row navigation line recognition method based on deep learning. This method can be implemented by a crop row navigation line recognition device based on deep learning. The crop row navigation line recognition device based on deep learning can be a terminal or a server. The processing flow of the crop row navigation line recognition method based on deep learning can include the following steps:
[0029] S1: Obtain crop image data.
[0030] In a possible implementation manner, S1 is specifically:
[0031] Use the Senyun SG2-IMX390C camera to obtain crop image data.
[0032] Among them, Senyun SG2-IMX390C is a high-performance industrial camera that uses Sony's IMX390C sensor, has high resolution, high sensitivity, and excellent performance in low-light environments, and is widely used in machine vision, automated inspection, agriculture and other fields, and can capture high-quality image data.
[0033] It should be noted that using the Senyun SG2-IMX390C camera to obtain crop image data has high resolution and low-light performance, can provide clear images under different lighting conditions, and the excellent imaging ability of the Senyun SG2-IMX390C camera ensures the accuracy of the data and the richness of details.
[0034] In the present invention, taking corn as an example, at the 3 - 5 leaf stage of corn, different weather conditions such as sunny and cloudy days are selected, and at different times of the day (such as 8 am, 12 noon, 6 pm, and 8 pm), videos are taken using the Senyun SG2 - IMX390C camera. The resolution of this camera is 1080P, the frame rate is 30fps, and GMSL2 is used as the interface. The camera is installed at a height of 1.5 meters from the agricultural machinery, and the shooting angle is between 30° and 60°. During the shooting process, factors such as leaf occlusion, straw residue coverage, plant gaps, weed interference, corn close planting, and night-time supplementary lighting are fully considered to obtain rich and diverse image data. The captured videos are converted into pictures frame by frame, and the original images are processed in terms of color, brightness, contrast, noise, rotation, and mirroring. Then, the professional image annotation software Labelme is used to annotate the pre - processed images, taking the corn seedlings as the annotation objects and marking their positions and ranges. After the annotation is completed, the data set is divided into a training set, a validation set, and a test set according to the ratio of 8:1:1 for subsequent model training and evaluation.
[0035] In a possible implementation manner, after S1, it further includes:
[0036] Pre - process the crop image data.
[0037] The pre - processing specifically includes: color processing, brightness adjustment, contrast adjustment, noise processing, rotation processing, and mirror processing.
[0038] Among them, color processing refers to adjusting the color information of the image, including changing color saturation, hue, or color temperature to better present the appearance of the crop. Brightness adjustment refers to increasing or decreasing the brightness of the image to improve the display effect of the image under different lighting conditions and ensure that details are visible. Contrast adjustment refers to adjusting the difference between the bright and dark parts of the image to make the details of the image more prominent and increase the visual impact of the image. Noise processing refers to removing noise points or unnecessary interferences in the image to improve the image quality and ensure the accuracy of subsequent processing and analysis. Rotation processing refers to rotating the image to adjust the angle or direction of the image to meet the analysis requirements. Mirror processing refers to performing horizontal or vertical flipping on the image to help generate data at different angles and enhance the diversity of training data.
[0039] It should be noted that pre - processing the crop image data can effectively improve the data quality. Through the adjustment of color, brightness, and contrast, problems with image quality under different lighting conditions can be addressed. Noise processing can remove unnecessary interferences, and rotation and mirror processing increase the diversity of the images, thereby improving the generalization ability of the deep learning model.
[0040] Refer to the attached Figure 2, showing a schematic diagram of the structure of a crop detection model provided by an embodiment of the present invention.
[0041] As Figure 2 shown, the structure of the crop detection model mainly includes three parts: the backbone network (Backbone), the feature pyramid network (NAS-FPN), and the detection head (Head). The backbone network consists of Fous units, CSP1_1 units, multiple CBS units, multiple CSP1_3 units, MobileViT units, CSP2_1 units, SPPF units, and Attention Block units, and is responsible for extracting multi-scale features from the input image. The feature pyramid network (NAS-FPN) is used to generate feature maps of different scales and transmit them to the classification box subnet to predict the category and location of the crops. The detection head processes the extracted features and outputs the detection results of the crops, such as the category and bounding box. Through this structure, the model can accurately detect crops in complex agricultural scenarios.
[0042] S2: Construct a crop detection model based on the improved YOLOX-Tiny, where the crop detection model includes a backbone network module, a NAS-FPN module, and a Head module connected in sequence.
[0043] Among them, YOLOX-Tiny is a lightweight version of the YOLO series, designed specifically for real-time object detection tasks, with relatively low computational complexity and fast inference speed. By improving the structure in the YOLO series, it can achieve high detection accuracy with limited computational resources. The backbone network module (Backbone) is the core part of the deep learning model and is responsible for extracting the basic features of the input image. The NAS-FPN module (Neural Architecture Search Feature Pyramid Network) is a feature pyramid network optimized by neural architecture search technology, used to effectively extract and fuse image features at multiple scales and improve the object detection accuracy. The Head module is the last part of the detection network and is responsible for classifying and regressing (such as bounding box prediction) based on the extracted features and outputting the final results of object detection.
[0044] It should be noted that by constructing a crop detection model based on the improved YOLOX-Tiny, combined with an efficient backbone network, NAS-FPN, and Head module, efficient crop detection can be achieved with limited computational resources. The backbone network module extracts image features, the NAS-FPN module enhances the fusion ability of multi-scale features, and the Head module accurately outputs the category and location of the crops, enabling the model to perform fast inference with high precision.
[0045] In a possible implementation, the backbone network module includes a Fous unit, a first CBS unit, a CSP1_1 unit, a second CBS unit, a first CSP1_3 unit, a MobileViT unit, a third CBS unit, a second CSP1_3 unit, a fourth CBS unit, a CSP2_1 unit, an SPPF unit, and an Attention Block unit, which are connected in sequence.
[0046] The SPPF unit includes a 5×5 max pooling operation subunit, a 9×9 max pooling operation subunit, a 13×13 max pooling operation subunit, and a basic convolution operation subunit.
[0047] Specifically, the Fous unit refers to a convolution unit with a specific design that may be responsible for feature extraction in this network. The CBS (Convolution, Batch Normalization, and Activation) unit refers to a common convolution layer structure that includes convolution operations, batch normalization, and activation functions (such as ReLU). The CSP (Cross-Stage Partial) unit is a module used to reduce computational complexity. By dividing the feature map into multiple parts, processing them separately, and then merging them, it can improve computational efficiency while maintaining performance. The MobileViT unit is a lightweight module designed based on VisionTransformer (ViT), which combines the advantages of convolutional neural networks (CNNs) and transformer models and can process more complex image features at low computational costs. The SPPF (Spatial Pyramid Pooling and Fusion) unit is used to extract information at different scales, can process input images of different sizes and fuse them, and enhance the model's ability to recognize multi-scale targets. The Attention Block unit uses the attention mechanism to help the model automatically focus on important parts of the image. Through this mechanism, the model can more effectively identify and process valuable feature information. The max pooling operation subunit refers to the max pooling operation applied to images. By using windows of different sizes (5×5, 9×9, 13×13) respectively, it extracts feature information at different scales from the image. The pooling operation helps to reduce the dimension, reduce the amount of computation, and at the same time retain important features. The basic convolution operation subunit refers to the basic operation in a convolutional neural network, which is used to extract local features of the input image. The convolution operation slides a filter (convolution kernel) over the input image and calculates the feature response.
[0048] In the present invention, after the first CSP1_3 module and the second CSP1_3 module, a part of the convolutional layers are selected and replaced with the lightweight Transformer of MobileViT. The MobileViT unit first divides the input feature map into multiple small patches, performs convolutional operations within each patch to extract local features. Then, the features of these patches are fed into the Transformer unit to capture global dependencies using the self-attention mechanism. Finally, the feature map size is restored through convolutional operations to match the number of channels in the subsequent part of the original network. In this way, without significantly increasing the computational load, the model's ability to capture the global features of crop rows is enhanced, and the detection accuracy in complex backgrounds is improved.
[0049] For example, assume that the number of input channels of the original convolutional layer is C in , and the number of output channels is C out . When replacing it with a MobileViT unit, first ensure that the number of input channels of the MobileViT unit is the same as C in . Inside the MobileViT unit, the input feature map is partitioned, the patch size is set to P×P, convolutional operations are performed within each patch, the convolutional kernel size is K×K, and the stride is . After the convolutional operations within the patches, the features are fed into the Transformer, the number of heads in the Transformer is set to H, and the dimension of the attention heads is set to D. After being processed by the Transformer, the feature map size is restored through convolutional operations so that its number of output channels is the same as C out , ensuring a smooth connection with the subsequent part of the original network.
[0050] In the present invention, a fast spatial pyramid pooling (SPPF) unit is added at the end of the backbone network to expand the receptive field and improve the target localization accuracy. The fast spatial pyramid pooling (SPPF) unit added at the end of the backbone network is composed of max-pooling operator sub-units with kernel sizes of 5×5, 9×9, and 13×13 and basic convolutions, which can expand the receptive field without significantly increasing the model size, enhance the expression ability of the feature map, and improve the target localization accuracy. To improve the multi-scale feature fusion effect, the relevant structure for generating the P2 feature map based on the FPN concept in the original model is abandoned, and the NAS-FPN module is adopted. The feature maps of different scales output by the backbone network are directly input into the NAS-FPN module. The NAS-FPN module automatically performs the fusion and processing of different-scale features through the optimal structure obtained by neural architecture search. Before inputting into the NAS-FPN module, the size and number of channels of the feature map are adjusted to meet the input requirements of the NAS-FPN module. After being processed by the NAS-FPN module, the output multi-scale feature map is directly connected to the subsequent prediction head for the detection of crop rows. This improved structure can more effectively fuse features of different scales and enhance the detection ability for crop rows of different sizes.
[0051] For example, the size and number of channels of the P3, P4, and P5 feature maps output by the backbone network often do not match the input requirements of the subsequent NAS-FPN module, so targeted adjustments are required. For the P3 feature map, size adjustment is usually achieved with the help of convolution operations. For example, if a convolution layer with a convolution kernel of 3×3 is used, if the size needs to be reduced, the step size is set to be greater than 1. For example, when the step size is set to 2, the height and width of the feature map can be reduced to half of the original size. If the size needs to remain unchanged, the step size is set to 1 and the padding is set to 1. In terms of channel number adjustment, 1×1 convolution is used. According to the input requirements of the NAS-FPN module, if the expected number of channels is less than the original number of channels of the P3 feature map, a smaller number of 1×1 convolution kernels are set for channel compression. Otherwise, the number of convolution kernels is increased to expand the number of channels. Similar processing is performed on the P4 and P5 feature maps. According to their respective original sizes and number of channels, as well as the input requirements of NAS-FPN, appropriate convolution operations and parameters are selected for adjustment. The P4 feature map is first reduced in size using a convolution layer with a stride of 2 and a convolution kernel of 3×3, and then the number of channels is adjusted using a 1×1 convolution layer. The P5 feature map itself is small in size, and the focus of adjustment is on optimizing the number of channels. The size can be kept stable through padding and convolution operations, and then the number of channels can be accurately adjusted using 1×1 convolution. After the adjusted P3, P4, and P5 feature maps are input into the NAS-FPN module, NAS-FPN will efficiently fuse and process these feature maps with the optimal structure obtained by its neural architecture search, intelligently integrate the information in feature maps of different scales, and enhance the feature expression of crop rows of different sizes. Finally, the NAS-FPN module outputs optimized multi-scale feature maps, which are directly connected to the subsequent prediction head. Based on these rich and well-integrated features, the prediction head detects crop rows and accurately identifies the location, size, and other information of crop rows, thereby providing accurate data support for subsequent navigation line extraction and fitting, effectively improving the performance of the entire crop row detection system.
[0052] S3: Input crop image data into the crop detection model for training.
[0053] It should be noted that by inputting crop image data into the crop detection model for training, the model can gradually learn how to identify crop characteristics through a large amount of labeled data. During the training process, the model will continuously optimize parameters and improve its ability to identify crops of different types and in different environments. This process enables the model to perform crop detection more accurately and effectively in practical applications, and has strong adaptability and generalization capabilities.
[0054] In a possible implementation, S3 is specifically:
[0055] The preprocessed crop image data is input into the crop detection model for training until the value of the loss function is less than the preset loss function value.
[0056] It should be noted that by inputting the processed crop image data into the crop detection model for training and continuously adjusting the model parameters until the loss function value reaches the preset target, the accuracy of the model can be ensured to continuously improve. The loss function is used to measure the difference between the model prediction value and the actual value. By continuously optimizing the model, the error can be effectively reduced, enabling the model to accurately identify crops in practical applications.
[0057] Among them, the loss functions specifically include: EIOU loss function, IoU-aware classification loss function, IoU-aware category loss function, cross-entropy loss function, and regularized bounding box loss function.
[0058] It should be noted that those skilled in the art can set the size of the preset loss function value according to the actual situation, and the present invention does not make any limitations.
[0059] S4: Obtain real-time crop image data.
[0060] Among them, the EIOU (Enhanced Intersection over Union) loss function is a loss function used for object detection, aiming to improve the overlap between the predicted bounding box and the ground truth bounding box. The IoU-aware classification loss function is a classification loss function that combines IoU (Intersection over Union) information, aiming to optimize the classification task of the detection model. The IoU-aware category loss function takes into account the impact of IoU information on category prediction. By incorporating IoU into the calculation of category prediction, it improves the accuracy of category prediction, especially having advantages in small object detection and precise localization tasks. The cross-entropy loss function is commonly used in classification problems and measures the difference between the predicted category and the actual category. For each sample, the cross-entropy calculates the difference between the predicted probability distribution and the true distribution, with the goal of minimizing this loss value to optimize the classification performance of the model. The regularized bounding box loss function is used for the bounding box regression task in object detection, aiming to reduce the difference between the bounding box predicted by the model and the ground truth bounding box. The purpose of regularization is to constrain the output of the model, avoid overfitting, and ensure the accuracy of the bounding box.
[0061] In the present invention, the EIOU loss function is adopted to replace the original loss function. Meanwhile, combining the IoU-aware classification loss and the IoU-aware category loss, the cross-entropy loss function is used for calculation, and the regularization bounding box loss function is added in the later stage of training.
[0062] S5: Input the real-time crop image data into the trained crop detection model, and output the detection box data of the crop, where the detection box data includes the coordinates of the upper left corner and the lower right corner of the detection box.
[0063] It should be noted that by inputting the real-time crop image data into the trained crop detection model, the position and size information of the crop in the image can be obtained in real time, the detection box data can be accurately output, the changes of the crop can be recognized in time and located, which greatly improves the efficiency and accuracy of agricultural operations.
[0064] S6: According to the detection box data, calculate the center point coordinates of the detection box to determine the characteristic points of the crop row.
[0065] Among them, the crop row refers to the row-column structure formed by crops arranged according to certain rules in agricultural planting. Usually, it has a consistent spacing and arrangement method. This arrangement method can improve the planting efficiency, facilitate management and mechanized operations, such as irrigation, fertilization, weeding and harvesting, etc. In modern agriculture, the identification and management of crop rows become particularly important. Especially when conducting precision agriculture operations through automated equipment, accurately identifying the position of crop rows is the key to ensuring crop growth and improving production efficiency.
[0066] Specifically, by calculating the center point coordinates of the detection box, the center position of each crop can be accurately determined, and the characteristic points of the crop row can be extracted from multiple detection boxes, improving the accuracy of crop row positioning and being able to capture the position of the crop row more stably.
[0067] Refer to the attached Figure 3 illustrates a schematic diagram of the fitting process of a crop row navigation line provided by an embodiment of the present invention.
[0068] As Figure 3 shown, it shows the fitting process of the crop row navigation line. First, the crop detection model is used to identify the crop area in the image and mark the detection box. Then, calculate the center point coordinates of each detection box to obtain the characteristic points of the crop row. Then, fit these characteristic points with a cubic polynomial curve to generate the fitting curve equation of the crop row. Based on this curve equation, further calculate the slope and intercept of the navigation line to determine the specific position and direction of the crop row, and finally draw the navigation line of the crop row. This process helps agricultural machinery accurately track the crop row and perform automated operations.
[0069] S7: Use a third-order polynomial curve to fit the feature points and determine the curve equation of the crop row.
[0070] Among them, the third-order polynomial curve is a mathematical expression used to fit a set of data points and can flexibly describe the non-linear relationship of the data.
[0071] It should be noted that using the third-order polynomial curve to fit the feature points can accurately capture the change trend of the crop row, generate an accurate curve equation, effectively reduce the fitting error, provide a stable navigation path, ensure that the agricultural machinery accurately tracks the crop row, and improve the efficiency of crop management and harvesting.
[0072] In a possible implementation manner, the curve equation is specifically:
[0073]
[0074] Among them, l i ( ) represents the fitting curve equation of the i-th crop row, x represents the feature point coordinates, and w 0,i , w 1,i , w 2,i and w 3,i all represent polynomial coefficients.
[0075] S8: Based on the curve equation, determine the slope and intercept of the navigation line for tracking the crop row.
[0076] Among them, the slope of the navigation line refers to the degree of inclination of the navigation line relative to the horizontal line, usually expressed as the ratio of the vertical change to the horizontal change of the line segment, and the slope determines the direction of the crop row. The intercept of the navigation line refers to the coordinates of the intersection point of the line and the vertical axis (usually the y-axis), indicating the starting position of the crop row.
[0077] It should be noted that determining the slope and intercept of the navigation line based on the curve equation can accurately describe the path direction and position of the crop row, provide the key parameters of the navigation line, help the agricultural equipment accurately track the crop row, thereby improving the accuracy and efficiency of the operation, and ensuring that the crop management and harvesting are more efficient and accurate.
[0078] In a possible implementation manner, S8 specifically includes:
[0079] S801: Simultaneously solve the curve equations of adjacent crop rows to determine the intersection coordinates of the curve equations.
[0080] S802: Determine the slope and intercept of the navigation line according to the intersection coordinates.
[0081] In a possible implementation manner, the calculation method of the slope of the navigation line is specifically:
[0082]
[0083] Among them, m c represents the slope of the navigation line, and b c represents the slope of the navigation line, x i represents the abscissa of the intersection point of the curve equation, and y i represents the ordinate of the intersection point of the curve equation.
[0084] The calculation method of the navigation line intercept is specifically as follows:
[0085]
[0086] Among them, tan represents the tangent function, arctan represents the arctangent function, m1 represents the slope of the first crop row curve, and m2 represents the slope of the second crop row curve.
[0087] S9: Determine the position and direction of the navigation line according to the slope of the navigation line and the navigation line intercept.
[0088] It should be noted that by determining the position and direction of the navigation line according to the slope and intercept of the navigation line, the accurate path of the crop row can be accurately calculated, enabling the agricultural machinery to accurately locate the trajectory of the crop row and operate efficiently along the predetermined path, avoiding deviation from the target, ensuring the accuracy and efficiency of the crop management and harvesting processes, and greatly improving the accuracy and stability of automated agricultural operations.
[0089] The beneficial effects brought by the technical solution provided by the embodiment of the present invention at least include:
[0090] In the embodiment of the present invention, by constructing a crop detection model based on the improved YOLOX-Tiny, the global feature capture ability of the crop row recognition model in complex backgrounds is enhanced, enabling the model to better adapt to different environmental conditions, effectively improving the detection accuracy of the crop row. Through targeted training of the model, the high accuracy of the model in actual detection is ensured, and the detection frame of the crop row can be accurately output, ensuring the accurate recognition of the position and range of the crop. By fitting the feature points of the crop row with a third-order polynomial, the non-linear structure of the crop row can be more accurately described, effectively handling the problem of curved crop rows, improving the fitting accuracy and practical adaptability of the navigation line. By determining the slope and intercept of the navigation line through the curve equation, the navigation line of the crop row can be accurately positioned, improving the navigation accuracy.
[0091] Refer to the attached Figure 4 illustrates the structural schematic diagram of a crop row navigation line recognition system provided by the present invention based on deep learning.
[0092] The present invention also provides a crop row navigation line recognition system 20 based on deep learning, which is applied to the above-mentioned crop row navigation line recognition method based on deep learning, and includes:
[0093] Processor 201.
[0094] Memory 202, on which computer-readable instructions are stored. When the computer-readable instructions are executed by the processor 201, the method for identifying crop row navigation lines based on deep learning as in the method embodiment is implemented.
[0095] The crop row navigation line identification system 20 provided by the present invention can execute the above-mentioned method for identifying crop row navigation lines based on deep learning and achieve the same or similar technical effects. To avoid repetition, the present invention will not elaborate further.
[0096] The beneficial effects brought by the technical solution provided by the embodiments of the present invention at least include:
[0097] In the embodiments of the present invention, by constructing a crop detection model based on the improved YOLOX-Tiny, the global feature capture ability of the crop row recognition model in complex backgrounds is enhanced, enabling the model to better adapt to different environmental conditions, effectively improving the detection accuracy of crop rows. Through targeted training of the model, the high accuracy of the trained model in actual detection is ensured, and the detection frame of the crop row can be accurately output, ensuring the accurate identification of the crop position and range. By fitting the feature points of the crop row with a third-order polynomial, the non-linear structure of the crop row can be more accurately described, effectively handling the problem of curved crop rows, improving the fitting accuracy and actual adaptability of the navigation line. By determining the slope and intercept of the navigation line through the curve equation, the navigation line of the crop row can be accurately positioned, improving the navigation accuracy.
[0098] It should be understood that the processor in the embodiments of the present invention may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.
[0099] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0100] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions according to the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0101] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context before and after.
[0102] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0103] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0104] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0105] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described devices, apparatuses, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0106] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be electrical, mechanical, or other forms.
[0107] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0108] In addition, the functional units in various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0109] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0110] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the crop row navigation line recognition method based on deep learning as in the method embodiment.
[0111] The computer-readable storage medium provided by the present invention can implement the steps and effects of the crop row navigation line recognition method based on deep learning in the above method embodiment. To avoid repetition, the present invention will not elaborate further.
[0112] The beneficial effects brought by the technical solution provided by the embodiments of the present invention at least include:
[0113] In the embodiments of the present invention, by constructing a crop detection model based on the improved YOLOX-Tiny, the global feature capture ability of the crop row recognition model in complex backgrounds is enhanced, enabling the model to better adapt to different environmental conditions, effectively improving the detection accuracy of crop rows. Through targeted training of the model, the high accuracy of the trained model in actual detection is ensured, and the detection frame of the crop row can be accurately output, ensuring the accurate recognition of the position and range of the crop. By fitting the feature points of the crop row with a third-order polynomial, the non-linear structure of the crop row can be more accurately described, effectively handling the problem of curved crop rows, improving the fitting accuracy and actual adaptability of the navigation line. By determining the slope and intercept of the navigation line through the curve equation, the navigation line of the crop row can be accurately positioned, improving the navigation accuracy.
[0114] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
[0115] The following points need to be explained:
[0116] (1)The accompanying drawings of the embodiments of the present invention only relate to the structures involved in the embodiments of the present invention, and other structures can refer to the general design.
[0117] (2)For clarity, in the accompanying drawings used to describe the embodiments of the present invention, the thickness of layers or regions is enlarged or reduced, that is, these drawings are not drawn to actual scale. It can be understood that when an element such as a layer, film, region or substrate is referred to as being "on" or "under" another element, the element can be "directly" on or under the other element or there can be intermediate elements.
[0118] (3)Without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other to obtain new embodiments.
[0119] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. The protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A crop row navigation line recognition method based on deep learning, characterized in that: include: S1: Acquire crop image data; S2: constructing a crop detection model based on the improved YOLOX-Tiny, wherein the crop detection model includes a backbone network module, a NAS-FPN module, and a Head module connected in sequence; S3: inputting the crop image data into the crop detection model for training; S4: Acquire real-time crop image data; S5: inputting the real-time crop image data into the trained crop detection model, and outputting crop detection frame data, wherein the detection frame data includes the coordinates of the upper left corner of the detection frame and the coordinates of the lower right corner of the detection frame; S6: Calculate the center point coordinates of the detection frame according to the detection frame data, and determine the feature points of the crop row; S7: using a third-order polynomial curve to fit the characteristic points and determine a curve equation of the crop row; S8: Determine a slope and an intercept of a navigation line for tracking the crop row based on the curve equation; S9: Determine the position and direction of the navigation line according to the slope of the navigation line and the intercept of the navigation line.
2. The crop row navigation line recognition method based on deep learning according to claim 1, characterized in that: The S1 is specifically: The crop image data is obtained using the Senyun SG2-IMX390C camera.
3. The crop row navigation line recognition method based on deep learning according to claim 1, characterized in that: After S1, it also includes: Preprocessing the crop image data; The preprocessing specifically includes: color processing, brightness adjustment, contrast adjustment, noise processing, rotation processing and mirror processing.
4. The crop row navigation line recognition method based on deep learning according to claim 1, characterized in that: The backbone network module includes a Fous unit, a first CBS unit, a CSP1_1 unit, a second CBS unit, a first CSP1_3 unit, a MobileViT unit, a third CBS unit, a second CSP1_3 unit, a fourth CBS unit, a CSP2_1 unit, a SPPF unit, and an Attention Block unit connected in sequence; The SPPF unit includes a 5×5 maximum pooling operation subunit, a 9×9 maximum pooling operation subunit, a 13×13 maximum pooling operation subunit and a basic convolution operation subunit.
5. The crop row navigation line recognition method based on deep learning according to claim 1, characterized in that: The S3 is specifically: Inputting the crop image data into the crop detection model for training until the value of the loss function is less than a preset loss function value; Among them, the loss functions specifically include: EIOU loss function, IoU-aware classification loss function, IoU-aware category loss function, cross entropy loss function and regularized bounding box loss function.
6. The crop row navigation line recognition method based on deep learning according to claim 1, characterized in that: The curve equation is specifically: ; Among them, l i ( ) represents the curve equation of the i-th crop row, x represents the coordinates of the feature point, w 0,i 、w 1,i 、w 2,i and w 3,i Both represent polynomial coefficients.
7. The crop row navigation line recognition method based on deep learning according to claim 1, characterized in that: The S8 specifically includes: S801: combining the curve equations of adjacent crop rows to determine the coordinates of the intersection of the curve equations; S802: Determine the slope of the navigation line and the intercept of the navigation line according to the coordinates of the intersection point.
8. The crop row navigation line recognition method based on deep learning according to claim 7, characterized in that: The calculation method of the slope of the navigation line is specifically as follows: ; Among them, m c represents the slope of the navigation line c, b c represents the slope of the navigation line c, x i The horizontal coordinate of the intersection point representing the curve equation, y i The ordinate of the intersection point representing the curve equation; The calculation method of the navigation line intercept is specifically as follows: ; Among them, tan represents the tangent function, arctan represents the inverse tangent function, m1 represents the slope of the first crop row curve, and m2 represents the slope of the second crop row curve.
9. A crop row navigation line recognition system based on deep learning, characterized in that: include: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the crop row navigation line recognition method based on deep learning as described in any one of claims 1 to 8 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, a crop row navigation line recognition method based on deep learning as described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Remote sensing target detection based on YOLOX-Tiny biased feature fusion network
CN116645608A
Self-adaptive navigation line direction prediction method and device, storage medium and agricultural machine
CN116839584A
Agricultural implement automatic navigation method used in field crop planting environment
CN117053808A
Target rapid detection method for improving YOLOX in complex environment and application
CN117253268A
Improved Yolov5 pedestrian detection method and system and storage medium
CN117542013A