Farmland obstacle detection method and device, storage medium and program product
Through the improved lightweight YOLO model and multi-level warning signals, the problems of low efficiency and low accuracy of traditional farmland obstacle detection are solved, and efficient and accurate obstacle identification and safety warning are achieved in complex farmland environments.
Patent Information
- Application Number
- CN202510772836.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-30
AI Technical Summary
Traditional farmland obstacle detection relies on manual inspections, which is inefficient and inaccurate, and is difficult to adapt to the complex and changing farmland environment, especially since obstacles are easily obscured, making detection difficult.
An improved lightweight YOLO model is used, combined with a depth camera, data enhancement, a dynamic location perception module, and a separation and enhancement attention module. Through deep separable convolution and dynamic deformable convolutional networks, obstacle features are identified and multi-level warning signals are generated.
It significantly improves the speed and accuracy of farmland obstacle detection, reduces the computing power requirements of edge devices, and can accurately identify and avoid obstacles in complex environments, ensuring the safe operation of harvesters.
Smart Images

Figure CN120726599A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a method, device, and storage medium for detecting farmland obstacles. Background Art
[0002] Currently, mechanized harvesting of crops, such as cotton, has become the primary method of harvesting in my country. With the rapid development of cutting-edge technologies such as artificial intelligence, big data, and the Internet of Things, crop production methods are moving from mechanization to automation and intelligence. The application of advanced technologies such as communications and intelligent control is also driving the evolution of precision agriculture technology systems for crop production. Farmland obstacle detection, a crucial component of harvesting, is the technical foundation for ensuring the safe and autonomous operation of harvesters.
[0003] Traditional farmland obstacle detection relies heavily on manual inspections, which are inefficient. Farmland obstacles are subject to complex environmental influences, often obscured and difficult to detect. Obstacles can also change dynamically, resulting in low manual detection accuracy and poor adaptability to the complex and ever-changing farmland environment. Therefore, overcoming the complex farmland operating environment and effectively identifying and avoiding multiple obstacle categories are crucial to improving operational efficiency and safety. Summary of the Invention
[0004] In response to the shortcomings of the existing technology, the present invention proposes a farmland obstacle detection method, device, storage medium, and program product. This method significantly improves the speed and accuracy of farmland obstacle detection and provides safety warnings.
[0005] In one aspect, the present invention provides a method for detecting farmland obstacles, comprising:
[0006] The depth camera on the harvester dynamically adjusts the shooting angle to collect images of obstacles in the farmland;
[0007] Based on the environmental characteristics of farmland obstacles, data enhancement processing is performed on the obstacle image and the obstacle boundary box and category information are marked to construct an obstacle dataset;
[0008] The obstacle dataset is input into a target detection model to identify target obstacle information, wherein the target detection model adopts an improved lightweight YOLO model, including:
[0009] A lightweight network is a main feature extraction network for enhancing implicit features of obstacles in the obstacle image;
[0010] A dynamic position sensing module, used to dynamically adjust the position deviation of each area in the obstacle image;
[0011] An improved dynamic deformable convolutional network, used as the bottleneck structure of the baseline model, is used to identify the common and differentiated features of different obstacles.
[0012] A separation and enhancement attention module is added to the detection head network to enhance the feature representation of occluded obstacles.
[0013] In one embodiment of the present invention, the obstacle includes at least one of a water spray pile, a weed pile, a woven bag, and a pedestrian.
[0014] In one embodiment of the present invention, data enhancement processing is performed on the obstacle image based on the environmental characteristics of the farmland obstacle, including:
[0015] In view of the uneven illumination characteristics of farmland, the obstacle image is subjected to random brightness to simulate the illumination changes when it is sunny, cloudy or blocked.
[0016] In view of the occlusion characteristics of farmland, the obstacle image is subjected to artifact enhancement technology to simulate the partial occlusion effect of crops on the obstacle;
[0017] In view of the multi-scale obstacle characteristics of farmland, the obstacle image is randomly cropped and randomly flipped to enhance the detection capability of multi-scale and multi-directional obstacles;
[0018] According to the interference characteristics of the farmland environment, Gaussian noise is added to the obstacle image to simulate sensor and environmental interference;
[0019] Among them, the environmental characteristics include: uneven farmland lighting characteristics, farmland occlusion characteristics, farmland multi-scale obstacle characteristics, and farmland environmental interference characteristics.
[0020] In one embodiment of the present invention, the lightweight feature extraction network is StarNet, which includes multiple Star Block modules stacked recursively in multiple layers, and uses star operations to perform element-by-element method on obstacle image features to exponentially increase the implicit feature dimension in the obstacle image.
[0021] In one embodiment of the present invention, in each Star Block module, basic features are extracted from the input obstacle image through depthwise separable convolution to obtain a first output feature;
[0022] Using batch normalization (BN) to normalize the first output feature;
[0023] The first output feature after normalization is upgraded through two convolutions. One convolution performs linear dimensionality upgrade to obtain the second output feature, and the other convolution performs nonlinear transformation through the ReLU activation function to obtain the third output feature.
[0024] Performing feature fusion on the second output feature and the third output feature by element-wise multiplication to obtain a fourth fused feature;
[0025] The fourth fusion feature is subjected to a depthwise separable convolution to perform final feature extraction, capturing the global semantic information of the obstacle and obtaining the fifth output feature.
[0026] In one embodiment of the present invention, the dynamic location awareness module is a DBA module, which is composed of a self-attention mechanism module and a multi-layer perceptron module, wherein:
[0027] The multi-layer perceptron module calculates the relative position deviation between each area and the harvester based on the position information of each area of the obstacle image;
[0028] The self-attention mechanism module assigns different spatial weights to each relative position deviation based on the local features of each region.
[0029] In one embodiment of the present invention, the multi-layer perceptron module performs nonlinear mapping on the position information of each region of the obstacle image through a fully connected layer;
[0030] Combined with layer normalization and ReLU activation function, the relative position deviation between each region and the harvester is calculated to generate a relative position deviation matrix.
[0031] In one embodiment of the present invention, the separation and enhancement attention module is a SEAM module, which includes: a depthwise separable convolution with residual connections and two fully connected layers, wherein:
[0032] The original feature map of the obstacle image is globally averaged pooled in each channel by depthwise separable convolution with residual connection to generate description vectors for each channel;
[0033] The description vectors of each channel are aggregated and compressed and expanded through two fully connected layers to capture long-range correlations in spatial and channel dimensions, establish dependencies between occluded and non-occluded objects, and generate channel attention weights.
[0034] The channel attention weight is multiplied by the original feature map to compensate for the loss of obstacle features caused by occlusion.
[0035] In one embodiment of the present invention, the learnable activation function in the KAN model is integrated with the fixed activation function in the traditional convolution to construct an improved dynamic deformable convolutional network; wherein:
[0036] The activation function of the improved dynamic deformable convolutional network is a linear combination of basis functions and spline functions. The basis function is determined by the fixed activation function in traditional convolution and is used to describe the common characteristics of obstacles. The spline function is determined by the learnable activation function in the KAN model and is used to dynamically adapt to the differentiated characteristics of different obstacles through piecewise polynomial parameterization.
[0037] In one embodiment of the present invention, a multi-level warning signal is generated based on the target obstacle information, and the multi-level warning signal is used to warn different safety levels, wherein the obstacle information includes the obstacle position, type, and size.
[0038] In one embodiment of the present invention, the multi-level warning signal at least includes:
[0039] In the first level of warning, if the obstacle is identified as a water spray pile or a pedestrian, and the obstacle position is smaller than the first safety threshold in front of the harvester, emergency braking is triggered;
[0040] Level 2 warning: If the obstacle is identified as a weed pile or woven bag, and the obstacle position is smaller than the second safety threshold in front of the harvester, and / or the obstacle size is larger than the size threshold, an alarm is triggered and the speed is reduced;
[0041] The priority of the first-level warning is higher than that of the second-level warning.
[0042] Another aspect of the present invention further provides a farmland obstacle detection device, comprising:
[0043] The acquisition module is used to dynamically adjust the shooting angle through the depth camera on the harvester to collect images of obstacles in the farmland;
[0044] A data preprocessing module is used to perform data enhancement processing on the obstacle image based on the environmental characteristics of the farmland obstacles and mark the obstacle bounding box and category information to construct an obstacle dataset;
[0045] The target detection module is used to input the obstacle dataset into the target detection model to identify target obstacle information, wherein the target detection model adopts an improved lightweight YOLO model, including:
[0046] A lightweight network is a main feature extraction network for enhancing implicit features of obstacles in the obstacle image;
[0047] A dynamic position sensing module, used to dynamically adjust the position deviation of each area in the obstacle image;
[0048] An improved dynamic deformable convolutional network, used as the bottleneck structure of the baseline model, is used to identify the common and differentiated features of different obstacles.
[0049] A separation and enhancement attention module is added to the detection head network to enhance the feature representation of occluded obstacles.
[0050] Another aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the steps of the above-mentioned farmland obstacle detection method when executed by a processor.
[0051] Another aspect of the present invention provides a computer program product, comprising a computer program, which implements the steps of the above-mentioned farmland obstacle detection method when executed by a processor.
[0052] From the above scheme, it can be seen that the advantages of the present invention are:
[0053] The proposed farmland obstacle detection method utilizes an improved lightweight YOLO model, effectively resolving the difficulty of accurate detection due to occlusion of obstacles in complex farmland environments. This significantly improves the speed and accuracy of farmland obstacle detection. Furthermore, the optimized network structure reduces the computing power required for deploying edge devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 The figure shows the overall process of the farmland obstacle detection method provided by one embodiment of the present invention;
[0055] Figure 2 This is a schematic diagram of the structure of the Star Block module in the lightweight network StarNet in the present invention;
[0056] Figure 3 Schematic diagram of the structure of the DBA module in the present invention;
[0057] Figure 4 Schematic diagram of the structure of the SEAM module in the present invention;
[0058] Figure 5 This is a performance comparison chart of the method of the present invention and different detection models in the prior art;
[0059] Figure 6 The figure shows the overall structure of a farmland obstacle detection device provided by one embodiment of the present invention.
[0060] Wherein, the accompanying drawings are marked as follows:
[0061] 300: Farmland obstacle detection device;
[0062] 310: acquisition module;
[0063] 320: data preprocessing module;
[0064] 330: target detection module;
[0065] 340: Early warning module. DETAILED DESCRIPTION
[0066] It should be noted that, in this application, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0067] Without further constraints, an element defined by the phrase "comprises a..." does not preclude the existence of additional identical elements in the process, method, article or apparatus that includes the element.
[0068] It should be noted that, in this application, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0069] Without further constraints, an element defined by the phrase "comprises a..." does not preclude the existence of additional identical elements in the process, method, article or apparatus that includes the element.
[0070] See also Figure 1 , Figure 1 The figure shows the overall process of the farmland obstacle detection method provided by one embodiment of the present invention.
[0071] A method for detecting farmland obstacles comprises the following steps:
[0072] Step S1: Dynamically adjust the shooting angle through the depth camera carried by the harvester to collect images of obstacles in the farmland.
[0073] In this example, a harvester's onboard depth camera dynamically adjusts its shooting angle to capture images of obstacles in the field under varying weather, lighting, and occlusion conditions. This generates multiple obstacle images and selects high-quality images to improve data quality and annotation efficiency. In practice, obstacles are common and can easily affect the harvester's normal operation, such as sprinkler piles, weed piles, woven bags, and pedestrians.
[0074] In one specific implementation, using cotton harvesting as an example, a cotton picker was used as a carrier equipped with a depth camera. The camera's mounting angle was adjusted based on the picker's 5m safe braking distance. Four common obstacles that can affect the normal operation of the cotton picker were captured in different weather conditions (sunny, cloudy), lighting conditions (front light, back light, fill light), and occlusion conditions. These obstacles included water spray piles, weed piles, woven bags, and pedestrians. 5,000 high-quality images were manually selected to improve data quality and annotation efficiency.
[0075] Step S2: Based on the environmental characteristics of the farmland obstacles, data enhancement processing is performed on the obstacle image and the obstacle boundary box and category information are marked to construct an obstacle dataset.
[0076] In this embodiment, based on the environmental characteristics of farmland obstacles, including uneven farmland lighting characteristics, farmland occlusion characteristics, farmland multi-scale obstacle characteristics, farmland environmental interference characteristics, etc., data enhancement operations are performed on several high-quality obstacle images after screening. Data enhancement techniques include random cropping, random flipping, random brightness, Gaussian noise, artifact enhancement, etc. After data enhancement, the data set of the initial several obstacle images is expanded to obtain an obstacle data set.
[0077] In a specific embodiment, in view of the uneven lighting characteristics of farmland, the obstacle image is subjected to random brightness to simulate the lighting changes when it is sunny, cloudy or blocked, which effectively improves the problem of uneven distribution of lighting conditions in the dataset.
[0078] In a specific embodiment, in view of the occlusion characteristics of farmland, the obstacle image is subjected to artifact enhancement technology to simulate the partial occlusion effect of crops on the obstacle, so that the model can better handle feature missing situations, thereby improving the detection performance in the actual working environment.
[0079] In a specific embodiment, in view of the multi-scale obstacle characteristics of farmland, the obstacle image is randomly cropped and randomly flipped to enhance the detection capability of multi-scale and multi-directional obstacle targets.
[0080] In a specific embodiment, according to the interference characteristics of the farmland environment, Gaussian noise is added to the obstacle image to simulate sensor and environmental interference, thereby further improving the adaptability of the model to real environmental noise.
[0081] Furthermore, the obstacle images after data augmentation are annotated with the bounding boxes and category information of the obstacles. In one specific embodiment, Labelme software is used for manual annotation. The annotation content includes the bounding boxes and category information of the obstacles in the YOLO format. The annotated objects are described using five numerical values: the target category, the x and y coordinates of the center of the bounding box, and the height and width of the bounding box. The annotated obstacle images are divided into a training set: a validation set: a test set in an 8:1:1 ratio.
[0082] Step S3: input the obstacle dataset into the target detection model to identify target obstacle information.
[0083] The target detection model adopts an improved lightweight YOLO model. In a specific embodiment, YOLO 11n is used as the baseline model for improvement, including: a lightweight network StarNet, which serves as the main feature extraction network to enhance the implicit features of obstacles in the obstacle image; a dynamic position perception module DBA, which introduces the dynamic position perception module DBA to reconstruct C2PSA to dynamically adjust the position deviation of each area in the obstacle image; an improved dynamic deformable convolutional network KAGNConv, which replaces the bottleneck structure in the baseline model C3k2 module to identify the common features and differentiated characteristics of different obstacles; a separation and enhancement attention module SEAM, which integrates the separation and enhancement attention module SEAM and adds it to the detection head network to enhance the feature expression of occluded obstacles.
[0084] In a specific embodiment, if Figure 2 As shown in Figure 2 A schematic diagram of the StarBlock module in the lightweight network StarNet is shown. The lightweight feature extraction network, StarNet, comprises multiple recursive StarBlock modules stacked in multiple layers. It uses a star-shaped operation to perform an element-by-element approach on obstacle image features, exponentially increasing the implicit feature dimensions in the obstacle image.
[0085] First, in each StarBlock module, the basic features of the input obstacle image are extracted through depthwise separable convolution to obtain the first output feature; and the first output feature is standardized using batch normalization BN. Here, the depthwise separable convolution preliminarily extracts low-level visual features of the obstacle, such as edge contours, color distribution, local contrast, etc. For example, for a sprinkler pile, it is mainly the straight edges of the metal material and the reflective area on the surface; for a weed pile, it is mainly the contours of irregular shapes and the dense textures of plant stems and leaves; for a woven bag, it is mainly the wrinkled boundaries of the flexible material, surface stains or damaged areas; for a pedestrian, it is mainly the human body contour (head, torso) and local features of dynamic posture (such as arm swing). Thus, the above features are aggregated to obtain the first output feature.
[0086] The first output features after normalization are then subjected to two 1×1 convolutions with an expansion factor of 4 for feature dimensionality increase. One convolution performs linear dimensionality increase to produce the second output features, while the other performs nonlinear transformation using the ReLU activation function to produce the third output features. The linear dimensionality increase through the 1×1 convolution preserves the spatial structure of the original features and enhances inter-channel correlation. For example, for a sprinkler, the feature represents the geometric shape of the metal structure (cylinder, linear support). For a weed pile, the feature represents the spatial distribution of plant clusters (density, coverage). For a woven bag, the feature represents the topology of the wrinkled area (continuous or broken texture). For a pedestrian, the feature represents the relative position of body parts (such as the connection between the head and torso). These features are aggregated to produce the second output features. The other convolution introduces nonlinearity using the ReLU activation function to highlight significant features (such as high-frequency texture and areas of contrast abrupt changes). For example, for a sprinkler, the feature represents reflective highlights and rust or wear details on the metal surface. For a weed pile, the feature represents the jagged texture of leaf edges and the longitudinal stripes of the stem. For woven bags, this represents the shadow changes caused by wrinkles and the light and dark contrast of surface stains. For pedestrians, this represents the texture of clothing patterns and the blurred edges caused by dynamic movement. By aggregating these features, we obtain the third output feature.
[0087] The second output feature and the third output feature are further fused by element-wise multiplication to obtain a fourth fused feature. By fusing linear and nonlinear features through element-wise multiplication, the feature expression of key areas is strengthened and redundant information is suppressed. For example, for sprinkler piles, the geometric shape of the metal edge and the reflective details are fused to enhance positioning accuracy. For weed piles, the densely distributed contours and leaf textures are combined to distinguish weed piles from background vegetation. For woven bags, the topological structure of folds and shadow changes are integrated to improve the recognition robustness of flexible materials. For pedestrians, the human body contour and dynamic texture are associated to reduce false detections caused by motion blur. The above features are aggregated to obtain the fourth output feature.
[0088] Finally, the fourth fusion feature is subjected to the final feature extraction through depthwise separable convolution to capture the global semantic information of the obstacle and obtain the fifth output feature. High-order abstract features are further extracted through depthwise separable convolution to capture the global semantic information of the obstacle. For example, for a sprinkler pile, the integrity of the metal structure (whether it is tilted or broken) and the contrast with the farmland background are obtained. For a weed pile, the growth status of the plant community (sparse or dense) and the degree of occlusion of the harvester path are obtained. For woven bags, the degree of damage of the material, the overall integrity, and the abnormal position in the farmland (such as deviation from the working area) are obtained. For pedestrians, the movement direction prediction (close to or away from the harvester) and the safety distance judgment are obtained. The above features are aggregated to obtain the fifth output feature.
[0089] In a specific embodiment, in view of the uneven distribution of obstacles during the harvesting process, especially weed piles, a dynamic position awareness module DBA is introduced to reconstruct C2PSA and dynamically adjust the position deviation of each area in the obstacle image. Figure 3 As shown in Figure 3 The structural diagram of the DBA module is shown. The dynamic position perception module consists of a self-attention mechanism module and a multi-layer perceptron module, and the dynamic position deviation is embedded in it to obtain the relative position information of each obstacle. The multi-layer perceptron module calculates the relative position deviation between each region and the harvester based on the position information of each region of the obstacle image; the self-attention mechanism module assigns different spatial weights to each relative position deviation based on the local features of each region. In one embodiment, the multi-layer perceptron module performs nonlinear mapping on the position information of each region of the obstacle image through a fully connected layer; combined with layer normalization and ReLU activation function, the relative position deviation between each region and the harvester is calculated to generate a relative position deviation matrix. In this embodiment, in view of the uneven distribution of obstacles, the position deviation can be dynamically adjusted according to the position and local features of the obstacles, thereby assigning different weights to the regions of interest and reducing the waste of computing resources.
[0090] In a specific embodiment, an improved dynamic deformable convolutional network KAGNConv is used to replace the bottleneck structure in the baseline model C3k2 module to identify the common features and differentiated characteristics of different obstacles. KAGNConv convolution is an innovative convolution operation proposed based on the KAN network theoretical framework, which combines the learnable activation function in the KAN model with the fixed activation function in the traditional convolution; the KAGN convolution layer is selected to balance the model complexity and performance. The KAN model uses a learnable activation function at the edge of the network, so that each weight parameter can be replaced by a set of learnable univariate functions. These functions are usually parameterized in the form of spline functions and can simulate complex functions with fewer parameters, thereby providing the model with higher flexibility and interpretability.
[0091] The activation function of the improved dynamic deformable convolutional network is a linear combination of a basis function and a spline function. The basis function is determined by the fixed activation function in traditional convolution and is used to describe the common characteristics of obstacles. The spline function is determined by the learnable activation function in the KAN model and is used to dynamically adapt to the differentiated characteristics of different obstacles through piecewise polynomial parameterization. The activation function φ(x) can be decomposed into a linear combination of the basis function b(x) and the spline function spline(x), expressed as follows:
[0092] φ(x)=w b b(x)+w s spline(x)
[0093]
[0094] spline(x)=∑c i B i (x)
[0095] Among them, w b 、w s represents the weight coefficient, c i is the weight coefficient, B i (x) represents the piecewise function corresponding to different obstacles.
[0096] In this embodiment, an improved dynamic deformable convolutional network KAGNConv is used to replace the bottleneck structure in the baseline model C3k2 module to identify the common features and differentiated features of different obstacles. The common features of different obstacles, such as shape and size, KAGNConv uses basis functions to enable the network to adaptively capture the contours of obstacles of different shapes and sizes. Differentiated features of different obstacles, such as sprinkler piles, where metal materials cause strong reflections, sharp edges but are easily affected by light. Through KAGNConv, a spline function is used to smooth the reflective area. For weed piles, plant stems and leaves are dense and have complex textures, with blurred boundaries. Through KAGNConv, a spline function is used to highlight high-frequency textures and distinguish weed piles from background vegetation. For woven bags, wrinkles and deformations are prone to cause local gradient mutations. Through KAGNConv, a spline function is used to highlight the local gradients in the wrinkled area.
[0097] In a specific embodiment, if Figure 4 As shown, Figure 4 A schematic diagram of the SEAM module is shown. The integrated Separation and Enhanced Attention Module (SEAM) is added to the detection head network to enhance the feature representation of occluded obstacles, addressing the difficulty in accurately detecting obstacles in complex farmland environments due to occlusion. The SEAM module comprises a depthwise separable convolution with residual connections and two fully connected layers. The depthwise separable convolution with residual connections performs global average pooling on the original feature map of the obstacle image across all channels to generate a description vector for each channel, effectively learning the importance of each channel and reducing the number of model parameters. The description vectors of each channel are then aggregated, compressed, and expanded by two fully connected layers to capture long-range correlations in the spatial and channel dimensions, establish dependencies between occluded and unoccluded objects, and generate channel attention weights. This process helps propagate contextual information, providing the detection head with richer and more comprehensive decision-making information. Finally, the channel attention weights are multiplied with the original feature map to compensate for the loss of obstacle features caused by occlusion, allowing the model to more accurately focus on partially occluded obstacles when processing occlusions.
[0098] In one embodiment, based on the obtained obstacle information, including the location, type, and size of the obstacle, a multi-level warning signal is further generated. The multi-level warning signal is used to warn of different safety levels, and corresponding safety warning levels and avoidance measures can be formulated based on the risks caused by the detected obstacles to the harvesting operation of the harvester. Based on the risks caused by the detected obstacles to the harvesting operation of the harvester, corresponding safety warning levels and avoidance measures can be formulated. For example, the multi-level warning signal can include: a first-level warning, a second-level warning, and a third-level warning, with the danger level gradually decreasing from level one to level three. The first-level warning indicates "the most dangerous", and the obstacle affects the normal operation of the harvester, and the harvester stops immediately. For example, the first-level warning is configured as follows: if the obstacle is identified as a water spray pile or a pedestrian, and the obstacle position is less than the first safety threshold in front of the harvester, emergency braking is triggered. The second-level warning indicates "relatively dangerous", and the obstacle may affect the normal operation of the harvester, and an alarm will be sounded immediately and the harvester will slow down. For example, the second-level warning configuration is: if the obstacle is identified as a weed pile or a woven bag, and the obstacle position is smaller than the second safety threshold in front of the harvester, and / or the obstacle size is larger than the size threshold, an alarm reminder is triggered and the speed is reduced. The third-level warning indicates "relatively safe", and the obstacle does not affect the normal operation of the harvester, and normal harvesting operations are carried out. For example, the third-level warning configuration is: if the obstacle is located on the side or rear of the harvester, only an alarm reminder is triggered, but normal operation is maintained. Therefore, during the harvesting process, according to the obstacle information obtained by dynamic detection, changes in obstacle type, size, position, etc., corresponding safety warnings and avoidance measures are executed.
[0099] The following takes cotton field obstacle detection as an example to verify the effectiveness of the method of the present invention.
[0100] The experiment used a self-built cotton field obstacle dataset and used the same training environment and parameter configuration during the experiment. The model was evaluated using evaluation indicators such as precision, recall, mean average precision (mAP50), and floating-point operations. In order to verify the reliability of the improved model of the present invention, different detection models were selected for comprehensive comparative tests. The detection results of each model are shown in Table 1, and the effect comparison diagram is shown in Figure 1. Figure 5 shown.
[0101] Table 1 Comparative test results of detection models
[0102]
[0103] According to the above Table 1 and the attached Figure 5As can be seen, the improved model of the present invention has the highest precision, recall, and mean average precision (mAP50). This model requires only 2.14×106 parameters, making it more suitable for deployment on edge devices with limited resources. The detection speed of the improved model decreases slightly, but still reaches 85 f / s, which can meet the real-time obstacle detection requirements during cotton picker operation. This demonstrates that the improved model effectively addresses issues such as difficulty in detecting objects accurately due to obstructions in the complex environment of cotton fields and the limited computing power of edge devices, while achieving a good balance between performance and computational complexity.
[0104] In summary, the farmland obstacle detection method provided by the present invention adopts an improved lightweight YOLO model, wherein the lightweight network StarNet serves as the main feature extraction network to enhance the implicit features of obstacles in the obstacle image; the dynamic position perception module DBA is introduced to reconstruct the C2PSA to dynamically adjust the position deviation of each area in the obstacle image; the improved dynamic deformable convolutional network KAGNConv replaces the bottleneck structure in the baseline model C3k2 module to identify the common features and differentiated characteristics of different obstacles; the separation and enhancement attention module SEAM is integrated and added to the detection head network to enhance the feature expression of occluded obstacles. This method uses the improved lightweight YOLO model as the target detection model, effectively solving the problem of difficult accurate detection caused by occlusion of obstacles in complex environments between farmlands, and significantly improving the speed and accuracy of farmland obstacle detection. At the same time, the optimized network structure reduces the computing power requirements for deploying edge devices.
[0105] It should also be understood that in the above-mentioned embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0106] The following are device embodiments corresponding to the above method embodiments, such as Figure 6 As shown, Figure 6 The following figure shows a schematic diagram of the structure of a farmland obstacle detection device 300 provided by one embodiment of the present invention. This method embodiment can be implemented in conjunction with the aforementioned device embodiment. The relevant technical details mentioned in the aforementioned method embodiment are still valid in this device embodiment and are not repeated here to reduce repetition.
[0107] A farmland obstacle detection device 300, comprising:
[0108] The acquisition module 310 is used to dynamically adjust the shooting angle through the depth camera carried by the harvester to collect images of obstacles in the farmland.
[0109] A data preprocessing module 320 is used to perform data enhancement processing on the obstacle image based on the environmental characteristics of the farmland obstacles and mark the obstacle bounding box and category information to construct an obstacle dataset;
[0110] The target detection module 330 is used to input the obstacle dataset into the target detection model to identify target obstacle information, wherein the target detection model adopts an improved lightweight YOLO model, including:
[0111] A lightweight network is a main feature extraction network for enhancing implicit features of obstacles in the obstacle image;
[0112] A dynamic position sensing module, used to dynamically adjust the position deviation of each area in the obstacle image;
[0113] An improved dynamic deformable convolutional network, used as the bottleneck structure of the baseline model, is used to identify the common and differentiated features of different obstacles.
[0114] A separation and enhancement attention module is added to the detection head network to enhance the feature representation of occluded obstacles.
[0115] In one embodiment, the device further includes an early warning module 340 for generating a multi-level early warning signal based on the target obstacle information. The multi-level early warning signal is used to warn of different safety levels, wherein the obstacle information includes the obstacle's location, type, and size. In one specific embodiment, the multi-level early warning signal includes at least: a first-level early warning, which triggers emergency braking if the obstacle is identified as a water spray pile or a pedestrian, and the obstacle's location is less than a first safety threshold in front of the harvester; a second-level early warning, which triggers an alarm and reduces the speed if the obstacle is identified as a weed pile or a woven bag, and the obstacle's location is less than a second safety threshold in front of the harvester, and / or the obstacle's size is greater than a size threshold; wherein the first-level early warning has a higher priority than the second-level early warning.
[0116] The embodiment of this device can be implemented in conjunction with the implementation of the above method embodiment. The relevant technical details mentioned in the implementation of the above method embodiment are still valid in the implementation of this device embodiment, and will not be repeated here to reduce repetition.
[0117] In addition, the present invention also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the farmland obstacle detection method disclosed in the above embodiment are implemented.
[0118] In addition, the present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the farmland obstacle detection method disclosed in the above embodiment.
[0119] In addition, it should be understood that the storage medium in the embodiments of the apparatus of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus RAM (DR RAM).
[0120] Although the embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the description and implementation methods. They can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily implemented. Therefore, without departing from the general concept defined by the claims and the scope of equivalents, the present invention is not limited to the specific details and illustrations shown and described herein.
Claims
1. A method for detecting farmland obstacles, characterized in that: include: The depth camera on the harvester dynamically adjusts the shooting angle to collect images of obstacles in the farmland; Based on the environmental characteristics of farmland obstacles, data enhancement processing is performed on the obstacle image and the obstacle boundary box and category information are marked to construct an obstacle dataset; The obstacle dataset is input into a target detection model to identify target obstacle information, wherein the target detection model adopts an improved lightweight YOLO model, including: A lightweight network is a main feature extraction network for enhancing implicit features of obstacles in the obstacle image; A dynamic position sensing module, used to dynamically adjust the position deviation of each area in the obstacle image; An improved dynamic deformable convolutional network, used as the bottleneck structure of the baseline model, is used to identify the common and differentiated features of different obstacles. A separation and enhancement attention module is added to the detection head network to enhance the feature representation of occluded obstacles.
2. The method according to claim 1, characterized in that The obstacle includes at least one of a water spray pile, a weed pile, a woven bag and a pedestrian.
3. The method according to claim 1, characterized in that Performing data enhancement processing on the obstacle image based on the environmental characteristics of the farmland obstacles, including: In view of the uneven illumination characteristics of farmland, the obstacle image is subjected to random brightness to simulate the illumination changes when it is sunny, cloudy or blocked. In view of the occlusion characteristics of farmland, the obstacle image is subjected to artifact enhancement technology to simulate the partial occlusion effect of crops on the obstacle; In view of the multi-scale obstacle characteristics of farmland, the obstacle image is randomly cropped and randomly flipped to enhance the detection capability of multi-scale and multi-directional obstacles; According to the interference characteristics of the farmland environment, Gaussian noise is added to the obstacle image to simulate sensor and environmental interference; Among them, the environmental characteristics include: uneven farmland lighting characteristics, farmland occlusion characteristics, farmland multi-scale obstacle characteristics, and farmland environmental interference characteristics.
4. The method according to claim 2, characterized in that The lightweight feature extraction network is StarNet, which includes multiple Star Block modules stacked recursively in multiple layers. Star operations are used to perform element-by-element method on obstacle image features to exponentially increase the implicit feature dimension in the obstacle image.
5. The method according to claim 2, characterized in that In each StarBlock module, the basic features of the input obstacle image are extracted through depth-wise separable convolution to obtain the first output feature; Using batch normalization (BN) to normalize the first output feature; The first output feature after normalization is upgraded through two convolutions. One convolution performs linear dimensionality upgrade to obtain the second output feature, and the other convolution performs nonlinear transformation through the ReLU activation function to obtain the third output feature. Performing feature fusion on the second output feature and the third output feature by element-wise multiplication to obtain a fourth fused feature; The fourth fusion feature is subjected to a depthwise separable convolution to perform final feature extraction, capturing the global semantic information of the obstacle and obtaining the fifth output feature.
6. The method according to claim 1, characterized in that The dynamic location awareness module is a DBA module, which consists of a self-attention mechanism module and a multi-layer perceptron module, wherein: The multi-layer perceptron module calculates the relative position deviation between each area and the harvester based on the position information of each area of the obstacle image; The self-attention mechanism module assigns different spatial weights to each relative position deviation based on the local features of each region.
7. The method according to claim 6, characterized in that The multi-layer perceptron module performs nonlinear mapping on the position information of each area of the obstacle image through a fully connected layer; Combined with layer normalization and ReLU activation function, the relative position deviation between each region and the harvester is calculated to generate a relative position deviation matrix.
8. The method according to claim 1, characterized in that The separation and enhancement attention module is a SEAM module, which includes: depthwise separable convolution with residual connection and two fully connected layers, where: The original feature map of the obstacle image is globally averaged pooled in each channel by depthwise separable convolution with residual connection to generate description vectors for each channel; The description vectors of each channel are aggregated and compressed and expanded through two fully connected layers to capture long-range correlations in spatial and channel dimensions, establish dependencies between occluded and non-occluded objects, and generate channel attention weights. The channel attention weight is multiplied by the original feature map to compensate for the loss of obstacle features caused by occlusion.
9. The method according to claim 1, characterized in that The learnable activation function in the KAN model is combined with the fixed activation function in traditional convolution to construct an improved dynamic deformable convolutional network; The activation function of the improved dynamic deformable convolutional network is a linear combination of basis functions and spline functions. The basis function is determined by the fixed activation function in traditional convolution and is used to describe the common characteristics of obstacles. The spline function is determined by the learnable activation function in the KAN model and is used to dynamically adapt to the differentiated characteristics of different obstacles through piecewise polynomial parameterization.
10. The method according to claim 2, characterized in that A multi-level warning signal is generated based on the target obstacle information, and the multi-level warning signal is used to warn different safety levels, wherein the obstacle information includes the obstacle location, type, and size.
11. The method according to claim 10, characterized in that The multi-level warning signal at least includes: In the first level of warning, if the obstacle is identified as a water spray pile or a pedestrian, and the obstacle position is smaller than the first safety threshold in front of the harvester, emergency braking is triggered; Level 2 warning: If the obstacle is identified as a weed pile or woven bag, and the obstacle position is smaller than the second safety threshold in front of the harvester, and / or the obstacle size is larger than the size threshold, an alarm is triggered and the speed is reduced; The priority of the first-level warning is higher than that of the second-level warning.
12. A farmland obstacle detection device, characterized in that: include: The acquisition module is used to dynamically adjust the shooting angle through the depth camera on the harvester to collect images of obstacles in the farmland; A data preprocessing module is used to perform data enhancement processing on the obstacle image based on the environmental characteristics of the farmland obstacles and mark the obstacle bounding box and category information to construct an obstacle dataset; The target detection module is used to input the obstacle dataset into the target detection model to identify target obstacle information, wherein the target detection model adopts an improved lightweight YOLO model, including: A lightweight network is a main feature extraction network for enhancing implicit features of obstacles in the obstacle image; A dynamic position sensing module, used to dynamically adjust the position deviation of each area in the obstacle image; An improved dynamic deformable convolutional network, used as the bottleneck structure of the baseline model, is used to identify the common and differentiated features of different obstacles. A separation and enhancement attention module is added to the detection head network to enhance the feature representation of occluded obstacles.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the farmland obstacle detection method described in any one of claims 1 to 11 are implemented.
14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the farmland obstacle detection method described in any one of claims 1 to 11 are implemented.
Citation Information
Patent Citations
Rail transit obstacle detection method based on improved convolutional neural network
CN113486726A
Vehicle distance early warning method, system and equipment based on field obstacle detection
CN117351461A
Control method and control device for collision prevention of agricultural vehicle and agricultural vehicle
CN118457571A
Obstacle detection and avoidance method and system for automatic driving agricultural machinery
CN118587677A
Obstacle detection method in operation process of facility agricultural robot
CN118609093A
Cited By
Intelligent tractor closed-loop obstacle avoidance method based on improved YOLOv10
CN122691285A